Projects

Software I specified for the lab, and side projects in data and machine learning: the problem, the decisions and the measured result.

3 tools
In daily lab use
2 h → 5 min
Polarisation measurement
R² 0.92
Solar forecast (side project)
RMSE 18.8
Engine-life prediction (side project)

Instrument software I specified

In daily use on the Kassel PEPICO spectrometer and laser.

Click to enlarge
In daily lab use

VAJRA: PEPICO Data Acquisition System

Software that records every electron and ion our spectrometer detects, and turns those timestamps into calibrated results.

1 kHz
Shot rate
every laser shot recorded
415.5k
Shots verified
end to end, test run
429
Unit tests
no hardware required
  • +One run showed the rotating waveplate was mislabelling 28% of shots. A hardware gate I specified now drops every shot taken while the stage moves, checked on every shot of a test run
  • +Showed by measurement that the converter gives no falling-edge marker events, although its programming interface lists them. The analysis relies on rising edges only
  • +429 unit tests that need no hardware attached, so the analysis can be checked from a laptop. The hardware tests are a separate suite
The full story: problem, approach, tools →

Problem

Our experiment fires a UV laser at a molecular beam a thousand times a second. Each shot can knock out an electron and leave an ion behind, and a time-to-digital converter records when each one arrived. The software has to hold on to every one of those timestamps through hours of running, on hardware that does not always behave. The old setup lost data when it was pushed. Worse, the raw capture and the processed output were tangled together, so reanalysing a run put the original at risk.

Approach

I specified the architecture and own the system. One core library does the hardware talking and the file writing, with no interface code in it at all. Three thin applications sit on top: a headless command-line runner as the safety net, the live acquisition GUI, and a separate offline analysis tool. The rule underneath all of it is that the raw capture gets written once and never touched again; everything else is derived from it. I also measured how the converter actually behaves on the bench instead of trusting the manual, and those measurements are in the repository.

More results

  • +Verified end to end on all 415,500 shots of a test run: the waveplate angles saved in the file match the raw capture
  • +Every run writes two files: the untouched raw capture, and an HDF5 product holding the histograms, the metadata and a SHA-256 checksum. The HDF5 is written under a temporary name and renamed only once it is complete, so a crash cannot leave a half-written file that looks finished
  • +Operator manuals for the acquisition and the analysis side, plus an API reference that CI rebuilds on every push
PythonPyQt6quTAG TDC SDKNumPyHDF5ParquetNewport ESP301pytestSphinxGitHub Actions
Click to enlarge
In active use

VAJRA Analysis

The offline half of VAJRA. It reads a recorded run and turns it into physics: calibrated spectra, coincidences, and dichroism yields.

3
Interfaces
one shared codebase
  • +Three interfaces over one implementation, so a number from the notebook and a number from the GUI cannot disagree
  • +Strict 1+1 coincidence sorting with the Stert false-background correction. That is the difference between counting real events and counting pairs that landed in the same window by chance
  • +Reads a whole scan from one file, so a power or temperature series is a single dataset instead of forty loose runs
The full story: problem, approach, tools →

Problem

A recorded run is a few hundred megabytes of timestamps and nothing else. Getting physics out of it means calibrating two different time axes, working out which electron belongs with which ion, and correcting for the pairs that only look related. Three different people want to do that three different ways: click through one run, explore it in a notebook, or leave it running over a whole folder overnight. Write those as three codebases and they drift, and then two analyses of the same file disagree with each other.

Approach

One set of functions with three ways in. A desktop application for working through a single run, Jupyter notebooks for the exploratory pass, and a command line for batch work. All three call the same code, so the same file gives the same numbers whichever route you take. It handles the electron-energy and ion-mass calibration, strict one-to-one electron-ion coincidences with the Stert correction for accidental pairs, and dichroism yields grouped by waveplate angle. It grew out of the notebook pipeline I had been running by hand, which is where the requirements came from.

More results

  • +Generates physically plausible test data in the real file formats, so the analysis can be checked without spending beamtime on it
  • +Knows how a run was marked, and analyses hardware-marked and software-marked runs the same way
  • +Analysis now runs at acquisition speed; the experiment sets the pace now, not the analysis
PythonPyQt6JupyterNumPyPandasHDF5ParquetMatplotlib
Click to enlarge
Completed and in use

Automated Laser Polarimetry Platform

Measures the full polarisation state of our laser automatically. What took two hours by hand now takes five minutes.

96%
Time reduction
2h to 5 min
  • +What took more than two hours by hand now takes about five minutes
  • +Full Stokes vector with uncertainties propagated from the fit covariance matrix, not quoted without them
  • +Swapping the motion controller costs a config change, not a rewrite
The full story: problem, approach, tools →

Problem

Characterising the polarisation at each focus of a Twin-Foci ultrafast laser setup required manually rotating a motorized waveplate, reading power at each angle, and fitting data in a spreadsheet. The process took over 2 hours per dataset and was error-prone.

Approach

I specified and own a PyQt5 platform that drives the rotation stage and an Ophir NOVAII power meter together, then pulls out the full Stokes vector using the Fourier method of Schaefer et al. The stage is not hardwired into the code: the same build runs a Newport ESP301 or the newer Trinamic TMCM-6110, picked by a config setting. There is also a simulation mode, so the interface and the fitting can be tested without occupying the laser.

More results

  • +Flags near-perfect circular polarisation on the Ring criterion, and reports the fit quality next to every result
  • +Publication-quality export at 300 DPI in PNG, PDF, and SVG
PythonPyQt5NumPySciPyMatplotlibPySerialPyVISAPyTrinamic

Run the measurement yourself

The same method the platform uses, running in your browser on simulated light. A quarter-wave plate turns in steps in front of a fixed polarizer, the detector reading traces I(θ), and a Fourier fit returns the four Stokes parameters. Pick the incoming polarization, turn the plate, or run a full scan.

Method: rotating quarter-wave-plate polarimetry with Fourier analysis, after Schaefer, Collett, Smyth, Barrett and Fraher, American Journal of Physics 75, 163 (2007). The light is the platform's 309 nm ultraviolet line, which is invisible, so it is drawn in violet. The field, the waveplate and the fit are exact; only the scatter on the measured points is simulated, at 1% of the incoming intensity.

Click to enlarge
Running around the clock

Lab Dashboard: Live Vacuum and Temperatures

Live vacuum pressures and temperatures for our lab, on any PC in the lab.

24/7
Monitoring
live on the lab network
  • +No gaps after a network outage: readings are kept locally and sent on later
  • +A clear warning when the readings stop updating, so a stopped logger cannot pass for a good vacuum
  • +Shows a capacitance gauge on the source next to the standard gauges. It reads the true pressure for any gas within its range, which matters when the source runs with water vapour; below that range it shows its zero offset
The full story: problem, approach, tools →

Problem

The vacuum and the source temperatures could only be read on one PC in the lab. And a frozen reading looks exactly like a healthy vacuum.

Approach

I specified a small monitoring system. It reads the lab's pressure and temperature controllers, keeps every sample locally first, and shows live values and their history to anyone on the lab network.

PythonSQLiteMQTT

Data and AI side projects

Built on public data, to learn machine learning end to end.

Click to enlarge

Probabilistic Solar Forecasting: Physics-Informed

Day-ahead solar generation forecasts for the German grid. Physics handles geometry; XGBoost learns only the residual.

R2=0.92
Physics + XGBoost
vs R2=0.78 physics only
60%
MAE reduction
1,552 vs 3,856 MW
CRPS 514.6
Uncertainty
P90 coverage 0.869 (target 0.90)
  • +Physics baseline alone: R2=0.78; Physics plus XGBoost: R2=0.92, MAE reduced 60%
  • +physics_pred is XGBoost's top feature by importance; model amplifies physics, not ignores it
  • +Split conformal prediction: P90 empirical coverage 0.869 against a 0.90 target; P50 under-covers after a seasonal shift (ADR-003)
The full story: problem, approach, tools →

Problem

Germany's Energiewende targets 80% of electricity from renewables by 2030. Solar is volatile and weather-dependent. A 10% forecasting error at midday peak costs real money on the balancing market. Operators need calibrated probabilistic forecasts, not point estimates: 12 GW plus or minus 2 GW, with a coverage guarantee.

Approach

Built a two-layer physics-informed architecture on 3 years of public SMARD and Open-Meteo data (26,000+ hourly records in TimescaleDB). The physics layer uses pvlib to compute what solar output should be on a geometrically perfect day. XGBoost learns only the residual: what physics cannot see, namely clouds, curtailments, and measurement noise. P10/P50/P90 intervals via split conformal prediction. Managed as a 22-day research sprint with 5 ADRs, 3 weekly reports, 2 retrospectives, and a public Kanban board.

More results

  • +CRPS=514.6 MW evaluated with reliability diagrams; calibration reported honestly
  • +Full stack: TimescaleDB, FastAPI, and Streamlit dashboard starts with one command
  • +GitHub Actions CI: ruff and pytest on every push; 5 Architecture Decision Records
PythonpvlibXGBoostMAPIEproperscoringTimescaleDBFastAPIStreamlitDockerMLflowGitHub Actions
Click to enlarge

Intelligent Predictive Maintenance System

Predicts turbofan engine remaining useful life with SHAP explainability, served through a FastAPI service with auto-generated diagnostic reports.

18.83
RMSE (cycles)
vs 51.33 baseline
62%
Error reduction
from domain insight
2.0
Train-test gap
vs 36.8 baseline RF
  • +XGBoost RMSE: 18.83 cycles with a train/test gap of only 2.0 (vs 12.6 for RF with RUL clipping)
  • +SHAP layer: top sensor drivers ranked per prediction with plain-language explanation
  • +PDF diagnostic report auto-generated in German DIN format from both API and dashboard
The full story: problem, approach, tools →

Problem

Industrial turbofan engines degrade gradually across 21 sensor channels. Threshold-based monitoring catches failures too late. Engineers need to know not just when failure occurs, but why: which sensors are driving the risk and what to do about it.

Approach

Built a complete end-to-end ML pipeline on NASA's CMAPSS dataset. The core insight: capping RUL targets at 125 cycles (treating early healthy cycles as interchangeable) dropped RMSE from 51 to 19, a domain decision that outperformed any algorithm choice. XGBoost with SHAP explainability, a 3-page Streamlit dashboard, FastAPI REST endpoint, and auto-generated PDF diagnostic reports in German DIN format. Everything runs with one command via Docker Compose.

More results

  • +FastAPI: POST /predict returns RUL, anomaly score, and status label in under 100ms
  • +Full MLflow experiment tracking: 162 logged runs, fully reproducible
  • +Docker Compose: API and dashboard start with one command
PythonXGBoostSHAPFastAPIStreamlitMLflowfpdf2Dockerscikit-learnPlotly

In development

LangGraph · Claude API · ChromaDB

Tool: RUL APITool: SHAPTool: RAG
In developmentView on GitHub

Sensor Intelligence Assistant

An AI assistant, in development, designed to reason across the predictive maintenance tools and diagnose turbofan engine anomalies from real API responses.

P3
Agentic AI Sprint
In progress
3
Reasoning tools
RUL + SHAP + RAG
  • +Planned: human-in-the-loop confirmation before any action
  • +Tool calls go to the real FastAPI endpoints from P1
  • +ChromaDB vector store for maintenance documentation retrieval
The full story: problem, approach, tools →

Problem

An operator types: What is happening with engine 14? A useful AI system should reason across multiple tools, deciding which to call and in what order based on what it finds at each step, and cite real sources rather than inventing them.

Approach

Designed as a LangGraph state machine with three tools, two of them from P1: Tool A calls the FastAPI prediction endpoint for remaining useful life. Tool B returns SHAP-grounded sensor explanations. Tool C retrieves relevant maintenance documentation from a ChromaDB vector store. The model is meant to write a diagnostic report citing the API responses, then pause for human confirmation before any action. Read-only by design, with LangSmith tracing.

More results

  • +LangSmith tracing of each step
LangGraphAnthropic Claude APIChromaDBsentence-transformersFastAPIStreamlitDockerLangSmith

Fellowship

How my PhD was funded, and what running it involved.

Click to enlarge
Completed

INSPIRE Fellowship Research Program

National PhD fellowship (DST, Government of India), plus hands-on funding work under my PI: fellowship reporting, a full proposal draft and project procurement.

5 yrs
Program duration
INSPIRE Fellowship
1
Granted patent
co-inventor
  • +Annual technical progress reports to DST, with the supervisor's continuation recommendation
  • +Contributed to one granted Indian patent for gas-separation membranes
  • +Procurement and closure documentation for a funded SERB project, prepared for PI certification
The full story: problem, approach, tools →

Problem

Hold a competitive national PhD fellowship while doing the research itself, and support the group's funded projects and proposals, with my supervisor as principal investigator.

Approach

Prepared my fellowship application and the annual technical progress reports to DST. On the group's SERB-funded project I specified consumables, compared vendor quotations, prepared purchasing material and assembled annual and closure reports for the PI to certify. Drafted a full research proposal (BRNS, 2021) under PI supervision.

More results

  • +Research visit to Jacobs University Bremen, Germany: Raman and UV-Vis work that led to two papers
Proposal WritingTechnical ReportingProcurement SupportScientific WritingStakeholder Communication

Questions about any of these?

The side projects are on GitHub. The lab software stays with the group, but I am happy to walk you through its design and how I validated it.