Projects
Software I specified for the lab, and side projects in data and machine learning: the problem, the decisions and the measured result.
Instrument software I specified
In daily use on the Kassel PEPICO spectrometer and laser.
VAJRA: PEPICO Data Acquisition System
Software that records every electron and ion our spectrometer detects, and turns those timestamps into calibrated results.
- +One run showed the rotating waveplate was mislabelling 28% of shots. A hardware gate I specified now drops every shot taken while the stage moves, checked on every shot of a test run
- +Showed by measurement that the converter gives no falling-edge marker events, although its programming interface lists them. The analysis relies on rising edges only
- +429 unit tests that need no hardware attached, so the analysis can be checked from a laptop. The hardware tests are a separate suite
The full story: problem, approach, tools →Hide the full story ↑
Problem
Our experiment fires a UV laser at a molecular beam a thousand times a second. Each shot can knock out an electron and leave an ion behind, and a time-to-digital converter records when each one arrived. The software has to hold on to every one of those timestamps through hours of running, on hardware that does not always behave. The old setup lost data when it was pushed. Worse, the raw capture and the processed output were tangled together, so reanalysing a run put the original at risk.
Approach
I specified the architecture and own the system. One core library does the hardware talking and the file writing, with no interface code in it at all. Three thin applications sit on top: a headless command-line runner as the safety net, the live acquisition GUI, and a separate offline analysis tool. The rule underneath all of it is that the raw capture gets written once and never touched again; everything else is derived from it. I also measured how the converter actually behaves on the bench instead of trusting the manual, and those measurements are in the repository.
More results
- +Verified end to end on all 415,500 shots of a test run: the waveplate angles saved in the file match the raw capture
- +Every run writes two files: the untouched raw capture, and an HDF5 product holding the histograms, the metadata and a SHA-256 checksum. The HDF5 is written under a temporary name and renamed only once it is complete, so a crash cannot leave a half-written file that looks finished
- +Operator manuals for the acquisition and the analysis side, plus an API reference that CI rebuilds on every push
VAJRA Analysis
The offline half of VAJRA. It reads a recorded run and turns it into physics: calibrated spectra, coincidences, and dichroism yields.
- +Three interfaces over one implementation, so a number from the notebook and a number from the GUI cannot disagree
- +Strict 1+1 coincidence sorting with the Stert false-background correction. That is the difference between counting real events and counting pairs that landed in the same window by chance
- +Reads a whole scan from one file, so a power or temperature series is a single dataset instead of forty loose runs
The full story: problem, approach, tools →Hide the full story ↑
Problem
A recorded run is a few hundred megabytes of timestamps and nothing else. Getting physics out of it means calibrating two different time axes, working out which electron belongs with which ion, and correcting for the pairs that only look related. Three different people want to do that three different ways: click through one run, explore it in a notebook, or leave it running over a whole folder overnight. Write those as three codebases and they drift, and then two analyses of the same file disagree with each other.
Approach
One set of functions with three ways in. A desktop application for working through a single run, Jupyter notebooks for the exploratory pass, and a command line for batch work. All three call the same code, so the same file gives the same numbers whichever route you take. It handles the electron-energy and ion-mass calibration, strict one-to-one electron-ion coincidences with the Stert correction for accidental pairs, and dichroism yields grouped by waveplate angle. It grew out of the notebook pipeline I had been running by hand, which is where the requirements came from.
More results
- +Generates physically plausible test data in the real file formats, so the analysis can be checked without spending beamtime on it
- +Knows how a run was marked, and analyses hardware-marked and software-marked runs the same way
- +Analysis now runs at acquisition speed; the experiment sets the pace now, not the analysis
Automated Laser Polarimetry Platform
Measures the full polarisation state of our laser automatically. What took two hours by hand now takes five minutes.
- +What took more than two hours by hand now takes about five minutes
- +Full Stokes vector with uncertainties propagated from the fit covariance matrix, not quoted without them
- +Swapping the motion controller costs a config change, not a rewrite
The full story: problem, approach, tools →Hide the full story ↑
Problem
Characterising the polarisation at each focus of a Twin-Foci ultrafast laser setup required manually rotating a motorized waveplate, reading power at each angle, and fitting data in a spreadsheet. The process took over 2 hours per dataset and was error-prone.
Approach
I specified and own a PyQt5 platform that drives the rotation stage and an Ophir NOVAII power meter together, then pulls out the full Stokes vector using the Fourier method of Schaefer et al. The stage is not hardwired into the code: the same build runs a Newport ESP301 or the newer Trinamic TMCM-6110, picked by a config setting. There is also a simulation mode, so the interface and the fitting can be tested without occupying the laser.
More results
- +Flags near-perfect circular polarisation on the Ring criterion, and reports the fit quality next to every result
- +Publication-quality export at 300 DPI in PNG, PDF, and SVG
Run the measurement yourself
The same method the platform uses, running in your browser on simulated light. A quarter-wave plate turns in steps in front of a fixed polarizer, the detector reading traces I(θ), and a Fourier fit returns the four Stokes parameters. Pick the incoming polarization, turn the plate, or run a full scan.
Method: rotating quarter-wave-plate polarimetry with Fourier analysis, after Schaefer, Collett, Smyth, Barrett and Fraher, American Journal of Physics 75, 163 (2007). The light is the platform's 309 nm ultraviolet line, which is invisible, so it is drawn in violet. The field, the waveplate and the fit are exact; only the scatter on the measured points is simulated, at 1% of the incoming intensity.
Lab Dashboard: Live Vacuum and Temperatures
Live vacuum pressures and temperatures for our lab, on any PC in the lab.
- +No gaps after a network outage: readings are kept locally and sent on later
- +A clear warning when the readings stop updating, so a stopped logger cannot pass for a good vacuum
- +Shows a capacitance gauge on the source next to the standard gauges. It reads the true pressure for any gas within its range, which matters when the source runs with water vapour; below that range it shows its zero offset
The full story: problem, approach, tools →Hide the full story ↑
Problem
The vacuum and the source temperatures could only be read on one PC in the lab. And a frozen reading looks exactly like a healthy vacuum.
Approach
I specified a small monitoring system. It reads the lab's pressure and temperature controllers, keeps every sample locally first, and shows live values and their history to anyone on the lab network.
Data and AI side projects
Built on public data, to learn machine learning end to end.
Probabilistic Solar Forecasting: Physics-Informed
Day-ahead solar generation forecasts for the German grid. Physics handles geometry; XGBoost learns only the residual.
- +Physics baseline alone: R2=0.78; Physics plus XGBoost: R2=0.92, MAE reduced 60%
- +physics_pred is XGBoost's top feature by importance; model amplifies physics, not ignores it
- +Split conformal prediction: P90 empirical coverage 0.869 against a 0.90 target; P50 under-covers after a seasonal shift (ADR-003)
The full story: problem, approach, tools →Hide the full story ↑
Problem
Germany's Energiewende targets 80% of electricity from renewables by 2030. Solar is volatile and weather-dependent. A 10% forecasting error at midday peak costs real money on the balancing market. Operators need calibrated probabilistic forecasts, not point estimates: 12 GW plus or minus 2 GW, with a coverage guarantee.
Approach
Built a two-layer physics-informed architecture on 3 years of public SMARD and Open-Meteo data (26,000+ hourly records in TimescaleDB). The physics layer uses pvlib to compute what solar output should be on a geometrically perfect day. XGBoost learns only the residual: what physics cannot see, namely clouds, curtailments, and measurement noise. P10/P50/P90 intervals via split conformal prediction. Managed as a 22-day research sprint with 5 ADRs, 3 weekly reports, 2 retrospectives, and a public Kanban board.
More results
- +CRPS=514.6 MW evaluated with reliability diagrams; calibration reported honestly
- +Full stack: TimescaleDB, FastAPI, and Streamlit dashboard starts with one command
- +GitHub Actions CI: ruff and pytest on every push; 5 Architecture Decision Records
Intelligent Predictive Maintenance System
Predicts turbofan engine remaining useful life with SHAP explainability, served through a FastAPI service with auto-generated diagnostic reports.
- +XGBoost RMSE: 18.83 cycles with a train/test gap of only 2.0 (vs 12.6 for RF with RUL clipping)
- +SHAP layer: top sensor drivers ranked per prediction with plain-language explanation
- +PDF diagnostic report auto-generated in German DIN format from both API and dashboard
The full story: problem, approach, tools →Hide the full story ↑
Problem
Industrial turbofan engines degrade gradually across 21 sensor channels. Threshold-based monitoring catches failures too late. Engineers need to know not just when failure occurs, but why: which sensors are driving the risk and what to do about it.
Approach
Built a complete end-to-end ML pipeline on NASA's CMAPSS dataset. The core insight: capping RUL targets at 125 cycles (treating early healthy cycles as interchangeable) dropped RMSE from 51 to 19, a domain decision that outperformed any algorithm choice. XGBoost with SHAP explainability, a 3-page Streamlit dashboard, FastAPI REST endpoint, and auto-generated PDF diagnostic reports in German DIN format. Everything runs with one command via Docker Compose.
More results
- +FastAPI: POST /predict returns RUL, anomaly score, and status label in under 100ms
- +Full MLflow experiment tracking: 162 logged runs, fully reproducible
- +Docker Compose: API and dashboard start with one command
In development
LangGraph · Claude API · ChromaDB
Sensor Intelligence Assistant
An AI assistant, in development, designed to reason across the predictive maintenance tools and diagnose turbofan engine anomalies from real API responses.
- +Planned: human-in-the-loop confirmation before any action
- +Tool calls go to the real FastAPI endpoints from P1
- +ChromaDB vector store for maintenance documentation retrieval
The full story: problem, approach, tools →Hide the full story ↑
Problem
An operator types: What is happening with engine 14? A useful AI system should reason across multiple tools, deciding which to call and in what order based on what it finds at each step, and cite real sources rather than inventing them.
Approach
Designed as a LangGraph state machine with three tools, two of them from P1: Tool A calls the FastAPI prediction endpoint for remaining useful life. Tool B returns SHAP-grounded sensor explanations. Tool C retrieves relevant maintenance documentation from a ChromaDB vector store. The model is meant to write a diagnostic report citing the API responses, then pause for human confirmation before any action. Read-only by design, with LangSmith tracing.
More results
- +LangSmith tracing of each step
Fellowship
How my PhD was funded, and what running it involved.
INSPIRE Fellowship Research Program
National PhD fellowship (DST, Government of India), plus hands-on funding work under my PI: fellowship reporting, a full proposal draft and project procurement.
- +Annual technical progress reports to DST, with the supervisor's continuation recommendation
- +Contributed to one granted Indian patent for gas-separation membranes
- +Procurement and closure documentation for a funded SERB project, prepared for PI certification
The full story: problem, approach, tools →Hide the full story ↑
Problem
Hold a competitive national PhD fellowship while doing the research itself, and support the group's funded projects and proposals, with my supervisor as principal investigator.
Approach
Prepared my fellowship application and the annual technical progress reports to DST. On the group's SERB-funded project I specified consumables, compared vendor quotations, prepared purchasing material and assembled annual and closure reports for the PI to certify. Drafted a full research proposal (BRNS, 2021) under PI supervision.
More results
- +Research visit to Jacobs University Bremen, Germany: Raman and UV-Vis work that led to two papers
Questions about any of these?
The side projects are on GitHub. The lab software stays with the group, but I am happy to walk you through its design and how I validated it.