Shubham Shinde

Shubham Shinde

Research software and machine learning engineer, Bristol

I build scientific and machine learning software, and I check whether its numbers can be trusted.

Most of my work comes down to one question: is a result a real effect, or an artefact of the code that produced it? Answering that well takes controlled test cases, honest uncertainties, and software a colleague can run without me in the room.

MSc, University of BristolScientific Computing with Data Science, completed 2026, Distinction expected

RSECon26 speakerLightning talk on silent numerical failure, Sheffield, September 2026

Available nowFor research software, scientific computing and machine learning roles in the UK

Selected work

Each result below is marked by where it stands. I would rather show an unfinished result honestly than a finished-looking one that isn't.

Measured result In progress or not yet validated

MSc dissertation, University of Bristol, 2025 to 2026

Why kinisi’s covariance matrices fail without warning

kinisi extracts diffusion coefficients from molecular dynamics trajectories. It does something unusually careful: it carries the full covariance matrix through its fit instead of treating each lag time as independent. That care is also the source of a fragility. In a minority of runs the matrix stops being positive definite, which quietly degrades the fitted coefficient and its uncertainty without raising any error.

Rather than patch the symptom, I built a random-walk testbed with a known analytical answer, so that a numerical artefact could be separated from a real physical signal. The work is supervised by Dr Andrew McCluskey, one of kinisi’s developers.

  • Matrix failure rate fell from 90% to 2% as sampling increased, showing the matrices were under-determined rather than wrong.
  • Root cause isolated to sample noise in long lag-time variance estimates, through a systematic convergence study.
  • Minimum-eigenvalue, ridge, linear and non-linear shrinkage reconditioning benchmarked against the analytical matrix.
  • A data-adaptive eigenvalue floor recovered parameters better than the library’s fixed default on the testbed.
  • Validation across noise levels is ongoing, and the adaptive rule has only been tested offline.
  • A consistency check is open upstream as pull request #222, awaiting maintainer review.

Python, NumPy, MDAnalysis, eigenvalue analysis, matrix conditioning

Generalised least squares needs the inverse covariance

$$\hat{\beta} = \left(X^{\top}\Sigma^{-1}X\right)^{-1}X^{\top}\Sigma^{-1}y$$

Σ is the MSD covariance. Inversion fails once the smallest eigenvalue reaches zero.

eigenvalue index, largest to smallest log λ numerical floor converged sampling low sampling
Schematic. Under low sampling the trailing eigenvalues drop below the floor and the matrix stops being invertible. The failure is in the estimate, not the physics.

Author and maintainer, 2025 to present

An open source library for quasi-elastic neutron scattering

A library I designed and wrote that takes the full analysis chain from raw HDF5 files through resolution correction, model fitting and Bayesian parameter estimation, with a FastAPI and React interface on top. It follows ISIS Neutron and Muon Source workflows and is validated on benzene data from the LET spectrometer.

The core model is Chudley–Elliott jump diffusion, which describes molecules hopping between sites. Getting the deconvolution of the instrument resolution right was the decisive problem, because a naive approach inflates the diffusion coefficient substantially.

  • About 16% less systematic error on the diffusion coefficient from physics-correct resolution correction.
  • 3× better χ² goodness of fit for Chudley–Elliott than a Fickian diffusion model.
  • Four independent MCMC chains, convergence verified at Gelman–Rubin R̂ < 1.01.
  • Calibrated uncertainties on every fitted parameter rather than bare point estimates.

Python, NumPy, SciPy, h5py, emcee, pytest

incident neutrons sample detectors Q each detector angle gives one spectrum S(Q, ω) the linewidth against Q carries the diffusion coefficient
Schematic. Quasi-elastic neutron scattering geometry, the measurement this library analyses.

Chudley–Elliott linewidth

$$\Gamma(Q)=\frac{D\,Q^2}{1+Q^2 l^2}$$

D is the diffusion coefficient and l the jump length.

# Bayesian parameter estimation with emcee
sampler = emcee.EnsembleSampler(
    nwalkers, ndim, log_posterior,
    args=[Q_vals, S_exp, R_exp])
sampler.run_mcmc(p0, 5000, progress=True)

Personal project, 2026 · 204 papers indexed, 224 tests, CI on every push

A retrieval system that refuses to answer without evidence

A retrieval-augmented generation system over machine learning papers, which turned into a question about trust. During testing I asked it something with no answer in the corpus. Instead of abstaining, it answered confidently from the model’s own knowledge and attached a real citation to a passage that didn’t support the claim.

The hallucination wasn’t the surprise. The citation was: it made a failure look like success. So I built a deterministic defence instead of a prompt-based one. Every answer must quote its source verbatim, the quote is checked against the retrieved text, and the system refuses when no citation survives. Then I built an adversarial test set to try to break it.

  • The defence held even when retrieval was deliberately steered into the trap passage, because it checks real text rather than the model’s account of it.
  • Cross-encoder reranking lifted MRR@10 from 0.801 to 0.955 across six ablated configurations.
  • Multi-query expansion gave no improvement.
  • HyDE lowered accuracy by 5 points on this corpus. Both negative results are published in the repository.

Python, LanceDB, BM25 and dense hybrid retrieval, cross-encoder reranking, pytest, mypy strict, GitHub Actions

query hybrid retrieval generator verbatim quoteverification answer with averified citation refuse to answer,no support found passfail
The verification gate. An answer ships only if at least one citation survives verbatim checking against the retrieved text.

Open source, MIT

scisolve: numerical answers that have to be computed, not generated

An agentic system that solves ODE, optimisation, curve-fitting and linear algebra problems by running NumPy and SciPy code. No number can be reported unless it traces back to a value the session actually computed, enforced in code rather than by prompting, and each answer is checked by one of six independent verification methods. 122 tests on Python 3.10 to 3.13.

Five of seven milestones complete. Tested through a scripted harness; not yet run against a live model API.

Performance study, 2025

Lebwohl–Lasher: seven implementations of the same physics

A Monte Carlo liquid crystal simulation taken from pure Python through NumPy, Numba, Cython, OpenMP and MPI to hybrid MPI with Cython. The fastest version runs 4.7 times faster than the vectorised NumPy one. The part that mattered most was the cross-version tests confirming every faster version still produced the same physics.

Distributed computing, 2026

Distributed ATLAS Higgs analysis

A reproduction of the ATLAS H → ZZ* → 4ℓ Higgs analysis that reaches the official 4.5σ, distributing 40 ROOT files across horizontally scaled Docker workers through RabbitMQ. Built for failure: exponential-backoff retries, a dead-letter queue, coordinator checkpointing and Prometheus metrics.

Physics-informed machine learning

Where machine learning fills in missing physics

The Washburn equation predicts capillary rise almost exactly when pore radius is known. Inverting it into a pore-radius feature lifted a random-forest classifier from 81% to 96%, and a hybrid pipeline raised R² from 0.79 to 0.985. The model supplies the missing input rather than replacing the physics.

Earlier work

Lunar terrain classifier
A hybrid CNN and Transformer classifier reaching 89.7% accuracy on NASA lunar terrain data, packaged with Docker, DVC and CI so every result could be reproduced from a single command. PyTorch, OpenCV.
Screening data portal
DICOM image ingestion and structured report storage for cervical cancer screening, with role-based access control protecting patient confidentiality and tests across every API route. Flask, PostgreSQL, Docker.

Experience

Three years across a defence research programme, a machine learning company and an engineering firm, plus a year of robotics research.

  1. Jun 2024 to Aug 2025Pune, India

    Data scientist and project engineer

    ARDE research programme, Raj Security and Facility Management

    • Designed a platform of FastAPI microservices with authentication and versioned endpoints, serving processed sensor data to dashboards at typically sub-100 ms response times on a restricted offline network.
    • Built a PySpark and SLURM pipeline ingesting over 20 GB of sensor data a day, with failure recovery, schema validation and audit logging into PostgreSQL for full traceability from raw measurement to result.
    • Built CI/CD on GitHub Actions and Docker that cut release and validation time from hours to under ten minutes.
    • Trained models on multi-GPU compute; an XGBoost and LightGBM ensemble reduced prediction error by roughly 42% against the legacy analytical baseline.
    • Provided on-call production support, including fault diagnosis, root-cause analysis and post-mortems for services scientists relied on.
  2. Jun 2023 to May 2024Satara, India

    Robotics research experience

    Rayat Science and Innovation Activity Centre

    • Developed AI features for educational robotics projects, integrating machine learning into robotic systems.
    • Trained machine learning models for object recognition from robot sensor data, and implemented Python algorithms for path planning and navigation.
    • Built pipelines to process real-time sensor data, tuned models through experiments, and demonstrated the systems to researchers and students.
  3. Feb 2023 to May 2023Remote

    Machine learning research engineer

    Spartifical

    • Developed a hybrid CNN and Transformer classifier reaching 89.7% accuracy on NASA lunar terrain data.
    • Built a full-stack data service with FastAPI inference endpoints and a React monitoring dashboard, with pytest and Jest suites at over 90% coverage.
    • Delivered a reproducible research artefact with Docker, DVC and CI, so colleagues could replicate every result with one command.
  4. Nov 2020 to May 2022Satara, India

    Design engineer

    Gurudatta Enterprises

    • Built 3D models, assemblies and engineering drawings in SolidWorks and CATIA V5, with tolerance stack-up analysis across more than 200 prototype test cycles.
    • Automated test-bench data logging in Python, cutting post-processing time by roughly 60%.

Skills

Scientific computing

Python, NumPy, SciPy, HDF5 and h5py, MDAnalysis, emcee, Matplotlib

Numerical methods

Bayesian MCMC, numerical linear algebra, Monte Carlo methods, uncertainty quantification, model validation

High-performance computing

SLURM, MPI, OpenMP, Cython, Numba, multi-GPU training, PySpark, Dask

Machine learning

PyTorch, scikit-learn, XGBoost, LightGBM, OpenCV, MLflow, DVC, retrieval-augmented generation

Engineering practice

FastAPI, PostgreSQL, Docker, Kubernetes, RabbitMQ, Prometheus, GitHub Actions, pytest, mypy, Git

Languages

Python, SQL, Bash, JavaScript and TypeScript, MATLAB

Publications and talks

  • Conference talk, 2026

    When a covariance matrix silently breaks your diffusion analysis

    Lightning talk, RSECon26, the UK Research Software Engineering Conference, University of Sheffield, September 2026

  • Journal article, 2023

    Building shadow detection using aerial imagery with attention-enhanced YOLOv4

    Shinde, S., et al. JETIR, volume 10, issue 9. Read the paper · Best Research Paper Award, Savitribai Phule Pune University

Education

  • 2025 to 2026Bristol, UK

    MSc Scientific Computing with Data Science

    University of Bristol · Distinction expected

    Dissertation on numerical stability in diffusion analysis, with coursework in high-performance computing, numerical methods and statistical data analysis. Funded by the SARTHI Maharashtra Government Foreign Scholarship.

  • 2022 to 2023Pune, India

    PG Diploma in Data Science and Artificial Intelligence

    Savitribai Phule Pune University · GPA 8.5 / 10

    Thesis on building shadow detection from aerial imagery, which received the Best Research Paper Award.

  • 2017 to 2020Pune, India

    BE Mechanical Engineering

    Savitribai Phule Pune University

    Capstone on optimising wind turbine blades using CFD. University Merit Scholarship, 2018.

About

I trained first as a mechanical engineer, and partway through simulating wind turbine blades in CFD realised the problems I cared about most were the computational ones. That led to a PG Diploma in data science, research in aerial image analysis, and an MSc at Bristol.

My current work sits in neutron scattering analysis and inside kinisi, an established scientific library. Both taught me the same thing from opposite directions: the interesting decisions in scientific software are usually about what the code refuses to approximate. I’m also quietly convinced that a good README is an act of kindness.

Service

  • Student academic representative, MSc cohort, University of Bristol
  • Digital platforms for the Pune chapters of the Aeronautical Society of India and the Indian Society of Systems for Science and Engineering

Looking for someone to build software you can trust?

I’m open to research software, scientific computing and machine learning engineering roles, particularly where computational science and production software meet.

Shubham Shinde, Bristol, 2026 Set in Newsreader and IBM Plex Sans