Hrishikesh JadhavAI Engineer
I build LLM extraction pipelines, RAG and MCP systems, and the evaluation gates that decide what ships.

I'm Hrishi, and at Toku I built the LLM extraction pipeline, RAG assistant and MCP server that run in daily production inside a global payroll platform. Before that I was a data scientist at GfK / NIQ in Nuremberg, building LLM classification, semantic retrieval and prediction models and running them in production with Python, AWS, Docker and GitLab CI/CD.
4+ years shipping production AI and backend systems in Germany.
Experience
Internal AI tooling
- Shipped AI tooling into daily production across product, engineering and sales teams: an MCP server exposing internal systems to LLM clients, a RAG assistant (FastAPI, LangChain, pgvector), and n8n automations for onboarding, payslip processing and support triage. Owned deployment, documentation, onboarding and hands-on user support.
- Moved security review to pull-request time across thousands of pull requests by integrating automated LLM-assisted security review into GitHub CI for Toku's financial platform.
LLM pipelines and evaluation
- Cut payroll data intake from hours of manual work per cycle to minutes across production records by building a self-hosted LLM extraction pipeline (Python, FastAPI, PostgreSQL, DigitalOcean) that masks personal data before inference.Self-hosted so that no employee data leaves controlled infrastructure.
- Kept field-level extraction accuracy above the agreed threshold across every extracted field, checked against a large golden evaluation set, by building the evaluation harness and a CI regression gate that blocks any deploy below it.So a change that lowers extraction accuracy cannot reach payroll data.
- Brought LLM inference costs down to a fraction of the estimated cost of equivalent third-party API usage by deploying open-weight models on self-hosted GPU infrastructure.
Personal data is masked before inference, and a golden-set gate decides what deploys. Customer-facing platforms
- Built status.toku.com, Toku's public status platform: production components probed from the Cloudflare edge on a short fixed interval, automated incident detection and resolution, dependency rollups, edge-vs-origin failure isolation, and a JSON API and RSS feed. New incidents notify the team on Slack, and failures that need a code fix are picked up by a coding agent.Probing the edge and the origin separately localises a failure to the CDN or the application before anyone looks at it.
- Built Toku's customer-facing help centre and AI assistant (toku.com/help, Next.js) on Toku's Notion knowledge base, covering payroll, benefits, policy and country-guide articles across dozens of countries, with general-help and visa-support content in separate retrieval corpora and multi-model fallback across providers.Separate corpora so immigration and work-authorisation questions never resolve against general payroll content.
- Reduced manual effort in product-taxonomy classification by around 30% by building an LLM classification service on Amazon Bedrock using prompt engineering and RegEx-based post-processing.
- Reduced catalogue lookup time by around 40% versus keyword search by building vector-embedding semantic retrieval over internal product catalogues.
- Improved simultaneous-viewer prediction accuracy by around 20% over the incumbent model by training and deploying a CatBoost model using 30+ engineered features from sociodemographics, temporal patterns and programme metadata.
- Kept production scoring pipelines running 24/7 without manual intervention by orchestrating S3 ingestion, feature generation, model scoring and automated integration tests in GitLab CI/CD, containerised with Docker on Linux.
- Built a TV-audience data-fusion pipeline using K-Nearest Neighbors and a genetic algorithm to match and clone panel households against large-scale return-path data.
- Built COVID-19 decision-support models in Python (SEIRD, scikit-learn, SciPy) and forecasting dashboards (Plotly, AWS) used to inform Government of India lockdown and testing-strategy decisions, with a 4.8% RMSE reduction against the prior baseline.
Research
2025
Ontology Evolution in Invasion Biology Using Large Language Models: A Hybrid Approach
Hrishikesh Jadhav, Tina Heger, Birgitta König-Ries, Alsayed Algergawy
LLM-TEXT2KG 2025, CEUR Workshop Proceedings Vol. 4020, pp. 195-206
A hybrid pipeline that combines GPT-4 prompting and zero-shot extraction with classical ontology engineering to build and evolve INBIO, a core ontology for invasion biology, with domain experts validating new classes.
2021
A Deep Learning Mobile Application based Sign Language Recognition for Aphasic Person
Hrishikesh Jadhav, Pushkar Dounde, Akash Pawar, Abhishek Muthange
Journal of Emerging Technologies and Innovative Research (JETIR)
An Android app that recognises sign-language gestures using Histogram of Oriented Gradients features with CNN and multiclass SVM classifiers.
Awards
- 2024
NIQ/GfK HACKFEST
Top 3, with a RAG learning assistant built and demoed in 24 hours
- 2024
BMW Innovation Challenge
Selected participant, 24-hour challenge at BMW iFactory Dingolfing (DocCheck use case)
- 2020
IEEE Machine Learning Hackathon
1st place
- 2020
HackCovid-19
Winner among 130 teams
- 2017
Smart India Hackathon
Winner (Ministry of Defence)
Contact
knowhrishi.de@gmail.comOpen to AI engineering roles in Germany from October 2026. EU Blue Card holder, authorised to work in Germany.