Allauddin Shaik · Data Engineer
Reliable data pipelines.Orchestrated and monitored.Built and run in production.
What I've shipped
Products
Invoice Audit Pipeline
The problem
Overcharges and tampered invoices slip through manual review, and money leaks quietly.
The solution
Upload an invoice: every line is checked against your agreed rates, the document inspected for tampering, and you're told exactly what to dispute.
The result
Sample run: 24 line items audited against base rates, $377.80 in overcharges flagged in seconds.
Built on FastAPI · Cloud Run · Firestore · GCS · React
Capabilities
Rate audit
Every line item checked against your agreed rates; overbilling surfaces automatically.
Tamper detection
Catches edits, altered figures, and regenerated pages in the document itself.
Ask any invoice
Plain-language answers about any charge, drawn from the invoice's own data (RAG).
Answers you can check
Every answer cites the line it came from, and the system says it doesn't know rather than guessing. Flag a bad answer and the retrieval behind it is traceable.
Want this hands-off? I can build the full orchestration layer: the moment an invoice lands in your inbox, it's audited automatically and your team is flagged the instant something looks wrong or fraudulent.
Shipping soon
Reddit Radar
Web · SoonMarket & keyword research that mines Reddit to find what a niche actually wants.
Clinical Flow
Web · SoonTurning core clinical datasets into decisions care teams can act on.
Track record
Systems that run in production today
Dentsu Global Services
Data Engineer
Remote · 2022–present
Owning the full lifecycle: data pipelines → backend & APIs → application → deployment.
80%
less time on manual reporting once the pipelines took over the recurring builds
70%
cut in annual service costs via a data & dashboard audit
24/7
pipelines running unattended in production on GCP & Azure
- Own a Bayesian media-simulation platform end to end: data architecture, serverless orchestration, backend services, and the React frontend. Took it from analysts loading CSVs and running scripts by hand to a product external clients now run their own budget plans on.
- Split orchestration deliberately by workload: scheduled data feeds on Apache Airflow, user-triggered optimization runs as asynchronous Cloud Run Jobs behind a service, since that work only runs when someone asks for it and an always-on control plane bills whether or not anything runs.
- Designed the analytics data models (star schema over contribution, adstock and efficacy-ratio facts) alongside the normalized transactional schema, rebuilt in line with ingestion so summaries cannot drift from source.
- Caught and closed a silent data-corruption class: unregistered model variables resolved to NULL on join, itself a reserved business value in this domain, so bad references became valid-looking attribution rather than errors. Enforced a fail-fast ingestion contract that rejects unresolvable keys and notifies the model owner.
- Rolled the platform out across three regions in the EU and Middle East from a single codebase, with per-region databases, service accounts and services wired at deploy time through environment config. Data residency requirements drove the design, and access is enforced on two axes: region isolation in the infrastructure, per-client scoping in the application.
- Automated the platform's model builds end to end, and extended the Bayesian simulation logic to support comparative scenario analysis (beyond the original forward-projection model) to meet a new business requirement.
- Built a natural-language query interface (text-to-SQL over the product database) so teams get answers from executed queries in seconds instead of waiting on an analyst, with figures coming from the query rather than from a model.
- Led the migration off a legacy R and R Shiny codebase to a Python backend and React front end, defining the API contract in Swagger ahead of implementation so the front-end team could build against stubs in parallel. CI/CD runs unit and integration tests on every change, promoted through dev, pre-production and production behind a business validation gate.
What I bring
How I work
I don't hand off pieces. I take a problem from raw data to deployed and used: scoped with you, shipped in small increments, and owned through deployment and monitoring.
Data Engineering & Platform
The foundation: data that moves reliably, on infrastructure that runs itself.
- Batch and event-driven pipelines, orchestrated with Airflow, running as serverless jobs on GCP and Azure
- Data architecture: Postgres and cloud storage underneath, with the analytics models (OLAP) that make the data usable
- Distributed processing and lakehouse storage: Spark for work that outgrows one machine, Parquet and Iceberg underneath so tables version, evolve their schema, and stay readable by any engine
- ML infrastructure: automated model runs, output pipelines, and monitoring around the models; the data scientists own the algorithms
- Deployment ownership: Docker, CI/CD, and monitoring treated as part of the build, not an afterthought
- APIs and the serving layer: FastAPI backends on top of the data, contract-first so front-end teams can build in parallel, plus production RAG where it adds value
- Making AI output trustworthy: grounding rules, citation checks, and evaluation harnesses, because an answer nobody can verify is not an answer
Working together
Step 1
We talk
You describe the problem in a 30-minute call. I tell you honestly what will actually solve it, and what it would take.
Step 2
I design and build
A clear plan first, then the system: shipped in small increments you can see, not a black box that goes quiet for weeks.
Step 3
It runs in production
Deployed, documented, monitored. I operate what I build: no handover pain, no babysitting needed.
Contact
Open to new roles, and open for collaboration.
If you're hiring, or you have a system you want built, let's talk.
Based in India · open to relocation, or remote worldwide
