Allauddin Shaik · Data Engineer

Reliable data pipelines.Orchestrated and monitored.Built and run in production.

Data architectureOrchestration (Airflow)ETL / ELTSpark (PySpark)Lakehouse (Iceberg · Parquet)Cloud (GCP · Azure)APIs (FastAPI)RAG & evaluation

What I've shipped

Products

Invoice Audit Pipeline

Try it live →No signup needed

The problem

Overcharges and tampered invoices slip through manual review, and money leaks quietly.

The solution

Upload an invoice: every line is checked against your agreed rates, the document inspected for tampering, and you're told exactly what to dispute.

The result

Sample run: 24 line items audited against base rates, $377.80 in overcharges flagged in seconds.

Built on FastAPI · Cloud Run · Firestore · GCS · React

Capabilities

Rate audit

Every line item checked against your agreed rates; overbilling surfaces automatically.

Tamper detection

Catches edits, altered figures, and regenerated pages in the document itself.

Ask any invoice

Plain-language answers about any charge, drawn from the invoice's own data (RAG).

Answers you can check

Every answer cites the line it came from, and the system says it doesn't know rather than guessing. Flag a bad answer and the retrieval behind it is traceable.

Want this hands-off? I can build the full orchestration layer: the moment an invoice lands in your inbox, it's audited automatically and your team is flagged the instant something looks wrong or fraudulent.

Shipping soon

Reddit Radar

Web · Soon

Market & keyword research that mines Reddit to find what a niche actually wants.

Clinical Flow

Web · Soon

Turning core clinical datasets into decisions care teams can act on.

Track record

Systems that run in production today

Dentsu Global Services

Data Engineer

Remote · 2022–present

Owning the full lifecycle: data pipelines → backend & APIs → application → deployment.

80%

less time on manual reporting once the pipelines took over the recurring builds

70%

cut in annual service costs via a data & dashboard audit

24/7

pipelines running unattended in production on GCP & Azure

  • Own a Bayesian media-simulation platform end to end: data architecture, serverless orchestration, backend services, and the React frontend. Took it from analysts loading CSVs and running scripts by hand to a product external clients now run their own budget plans on.
  • Split orchestration deliberately by workload: scheduled data feeds on Apache Airflow, user-triggered optimization runs as asynchronous Cloud Run Jobs behind a service, since that work only runs when someone asks for it and an always-on control plane bills whether or not anything runs.
  • Designed the analytics data models (star schema over contribution, adstock and efficacy-ratio facts) alongside the normalized transactional schema, rebuilt in line with ingestion so summaries cannot drift from source.
  • Caught and closed a silent data-corruption class: unregistered model variables resolved to NULL on join, itself a reserved business value in this domain, so bad references became valid-looking attribution rather than errors. Enforced a fail-fast ingestion contract that rejects unresolvable keys and notifies the model owner.
  • Rolled the platform out across three regions in the EU and Middle East from a single codebase, with per-region databases, service accounts and services wired at deploy time through environment config. Data residency requirements drove the design, and access is enforced on two axes: region isolation in the infrastructure, per-client scoping in the application.
  • Automated the platform's model builds end to end, and extended the Bayesian simulation logic to support comparative scenario analysis (beyond the original forward-projection model) to meet a new business requirement.
  • Built a natural-language query interface (text-to-SQL over the product database) so teams get answers from executed queries in seconds instead of waiting on an analyst, with figures coming from the query rather than from a model.
  • Led the migration off a legacy R and R Shiny codebase to a Python backend and React front end, defining the API contract in Swagger ahead of implementation so the front-end team could build against stubs in parallel. CI/CD runs unit and integration tests on every change, promoted through dev, pre-production and production behind a business validation gate.

What I bring

How I work

I don't hand off pieces. I take a problem from raw data to deployed and used: scoped with you, shipped in small increments, and owned through deployment and monitoring.

Data Engineering & Platform

The foundation: data that moves reliably, on infrastructure that runs itself.

  • Batch and event-driven pipelines, orchestrated with Airflow, running as serverless jobs on GCP and Azure
  • Data architecture: Postgres and cloud storage underneath, with the analytics models (OLAP) that make the data usable
  • Distributed processing and lakehouse storage: Spark for work that outgrows one machine, Parquet and Iceberg underneath so tables version, evolve their schema, and stay readable by any engine
  • ML infrastructure: automated model runs, output pipelines, and monitoring around the models; the data scientists own the algorithms
  • Deployment ownership: Docker, CI/CD, and monitoring treated as part of the build, not an afterthought
  • APIs and the serving layer: FastAPI backends on top of the data, contract-first so front-end teams can build in parallel, plus production RAG where it adds value
  • Making AI output trustworthy: grounding rules, citation checks, and evaluation harnesses, because an answer nobody can verify is not an answer

Working together

Step 1

We talk

You describe the problem in a 30-minute call. I tell you honestly what will actually solve it, and what it would take.

Step 2

I design and build

A clear plan first, then the system: shipped in small increments you can see, not a black box that goes quiet for weeks.

Step 3

It runs in production

Deployed, documented, monitored. I operate what I build: no handover pain, no babysitting needed.

Contact

Open to new roles, and open for collaboration.

If you're hiring, or you have a system you want built, let's talk.

Based in India · open to relocation, or remote worldwide

Allauddin Shaik