Job Description
**ABOUT THE ROLE**
We are building a multi\-tenant AI agent platform: users chat with an agent that answers from their own documents, uploaded directly or synced from Google Drive and OneDrive. The platform is built and running on our own infrastructure, and we are getting it ready to launch.
We are hiring an engineer to own it end to end. You will run the product independently: build features, keep the infrastructure healthy, deploy to production and decide how the system grows. This is not a ticket\-taking role. You will be the most senior technical person on the product.
**WHAT YOU'LL DO**
\- Own the full stack: the React/TypeScript frontend, the Python/FastAPI backend, Postgres/Supabase and the Docker infrastructure.
\- Build and tune the retrieval pipeline: hybrid dense and sparse search (pgvector \+ OpenSearch), reranking, query classification and an agentic multi\-step retrieval loop.
\- Run the document ingestion pipeline: parsing (Docling), chunking, embeddings, metadata extraction and background workflows on Hatchet.
\- Maintain the cloud\-drive sync connectors (Google Drive, OneDrive) and the bring\-your\-own\-LLM feature (Claude, OpenAI/Codex, Gemini).
\- Manage LLM routing and costs through our LiteLLM gateway (OpenRouter, OpenAI, DeepInfra).
\- Run three environments (development, staging, production) on a Hetzner server with Docker Compose and self\-hosted Supabase, with deploys through GitHub Actions.
\- Keep multi\-tenant data secure: Row\-Level Security on every tenant table, plus authentication and data isolation between users and organizations.
\- Measure quality and performance: Opik/OpenTelemetry tracing, RAGAS evaluations, latency budgets and regression tests.
\- Write specs before you build (requirements → design → tasks) and keep pytest, Vitest and Playwright suites green.
\- Work effectively with AI coding agents (e.g. Claude Code). Much of our development runs on an agent\-assisted, spec\-driven workflow.
**WHAT YOU BRING (REQUIRED)**
\- 5\+ years of professional software engineering, including significant time owning production systems end to end.
\- Strong Python (3\.11\+), FastAPI or a similar async framework, and Pydantic.
\- Solid PostgreSQL: schema design, migrations, indexing and Row\-Level Security.
\- Hands\-on experience shipping LLM features to production. You call model APIs directly with raw SDKs (not LangChain/LangGraph), use structured outputs and streaming (SSE), and handle prompt, cost and latency tradeoffs.
\- Practical experience with RAG: embeddings, vector search, chunking strategies, reranking and judging retrieval quality.
\- Comfort with Linux servers, Docker and Docker Compose, CI/CD (GitHub Actions) and debugging production issues over SSH.
\- React \+ TypeScript proficiency. You can ship UI features, not just APIs.
\- Good engineering habits: automated tests, clear commits and written design docs.
\- You can work independently: you can take an ambiguous goal, break it down, decide, ship and report back without daily direction.
**NICE TO HAVE**
\- Supabase (Auth, Storage, Realtime) or self\-hosting it.
\- OpenSearch/Elasticsearch, BM25 or hybrid search tuning.
\- Workflow engines such as Hatchet, Temporal or Celery.
\- Document parsing (Docling, Unstructured, OCR) for PDFs and Office files.
\- LLM observability and evaluation (Opik, Langfuse, RAGAS, OpenTelemetry).
\- Google Drive / Microsoft Graph APIs and OAuth flows.
\- Multi\-tenant SaaS security and data isolation.
\- Startup or founding\-engineer experience.
**OUR STACK**
Frontend: React, Vite, TypeScript, Tailwind, shadcn/ui
Backend: Python 3\.12, FastAPI, Pydantic, Hatchet workers
Data: self\-hosted Supabase (Postgres \+ pgvector), OpenSearch
AI: LiteLLM gateway, OpenRouter / OpenAI / DeepInfra, Qwen3 embeddings, ZeroEntropy reranking
Infra: Hetzner, Docker Compose, GitHub Actions, Tailscale
Quality: pytest, Vitest, Playwright, ruff, pyright, Opik
**WHAT WE OFFER**
\- Real ownership of a modern AI product, with a direct line to the founder.
\- Freedom to make architecture decisions and shape the roadmap.
Pay: ₹1,200,000\.00 \- ₹2,400,000\.00 per year
Benefits:
* Health insurance
* Paid sick time
* Paid time off
* Work from home
Experience:
* RAG: embeddings, vector search, chunking, reranking : 2 years (Required)
Work Location: Hybrid remote in Pune, Maharashtra