By Lokesh, Lead FDE Trainer at FDE Masters · Updated 7 October 2026 · 11 min read
What do forward deployed engineers build? Eight things, repeatedly: permission-aware enterprise RAG, agents with a human in the loop, MCP servers into ServiceNow, Salesforce and SAP, eval suites wired into CI, SSO integration, deployments inside the customer’s VPC, security evidence packs, and the SOW and executive demos around them. This post explains each for developers, with a worked NBFC example.
Key takeaways
- FDE deliverables are integration artefacts, not models. Anthropic’s FDE job description names MCP servers, sub-agents and agent skills; OpenAI’s engagements gate delivery on evals.
- The hard parts are permissions, identity and evidence, not prompting.
- The Deccan Finance capstone builds all eight for one fictional NBFC.
What are the deliverables of a forward deployed engineer?
A forward deployed engineer (FDE) builds and deploys software inside a customer’s environment and owns whether it gets used; Palantir created the role and OpenAI, Anthropic and most AI vendors now run it. For the history and career view, start with what a forward deployed engineer is. This post is about the output.
The output is rarely a model. It is the layer between a model and a business: retrieval over the customer’s documents, agents over their workflows, connectors into their systems, and the tests and documents that let a security team and an executive say yes. The Anthropic FDE job description lists MCP servers, sub-agents and agent skills as the job; the Pragmatic Engineer describes OpenAI engagements moving from scoping to eval-based validation to delivery.
| Artefact | What it is | Why the customer needs it | Typical tools |
|---|---|---|---|
| Enterprise RAG | Retrieval over internal documents, respecting who may see what | Answers grounded in their data, no leaks | Postgres + pgvector, hybrid search, rerankers |
| Agents with human in the loop | Multi-step workflows that pause for approval before acting | Automation without unreviewed writes | LangGraph, Claude Agent SDK |
| MCP servers and connectors | Standard tool interfaces into ServiceNow, Salesforce, SAP, databases | Models can act on real systems safely | MCP, OAuth2, developer sandboxes |
| Eval suites and CI gates | Golden datasets and judges that block bad deploys | Proof it works, protection against regressions | promptfoo, RAGAS, Langfuse, GitHub Actions |
| Identity and SSO | OIDC or SAML login, role mapping, RBAC | Only the right people, with the right rights | OIDC, SAML, customer IdP |
| Deployment inside the VPC | The system running in the customer’s cloud or on-prem | Data residency, network control | Terraform, Kubernetes, AWS, model gateways |
| Security evidence pack | Documents and test results for a security review | Sign-off from security and compliance | OWASP LLM Top 10 checks, DPDP, SOC 2 mappings |
| SOW and executive demos | Scope, success metrics, and the demo that proves them | Budget, alignment, expansion | Scoping PRD, pyramid-principle notes |
1. Permission-aware enterprise RAG
Retrieval-augmented generation (RAG) fetches relevant chunks of the customer’s documents and gives them to the model as context. The enterprise version must answer a harder question first: which chunks is this specific user allowed to see?
Permission-aware retrieval means every chunk carries the access control list of its source document, and the query filters on the caller’s identity before similarity search, not after. Add hybrid search (keyword plus vector) because enterprise queries contain product codes that embeddings handle badly, and a reranker to lift precision.
What you hand over: an ingestion pipeline with versioning and dedup, a pgvector schema with ACL columns, a retrieval API, and a golden set of 50 to 100 questions the customer helped write. The Brolly RAG Chatbot product follows this pattern; document versioning, not the model, was the hardest part of each deployment.
2. Agents with a human in the loop
An agent is a loop where the model decides which tool to call next until a task is done. In an enterprise the task usually ends in a write: a ticket, a record update, an email. Human-in-the-loop (HITL) means the agent pauses before that write, shows a person what it intends to do, and proceeds only on approval.
The engineering is in the pause: persisted state so approval can arrive hours later, enough context for the approver to decide in one screen, and guardrails such as allowed tools per role, spend limits and a step cap. What you hand over: a LangGraph or Claude Agent SDK workflow with checkpointing, an approval step inside the customer’s existing tool, and Langfuse traces for every run. Sub-agents, where a supervisor delegates bounded tasks to specialists, are the artefact Anthropic’s job description names.
3. MCP servers and connectors (ServiceNow, Salesforce, SAP)
MCP (Model Context Protocol) is an open standard for exposing tools and data to models. An MCP server wraps a system, say ServiceNow, and publishes typed tools such as “search incidents” or “create change request”, so any MCP-capable client, including Claude, can use it without custom glue.
Building one for an enterprise means handling what the quickstart skips: OAuth2 against the customer’s tenant, per-user credentials rather than one service account, rate limits, and tool descriptions precise enough that the model does not call “delete” when it means “close”. SAP and legacy ERPs add SOAP and batch-file interfaces that need a thin adapter underneath.
What you hand over: one MCP server per system, tested against the vendor’s developer sandbox, with a tool-level permission matrix and an audit log.
4. Eval suites and CI gates
An eval suite is the test suite for an LLM system: a golden dataset with expected outputs, scorers (exact match, LLM-as-judge rubrics, RAGAS metrics such as faithfulness), and thresholds. A CI gate runs it on every change and blocks the deploy when a threshold fails.
This turns “does it work?” into a number both sides trust, which is why OpenAI validates with evals before delivery. What you hand over: the golden set in the customer’s repo, promptfoo or RAGAS configuration, a GitHub Actions job that fails below threshold, and Langfuse dashboards so production traces become new eval cases.
Want to become a Forward Deployed Engineer in Hyderabad?
The FDE Career Program builds all eight artefacts for one client, Deccan Finance, across four graded projects and a live 2-week forward deployment. Classroom near JNTU Metro or live online.
5. Identity and SSO integration
Nothing ships in an enterprise without single sign-on. The FDE wires the application to the customer’s identity provider over OIDC or SAML, maps groups to roles, and enforces role-based access control (RBAC) on every endpoint and every retrieved chunk.
Identity integration often takes longer than the AI work, which surprises most new FDEs. What you hand over: a working login against the customer’s IdP, a role matrix signed off by their IT team, and tests proving a user in group A cannot retrieve group B’s documents.
6. Deployment inside the customer’s VPC
Most enterprise customers, and every regulated one, want the system in their own cloud account or data centre. The FDE writes Terraform for network, databases and compute, runs containers on Kubernetes or a managed equivalent, and routes model calls through a gateway the customer controls.
You rarely have admin rights, so changes go through the customer’s pipeline and change process. What you hand over: Terraform modules, Kubernetes manifests, a CI/CD pipeline the customer’s team can run, observability dashboards, and a runbook for the on-call engineer who is not you.
7. Security evidence packs
Before go-live a security team will ask for proof. An evidence pack answers in one folder: architecture and data-flow diagrams, prompt-injection and exfiltration test results mapped to the OWASP LLM Top 10, a PII statement aligned with the DPDP Act 2023 and any SOC 2, GDPR or HIPAA controls, dependency and secrets scans, and the access matrix.
FDEs who arrive without this lose weeks; FDEs who bring it get a faster yes. At FDE Masters the pack is reviewed by SOC Masters faculty before every deployment.
8. SOW, status notes and executive demos
The last artefacts are documents, and they decide whether the engineering gets funded and extended. A scoping PRD turns a vague request into a bounded problem with a success metric; a statement of work (SOW) fixes scope, timeline and acceptance criteria. Weekly 150-word status notes keep the sponsor informed, and the executive demo proves the metric was hit and sets up the next phase.
55% of FDE postings in Bloomberry’s 1,000-posting analysis list customer work before any technology; these documents are that work made concrete. Our day in the life of a forward deployed engineer shows where they fit in a working day.
Worked example: the Deccan Finance capstone
Deccan Finance is the fictional mid-size NBFC that every FDE Masters exercise runs against. The brief: loan officers spend hours answering policy questions from 400 internal documents, and the collections team wants help drafting follow-ups. Here is how the eight artefacts come together.
- Discovery and scoping (Customer Engagement Lab). Two discovery calls with faculty playing the head of operations and the CISO. Output: a scoping PRD with one metric, answer accuracy above 85% on a 60-question set the client owns, and a phase-one SOW.
- Permission-aware RAG. Ingest the 400 documents, including duplicated policy versions and scanned circulars. Each chunk carries a department ACL, so branch staff cannot retrieve credit-committee minutes. Hybrid search plus a reranker on pgvector.
- Evals as the gate. The 60-question golden set plus 20 adversarial questions runs in promptfoo and RAGAS on every pull request; below 85% faithfulness, GitHub Actions blocks the deploy.
- Agent with HITL for collections. A LangGraph workflow reads an overdue account, drafts a follow-up in English or Telugu, and pauses. A collections officer approves, edits or rejects in a small Next.js screen.
- MCP servers. One over a mocked loan management system (SQL Server) and one over a ServiceNow developer sandbox so the agent can open a ticket on a disputed balance. Per-user OAuth2, tool-level permissions, audit log.
- Identity. OIDC login against a sandbox identity provider, groups mapped to branch, credit and collections roles, tests proving cross-role retrieval fails.
- Deployment inside the VPC. Terraform for a private subnet, RDS Postgres, containers on Kubernetes, a model gateway with per-team cost caps, a GitHub Actions pipeline and a runbook.
- Evidence pack and executive demo. OWASP LLM Top 10 results, DPDP Act 2023 PII statement, access matrix and data-flow diagram reviewed by SOC Masters faculty, then a 15-minute demo to the COO in English and in Telugu or Hindi, leading with the accuracy number and ending with the phase-two proposal.
Weeks 17 and 18 repeat the pattern with a real customer, a Brolly Group business unit or a Digital Brolly SME client. The week-by-week sequence is on the forward deployed engineer curriculum page.
How do you learn to build these?
In this order: Python, SQL and APIs; cloud and containers; LLM fundamentals; RAG; evals; agents and MCP; integration and identity; security; then customer craft. Our forward deployed engineer skills guide covers each layer with specific tools.
The FDE Career Program walks beginners through all eight artefacts in 18 weeks for ₹45,000 online or ₹55,000 classroom in Kukatpally; engineers with 2 to 7 years take the 10-week Advanced Program. Both end with what employers at Palantir, OpenAI, Anthropic and their Indian counterparts ask about: a deployed system, an eval report and a demo you ran yourself.
Frequently asked questions
What do forward deployed engineers build most often?
Permission-aware RAG systems and agents with a human approval step are the two most common builds, followed by MCP servers or connectors into enterprise systems such as ServiceNow, Salesforce and SAP. Around every build sit eval suites in CI, SSO integration, deployment inside the customer’s VPC, a security evidence pack and the scoping and demo documents that keep the project funded.
Do forward deployed engineers train or fine-tune models?
Rarely. Most FDE work uses hosted models through the Claude or OpenAI APIs and focuses on retrieval, tool integration, evals and deployment. Fine-tuning appears in a minority of engagements where a customer has large labelled datasets and strict latency or cost targets. Across 1,000 FDE postings, LLM application skills and AI agents are far more common requirements than model training.
What is an MCP server and why do FDEs build them?
An MCP (Model Context Protocol) server exposes a system’s actions and data as typed tools that any MCP-capable model client can use. FDEs build them so agents can read and write to ServiceNow, Salesforce, SAP or internal databases through one standard interface, with per-user authentication, tool-level permissions and an audit log. Anthropic’s FDE job description names MCP servers as a core deliverable.
What is an eval suite in an FDE project?
A golden dataset of inputs and expected outputs, scorers such as exact match, RAGAS metrics or an LLM-as-judge rubric, and thresholds that define pass or fail. Run in CI with promptfoo or similar tools, it blocks deployments that regress and gives the customer a number to trust. OpenAI’s FDE engagements validate with evals before delivery for exactly this reason.
Want to become a Forward Deployed Engineer in Hyderabad?
See the Deccan Finance capstone in a free demo class: the RAG pipeline, the eval gate, the MCP servers and the evidence pack, built by learners like you. Classroom near JNTU Metro or live online.