We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible.
As a Lead Software Engineer at JPMorganChase within the Commercial and Investment Bank, Payments Technology, you are an integral part of an agile team that works to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. As a core technical contributor in London, you own major components of our B2B agentic commerce agents end to end, from the multi-agent system that negotiates and onboards corporate suppliers to the production path that trains, serves, and monitors the ML models those agents use as tools.
Job responsibilities
- Executes creative software solutions, design, development, and technical troubleshooting with ability to think beyond routine or conventional approaches to build solutions or break down technical problems
- Owns the design and delivery of one or more production agents (e.g., orchestration, negotiation, supplier onboarding, or outreach) running on NEO, including their tools, memory, guardrails, and evaluations
- Builds the production path for optimization and prediction models: training on Databricks, automated testing in each environment, promotion from development through UAT to production, and serving on Kubernetes as APIs and MCP tools agents can call
- Designs agent workflows that keep pricing, eligibility, and policy decisions in deterministic services, with the LLM limited to conversation, extraction, and coordination
- Develops secure and high-quality production code, and reviews and debugs code written by others
- Establishes evaluation and observability for agents and models: regression suites in CI/CD, LLM-as-judge scoring, OpenTelemetry traces, and model drift and performance monitoring
- Prepares agents and models for the firm's model risk review, producing the documentation, test evidence, and controls required to go live
- Drives team adoption of enterprise-authorized AI-assisted engineering practices to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review and refactoring, test strategy acceleration, incident and root-cause analysis support), while establishing consistent validation standards and promoting reuse of effective patterns across the team
- Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation
- Identifies opportunities to eliminate or automate remediation of recurring issues to improve overall operational stability of software applications and systems
- Leads evaluation sessions with external vendors, startups, and internal teams to drive outcomes-oriented probing of architectural designs, technical credentials, and applicability for use within existing systems and information architecture
Required qualifications, capabilities, and skills
- Formal training or certification on software engineering concepts and advanced applied experience
- Hands-on practical experience delivering system design, application development, testing, and operational stability
- Advanced, strong hands-on development experience in Python and/or Java, and proficient in one or more additional languages (e.g., TypeScript, SQL)
- Strong understanding of concurrency and distributed computation: multithreading and the async model, process and worker architectures, task queues and backpressure, and the failure modes of concurrent distributed systems (races, deadlocks, retries, idempotency)
- Hands-on experience shipping LLM-based applications or agents to production, including tool calling, retrieval, and evaluation
- Hands-on experience productionizing ML models: training pipelines, model registries, CI/CD for models, and serving behind APIs
- Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security
- Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs and outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices
- Proficient in all aspects of the Software Development Life Cycle
- Advanced understanding of agile methodologies such as CI/CD, Application Resiliency, and Security
- In-depth knowledge of the financial services industry and their IT systems
- Practical cloud native experience
- Proficiency with Kubernetes and Amazon EKS, micro-VM isolation (e.g., Firecracker, Kata Containers, gVisor), and sidecar patterns (service mesh proxies, policy and telemetry agents), with hands-on experience applying security at multiple layers of the stack: network and mTLS, workload identity, authorization, runtime isolation, and application
Preferred qualifications, capabilities, and skills
- Experience with agent frameworks and protocols: Google ADK, LangGraph, MCP, A2A, AG-UI
- Experience with Databricks, MLflow, Spark, and Delta Lake or Apache Iceberg
- Experience serving models on Kubernetes (e.g., EKS, KServe, Triton, vLLM), including GPU workloads and fine-tuned small language models
- Experience with optimization or pricing models, and with exposing them as services to downstream systems
- Experience with fine-grained authorization (OpenFGA, OPA/Rego) and OpenTelemetry
- Experience with AWS Step Functions, EventBridge, and Lambda for long-running, event-driven workflows
- Domain exposure to commercial card, accounts payable, supplier enablement, or KYC onboarding
