Manager, Software Engineering - AI Observability

FigmaSan Francisco, CA • New York, NY • United StatesEngineering

From another job board

NEL has not scanned this employer's domain, and NEL payment protection does not apply. You apply on the employer's own site.

Listed on Figma’s public greenhouse board and shown here for discovery. NEL is not involved in this hiring process.

About this role

Figma is growing our team of passionate creatives and builders on a mission to make design accessible to all. Figma’s platform helps teams bring ideas to life—whether you're brainstorming, creating a prototype, translating designs into code, or iterating with AI. From idea to product, Figma empowers teams to streamline workflows, move faster, and work together in real time from anywhere in the world. If you're excited to shape the future of design and collaboration, join us! As the Engineering Manager for Observability, you'll lead the team building the systems that show Figma how its platform and its AI products actually behave in production. That spans our core observability stack, meaning metrics, logs, and distributed tracing, as well as a fast-growing new frontier: the AI observability pipeline that turns traces, conversations, and model outputs into the classification and evals that make Figma's AI features measurably better. You'll set strategy across both, raise the bar on instrumentation and data quality, and push into AI-driven approaches to anomaly detection and operational automation. Observability was an unowned space at Figma until recently, so this is a rare chance to shape an early-stage, high-leverage platform that every engineering team here depends on. This is a full time role that can be held from one of our US hubs or remotely in the United States. What you’ll do at Figma: • Lead and grow a 6 engineer team responsible for the reliability, scalability, and evolution of Figma's observability and AI observability platforms. • Own and evolve the AI observability ecosystem: a privacy-safe telemetry pipeline that aggregates trace, conversation, and model data, runs classification, and lets teams run evals to measure and improve AI feature quality. • Own Figma's core observability stack, including platforms like Datadog, ensuring high availability, strong data quality, and a healthy signal-to-noise ratio across metrics, logs, and traces. • Set the technical strategy for the instrumentation standards, libraries, agents, and operators that monitor services across the company. • Explore and ship AI-driven approaches to anomaly detection, root cause analysis, signal correlation, and operational automation. • Partner across infrastructure, product engineering, finance, and security to give teams clear visibility into system health at scale, including cost efficiency as the team's scope grows. • Coach and develop engineers through career growth, feedback, and technical leadership, building a culture of ownership and high-quality execution. We'd love to hear from you if you have: • 4+ years of experience leading and growing engineering teams in infrastructure, observability, platform, or AI systems, with a track record of delivering reliable production systems at scale. • A strong foundation in distributed systems and a platform mindset: you build systems other teams adopt and rely on, not just for your own team. • Hands-on depth in observability, AI systems telemetry, or both, whether that means metrics, logs, and distributed tracing, or the traces, classification, and evals that power AI products. • Fluency with the modern tooling in your area, such as Datadog, OpenTelemetry, and Prometheus, or LLM tracing, prompt and model versioning, and eval workflows. • A high operational bar (instrumentation and data quality, SLO design, incident response) and the ability to set technical direction and drive cross-functional alignment in complex environments. While not required, it’s an added plus if you also have: • Breadth across both classic observability and AI observability, or the curiosity to grow into whichever is newer to you. • Experience creating AI Telemetry-assisted improvement feedback loops for services, in particular for AI products and features. • Experience improving infrastructure efficiency, including cost attribution, forecasting, vendor negotiation, and usage modeling. • A track record…

Sign in to open the employer’s application

This role is hosted on Figma’s own site, and NEL is not part of that hiring process. An account is free, takes a minute, and lets you keep the roles you are following in one place.

You can sign up as a professional looking for work or as a client hiring for one. Either opens this application.

Looking for escrow-protected work?

Roles posted directly on NEL are paid through escrow, with funds held and released on agreed milestones.