6,600 results
Technical Architect
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role Anthropic is hiring Technical Architects: hands-on builders and live technical communicators who help our customers go from "we bought Claude" to "our people use it every day." You'll be the technical partner our CSMs pull into the room whenever the work gets technical. For our customers, that's most of the time. One week you're running an advanced Claude Code hackathon for 1000 engineers at a global bank. The next you're on a call with their CISO walking through data handling, SSO, and governance controls. The week after, you're pair-building a plugin with a customer's platform team and debugging a SCIM sync live while thirty people watch. You are the helper people look for. This is a post-sale role. You'll work alongside CSMs, Applied AI, Support, FDEs, and more. We're looking for technologists who could have been researchers or staff engineers, but who get the most energy from helping customers get unstuck by teaching them something new. What you'll do - Be the technical lead in customer-facing moments across the post-sale journey: onboarding workshops, developer hackathons, live troubleshooting, executive security reviews, adoption working sessions. - Build and ship working software (Claude Code plugins, agents, sub-agents, MCP integrations, Cowork skills) both as customer deliverables and as reusable field assets. - Drive adoption of the Claude Code capabilities that make the product stick (sub-agents, skills, hooks, MCP, headless mode, managed settings) and turn new capabilities into field-ready demos and guides within days of release. - Own enterprise deployment and identity conversations: SSO (SAML/OIDC), SCIM provisioning, role, seat, and spend-limit management, sandbox and permission design. Sit credibly across from CISOs and security architects. - Advise and unblock customers running production workloads on the Anthropic API and on Bedrock / Vertex, in close partnership with Applied AI. - Teach. Design and deliver enablement that turns users into daily active developers, for audiences ranging from senior staff engineers to business users new to AI and develop the customer champions who carry adoption after you leave the room. - Represent Anthropic at customer all-hands and builder events, and bring structured field signal back to the Claude Code and Cowork product teams. - Partner tightly with CSMs as their technical counterpart, and with Applied AI, security specialists, Product Support, Engineering, and Sales as the connective technical tissue of the account. What we're looking for - You build. You've shipped real software, you've used Claude Code or comparable AI coding tools yourself, and you can stand behind your engineering choices. - You hold the room. You're at ease being the technical voice in front of senior developers, IT and security leaders, and non-technical stakeholders (often in the same hour), and you stay calm and useful when something breaks live. - You like the messy middle. You’re comfortable being pulled into an ambiguous customer situation and developing a solution in real time.</li&g
Engineering Manager, Research Productivity
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role: Anthropic’s Research Tools team builds systems that support our large-scale, distributed finetuning runs and improve the productivity of researchers. As a manager, you’ll support a team of machine learning and distributed systems experts to make these systems and tools highly efficient, support fast iteration on model development and research, and evolve the infrastructure continuously to incorporate new research advances. Our Research Tooling sits at the intersection of almost every technical group at Anthropic. You’ll work with research teams to incorporate their innovations into our production finetuning pipeline, product teams to help us iterate quickly on customer-oriented model improvements, and infrastructure teams to make sure our training runs and data pipelines are as efficient as possible. About Anthropic: Anthropic is an AI safety and research company working to build reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our customers and society as a whole. Our interdisciplinary team has experience across ML, physics, policy, business, and product. Responsibilities: - Prioritize the team’s work in collaboration with the technical lead, research teams, and product teams to support fast iteration on research projects and training runs. - Design processes (e.g., postmortem review, incident response, on-call rotations) that help the team operate effectively. - Coach and support your reports to understand and pursue their professional growth. - Run the team’s recruiting efforts efficiently, ensuring we can grow as quickly as we need through a period of rapid growth. You may be a good fit if you: - Believe that advanced AI systems could have a transformative effect on the world and are interested in helping make sure that transformation goes well. - Are an experienced manager (at least 2 years) and actively enjoy people management. - Are a quick study: this team sits at the intersection of a large number of different complex technical systems that you’ll need to understand (at a high level) to be effective. Strong candidates may also have: - Experience working with research teams, especially as part of a “research to production” pipeline - Strong people management experience: Coaching, performance evaluation, mentorship, career development - Strong project management skills: Prioritization, communicating across team/org boundaries - Experience recruiting for your team: Predicting staffing needs, designing interview loops, evaluating candidates, and closing them Deadline to apply: None. Applications will be reviewed on a rolling basis. <div cla
AI Deployment Engineer - Startups
About the team The AI Deployment Engineering team works closely with frontier startups. We are trusted advisors to, and thought partners with, startups to ensure that OpenAI’s technology is deployed safely and effectively, whilst also partnering with engineering, research, and product to turn those insights into evaluation systems, product improvements, and better model behavior. This team sits at the intersection of customer reality and model quality. We combine hands-on technical depth with strong product judgment, helping translate complex, high-value use cases into clear signals that can improve both the customer experience and the underlying systems. This role is based in Paris. We use a hybrid work model of 3 days in the office per week. We offer relocation assistance. We require fluency in French for this role. About the role We are seeking a technically proficient, product-minded engineer to help push the frontier of advanced AI with our strategic startup customers. You'll work with some of the most exciting AI startups in the world, helping them optimize their own systems and turning those learnings into durable improvements across OpenAI’s research and products. You will partner deeply on complex workflows, identify the gaps that matter, and help transform those gaps into reproducible evaluations, technical insights - helping shape OpenAI's research and product direction. This role is well suited to engineers who are equally comfortable debugging a workflow, iterating on prompts or agents, designing evaluations, and collaborating across research and product. You should be excited by ambiguous, high-impact problems and motivated by the opportunity to shape how advanced AI systems improve in practice. In this role, you will: - Work directly with strategic startup customers to understand critical workflows, uncover failure modes, and identify high-impact opportunities for improvement. - Prototype and iterate on prompts, agents, and workflow designs to better understand system behavior and unlock customer value. - Synthesize and deliver valuable feedback to the Product and Research teams, turning real usage patterns into clear, reproducible evals, benchmarks, and technical artifacts that improve model and product quality and ensure customer-grounded learnings influence roadmap and model development. - Build repeatable tools, patterns, and evaluation approaches that raise the quality bar across multiple use cases. - Operate with strong judgment in ambiguous environments, balancing immediate technical problem-solving with longer-term system improvement. - Build relationships within the startup ecosystem, serving as a technical partner to both individual customers and the broader community. You’ll thrive in this role if you: - Have strong software engineering & AI fundamentals. For example, experience as a startup CTO, software engineer, ML engineer, Data Scientist or equivalent. Experience shipping production systems end-to-end is a strong plus. - Have experience as a technical founder, or engineer at an early stage startup - Have familiarity with, or interest in, model training pipelines and reinforcement learning. - Have experience building AI applications, agents, or evaluation systems, and can reason clearly about model behavior in complex workflows. - Are comfortable working directly with highly technical users and translating their challenges into concrete technical signals. - Can move fluidly between prototyping, debugging, evaluation design, and cross-functional collaboration. - Communicate clearly across technical and non-technical audiences. - Bring high agency, strong product sense, and a bias toward building durable improvements rather than one-off fixes About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world throug
Applied AI Architect, Public Sector
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role As an Applied AI team member at Anthropic, you will be a pre-sales architect focused on becoming a trusted technical advisor to the UK and Northern Europe public sector, with a primary focus on UK Central Government departments, executive agencies, and arm's length bodies, and a reach extending across devolved administrations, local government, the NHS, and Northern European public sector markets. This includes a growing focus on defence and national security, working with the MOD and intelligence agencies on some of the UK's most sensitive and mission-critical challenges. You will help these organisations understand the value of Claude and paint the vision for how they can successfully integrate and deploy Claude into their technology estates to modernise operations, improve policy delivery, and transform citizen services. You’ll combine your deep technical expertise with customer-facing skills to architect innovative LLM solutions that address complex mission challenges while maintaining our high standards for safety and reliability. Working closely with our Sales, Product, Engineering, and Partnerships teams, you’ll guide customers from initial technical discovery through successful deployment. You’ll leverage your expertise to help customers understand Claude’s capabilities, develop evals, and design scalable, compliant architectures that maximise the value of our AI systems within the constraints that public sector organisations operate under. Responsibilities - Partner with account executives to deeply understand customer requirements and translate them into technical solutions, ensuring alignment between departmental outcomes, policy objectives, and technical implementation. - Serve as the primary technical advisor to public sector customers throughout their Claude adoption journey, from discovery to initial evaluation through deployment. You will need to coordinate internally across multiple teams and stakeholders to drive customer success. - Support customers building with Claude Code, the Claude API, and Claude for Enterprise. - Create and deliver compelling technical content tailored to different audiences. You will need to span the gamut from technical deep dives for engineering and delivery teams up to business-value conversations with senior civil servants and C-suite executives (Permanent Secretaries, Directors General, SROs, CDIOs). - Support defence and national security engagements, including with the MOD and intelligence agencies, designing solutions that work within the security, classification, and accreditation constraints of these environments. - Guide technical architecture decisions and help customers integrate Claude effectively into their existing technology stack, with alignment to NCSC guidelines, Cyber Essentials Plus, the Government Security Classifications framework, the Technology Code of Practice, and the Service Standard. - Help customers develop evaluation frameworks to measure Claude’s performance for their specific use cases. - Identify common integration patterns across the UK publi
Research Engineer, AI Safety & Alignment
ABOUT THE ROLE AND TEAM Joining us as a Research Engineer, you'll be at the forefront of tackling one of the most critical challenges in AI today: safety and alignment. Your work will be pivotal in understanding and mitigating the risks of advanced AI, conducting foundational research to make our models safer, and solving the core technical problems of AI alignment—ensuring our models behave in accordance with human values and intentions. The Safety team is dedicated to pioneering and implementing techniques that make our models more robust, honest, and harmless. As a Research Engineer, you will bridge the gap between theoretical research and practical application, writing high-quality code to test hypotheses and integrating successful safety solutions directly into our products. Your research will not only protect millions of users but also contribute to the broader scientific community's understanding of how to build safe, beneficial AI. WHAT YOU'LL DO - Develop and implement novel evaluation methodologies and metrics to assess the safety and alignment of large language models. - Research and develop cutting-edge techniques for model alignment, value learning, and interpretability. - Conduct adversarial testing to proactively uncover potential vulnerabilities and failure modes in our models. - Analyze and mitigate biases, toxicity, and other harmful behaviors in large language models through techniques like reinforcement learning from human feedback (RLHF) and fine-tuning. - Collaborate with engineering and product teams to translate safety research into practical, scalable solutions and best practices. - Stay abreast of the latest advancements in AI safety research and contribute to the academic community through publications and presentations. WHO YOU ARE - Hold a PhD (or equivalent experience) in a relevant field such as Computer Science, Machine Learning, or a related discipline. - Write clear and clean production-facing and training code - Experience working with GPUs (training, serving, debugging) - Experience with data pipelines and data infrastructure - Strong understanding of modern machine learning techniques, particularly transformers and reinforcement learning, with a focus on their safety implications. - Are passionate about the responsible development of AI and dedicated to solving complex safety challenges. NICE TO HAVE - Experience with product experimentation and A/B testing - Experience training large models in a distributed setting - Familiarity with ML deployment and orchestration (Kubernetes, Docker, cloud) - Experience with explainable AI (XAI) and interpretability techniques. - Have research in AI safety, alignment, ethics, or a related area. - Knowledge of the broader societal and ethical implications of AI, including policy and governance. - Publications in relevant academic journals or conferences in the field of machine learning ABOUT CHARACTER.AI Character.AI http://Character.AI empowers people to connect, learn and tell stories through interactive entertainment. Over 20 million people visit Character.AI http://Character.AI every month, using our technology to supercharge their creativity and imagination. Our platform lets users engage with tens of millions of characters, enjoy unlimited conversations, and embark on infinite adventures. In just two years, we achieved unicorn status and were honored as Google Play's AI App of the Year—a testament to our innovative technology and visionary approach. Join us and be a part of establishing this new entertainment paradigm while shaping the future of Consumer AI! At Character, we value diversity and welcome applicants from all backgrounds. As an equal opportunity employer, we firmly uphold a non-discrimination policy based on race, religion, national origin, gender, sexual orientation, age, veteran status, or disability. Your unique perspectives are vital to our success.
Solutions Engineer (Clearance Required)
Our customer base is growing exponentially, and you will be on the front lines of ensuring that the world's most innovative companies become passionate, lifelong, Scale customers. Our Solutions Engineers ensure customers' first experiences with Scale's technology are flawless and lead to a successful long-term partnership. The work will vary daily, and we’re looking for technical experts excited to solve tough problems. As a solution engineer, you will be a part of helping shape our early-stage federal business by re-envisioning our commercial product offerings for our federal clients. What you'll do: - Become an expert on the end-to-ends of Scale Products - Create tailored demonstrations and collateral for federal stakeholders at both the executive and analyst level. - Partner with Scale Account Executives to deliver customer pilots according to requirements agreed by the customer. - Integrate and ingest a variety of external datasets to solve government use cases. - Interact with customers on a day-to-day basis to understand their pain points and design solutions - Work with internal product and engineering teams to turn customer requirements into Scale capabilities - Understand public sector mission sets and strategic objectives to better showcase Scales products. Ideally you'd have: - Strong engineering background, preferably in computer science, mathematics, or other quantitative fields - Strong communication skills - ability to interact with both technical and non-technical customers at all levels - At ease with technology, able to quickly pick up new tech stacks and troubleshoot - Previous experience working with Public Sector customers - our business is diverse and growing across both National Security and Federal Civilian communities. - Proficiency in scripting languages such as Python, Javascript/Typescript, Bash scripts, or programming languages. - A strong desire to roll up your sleeves and help build a business in an extremely fast-paced environment - Active US Government Security Clearance (TS / SCI required) - Based in the Washington, DC area or willing to relocate - Background working in AI/ML, particularly Generative AI and Large Language Models Compensation packages at Scale for eligible roles include base salary, equity, and benefits. The range displayed on each job posting reflects the minimum and maximum target for new hire salaries for the position and may be inclusive of several career levels at Scale; it will be determined during the interview process based on work location and additional factors, including job-related skills, experience, qualifications, interview performance, and relevant education or training. Scale employees in eligible roles are also granted equity based compensation, subject to Board of Director approval. Your recruiter can share more about the specific salary range for your preferred location during the hiring process, and confirm whether the hired role will be eligible for equity grant. You'll also receive benefits including, but not limited to: comprehensive health, dental and vision coverage, retirement benefits, a learning and development stipend, and generous PTO. Additionally, this role may be eligible for additional benefits such as a commuter stipend. </
Agent Post-Training, Context Research
ABOUT THE TEAM The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. ABOUT THE ROLE We believe that the final enabler for AGI is spending compute on context. As a Context Researcher on Agent Post-Training, you will scale compute spent on context. You will get to work in our frontier training stack on enabling the next paradigm of model training with a clear product interface for iterative deployment (Codex Chronicle). You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. IN THIS ROLE, YOU MIGHT - Design and run experiments that improve scaling of compute on context. - Own end-to-end improvements to the post-training stack, including RL, data pipelines, graders, reward signals, evals, diagnostics, and model-behavior analysis. - Build evals and environments that expose the next set of model failures, then turn those failures into training data, product fixes, or new research directions. - Partner with Codex and ChatGPT product teams to understand what users need and translate product signal into model improvements. - Work on early-training and alignment interventions, including data mixtures, objectives, synthetic data, and eval loops that shape downstream agent behavior. - Help decide which integrations, capabilities, and fixes are ready for inclusion in major model runs. - Improve the machinery for large-scale training and launch: experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness. - Take on cross-functional projects that touch model training, product infrastructure, and the production agent harness, such as multi-agent systems or training directly against production-like environments. - Debug hard failures in shipped or near-shipped models and turn messy qualitative behavior into concrete hypotheses, experiments, and fixes. YOU MIGHT THRIVE IN THIS ROLE IF YOU - Have strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, and can learn quickly across the parts you have not worked in before. - Have hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems. - Are excited by open-ended problems where the path is unclear, the signal is noisy, and the right answer requires both research taste and engineering execution. - Care about product impact and model behavior, not just benchmark movement. You have opinions about what makes an agent useful, reliable, honest, tasteful, and easy to work with. - Can move from a vague behavioral problem to a concrete experiment: define the hypothesis, build the pipeline, run the model, analyze the result, and decide what to do next. - Are comfortable working across research, product, infrastructure, data
People Research Scientist, People
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the Role: We are seeking a People Research Scientist to join our People Data Solutions team. You’ll be the research expert supporting our broader People organization, using rigorous scientific methods to advance our understanding of the employee experience, manager effectiveness, organizational health, and workforce dynamics. This role sits at the intersection of organizational science, behavioral research, and people strategy – developing novel frameworks and conducting systematic research that drives evidence-based people decisions across our growing organization. This role offers the opportunity to make a significant impact on both our people practices and the broader field of people science at a leading AI safety company. Responsibilities: Research Design & Scientific Inquiry - Design and execute systematic research studies to answer fundamental questions about employee experience, manager effectiveness, and organizational health - Generate and test hypotheses about people programs, employee behavior, and workforce outcomes using rigorous experimental and quasi-experimental methods - Conduct longitudinal studies tracking employee cohorts to understand long-term workforce trends and the impact of people initiatives - Perform meta-analyses of people interventions across the industry to identify best practices and knowledge gaps - Navigate research ethics considerations when studying employee data, ensuring responsible research practices Employee listening & survey research - Design, analyze, and iterate on employee listening programs including engagement surveys, pulse surveys, and lifecycle surveys - Apply psychometric methods to validate survey instruments and ensure measurement reliability - Translate survey findings into strategic recommendations that drive meaningful organizational change Manager research & organizational effectiveness - Conduct research on manager behaviors, competencies, and their impact on team outcomes - Build measurement frameworks to evaluate and improve manager effectiveness programs - Study organizational dynamics including team composition, collaboration patterns, and their relationship to performance outcomes Visualization & communication - Build compelling visualizations and dashboards that make complex research findings accessible to diverse audiences - Present research findings to senior leadership with clear, actionable recommendations - Develop self-service analytics capabilities that empower People team partners Minimu
AI Deployment Engineer, Startups
About the team The AI Deployment Engineering team works closely with frontier startups. We are trusted advisors to, and thought partners with, startups to ensure that OpenAI’s technology is deployed safely and effectively, whilst also partnering with engineering, research, and product to turn those insights into evaluation systems, product improvements, and better model behavior. This team sits at the intersection of customer reality and model quality. We combine hands-on technical depth with strong product judgment, helping translate complex, high-value use cases into clear signals that can improve both the customer experience and the underlying systems. About the role We are seeking a technically proficient, product-minded engineer to help push the frontier of advanced AI with our strategic startup customers. You'll work with some of the most exciting AI startups in the world, helping them optimize their own systems and turning those learnings into durable improvements across OpenAI’s research and products. You will partner deeply on complex workflows, identify the gaps that matter, and help transform those gaps into reproducible evaluations, technical insights - helping shape OpenAI's research and product direction. This role is well suited to engineers who are equally comfortable debugging a workflow, iterating on prompts or agents, designing evaluations, and collaborating across research and product. You should be excited by ambiguous, high-impact problems and motivated by the opportunity to shape how advanced AI systems improve in practice. This role is based in Stockholm. In this role, you will: - Work directly with strategic startup customers to understand critical workflows, uncover failure modes, and identify high-impact opportunities for improvement. - Prototype and iterate on prompts, agents, and workflow designs to better understand system behavior and unlock customer value. - Synthesize and deliver valuable feedback to the Product and Research teams, turning real usage patterns into clear, reproducible evals, benchmarks, and technical artifacts that improve model and product quality and ensure customer-grounded learnings influence roadmap and model development. - Build repeatable tools, patterns, and evaluation approaches that raise the quality bar across multiple use cases. - Operate with strong judgment in ambiguous environments, balancing immediate technical problem-solving with longer-term system improvement. - Build relationships within the startup ecosystem, serving as a technical partner to both individual customers and the broader community. You’ll thrive in this role if you: - Have strong software engineering & AI fundamentals. For example, experience as a startup CTO, software engineer, ML engineer, Data Scientist or equivalent. Experience shipping production systems end-to-end is a strong plus. - Have experience as a technical founder, or engineer at an early stage startup - Have familiarity with, or interest in, model training pipelines and reinforcement learning. - Have experience building AI applications, agents, or evaluation systems, and can reason clearly about model behavior in complex workflows. - Are comfortable working directly with highly technical users and translating their challenges into concrete technical signals. - Can move fluidly between prototyping, debugging, evaluation design, and cross-functional collaboration. - Communicate clearly across technical and non-technical audiences. - Bring high agency, strong product sense, and a bias toward building durable improvements rather than one-off fixes. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mi
Research Engineer, Cybersecurity RL (Reinforcement...
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About Horizons The Horizons team leads Anthropic's reinforcement learning (RL) research and development, playing a critical role in advancing our AI systems. We've contributed to every Claude release, with significant impact on the autonomy, coding, and reasoning capabilities of Anthropic's models. About the role We're hiring for the Cybersecurity RL team within Horizons. As a Research Engineer, you'll help to safely advance the capabilities of our models in secure coding, vulnerability remediation, and other areas of defensive cybersecurity. This role blends research and engineering, requiring you to both develop novel approaches and realize them in code. Your work will include designing and implementing RL environments, conducting experiments and evaluations, delivering your work into production training runs, and collaborating with other researchers, engineers, and cybersecurity specialists across and outside Anthropic. The role requires domain expertise in cybersecurity paired with interest or experience in training safe AI models. For example, you might be a white hat hacker who's curious about how LLMs could augment or transform your work, a security engineer interested in how AI could help harden systems at scale, or a detection and response professional wondering how models could enhance defensive workflows. You may be a good fit if you: - Have experience in cybersecurity research. - Have experience with machine learning. - Have strong software engineering skills. - Can balance research exploration with engineering implementation. - Are passionate about AI's potential and committed to developing safe and beneficial systems. Strong candidates may also have: - Professional experience in security engineering, fuzzing, detection and response, or other applied defensive work. - Experience participating in or building CTF competitions and cyber ranges. - Academic research experience in cybersecurity. - Familiarity with RL techniques and environments. - Familiarity with LLM training methodologies. The annual compensation range for this role is listed below. For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. Annual Salary: $300,000 - $405,000 USD <strong>
Research Engineer, Universes
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the Team The Universes team within Research is responsible for training AI models to perform complex, difficult, long-horizon agentic tasks in ultra-realistic settings. We design and implement novel training environments that go far beyond what models can do today — environments where models learn to navigate ambiguity, handle interruptions, maintain context over extended interactions, and exercise judgment in open-ended scenarios. About the Role We're looking for Research Engineers to help us build the next generation of training environments for capable and safe agentic AI. This role blends research and engineering responsibilities, requiring you to both implement novel approaches and contribute to research direction. You'll work on fundamental research in reinforcement learning, designing training environments and methodologies that push the state of the art, and building evaluations that measure genuine capability. Responsibilities: - Build the next generation of agentic environments - Build rigorous evaluations that measure real capability - Collaborate across research and infrastructure teams to ship environments into production training - Debug and iterate rapidly across research and production ML stacks - Contribute to research culture through technical discussions and collaborative problem-solving You may be a good fit if you: - Are highly impact-driven — you care about outcomes, not activity - Operate with high agency - Have good research taste or senior technical experience, demonstrating good judgment in identifying what actually matters in complex problem spaces - Can balance research exploration with engineering implementation - Are passionate about the potential impact of AI and are committed to developing safe and beneficial systems - Are comfortable with uncertainty and adapt quickly as the landscape shifts - Have strong software engineering skills and can build robust infrastructure - Enjoy pair programming (we love to pair!) Strong candidates may also have one or more of the following: - Have industry experience with large language model training, fine-tuning or evaluation - Have industry experience building RL environments, simulation systems, or large-scale ML infrastructure - Senior experience in a relevant technical field even if transitioning domains - Deep expertise in sandboxing, containerization, VM infrastructure, or distributed systems - &
Agent Post-Training, Frontier Evals and Environmen...
ABOUT THE TEAM The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. ABOUT THE ROLE As a researcher working on Frontier Evals & Environments, you will help build north star model environments to drive progress towards safe AGI/ASI. Your work will directly guide the research programs of the most ambitious training runs happening at OpenAI. Some prior open-sourced evaluations built by researchers in this role include GDPval https://openai.com/index/gdpval/, SWE-bench Verified https://openai.com/index/introducing-swe-bench-verified/, MLE-bench https://openai.com/index/mle-bench/, PaperBench https://openai.com/index/paperbench/, and SWE-Lancer https://openai.com/index/swe-lancer/. If you are interested in feeling firsthand the fast progress of our models, and steering them towards good outcomes, this is the role for you. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. IN THIS ROLE, YOU MIGHT - Create ambitious RL environments to push our models to their limits, and measure frontier model capabilities, skills, and behaviors - Develop new methodologies for automatically exploring the behavior of these models - Dive deep into the science of measurement, including understanding scalability, reliability, and variance of our evaluation methodology - Help steer training for our largest training runs, and see the future first - Design scalable systems and processes to support continuous evaluation - Build self-improvement loops to automate model understanding YOU MIGHT THRIVE IN THIS ROLE IF YOU - Have strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, and can learn quickly across the parts you have not worked in before. - Have hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems. - Are excited by open-ended problems where the path is unclear, the signal is noisy, and the right answer requires both research taste and engineering execution. - Care about product impact and model behavior, not just benchmark movement. You have opinions about what makes an agent useful, reliable, honest, tasteful, and easy to work with. - Can move from a vague behavioral problem to a concrete experiment: define the hypothesis, build the pipeline, run the model, analyze the result, and decide what to do next. - Are comfortable working across research, product, infrastructure, data, evals, and safety boundaries, and can communicate clearly with each group. - Like building load-bearing systems and processes when that is what the team needs, even if the work is not glamorous. - Want to train and ship the models that make agents genuinely useful for developers, enterprises, researchers, and everyday users. Abo
Research Engineer, Domain Scaling
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role The Domain Scaling team has the goal to make Claude world-class at real-world knowledge work in domains like finance, healthcare, and legal. This is a unique role that combines executing directly on applied research and data sourcing (real-world and synthetic) to improve our models. You'll own the end-to-end process of creating RL environments for new capabilities: identifying high-value tasks, designing reward signals, managing vendor relationships, and measuring impact on model performance. Responsibilities - Own the data strategy for knowledge work verticals end-to-end, from task sourcing through RL training - Manage technical relationships with external data vendors, including evaluation of data quality and reward design - Collaborate with domain experts to design data pipelines and evaluations - Explore novel ways of creating RL envs for high value tasks - Develop and improve QA frameworks to catch reward hacking and ensure env quality - Run generalization experiments to measure how data strategy changes improve model capabilities - Partner with other RL research teams and product teams to translate capability goals into training envs and evals You may be a good fit if you - Have experience with fine-tuning large language models for specific domains or real-world use cases - Have experience with reinforcement learning, reward design, or training data curation for LLMs - Are comfortable managing technical vendor relationships and iterating quickly on feedback - Find value in reading through datasets to understand them and spot issues - Have strong cross-functional collaboration skills - Are passionate about making AI more useful and accessible across different industries - Are excited about a role that includes a combination of applied research and hands-on data work Strong candidates may also - Have experience training production ML systems - Have experience designing evals or benchmarks for LLMs - Have domain expertise in a vertical where we would like to make our models more useful - Have experience working with external vendors or technical partners The annual compensation range for this role is listed below. For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. Annual Salary: <div class
Research Scientist, Interpretability
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role: When you see what modern language models are capable of, do you wonder, "How do these things work? How can we trust them?" The Interpretability team at Anthropic is working to reverse-engineer how trained models work because we believe that a mechanistic understanding is the most robust way to make advanced systems safe. We’re looking for researchers and engineers to join our efforts. People mean many different things by "interpretability". We're focused on mechanistic interpretability, which aims to discover how neural network parameters map to meaningful algorithms. Some useful analogies might be to think of us as trying to do "biology" or "neuroscience" of neural networks using “microscopes” we build, or as treating neural networks as binary computer programs we're trying to "reverse engineer". A few places to learn more about our work and team at a high level are this introduction to Interpretability from our research lead, Chris Olah ; a discussion of our work on the Hard Fork podcast produced by the New York Times, and this blog post (and accompanying video) sharing more about some of the engineering challenges we’d had to solve to get these results. Some of our team's notable publications include A Mathematical Framework for Transformer Circuits , In-context Learning and Induction Heads , Toy Models of Superposition , Scaling Monosemanticity , and our Circuits’ Methods and Biology papers. This work builds on ideas from members' work prior to Anthropic such as the original circuits thread , Multimodal Neurons , <a class="text-accent-seco
Biological Safety Research Scientist
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the Role We are looking for biological scientists to help build safety and oversight mechanisms for our AI systems. As a Safeguards Biological Safety Research Scientist, you will apply your technical skills to design and develop our safety systems which detect harmful behaviors and to prevent misuse by sophisticated threat actors. You will be at the forefront of defining what responsible AI safety looks like in the biological domain, working across research, policy, and engineering to translate complex biosecurity concepts into concrete technical safeguards. This is a unique opportunity to shape how frontier AI models handle dual-use biological knowledge—balancing the tremendous potential of AI to accelerate legitimate life sciences research while preventing misuse by sophisticated threat actors. In this role, you will: - Design and execute capability evaluations ("evals") to assess the capabilities of new models - Collaborate closely with internal and external threat modeling experts to develop training data for our safety systems, and with ML engineers to train these safety systems, optimizing for both robustness against adversarial attacks and low false-positive rates for legitimate researchers - Analyze safety system performance in traffic, identifying gaps and proposing improvements - Develop rigorous stress-testing of our safeguards against evolving threats and product surfaces - Partner with Research, Product, and Policy teams to ensure biological safety is embedded throughout the model development lifecycle - Contribute to external communications, including model cards, blog posts, and policy documents related to biological safety - Monitor emerging technologies for their potential to contribute to new risks and new mitigation strategies, and strategically address these Minimum Qualifications: - A PhD in molecular biology, virology, microbiology, biochemistry, systems or computational biology, or a related life sciences field, OR equivalent professional experience - Extensive experience in scientific computing and data analysis, with proficiency in programming (Python preferred) - Deep expertise in modern biology, including both "reading" (e.g. high-throughput measurement, functional assays) and "writing" (gene synthesis, genome editing, strain construction, protein engineering) techniques in biology - Familiarity with dual-use research concerns, select agent regulations, and biosecurity frameworks (e.g., Biological Weapons Convention, Australia Group guidelines) - Strong analytical and writing skills, with the ability to navigate ambiguity and explain complex technical concepts to non-technical stakeholders - Have a passion for learning new skills and an ability to rapidly adapt to changing techniques and technologies - Comfort working in a fast-paced environment where priorities may shift as AI capabilities evolve Preferred Qualifications - Background in AI/ML systems, particularly experience with large language models - Experience in
Research Operations, Discovery
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the Team Our team is organized around the north star goal of building an AI scientist—a system capable of solving the long-term reasoning challenges and basic capabilities necessary to push the scientific frontier. About the Role We're seeking a Science Research Operations team member to build and own the operational infrastructure that keeps our research organization running at full speed. Our science teams are working on some of the hardest and most consequential problems in AI—training large-scale models, running complex experiments, and building novel products at the frontier. What makes that possible isn't just talent: it's the coordination, systems, and programs that let researchers spend their time on the science rather than the overhead around it. This role sits at the intersection of research operations, technical program management, and product strategy. You'll work directly with research scientists and research engineers, doing a mix of tasks including running research partnerships, managing complex internal programs, and helping run the team’s day-to-day operations. You'll also contribute to science product development—helping translate research directions into product strategy and ensuring our production deployment environments reflect our best configurations. This is not a pure coordination role. The best candidates will engage substantively with what the team is building, have a role in determining our strategy, spot problems before they surface, and bring genuine ownership to the systems and programs they run. Responsibilities: - Build and manage custom expert contractor networks, sourcing domain specialists for eval and training data work that requires expertise beyond standard channels - Run research partnerships with external partners, from scoping through delivery - Provide end-to-end TPM support for major research pushes—coordinating across teams, tracking dependencies, and keeping stakeholders aligned - Ensure that our research progress is complemented by products that enable scientists to make maximal use of model capabilities. - Support recruiting efforts. - Coordinate external communications for the team, including supporting blog posts and preparing public-facing materials - Partner with product teams to contribute to science product strategy, product design, and novel product integrations where research and product intersect - Own team logistics including onboarding coordination, team events, and operational programs that improve team efficiency You may be a good fit if you: - Have experience in research operations, technical program management, or a related role in a fast-moving technical environment - Can context-switch fluidly between operational work (logistics, tracking, coordination) and higher-order work (strategy, partnerships, product thinking) - Have a technical background, with experience in software development, machine learning, or biology R&D. - Are comfortable working directly with research scientis
Research Engineer, Code RL (Reinforcement Learning...
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the RL Teams Our Reinforcement Learning teams play a critical role in advancing our AI systems. We've contributed to all Claude models, with significant impacts on the autonomy and coding capabilities of our latest Claude models. Our work spans several key areas: - Developing systems that enable models to use computers effectively - Advancing code generation through reinforcement learning - Pioneering fundamental RL research for large language models - Building scalable RL infrastructure and training methodologies - Enhancing model reasoning capabilities We collaborate closely with Anthropic's alignment and frontier red teams to ensure our systems are both capable and safe. We partner with the applied production training team to bring research innovations into deployed models, and are dedicated to implement our research at scale. Our Reinforcement Learning teams sit at the intersection of cutting-edge research and engineering excellence, with a deep commitment to building high-quality, scalable systems that push the boundaries of what AI can accomplish. About the Role We're hiring for the Code RL team within the RL organization. As a Research Engineer, you'll advance our models' ability to write, edit, test, debug, and ship real software — end to end, on real codebases, with real tools — and to do it correctly, fast, and safely. This role blends research and engineering. You'll design RL environments and coding tasks, build the reward signals and verifiers that capture what "good code" means, run training experiments on frontier models, diagnose why a model does (or doesn't) get better at a class of software-engineering work, and improve the speed and reliability of the pipelines that make all of that iterate fast. Code RL spans several focus areas — from agentic coding behaviors and code correctness, to long-horizon autonomous engineering, to high-performance code for accelerators — and we'll match you to the area where you'll have the most impact. You may be a good fit if you: - Have strong software-engineering skills and deep Python expertise, including async/concurrent programming - Are comfortable owning systems end to end and debugging across the stack - Can balance research exploration with engineering implementation, and engage rigorously in shaping experimental design and interpreting results - Care about code quality, testing, and performance - Are passionate about the potential impact of AI and are committed to developing safe and beneficial systems Strong candidates may also have: - Experience with reinforcement learning, RLHF, post-training, or LLM finetuning - Built coding agents, code-execution sandboxes, eval harnesses, veri
Data Scientist, Supply
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About Anthropic Anthropic is an AI safety and research company. We build reliable, interpretable, and steerable AI systems, and we believe AI will have a vast impact on the world — our goal is to ensure that impact is positive. About the role Anthropic is compute-constrained, and how we allocate that compute is one of the highest-leverage decisions we make as a company. Today, allocation choices are only loosely tied to the user outcomes we ultimately care about — retention, lifetime value, and the experience of people relying on Claude. This role exists to change that by addressing two intertwined problems at the heart of how we allocate compute. The first is an allocation problem: matching a volatile, heterogeneous stream of demand to a finite, heterogeneous fleet of chips. Which models run on which hardware, in which regions, under what serving configurations — with demand shifting and capacity bounded — is a problem the team navigates continuously today, with more intuition than rigor. You will bring structure to it: building the metrics and analytical frameworks that make the trade-offs legible, and partnering with the infrastructure teams that own these systems to turn that understanding into better decisions. The second is a causal-inference problem: there are many levers — rate limits, pricing, cache behavior, capacity shifts, routing changes — and only a partial picture of what pulling each one actually does to the users on the other end. You will build the causal understanding that closes that gap, choosing whatever approach the question calls for, so allocation decisions are made on expected user impact rather than intuition. This role is a fit for someone who thinks natively in terms of constrained allocation and queueing, who treats "what would happen if we changed X" as an identification problem rather than a dashboard query, and who wants their work to translate into operational and productionized change. You will work closely with the infrastructure engineers who run our compute, and your findings will be presented to senior leadership. Key responsibilities - Build and run testing frameworks — observational and synthetic — to quantify how different inputs affect compute allocation outcomes - Connect compute allocation decisions to downstream user outcomes (retention, lifetime value, revenue) - Partner closely with infrastructure engineers, product, and research to instrument systems, measure what matters, and ship operational changes - Develop the metric hierarchies, dashboards, and reporting that turn supply decisions into shared understanding across the company - Contribute analyses and recommendations to executive forums, and co-author the supply narrative shared with the CTO and staff Minimum qualifications - Strong technical individual-contributor background in data science, analytics, or operations research - Demonstrated comfort reasoning about resource allocation and trade-offs under constraints — drawn to systems problems, not just dashboards - Working fluency with causal inference — able to recogni
Software Engineer, Enterprise
At Scale AI, we’re not just building AI tools—we’re pioneering the next era of enterprise AI. As businesses race to harness the power of Generative AI, Scale is at the forefront, delivering cutting-edge solutions that transform workflows, automate complex processes, and drive unparalleled efficiency for the largest enterprises. Our Scale Generative AI Platform (SGP) provides foundational services and APIs, enabling businesses to seamlessly integrate AI into their operations at production scale. We’re looking for a Backend Engineer to help bring large-scale GenAI systems to production. In this role, you’ll build the core infrastructure that powers AI products for some of the world’s largest enterprises—designing scalable APIs, distributed data systems, and robust deployment pipelines that enable production-grade reliability and performance. This is a rare opportunity to be at the center of the GenAI revolution, solving hard backend and infrastructure challenges that make AI truly work at enterprise scale. If you're excited about shaping how AI systems are deployed and scaled in the real world, we want to hear from you. At Scale, we don’t just follow AI advancements — we lead them. Backed by deep expertise in data, infrastructure, and model deployment, we are uniquely positioned to solve the hardest problems in AI adoption. Join us in shaping the future of enterprise AI, where your work will directly impact how businesses operate, innovate, and grow in the age of GenAI. You Will: - Design, build, and scale backend systems that power enterprise GenAI products, focusing on reliability, performance, and deployment across both Scale’s and customers’ infrastructure. - Develop core services and APIs that integrate AI models and enterprise data sources securely and efficiently, enabling production-scale AI adoption. - Architect scalable distributed systems for data processing, inference, and orchestration of large-scale GenAI workloads. - Optimize backend performance for latency, throughput, and cost—ensuring AI applications can operate at enterprise scale across hybrid and multi-cloud environments. - Manage and evolve cloud infrastructure (AWS, Azure, or GCP), driving automation, observability, and security for large-scale AI deployments. - Collaborate with ML and product teams to bring cutting-edge GenAI models into production through efficient APIs, model serving systems, and evaluation frameworks. - Continuously improve reliability and scalability , applying strong engineering practices to make AI systems robust, maintainable, and enterprise-ready. Ideally, You Have: - 4+ years of experience developing large-scale backend or infrastructure systems, with a strong emphasis on distributed services, reliability, and scalability. - Proficiency in Python or TypeScript , with experience designing high-performance APIs and backend architectures using frameworks such as FastAPI, Flask, Express, or NestJS. - Deep familiarity with cloud infrastructure (AWS and Azure preferred), including container orchestration (Kubernetes, Docker) and Infrastructure-as-Code tools like Terraform. - Experience managing data systems such as relational and NoSQL databases (PostgreSQL, Dyna
Technical Program Manager, Hardware Chips Developm...
About the Team We believe that increasing compute is a huge lever to AI progress. The Hardware team owns the design and/or sourcing of the compute, storage and interconnect hardware needed to build OpenAI’s supercomputers at the scale needed to deliver AGI that is beneficial to humanity. This includes: - Optimizing the processing hardware for AI models - Designing memory systems to keep up with the needs of training and inference - Enabling the scale-up and scale-out interconnect fabrics create the world’s most power supercomputers We work at the very cutting edge of speed and scale. You won’t encounter another organization with as much compute per employee. We are a small team that moves quickly, with access to huge resources, working with a direct impact on the success of OpenAI and, by extension, the field of AI as a whole. About the Role As a Hardware Chips Programs Manager at OpenAI, you will help bring our chips hardware roadmap to life, navigating an array of technical and partnership challenges. We’re looking for people excited to push the frontiers of computing by navigating technical explorations and are passionate about building. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. "To comply with U.S. export control laws and regulations, candidates for this role may need to meet certain legal status requirements as provided in those laws and regulations." In this role, you will: - Manage the design and implementation planning of our ML acceleration hardware, working across technical, cross-functional and external stakeholders - Lead planning and scheduling of chip hardware designs with our strategic partners and vendors - Coordinate and marshal internal resources and communication for efficient interaction with partners and vendors. You might thrive in this role if you: - Have experience as a technical program manager for data center hardware products (server, GPU, TPU, networking, storage and so on) - Know the whole end-to-end system program management from concept, design, production, deployment into the data center - Have some experience with System SW programs through NPI - Want to help design some of the world’s largest supercomputing systems, working at the edge of complex hardware challenges - Enjoy working with and enabling world-class AI Researchers and Engineers - Are passionate about the technical program function, and enjoy independently owning and delivering on your teams’ goals and cutting-edge problems in AI compute About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement https://cdn.openai.com/policies/eeo-policy-statement.pdf. Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that crim
Research Economist, Economic Research
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the Role As an Economist at Anthropic, you will work to measure and understand AI's effects on the global economy. You will make fundamental contributions to the development of the Anthropic Economic Index, establishing new methodologies to measure the usage, diffusion, and impact of AI throughout the economy using privacy-preserving tools and novel data sources. You will use frontier methods in econometrics, machine learning, and structural estimation. Such rigour will drive impact, shaping both policy discussions externally and informing Anthropic’s internal business and product decisions. Our team combines rigorous empirical methods with novel measurement approaches. We're building first-of-its-kind datasets tracking AI's impact on labor markets, productivity, and economic transformation. Using our privacy-preserving measurement system ( Clio ), we analyze millions of real-world AI interactions to understand how AI augments and automates work across different occupations and tasks. Responsibilities - Make fundamental contributions to the development and expansion of the Anthropic Economic Index , including quarterly reports and industry-specific deep dives - Design and conduct empirical research on AI's economic effects, drawing on external data sources and the privacy-preserving measurement systems internally - Develop new methodological approaches for studying AI's impact on: - Labor markets and the future of work - Productivity and task transformation - Economic inequality and displacement - Industry-specific disruption and adaptation - Aggregate economic trajectories (GDP, productivity, unemployment) under varying AI-adoption scenarios - Develop causal-inference tooling — e.g. surrogate indexes, heterogeneous-effect pipelines — to help Anthropic evaluate the downstream economic consequences of its own compute, product, and pricing decisions - Build and maintain relationships with academic institutions, policy think tanks, and other research partners - Work cross-functionally with other technical teams to improve our measurement infrastructure and data collection - Translate research insights into actionable recommendations for both product decisions and policy discussions - Amplify external engagement through research publications, policy briefs, and presentations to diverse stakeholders You May Be a Good Fit If You Have - PhD in Economics - Strong track record of empirical research, particularly studies combining novel data sources and economic theory or those implementing frontier methods in causal inference and machine learning - Experience relevant to the study of AI’s impact on the economy, including: - Labor market analysis and occupational change - Task-based appr
Hardware / Software CoDesign Engineer - 3P
About the Team OpenAI’s Hardware organization develops silicon and system-level solutions designed for the unique demands of advanced AI workloads. The team is responsible for building the next generation of AI-native silicon while working closely with software and research partners to co-design hardware tightly integrated with AI models. In addition to delivering production-grade silicon for OpenAI’s supercomputing infrastructure, the team also creates custom design tools and methodologies that accelerate innovation and enable hardware optimized specifically for AI. About the Role As an Engineer on our hardware optimization and co-design team, you will co-design future hardware from different vendors for programmability and performance. You will work with our kernel, compiler and machine learning engineers to understand their unique needs related to ML techniques, algorithms, numerical approximations, programming expressivity, and compiler optimizations. You will evangelize these constraints with various vendors to develop and influence future hardware architectures towards efficient training and inference on our models. If you are excited about efficiently distributing a large language model across devices, dealing with and optimizing system-wide/rack-wide networking bottlenecks and eventually tailoring the compute pipe and memory hierarchy of the hardware platform, simulating workloads at different abstractions and working closely with our partners, this is the perfect opportunity! This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. Key Responsibilities - Co-design future hardware for programmability and performance with our hardware vendors - Assist hardware vendors in developing optimal kernels and add support for it in our compiler - Develop performance estimates for critical kernels for different hardware configurations and drive decisions on compute core and memory hierarchy features - Build system performance models at different abstraction levels and carry out analysis to drive decisions on scale up, scale out, front end networking - Work with machine learning engineers, kernel engineers and compiler developers to understand their vision and needs from high performance accelerators - Manage communication and coordination with internal and external partners - Influence the roadmap of hardware partners to optimize them for OpenAI’s workloads. - Evaluate potential partners’ accelerators and platforms. - As the scope of the role and team grows, understand and influence roadmaps for hardware partners for our datacenter networks, racks, and buildings. Qualifications - 4+ years of industry experience, including experience harnessing compute at scale and optimizing ML platform code to run efficiently on target hardware. - Strong experience in software/hardware co-design - Deep understanding of GPU and/or other AI accelerators - Experience with CUDA, Triton or a related accelerator programming language - Experience driving Machine Learning accuracy with low precision formats - Experience with system performance modeling and analysis to optimize ML model deployment - Strong coding skills in C/C++ and Python - Are familiar with the fundamentals of deep learning computing and chip architecture/microarchitecture. - Able to actively collaborate with ML engineers, kernel writers, compiler developers, system engineers, chip architects/microarchitects Preferred Skills - PhD in Computer Science and Engineering with a specialization in Computer Architecture, Parallel Computing. Compilers or other Systems - Strong understanding of LLMs and challenges related to their training and inference About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world
Threat Modeler, Preparedness
ABOUT THE TEAM Preparedness is a critical Safety Research team at OpenAI, which is focused on mitigating AI threats https://openai.com/index/updating-our-preparedness-framework/ that could scale to an extreme level of severity. Our work involves: - Tracking and prediction. Monitoring https://openai.com/index/how-we-monitor-internal-coding-agents-misalignment/ and predicting the evolving misalignment propensities and capabilities of frontier AI systems. - Mitigation. Keeping misuse safeguards, alignment tools, and security measures on track to adequately address extreme threats that might arise in the future. - Coordination. Setting mitigation targets by maintaining OpenAI’s preparedness framework, and partnering with other staff to achieve these targets. This is urgent, fast-paced work that has far-reaching implications for the company and for society. ABOUT THE ROLE As a threat modeler, you will own OpenAI’s holistic approach to identifying, modeling, and forecasting frontier risks from frontier AI systems. This role ensures that our evaluation frameworks, safeguards, and taxonomies are robust, high-coverage, and forward-looking. You will help the company answer the “why” behind our most stringent risk-prevention efforts, shaping the rationale for prioritizing and mitigating risks across domains. You will serve as a central node connecting technical, governance, and policy perspectives on prioritization, focus and rationale on our approach to frontier risks from AI. IN THIS ROLE, YOU WILL: - Develop and maintain comprehensive threat models across all misuse areas (bio, cyber, attack planning, etc.). - Develop plausible and convincing threat models across loss of control, self-improvement, and other possible alignment risks from frontier AI systems - Forecast risks by combining technical foresight, adversarial simulation, and emerging trends. - Pair closely with technical partners on capability evaluations to ensure these map to and cover the gambit of severe risks differentially enabled by frontier AI systems. - Pair closely with Bio and Cyber Leads to size the remaining risk of the designed safeguards and translate threat models into actionable mitigation designs. - Act as the thought partner and explainer of “why” and “when” for high-investment mitigation efforts—helping stakeholders understand the rationale behind prioritization. - Serve as the central node connecting technical, governance, and policy perspectives on prioritization, focus and rationale on our approach to misuse risk. YOU MIGHT THRIVE IN THIS ROLE IF YOU: - Understand risks from frontier AI systems and have a strong grasp of AI alignment literature. Bring deep experience in threat modeling, risk analysis, or adversarial thinking (e.g., security, national security, or safety). - Know how AI evaluations work and can connect eval results to both capability testing and safeguard sufficiency. - Enjoy working across technical and policy domains to drive rigorous, multidisciplinary risk assessments. - Communicate complex risks clearly and compellingly to both technical and non-technical audiences. - Think in systems and naturally anticipate second-order and cascading risks. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional informatio
Machine Learning Engineer, Integrity
About the Team The Integrity team at OpenAI is dedicated to ensuring that our cutting-edge technology is not only revolutionary, but also secure from a myriad of adversarial threats. We strive to maintain the integrity of our platforms as they scale. The Integrity team is at the front lines of defending against misuse in all its forms: content abuse, scaled attacks, and other actions that could undermine the user experience or harm our operational stability. About the Role As a Machine Learning Engineer in OpenAI's Integrity team, you will have the opportunity to work with some of the brightest minds in AI. You’ll work on state-of-the-art models and classifiers, experiment with new architecture and approaches, and push forward our abilities in content and user understanding. You’ll help turn research breakthroughs into tangible solutions that improve the trust and safety of our platform. If you're excited about training LLMs and building ML models, this role is your chance to make a significant mark. In this role, you will: - Innovate and Deploy: Design and deploy advanced machine learning models that solve real-world problems. Bring OpenAI's research from concept to implementation, creating AI-driven applications with a direct impact. - Collaborate with the Best: Work closely with researchers, software engineers, and product managers to understand complex business challenges and deliver AI-powered solutions. Be part of a dynamic team where ideas flow freely and creativity thrives. - Optimize and Scale: Implement scalable data pipelines, optimize models for performance and accuracy, and ensure they are production-ready. Contribute to projects that require cutting-edge technology and innovative approaches. - Learn and Lead: Stay ahead of the curve by engaging with the latest developments in machine learning and AI. Take part in code reviews, share knowledge, and lead by example to maintain high-quality engineering practices. - Make a Difference: Monitor and maintain deployed models to ensure they continue delivering value. Your work will directly influence how AI benefits individuals, businesses, and society at large. You might thrive in this role if you: - Master's/ PhD degree in Computer Science, Machine Learning, Data Science, or a related field. - Demonstrated experience in deep learning and transformers models - Experience with content understanding or abuse prevention with LLMs is a plus - Proficiency in frameworks like PyTorch or Tensorflow - Strong foundation in data structures, algorithms, and software engineering principles. - Are familiar with methods of training and fine-tuning large language models, such as distillation, supervised fine-tuning, and policy optimization - Excellent problem-solving and analytical skills, with a proactive approach to challenges. - Ability to work collaboratively with cross-functional teams. - Ability to move fast in an environment where things are sometimes loosely defined and may have competing priorities or deadlines - Enjoy owning the problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employm
Research Engineer/Research Scientist, Personal AGI...
About the Team The Personal AGI team is responsible for training and improving pre-trained models to be deployed into ChatGPT, the API, and potential future products. In the Model Experience team, we shape the default character and behavior of ChatGPT: how the model communicates, responds to users, uses its capabilities, and behaves across different contexts and languages. Our goal is to make every interaction with ChatGPT thoughtful, helpful, and trustworthy. We take an opinionated view of what good human–AI interaction should look like, then turn that vision into real model behavior through human data, evaluations, reward models, and post-training. Our work sits at the intersection of research, product, and model design. We partner closely with teams across OpenAI to conduct research and ensure our models are thoughtful, safe, reliable to serve millions of users. About the Role As a Research Engineer / Scientist, you will research and develop improvements to our models. Our team works in research areas combining reinforcement learning and products. We're looking for individuals with strong ML engineering skills and research experience, especially with novel and highly capable models. An ideal candidate is passionate about product-driven research and the quality of human-AI interaction. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: - Own and pursue a research agenda to improve model capability and performance. - Collaborate closely with the other research and product teams, allowing customers to optimize their own models. - Build robust evaluations for tracking modeling improvements. - Design, implement, test, and debug code across our research stack. You might thrive in this role if you: - Have a deep understanding of machine learning and machine learning applications. - Have good judgment about model behavior and can communicate this judgment effectively. - Enjoy taking ambitious, qualitative problems and turning them into concrete training interventions. - Have a working knowledge of relevant models, and building evaluations for model capability improvement. - Are comfortable diving into a large ML codebase to debug. - Thrive in a dynamic, technically complex, and collaborative environment. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement https://cdn.openai.com/policies/eeo-policy-statement.pdf. Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from
Agent Post-Training, Computer Use Research
ABOUT THE TEAM The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. ABOUT THE ROLE As a member of Agent Post-Training, Computer Use, you will teach models to operate computers. You will help train models that can navigate browsers and desktops, use tools and applications, reason through complex workflows, collaborate with users and other agents, and complete long-horizon tasks with reliability and judgment. This work sits at the intersection of frontier model training, product behavior, evaluation, and systems engineering, and will directly shape the computer-use capabilities shipped in OpenAI’s next generation of agents. Currently, our models are the best in the world at this behavior! You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. IN THIS ROLE, YOU MIGHT - Design and run experiments that improve agentic model behavior for complex computer use https://openai.com/index/codex-for-almost-everything/, including desktop and browser. - Own end-to-end improvements to the post-training stack, including RL, data pipelines, graders, reward signals, evals, diagnostics, and model-behavior analysis. - Build evals and environments that expose the next set of model failures, then turn those failures into training data, product fixes, or new research directions. - Partner with Codex and ChatGPT product teams to understand what users need and translate product signal into model improvements. - Work on early-training and alignment interventions, including data mixtures, objectives, synthetic data, and eval loops that shape downstream agent behavior. - Help decide which integrations, capabilities, and fixes are ready for inclusion in major model runs. - Improve the machinery for large-scale training and launch: experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness. - Take on cross-functional projects that touch model training, product infrastructure, and the production agent harness, such as multi-agent systems or training directly against production-like environments. - Debug hard failures in shipped or near-shipped models and turn messy qualitative behavior into concrete hypotheses, experiments, and fixes. YOU MIGHT THRIVE IN THIS ROLE IF YOU - Have strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, and can learn quickly across the parts you have not worked in before. - Have hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems. - Are excited by open-ended problems where the path is unclear, the signal is noisy, and the right answer requires both research taste and engineering execution. - Care about product impact and model behavior, n
Research Engineer, Pretraining
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. Anthropic is at the forefront of AI research, dedicated to developing safe, ethical, and powerful artificial intelligence. Our mission is to ensure that transformative AI systems are aligned with human interests. We are seeking a Research Engineer to join our Pretraining team, responsible for developing the next generation of large language models. In this role, you will work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems. Key Responsibilities: - Conduct research and implement solutions in areas such as model architecture, algorithms, data processing, and optimizer development - Independently lead small research projects while collaborating with team members on larger initiatives - Design, run, and analyze scientific experiments to advance our understanding of large language models - Optimize and scale our training infrastructure to improve efficiency and reliability - Develop and improve dev tooling to enhance team productivity - Contribute to the entire stack, from low-level optimizations to high-level model design Qualifications: - Advanced degree (MS or PhD) in Computer Science, Machine Learning, or a related field - Strong software engineering skills with a proven track record of building complex systems - Expertise in Python and experience with deep learning frameworks (PyTorch preferred) - Familiarity with large-scale machine learning, particularly in the context of language models - Ability to balance research goals with practical engineering constraints - Strong problem-solving skills and a results-oriented mindset - Excellent communication skills and ability to work in a collaborative environment - Care about the societal impacts of your work Preferred Experience: - Work on high-performance, large-scale ML systems - Familiarity with GPUs, Kubernetes, and OS internals - Experience with language modeling using transformer architectures - Knowledge of reinforcement learning techniques - Background in large-scale ETL processes You'll thrive in this role if you: - Have significant software engineering experience - Are results-oriented with a bias towards flexibility and impact - Willingly take on tasks outside your job description to support the team - Enjoy pair programming and collaborative work &l
Research Engineer, Knowledge Team
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role: We are looking for Research Engineers to help us redesign how Claude interacts with external data sources. Many of the paradigms for how data and knowledge bases are organized assume human consumers and constraints. This is no longer true in a world of LLMs! Your job will be to design new architectures for how information is organized, and train language models to optimally use those architectures. Responsibilities: - Designing and implementing from scratch new information architecture strategies - Performing finetuning and reinforcement learning to teach language models how to interact with new information architectures - Building “hard” knowledge base eval sets to help identify failure modes of how language models work with external data - Designing and evaluating advanced agentic search capabilities. You may be a good fit if you: - Are a very experienced Python programmer who can quickly produce reliable, high quality code that your teammates love using - Have good machine learning research experience - Have experience developing software that utilizes Large Language Models such as Claude - Are results-oriented, with a bias towards flexibility and impact - Pick up slack, even if it goes outside your job description - Enjoy pair programming (we love to pair!) - Want to partner with world-class ML researchers to develop new LLM capabilities - Care about the societal impacts of your work - Have clear written and verbal communication Strong candidates will also have experience with: - Collaborating with product teams to quickly prototype and deliver innovative solutions - Building complex agentic systems that utilize LLMs - Developing scalable distributed information retrieval systems, such as search engines, knowledge graphs, RAG, indexing, ranking, query understanding, and distributed data processing The annual compensation range for this role is listed below. For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. Annual Salary: $350,000 - $850,000 USD Logistics Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience Minimum years of experience: Years of expe
Research Engineer, Model Evaluations
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role We're looking for Research Engineers to build the evaluations that tell us — and the world — what Claude can actually do. Your work will turn ambiguous notions of "intelligence" into clear, defensible metrics that researchers, leadership, and the public can rely on. You'll design and implement evaluations across the full spectrum of Claude's capabilities and personality, and build the infrastructure that runs them reliably at scale. You'll partner closely with researchers throughout the lifecycle of a new capability — from defining what to measure, to running the eval against live training checkpoints, to interpreting the results. The goal is to make Anthropic the leader in extremely well-characterized AI systems, with performance that is exhaustively measured and validated across the tasks that matter. Key responsibilities - Design and run new evaluations of Claude's capabilities — reasoning, agentic behavior, knowledge, safety properties — and produce visualizations that make the results legible to researchers and decision-makers - Build and harden the distributed eval execution platform so hundreds of evals run reliably against checkpoints throughout production RL training runs - Own the dashboards researchers and leadership use to monitor model health during training, improving signal-to-noise, reducing latency, and making regressions impossible to miss - Debug anomalous eval results mid-training-run, determine whether the cause is a model change or an infrastructure issue, and communicate the answer clearly under time pressure - Improve the tooling, libraries, and workflows researchers use to implement and iterate on evaluations - Partner with research teams across the full lifecycle of a new capability — from defining what to measure to interpreting results as training progresses - Run experiments to characterize how prompting, sampling, and scaffolding choices affect results on internal and industry benchmarks - Communicate evaluations and their results to internal stakeholders and, where approp
Research Engineer, Interpretability
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role: When you see what modern language models are capable of, do you wonder, "How do these things work? How can we trust them?" The Interpretability team at Anthropic is working to reverse-engineer how trained models work because we believe that a mechanistic understanding is the most robust way to make advanced systems safe. Think of us as doing "neuroscience" of neural networks using "microscopes" we build - or reverse-engineering neural networks like binary programs. More resources to learn about our work: - Our research blog - covering advances including Monosemantic Features and Circuits - An Introduction to Interpretability from our research lead, Chris Olah - The Urgency of Interpretability from CEO Dario Amodei - Engineering Challenges Scaling Interpretability - directly relevant to this role - 60 Minutes segment - Around 8:07, see a demo of tooling our team built - New Yorker article - what it's like to work on one of AI's hardest open problems Even if you haven’t worked on interpretability before, the infrastructure expertise is similar to what's needed across the lifecycle of a production language model: - Pretraining: Training dictionary learning models looks a lot like model pretraining - creating stable, performant training jobs for massively parameterized models across thousands of chips - Inference: Interp runs a customized inference stack. Day-to-day analysis requires services that allow editing a model's internal activations mid-forward-pass - for example, adding a "steering vector" - Performance: Like all LLM work, we push up against the limits of hardware and software. Rather than squeezing the last 0.1%, we are focused on finding bottlenecks, fixing them and moving ahead given rapidly evolving research and safety mission The science keeps scaling - and it's now applied directly in safety audits on frontier models, with real deadlines. As our research has matured, engineering and infrastructure have become a bottleneck. Your work will have a direct impact on one of the most important open problems in AI. Responsibilities: - Build and maintain the specialized inference and training infrastructure that powers interpretability research - including instrumented forward/backward passes, activation extraction, and steering vector a
Anthropic Fellows Program, AI Security
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. Apply using this link . Applications for the next cohort of Anthropic Fellows close at 11:59pm PT on July 26 . The cohort is expected to start November 2 . In some circumstances, we can accommodate fellows starting outside the usual cohort timelines — please note in your application if the November start date doesn't work for you. This page is specific to one of the Anthropic Fellows Workstreams, see also the main Anthropic Fellows posting . Anthropic Fellows Program overview The Anthropic Fellows Program is designed to foster AI research and engineering talent. We provide funding and mentorship to promising technical talent - regardless of previous experience. Fellows will primarily use external infrastructure (e.g. open-source models, public APIs) to work on an empirical project aligned with our research priorities, with the goal of producing a public output (e.g. a paper submission). In one of our earlier cohorts, over 80% of fellows produced papers. We run multiple cohorts of Fellows each year and review applications on a rolling basis. What to expect - 4 months of full-time research - Direct mentorship from Anthropic researchers - Access to a shared workspace (in either Berkeley, California or London, UK) - Connection to the broader AI safety and security research community - Weekly stipend of 3,850 USD / 2,310 GBP / 4,300 CAD + benefits (these vary by country) - Funding for compute (~$15k/month) and other research expenses Interview process The interview process will include an initial application & reference check, technical assessments & interviews, and a research discussion. We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Compensation The expected base stipend for this role is 3,850 USD / 2,310 GBP / 4,300 CAD per week, with an expectation of 40 hours per week for 4 months (with possible extension). Fellows workstreams Due to the success of the Anthropic Fellows for AI Safety Research progra
Software Engineer, Research - Human Data
ABOUT THE TEAM OpenAI’s mission is to ensure that artificial general intelligence (AGI) benefits all of humanity. A key part of achieving that mission is training models that deeply understand and reflect human preferences — the Human Data team is at the heart of that effort. The Human Data engineering team creates the systems that enable scalable, high-quality human feedback. These systems are essential to how OpenAI trains and improves its most advanced models. Engineers on this team collaborate closely with world-class researchers to bring alignment techniques to life — from experimental ideas to production-ready feedback loops. ABOUT THE ROLE We’re looking for software engineers to join the Human Data team and build the platforms, prototypes, tools, and infrastructure that power how our AI models are trained, aligned, and evaluated. You’ll partner with researchers and cross-functional teams to bring alignment ideas to life, influence future model training, and shape how models interact with the real world. We’re looking for people who are excited by technical ownership, enjoy working across the stack, and are eager to solve ambiguous problems in a high-impact, fast-paced environment. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. IN THIS ROLE, YOU WILL: - Build and maintain robust full-stack systems for feedback collection, data labeling, and evaluation pipelines, while maintaining high levels of security. - Translate experimental alignment research into scalable production infrastructure, including inference and model training stacks. - Design and iterate on user-facing tools and backend services to support high-quality data workflows - Partner with researchers, engineers, and program leads to shape feedback loops and model interaction paradigms - Drive infrastructure improvements that enable faster iteration and scaling across OpenAI’s frontier models, from internal research tooling all the way to production ChatGPT. YOU MIGHT THRIVE IN THIS ROLE IF YOU: - Have strong software engineering fundamentals and experience building production systems at scale - Enjoy full-stack development with end-to-end ownership — from backend pipelines to user interfaces - Are motivated by high-impact collaboration with research teams and solving novel, ambiguous problems - Are excited to shape how AI systems learn from human preferences and reflect a broad range of human values - Care deeply about inclusive tooling and building systems that enhance model safety, reliability, and usefulness About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement https://cdn.openai.com/policies/eeo-policy-statement.pdf. Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County
Research Engineer, Machine Learning (Reinforcement...
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the teams Our Reinforcement Learning teams lead Anthropic's reinforcement learning research and development, playing a critical role in advancing our AI systems. We've contributed to all Claude models, with significant impacts on the autonomy and coding capabilities of Claude Sonnet 4.5 and Opus 4.5. Our work spans several key areas: - Developing systems that enable models to use computers effectively - Advancing code generation through reinforcement learning - Pioneering fundamental RL research for large language models - Building scalable RL infrastructure and training methodologies - Enhancing model reasoning capabilities We collaborate closely with Anthropic's alignment and frontier red teams to ensure our systems are both capable and safe. We partner with the applied production training team to bring research innovations into deployed models, and are dedicated to implement our research at scale. Our Reinforcement Learning teams sit at the intersection of cutting-edge research and engineering excellence, with a deep commitment to building high-quality, scalable systems that push the boundaries of what AI can accomplish. About the Role As a Research Engineer within Reinforcement Learning, you will collaborate with a diverse group of researchers and engineers to advance the capabilities and safety of large language models. This role blends research and engineering responsibilities, requiring you to both implement novel approaches and contribute to the research direction. You'll work on fundamental research in reinforcement learning, creating 'agentic' models via tool use for open-ended tasks such as computer use and autonomous software generation, improving reasoning abilities in areas such as mathematics, and developing prototypes for internal use, productivity, and evaluation. Representative projects: - Architect and optimize core reinforcement learning infrastructure, from clean training abstractions to distributed experiment management across GPU clusters. Help scale our systems to handle increasingly complex research workflows. - Design, implement, and test novel training environments, evaluations, and methodologies for reinforcement learning agents which push the state of the art for the next generation of models. - Drive performance improvements across our stack through profiling, optimization, and benchmarking. Implement efficient caching solutions and debug distributed systems to accelerate both training and evaluation workflows. - Collaborate across research and engineering teams to develop automated testing frameworks, design clean APIs, and build scalable infrastructure that accelerates AI research. You may be a good fit if you: - Are proficient in Python and async/concurrent programming with frameworks like Trio - Have experience with machine learning frameworks (PyTorch, TensorFlow, JAX) - Have industry experience in machine learning research - Can balance research exploration with engineering implementation<
ML/Research Engineer, Safeguards
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role We are looking for ML Engineers and Research Engineers to help detect and mitigate misuse of our AI systems. As a member of the Safeguards ML team, you will build systems that identify harmful use—from individual policy violations to sophisticated, coordinated attacks—and develop defenses that keep our products safe as capabilities advance. You will also work on systems that protect user wellbeing and ensure our models behave appropriately across a wide range of contexts. This work feeds directly into Anthropic's Responsible Scaling Policy commitments. Responsibilities - Develop classifiers to detect misuse and anomalous behavior at scale. This includes developing synthetic data pipelines for training classifiers and methods to automatically source representative evaluations to iterate on - Build systems to monitor for harms that span multiple exchanges, such as coordinated cyber attacks and influence operations, and develop new methods for aggregating and analyzing signals across contexts - Evaluate and improve the safety of agentic products—developing both threat models and environments to test for agentic risks, and developing and deploying mitigations for prompt injection attacks - Conduct research on automated red-teaming, adversarial robustness, and other research that helps test for or find misuse You may be a good fit if you - Have 4+ years of experience in ML engineering, research engineering, or applied research, in academia or industry - Have proficiency in Python and experience building ML systems - Are comfortable working across the research-to-deployment pipeline, from exploratory experiments to production systems - Are worried about misuse risks of AI systems, and want to work to mitigate them - Have strong communication skills and ability to explain complex technical concepts to non-technical stakeholders Strong candidates may also have experience with - Language modeling and transformers - Building classifiers, anomaly detection systems, or behavioral ML - Adversarial machine learning or red-teaming - Interpretability or probes - Reinforcement learning - High-performance, large-scale ML systems The annual compensation range for this role is listed below. For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. Annual Salary: $350,000 - $500,000 USD Logistics Minimum education: Bac
[Expression of Interest] Research Manager, Interpr...
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. Note: we don't have open Research Manager positions on the Interpretability team at this time. However, we're actively growing our team of Research Engineers and Research Scientists . If you're excited about interpretability research and open to an individual contributor role, we encourage you to apply. About the Interpretability team When you see what modern language models are capable of, do you wonder, "How do these things work? How can we trust them?" The Interpretability team’s mission is to reverse engineer how trained models work, and Interpretability research is one of Anthropic’s core research bets on AI safety. We believe that a mechanistic understanding is the most robust way to make advanced systems safe. People mean many different things by "interpretability". We're focused on mechanistic interpretability, which aims to discover how neural network parameters map to meaningful algorithms. Some useful analogies might be to think of us as trying to do "biology" or "neuroscience" of neural networks, or as treating neural networks as binary computer programs we're trying to "reverse engineer". We aim to create a solid scientific foundation for mechanistically understanding neural networks and making them safe (see our vision post ). We have focused on resolving the issue of "superposition" (see Toy Models of Superposition , Superposition, Memorization, and Double Descent , and our May 2023 update ), which causes the computational units of the models, like neurons and attention heads, to be individually uninterpretable, and on finding ways to decompose models into more interpretable components. Our subsequent work which found millions of features in Claude 3.0 Sonnet, one of our production language models, represents progress in this direction. In our most recent work , we developed methods that allow us to build circuits using features and use these circuits to understand the mechanisms associated with a model's computation and study specific examples of multi-hop reasoning, planning, and chain-of-thought faithfulness on Claude Haiku 3.5, one of our production models.” This is a stepping stone towards our overall goal of mechanistically understanding neural networks. A few places to learn more about our work and team are this introduction to Interpretability from our research lead, Chris Olah, Stanford CS25 lecture given by Josh Batson, and TWIML AI podcast with E
Research Engineer / Research Scientist, Pre-traini...
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the team We are seeking passionate Research Scientists and Engineers to join our growing Pre-training team in Zurich. We are involved in developing the next generation of large language models. The team primarily focuses on multimodal capabilities: giving LLMs the ability to understand and interact with modalities other than text. In this role, you will work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems. Responsibilities In this role you will interact with many parts of the engineering and research stacks. - Conduct research and implement solutions in areas such as model architecture, algorithms, data processing, and optimizer development - Independently lead small research projects while collaborating with team members on larger initiatives - Design, run, and analyze scientific experiments to advance our understanding of large language models - Optimize and scale our training infrastructure to improve efficiency and reliability - Develop and improve dev tooling to enhance team productivity - Contribute to the entire stack, from low-level optimizations to high-level model design Qualifications & Experience We encourage you to apply even if you do not believe you meet every single criterion. Because we focus on so many areas, the team is looking for both experienced engineers and strong researchers, and encourage anyone along the researcher/engineer spectrum to apply. - Degree (BA required, MS or PhD preferred) in Computer Science, Machine Learning, or a related field - Strong software engineering skills with a proven track record of building complex systems - Expertise in Python and deep learning frameworks - Have worked on high-performance, large-scale ML systems, particularly in the context of language modeling - Familiarity with ML Accelerators, Kubernetes, and large-scale data processing - Strong problem-solving skills and a results-oriented mindset - Excellent communication skills and ability to work in a collaborative environment You'll thrive in this role if you - Have significant software engineering experience - Are able to balance research goals with practical engineering constraints - Are happy to take on tasks outside your job description to support the team - Enjoy pair programming and collaborative work - Are eager to learn more about machine learning research &l
Technical Recruiter, Infrastructure
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role Anthropic's Infrastructure organization is foundational to our mission of developing AI systems that are reliable, interpretable, and steerable. The systems we build determine how quickly we can train new models, how reliably we can run safety experiments, and how effectively we can scale Claude to millions of users — demonstrating that safe, reliable infrastructure and frontier capabilities can go hand in hand. As Technical Recruiter, Infrastructure, you'll join the small team of recruiters who hire for that organization, owning full lifecycle recruiting for your searches and partnering with infrastructure leaders to turn ambiguous needs into clear search strategies. Key responsibilities - Own full lifecycle recruiting for a portfolio of roles across the Infrastructure organization, from intake through offer and close - Run structured intakes with infrastructure hiring managers, translating ambiguous needs into scoped requirements, calibrated bars, and search strategies - Build and maintain pipelines of specialized infrastructure talent, with an emphasis on passive candidates - Refine infrastructure interview loops, take-home assignments, and scorecards alongside hiring managers, your recruiting counterparts, and Recruiting Operations - Develop deep domain knowledge aligned with the teams you support, so you can identify niche talent with the right specific domain fit - Advise hiring managers with market data and candid calibration feedback, and influence decisions through credibility rather than volume - Partner with Compensation, People Partners, and Mobility to structure equitable offers and guide candidates to close - Handle sensitive role and candidate information with discretion, including for searches whose scope is confidential Minimum qualifications - Deep full lifecycle recruiting experience, with substantial time supporting infrastructure, platform, or comparably technical engineering organizations - Ability to hold a substantive technical conversation about infrastructure domains such as Kubernetes and container orchestration, cloud networking, cluster networking, and systems languages, and to evaluate technical qualifications rather than match keywords - Proficiency with an applicant tracking system like Greenhouse and other modern sourcing tools - Experience partnering directly with hiring managers on intake, bar calibration, and interview loop design - Sound independent judgment on candidate quality, and the ability to independently partner with multiple hiring managers on complex searches - A strong sense of ownership over your work, and the adaptability to adjust as priorities and hiring needs shift - Genuine interest in Anthropic's mission and in the role a strong infrastructure function plays in achieving it <h2>
Research Engineer, Pretraining Scaling - London
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the Role: Anthropic's ML Performance and Scaling team trains our production pretrained models, work that directly shapes the company's future and our mission to build safe, beneficial AI systems. As a Research Engineer on this team, you'll ensure our frontier models train reliably, efficiently, and at scale. This is demanding, high-impact work that requires both deep technical expertise and a genuine passion for the craft of large-scale ML systems. This role lives at the boundary between research and engineering. You'll work across our entire production training stack: performance optimization, hardware debugging, experimental design, and launch coordination. During launches, the team works in tight lockstep, responding to production issues that can't wait for tomorrow. Responsibilities: - Own critical aspects of our production pretraining pipeline, including model operations, performance optimization, observability, and reliability - Debug and resolve complex issues across the full stack—from hardware errors and networking to training dynamics and evaluation infrastructure - Design and run experiments to improve training efficiency, reduce step time, increase uptime, and enhance model performance - Respond to on-call incidents during model launches, diagnosing problems quickly and coordinating solutions across teams - Build and maintain production logging, monitoring dashboards, and evaluation infrastructure - Add new capabilities to the training codebase, such as long context support or novel architectures - Collaborate closely with teammates across SF and London, as well as with Tokens, Architectures, and Systems teams - Contribute to the team's institutional knowledge by documenting systems, debugging approaches, and lessons learned You May Be a Good Fit If You: - Have hands-on experience training large language models, or deep expertise with JAX, TPU, PyTorch, or large-scale distributed systems - Genuinely enjoy both research and engineering work—you'd describe your ideal split as roughly 50/50 rather than heavily weighted toward one or the other - Are excited about being on-call for production systems, working long days during launches, and solving hard problems under pressure - Thrive when working on whatever is most impactful, even if that changes day-to-day based on what the production model needs - Excel at debugging complex, ambiguous problems across multiple layers of the stack - Communicate clearly and collaborate effectively, especially when coordinating across time zones or during high-stress incidents - Are passionate about the work itself and want to refine your craft as a research engineer - Care about the societal impacts of AI and responsible scaling Strong Candidates May Also Have: - Previous experience training LLM’s or working extensively with JAX/TPU, PyTorch, or other ML frameworks at scale - Contributed to open-source LLM frame
Research Scientist, Life Sciences (Experimental Bi...
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the team Anthropic's Life Sciences team is building a world-class research group focused on making fundamental biological discoveries. The team combines cutting-edge AI with hands-on biological research, positioning Anthropic at the forefront of AI-accelerated scientific discovery. About the role We're seeking an exceptional Research Scientist to join the team. As a founding member of Life Sciences, you'll work in a high-impact group that operates at the intersection of computational and experimental biology. You'll help establish Anthropic as a leader in biology research while developing product intuition through direct engagement with the challenges and opportunities of laboratory science. Key responsibilities - Design, execute, and iterate on the experimental programs at the core of the team's research: molecular biology, biochemistry, protein and nucleic acid characterization, high-throughput functional screens, and the assay development that makes new questions answerable - Partner directly with computational biologists to design experiments that produce high-quality, analysis-ready data, and feed results back fast enough to immediately inform the next round of analysis - Generate and prioritize hypotheses by combining your experimental judgment with the literature, curated biological knowledge bases, and the team's computational predictions - Use Claude and our internal agent frameworks heavily in your own work — for experimental planning, protocol development, and data interpretation — and feed what you learn back to the model-improvement and product teams as evaluations, datasets, and concrete failure cases Minimum qualifications - Have a Ph.D. in a biological science (molecular biology, biochemistry, bioengineering, computational biology) or a related field - Have a track record of bridging biological domain knowledge with computational approaches to solve real scientific problems - Have basic proficiency in Python and are familiar with ML development practices <h2 class="text-text-100 mt-3 -mb-1 text-[1.125rem] font-bold
Research Engineer, Performance RL (Reinforcement L...
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the RL Teams Our Reinforcement Learning teams lead Anthropic's reinforcement learning research and development, playing a critical role in advancing our AI systems. We've contributed to all Claude models, with significant impacts on the autonomy and coding capabilities of Claude Sonnet 4.6 and Opus 4.6. Our work spans several key areas: - Developing systems that enable models to use computers effectively - Advancing code generation through reinforcement learning - Pioneering fundamental RL research for large language models - Building scalable RL infrastructure and training methodologies - Enhancing model reasoning capabilities We collaborate closely with Anthropic's alignment and frontier red teams to ensure our systems are both capable and safe. We partner with the applied production training team to bring research innovations into deployed models, and are dedicated to implement our research at scale. Our Reinforcement Learning teams sit at the intersection of cutting-edge research and engineering excellence, with a deep commitment to building high-quality, scalable systems that push the boundaries of what AI can accomplish. About the Role We're hiring for the Code RL team within the RL organization. As a Research Engineer, you'll advance our models' ability to safely write correct, fast code for accelerators. You'll need to know accelerator performance well to turn it into tasks and signals models can learn from. Specifically, you will: - Invent, design and implement RL environments and evaluations. - Conduct experiments and shape our research roadmap. - Deliver your work into training runs. - Collaborate with other researchers, engineers, and performance engineering specialists across and outside Anthropic. You may be a good fit if you: - Have expertise with accelerators (CUDA, ROCm, Triton, Pallas), ML framework programming (JAX or PyTorch). - Have worked across the stack – kernels, model code, distributed systems. - Know how to balance research exploration with engineering implementation. - Are passionate about AI's potential and committed to developing safe and beneficial systems. Strong candidates may also have: - Experience with reinforcement learning. - Experience porting ML workloads between different types of accelerators. - Familiarity with LLM training methodologies. The annual compensation range for this role is listed below. For sales roles, the range provided is th
Software Engineer, Privacy
About the Team The Privacy Team at OpenAI is committed to building a secure and trustworthy platform. Our area of responsibility encompasses all OpenAI products and systems that process user data. We provide cross-functional partners with the tools needed to ensure that all products adhere to the highest standards of data privacy and legal compliance. Our approach to prioritizing responsible data use is integral to OpenAI's mission of safely introducing Artificial General Intelligence (AGI) that offers widespread benefits. About the Role We’re in search of a Software Engineer with experience building data pipelines and working closely with members of the Legal team. This role is perfect for someone who's passionate about the intersection of systems, privacy, and legal compliance. You will architect, design, and write backend systems responsible for handling some of the most sensitive data at OpenAI. This role is based in Dublin, Ireland. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: - Design, build, and maintain back-end systems and services that power privacy and data compliance functions within our API products and consumer applications. - Work closely with legal advisors and other engineers to respond to court orders and other legal processes, all while upholding strict data privacy and legal standards. - Identify opportunities for automation and build the tools that enable other teams to automate tasks involving customer data. - Develop and implement data handling policies and procedures in compliance with legal and ethical standards, ensuring the integrity and confidentiality of user data. You might thrive in this role if you: - Have experience building data pipelines, especially for legal processes and investigative workflows. - Can translate legal requirements into technical solutions and explain technical solutions to a non-technical audience. - Take responsibility for problems from beginning to end, and are prepared to acquire any missing knowledge necessary to get the job done. - Create tools to speed up your own and your colleagues’ workflows, particularly when pre-existing solutions are inadequate. - Deeply care about user experience and take pride in developing products that meet customer needs while ensuring privacy. - Have a background in security investigations or experience working in collaboration with trust and safety, legal, and engineering teams. Compensation, Benefits and Perks This is a position with OpenAI Ireland Ltd., which controls the hiring and management of this position. Total compensation includes an annual salary, generous equity, and benefits. - Medical, dental, and vision insurance for you and your family - Mental health and wellness support - PRSA plan with 6% employer matching - Unlimited time off - Annual learning & development stipend (€1,400 per year) About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement https://cdn.openai.com/policies/eeo-policy-statement.pdf. Background checks for applicants will be administered in accordance with applicable law, and qualified applicants w
Research Engineer, Codex
ABOUT THE TEAM The Codex Research team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. ABOUT THE ROLE As a member of the Codex Research team, you will improve the capabilities, reliability, and product fit of OpenAI's agentic models. You might own a research direction, build the infrastructure that makes large training runs faster and more trustworthy, create evals that reveal where models fail, or drive a capability from an idea through experimentation, integration, and launch. This role is intentionally broad. The strongest candidates are not defined by one method or subfield; they are people who can take an ambiguous capability problem and make progress across research, engineering, data, evals, and product. You should be excited to work on models that act in the world: writing and debugging code, using tools, calling functions, operating computers, collaborating with other agents, and completing valuable work on behalf of users. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. IN THIS ROLE, YOU MIGHT: - Design and run experiments that improve agentic model behavior across coding, tool use, function calling, computer use, multi-agent collaboration, long-horizon tasks, factuality, instruction following, and calibrated reasoning. - Own end-to-end improvements to the post-training stack, including RL, data pipelines, graders, reward signals, evals, diagnostics, and model-behavior analysis. - Build evals and environments that expose the next set of model failures, then turn those failures into training data, product fixes, or new research directions. - Partner with Codex, API/platform, ChatGPT, and general-agent product teams to understand what users need and translate product signal into model improvements. - Work on early-training and alignment interventions, including data mixtures, objectives, synthetic data, and eval loops that shape downstream agent behavior. - Help decide which integrations, capabilities, and fixes are ready for inclusion in major model runs. - Improve the machinery for large-scale training and launch: experiment velocity, reliability, observability, reproducibility, cost, latency, and production readiness. - Take on cross-functional projects that touch model training, product infrastructure, and the production agent harness, such as multi-agent systems or training directly against production-like environments. - Debug hard failures in shipped or near-shipped models and turn messy qualitative behavior into concrete hypotheses, experiments, and fixes. YOU MIGHT THRIVE IN THIS ROLE IF YOU: - Have strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field, and can learn quickly across the parts you have not worked in before. - Have hands-on experience with LLMs, RL, RLHF/RLAIF, post-training, evals,
Research Engineer, Pretraining Scaling
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About the Role: Anthropic's ML Performance and Scaling team trains our production pretrained models, work that directly shapes the company's future and our mission to build safe, beneficial AI systems. As a Research Engineer on this team, you'll ensure our frontier models train reliably, efficiently, and at scale. This is demanding, high-impact work that requires both deep technical expertise and a genuine passion for the craft of large-scale ML systems. This role lives at the boundary between research and engineering. You'll work across our entire production training stack: performance optimization, hardware debugging, experimental design, and launch coordination. During launches, the team works in tight lockstep, responding to production issues that can't wait for tomorrow. Responsibilities: - Own critical aspects of our production pretraining pipeline, including model operations, performance optimization, observability, and reliability - Debug and resolve complex issues across the full stack—from hardware errors and networking to training dynamics and evaluation infrastructure - Design and run experiments to improve training efficiency, reduce step time, increase uptime, and enhance model performance - Respond to on-call incidents during model launches, diagnosing problems quickly and coordinating solutions across teams - Build and maintain production logging, monitoring dashboards, and evaluation infrastructure - Add new capabilities to the training codebase, such as long context support or novel architectures - Collaborate closely with teammates across SF and London, as well as with Tokens, Architectures, and Systems teams - Contribute to the team's institutional knowledge by documenting systems, debugging approaches, and lessons learned You May Be a Good Fit If You: - Have hands-on experience training large language models, or deep expertise with JAX, TPU, PyTorch, or large-scale distributed systems - Genuinely enjoy both research and engineering work—you'd describe your ideal split as roughly 50/50 rather than heavily weighted toward one or the other - Are excited about being on-call for production systems, working long days during launches, and solving hard problems under pressure - Thrive when working on whatever is most impactful, even if that changes day-to-day based on what the production model needs - Excel at debugging complex, ambiguous problems across multiple layers of the stack - Communicate clearly and collaborate effectively, especially when coordinating across time zones or during high-stress incidents - Are passionate about the work itself and want to refine your craft as a research engineer - Care about the societal impacts of AI and responsible scaling Strong Candidates May Also Have: - Previous experience training LLM’s or working extensively with JAX/TPU, PyTorch, or other ML frameworks at scale - Contributed to open-source LLM frame
Research Engineer/Research Scientist, RL/Reasoning
About the Team The RL and Reasoning team drives the core reasoning paradigm and has created groundbreaking innovations such as o1 and o3. They focus on pushing the boundaries of reinforcement learning research, building next-generation generative models, and deploying them at scale. About the Role As a Research Engineer/Research Scientist at OpenAI, you will advance the frontier of AI alignment and capabilities through cutting-edge RL methods. Your work will sit at the heart of training intelligent, aligned, and general-purpose agents, including the systems that power various models. We’re looking for people who have a background in reinforcement learning research, are able to iterate quickly, and are proficient at coding. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. You might thrive in this role if: - You love being on the cutting edge of RL and language model research. - You’re a self-starter who takes initiative and ownership of ideas, driving them to completion. - You value principled approaches, simple experiments in tightly-controlled settings, and reaching trustworthy conclusions which stand the test of time. - You thrive in a fast-paced, dynamic, and technically complex environment where rapid iteration is key. - You’re comfortable diving into a large ML codebase to debug and improve it. - You have a deep understanding of machine learning and machine learning applications. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement https://cdn.openai.com/policies/eeo-policy-statement.pdf. Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data security obligations. To notify OpenAI that you believe this job posting is non-compliant, please submit a report through this form https://form.asana.com/?d=57018692298241&k=5MqR40fZd7jlxVUh5J-UeA. No response will be provided to inquiries unrelated to job posting compliance. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via this link https://form.asana.com/?k=bQ7w9h3iexRlicUdWRiwvg&d=57018692298241. OpenAI Global
Research Engineer/Research Scientist, Pre-training
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. Anthropic is at the forefront of AI research, dedicated to developing safe, ethical, and powerful artificial intelligence. Our mission is to ensure that transformative AI systems are aligned with human interests. We are seeking a Research Engineer to join our Pre-training team, responsible for developing the next generation of large language models. In this role, you will work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems. Key Responsibilities: - Conduct research and implement solutions in areas such as model architecture, algorithms, data processing, and optimizer development - Independently lead small research projects while collaborating with team members on larger initiatives - Design, run, and analyze scientific experiments to advance our understanding of large language models - Optimize and scale our training infrastructure to improve efficiency and reliability - Develop and improve dev tooling to enhance team productivity - Contribute to the entire stack, from low-level optimizations to high-level model design Qualifications: - Advanced degree (MS or PhD) in Computer Science, Machine Learning, or a related field - Strong software engineering skills with a proven track record of building complex systems - Expertise in Python and experience with deep learning frameworks (PyTorch preferred) - Familiarity with large-scale machine learning, particularly in the context of language models - Ability to balance research goals with practical engineering constraints - Strong problem-solving skills and a results-oriented mindset - Excellent communication skills and ability to work in a collaborative environment - Care about the societal impacts of your work Preferred Experience: - Work on high-performance, large-scale ML systems - Familiarity with GPUs, Kubernetes, and OS internals - Experience with language modeling using transformer architectures - Knowledge of reinforcement learning techniques - Background in large-scale ETL processes You'll thrive in this role if you: - Have significant software engineering experience - Are results-oriented with a bias towards flexibility and impact - Willingly take on tasks outside your job description to support the team - Enjoy pair programming and collaborative work - Are eager to learn more about machine learning research - Are enthusiastic to work at an organization that functions as a single, cohesive team pursuing large-scale AI research projects - Are working to align state of the art models with human values and preferences, understand and interpret deep neural networks, or develop new models to support these areas of research - View research and engineering as
Researcher, Frontier Cybersecurity Risks
ABOUT THE TEAM Preparedness is a critical Safety Research team at OpenAI, which is focused on mitigating AI threats to global security https://openai.com/index/updating-our-preparedness-framework/ that could scale to an extreme level of severity. Our work involves: 1. Measurement. Monitoring and predicting the evolving capabilities of frontier AI systems. 2. Mitigation. Keeping misuse safeguards, alignment tools, and security measures on track to adequately address extreme threats that might arise in the future. 3. Coordination. Setting mitigation targets by maintaining OpenAI’s preparedness framework https://openai.com/index/updating-our-preparedness-framework/, and partnering with other staff to achieve these targets. This is urgent, fast-paced work that has far-reaching implications for the company and for society. ABOUT THE ROLE Models are becoming increasingly capable—moving from tools that assist humans to agents that can plan, execute, and adapt in the real world. As we push toward AGI, cybersecurity becomes one of the most important and urgent frontiers: the same systems that can accelerate productivity can also accelerate exploitation. As a Researcher for cybersecurity risks, you will help design and implement an end-to-end mitigation stack to reduce severe cyber misuse across OpenAI’s products. This role requires strong technical depth and close cross-functional collaboration to ensure safeguards are enforceable, scalable, and effective. You’ll contribute directly to building protections that remain robust as products, model capabilities, and attacker behaviors evolve. IN THIS ROLE, YOU WILL: - Design and implement mitigation components for model-enabled cybersecurity misuse—spanning prevention, monitoring, detection, and enforcement—under the guidance of senior technical and risk leadership. - Integrate safeguards across product surfaces in partnership with product and engineering teams, helping ensure protections are consistent, low-latency, and scale with usage and new model capabilities. - Evaluate technical trade-offs within the cybersecurity risk domain (coverage, latency, model utility, and user privacy) and propose pragmatic, testable solutions. - Collaborate closely with risk and threat modeling partners to align mitigation design with anticipated attacker behaviors and high-impact misuse scenarios. - Execute rigorous testing and red-teaming workflows, helping stress-test the mitigation stack against evolving threats (e.g., novel exploits, tool-use chains, automated attack workflows) and across different product surfaces—then iterate based on findings. YOU MIGHT THRIVE IN THIS ROLE IF YOU: - Have a passion for AI safety and are motivated to make cutting-edge AI models safer for real-world use. - Bring demonstrated experience in deep learning and transformer models. - Are proficient with frameworks such as PyTorch or TensorFlow. - Possess a strong foundation in data structures, algorithms, and software engineering principles. - Are familiar with methods for training and fine-tuning large language models, including distillation, supervised fine-tuning, and policy optimization. - Excel at working collaboratively with cross-functional teams across research, security, policy, product, and engineering. - Have significant experience designing and deploying technical safeguards for abuse prevention, detection, and enforcement at scale. - (Nice to have) Bring background knowledge in cybersecurity or adjacent fields. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectr
Software Engineer, Fleet Infrastructure
This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running the world’s largest, most reliable, and frictionless GPU fleet to support OpenAI’s general purpose model training and deployment. Work on this team ranges from - Maximizing GPUs doing useful work by building user-friendly scheduling and quota systems - Running a reliable and low maintenance platform by building push-button automation for kubernetes cluster provisioning and upgrades - Supporting research workflows with service frameworks and deployment systems - Ensuring fast model startup times though high performance snapshot delivery across blob storage down to hardware caching - Much more! About the Role As an engineer within Fleet infrastructure, you will design, write, deploy, and operate infrastructure systems for model deployment and training on one of the world’s largest GPU fleet. The scale is immense, the timelines are tight, and the organization is moving fast; this is an opportunity to shape a critical system in support of OpenAI's mission to advance AI capabilities responsibly. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: - Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems. - Interface with researchers and product teams to understand workload requirements - Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service You might thrive in this role if you: - Have experience with hyperscale compute systems - Possess strong programming skills - Have experience working in public clouds (especially Azure) - Have experience working in Kubernetes - Execution focused mentality paired with a rigorous focus on user requirements - As a bonus, have an understanding of AI/ML workloads About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional information, please see OpenAI’s Affirmative Action and Equal Employment Opportunity Policy Statement https://cdn.openai.com/policies/eeo-policy-statement.pdf. Background checks for applicants will be administered in accordance with applicable law, and qualified applicants with arrest or conviction records will be considered for employment consistent with those laws, including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act, for US-based candidates. For unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. In addition, job duties require access to secure and protected information technology systems and related data sec
Recruiter, AI/ML Research EMEA
About the Team OpenAI’s mission is to build safe artificial general intelligence (AGI) that benefits all of humanity. Achieving this requires bringing the world’s most exceptional talent under one roof to push the boundaries of what’s possible. Our Research Recruiting team plays a critical role in this effort. We are an embedded part of the research organization, working side by side with our research staff to deeply understand evolving priorities, build trust, and strategically shape the future of OpenAI’s talent. About the Role You will own and execute long-term talent strategies to identify, engage, and recruit many of the world’s leading and emerging AI researchers, research engineers, and technical scientists working at the frontier of machine learning. This is not a traditional execution-focused recruiting role. You will operate as a strategic partner to OpenAI’s research staff, helping define hiring priorities, shape search strategy, influence candidate evaluation, and guide hiring decisions that directly impact the direction and quality of our frontier-model research and fulfillment of our mission. In this role, you will: - Partner directly with research and technical staff to define hiring priorities, shape search strategies, and anticipate future talent needs as technical roadmaps evolve. - Proactively identify and cultivate exceptional AI/ML research talent across industry, academia, and emerging labs, often before formal hiring needs exist. - Use market insights and candidate signals to influence hiring decisions, leveling, and compensation strategy for highly specialized research roles. - Serve as a trusted advisor throughout candidate evaluation and closing — helping leaders calibrate for research excellence, long-term potential, and organizational fit. - Collaborate closely with your sourcing partner to execute complex, high-impact searches in ambiguous or rapidly evolving technical domains. You might thrive in this role if you: - Significant experience recruiting within highly technical or specialized environments. - Deep interest in AI research and a desire to engage directly with global research communities. - Experience recruiting within highly technical or specialized environments such as ML/AI, distributed systems, infrastructure, scientific computing, or quantitative research. - Track record of leading complex, ambiguous technical searches from early talent mapping through close. - Experience navigating high-stakes negotiations with senior technical or research candidates. - Comfort operating in fast-moving environments where hiring priorities and role definitions may evolve over time. Workplace & Location This role is based in our London office and we aren’t considering remote applications at this time. We use a hybrid work model of 3 days in the office with optional work from home on Thursdays and Fridays. We also offer relocation assistance to new employees. Our open-plan offices have height-adjustable desks, conference rooms, phone booths, well-stocked kitchens full of snacks and drinks, three in-house prepared meals daily, outdoor space for working and socializing, wellness rooms, private bike storage, and more. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. We are an equal opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. For additional informat
Get new jobs by email
A weekly edit of the newest roles. No spam.