Job Title: AI Safety Evaluation & Monitoring – Contract
Location: REMOTE – must work EST
W2 ONLY
We are looking for an experienced AI Safety Evaluation & Monitoring professional to measure and monitor risks within conversational and agentic AI products.
This is not a traditional Data Scientist role. This position focuses specifically on AI safety measurement, evaluation, and production monitoring—identifying AI risks, determining how those risks should be measured, and building monitoring systems that provide ongoing safety visibility as new AI features launch.
Responsibilities
- Develop safety monitoring for conversational, recommendation, and agentic/tool-using AI systems.
- Design safety measurements for single- and multi-turn AI experiences.
- Define risk metrics, thresholds, decision criteria, and model performance measurements.
- Build evaluation datasets, scoring rubrics, and safety evaluation frameworks.
- Build reusable Python pipelines and LLM-as-a-judge evaluation workflows.
- Develop monitoring and reporting dashboards for AI safety risks.
- Evaluate production AI behavior and identify emerging safety risks.
- Translate findings into safety mitigations and improvements to AI models, policies, and products.
- Establish monitoring coverage as new AI features are released.
Required Skills & Experience
- AI Safety Experience: Must have hands-on experience evaluating, measuring, and monitoring AI/LLM safety risks in a production AI product.
- AI Safety Measurement: Must have built AI safety evaluation frameworks, risk metrics, scoring rubrics, and production monitoring systems to identify and track AI safety risks.
- Conversational & Agentic AI Evaluation: Must have hands-on experience evaluating conversational AI/LLMs, including multi-turn and agentic/tool-using AI systems.
- Experience evaluating conversational AI and LLM-based products.
- Experience evaluating multi-turn conversational AI systems.
- Experience evaluating agentic or tool-using AI systems.
- Hands-on experience building and calibrating LLM-as-a-judge evaluations.
- Experience with human-in-the-loop AI evaluations.
- Experience with multilingual AI evaluation.
- Strong hands-on Python skills for building evaluation and monitoring pipelines.
- Strong SQL skills for analyzing evaluation and production monitoring data.
About SSI People: With over 26 years of industry experience, SSi People has built its reputation and expertise on putting people first. Everything we do works toward delivering exceptional experience for our consultants, our clients, and our internal team. Through a genuine commitment to people in everything we do. We have developed refined processes and a stellar internal team to deliver talent quickly. More importantly, we focus on building long-term relationships, not transactions. Putting people first is just what we do well.
By applying for this job, you agree to receive calls, AI-generated calls, text messages, or emails from SSi People and its affiliates, and contracted partners. Frequency varies for text messages. Message and data rates may apply. Carriers are not liable for delayed or undelivered messages. You can reply STOP to cancel and HELP for help. You can access our privacy policy here: SSi People Privacy Policy