AI Quality Analyst – English
Remote | Contractor | 3-Month Engagement
4–40 hours per week | 4 hours of overlap with PST required
About the Role
As an AI Quality Analyst – English, you will evaluate a new personalization feature for an advanced AI assistant. You will assess how effectively the model uses information from previous conversations, email, search activity, and video activity to make responses more relevant, natural, and helpful.
This role combines creative prompt design, analytical evaluation, attention to detail, and strong written communication. You will create prompts based on your own personal experiences and context, then evaluate how accurately and naturally the AI incorporates that information into its responses.
You will assess dimensions such as Grounding, Integration, and Helpfulness, while identifying incorrect personalization, unsupported inferences, hallucinations, and unnatural use of personal information.
What You’ll Do
Conversational Prompt Design
Design and execute multi-turn conversational prompts, typically involving 1–5 turns.
Create prompts that require the AI to appropriately use personal information and experiences.
Approach evaluations creatively to thoroughly test the model’s personalization capabilities.
Develop prompts based on realistic personal context and intended outcomes.
AI Response Evaluation
Evaluate model responses based on the intent established in the starting prompt.
Determine whether personalization was applied appropriately and meaningfully.
Identify incorrect personalization, poor inferences, forced connections, and other quality issues.
Assess whether responses are natural, relevant, useful, and easy to use.
Grounding & Personalization Analysis
Analyze responses for grounding issues.
Verify that claims about you are supported by available evidence.
Identify flawed inferences, unsupported claims, and hallucinations.
Evaluate whether personal information is incorporated accurately rather than simply referenced unnecessarily.
Integration & Naturalness
Assess how naturally personal information is integrated into responses.
Identify robotic behavior, excessive explanation, or unnecessary "overnarrating."
Evaluate whether personalization improves the overall quality of the response.
Side-by-Side Evaluation
Rigorously compare two model responses side-by-side.
Stack-rank responses based on overall helpfulness, usability, naturalness, and enjoyment.
Identify subtle differences between responses that may affect the user experience.
Provide clear reasoning for your rankings.
Evaluation Documentation
Write concise, structured, and defensible rationales for model comparisons.
Explicitly reference relevant turn numbers when explaining issues or positive aspects.
Provide constructive feedback and detailed annotations.
Extract and verify available debugging information to confirm that chat summaries and relevant data sources were properly utilized.
Data Hygiene
Maintain strict data hygiene throughout evaluation activities.
Delete evaluation conversations as required to prevent them from affecting future chat history.
Required Qualifications
Strong English reading and writing skills, with a high degree of comprehension.
Ability to evaluate nuanced and ambiguous AI responses.
Strong analytical thinking and judgment.
Experience designing creative, multi-turn prompts based on personal context.
Ability to identify incorrect personalization, poor inferences, and forced connections.
Exceptional attention to detail when comparing model responses.
Ability to write clear, concise, and structured evaluation rationales.
Ability to provide constructive feedback and detailed annotations.
Strong communication and collaboration skills.
Ability to work independently in a remote environment.
Willingness to use your primary personal Google account, rather than a testing account, and enable the required personal data sources for genuine evaluation.
Desktop or laptop with a reliable internet connection.
Full-time availability within your local time zone is required, as the team operates across a global 24-hour schedule.
Education & Experience
BS/BA degree or equivalent experience in a relevant field such as:
Policy
Law
Ethics
Linguistics
Journalism
Computer Science
A related analytical field
Experience in data annotation, AI quality evaluation, content moderation, or a related role is strongly preferred.
Core Skill
Domain-Specific Languages
Engagement Details
Engagement Type: Contractor
Engagement Length: 3 months
Availability: At least 4 hours per day, up to 40 hours per week
Schedule Requirement: 4 hours of overlap with PST
Work Arrangement: Remote
What Success Looks Like
Success in this role means consistently producing accurate, evidence-based, and well-documented AI evaluations. You will demonstrate strong judgment when assessing personalization quality, distinguish genuine contextual relevance from forced or incorrect personalization, and clearly explain why one model response performs better than another.
Your evaluations should be precise, defensible, and grounded in specific evidence from the conversations being assessed.
Evaluation Process
Shortlisted candidates will receive a Job Interest Form.
Following profile review, an assessment will be shared.
The assessment must be completed within 24 hours.
Candidates selected based on assessment results will be contacted regarding pre-onboarding requirements.