AI QA Tester
AI QA Tester
We are currently recruiting for an AI QA Tester with automation testing experience to join one of our Insurance clients on a 6-month contract.
Inside IR35
Remote
Key responsibilities
- Own and maintain the test strategy for the agentic platform, spanning the agent orchestration layer, the data APIs, the retrieval pipeline and the underlying Azure services.
- Build and curate golden test datasets and expected-outcome sets with input from business subject-matter experts, covering happy paths, edge cases, ambiguous questions and out-of-scope requests.
- Design and automate LLM evaluation: groundedness / faithfulness, answer relevance, retrieval precision and recall, citation correctness, task completion, tone and consistency - using Azure AI Foundry evaluation, LLM-as-judge and deterministic checks as appropriate.
- Wire evaluations into CI/CD as release gates, so prompt, model, index or code changes cannot regress quality unnoticed; report trends over time.
- Carry out adversarial and red-team testing: prompt injection, jailbreaks, indirect injection via knowledge-base content, data-exfiltration attempts, tool misuse and privilege escalation through agent tool calls.
- Validate the safety and monitoring controls - Azure AI Content Safety categories and thresholds, refusal and escalation behaviour, and Azure AI Anomaly Detector signals - including deliberate false-negative and false-positive probing.
- Test authorisation rigorously: confirm that agents and APIs never return member or document data outside the requesting user's entitlement, including via retrieval or summarisation side channels.
- Build automated functional and contract test suites for the Agent Data API and Member Data API (for example pytest, Playwright, Postman/Newman, RestSharp) and for the end-to-end conversational flows.
- Run non-functional testing: latency and response-time budgets, throughput and concurrency, token cost per interaction, rate-limit and failover behaviour, and resilience when BCUK or a dependency degrades.
experience
- Strong QA engineering background with hands-on test automation, not manual scripting alone.
- Automation coding ability in Python, C# and/or JavaScript/TypeScript.
- API testing experience, including contract testing, authentication, negative testing and mocked dependencies.
- Demonstrable experience testing AI, machine learning or LLM-based systems, or a clear grasp of how to assure non-deterministic output using evaluation metrics and statistical thresholds.
- Understanding of RAG architectures and where they fail - retrieval misses, stale indexes, chunk boundary loss, hallucination, ungrounded citations.
- Awareness of LLM security risks (for example the OWASP Top 10 for LLM applications) and practical adversarial testing technique.
- CI/CD integration of automated test suites (GitHub Actions or Azure DevOps).
- Test data management, including handling of personal data safely in test environments.
- Confidence challenging engineers and stakeholders on release readiness.
- Experience with Azure AI Foundry evaluations, Azure AI Content Safety testing or PyRIT / other red-teaming tooling.
- Performance and load testing tools (k6, JMeter, Azure Load Testing).
- Azure Monitor / Application Insights and KQL for investigating failures from telemetry.
- Accessibility and conversational UX testing experience.
- Familiarity with UK GDPR, DPIA processes or regulated-sector assurance evidence.
Guidant, Carbon60, Lorien & SRG - The Impellam Group Portfolio are acting as an Employment Business in relation to this vacancy.
Similar Jobs
Apply to this Job
Share this Job
