Discovery Internship: Prompt & Evaluation Engineer Intern (Sum Vivas)

Details


1117791


Liverpool John Moores University


19/10/2026


Other - See Job


100 Hours


Smart Casual

Pay


£13.45


£1.62

Description

Role

***Ringfenced to Level 5 & 6 LJMU students only***

***Applications must be submitted using the application form linked below. Please upload it instead of a CV. Any other method will not be accepted***

Download application form

Duties and responsibilities

The role:

Sum Vivas builds Digital Employees - AI agents deployed across the visitor economy, events, hospitality and healthcare.

This role owns both halves of that loop. You will write and iterate the prompts that define agent behaviour, and you build the evaluation and analytics that prove whether the changes worked. It is a hands-on engineering role.

What you will own:

Prompt engineering across our live agent estate writing, versioning and iterating multimodule system prompts covering identity, tool routing, escalation logic, guardrails and channel behaviour Diagnosing behavioural failures from real transcripts: escalation bugs, channel misdetection, tool misfires, intent misclassification then fixing them at the prompt level and proving the fix Our LLM-as-judge evaluation pipeline: designing scoring rubrics, writing judge prompts with strict JSON schemas, running batch evaluations and validating judge output against human review. Session analytics across the platform surfacing drop-off points, unanswered queries and coverage gaps, and feeding them back into the next prompt version Clientfacing performance reporting, built as a repeatable pipeline rather than hand-crafted each cycle.

Knowledge base quality: identifying what the agent could not answer and closing the gap

Essential

Demonstrable advanced prompt engineering - you have written and maintained complex production system prompts, not just used a chat interface. You can explain a behavioural bug you diagnosed and how you fixed it

Strong Python and SQL; comfortable with messy JSON and API data

Sound evaluation instinct you know the difference between a metric that moves and a metric that matters, and you are sceptical of your own judge scores

Clear written English for a non-technical client audience

Real attention to detail; you notice when a number or an output looks wrong

Skills and experience

Desirable

Experience evaluating LLM or conversational agent output at scale

Voice, kiosk or social channel agents (TTS constraints, character limits)

Neo4j or other graph databases

Node.js scripting for document generation

RAG and knowledge base design

Location
L24 9HJ

Additional information

***Ringfenced to Level 5 & Level 6 LJMU students only*** 

***Applications must be submitted using the application form linked below. Please upload it instead of a CV. Any other method will not be accepted*** 

Download application form  

Applications close at 11:59pm on Sunday 27th of September 2026 

Apply for job