The role of a Shopify Shop App Agent Annotator at Mercor
About the role
Shopify’s Shop App Agent is the AI assistant buyers talk to inside the Shop app — product discovery and search, orders, shipping and tracking, account questions, and Shop Cash.
When a shopper is unhappy with one of its answers, they tap thumbs-down — and that’s all they leave. No comment, no reason, no survey. Your job is to work out why, from the transcript alone.
What you’ll do
- Read a full shopper ↔ agent conversation, including the tool calls the agent made and what came back
- Find the specific turn that received the thumbs-down
- Classify it with one primary tag — the domain of the complaint: a disliked recommendation, a failed order lookup, an action the agent couldn’t take, a rejection of AI itself, or a genuinely broken response — and one specific secondary tag within it
- Write a short comment: what the shopper wanted, what the agent did or failed to do, and why those tags fit
- Commitment: ~5 hours/week
What makes someone good at this
The hardest part is not guessing. Plenty of thumbs-downs have no visible cause — the agent did nothing wrong, or the shopper simply didn’t like the answer. There is an explicit unknown label for exactly that, and using it honestly matters more than producing a confident-sounding reason. We would rather record “we don’t know why” than invent a cause that sends the wrong signal to the model.
Qualifications
Required:
- Rule discipline — you can apply a fixed taxonomy the same way across hundreds of conversations, and you notice when a case sits between two labels rather than forcing it.
- Comfort reading structured traces — tool calls, their inputs, and their raw outputs, so you can tell what the agent actually did from what it merely claimed.
- Restraint under ambiguity — you are willing to label something
unknown rather than reach for a plausible-sounding cause.
- Clear, concise written English — enough to explain in two or three sentences why a conversation got the labels you gave it.
Strongly preferred:
- Consumer e-commerce fluency — orders, tracking, returns, refunds, and a working sense of where a merchant’s responsibility ends and the platform’s begins.
- Prior annotation, labelling, or model-evaluation work against a defined rubric or taxonomy.
Preferred (nice to have — we’ll ramp you on the specifics):
- Hands-on experience with the Shop app or other AI shopping assistants as a shopper.
- Familiarity with customer-support operations and escalation paths.
What this is not
This is not a customer-support role and not an engineering role. You are not fixing the agent and not replying to shoppers — you are diagnosing why a real shopper was unhappy, precisely and repeatably, so the team can measure where the agent falls short.
We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.
Contract and Payment Terms
- You will be engaged as an independent contractor.
- This is a fully remote role that can be completed on your own schedule.
- Projects can be extended, shortened, or concluded early depending on needs and performance.
- Your work at Mercor will not involve access to confidential or proprietary information from any employer, client, or institution.
- Payments are weekly on Stripe or Wise based on services rendered.
- Please note: We are unable to support H1-B or STEM OPT candidates at this time.
About Mercor
Mercor partners with leading AI labs and enterprises to train frontier models using human expertise. You will work on projects that focus on training and enhancing AI systems. You will be paid competitively, collaborate with leading researchers, and help shape the next generation of AI systems in your area of expertise.
To learn more about Mercor visit their website
Please let Mercor know you found this job position on Remote Career Africa as a way to support us to continue providing you with quality remote jobs
Always read and understand the full job requirements before you apply