MOLLKOM / AI & machine learning

Research Engineer, AI Evaluation & Agent Reliability

Turn AI quality from an impression into evidence for better decisions.

Riyadh, Saudi ArabiaHybridSeniorApplications open

The mission

Commerce outcomes require more than fluent answers. Design evaluations that distinguish good writing, correct information and completed action, helping the team release improvements with measurable effects.

What you’ll work on

  • Build benchmarks for catalog, content, conversations and tool-use tasks.
  • Design rubrics and human/automated evaluation with agreement, bias and coverage analysis.
  • Develop regression and adversarial tests for permissions and unsupported claims.
  • Analyze experiments and model comparisons, communicating actionable findings and statistical limits.

What you bring

  • Experience in ML evaluation, applied research or AI quality engineering.
  • Strong Python, data analysis and experiment design.
  • Understanding of LLM-as-judge limitations, data leakage and evaluation bias.
  • Ability to communicate results clearly and translate findings into engineering improvements.

SHOW YOUR THINKING

Let your work start the conversation.

Share an evaluation you designed and how it changed a release or model-selection decision.

Share public links or a short explanation without disclosing confidential information from previous work.

About Mollkom

Mollkom is an AI-enabled commerce operating environment. We connect product preparation, content, marketing and customer service with orders, inventory, point of sale and delivery, so the steps work from shared knowledge.

Meet the platform
Apply for this role