Applied AI · Trust & Safety · AI Governance
Mateen
Rahmat
Staff AI Prompt Engineer at LinkedIn.
He designs AI for large-scale decision-making, building the evaluation frameworks, governance standards and human oversight that make its performance measurable and its decisions accurate and reliable.
This is Mateen's AI. It answers from this page and from notes he wrote and it shows where each answer came from. When it does not know something about him, it says that plainly.
About

Mateen works where applied AI meets Trust & Safety and large-scale platform operations. He defines how AI systems are built, evaluated and governed when they run at scale under human oversight.
Today he sets the technical standard for prompt architecture and the evaluation frameworks that decide production readiness.
He has moved advanced AI research into production and led a global team of 30 annotators.
The through-line of his work is making AI performance measurable and grounding every automation decision in evidence.
- 30+
- annotators led
- 5
- years in AI Trust and Safety
Areas of work
01
Evaluation frameworks
The frameworks that decide whether a generative AI system is ready for production. They rest on representative sampling and on measuring precision and recall separately, because the cost of the two errors is rarely the same. Regression testing and drift monitoring keep it honest after launch. The question underneath is whether the system can be trusted with a decision and how anyone would know.
02
AI governance and operating models
Governance standards across the whole AI lifecycle, from versioning and change control to approval pathways, responsible-AI practice in production and clear decision rights across product, engineering, data science and operations.
03
Prompt architecture
The technical standard for prompt design, translating complex requirements into precise, testable model behaviour with zero-shot, few-shot, hybrid and retrieval-augmented approaches. A prompt, in his view, is a specification and a specification has to be measured.
04
Data and annotation operations
He has run the AI operations behind generative-AI models, leading a team of more than 30 annotators and overseeing training pipelines from data selection and preprocessing to fine-tuning. Quality controls kept the labelled data reliable.
05
Agentic AI under human oversight
AI agents that act within written guidelines, hand uncertain or high-stakes decisions to a person and are measured continuously. He defines how such agents operate and scale in high-stakes settings, including where a person sits in the loop and how their corrections feed the next evaluation.
Perspectives
Perspective 01
On metrics
Every AI project should begin with the decision it is meant to support and a written definition of what a correct one looks like. Once that definition exists, everything else becomes measurable. The team knows which examples to sample, which errors are expensive and what precision and recall the system has to reach before anyone trusts it. The model is then a candidate to be judged against that standard.
Perspective 02
On ground truth
A ground-truth set should be small, built from real and representative examples and re-run after every prompt change. That is what tells you how the system behaves today. The discipline is in the re-running. Prompts evolve, model versions change and requirements move with the product. Regression testing after each of those changes catches the quiet failures that would otherwise go unseen. The set has to stay honest, stay current and stay small enough that nobody is tempted to skip it.
Perspective 03
On oversight
Human oversight belongs in the design from the first day. It decides which decisions the model may make on its own and which ones it must hand to a person. It also decides how confidence is expressed and what happens to a correction once someone makes one. At scale the design question is where people sit in the loop and whether the loop is fast enough to matter. The people in that loop are the most valuable sensor the system has and their attention should go where the model is weakest.
Model training
Training a generative-AI model well is a loop. Every round of evaluation shows what to collect, label and fine-tune next.
Hover or tap a step to see what happens there.
01
Data selection
Which real examples the model must learn from, sampled to represent the decisions it will actually face.
02
Preprocessing
Cleaning and structuring the data so labels attach to exactly what the model will see in production.
03
Fine-tuning
Turning the labelled data into model behaviour, tuned for accuracy and scalability.
04
Evaluation
Accuracy, efficiency and bias, measured against held-out ground truth before anything ships.
05
Retraining
Evaluation metrics decide the next retraining cycle and the fine-tuning strategy.
Agentic AI
An agent is a model that takes actions. The system built around it is what makes it trustworthy. That system decides which actions the agent can take on its own, when it hands over to a person and how every outcome is measured. This diagram shows the general pattern as it is practised across the industry.
Hover or tap a step to see what happens there.
01
Requirements become prompt architecture
Written requirements are translated into precise, testable model behaviour, versioned and approved as a specification.
02
The agent and its tools
The model, retrieval, classifiers and tools together produce a decision and a confidence in it.
03
Take the action, or hand it to a person
Confident, low-stakes decisions are actioned by the agent. Uncertain or high-stakes decisions go to a person and that person's correction becomes new data for the system.
04
Every decision feeds back into the system
Every decision is monitored and evaluated. What the evaluation finds flows back into the requirements and the prompts, within the same governance standards of versioning, approvals and decision rights.
Building
Mateen is building a set of agents on one architecture. The first handles Irish admin, where not knowing a rule quietly costs people money. It finds the tax credits nobody claimed and works out what a person can really borrow, then prepares the paperwork for them to submit.
Tap any point to see what it is.
01
One architecture, many domains
Every agent runs the same six stages, so the only thing that changes between them is the domain. It reads the person's documents and looks up the rules, then does the arithmetic in code so that no figure depends on a model remembering it.
02
The person stays in the loop
Reading, computing and drafting run on their own. Anything that submits, sends, books or pays waits for the person. What they change is kept as data.
03
Every figure has a source
Rules live in tables where each figure carries its source and the date it was checked. A figure without one fails the tests, so nothing rests on a model's recollection.
04
Everything is measured
A fixed set of worked examples runs against the live site and scores each answer for whether the decision was right and whether every figure can be traced back. The score gets published, including the runs that went badly.
Selected work
01
ireland.mateenrahmat.com, 2026
Ireland agent stack
A stack of agents for Irish admin, covering mortgages, tax back, renting, immigration and health. One architecture runs all of them. The agent reads the person's documents, looks up a rule table where every figure carries its source and the date it was checked, does the arithmetic in code and drafts what needs doing. Reading, computing and drafting run on their own. Anything that submits, sends, books or pays waits for the person in an approvals inbox. What they change is kept as labelled data. Every run ends with one structured decision and is traced beside the conversation.
ResultA working demonstration of agents under human oversight, measured by a fixed evaluation that plays every worked example against production and scores whether the decision was right and whether every figure can be traced.
- Agents
- Evaluation
- Human oversight
- Cloudflare Workers
- Claude
02
Personal site, 2026
Mateen's AI on this site
An assistant that answers only from written sources, the CV as numbered facts, a notes file and nothing else about him. Every claim carries a citation that jumps to its source, private topics are declined and a fixed eval checks grounded, general and declined behaviour before every deploy.
ResultEvery statement about Mateen on this page can be traced to a source and the assistant says when it cannot.
- Claude
- Cloudflare Workers
- Evaluation
03
University College Dublin, MSc project
Explainable financial forecasting platform
Planned and built a forecasting platform in which every prediction comes with an explanation a non-specialist can act on, as part of the MSc in Data Science and AI.
ResultA forecasting system judged on its accuracy and on whether its reasoning could be inspected.
- Explainability
- Forecasting
- Python
Experience
May 2026 to Present
Dublin
Staff AI Prompt Engineer
LinkedIn·He owns how generative AI is designed, validated and governed for high-stakes decisions at scale.
- 01Shapes how AI agents are adopted and scaled for high-stakes decisions, working with senior leadership on priorities.
- 02Leads generative-AI oversight from design through validation to production.
- 03Builds the evaluation frameworks that decide whether generative AI is ready for production.
- 04Sets governance standards for AI in production, covering versioning, approvals and decision rights.
- 05Sets the technical standard for prompt design, turning complex requirements into precise, testable model behaviour.
Apr 2025 to Apr 2026
Dublin
Senior GenAI Project Manager
Meta·He took generative-AI research into production and ran the data operation behind it.
- 01Oversaw training pipelines for generative-AI models end to end.
- 02Led and mentored a team of more than 30 data annotators producing high-quality labelled datasets.
- 03Defined and tracked the metrics used to judge model accuracy, efficiency and bias.
- 04Partnered with engineering, research and product teams on generative-AI initiatives.
- 05Managed several AI and machine-learning initiatives at once, from concept to deployment.
- 06Drove experimentation with LLM applications and AI-powered productivity tools.
Sep 2021 to Mar 2025
Dublin
Risk Prevention Specialist
Meta·He did data-driven risk prevention for three and a half years, which is where the habit of measuring something before trusting it started.
- 01Analysed large-scale datasets with SQL and Python to find the patterns behind product improvements.
- 02Delivered data-backed recommendations to leadership and cross-functional partners.
- 03Planned and implemented AI-assisted monitoring that improved detection accuracy and reduced false positives.
- 04Managed end-to-end projects across AI, data and risk.
Education
2024 to 2026
MSc, Data Science and AI
University College Dublin
2017 to 2020
BSc, Information Technology (Data Science)
American University of AfghanistanCum Laude
Languages
English (fluent) and Dari (native)
Skills
AI
- Prompt engineering (zero-shot, few-shot, hybrid, RAG)
- AI agents & assistants
- Human-centric AI
- AI ethics
- Fine-tuning data curation & annotation
Evaluation & Governance
- Ground-truth dataset design
- Sampling methodology
- Precision and recall analysis
- Error taxonomies
- Regression testing
- Drift monitoring
- Versioning & approval pathways
Product & Program
- Cross-functional leadership
- Stakeholder partnership
- Leading global teams
- Operational strategy
Programming
- Python
- Java
- SQL
- R
- MATLAB
Data & Analytics
- Statistical analysis
- Data cleaning & modelling
- Data ethics
- Tableau
- Power BI
- SAP BusinessObjects
Trust & Safety
- Platform integrity
- Crisis management
- Risk prevention & anti-abuse
- Content review & enforcement
Contact
Happy to connect.