Critical Judgment as a Core Competency in the AI Age: Prompt Engineering for Quality Assurance
Generative AI makes first drafts nearly free. But evaluating the results? That still takes expertise and critical thinking. Prompt engineering must adapt.
Overview
In ‘The First Draft Is Free. The Final Judgment Is Not.’, Ara argues that AI lowers the cost of first drafts, but human judgment still carries the final evaluation. She cites data: the Stanford AI Index 2025 shows LLM inference is 280 times cheaper in 18 months. OpenAI’s analysis of 1.5 million ChatGPT conversations lists writing and planning as the main uses. Production is becoming cheap.
But the BCG experiment reveals a flip side. On clearly defined tasks, AI boosts productivity. On complex, real-world problems it falls behind – and the output still looks convincing. The METR study found similar: developers using AI were slower, but thought they were faster. So human evaluation remains essential. We need to strengthen it.
Prompt engineering shouldn’t just generate. It should craft prompts that encourage reflection, verification, and quality checks. Here I analyze a prompt built from the article’s requirements, with built-in evaluation steps.
Prompt Analysis
The article doesn’t include full prompts, but mentions use cases like asking for 20 headlines or three positioning directions. Let’s build a prompt for headline creation, adding an evaluation step. That turns the AI into an analysis tool, not just a generator.
The Prompt
As an experienced marketing editor for a B2B technology company focused on AI applications, develop 20 headlines for a blog article on the topic 'Productivity and Judgment in the Age of Generative AI.' The headlines should spark curiosity and clearly communicate that the article offers practical tips for quality assurance of AI-generated content. Avoid clickbait and exaggerated promises. After the list, briefly evaluate each headline for its suitability for the target audience (C-level executives and product managers) and recommend the three best with justification.
Components
Role/Persona: The role ‘experienced marketing editor’ sets quality expectations. Without it, the AI produces generic headlines.
Context: Context includes the company, subject, and audience. That narrows the output and boosts relevance.
Task: The task is straightforward: develop 20 headlines. That gives enough options to evaluate later. It fits standard content marketing workflows.
Output Format: The output must include a list, evaluation, and recommendation. This structure forces reflection, an early step toward judgment.
Constraints: Constraints like ‘avoid clickbait and exaggerated promises’ keep the output credible. They push the AI toward substantive phrasing.
These elements combine generation with evaluation. The prompt builds critical review into the process, which helps human reviewers.
Frequently Asked Questions
Why is human judgment so important despite AI?
AI gives plausible but frequently wrong results. The jagged frontier shows a model may excel at one task and fail at another. When output looks professional, errors go unnoticed without human judgment.
How can prompt engineering improve the quality of AI outputs?
Prompts should generate and also prompt reflection. Ask the AI to question its answer, consider alternatives, or critique its own suggestions. This turns it into a reviewing partner.
What role does domain knowledge play in evaluating AI results?
Domain knowledge helps spot fallacies and gaps. An expert sees when a strategy sounds logical but misses the real problem.
What does ‘First draft is free’ concretely mean for everyday work?
Draft texts, code, or analyses become nearly free. Value moves to revising, verifying, and deciding – tasks that need human judgment.
How do you know if an AI output is really good?
Don’t just check style and logic. Question assumptions, look for missing information, consider alternatives, and trace consequences. Ask the AI about its weaknesses – that involves it in review.
Can AI learn to develop judgment?
AI recognizes patterns and produces plausible answers. But genuine judgment, with situational understanding and values, stays human. Skillful prompting can make AI a tool for critical reflection.
Source
Based on this article.