Here is a statement that sounds convincing, yet is largely false: the more you work with AI, the sharper your judgment becomes.
The logic appears compelling. If you review enough outputs, study enough strong examples, catch enough errors, your critical judgment will naturally improve.
If you spend a year working closely with a capable model, you should emerge with better judgment than when you began.
Unfortunately, sound judgment develops under specific, well-documented conditions, and the ways people commonly use AI undermine those conditions before anyone realizes what has been lost.
In the latest edition of The Outlearn Advantage on LinkedIn, “Building Judgment in the Age of AI,” I argued that judgment is not built by reviewing answers. It’s built by forming your own point of view, testing your reasoning against reality, and refining your mental model through feedback.
This article unpacks the science behind that claim: what the research tells us about how skilled intuition develops, and why passive AI use short-circuits the very learning loop expertise depends on.
Before diving into the evidence, I want to clarify a crucial distinction: reviewing and deciding are different skills. Reviewing asks, “Does this output look right?” Deciding asks, “What do I believe is right before anyone or anything gives me the answer?”
Both are valuable but reviewing depends on the judgment you have already built.
The eye that catches AI’s mistakes has to be trained somewhere.
The Two Conditions For Skilled Intuition
For decades, two of the most prominent figures in decision science found themselves at odds over a deceptively simple question: Can expert intuition be trusted?
Gary Klein, who studied firefighters, nurses, and military commanders, believed the answer was a resounding yes. He documented how seasoned professionals made fast, accurate decisions under intense pressure and argued that expert intuition was not only real but remarkably powerful.
Daniel Kahneman saw things differently. Having spent his career cataloging the systematic ways human judgment breaks down, he contended that intuition was frequently compromised by bias and overconfidence.
In 2009, they did something remarkable: they collaborated on a paper to pinpoint the exact source of their disagreement. The result, Conditions for Intuitive Expertise: A Failure to Disagree, remains one of the most illuminating pieces ever written on how human judgment is formed.
What they found was that skilled intuition is genuine, yet it emerges only when two essential conditions converge.
First, the environment needs to be consistent and predictable, offering reliable cues that are genuinely linked to specific outcomes. Second, the individual must have ample opportunity to recognize and learn from those cues over time, building their understanding through repeated practice and feedback.
An environment with valid cues but no feedback can’t foster true expertise, while feedback without reliable cues merely breeds confident superstition. To truly learn, you need a predictable environment paired with consistent feedback on your decisions.
Anders Ericsson’s landmark research on deliberate practice, the work behind the oft-misquoted 10,000 rule, outlines four key requirements for truly effective practice: a clear goal, an independent attempt, immediate feedback, and error correction before the next repetition.
Keep that list in mind since we will return to it as we explore which of those conditions passive AI use eliminates.
Intuition can be a remarkably dependable guide, but only when the conditions that build it are in place. Without those conditions, confidence can become dangerously uncalibrated, creating the illusion of expertise without the capability to back it up.
Kind Environments & Wicked Ones
Robin Hogarth, a decision scientist, gave these conditions a more memorable frame. He divided learning environments into two types.
In a kind learning environment, feedback is clear, immediate, and directly tied to your decision. You take a position, and reality tells you quickly and unambiguously whether you were right.
Chess is kind: a poor move weakens your position within a few turns.
Surgery is relatively kind: the consequences are visible on the table.
Weather forecasting is kind: the forecast is confirmed or corrected by tomorrow.
In a wicked learning environment, feedback is delayed, noisy, absent, or disconnected from the decision that produced it. You take a position, but reality doesn’t answer quickly, cleanly, or directly.
Hiring is wicked: you may not know whether you made the right decision for months or years, and even then, the signal is filtered through culture, management, timing, incentives, and circumstance.
Strategy is wicked: outcomes unfold over years and are shaped by forces far beyond your original reasoning.
Investing is wicked: you can be right for the wrong reasons and wrong for the right ones, and the market will not tell you which.
Hogarth’s central insight is the one that matters here: in wicked environments, people can develop strong intuitions that are confidently wrong and never clearly discover it, because the feedback needed to correct them never arrives in a clean, timely, or reliable form. Experience accumulates. Confidence grows. Accuracy does not.
One additional distinction sharpens the picture, because not all feedback teaches equally. Research on feedback interventions, most notably Kluger and DeNisi’s landmark meta-analysis and Hattie and Timperley’s synthesis in education, shows that outcome feedback is a weak teacher, while process feedback is a stronger one. “It worked” or “it did not work” tells you far less than “this specific assumption failed.”
What builds judgment is not merely hearing from reality. It’s hearing from reality about your reasoning. You can call this the diagnosticity of feedback: how clearly the signal points back to the belief that produced the decision.
This is why years of experience in a wicked domain can produce someone who is more certain, but no more correct, than they were at the start. Why? Because the environment never graded their judgment cleanly enough for them to learn.
What AI Does To The Feedback Loop
Now the part that should concern anyone building a career alongside these tools.
AI inserts itself directly between your judgment and its consequences. And if used passively, it can degrade kind learning environments into wicked ones. AI doesn’t eliminate feedback. And it doesn’t make an environment wicked simply by being present.
What it does do is weaken the connection between your judgment and the outcome unless you deliberately preserve that connection. The same tool that can sever your feedback loop can, when used in a different sequence, strengthen it.
In the normal learning loop, you form a judgment by first forming a recommendation, diagnosis, forecast, or strategic position. Reality responds, and you compare what actually happened with what you expected. You update your mental model, then you do it again. That’s the kind-environment loop. It’s how calibration is built: one prediction, one correction, one refinement at a time.
Now insert AI, used the default way.
Instead of forming the initial judgment yourself, you ask the model for the recommendation. You review it, refine it, approve it, and send the work forward.
Eventually, the outcome arrives, but whose judgment did reality just grade?
You didn’t put your own point of view on record before the model spoke. You can’t trace a good outcome to your reasoning because the reasoning began with the model. You can’t trace a bad outcome to a specific belief you held because you never committed to one.
And notice what happened to diagnosticity. Even when the outcome is clear, it’s only outcome feedback about the final work, not process feedback about your reasoning.
Passive AI use doesn’t merely weaken the feedback loop; it reallocates cognitive effort.
The effort that once went into forming a judgment (framing the problem, weighing trade-offs, making assumptions explicit, and taking a position) gets redirected toward evaluating an output.
That may feel similar, but it’s not the same cognitive rep. The work that once built your decision-making machinery is replaced by work that exercises a narrower skill: reviewing, editing, and approving. And because both activities feel like “working on the problem,” the shift is easy to miss.
Run the Ericsson checklist against this workflow.
Clear goal? Usually.
Self-generated attempt? Gone. The model generated it.
Immediate feedback on your attempt? Gone. There was no attempt on your part to grade.
Two of the core conditions for deliberate practice are removed by default, silently, at scale. A year of “working closely with AI” under these conditions can produce what wicked environments often produce: more confidence and no more accuracy. The illusion of a sharpened eye without the deeper sharpening.
Why Reviewing Feels Like Learning
The reason this is so easy to miss is that reviewing AI output feels productive. You are engaged. You are seeing how a strong answer is constructed. You occasionally catch an error.
Research on worked examples from Sweller’s cognitive load work and Renkl’s work on example-based learning shows that studying high-quality solutions genuinely can build knowledge, especially for novices entering a domain. Reviewing is not useless, but the learning it produces has a ceiling.
Across learning research, the hierarchy is clear: passive review is weak, comparative evaluation is stronger, and generative practice, where you are producing the answer yourself, is the strongest.
And in professional judgment, generation is often the bottleneck. Nobody pays a premium for someone who can nod at a good strategy once it appears. They pay for someone who can produce sound judgment when no answer is on the screen yet.
The illusion of competence helps explain why reviewing AI output can feel like a higher-order skill. Robert Bjork’s work shows that fluency, the ease with which information is processed, often masquerades as understanding. When you read a polished AI response, and its logic feels clear, coherent, and easy to follow, your brain treats that smoothness as evidence of mastery. However, recognition is not generation because following a line of thought is not the same as forming one.
Studying learners working alongside AI tools, Fan and Gašević documented what they called metacognitive laziness: people offloaded the monitoring and regulation of their own thinking to the tool, and their metacognitive engagement declined even as their immediate output improved. Performance up. Self-regulation down.
Bjork explains why the review session feels like growth. Fan and Gašević measured what is actually happening underneath it. When you choose to review rather than decide, three essential elements are lost.
There are no stakes. You committed to nothing, so being wrong costs little and teaches little. The discomfort of discovering where your own reasoning was flawed never arrives.
There is limited ownership. The reasoning largely belongs to the model, so your own reasoning is not fully tested or corrected.
There is no clear error attribution. When an output turns out wrong, you can’t trace the failure to a specific belief you held because you never put one on record. Without a belief to attribute the failure to, there is nothing to update. In other words, the lesson has nowhere to land.
Review hands you the feeling of pattern recognition without the mechanism that builds it. This is how someone can spend a year immersed in AI and end it exactly as sharp as they began, because their confidence rose while their accuracy did not.
The Fix: Commit Before You See the Answer
The solution lies in shifting the order of operations. You need to form your own judgment before the model decides for you.
Commit your decision on paper (your prediction, recommendation, diagnosis, or position) with enough precision that it could be proven wrong, Now there are stakes, because you have taken a stand. There is ownership, because the reasoning originates with you. And there is something for the outcome to anchor to, because your belief has been stated before you reach for the tool.
Four separate lines of research converge on why this works.
First is the generation effect, one of the most replicated findings in memory research: we understand and retain what we produce ourselves far better than what we merely read.
Second is retrieval practice. Grounded in Roediger and Karpicke’s seminal research on the “testing effect,” this concept demonstrates that actively retrieving information from memory does more than just boost retention; it sharpens “discrimination,” the ability to select the right knowledge under pressure.
Third, is Robert Bjork’s concept of desirable difficulties, the idea that conditions which slow us down in the moment strengthen learning over time. Committing to your own reasoning before consulting AI is slower, more effortful, and less comfortable than simply prompting first. However, that friction isn’t the cost of the practice; it’s the practice. The discomfort you feel is precisely the condition under which durable judgment is forged. A frictionless workflow, by contrast, offers the feeling of productivity while quietly building very little.
Fourth, stems from Philip Tetlock’s decades of forecasting research, which revealed that people who successfully refine their judgment over time share a single habit: they make specific, written predictions and score them against reality. Recording a falsifiable forecast is what transforms raw experience into precise calibration. It’s also the only way to truly understand your own decision-making thresholds, revealing not just how often you’re right, but how well your confidence aligns with your actual success rate, and where you’re most prone to false alarms or missed opportunities.
You can’t calibrate your confidence with correctness without first explicitly declaring your level of certainty.
The Practice: The 3-line Judgment Protocol
A decision journal is far more than just another productivity habit. It is the vital mechanism that transforms a wicked environment into a kind one.
For the choices that actually matter, pause before you open the tool and write three lines.
The judgment. What do you predict will happen, and what do you think should be done about it? Be specific enough that you could be proven wrong. A judgment that carries no risk of being incorrect teaches you nothing.
The confidence. Assign it a number between 0 and 100. It will feel unnatural at first, that’s the point. Quantifying the feeling is what transforms a vague gut instinct into a measurable track record. You can’t calibrate what you’ve never bothered to measure.
The key assumption. What assumptions must hold for your conclusion to stand? Identify them now, they’ll be the very premises you scrutinize later. Being right for the wrong reasons is one of the most costly lessons for professionals. It validates a flawed model and all but ensures you’ll rely on it again when the stakes are higher.
Finally, when reality reports back: compare, attribute, update. Did the prediction hold? Did the outcome match the confidence? Which assumption stood, and which quietly failed?
That update is the moment judgment is actually built.
What It Looks Like When It Runs
A product leader is setting pricing for next quarter. The metric she’s watching is churn, the percentage of customers who cancel and leave. Raise prices too aggressively, and churn spikes. This means the revenue gained is wiped out by customers walking out the door.
Without the protocol: She asks an AI for a pricing strategy, reviews the logical output, refines it, and signs off. The price goes live. Three months later, churn has shifted, yet she has learned nothing she can trace back to her own thinking because she never committed to her own beliefs on the record. The outcome arrived with nowhere to land.
With the protocol: Before writing a single prompt, she notes, “An 8% price increase will raise churn by less than 2%. Confidence: 70%. Key assumption: high switching costs make our customers price-insensitive because leaving is both costly and disruptive, most will absorb the hike rather than endure the friction of migrating.” She then queries the model, stress-tests its reasoning against her own, and ships.
Three months later, churn has climbed by 5%, more than double her forecast. Customers are leaving at a pace her models deemed impossible. The outcome itself hasn’t changed, but its implications have. Her key assumption about high switching costs has been exposed: leaving was far less painful than she believed. Now, she is forced to recalibrate her customer model; her 70% confidence level is marked as a definitive miss. A result she might once have dismissed as mere noise now lands as a clear misjudgment that needs recalibration.
One logged decision. One corrected mental model. That is the engine of judgment running, and it ran on just three sentences.
What You’re Actually Building
Here is the economic version of everything above.
AI has turned execution into a commodity. Drafts, analyses, summaries, and first-pass strategies are now cheap, fast, and universally accessible, compressing the competitive advantage they once provided to nearly zero. In an era of abundant output, the truly scarce resource shifts to the one thing AI can’t replicate: calibrated judgment under uncertainty.
The market is already pricing this in: generation is cheap, but judgment is becoming more valuable. The person whose instincts have been tempered by reality is becoming more valuable, not less. And that person is not built by reviewing answers.
They are built by making, owning, testing, and deliberately refining their judgments. And through intentional practice, they foster a kind learning environment, actively resisting the default workflows that quietly make it wicked.
This is where the “reviewer’s eye” actually comes from. The ability to catch what AI misses, and to know exactly when work is ready to execute. It’s not a standalone skill you can develop through sheer repetition but a visible expression of deeper judgment, forged one deliberate decision at a time.
Develop your judgment, and your eye will naturally follow. Neglect it, and whatever confidence you project is merely borrowed. So before you close this, one question worth sitting with:
On the decisions that actually mattered this past month, how many did you commit to before the model did the work, and how many can you still trace back to a belief that was yours?
If the honest answer is “few,” don’t mistake this for a failure of discipline. It’s simply a wicked environment behaving exactly as expected. The remedy is structural, and it begins with writing just three statements before you input a single prompt.
Your judgment. Your confidence level. Your assumptions.
Form your own opinion first. Let the machine challenge it second. Then measure both against reality because, in the end, the score is the real teacher.
You can’t outperform what you haven’t outlearned and you can’t outlearn a judgment you never formed, a decision you never owned, or a consequence you never connected back to your own thinking.
Charles Good helps organizations use AI to build human capability, not just improve productivity. As President of the Institute for Management Studies, he reaches more than 20,000 professionals annually. His Outlearn Loop framework, grounded in behavioral and learning science, helps leaders design AI integration that strengthens judgment, adaptability, and performance. He writes The AI Capability Playbook and The Performance Playbook on Substack and hosts The Good Leadership Podcast (Apple / Spotify / YouTube).
References
Kahneman & Klein — the two conditions. Kahneman, D., & Klein, G. (2009). Conditions for intuitive expertise: A failure to disagree. American Psychologist, 64(6), 515–526.
Hogarth — kind vs. wicked learning environments. Hogarth, R. M. (2001). Educating Intuition. University of Chicago Press. Hogarth, R. M., Lejarraga, T., & Soyer, E. (2015). The two settings of kind and wicked learning environments. Current Directions in Psychological Science, 24(5), 379–385.
Ericsson — deliberate practice. Ericsson, K. A., Krampe, R. T., & Tesch-Römer, C. (1993). The role of deliberate practice in the acquisition of expert performance. Psychological Review, 100(3), 363–406.
Kluger & DeNisi — feedback intervention theory. Kluger, A. N., & DeNisi, A. (1996). The effects of feedback interventions on performance: A historical review, a meta-analysis, and a preliminary feedback intervention theory. Psychological Bulletin, 119(2), 254–284.
Hattie & Timperley — the power of feedback. Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81–112.
Sweller & Renkl — worked examples. Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285. Renkl, A. (2014). Toward an instructionally oriented theory of example-based learning. Cognitive Science, 38(1), 1–37.
Bjork — fluency, illusions of competence, and desirable difficulties. Bjork, R. A., Dunlosky, J., & Kornell, N. (2013). Self-regulated learning: Beliefs, techniques, and illusions. Annual Review of Psychology, 64, 417–444. Bjork, E. L., & Bjork, R. A. (2011). Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. In Psychology and the Real World (pp. 56–64). Worth Publishers.
Fan & Gašević — metacognitive laziness. Fan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., Shen, Y., Li, X., & Gašević, D. (2024). Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology.
Slamecka & Graf — the generation effect. Slamecka, N. J., & Graf, P. (1978). The generation effect: Delineation of a phenomenon. Journal of Experimental Psychology: Human Learning and Memory, 4(6), 592–604.
Roediger & Karpicke — retrieval practice / the testing effect. Roediger, H. L., III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255.
Tetlock — written predictions and calibration. Tetlock, P. E., & Gardner, D. (2015). Superforecasting: The Art and Science of Prediction. Crown.


The distinction between reviewing and deciding is a strong design test for AI-assisted research. If the model speaks first, the human's independent hypothesis may never exist, so there is nothing clean to compare against later evidence. A lightweight version could be three fields before prompting: expected finding, key assumption, and disconfirming evidence. After the model responds, record what changed and why. That creates a learning trail rather than a polished answer. I wonder which of those fields people actually sustain after the novelty wears off.
I really like your “commit before you see the answer” principle, and it is something I am trying to apply regularly.
I would add one nuance: “being” certain or confident is probably better described as "feeling" certain or confident. Which manager or leader does not behave very confident these days (without being certain)? Whenever I encounter someone who is very certain, the interesting question for me is: what evidence would actually justify that "feeling"?
If I write down my judgement before consulting AI, I can also state in advance what data would show that I was right, only partly right, or wrong. That creates a useful reference point later and makes it harder to rewrite my own history through hindsight bias, especially after I have changed my mind and then prompt that revised view back into the system.
AI could then play very different roles. It can reinforce confirmation bias if I use it to validate my position. But I can also use it deliberately as an adversarial tool: ask it to take opposing perspectives, war-game objections, or build the strongest case against my judgement. I would still treat it as a black box, although humans often are too.
The limitation, of course, is that this is still simulated opposition. It may prepare me for some of the headwind I will face, but not for everything reality will produce.
So my questions:
1. If we use AI to challenge a pre-committed judgement, how do we know that we are improving our judgement rather than simply becoming better at defending the same judgement against anticipated objections?
2. What kind and amount of real-world evidence would have to contradict us before we should say that our original confidence was not merely too high, but actually unjustified?