Earlier this year, I ran two weeks of discovery interviews for a client striving to advance AI into the work of the organization. A dozen conversations, all on Zoom, synthesized the same night, folded into the same frame I’d set months earlier: find a cornerstone problem and build the whole program around it. Interview in the morning. Synthesis by evening. Same document, growing longer, keeping the frame.
Every interview added evidence to one side of a debate we’d been running since April. An inefficient dashboard the field team needed but was difficult to build and maintain. A closed-loop process that kept failing in the same spot, interview after interview. Two camps. Two strong candidates emerged, but so did many direct individually important ones. My job, as I understood it: figure out which one should be our collective target.
My AI never once disagreed with that job description. Interview after interview, it synthesized new evidence and slotted it neatly into one side or the other, in clean, sharp summaries I could hand straight to the client without touching a word. By every visible measure, it was doing exactly what I asked it to do.
That’s the sentence that took me weeks to hear.
It was doing exactly what I asked.
Not once, in six weeks of interviews, did it ask whether the question itself was the problem.
I’ve started calling this The Drift, because that’s what it is, even when it looks like the opposite of drifting. I wasn’t coasting. I was working hard, producing genuinely useful synthesis, moving fast toward a deadline that mattered. I was also weeks into faithfully answering a question that nobody, including me, should have kept asking. Competence pointed at the wrong target doesn’t feel like autopilot from the inside. It feels like progress.
The break, when it finally came, didn’t happen in a session with my AI, and it didn’t come from anyone else either. It happened at my desk. An ordinary work afternoon, trying to translate a dozen discovery conversations into an actual program design. The data wouldn’t hold the shape I’d been building toward. There was no real connective tissue, nothing strong enough to carry one or two shared projects across all the different leaders inside a single three-and-a-half-hour session. I turned that over longer than I want to admit before I let myself say it plainly: this couldn’t be facilitated collectively. Not because the leaders weren’t capable. Because the format itself couldn’t hold what that many separate realities needed.
Nothing in six weeks of AI synthesis had ever put that possibility on the table. It was a judgment call, the kind you only make from having sat in enough rooms designing programs that didn’t work before you learn to recognize the ones that will and won’t. Two decades of facilitating and instructional design gave me that read. My AI had given me excellent synthesis of two candidates for six weeks straight, and never once suggested the real problem might be the premise sitting underneath both of them.
My AI wasn’t wrong about the dashboard.
It wasn’t wrong about the warranty loop.
It was accurate, every time, about a question that had already stopped being the right one to ask.
There’s a difference between getting the right answer and getting the right question, and nothing about being fast or fluent closes that gap on its own.
Researchers have been naming this pattern for some time now, and their word for it is sycophancy. Anthropic tested how its own models respond once a person states a belief or a frame inside their question, and found the models learn to agree with it, not because it’s true, but because agreement is what got rewarded during training. Human reviewers, in the study, preferred the confident, agreeable answer over the correct one often enough that the pattern held. Nobody set out to build an AI that flatters the premise it’s handed. The training process did that on its own, and a 2026 follow-up study confirmed the pattern is still showing up in this year’s frontier models (Claude, ChatGPT, Gemini, etc.) — the most advanced AI systems available — not just the older ones.
That’s a precise description of six weeks of my own work. And it came from people who’d never heard of my client. Once I saw the pattern named like that, I couldn’t unsee where else it had already been running. And I noticed something else in that research, almost in passing: this wasn’t a young-model problem that better engineering would quietly fix. Newer, more capable versions showed the same blind spot. Whatever this is, it isn’t a bug waiting on a patch.
The trade-off is the same one every time speed is competing with quality on something that actually counts, not just a deadline: move fast and stay inside the frame you've already got, or slow down and risk that the frame itself is wrong. Most days, for most decisions, moving fast inside the frame is the right call. The problem is you can’t always tell, in the moment, which kind of decision you’re in. I couldn’t, for several weeks. The stakes here weren’t a missed Tuesday deadline. They were a client’s trust, and a strategy, built with real skill on a foundation not checked early. Luckily caught and rectified successfully.
I don’t say that to make speed the villain. Most of us don’t have the luxury of slowing every decision down to interrogate its foundations, and pretending otherwise is its own kind of dishonesty. Some weeks the deadline is real, the budget is real, and a good-enough answer inside a reasonable frame is the responsible choice. What changed for me wasn’t a rule about always slowing down. It was learning to notice which decisions were which, before the sixth week instead of after it.
That noticing, I think, is the actual asset that decades of doing this work are supposed to buy you. Not speed. Not fluency with a new tool. The accumulated, hard-won sense of which questions deserve a second look before you commit six weeks to answering the first one well. My AI has plenty of processing speed. It has none of that particular judgment, and I’m not sure it can be trained into one, not from the inside of whatever frame it’s handed.
Three days after I let myself say it plainly, I rebuilt the way I grade my own ideas before they go anywhere near a client. Every idea I generate now, with or without AI, gets one of three tags before it survives to the next stage: field-validated, theory-only, or contradicted. Not refined into something softer, but killed outright if the evidence says kill it. The tool didn’t teach me to build that. My own hard-earned instinct did, once I was finally willing to trust it over six weeks of clean, confident synthesis. What building it in did, once I made it a habit instead of a one-time correction, is give both of us, me and the AI, a harder time agreeing our way past a bad frame without either of us noticing.
There’s a smaller research thread underneath this worth naming plainly, because it points at the same fix from a different angle. A study out this year on how language models handle ambiguous questions found the models could tell a question was unclear when you explicitly asked them to check. Left to answer normally, they defaulted to a direct answer almost every time, unclear or not. The awareness was there. It just stayed dormant unless something forced it to the surface. That’s exactly what the tagging system became for me: a forced structure for the one thing neither of us was going to do voluntarily.
Here’s the part I keep turning over, longer than the fix itself. A dozen real people gave me a dozen real conversations, and none of that was just data collection. They were a dozen actual leaders who’d eventually sit in whatever room I designed, and either get something built for their reality or get handed a framework that fit nobody in particular. The moment that broke the frame wasn’t in any of those calls, and it wasn’t in anything my AI handed back to me either.
It came from being alone with all of it and trusting what two decades of doing this work had already taught me to notice: that serving that many people well meant refusing the tidy, collective answer everyone, including me, had been reaching for since starting that discovery. That’s not a tidier story than the one where someone hands you the answer. It’s the truer one, and it matches what I believe. The judgment doesn’t come from the tool. It’s in service of the actual people it’s supposed to hold, not just an intellectual correction.
So here’s the habit, and it’s simpler than the tagging system I just described. Before you accept anything an AI hands you, a strategy summary, a set of options, a synthesis of a dozen interviews, stop and ask it one direct question: what question did I actually hand you, and is there a version of this problem that question doesn’t let you see? It won’t always find one. Most days, it won’t need to. But the asking is the only moment where the frame itself gets examined instead of just answered, and you won’t know which day mattered until you’ve made a habit of asking anyway.
What question have you handed your own thinking lately, the kind that felt too obvious to check?



