As predicted: this cycle’s summer analyst return-offer conversion is looking uneven. Wall Street Oasis threads describe some banks reportedly landing well below 70% — down from what’s typically a 70-90% range — and the explanation candidates say they’re hearing isn’t “the intern class was weak.” It’s “we don’t have the seats.” That’s headcount allocation, not performance, and it lines up with reporting that some banks are cutting incoming junior classes by as much as two-thirds as AI tools like Rogo absorb more of the modeling and deck-formatting grunt work.
Here’s the connection to hiring: as junior classes shrink, recruiting is shifting toward lateral hires — seasoned associates and VPs who already have the skillset to catch a model’s mistakes. Banks would rather pay up for someone who’s proven that judgment than bet on training it into a smaller junior class. Clients tell me that shift is exactly where this is showing up: in how technicals interviews are being run, and in how take-home models are being graded, for junior and lateral candidates alike.
The core issue: banks aren’t hiring for people who can use AI. They already have it. What they’re actually screening for is the person who can look at AI output and catch what’s wrong with it — the wrong EV, the footnote that doesn’t tie, the number that’s off before it hits a client. That’s the skill: not “can you produce an answer,” but “can you catch the mistake.” Those are two very different things, and interviewers know exactly which one they’re testing for. It matters this much because the tools themselves aren’t reliable enough to skip that check. Independent benchmarking of AI agents on real banking workflows found even the best-performing model failed nearly half its evaluation criteria, and human bankers rated none of the unsupervised output as client-ready. Rogo’s own rollout inside banks bears this out day to day — the reporting shows it’s mostly MDs and VPs plugging in prompts, while associates spend real hours quietly fixing what comes back: a wrong EV pulled from a linked source, an off date buried in a footnote, slide language that reads a little too smooth to survive a client review untouched. Someone still has to catch it. Leaning on AI to get through a case study demonstrates the wrong skill — and it’s turning into a real problem in two very different parts of the process.
Take-home models. You get hours alone with a laptop, and it feels like the lowest-risk moment to lean on a chatbot — nobody’s watching. But the exercise was never really about the file you submit. It’s a setup for the walkthrough that follows, where you have to defend your own model. If you didn’t build it, you can’t defend it. I’ve heard this play out from clients on the hiring side more than once: candidate submits a clean-looking model, gets asked one follow-up question about an assumption, and the whole thing falls apart. That’s not a bad interview — that’s the interview working exactly as designed.
Live technicals screens. Same story, just faster. Banks are done pretending this isn’t happening. Goldman’s campus team has explicitly banned outside sources, including ChatGPT, during interviews, and tools like HireVue and TestGorilla are now built to flag tab-switching and unusual response timing — TestGorilla’s own data puts the share of candidates who’ve attempted to cheat on an assessment at roughly one in six. So even setting aside whether your interviewer catches it, the software might too.
And your interviewer will catch it — often within one answer. The tell isn’t subtle to someone who’s built a hundred models by hand: phrasing that’s a little too smooth, a textbook explanation that doesn’t quite match how the number was actually derived, a pause the second the question veers off-script. What’s changed recently isn’t just the ability to spot it — it’s the patience for it. The same associates and VPs running your interview are the ones spending real hours on the desk cleaning up AI-generated first drafts: chasing down the wrong EV, the off footnote, the slide copy that reads a little too generic. When they see a candidate leaning on that same crutch, it doesn’t read as clever prep. It reads as someone trying to skip the part of the job that’s currently eating their week.
Bottom line. Build it cold first. Use AI to check yourself after, not to generate the first draft, and not to rehearse your talking points instead of learning the logic. The value of practice isn’t the finished output — it’s the several wrong versions you build getting there, because that’s where the instinct interviewers are testing for actually gets formed. If the only version of your DCF is one a model produced, you’ve prepped for an interview you’ll never actually get.
Good luck out there.
