In the past five years, AI has become central to daily life. Most students have used ChatGPT, Claude, or Gemini at some point, whether it be for homework help or out of curiosity. By now, these models feel familiar. However, they are evolving quickly under the hood. For years, the recipe for improving AI was simple: build a larger model, feed it more text, and add more computer chips. That recipe still works, but over the past six months, the benefits have moved elsewhere.
A large language model (LLM) is a program that writes one word at a time. It selects each word by predicting what comes next, drawing from patterns absorbed from enormous amounts of text. An LLM is essentially an autocomplete that has read most of the internet. That design produced the first wave of the AI revolution, but it comes with a flaw. The program commits to each word as it goes, so a single wrong turn early on ruins everything that follows. Early models weren’t able to erase their work and start over. Newer models, however, can. Before answering, they work through a problem on a private scratchpad, testing several approaches, checking intermediate results, and discarding dead ends. They reply only after this effort. Researchers call this “test-time compute,” because the effort happens when a question is asked rather than during training. Answers take longer and cost more, but they are far more accurate.
Thinking is a massive component of AI, but there is still another untouched piece—doing. A chatbot from two years ago could only produce text. When asked to repair a broken program, it would describe a repair without opening a single file. Nowadays, we have systems called agents to close that gap. An agent reads the program files, runs the code, reads the error message, edits a line, and runs the code again, repeating the cycle until the program works. The difference resembles the gap between a friend who explains a repair over the phone and a mechanic who picks up a wrench. A good mechanic does not guess from the sound of the engine, but runs a diagnostic, replaces a part, starts the car, and listens again. That loop matters more than any single step—acting, observing the result, and correcting the course is what allows a long task to reach a conclusion.
Thinking and acting both consume electricity, creating a new problem in older versions of AI. Older models routed every question through the entire network, causing even small questions to result in outsized electricity usage. While feasible early on in the AI revolution, the amount of energy the test-time compute framework uses is not scalable. The fix is a design called “mixture of experts.” A routing layer reads the question and sends it only to the small portion of the network suited to it, much like a service desk assigns a single car to the transmission specialist. The model keeps all of its knowledge but only uses a fraction of it at a time. New chips are built for that same pattern. Their real limit is not calculation speed. Instead, it’s memory bandwidth, referring to how quickly data can be pulled into a processor. A skilled mechanic still works slowly if every tool sits in a storeroom down the hall—the trip back and forth will take them time. Shortening that trip is what makes reasoning and AI agents affordable.
Machines now handle the parts of work that are easy to check. What remains valuable is harder: breaking a messy problem into multiple steps, deciding whether an answer is right, and explaining the reasoning behind it. Those skills are applicable to math, statistics, and english classes as much as to any coding class. Additionally, a caution to keep in mind is that most of the information about these systems comes from the companies selling them. The models still make confident mistakes, and catching those mistakes remains a human job.





































































































