Fetch execute cycle

Traditional: fetch instruction. Based on bits in instruction, decide which gates need to be activated to cause data to be fetched, tested, etc. Activate data transfers and paths (e.g., from memory and register via ALU with operation (add, subtract, and, xor, compare) selected).

This traditional approach required complex logic to decode the instruction. Complex logic meant many levels of and/or/not logic. Each level added 5–10ns of delay. This led to…

Microcode: Each microcode instruction was a “long word” with dozens of fields. Each field controlled a set of gates that activated a data path or other action. A single instruction selected a microcode sequence to execute; the microcode could use bits of the instruction to select data paths. Logic decode was shallow. Easier to fix microcode than rewire a backplane. Which raised the question of “why all those CISC instructions anyway?” So RISC machines were born. Simple instructions, shallow decoding, fewer gates, lower power, faster instruction cycle. But bizarrely complicated to program; internal timings impinged the way the instructions had to be sequenced in the source. Classic example was that the results of a compare instruction were not available to the next instruction because it was already being fetched and decoded while the compare instruction was being executed, so the conditional branch had to be two instructions after the compare. Which led to suboptimal code (gratuitous no-ops being tossed in) or nearly-unmaintainable code. Even hard for compilers to manage. Which led to…

the x86/x64 architectures, where a CISC instruction is compiled into a sequence of RISC instructions. These are executed opportunistically and asynchronously, which then required a huge amount of logic to give the illusion that the CISC instructions were being executed sequentially, which led to…

The Itanic, I mean, Itanium, architecture, which appears to have been an aberration. It encoded up to three instructions in a longword, plus flag bits to say which could be executed in parallel, thus moving the complexity from simple linear complexity of circuitry to an NP-hard bin-packing problem in the compiler. While intending to be the best of both worlds, seems to have embodied the worst. Has been thankfully forgotten.

So there is no simple answer “under the hood”. But the simplest answer is the first, and while the implementation of that model is what we tell programmers is the model they can safely pretend is what is going on, the truth is far more complicated and is faster than the simple model suggests.

Comments

Popular posts from this blog

Thermodynamics equilibrium

Trappist 1e

π—₯π—Όπ—―π—Όπ˜π—Άπ—°s software