31 / 37 · Concept
Multiply-add pipelines and data alignment
Split arithmetic at register boundaries and keep data, valid, and side operands aligned to the same transaction.
Lessons are free to read. Enroll to save your learning progress.
From an expression to two stages
Prerequisites: signed arithmetic, nonblockingNonblocking assignment An assignment written as <= in sequential RTL. It evaluates the right-hand side and schedules the update, allowing registers at the same edge to compute from the previous state. Learn more assignments, setupSetup The minimum time that input data must be stable before the capturing clock edge. If it is violated, the stored result is not guaranteed. Learn more requirements, and valid signals.
Suppose each input transaction contains signed 8-bit a and b, plus signed 16-bit c, with the following result. This is a multiply-add that processes each transaction independently, not an accumulator.
An exact 8-bit-by-8-bit product requires 16 bits. When adding c, computing at 16-bit width and widening afterward cannot recover upper bits already lost. Sign-extend both operands to 17 bits before addition.
Stage 1 stores the product and c together. Stage 2 adds those two values from the same transaction. Without delaying c, the old transaction's product would be mixed with the current transaction's c.
module multiply_add_pipeline (
input logic clk, rst, in_valid,
input logic signed [7:0] a, b,
input logic signed [15:0] c,
output logic out_valid,
output logic signed [16:0] y
);
logic v1;
logic signed [15:0] p1, c1;
logic signed [16:0] p_ext, c_ext;
assign p_ext = {p1[15], p1};
assign c_ext = {c1[15], c1};
always_ff @(posedge clk) begin
if (rst) begin
v1 <= 1'b0;
out_valid <= 1'b0;
end else begin
v1 <= in_valid;
out_valid <= v1;
if (in_valid) begin
p1 <= a * b;
c1 <= c;
end
if (v1) y <= p_ext + c_ext;
end
end
endmoduleThis interface can accept input at every edge and assumes that the receiver can always accept output. Do not compare y when out_valid=0. Initializing valid distinguishes unused data even without resettingReset A control that returns state to a specified initial value. Define whether it is synchronous or asynchronous and whether it has priority over other controls. Learn more the data registersRegister A circuit that stores multiple bits of state. The synchronous registers in this course store their specified inputs at a clock edge. Learn more.
Distinguish “two stages” from the observation time
The table records register states immediately after each edge that accepts input. A and B are different transactions.
| Edge | Accepted input | Stage 1 | Output register |
|---|---|---|---|
| E0 | A | A's product and c | invalid |
| E1 | B | B's product and c | A's result, valid |
| E2 | None | invalid | B's result, valid |
| E3 | None | invalid | invalid |
A's result is ready immediately after E1. The next circuit using the same clock samples it at E2. Thus two cycles separate input acceptance at E0 and output reception at E2. Do not mix the time when the output register changes with the time when the next circuit receives it when explaining latency.
Each column records one rising edgeRising edge The instant at which the clock changes from 0 to 1. Distinguish it from a level, which refers to the entire interval during which CLK=1. Learn more. Inputs are sampled just before the edge; states whose names include “after” are shown immediately after the update. Column alignment indicates sample order, not physical propagation delayPropagation delay The time from an input change until the output settles to the correct value. Logical equivalence and timing behavior are separate properties. Learn more. The transactions are A=(2,3,10) and B=(4,5,100). Stage 1 stores the product and c together, and the next edge adds the previous stage's values. Dashes denote invalid intervals.
View waveform data
| Signal | Wave | Bus values |
|---|---|---|
| edge | 2345 | E0 → E1 → E2 → E3 |
| accepted | 234. | A → B → - |
| product1 after | 234. | 6 → 20 → - |
| c1 after | 234. | 10 → 100 → - |
| y after | 2345 | - → 16 → 120 → - |
| out_valid after | 01.0 |
How much faster does pipelining make it?
Consider a simple model with a multiplication delay of 3.2 ns, addition delay of 1.1 ns, and register overhead of 0.2 ns. These values are assumptions for the calculation.
The upper bound on maximum frequency changes from approximately 222 MHz to 294 MHz. Two stages do not necessarily double throughput. The slowest stage and register overhead determine the period. Actual results depend on synthesis, placement, and constraints.
Verify complete transactions
When an input is accepted, append a×b+c from an integer reference model to a queue. At the edge where output valid is received, remove and compare the expected value at the front. Insert intervals with in_valid=0 to check that bubbles move with the data. A pattern that changes only c substantially on every transaction exposes a missing delay effectively.
Further reading: MIT OpenCourseWare — Performance Measures
Try it yourself
Apply consecutive transactions A=(a=3,b=-4,c=10) and B=(a=2,b=5,c=100). Find the correct output sequence and the erroneous result for A that could occur if the current input c is added without delaying it.
Read the explanation
The correct results are -12+10=-2 for A and 10+100=110 for B. If the current c is B's 100 when A's product reaches stage 2, the result is 88. Correctly delaying valid alone cannot prevent this error when operand alignment is wrong.