Altifigence Academy

31 / 37 · Concept

Multiply-add pipelines and data alignment

Split arithmetic at register boundaries and keep data, valid, and side operands aligned to the same transaction.

From an expression to two stages

Prerequisites: signed arithmetic, nonblockingNonblocking assignment An assignment written as <= in sequential RTL. It evaluates the right-hand side and schedules the update, allowing registers at the same edge to compute from the previous state. Learn more assignments, setupSetup The minimum time that input data must be stable before the capturing clock edge. If it is violated, the stored result is not guaranteed. Learn more requirements, and valid signals.

Suppose each input transaction contains signed 8-bit a and b, plus signed 16-bit c, with the following result. This is a multiply-add that processes each transaction independently, not an accumulator.

y=a×b+cy=a\times b+c

An exact 8-bit-by-8-bit product requires 16 bits. When adding c, computing at 16-bit width and widening afterward cannot recover upper bits already lost. Sign-extend both operands to 17 bits before addition.

Stage 1 stores the product and c together. Stage 2 adds those two values from the same transaction. Without delaying c, the old transaction's product would be mixed with the current transaction's c.

SystemVerilog
module multiply_add_pipeline (
    input  logic               clk, rst, in_valid,
    input  logic signed [7:0]  a, b,
    input  logic signed [15:0] c,
    output logic               out_valid,
    output logic signed [16:0] y
);
    logic v1;
    logic signed [15:0] p1, c1;
    logic signed [16:0] p_ext, c_ext;
    assign p_ext = {p1[15], p1};
    assign c_ext = {c1[15], c1};

    always_ff @(posedge clk) begin
        if (rst) begin
            v1        <= 1'b0;
            out_valid <= 1'b0;
        end else begin
            v1        <= in_valid;
            out_valid <= v1;
            if (in_valid) begin
                p1 <= a * b;
                c1 <= c;
            end
            if (v1) y <= p_ext + c_ext;
        end
    end
endmodule

This interface can accept input at every edge and assumes that the receiver can always accept output. Do not compare y when out_valid=0. Initializing valid distinguishes unused data even without resettingReset A control that returns state to a specified initial value. Define whether it is synchronous or asynchronous and whether it has priority over other controls. Learn more the data registersRegister A circuit that stores multiple bits of state. The synchronous registers in this course store their specified inputs at a clock edge. Learn more.

Distinguish “two stages” from the observation time

The table records register states immediately after each edge that accepts input. A and B are different transactions.

EdgeAccepted inputStage 1Output register
E0AA's product and cinvalid
E1BB's product and cA's result, valid
E2NoneinvalidB's result, valid
E3Noneinvalidinvalid

A's result is ready immediately after E1. The next circuit using the same clock samples it at E2. Thus two cycles separate input acceptance at E0 and output reception at E2. Do not mix the time when the output register changes with the time when the next circuit receives it when explaining latency.

Each column records one rising edgeRising edge The instant at which the clock changes from 0 to 1. Distinguish it from a level, which refers to the entire interval during which CLK=1. Learn more. Inputs are sampled just before the edge; states whose names include “after” are shown immediately after the update. Column alignment indicates sample order, not physical propagation delayPropagation delay The time from an input change until the output settles to the correct value. Logical equivalence and timing behavior are separate properties. Learn more. The transactions are A=(2,3,10) and B=(4,5,100). Stage 1 stores the product and c together, and the next edge adds the previous stage's values. Dashes denote invalid intervals.

The product and c must belong to the same transaction
View waveform data
Wave data: each character is one interval; a dot holds the previous state; p is a clock cycle.
SignalWaveBus values
edge2345E0 → E1 → E2 → E3
accepted234.A → B → -
product1 after234.6 → 20 → -
c1 after234.10 → 100 → -
y after2345- → 16 → 120 → -
out_valid after01.0

How much faster does pipelining make it?

Consider a simple model with a multiplication delay of 3.2 ns, addition delay of 1.1 ns, and register overhead of 0.2 ns. These values are assumptions for the calculation.

Tone≥3.2+1.1+0.2=4.5 nsT_{\mathrm{one}}\ge3.2+1.1+0.2=4.5\ \mathrm{ns}
Tpipe≥max⁡(3.2,1.1)+0.2=3.4 nsT_{\mathrm{pipe}}\ge\max(3.2,1.1)+0.2=3.4\ \mathrm{ns}

The upper bound on maximum frequency changes from approximately 222 MHz to 294 MHz. Two stages do not necessarily double throughput. The slowest stage and register overhead determine the period. Actual results depend on synthesis, placement, and constraints.

Verify complete transactions

When an input is accepted, append a×b+c from an integer reference model to a queue. At the edge where output valid is received, remove and compare the expected value at the front. Insert intervals with in_valid=0 to check that bubbles move with the data. A pattern that changes only c substantially on every transaction exposes a missing delay effectively.

Further reading: MIT OpenCourseWare — Performance Measures

Try it yourself

Apply consecutive transactions A=(a=3,b=-4,c=10) and B=(a=2,b=5,c=100). Find the correct output sequence and the erroneous result for A that could occur if the current input c is added without delaying it.

Read the explanation

The correct results are -12+10=-2 for A and 10+100=110 for B. If the current c is B's 100 when A's product reaches stage 2, the result is 88. Correctly delaying valid alone cannot prevent this error when operand alignment is wrong.

Your choice applies to this browser. Change it any time using the footer.