DDR4 and LPDDR4 writes and reads: the roles of address, CS and DQS, and two-rank layouts
A memory access is just placing or fetching data at an address, but inside DDR4 and LPDDR4 the address is split into row and column, several devices are selected with CS, and data is captured against a reference signal called DQS. Following this also shows where length matching on the board matters.
Memory is a set of addresses and data
What a memory chip does comes down to one thing: specify an address and read or write the data there. A write is "put this value at this address" and a read is "give me the value at this address", and fast DRAMs such as DDR4 and LPDDR4 are no different in principle.
What differs is how the address and data are exchanged with the outside. The address is sent in two parts to save pins, a signal (CS) is needed to say who the current command is for when several chips share one bus, and a dedicated reference signal (DQS) travels with the data so it is not missed. This article centers on those three and follows DDR4 internals, two-rank layouts and the differences from LPDDR4. Choosing among memory types is covered in SRAM, DRAM, NOR, NAND, EEPROM, FRAM, MRAM and eMMC: how memories differ and when to use each.
Banks, rows, columns and sense amplifiers: why the address is split in two
Inside a DRAM chip, cells of one capacitor and one transistor (1T1C) sit in a grid (array). DDR4 divides it into bank groups and banks: 4 bank groups × 4 banks = 16 banks for x8, and 2 bank groups × 4 banks = 8 banks for x16. Alternating between bank groups lets the next command follow at a shorter interval (the difference between tCCD_S and tCCD_L).
A cell holds little charge and must be amplified by a sense amplifier before reading, so an access takes three steps. 1. ACT (activate) selects one row, puts the charge of that whole row on the bit lines, and amplifies and latches it in the sense amplifiers. 2. RD/WR selects a column and outputs only the needed column from the latched data to DQ, or writes from DQ. 3. PRE (precharge) closes the row and returns it to a state where the next row can be opened.
The address is split into row and column both to match these three steps and to cut the wiring to about half by time-multiplexing the same pins. The address pins (mainly A[13:0]) carry the row address on ACT and the column address on RD/WR.
Commands and timing: ACT / RD / WR / PRE set by CS_n, BG, BA and A
A command is encoded by the combination of CS_n (chip select), ACT_n, RAS_n/A16, CAS_n/A15 and WE_n/A14. With ACT_n low it is an ACT command and the A16-type pins are also part of the row address; with ACT_n high, the RAS_n/CAS_n/WE_n combination selects RD, WR, PRE and so on. BG[1:0] and BA[1:0] select the bank, and A[13:0] is the row or column address. Pins have two meanings because the same wires are shared between command and address.
On the time axis of Figure 3, tRCD is the time from opening a row with ACT until reads and writes are possible (RAS to CAS Delay), and CL is the time from RD to the first data (CAS Latency). For DDR4-3200 CL22, the clock is 1600MHz (data moves on both edges, so 3200 MT/s effective), the period is 1 ÷ 1600MHz ≈ 0.625ns, and 22 cycles of CL ≈ 13.75ns. For writes, CWL (CAS Write Latency) is a value shorter than CL, set by the standard for each speed bin. Data is basically BL8 (burst length 8): 8 bits are transferred on both edges over 4 clock cycles (the four humps of DQ in the figure).
| Parameter | Meaning | Cycles (guide) | Time (guide) |
|---|---|---|---|
| tRCD | ACT to column access allowed | 22 | 13.75ns |
| CL | RD to data output | 22 | 13.75ns |
| CWL | WR to DQS/DQ output | Less than CL | Set by the standard |
| tRP | PRE to row closed | 22 | 13.75ns |
| BL8 transfer | Time to transfer 8 data bits | 4 (both edges, DDR) | 2.5ns |
The role of DQS: source-synchronous, the sender distributes the reference too
CK (the system clock) is not used directly as the reference for capturing data. Instead, the sending side distributes a reference signal, DQS (data strobe), along with the data. This is source-synchronous operation. There is one DQS per byte lane (a group of 8 DQ bits).
For writes and reads, the phase relation between DQS and DQ is reversed. In a write the controller is the sender and shifts DQS by 90° (a quarter period) in advance so its edge falls at the middle of the DQ data, so the DRAM can capture directly on the DQS edge. In a read the DRAM is the sender and outputs DQS edge-aligned with the DQ edge, so the controller must delay DQS by 90° itself before using it, or it would catch the instant the data changes.
CK is not the reference because it is distributed one way, so the wiring delay to the DRAM, the internal delay and the wiring delay of DQ coming back add up to an amount that is not negligible against one period, and it also varies with trace length and temperature. DQS travels along the same path as the data, so the relation to the data is largely preserved even if the delay varies. Just before DQS rises, a preamble (a stable interval before the first toggle) is placed so the receiver knows data is coming.
Differences in length between DQ and DQS within a byte lane bear directly on the accuracy of this phase alignment. The approach to length matching in How tightly to match trace lengths applies as is.
The role of CS (chip select): a shared bus, and only the selected device answers
The command/address bus (ACT_n/RAS_n/CAS_n/WE_n/BG/BA/A and CK, excluding CS_n) is shared among several chips and ranks, because dedicated wiring would use too many pins and traces. Instead there is one CS_n (chip select, active low) per device, and only a device with CS_n low takes the command at that moment as addressed to it. For a device with CS_n high, the command and address in that cycle might as well not exist, which the DDR4 standard calls the DESELECT command.
If CS_n is miswired or floating, a device can take commands at unintended times, with symptoms such as skipped refresh or corrupted data.
Two ranks: CS0/CS1 are separate, DQ/DQS and address are shared
A module (DIMM, for instance) carrying two sets of DRAM chips is called two ranks. DQ, DQS and the command/address bus (including CK) are shared between ranks, while CS_n, CKE and ODT are wired separately per rank (CS0/CS1 and so on). The controller selects rank 0 by driving CS0 low and rank 1 by driving CS1 low, and shares the bus by time division.
When the DQ driver changes with the rank, if the next sender starts before the previous one has finished, signals collide on the bus (bus contention). The DDR4 controller ensures the turnaround gap needed between ranks (the time for rank-to-rank switching) before issuing a command to the next rank, and also makes DQS rise only after the previous rank's DQS has surely ended.
ODT (On-Die Termination) also makes this shared bus work. During a write the unselected rank's DRAM is on the same DQ wires, and unterminated it would reflect and disturb the waveform, so DDR4 enables the unselected rank's ODT to terminate the bus. In JEDEC terms the nominal termination is RTT_NOM, the termination during writes is RTT_WR, and the weak idle termination is RTT_PARK, all set in the mode registers.
Stacking ranks adds bus load (input capacitance, branches), increasing reflection and crosstalk, and the maximum speed that can be allowed tends to fall. One-rank configurations are often guaranteed to higher frequencies, and how far it falls depends on the board and parts, but the speed trade-off with rank count should be a design premise.
DDR4's DQ is terminated by ODT, but one-way signals fed by fly-by wiring, such as the command/address bus, still use external termination resistors pulled to VTT (mid-level). The VTT termination circuit is covered in the "DDR VTT termination" section of An LDO cannot sink current: where current injected through protection diodes goes.
How LPDDR4 differs: two channels per die, a 6-bit CA bus, low-voltage signaling
LPDDR4 makes different trade-offs from DDR4 for mobile devices. One die holds two independent 16-bit channels (channels A and B), each with its own CK, CA (command/address), CS, DQ[15:0] and DQS, and they operate as completely separate memories.
LPDDR4's command/address uses only the 6 lines CA[5:0], sampled as SDR, that is, on the rising edge of CK only. CS is active high (the opposite polarity of DDR4's CS_n) and is high only in the first cycle of a command. Most commands span two clocks: ACT, which opens a row, sends bank and row address in the two cycles ACT-1 and ACT-2, and reads and writes send the column address in the two cycles RD-1/WR-1 and CAS-2. Sending the row and column addresses takes ACT (2 clocks) + RD (2 clocks) = 4 clocks in total. The idea behind DDR4 splitting row and column to save pins is extended to the command itself.
The signaling is LVSTL (Low Voltage Swing Terminated Logic), a terminated scheme with an even smaller swing than DDR4's POD12. Termination is ODT with the same idea as DDR4, but DQ ODT for data and CA ODT for command/address are specified separately. Burst length and VDDQ are in the table.
| Item | DDR4 | LPDDR4 |
|---|---|---|
| Channel structure | A single x4/x8/x16 bus | Two independent channels per die (x16 × 2) |
| Command/address | Many pins (BG/BA/A etc.), fixed on one edge | 6 lines CA[5:0], SDR. Two clocks per command (ACT-1/2, RD-1 and CAS-2, etc.) |
| Banks | x8: 4BG×4 = 16 banks, x16: 2BG×4 = 8 banks | Equivalent to 8-16 banks, depending on die density |
| Burst length | BL8 is basic | BL16 is basic (BL32 selectable on some parts) |
| VDDQ | 1.2V | 1.1V |
| Signaling | POD12 (DQ) | LVSTL (lower swing and termination) |
| Termination | ODT (same idea for DQ and CA/ADDR) | DQ ODT and CA ODT specified separately |
| Chip select | CS_n per rank | CS per channel and per rank |
Training: fly-by routing and write leveling
A DDR4 DIMM routes CK and command/address to each chip in a fly-by (daisy chain): one trace runs end to end and each chip taps off it, which keeps a clean transmission line without stubs. The price is that CK arrives at a different time at each chip, later for chips farther downstream.
This skew would misalign the DQS the controller sends with each chip's CK, so DDR4 has an initialization procedure called write leveling. The controller sweeps the DQS delay for each chip while toggling it, and the DRAM samples the level of the CK reaching that chip on the DQS rising edge and returns the result over DQ. The controller finds the delay where the response flips from 0 to 1 and fixes it as the DQS timing for that chip. Because CK is later at downstream chips, DQS is also sent later for them.
After that come read training (delaying DQS so data is captured at the center of the DQ eye) and Vref training (fine-tuning the H/L threshold voltage to the signal quality). The controller performs these automatically right after power-up, but trace length and component placement affect their success and margin.
LPDDR4 has short traces and is often routed point to point rather than fly-by, so CK skew is smaller than in DDR4. Procedures equivalent to read training and Vref training are still specified, so short traces do not remove the need for adjustment.
Refresh: ACT/PRE running out of sight
A DRAM cell is a capacitor, so its charge leaks away and the data is lost if left alone. The controller issues the REF (refresh) command periodically, and the DRAM opens all rows in turn (repeating ACT/PRE internally) to recharge them. The standard interval (tREFI) is 7.8µs (at 85°C or below), and on some parts the standard requires halving the interval at the high end (85-95°C), meaning refresh twice as often.
While REF runs (tRFC, longer for higher density, hundreds of ns on large parts), that bank (or all banks, depending on the command) cannot be accessed. Refresh therefore takes a few percent of memory bandwidth for certain. LPDDR4 also has temperature-compensated self-refresh (TCSR), which lowers the refresh rate with temperature and saves wasted power when cold.
Implications for board design: what to length-match
There are two things to length-match. 1. DQ and DQS within a byte lane. Their relative timing is directly the read/write margin. 2. CK and CMD/ADDR. The DRAM latches CMD/ADDR on the CK edge, so skew between them cuts directly into setup and hold margin. On a DIMM both arrive at each chip together and in sequence over the fly-by wiring, so keep their relative timing from drifting along the way. DQ/DQS is source-synchronous and CMD/ADDR is CK-synchronous, so the references differ and the groups are matched separately.
For how tight to go (converting length differences to time skew), the How tightly to match trace lengths article and the length to delay and delay to length calculators work as is. For effective permittivity and characteristic impedance, use the microstrip impedance calculator too.
One rank has a light load and simple routing but too little capacity; two ranks add capacity but increase bus load, tightening reflection and crosstalk margin, and the fly-by skew needs care. The driver strength can be estimated with the I/O buffer drive calculator.
Frequently asked questions
Why capture data with DQS instead of CK?
Why can a two-rank module run slower?
What happens if CS_n is miswired or left floating?
Does LPDDR4 need write leveling too?
What is the difference between RTT_NOM, RTT_WR and RTT_PARK?
Standards and references
- JEDEC JESD79-4, DDR4 SDRAM Standard — Command definitions (ACT/RD/WR/PRE/REF), bank group structure, timing parameters such as tRCD/CL/CWL/tRP, and the write leveling procedure
- JEDEC JESD209-4, LPDDR4 Standard — Two-channel structure, command/address encoding on CA[5:0], BL16/32, VDDQ and LVSTL signaling
- JEDEC JESD8 series (I/O standards such as SSTL and POD) — Basics of signaling and termination for DDR-family memory
- Memory vendors' compatibility documents and speed bin tables — Guide to maximum operating frequency by rank count and configuration