A high-aging dynamic memory command bus training method
Patent Information
- Application Number
- CN202610966238.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2046-07-01
AI Technical Summary
[0006]本发明提供了一种高时效的动态存储器命令总线训练方法,解决了现有技术中动态存储器命令总线的训练时效低,交互开销大的问题
本发明的技术方案通过在动态存储器的命令地址总线接收侧,为命令地址总线的多个信号线分别配置计数器和寄存器;主芯片通过所述命令地址总线向所述动态存储器发送测试序列,动态存储器基于所述时钟信号记录所述多个信号线的跳变时间信息,并存储于对应的寄存器中;通过数据总线读取多个信号线对应的寄存器中存储的多个计数值;确定并设置所述命令地址总线中所述多个信号线的发送延时。
Smart Images

Figure CN122474095B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of memory access control, and more specifically to a highly efficient dynamic memory command bus training method. Background Technology
[0002] Dynamic Random Access Memory (DRAM) is the main memory of modern computer systems. It uses the amount of charge stored in a capacitor to represent binary data, offering advantages such as high storage density and low cost, and is widely used in various computing devices. With the continuous improvement of processor performance, the data transfer rate of the DRAM interface has evolved to the level of several gigabits per second, placing stringent requirements on the signal integrity between the memory and the controller.
[0003] In a DRAM system, the Command / Address Bus (CA bus) is responsible for transmitting read / write commands and address information issued by the memory controller to the DRAM chips. Command bus training is a mechanism used to calibrate the CA bus signal. Its core purpose is to adjust the phase relationship between the CA signal and the clock signal, thereby ensuring that the DRAM can correctly sample command and address information under high-speed transmission conditions.
[0004] In existing technologies, command bus training methods typically employ a "scanning" approach. This involves the controller traversing the delay value range with a fixed step size, sequentially sending test patterns and waiting for DRAM feedback. Pass or failure boundaries are determined through comparison, and the center of the window is ultimately selected as the optimal operating point. However, existing methods have low training efficiency, requiring multiple iterations to train a single signal line. If all signal lines of the command address bus are trained individually, the training time increases linearly with the number of signal lines, resulting in insufficient independent calibration capability for each signal line. For example, existing training schemes often adjust delays on a per-capable basis for the entire command address bus, making it difficult to achieve independent calibration for each signal line. Due to factors such as routing differences and on-chip deviations, the transmission delays of different command address signal lines vary, and the overall adjustment method cannot provide optimal timing margins for each signal line.
[0005] Therefore, how to provide a command bus training method with high timeliness and low interaction overhead, which can achieve independent and fast delay calibration for each signal line of the command address bus, is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0006] This invention provides a highly efficient dynamic memory command bus training method, which solves the problems of low training efficiency and high interaction overhead in the prior art.
[0007] In view of the above problems, the present invention provides a high-efficiency dynamic memory command bus training method, comprising: On the command address bus receiving side of the dynamic memory, counters and registers are configured for multiple signal lines of the command address bus, wherein the counters operate based on the clock signal received by the dynamic memory. The main chip sends a test sequence to the dynamic memory via the command address bus. The dynamic memory records the transition time information of the multiple signal lines based on the clock signal and stores it in the corresponding register. The multiple count values stored in the registers corresponding to the multiple signal lines are read through the data bus. Based on the multiple count values, determine and set the transmission delay of the multiple signal lines in the command address bus.
[0008] One or more technical solutions provided in this invention have at least the following technical effects or advantages: The technical solution of the present invention configures counters and registers for multiple signal lines of the command address bus on the receiving side of the dynamic memory; the main chip sends a test sequence to the dynamic memory through the command address bus; the dynamic memory records the transition time information of the multiple signal lines based on the clock signal and stores it in the corresponding registers; multiple count values stored in the registers corresponding to the multiple signal lines are read through the data bus; and the transmission delay of the multiple signal lines in the command address bus is determined and set.
[0009] The high-efficiency dynamic memory command bus training method provided by this invention first configures independent counters and registers for multiple signal lines of the command address bus on the dynamic memory receiving side, and makes the counters run based on the receiving clock, thereby achieving the effect of establishing an independent time quantization hardware foundation for each signal line on the dynamic memory side, providing hardware support for subsequent single acquisition of the transition times of all signal lines.
[0010] Furthermore, the main chip sends only one set of test sequences, and the dynamic memory detects the transitions of each signal line in parallel based on the clock signal, and latches the count value corresponding to the transition moment into its respective register. This achieves the effect of transforming the traditional trial-and-error training that requires multiple iterations into a single measurement-based acquisition, which greatly reduces the number of training interactions and thus improves training efficiency.
[0011] Furthermore, the data bus sends read commands to the dynamic memory, sequentially reads the count values stored in each register, and stores them according to the signal line correspondence. This achieves the effect of completely transmitting the quantized data of the independent and precise transition time of each signal line back to the main chip, providing an accurate and traceable data foundation for subsequent independent calculation of delay compensation for each signal line.
[0012] Finally, the count value of each signal line is converted into the receiving time, the transmission delay is calculated in combination with the predefined transmission time point, and the optimal transmission delay of each line is determined and configured independently according to the expected sampling point. This achieves the effect of independent delay calibration, completely eliminating the timing offset between lines caused by trace differences and on-chip deviations, so that each signal line can obtain the optimal timing margin.
[0013] In summary, the technical solution of this invention improves the timeliness of command bus training by decoupling the number of training interactions and training time from the number of signal lines. Simultaneously, it achieves independent delay calibration, effectively compensating for timing offsets between lines caused by factors such as routing differences and on-chip deviations. This ensures that each signal line obtains optimal timing margin, and the delay compensation amount can be determined through simple calculation, reducing the logic complexity and resource consumption of the controller and supporting rapid dynamic retraining during operation. It effectively solves the problems of low training timeliness and high interaction overhead in existing dynamic memory command buses. Attached Figure Description
[0014] Figure 1 This is a flowchart illustrating a high-efficiency dynamic memory command bus training method provided in an embodiment of the present invention.
[0015] Figure 2 This is a logical schematic diagram of a high-efficiency dynamic memory command bus training method provided in an embodiment of the present invention. Detailed Implementation
[0016] This invention provides a highly efficient dynamic memory command bus training method, which solves the problems of low training efficiency and high interaction overhead in the prior art.
[0017] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to these processes, methods, products, or devices.
[0018] Examples, such as Figure 1 , Figure 2 As shown, this embodiment of the invention provides a high-efficiency dynamic memory command bus training method, the method comprising: S100: On the command address bus receiving side of the dynamic memory, counters and registers are configured for multiple signal lines of the command address bus, wherein the counters operate based on the clock signal received by the dynamic memory.
[0019] In existing technologies, the command address bus receiving side of dynamic memory (DRAM) only includes basic logic such as input buffers and command decoders. Its structure is relatively simple; the DRAM is only responsible for sampling and decoding the command address signal at the rising edge of the clock and lacks the ability to measure signal transition moments. Therefore, during command bus training, the memory controller must traverse the delay value range with a fixed step size, sequentially sending test patterns and waiting for the DRAM to provide feedback on the sampling results via the data bus. After repeatedly iterating to determine the pass / fail boundary, the center of the window is selected as the optimal operating point. Traditional solutions result in numerous training interactions, long training times, and difficulty in achieving independent calibration of each signal line, failing to effectively compensate for timing offsets between lines caused by trace differences and on-chip deviations.
[0020] In this step, counters and registers are configured for each of the multiple signal lines on the command address bus, with each signal line forming a one-to-one mapping with its corresponding counter and register. Because different signal lines may exhibit different transition times due to routing differences and on-chip deviations, a corresponding counter and register are set for each signal line to provide a precise basis for subsequent independent transmission delay calibration of each signal line.
[0021] In this step, the counter operates based on the clock signal received from the dynamic memory. The clock signal is a periodic square wave signal used in a digital system to synchronize the operation of various components. It defines the time base for all operations in the system and is the smallest time unit for a digital circuit to perform a basic operation. For example, a rise edge occurs from a low level to a high level, and a fall edge occurs from a high level to a low level. One such rise and fall cycle constitutes one time period, i.e., the time interval between two rising edges. The counter operates according to the clock signal.
[0022] In this step, the register works in conjunction with the edge detector set on the command address bus receiving side of the dynamic memory to capture transition signals. When the edge detector detects a signal transition, that is, when it detects that the level of the current clock cycle is different from the level of the previous clock cycle, it outputs a pulse to the register, and the register locks and stores the value of the counter at the current moment.
[0023] In this step, the detection and storage operations corresponding to different signal lines across multiple signal lines are executed in parallel and independently. That is, the edge detectors of all command address signal lines simultaneously monitor the transition events on their respective signal lines. When each signal line transitions, only the register latch count value corresponding to that line is triggered, without affecting the detection and storage of other signal lines. When any line transitions, there is no need to wait for other lines to complete detection or storage, thereby achieving the goal of efficient training.
[0024] In this step, by configuring counters and registers for multiple signal lines of the command address bus on the receiving side of the dynamic memory, the traditional serial training process, which requires multiple iterations and line-by-line measurements, is transformed into a one-time acquisition process that can capture the transition moments of all signal lines in parallel with a single transmission of the test sequence. Simultaneously, each signal line has an independent measurement channel, which can accurately record the different transition moments caused by routing differences and on-chip deviations, providing a precise basis for subsequent independent transmission delay calibration of each signal line.
[0025] In summary, this step, by configuring independent counters and registers for each of the multiple signal lines of the command address bus on the dynamic memory receiving side and allowing the counters to run based on the receiving clock, achieves the effect of establishing an independent time quantization hardware foundation for each signal line on the dynamic memory side, providing hardware support for subsequent single acquisition of the transition moments of all signal lines.
[0026] S200: The main chip sends a test sequence to the dynamic memory through the command address bus. The dynamic memory records the transition time information of the multiple signal lines based on the clock signal and stores it in the corresponding register.
[0027] In traditional command bus training, the main chip typically needs to repeatedly send test sequences, fine-tuning the delay or phase with each transmission. The boundary is gradually approached by observing the pass / fail feedback from the dynamic memory, and adjustments are usually made on a byte-by-byte basis. This invention allows the main chip to send only one set of test sequences, while the dynamic memory side uses parallel hardware to directly record the actual transition times of each signal line. This not only compresses the training time from multiple iterations to a single acquisition but also captures the independent transition times of each line due to physical differences, providing a data foundation for subsequent fine-grained compensation and thus improving the overall timing margin.
[0028] Step S200 in the method provided by the present invention includes: The main chip continuously sends test sequences to the dynamic memory via the command address bus, wherein the main chip does not need to wait for the dynamic memory to return sampling results; The dynamic memory uses the clock signal as a reference to continuously sample the multiple signal lines and detects the signal transitions generated by the test sequence on the multiple signal lines. When any of the multiple signal lines undergoes a signal transition, the current count value of the counter corresponding to that signal line is stored in the register corresponding to that signal line.
[0029] In this step, the test sequence is a specific, predefined data pattern sent by the main chip (memory controller) to the dynamic memory via the command address bus during command bus training. This pattern is used to stimulate transition events on each signal line of the command address bus, thereby providing predictable and repeatable measurement samples for measuring transition moments on the dynamic memory side. Traditional solutions primarily involve sending test sequences with different delays and obtaining results through comparison by the main chip. In this invention, the main chip only needs to send one set of test sequences to complete the transition measurements of all signal lines, effectively improving measurement efficiency.
[0030] The test sequence contains multiple expected transition points predefined for each of the multiple signal lines. The dynamic memory only detects and records signal transitions that occur at the expected transition points to prevent transitions caused by glitches or crosstalk from causing measurement failures.
[0031] In this step, the counter is an edge counter, which increments on the edge of the clock signal. The dynamic memory samples multiple signal lines once according to the clock signal. That is, at the beginning of each clock cycle, the counter value increases once. Therefore, the count value corresponds to the number of cycles and represents the number of times it has been sampled.
[0032] In this step, the main chip first continuously sends test sequences to the dynamic memory via the command address bus. In traditional solutions, the sampling results of the test sequences need to be returned from the dynamic memory to the main chip for comparison. In this step, the main chip does not need to wait for the dynamic memory to return the sampling results.
[0033] Furthermore, the dynamic memory uses a clock signal as a reference to continuously sample multiple signal lines, detecting signal transitions generated on these lines by the test sequence. Simultaneous sampling of multiple signal lines allows for independent measurement of each line's transition timing.
[0034] Furthermore, when any of the multiple signal lines undergoes a signal transition, the current count value of the counter corresponding to that signal line is stored in the register corresponding to that signal line. This completes one detection process.
[0035] For example, consider a signal line CA0 with a test sequence length of three clock cycles T1, T2, and T3, each with a period of 1 ns. A predefined expected transition point for CA0 in the test sequence is assumed to be 2.2 ns, resulting in a detection window of 2.0 ns to 2.4 ns. First, the main chip continuously sends the test sequence to the dynamic memory via the command address bus. The dynamic memory uses the clock signal as a reference, sampling at 0 ns, 1 ns, and 2 ns respectively, with count values of 1, 2, and 3. When a signal transition occurs at 1.5 ns for CA0, since it is not within the detection window, it indicates the expected transition point is too far away and is not recorded. When a signal transition occurs at 2.1 ns for CA0, and 2.1 ns is within the detection window (2.0 ns to 3.0 ns), falling within cycle T3, the count value is 3. The current count value is then stored in a register, completing one detection cycle. Similarly, transition detection for other signal lines uses the same method, which will not be elaborated upon here.
[0036] It should be noted that the above values are for illustrative purposes only and do not constitute a limitation on the present invention.
[0037] In summary, this step achieves the effect of transforming the traditional iterative training that requires multiple iterations into a single measurement acquisition by having the main chip send only one set of test sequences and the dynamic memory detect the transitions of each signal line in parallel based on the clock signal and latch the corresponding count values at the transition moments into their respective registers. This greatly reduces the number of training interactions and thus improves training efficiency.
[0038] S300: Read the multiple count values stored in the registers corresponding to the multiple signal lines via the data bus.
[0039] After obtaining multiple count values stored in the register, this step directly reads these count values from the register via the data bus. Compared to traditional solutions where the controller needs to repeatedly send test patterns and wait for binary pass or failure feedback, requiring multiple iterative boundary searches to approximate the optimal delay, this step obtains the precise count values for the transition moments of each signal line with a single read. This significantly reduces the number of training interactions and elevates qualitative judgment to quantitative measurement. The controller can directly calculate the precise delay compensation required for each signal line without the need for complex boundary search algorithms.
[0040] Step S300 in the method provided by the present invention includes: After the test sequence is sent, it is confirmed that the dynamic memory has completed the recording of the signal transitions of the multiple signal lines; A read command is sent to the dynamic memory via the data bus to sequentially read the count values stored in the registers corresponding to the plurality of signal lines; Receive the plurality of count values returned by the dynamic memory through the data bus; The plurality of count values are stored according to their correspondence with the plurality of signal lines.
[0041] In this step, the read command is a type of read command trained by bypassing the memory array to directly access the registers, in order to prevent data corruption that may result from direct access to the memory array. This ensures that the training process is safe and controllable.
[0042] In this step, after the test sequence is sent, it is necessary to wait for the dynamic memory to complete the detection and latching of signal transitions for all signal lines. The waiting time is set to the length of the test sequence plus the processing time, with a safety margin. For example, if the test sequence length is 8 cycles, the dynamic memory processing time is 2 cycles, and a safety margin of 3 cycles is set, the waiting time is 8 + 2 + 3 = 13 cycles. After the waiting time expires, it is confirmed that the recording is complete. After confirming that the dynamic memory has completed recording the signal transitions of the multiple signal lines, the next operation is performed.
[0043] Furthermore, after confirming that the dynamic memory recording is complete, a read command is sent to the dynamic memory via the data bus to sequentially read the count values stored in each register. The count value stored in the register corresponding to each signal line is obtained; for example, if the count value stored in the register corresponding to the CA0 signal line is 3, then the read count value data is CA0-3.
[0044] Furthermore, the system receives multiple count values returned by the dynamic memory via the data bus. The received data includes multiple count values corresponding to the read signal lines, as well as information on whether the count values are valid. If a read count value is determined to be invalid, the next step is executed.
[0045] The method for determining the validity of a count value involves setting a validity flag. For example, if the count value corresponds to the expected transition point in a given period, and there is only one recorded transition count value, the count value is considered valid, and the validity flag is set to 1. Conversely, if the read count value is too large compared to the expected transition point, the count value corresponding to that signal line is considered invalid; if no count value is read, the count value corresponding to that signal line is considered invalid; if multiple pulses appear within the detection window, causing the count value to overlap, the count value corresponding to that signal line is considered invalid, and in all these cases, the validity flag is 0.
[0046] In this step, after reading multiple count values, if a valid count value is not stored in the register corresponding to a certain signal line, it is determined that the training process of that signal line has malfunctioned, and the system error reporting mechanism is triggered. The system then identifies the number of the erroneous signal line, the error message, and retrains. Error messages include: no transition, multiple transitions, and overflow.
[0047] Furthermore, the multiple count values are stored according to their correspondence with multiple signal lines to obtain the final count value data, including each signal line and its corresponding count value, providing a structured data foundation for subsequent delay calculations.
[0048] For example, suppose there are two signal lines, CA0 and CA1, and the test sequence length is 4 cycles, with the desired transition points configured in cycles T1 and T3 respectively. After the main chip sends the test sequence, it waits for a fixed delay covering the maximum possible transition offset, such as 13 clock cycles, to ensure that the dynamic memory has completed all transition detection and latching. Subsequently, the main chip sends dedicated training read commands sequentially through the data bus, such as address 1 corresponding to CA0 and address 2 corresponding to CA1. Under the read command at their respective addresses, the dynamic memory latches the count values in its registers, such as CA0 latching 5 and CA1 latching 12, and determines whether it is valid. If the valid flag is set to 1, it returns through the data bus. After receiving the data, the main chip saves it as the corresponding count value data, such as CA0-5, CA1-12, and simultaneously saves the hardware-automatically set valid flag. At this point, the main chip obtains independent and accurate quantized values of the transition time for each signal line, preparing the data for subsequent calculation of the required transmission delay compensation for each line.
[0049] It should be noted that the above values are for illustrative purposes only and do not constitute a limitation on the present invention.
[0050] In summary, this step sends a read command to the dynamic memory via the data bus, sequentially reads the count values stored in each register, and stores them according to the signal line correspondence. This achieves the effect of completely transmitting the independent and precise transition time quantization data of each signal line back to the main chip, providing an accurate and traceable data foundation for subsequent independent delay compensation calculations for each signal line.
[0051] S400: Determine and set the transmission delay of the plurality of signal lines in the command address bus based on the plurality of count values.
[0052] This step, after obtaining multiple count values corresponding to multiple signal lines, determines and sets the transmission delay for each signal line. Traditional solutions require multiple iterative trials to find a rough boundary window center on a bus-wide basis, failing to compensate for inter-line differences. In this step, the main chip directly calculates the precise delay compensation required for each line based on the difference between the independent count value and the expected value for each signal line, and independently configures its transmission delay. This independent calibration method for each signal line not only completely eliminates inter-line timing offsets caused by routing differences and on-chip deviations, ensuring optimal timing margins for each line, but also transforms delay settings from fuzzy trial results to deterministic measurement results, improving system timing reliability and calibration accuracy.
[0053] Step S400 in the method provided by the present invention includes: Obtain the transmission time points at which the signal transitions of the multiple signal lines occur when the main chip sends the test sequence; For each of the plurality of signal lines, the transmission delay of the signal line is calculated based on the count value corresponding to the signal line and the transmission time point corresponding to the signal line. Based on the transmission delay, determine the optimal transmission delay for the signal line; According to the optimal transmission delay, the transmission delay of each of the multiple signal lines in the command address bus is set respectively.
[0054] In this step, the method for calculating the transmission delay of the signal line is as follows: Multiply the count value by the period of the clock signal to obtain the reception time when the signal transition reaches the input terminal of the dynamic memory; The transmission delay is obtained by subtracting the sending time from the receiving time.
[0055] For example, if the clock signal period is 1.0ns and the count value corresponding to the CA0 signal line is read as 5, then the receiving time at the dynamic memory input is 1.0ns × 5 = 5.0ns. If the sending time is then obtained as 1.0ns, the transmission delay is 5.0ns - 1.0ns = 4.0ns.
[0056] It should be noted that the above values are for illustrative purposes only and do not constitute a limitation on the present invention.
[0057] In this step, the first step is to obtain the transmission time points when multiple signal lines experience signal transitions during the main chip's transmission of the test sequence. Since there is a transmission delay between the main chip transmitting the test sequence and the dynamic memory detecting the transition, calculating this delay requires first obtaining the time when the main chip transmits the test sequence, i.e., the transmission time points.
[0058] Furthermore, for each of the multiple signal lines, the transmission delay of the signal line is calculated based on the count value corresponding to the signal line and the transmission time point corresponding to the signal line.
[0059] Furthermore, based on the transmission delay, the optimal transmission delay of the signal line is determined. This is to adjust the transmission delay of the main chip so that when the signal arrives at the dynamic memory, the transition time is within the desired sampling window. Therefore, it is necessary to calculate the delay compensation value, which is the difference between the transmission delay calculated above and the desired transition point, in order to determine the optimal transmission delay.
[0060] Furthermore, according to the optimal transmission delay, the transmission delay of multiple signal lines in the command address bus is set separately. The calculated optimal transmission delay value is then written into the independent delay register of each signal line of the main chip to make the configuration effective.
[0061] For example, taking a clock period of 1ns and a counter that increments by 1 in each cycle starting from 0, assuming the test sequence is predefined, CA0 jumps at a transmission time of 1.0ns during transmission period T1, and CA1 jumps at a transmission time of 3.0ns during T3. The main chip reads the count values from the dynamic memory side and finds that CA0 is 5 and CA1 is 12. The receiving time of CA0 is calculated to be 5.0ns and that of CA1 to be 12.0ns, thus calculating the transmission delay: CA0 is 5.0ns - 1.0ns = 4.0ns, and CA1 is 12.0ns - 3.0ns = 9.0ns. If the desired optimal sampling point is when the jump reaches a count value of 4.0 in DRAM, i.e., a receiving time of 4.0ns, then CA0 needs to be compensated by 4.0 - 5.0 = -1.0ns, i.e., transmitted 1.0ns earlier, and CA1 needs to be compensated by 4.0 - 12.0 = -8.0ns, i.e., transmitted 8.0ns earlier. The main chip writes two independent delay compensation values into the transmit delay registers of CA0 and CA1 respectively, completing the independent calibration of each signal line.
[0062] It should be noted that the above values are for illustrative purposes only and do not constitute a limitation on the present invention.
[0063] In summary, this step achieves independent delay calibration by converting the count value of each signal line into the reception time, calculating the transmission delay in combination with the predefined transmission time point, and determining and configuring the optimal transmission delay for each line based on the desired sampling point. This completely eliminates the timing offset between lines caused by trace differences and on-chip deviations, ensuring that each signal line can obtain the optimal timing margin.
[0064] In summary, this invention first configures independent counters and registers for each of the multiple signal lines of the command address bus on the dynamic memory receiving side, providing hardware support for subsequent single acquisition of the transition times of all signal lines. Next, the main chip sends only one set of test sequences, and the dynamic memory detects the transitions of each signal line in parallel based on the clock signal, latching the corresponding count values at each transition time into their respective registers. This significantly reduces the number of training interactions, thereby improving training efficiency. Then, the data bus sends a read command to the dynamic memory, sequentially reading the count values stored in each register and storing them according to the signal line correspondence, achieving the effect of completely transmitting the independent and precise quantized transition time data of each signal line back to the main chip. Finally, the count value of each signal line is converted into a receiving time, the transmission delay is calculated based on a predefined transmission time point, and the optimal transmission delay is determined and independently configured according to the desired sampling point, achieving the effect of independent delay calibration. This completely eliminates the inter-line timing offset caused by trace differences and on-chip deviations, ensuring that each signal line obtains optimal timing margin. This method decouples training time from the number of training interactions, improving the timeliness of command bus training. It also implements independent delay calibration, effectively compensating for timing offsets between lines caused by routing differences and on-chip deviations. This ensures that each signal line receives optimal timing margin, and the delay compensation amount can be determined through simple calculation, reducing the logic complexity and resource consumption of the controller and supporting rapid dynamic retraining during operation. It effectively solves the problems of low training timeliness and high interaction overhead in existing dynamic memory command buses.
[0065] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0066] This specification and accompanying drawings are merely illustrative examples of the invention and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its scope. Therefore, if such modifications and modifications fall within the scope of the invention and its equivalents, the invention is intended to include these modifications and modifications.
Claims
1. A high-efficiency dynamic memory command bus training method, characterized in that, include: On the command address bus receiving side of the dynamic memory, counters and registers are configured for multiple signal lines of the command address bus, wherein the counters operate based on the clock signal received by the dynamic memory. The main chip sends a test sequence to the dynamic memory through the command address bus. The dynamic memory records the transition time information of the multiple signal lines based on the clock signal. The transition time information is the current count value of the counter when the signal transition occurs, and the current count value is stored in the corresponding register. The multiple count values stored in the registers corresponding to the multiple signal lines are read through the data bus. Based on the multiple count values, determine and set the transmission delay of the multiple signal lines in the command address bus.
2. The high-efficiency dynamic memory command bus training method as described in claim 1, characterized in that, Each of the plurality of signal lines forms a one-to-one mapping relationship with the corresponding counter and register.
3. The high-efficiency dynamic memory command bus training method as described in claim 1, characterized in that, The detection and storage operations corresponding to different signal lines among the multiple signal lines are executed in parallel and are independent of each other.
4. The high-efficiency dynamic memory command bus training method as described in claim 1, characterized in that, The main chip sends a test sequence to the dynamic memory via the command address bus. The dynamic memory records the transition time information of the multiple signal lines based on the clock signal. The transition time information is the current count value of the counter when the signal transition occurs, and the current count value is stored in the corresponding register, including: The main chip continuously sends test sequences to the dynamic memory via the command address bus, wherein the main chip does not need to wait for the dynamic memory to return sampling results; The dynamic memory uses the clock signal as a reference to continuously sample the multiple signal lines and detects the signal transitions generated by the test sequence on the multiple signal lines. When any of the multiple signal lines undergoes a signal transition, the current count value of the counter corresponding to that signal line is stored in the register corresponding to that signal line.
5. The high-efficiency dynamic memory command bus training method as described in claim 4, characterized in that, The counter is an edge counter, and the counter value increases once every time the dynamic memory samples the plurality of signal lines according to the clock signal.
6. The high-efficiency dynamic memory command bus training method as described in claim 4, characterized in that, The test sequence includes multiple expected transition points predefined for each of the multiple signal lines, and the dynamic memory only detects and records signal transitions occurring at the expected transition points.
7. The high-efficiency dynamic memory command bus training method as described in claim 1, characterized in that, Read multiple count values stored in the registers corresponding to the multiple signal lines via the data bus, including: After the test sequence is sent, it is confirmed that the dynamic memory has completed the recording of the signal transitions of the multiple signal lines; A read command is sent to the dynamic memory via the data bus to sequentially read the count values stored in the registers corresponding to the plurality of signal lines; Receive the plurality of count values returned by the dynamic memory through the data bus; The plurality of count values are stored according to their correspondence with the plurality of signal lines.
8. The high-efficiency dynamic memory command bus training method as described in claim 7, characterized in that, Also includes: After reading the multiple count values, if it is detected that no valid count value is stored in the register corresponding to a certain signal line, it is determined that the training process of that signal line is abnormal and the system error reporting mechanism is triggered.
9. The high-efficiency dynamic memory command bus training method as described in claim 1, characterized in that, Based on the plurality of count values, determine and set the transmission delay of the plurality of signal lines in the command address bus, including: Obtain the transmission time points at which the signal transitions of the multiple signal lines occur when the main chip sends the test sequence; For each of the plurality of signal lines, the transmission delay of the signal line is calculated based on the count value corresponding to the signal line and the transmission time point corresponding to the signal line. Based on the transmission delay, determine the optimal transmission delay for the signal line; According to the optimal transmission delay, the transmission delay of each of the multiple signal lines in the command address bus is set respectively.
10. The high-efficiency dynamic memory command bus training method as described in claim 9, characterized in that, Calculating the transmission delay of the signal line includes: Multiply the count value by the period of the clock signal to obtain the reception time when the signal transition reaches the input terminal of the dynamic memory; The transmission delay is obtained by subtracting the sending time from the receiving time.
Citation Information
Patent Citations
Memory device with low pin count interface and corresponding methods and systems
CN115964314A
HBM3 command address bus training method and device
CN119621617A