Command bus training method, electronic device, medium and command bus training apparatus

By combining command address remapping, data remapping, and delay training modules with the center point algorithm, the alignment and synchronization challenges caused by pin layout and routing in LPDDR5 command bus training are solved, achieving fast and reliable signal alignment and reducing training latency.

CN119847954BActive Publication Date: 2025-11-18XIN YAOHUI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411939199.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-11-18
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

Existing command bus training schemes are affected by the pin layout and routing after chip packaging, resulting in poor training performance and high overall training latency, making it difficult to meet the state alignment and synchronization requirements of LPDDR5.

Method used

The command address remapping module and the data remapping module are used to map the pins of the memory module to be trained to the pins of the standard memory module. The delay training module and the center point algorithm are used to adjust the delay of the command address signal to achieve alignment between the command address signal and the data signal.

Benefits of technology

It effectively overcomes the adverse effects of pin layout and wiring after chip packaging, quickly and reliably determines alignment relationships and locates the optimal center position, improves training effect, reduces overall training latency, and supports fast response and device adaptation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119847954B_ABST
    Figure CN119847954B_ABST
Patent Text Reader

Abstract

The application relates to the computer technical field and provides a command bus training method, an electronic device, a medium and a command bus training device. Through optimization on hardware, a command address remapping module, a data remapping module and a delay training module are provided, and through optimization on software, the time delay of a command address signal is adjusted based on a center point algorithm through the delay training module, alignment between the command address signal and a data signal is realized, the demand of low-power fifth-generation double-rate and related memory technology specifications in state alignment and synchronization can be met, the adverse effects of pin layout and wiring after chip packaging can be effectively overcome, the adverse effects caused by the existence of a metastable state interval can be effectively overcome, the alignment relationship and the optimal center position can be quickly and reliably determined, the training effect is improved and the overall training delay is reduced, and the application is helpful for quick response and equipment adaptation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a command bus training method, electronic device, medium, and command bus training apparatus. Background Technology

[0002] Double Data Rate 5 (DDR5) is a memory technology specification designed to improve data flow channels and data transfer efficiency. Low Power Double Data Rate 5 (LPDDR5) adds low power (LP) to DDR5, meaning lower power consumption and a smaller size. LPDDR5 retains the characteristics of DDR, namely double the data rate, meaning data is transmitted on both the rising and falling edges of the signal, i.e., twice the data transfer volume of synchronous dynamic random access memory (SDRAM) per clock cycle. LPDDR5 is used in power- and size-sensitive applications such as smartphones, smartwatches, and automotive intelligent systems. The LPDDR5 clocking scheme uses a dual-clock signal design, unlike the single-clock signal design of DDR5. This helps improve the effective data bus speed, but also brings challenges in state alignment and synchronization. To meet the state alignment and synchronization requirements of the LPDDR5 clocking scheme, command bus training (CBT) is used to achieve signal alignment. However, existing command bus training schemes are affected by the pin layout and routing after chip packaging. During the training process, they rely on asynchronous signal feedback to determine the alignment relationship. Furthermore, due to the existence of metastable regions, there are deviations in position determination, resulting in poor training performance and high overall training latency. This is not conducive to fast response and device adaptation.

[0003] To this end, this application provides a command bus training method, electronic device, medium, and command bus training apparatus, which can not only meet the requirements of low-power fifth-generation double-rate and related memory technology specifications in terms of state alignment and synchronization, but also effectively overcome the adverse effects of pin layout and wiring after chip packaging, effectively overcome the adverse effects caused by the existence of metastable regions, quickly and reliably determine the alignment relationship and locate the optimal center position, improve the training effect and reduce the overall training latency, and help with rapid response and device adaptation. Summary of the Invention

[0004] Firstly, this application provides a command bus training method. The command bus training method includes: mapping multiple command address bus pins of a storage module to be trained to standard command address bus pins of a standard storage module using a command address remapping module; mapping multiple data bus pins of the storage module to be trained to standard data bus pins of the standard storage module using a data remapping module, wherein the multiple command address bus pins of the storage module to be trained and the multiple data bus pins of the storage module to be trained satisfy a first connection relationship, and the standard command address bus pins of the standard storage module and the standard data bus pins of the standard storage module satisfy a second connection relationship; using a delay training module, based on the first and second connection relationships, determining the asynchronous signal feedback source corresponding to each of the multiple command address bus pins of the storage module to be trained when they are invoked; and then, adjusting the delay of the command address signals corresponding to each of the multiple command address bus pins of the storage module to be trained based on a center point algorithm, thereby achieving alignment between the command address signals corresponding to each of the multiple command address bus pins of the storage module to be trained and the data signals corresponding to each of the multiple data bus pins of the storage module to be trained.

[0005] The first aspect of this application not only meets the requirements of low-power fifth-generation double-rate and related memory technology specifications in terms of state alignment and synchronization, but also effectively overcomes the adverse effects of pin layout and wiring after chip packaging, effectively overcomes the adverse effects of the existence of metastable regions, quickly and reliably determines the alignment relationship and locates the optimal center position, improves training effect and reduces overall training latency, and helps with rapid response and device adaptation.

[0006] In one possible implementation of the first aspect of this application, the plurality of command address bus pins of the storage module to be trained includes a first command address bus pin. The first command address bus pin is connected to a first data bus pin among the plurality of data bus pins of the storage module to be trained based on the first connection relationship. The delay training module adjusts the delay of the first command address signal corresponding to the first command address bus pin based on the center point algorithm, thereby achieving alignment between the first command address signal corresponding to the first command address bus pin and the first data signal corresponding to the first data bus pin.

[0007] In one possible implementation of the first aspect of this application, the delay training module adjusts the delay of the first command address signal corresponding to the first command address bus pin based on the center point algorithm, thereby achieving alignment between the first command address signal corresponding to the first command address bus pin and the first data signal corresponding to the first data bus pin. This includes: the delay training module determining the left boundary of the first command address signal using a first function of the center point algorithm, and determining the right boundary of the first command address signal using a second function of the center point algorithm; then determining the center point position of the first command address signal based on the determined left and right boundaries; and then adjusting the delay of the first command address signal based on the determined center point position.

[0008] In one possible implementation of the first aspect of this application, the first function of the center point algorithm is to sequentially scan the error position of the left boundary of the first command address signal, the metastable position of the left boundary of the first command address signal, and the stable position of the left boundary of the first command address signal; and the second function of the center point algorithm is to sequentially scan the error position of the right boundary of the first command address signal, the metastable position of the right boundary of the first command address signal, and the stable position of the right boundary of the first command address signal.

[0009] In one possible implementation of the first aspect of this application, the delay training module determines the left boundary of the first command address signal using the first function of the center point algorithm, including: increasing the delay of the first command address signal when the first data signal feedback is low and the clock signal does not acquire the first command address signal, and maintaining the delay of the first command address signal when the first data signal feedback is high and the clock signal acquires the first command address signal.

[0010] In one possible implementation of the first aspect of this application, the delay training module determines the right boundary of the first command address signal using the second function of the center point algorithm, including: reducing the delay of the first command address signal when the first data signal feedback is low and the clock signal does not acquire the first command address signal, and maintaining the delay of the first command address signal when the first data signal feedback is high and the clock signal acquires the first command address signal.

[0011] In one possible implementation of the first aspect of this application, the delay training module, which uses the first function of the center point algorithm to determine the left boundary of the first command address signal, further includes: initially setting the delay of the first command address signal to a first initial delay, and then using the first function of the center point algorithm to perform a first loop search number; the delay training module, which uses the second function of the center point algorithm to determine the right boundary of the first command address signal, further includes: initially setting the delay of the first command address signal to a second initial delay, and then using the first function of the center point algorithm to perform a second loop search number, wherein the first initial delay is less than the second initial delay.

[0012] In one possible implementation of the first aspect of this application, both the first loop search count and the second loop search count are determined based on the maximum frequency of the clock signal, the first initial delay corresponds to a single step size unit, the step size unit is determined based on the maximum frequency of the clock signal, and the second initial delay corresponds to multiple step size units.

[0013] In one possible implementation of the first aspect of this application, the plurality of command address bus pins of the memory module to be trained include a command address bus pin for a first data channel and a command address bus pin for a second data channel, wherein the standard command address bus pin of the standard memory module corresponding to the command address bus pin for the first data channel is different from the standard command address bus pin of the standard memory module corresponding to the command address bus pin for the second data channel.

[0014] In one possible implementation of the first aspect of this application, the plurality of data bus pins of the training memory module include paired data bus pins for differential signal transmission, and the standard data bus pins of the standard memory module corresponding to the paired data bus pins are used for differential signal transmission.

[0015] In one possible implementation of the first aspect of this application, the first connection relationship and the second connection relationship are determined based on the computer memory specifications associated with the storage module to be trained.

[0016] In one possible implementation of the first aspect of this application, the computer memory specifications are Low Power Generation 5 Double Data Rate, Internet Protocol Version 4, data processing unit design specifications, or image processing unit design specifications.

[0017] Secondly, embodiments of this application also provide a computer device, the computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a method according to any of the above-mentioned implementations.

[0018] Thirdly, embodiments of this application also provide a computer-readable storage medium storing computer instructions that, when executed on a computer device, cause the computer device to perform a method according to any of the above-described implementations.

[0019] Fourthly, embodiments of this application also provide a computer program product, the computer program product including instructions stored on a computer-readable storage medium, which, when executed on a computer device, cause the computer device to perform a method according to any of the above-described aspects. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 A flowchart illustrating a command bus training method provided in an embodiment of this application;

[0022] Figure 2 A schematic diagram of a command bus training device provided in an embodiment of this application;

[0023] Figure 3 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Detailed Implementation

[0024] The embodiments of this application will now be described in further detail with reference to the accompanying drawings.

[0025] It should be understood that in the description of this application, "at least one" means one or more, and "multiple" means two or more. In addition, the words "first," "second," etc., unless otherwise stated, are used only for the purpose of distinguishing descriptions and should not be construed as indicating or implying relative importance or order.

[0026] Figure 1 This is a flowchart illustrating a command bus training method provided in an embodiment of this application. Figure 1As shown, the command bus training method includes the following steps.

[0027] Step S101: Through the command address remapping module, map the multiple command address bus pins of the memory module to be trained to the standard command address bus pins of the standard memory module one by one.

[0028] Step S103: Through the data remapping module, the multiple data bus pins of the storage module to be trained are mapped one by one to the standard data bus pins of the standard storage module. The multiple command address bus pins of the storage module to be trained and the multiple data bus pins of the storage module to be trained satisfy a first connection relationship, and the standard command address bus pins of the standard storage module and the standard data bus pins of the standard storage module satisfy a second connection relationship.

[0029] Step S105: Using the delay training module, based on the first connection relationship and the second connection relationship, determine the asynchronous signal feedback source corresponding to each of the multiple command address bus pins of the storage module to be trained when they are called. Then, based on the center point algorithm, adjust the delay of the command address signals corresponding to each of the multiple command address bus pins of the storage module to be trained, thereby achieving alignment between the command address signals corresponding to each of the multiple command address bus pins of the storage module to be trained and the data signals corresponding to each of the multiple data bus pins of the storage module to be trained.

[0030] Figure 1The command bus training method shown can be applied to memory technology specifications such as LowPower Double Data Rate 5 (LPDDR5). LPDDR5 is based on Double Data Rate 5 (DDR5) but adds low power (LP), meaning lower power consumption and smaller size. LPDDR5 retains the characteristics of DDR, namely double the data rate, which means data transmission occurs on both the rising and falling edges of the signal, i.e., twice the data transfer volume of synchronous dynamic random access memory (SDRAM) per clock cycle. LPDDR5 is used in power- and size-sensitive applications such as mobile devices like smartphones and smartwatches, or in intelligent automotive systems. LPDDR5's clocking scheme uses a dual-clock signal design, unlike the single-clock signal design used in DDR5. This helps improve the effective data bus speed, but also brings challenges in state alignment and synchronization. Specifically, LPDDR5 defines a command address (CA) bus and a clock (CK) signal. The command address bus is used for communication from the host to the device, and the clock signal is used to set the transfer rate from the host's CPU to the DDR. Prior to LPDDR5, a single clock signal was used, and the command address bus was a single-data-rate (SDR) bus. Therefore, in clock schemes prior to LPDDR5, the clock signal and the data strobe (DQS) signal had the same frequency. The maximum transfer rate of the command address bus corresponded to the maximum operating rate of the clock signal and DQS signal at their highest frequencies, while the maximum transfer rate of the data (DQ) bus was twice that of the command address bus (two data packets are transferred per clock cycle of the DDR bus). For example, if the maximum frequency of the clock signal was 2133 MHz, the maximum frequency of the DQS signal was 2133 MHz, the maximum effective rate of the command address bus was 2133 megabits per second (Mbps), and the maximum effective rate of the data bus was 4266 Mbps. LPDDR5 introduced a dual-clock signal design, providing two signals: a write clock (WCK) signal from the host to the device, and a read data (RDQS) signal from the device to the host. The LPDDR5 clock scheme decouples the data strobe signal, i.e., the RDQS signal, from the clock signal, which allows the clock signal frequency to be significantly lower than the RDQS signal frequency.For example, the maximum frequency of the clock signal is 800MHz, the maximum frequency of the RDQS signal is 3200MHz, the maximum effective rate of the command address bus is 1600Mbps, and the maximum effective rate of the data bus is 6400Mbps. Therefore, the clock scheme used in low-power fifth-generation double-rate and related memory technology specifications decouples the clock signal and data strobe signal, thus decoupling the clock signal and write clock signal. The frequency of the write clock signal can be two or four times the frequency of the clock signal. However, this also brings challenges in state alignment and synchronization. Furthermore, the routing between memory chips, such as dynamic random access memory (DRAM) chips, and the command address bus and data bus of the system chip may not be one-to-one. Therefore, the impact of the pin layout and routing after chip packaging on command bus training needs to be considered. Figure 1 The command bus training method shown, through hardware optimization, provides a command address remapping module, a data remapping module, and a delay training module. Through software optimization, the delay training module, based on a center point algorithm, adjusts the delay of the command address signal, achieving alignment between the command address signal and the data signal. This not only meets the requirements of low-power fifth-generation double-rate memory technology specifications in terms of state alignment and synchronization, but also effectively overcomes the adverse effects of pin layout and routing after chip packaging, effectively overcomes the adverse effects of metastable regions, and quickly and reliably determines the alignment relationship and locates the optimal center position, improving training effectiveness and reducing overall training latency, thus contributing to rapid response and device adaptation. The following section combines... Figure 1 Further details.

[0031] refer to Figure 1In step S101, the command address remapping module maps the multiple command address bus pins of the memory module to be trained to the standard command address bus pins of the standard memory module. In step S103, the data remapping module maps the multiple data bus pins of the memory module to be trained to the standard data bus pins of the standard memory module. The multiple command address bus pins of the memory module to be trained and the multiple data bus pins of the memory module to be trained satisfy a first connection relationship, and the standard command address bus pins of the standard memory module and the standard data bus pins of the standard memory module satisfy a second connection relationship. During command bus training, when it is necessary to determine the delay of a certain command address bus, it is necessary to determine which data signal line the feedback result comes from, and thus adjust the circuit delay of the command address bus according to the feedback result. Moreover, when finding the boundary of the address bus signal, it is necessary to consider the uncertainty of electronic circuit signals in boundary delay sampling. For example, when the clock signal is used to sample the command address signal, there are metastable intervals on the left and right boundaries, which will lead to inaccurate positioning of the left and right boundaries, and may result in deviation in the positioning of the optimal center point. Furthermore, the memory modules to be trained may employ various memory technology specifications, with different product models and packaging specifications. Therefore, the number and distribution of pins on the memory modules to be trained can be complex and varied. For example, LPDDR5 generally specifies 7 command address bus pins, while Internet Protocol version 4 (IPv4) generally specifies 8 command address bus pins. As another example, dedicated processors may have customized requirements for the pin distribution and bus design of their memory modules, such as graphics processing units (GPUs) and data processing units (DPUs). To address this, through hardware optimization, command address remapping modules and data remapping modules are provided. This is equivalent to mapping any pin distribution and packaging design specification on the memory modules to a hypothetical standard memory module. This includes mapping multiple command address bus pins of the memory modules to standard command address bus pins of the standard memory module through the command address remapping module, and mapping multiple data bus pins of the memory modules to standard data bus pins of the standard memory module through the data remapping module. Thus, the signal transmission and alignment problem that the training memory module needs to solve through the command bus training is transformed into the signal transmission and alignment problem of the standard memory module. This can be achieved through software optimization, by using a delay training module and adjusting the delay of the command address signal based on the center point algorithm, thereby realizing the alignment between the command address signal and the data signal.To address the inconsistency between command address bus pins and data bus pins, the aforementioned command address remapping module and data remapping module can overcome this issue through a first connection relationship and a second connection relationship. For example, for a specific command address bus CA0, it is first determined that this specific command address bus CA0 is connected to a standard command address bus pin (such as CAX) of the standard memory module through the command address remapping module. This standard command address bus pin (CAX) provides feedback through a standard data bus pin (such as DQX). After remapping by the data remapping module, the standard data bus pin (DQX) ultimately provides asynchronous signal feedback, which may be DQ0 or DQ5. The command address remapping module and data remapping module can flexibly adapt to the first connection relationship and can also combine with the second connection relationship to optimize the subsequent center point algorithm. Therefore, this not only facilitates adaptation to various product models and packaging specifications of the memory modules to be trained but also simplifies the difficulty of hardware and software co-design. Software-level optimization can be performed on the second connection relationship and the standard memory module, using various sampling optimization algorithms such as the center point algorithm to improve the training effect of the standard memory module. In addition, the delay training module can be implemented using programmable devices, firmware, or circuit algorithms.

[0032] Continue to refer to Figure 1In step S105, the delay training module determines the asynchronous signal feedback source corresponding to each of the multiple command address bus pins of the storage module to be trained being invoked, based on the first and second connection relationships. Then, the delay of the command address signals corresponding to each of the multiple command address bus pins of the storage module to be trained is adjusted based on the center point algorithm, thereby achieving alignment between the command address signals corresponding to each of the multiple command address bus pins of the storage module to be trained and the data signals corresponding to each of the multiple data bus pins of the storage module to be trained. Thus, based on the above-mentioned construction of a hypothetical standard storage module through the command address remapping module and the data remapping module, the training result of the command bus of the standard storage module can be used for signal transmission and alignment of the storage module to be trained. That is, the actual first connection relationship of the storage module to be trained is transformed into the second connection relationship of the standard storage module. Then, based on the first and second connection relationships, the asynchronous signal feedback source corresponding to each of the multiple command address bus pins of the storage module to be trained being invoked can be conveniently determined. For example, for a specific command address bus CA0, it is first determined that this specific command address bus CA0 is connected to a standard command address bus pin (such as CAX) of the standard memory module through a command address remapping module. This standard command address bus pin (CAX) provides feedback through a standard data bus pin (such as DQX). This standard data bus pin (DQX) is remapped by a data remapping module, and ultimately, the asynchronous signal feedback is provided by DQ0, or possibly DQ5. This means that through the command address remapping module and the data remapping module, the first connection relationship can be flexibly adapted, and the second connection relationship can be combined to optimize the subsequent center point algorithm. Therefore, it is not only beneficial for adapting to various product models and packaging specifications of the memory modules to be trained, but also simplifies the difficulty of hardware and software co-design. Software-level optimization can be performed on the second connection relationship and the standard memory module. In this way, the center point algorithm can better optimize for the deviation of the optimal center point position caused by the existence of metastable regions, which helps to improve training efficiency and training accuracy. Thus, through hardware optimization, a command address remapping module, a data remapping module, and a delay training module are provided. Through software optimization, the delay training module adjusts the delay of the command address signal based on the center point algorithm, achieving alignment between the command address signal and the data signal. This not only meets the requirements of low-power fifth-generation double-rate and related memory technology specifications in terms of state alignment and synchronization, but also effectively overcomes the adverse effects of pin layout and wiring after chip packaging, effectively overcomes the adverse effects of the existence of metastable regions, quickly and reliably determines the alignment relationship and locates the optimal center position, improves training effect and reduces overall training latency, and helps with rapid response and device adaptation.

[0033] Figure 2 This is a schematic diagram of a command bus training device provided in an embodiment of this application. Figure 2 As shown, the command bus training device 200 is applied to the memory module 201 to be trained. The command bus training device 200 includes a command address remapping module 211, a data remapping module 213, and a delay training module 215. The method for training the command bus associated with the memory module 201 using the command bus training device 200 includes: mapping multiple command address bus pins of the memory module 201 to standard command address bus pins of a standard memory module using the command address remapping module 211; and mapping multiple data bus pins of the memory module 201 to standard data bus pins of the standard memory module using the data remapping module 213. The multiple command address bus pins and multiple data bus pins of the memory module 201 satisfy a first connection relationship, and the standard command address bus pins of the standard memory module... The address bus pins and the standard data bus pins of the standard storage module are made to satisfy a second connection relationship. Based on the first connection relationship and the second connection relationship, the delay training module 215 determines the asynchronous signal feedback source corresponding to each of the multiple command address bus pins of the storage module 201 to be trained when they are called. Then, the delay of the command address signals corresponding to each of the multiple command address bus pins of the storage module 201 to be trained is adjusted based on the center point algorithm, thereby achieving alignment between the command address signals corresponding to each of the multiple command address bus pins of the storage module 201 to be trained and the data signals corresponding to each of the multiple data bus pins of the storage module 201 to be trained.

[0034] Figure 2 The command bus training device 200 shown provides a command address remapping module 211, a data remapping module 213, and a delay training module 215 through hardware optimization, and a delay training module 215 through software optimization. The delay training module 215 adjusts the delay of the command address signal based on the center point algorithm, realizing the alignment between the command address signal and the data signal. This not only meets the requirements of low-power fifth-generation double-rate and related memory technology specifications in terms of state alignment and synchronization, but also effectively overcomes the adverse effects of pin layout and wiring after chip packaging, effectively overcomes the adverse effects of the existence of metastable regions, quickly and reliably determines the alignment relationship and locates the optimal center position, improves the training effect and reduces the overall training latency, and helps with rapid response and device adaptation.

[0035] refer to Figure 1 and Figure 2In one possible implementation, the plurality of command address bus pins of the storage module to be trained includes a first command address bus pin. This first command address bus pin, based on the first connection relationship, is connected to a first data bus pin among the plurality of data bus pins of the storage module to be trained. The delay training module adjusts the delay of the first command address signal corresponding to the first command address bus pin based on the center point algorithm, thereby achieving alignment between the first command address signal corresponding to the first command address bus pin and the first data signal corresponding to the first data bus pin. Here, the first command address bus pin is any command address bus among the plurality of command address bus pins of the storage module to be trained. Based on the assumption of a standard storage module constructed through the command address remapping module and the data remapping module, the training result of the command bus of the standard storage module can be used for signal transmission and alignment of the storage module to be trained. That is, the actual first connection relationship of the storage module to be trained is converted into the second connection relationship of the standard storage module. Then, based on the first and second connection relationships, the asynchronous signal feedback source corresponding to each of the plurality of command address bus pins of the storage module to be trained being invoked can be conveniently determined. For example, for a specific command address bus CA0, it is first determined that this specific command address bus CA0 is connected to a standard command address bus pin (such as CAX) of the standard memory module through a command address remapping module. This standard command address bus pin (CAX) provides feedback through a standard data bus pin (such as DQX). This standard data bus pin (DQX) is remapped by a data remapping module, and ultimately, the asynchronous signal feedback is provided by DQ0, or possibly DQ5. This means that through the command address remapping module and the data remapping module, the first connection relationship can be flexibly adapted, and the second connection relationship can be combined to optimize the subsequent center point algorithm. Therefore, it is not only beneficial for adapting to various product models and packaging specifications of the memory modules to be trained, but also simplifies the difficulty of hardware and software co-design. Software-level optimization can be performed on the second connection relationship and the standard memory module. In this way, the center point algorithm can better optimize for the deviation of the optimal center point position caused by the existence of metastable regions, which helps to improve training efficiency and training accuracy.

[0036] In some embodiments, the delay training module adjusts the delay of the first command address signal corresponding to the first command address bus pin based on the center point algorithm, thereby achieving alignment between the first command address signal corresponding to the first command address bus pin and the first data signal corresponding to the first data bus pin. This includes: the delay training module determining the left boundary of the first command address signal using a first function of the center point algorithm, and determining the right boundary of the first command address signal using a second function of the center point algorithm; then determining the center point position of the first command address signal based on the determined left and right boundaries; and finally adjusting the delay of the first command address signal based on the determined center point position. Thus, by providing the first and second functions, optimization can be performed for both left and right boundary cases respectively, thereby better determining stable boundaries and ultimately determining the optimal center point position, improving training effectiveness.

[0037] In some embodiments, the first function of the center point algorithm is to sequentially scan the error position, the metastable position, and the stable position of the left boundary of the first command address signal. A second function of the center point algorithm is to sequentially scan the error position, the metastable position, and the stable position of the right boundary of the first command address signal. Thus, by scanning from the error position to the metastable position and then to the stable position, it means starting from both sides and moving towards the stable state of the optimal center point position to find the stable positions of the left and right boundaries. The stable positions of the left and right boundaries found in this way, along with the further determined center point position, are accurate. Therefore, better determination of stable boundaries and thus the optimal center point position improves training effectiveness.

[0038] In some embodiments, the delay training module determines the left boundary of the first command address signal using the first function of the center point algorithm, including: increasing the delay of the first command address signal when the first data signal feedback is low and the clock signal has not acquired the first command address signal; and maintaining the delay of the first command address signal when the first data signal feedback is high and the clock signal has acquired the first command address signal. This achieves a scanning sequence from the erroneous position to the metastable position and then to the stable position. This means that starting from both sides, the stable positions of the left and right boundaries are searched towards the stable state of the optimal center point position. The stable positions of the left and right boundaries found in this way, along with the further determined center point position, are accurate. Therefore, through the delay training module, the algorithm for finding the stable position of the left boundary increases or maintains the circuit delay based on the asynchronous feedback result, better determining the stable boundary and thus the optimal center point position, improving the training effect.

[0039] In some embodiments, the delay training module determines the right boundary of the first command address signal using the second function of the center point algorithm, including: reducing the delay of the first command address signal when the first data signal feedback is low and the clock signal has not acquired the first command address signal; and maintaining the delay of the first command address signal when the first data signal feedback is high and the clock signal has acquired the first command address signal. This achieves a scanning sequence from the error position to the metastable position and then to the stable position. This means that starting from both sides, the stable positions of the left and right boundaries are searched towards the stable state of the optimal center point position. The stable positions of the left and right boundaries found in this way, along with the further determined center point position, are accurate. Therefore, through the delay training module, the algorithm for finding the stable position of the right boundary reduces or maintains circuit delay based on asynchronous feedback results, better determines the stable boundary, and thus determines the optimal center point position, improving the training effect.

[0040] In some embodiments, the delay training module, using the first function of the center point algorithm to determine the left boundary of the first command address signal, further includes: initially setting the delay of the first command address signal to a first initial delay, and then performing a first loop search number using the first function of the center point algorithm; the delay training module, using the second function of the center point algorithm to determine the right boundary of the first command address signal, further includes: initially setting the delay of the first command address signal to a second initial delay, and then performing a second loop search number using the first function of the center point algorithm, wherein the first initial delay is less than the second initial delay. Thus, by distinguishing the steady-state positions of the left and right boundaries using a search algorithm, setting the first initial delay as a smaller circuit delay and setting the second initial delay as a larger circuit delay, stable boundaries are better determined, thereby determining the optimal center point position and improving training effectiveness.

[0041] In some embodiments, both the first and second loop search counts are determined based on the maximum frequency of the clock signal. The first initial delay corresponds to a single step size unit, which is determined based on the maximum frequency of the clock signal. The second initial delay corresponds to multiple step size units. Here, the maximum frequency of the clock signal determines the number of loop searches; the higher the maximum frequency of the clock signal, the more loop searches are performed because the step size becomes smaller, but the delay of the address bus signal relative to the clock signal does not decrease. Additionally, the maximum rate of the command address bus corresponds to the frequency of the command address signal, and protocols generally specify that the frequency of the command address signal is consistent with the maximum frequency of the clock signal. Regarding the initial delay, a first initial delay is set as a smaller circuit delay, and a second initial delay is set as a larger circuit delay. The first initial delay is used for the algorithm to find the steady-state position at the left boundary. The smaller value of the first initial delay corresponds to a step size unit, such as 1 unit interval (UI) or 0.1 UI. The second initial delay is used for the algorithm to find the steady-state position at the right boundary. The larger value of the second initial delay corresponds to multiple step size units, such as 7 UI. Here, the UI is determined based on the maximum frequency of the clock signal, and a step size can be set to 64 UIs or 32 UIs. Additionally, if the command address signal has a large clock skew relative to the clock signal, a larger scan range is required. The scan range is determined by the initial delay. Therefore, in some embodiments, when the command address signal has a large clock skew relative to the clock signal, a larger scan range can be obtained by setting a first initial delay and a second initial delay to offset the adverse effects of the large clock skew. This allows for better determination of stable boundaries and thus the optimal center point position, improving training effectiveness.

[0042] In one possible implementation, the multiple command address bus pins of the memory module to be trained include command address bus pins for a first data channel and command address bus pins for a second data channel. The standard command address bus pins of the standard memory module corresponding to the command address bus pins for the first data channel are different from the standard command address bus pins of the standard memory module corresponding to the command address bus pins for the second data channel. Thus, through hardware optimization, a command address remapping module, a data remapping module, and a delay training module are provided. Through software optimization, the delay training module adjusts the delay of the command address signal based on a center point algorithm, achieving alignment between the command address signal and the data signal. This not only meets the requirements of low-power fifth-generation double-rate and related memory technology specifications in terms of state alignment and synchronization, but also effectively overcomes the adverse effects of pin layout and wiring after chip packaging, effectively overcomes the adverse effects of the existence of metastable regions, quickly and reliably determines the alignment relationship and locates the optimal center position, improves training effect and reduces overall training latency, and facilitates rapid response and device adaptation.

[0043] In one possible implementation, the multiple data bus pins of the memory module to be trained include paired data bus pins for differential signal transmission, and the standard data bus pins of the standard memory module corresponding to the paired data bus pins are also used for differential signal transmission. Thus, through hardware optimization, a command address remapping module, a data remapping module, and a delay training module are provided. Through software optimization, the delay training module, based on a center point algorithm, adjusts the delay of the command address signal, achieving alignment between the command address signal and the data signal. This not only meets the requirements of low-power fifth-generation double-rate and related memory technology specifications in terms of state alignment and synchronization, but also effectively overcomes the adverse effects of pin layout and wiring after chip packaging, effectively overcomes the adverse effects of the existence of metastable regions, and quickly and reliably determines the alignment relationship and locates the optimal center position, improving training effect and reducing overall training latency, thus facilitating rapid response and device adaptation.

[0044] In one possible implementation, the first and second connection relationships are determined based on the computer memory specifications associated with the memory module to be trained. The number and distribution of pins on the memory module to be trained can be complex and varied. For example, LPDDR5 typically specifies 7 command address bus pins, while Internet Protocol version 4 (IPv4) typically specifies 8 command address bus pins. Furthermore, dedicated processors may have customized requirements for the pin distribution and bus design on their memory modules, such as graphics processing units (GPUs) and data processing units (DPUs). To address this, through hardware optimization, command address remapping modules and data remapping modules are provided. These are equivalent to mapping any pin distribution and package design specification on the memory module to a hypothetical standard memory module. This includes mapping multiple command address bus pins of the memory module to standard command address bus pins of the standard memory module through the command address remapping module, and mapping multiple data bus pins of the memory module to standard data bus pins of the standard memory module through the data remapping module. Thus, the signal transmission and alignment problem that the training memory module needs to solve through the command bus training is transformed into the signal transmission and alignment problem of the standard memory module. This can be achieved through software optimization, by using a delay training module and adjusting the delay of the command address signal based on the center point algorithm, thereby realizing the alignment between the command address signal and the data signal.

[0045] In some embodiments, the computer memory specifications are low-power 5th generation double speed, Internet Protocol version 4, data processing unit design specifications, or image processing unit design specifications. Thus, through the command address remapping module and data remapping module, the first connection relationship can be flexibly adapted, and the second connection relationship can be combined to optimize the subsequent centroid algorithm. Therefore, this not only facilitates adaptation to various product models and packaging specifications of the storage modules to be trained, but also simplifies the difficulty of hardware and software co-design, allowing software-level optimization to be performed on the second connection relationship and the standard storage module, improving the training effect of the standard storage module through various sampling optimization algorithms such as the centroid algorithm.

[0046] Figure 3This is a schematic diagram of a computing device 300 provided in an embodiment of this application. The computing device 300 includes one or more processors 310, a communication interface 320, and a memory 330. The processors 310, communication interface 320, and memory 330 are interconnected via a bus 340. Optionally, the computing device 300 may further include an input / output interface 350, which is connected to input / output devices for receiving user-set parameters, etc. The computing device 300 can be used to implement some or all of the functions of the device embodiment or system embodiment in the above-described embodiments of this application; the processor 310 can also be used to implement some or all of the operation steps of the method embodiment in the above-described embodiments of this application. For example, the specific implementation of various operations performed by the computing device 300 can be referred to the specific details in the above embodiments, such as the processor 310 being used to execute some or all of the steps or operations in the above-described method embodiments. For example, in the embodiments of this application, the computing device 300 can be used to implement some or all of the functions of one or more components in the above-described device embodiments. In addition, the communication interface 320 can be used specifically for communication functions necessary to implement the functions of these devices and components, and the processor 310 can be used specifically for processing functions necessary to implement the functions of these devices and components.

[0047] It should be understood that, Figure 3 The computing device 300 may include one or more processors 310, and the multiple processors 310 may collaboratively provide processing power in a parallel connection mode, a serial connection mode, a serial-parallel connection mode, or an arbitrary connection mode; or the multiple processors 310 may form a processor sequence or a processor array; or the multiple processors 310 may be divided into a main processor and an auxiliary processor; or the multiple processors 310 may have different architectures, such as adopting a heterogeneous computing architecture. Furthermore, Figure 3 The structural and functional descriptions of the computing device 300 shown are exemplary and non-limiting. In some exemplary embodiments, the computing device 300 may include... Figure 3 The diagram shows more or fewer components, or combinations of some components, or splitting of some components, or different arrangements of components.

[0048] The processor 310 can have various specific implementations. For example, it may include one or more combinations of a central processing unit (CPU), a graphics processing unit (GPU), a neural network processing unit (NPU), a tensor processing unit (TPU), or a data processing unit (DPU). This application does not impose specific limitations on these embodiments. The processor 310 can also be a single-core or multi-core processor. The processor 310 can be a combination of a CPU and hardware chips. These hardware chips can be application-specific integrated circuits (ASICs), programmable logic devices (PLDs), or combinations thereof. The PLDs can be complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), generic array logic (GALs), or any combination thereof. The processor 310 can also be implemented using logic devices with built-in processing logic, such as FPGAs or digital signal processors (DSPs). The communication interface 320 can be a wired interface or a wireless interface, used to communicate with other modules or devices. The wired interface can be an Ethernet interface, a local interconnect network (LIN), etc., and the wireless interface can be a cellular network interface or a wireless LAN interface, etc.

[0049] Memory 330 may be non-volatile memory, such as read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Memory 330 may also be volatile memory, which may be random access memory (RAM) used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). The memory 330 can also be used to store program code and data, so that the processor 310 can call the program code stored in the memory 330 to execute some or all of the operation steps in the above method embodiments, or to execute the corresponding functions in the above device embodiments. Furthermore, the computing device 300 may include, compared to... Figure 3 The number of components displayed may be more or less, or there may be different component configurations.

[0050] Bus 340 can be a Peripheral Component Interconnect Express (PCIe) bus, or an Extended Industry Standard Architecture (EISA) bus, a Unified Bus (Ubus or UB), a Compute Express Link (CXL) bus, a Cache Coherent Interconnect for Accelerators (CCIX) bus, etc. Bus 340 can be divided into address bus, data bus, control bus, etc. In addition to the data bus, bus 340 can also include a power bus, a control bus, and a status signal bus. However, for clarity, Figure 3 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0051] The methods and devices provided in this application are based on the same inventive concept. Since the principles by which the methods and devices solve problems are similar, the embodiments, implementation methods, examples, or methods of implementation of the methods and devices can be referred to each other, and repeated details will not be repeated. This application also provides a system comprising multiple computing devices, the structure of each computing device of which can refer to the structure of the computing devices described above. The functions or operations achievable by this system can refer to the specific implementation steps in the above method embodiments and / or the specific functions described in the above device embodiments, and will not be repeated here.

[0052] This application also provides a computer-readable storage medium storing computer instructions. When these computer instructions are executed on a computer device (such as one or more processors), they can implement the method steps described in the above method embodiments. The specific implementation of the above method steps by the processor of the computer-readable storage medium can refer to the specific operations described in the above method embodiments and / or the specific functions described in the above device embodiments, and will not be repeated here.

[0053] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. This application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Embodiments of this application can be implemented wholly or partially by software, hardware, firmware, or any other combination. When implemented in software, the above embodiments can be implemented wholly or partially as a computer program product. This application can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. Computer-readable storage media can be any available medium that a computer can access, or a data storage device such as a server or data center that contains one or more sets of available media. Available media can be magnetic media (such as floppy disks, hard disks, and magnetic tapes), optical media, or semiconductor media. Semiconductor media can be solid-state drives, random access memory, flash memory, read-only memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, registers, or any other suitable form of storage medium.

[0054] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. Each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0055] In the above embodiments, the descriptions of each embodiment have their own emphasis. Parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the spirit and scope of the embodiments of this application. The steps in the methods of the embodiments of this application can be adjusted in order, combined, or deleted according to actual needs; the modules in the systems of the embodiments of this application can be divided, combined, or deleted according to actual needs. If these modifications and variations of the embodiments of this application fall within the scope of the claims of this application and their equivalents, then this application also intends to include these modifications and variations.

Claims

1. A command bus training method, characterized in that, The command bus training method includes: The command address remapping module maps multiple command address bus pins of the memory module to be trained to the standard command address bus pins of the standard memory module. Through the data remapping module, multiple data bus pins of the storage module to be trained are mapped one by one to the standard data bus pins of the standard storage module. The multiple command address bus pins of the storage module to be trained and the multiple data bus pins of the storage module to be trained satisfy a first connection relationship, and the standard command address bus pins of the standard storage module and the standard data bus pins of the standard storage module satisfy a second connection relationship. The delay training module determines the asynchronous signal feedback source corresponding to each of the multiple command address bus pins of the storage module under training when they are called, based on the first connection relationship and the second connection relationship. Then, the delay of the command address signal corresponding to each of the multiple command address bus pins of the storage module under training is adjusted based on the center point algorithm, thereby achieving alignment between the command address signal corresponding to each of the multiple command address bus pins of the storage module under training and the data signal corresponding to each of the multiple data bus pins of the storage module under training.

2. The command bus training method according to claim 1, characterized in that, The multiple command address bus pins of the storage module to be trained include a first command address bus pin. Based on the first connection relationship, the first command address bus pin is connected to the first data bus pin among the multiple data bus pins of the storage module to be trained. The delay training module adjusts the delay of the first command address signal corresponding to the first command address bus pin based on the center point algorithm, thereby achieving alignment between the first command address signal corresponding to the first command address bus pin and the first data signal corresponding to the first data bus pin.

3. The command bus training method according to claim 2, characterized in that, The delay training module adjusts the delay of the first command address signal corresponding to the first command address bus pin based on the center point algorithm, thereby achieving alignment between the first command address signal corresponding to the first command address bus pin and the first data signal corresponding to the first data bus pin, including: The delay training module uses the first function of the center point algorithm to determine the left boundary of the first command address signal and the second function of the center point algorithm to determine the right boundary of the first command address signal. Then, based on the determined left and right boundaries of the first command address signal, it determines the center point position of the first command address signal. Finally, it adjusts the delay of the first command address signal based on the determined center point position of the first command address signal.

4. The command bus training method according to claim 3, characterized in that, The first function of the center point algorithm is to sequentially scan the error position of the left boundary of the first command address signal, the metastable position of the left boundary of the first command address signal, and the stable position of the left boundary of the first command address signal. The second function of the center point algorithm is to sequentially scan the error position of the right boundary of the first command address signal, the metastable position of the right boundary of the first command address signal, and the stable position of the right boundary of the first command address signal.

5. The command bus training method according to claim 3, characterized in that, The delayed training module uses the first function of the center point algorithm to determine the left boundary of the first command address signal, including: When the first data signal feedback is low and the clock signal does not acquire the first command address signal, the delay of the first command address signal is increased; when the first data signal feedback is high and the clock signal acquires the first command address signal, the delay of the first command address signal is maintained.

6. The command bus training method according to claim 5, characterized in that, The delayed training module uses the second function of the center point algorithm to determine the right boundary of the first command address signal, including: When the first data signal feedback is low and the clock signal does not acquire the first command address signal, the delay of the first command address signal is reduced; when the first data signal feedback is high and the clock signal acquires the first command address signal, the delay of the first command address signal is maintained.

7. The command bus training method according to claim 6, characterized in that, The delay training module, which uses the first function of the center point algorithm to determine the left boundary of the first command address signal, further includes: initially setting the delay of the first command address signal to a first initial delay, and then using the first function of the center point algorithm to perform a first loop search number; the delay training module, which uses the second function of the center point algorithm to determine the right boundary of the first command address signal, further includes: initially setting the delay of the first command address signal to a second initial delay, and then using the first function of the center point algorithm to perform a second loop search number, wherein the first initial delay is less than the second initial delay.

8. The command bus training method according to claim 7, characterized in that, The first and second loop search counts are both determined based on the maximum frequency of the clock signal. The first initial delay corresponds to a single step size unit, which is determined based on the maximum frequency of the clock signal. The second initial delay corresponds to multiple step size units.

9. The command bus training method according to claim 1, characterized in that, The multiple command address bus pins of the storage module to be trained include a command address bus pin for the first data channel and a command address bus pin for the second data channel. The standard command address bus pin of the standard storage module corresponding to the command address bus pin for the first data channel is different from the standard command address bus pin of the standard storage module corresponding to the command address bus pin for the second data channel.

10. The command bus training method according to claim 1, characterized in that, The multiple data bus pins of the storage module to be trained include paired data bus pins, which are used for differential signal transmission. The standard data bus pins of the standard storage module corresponding to the paired data bus pins are also used for differential signal transmission.

11. The command bus training method according to claim 1, characterized in that, The first connection relationship and the second connection relationship are determined based on the computer memory specifications associated with the storage module to be trained.

12. The command bus training method according to claim 11, characterized in that, The computer memory specifications are low-power 5th generation double speed, Internet Protocol version 4, data processing unit design specifications, or image processing unit design specifications.

13. An electronic device, characterized in that, The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method according to any one of claims 1 to 12.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed on a computer device, cause the computer device to perform the method according to any one of claims 1 to 12.

15. A command bus training device, characterized in that, The command bus training device is applied to the storage module to be trained. The command bus training device includes a command address remapping module, a data remapping module, and a delay training module. The method for training the command bus associated with the storage module to be trained using the command bus training device includes: The command address remapping module maps the multiple command address bus pins of the memory module to be trained to the standard command address bus pins of the standard memory module. Through the data remapping module, multiple data bus pins of the storage module to be trained are mapped one by one to the standard data bus pins of the standard storage module. The multiple command address bus pins of the storage module to be trained and the multiple data bus pins of the storage module to be trained satisfy a first connection relationship, and the standard command address bus pins of the standard storage module and the standard data bus pins of the standard storage module satisfy a second connection relationship. The delay training module determines the asynchronous signal feedback source corresponding to each of the multiple command address bus pins of the storage module under training being invoked based on the first connection relationship and the second connection relationship. Then, the delay of the command address signals corresponding to each of the multiple command address bus pins of the storage module under training is adjusted based on the center point algorithm, thereby achieving alignment between the command address signals corresponding to each of the multiple command address bus pins of the storage module under training and the data signals corresponding to each of the multiple data bus pins of the storage module under training.

Citation Information

Patent Citations

  • Method of data and address sharing pin for self-adaptively adjusting memory access granularity

    CN103246625A

  • Training for mapping swizzled data command / address signals

    CN104903877A