Storage controller, storage system, computer device, and storage access method
By tightly integrating the random access memory chip and the controller chip, zero-copy data interaction and high-speed address mapping are achieved, improving the data throughput efficiency of the storage controller, solving the performance deficiency of traditional storage architectures in large-scale neural network models, and meeting the needs of high-performance storage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-05-08
AI Technical Summary
Traditional storage architectures struggle to meet the stringent high-performance storage requirements of large-scale neural network models, particularly in terms of bandwidth, latency, and concurrent processing capabilities.
A storage controller architecture is provided, which achieves zero-copy data interaction by tightly integrating the random access memory chip and the controller chip and sharing address encoding. The logical address mapping table resides in the first random access memory space to support high-speed query and update. The command queue is distributed in the second random access memory space to support multi-channel parallel instruction scheduling. The third random access memory space is used as a data transfer buffer to coordinate the flow of read and write data.
It significantly reduces data access latency, improves overall response speed and throughput performance, reduces system power consumption, and is suitable for high-concurrency, low-latency artificial intelligence computing scenarios, supporting the acceleration of inference/training of large-scale neural network models.
Smart Images

Figure CN121996156A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a storage controller, storage system, computer device, and storage access method. Background Technology
[0002] With the rapid development of artificial intelligence technology, the number of parameters and computational demands of large-scale neural network models are growing exponentially. Current large-scale neural network models generate massive amounts of random access requests during training and inference, placing extremely high demands on the bandwidth, latency, and concurrent processing capabilities of storage systems. However, traditional storage architectures struggle to meet the stringent high-performance storage requirements of large-scale neural network models. Summary of the Invention
[0003] This invention provides a storage controller, storage system, computer device, and storage access method that can meet the stringent requirements of large-scale neural network models for high-bandwidth, low-latency storage access.
[0004] In a first aspect, embodiments of the present invention provide a storage controller, comprising: Controller chip; Random access memory (RAM) chips are bonded to controller chips and use the same address encoding rules as the controller chips. The RAM chips include: The first random access memory space is used to store the logical address mapping table; The second random access memory space is used to store command queues corresponding to different memory channels. Each memory channel is connected to at least one memory chip. The command queue is used to cache operation commands issued by the controller chip to the memory chip. The third random access memory space is used to store read data returned by the memory chip in response to operation commands, or to store write data that the memory chip needs to write in response to operation commands.
[0005] In a second aspect, embodiments of the present invention provide a storage system, including: The storage controller provided in this embodiment of the invention; Multiple memory chips are connected to the memory controller via multiple memory channels, wherein each memory channel is connected to at least one memory chip.
[0006] Thirdly, embodiments of the present invention provide a computer device, including: Computational unit; The storage system provided in this embodiment of the invention.
[0007] Fourthly, embodiments of the present invention provide a storage access method applicable to a storage controller. The storage controller includes a controller chip and a random access memory (RAM) chip. The RAM chip is bonded to the controller chip and uses the same address encoding rule as the controller chip. The RAM chip includes a first random access memory space, a second random access memory space, and a third random access memory space. The storage access method includes: The controller chip receives storage access requests from the computing unit. According to the logical address mapping table stored in the first random access memory space, the controller chip splits the storage access request into multiple operation commands and writes them into the command queues corresponding to different storage channels stored in the second random access memory space. Based on the third random access memory space, the controller chip sends operation commands to the connected memory chip via the storage channel for execution.
[0008] This invention provides a novel storage controller architecture, comprising: a controller chip; and a random access memory (RAM) chip, bonded to the controller chip and using the same address encoding rule as the controller chip. The RAM chip includes: a first RAM space for storing a logical address mapping table; a second RAM space for storing command queues corresponding to different storage channels, each storage channel connecting to at least one storage chip, the command queues used to cache operation commands issued by the controller chip to the storage chips; and a third RAM space for storing read data returned by the storage chips in response to operation commands, or write data required by the storage chips in response to operation commands. Thus, by tightly integrating the RAM chip and the controller chip and sharing address encoding, zero-copy data interaction between the controller chip and the RAM chip is achieved, significantly reducing data access latency. Meanwhile, by residing the logical address mapping table in the first random access memory (RAM) space, high-speed address mapping lookup and updates are achieved; distributing the command queue in the second RAM space enables multi-channel parallel instruction scheduling, enhancing concurrent processing capabilities; and using the third RAM space as a data transfer buffer effectively coordinates the efficient flow of read and write data between the controller and the memory chip, thereby comprehensively improving overall response speed and throughput performance. This memory controller architecture significantly improves data throughput efficiency and reduces overall system power consumption through tight hardware-level coupling, making it suitable for high-concurrency, low-latency artificial intelligence computing scenarios. In typical scenarios, this memory controller architecture can support accelerated inference / training of large-scale neural network models such as the Transformer architecture, meeting the stringent high-performance storage requirements of large-scale neural network models. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 This is a schematic diagram of the structure of a storage controller provided in an embodiment of the present invention; Figure 2 yes Figure 1 A schematic diagram of the internal control core of the controller chip 110; Figure 3 This is an example diagram showing the connection between the controller chip 110 and the memory chip; Figure 4 This is a detailed structural diagram of the controller chip 110 in an embodiment of the present invention; Figure 5 This is another structural schematic diagram of the storage controller provided in an embodiment of the present invention; Figure 6 This is another structural schematic diagram of the storage controller provided in an embodiment of the present invention; Figure 7 This is a schematic diagram of the structure of the storage system provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of the structure of the computer device provided in an embodiment of the present invention; Figure 9 This is a flowchart illustrating the storage access method provided in an embodiment of the present invention. Detailed Implementation
[0011] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.
[0012] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0013] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0014] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0015] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0016] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0017] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0018] This invention provides a storage controller, a storage system, a computer device, and a storage access method. The storage controller includes: a controller chip; and a random access memory (RAM) chip, bonded to the controller chip and using the same address encoding rule as the controller chip. The RAM chip includes: a first RAM space for storing a logical address mapping table; a second RAM space for storing command queues corresponding to different storage channels, each storage channel being connected to at least one storage chip, the command queues being used to cache operation commands issued by the controller chip to the storage chips; and a third RAM space for storing read data returned by the storage chip in response to operation commands, or storing write data required by the storage chip in response to operation commands.
[0019] Please refer to Figure 1 This is a schematic diagram of the structure of a storage controller disclosed in an embodiment of the present invention, as shown below. Figure 1 As shown, the storage controller may include: Controller chip 110; Random access memory (RAM) chip 120 is bonded to controller chip 110 and uses the same address encoding rule as the controller chip. The RAM chip includes: The first random access memory space is used to store the logical address mapping table; The second random access memory space is used to store command queues corresponding to different memory channels. Each memory channel is connected to at least one memory chip. The command queue is used to cache operation commands issued by the controller chip to the memory chip. The third random access memory space is used to store read data returned by the memory chip in response to operation commands, or to store write data that the memory chip needs to write in response to operation commands.
[0020] It should be noted that the storage controller is the core component responsible for managing the read, write, erase, bad block management, and data error correction operations of storage chips (such as NAND Flash, PCRAM, RRAM, MRAM, and other non-volatile storage chips). It implements the mapping management between logical addresses and physical addresses through a logical address mapping table and provides an efficient storage access interface to achieve efficient collaboration with external computing units (such as CPU, GPU, NPU, etc.).
[0021] To meet the stringent requirements of large-scale neural network models for high-performance storage, this invention provides a novel storage controller architecture, the functional modules of which will be described in detail below.
[0022] The controller chip 110 is the core control unit of the entire storage controller, integrating multiple heterogeneous controller cores, each responsible for handling critical tasks such as data transmission, address mapping, and error correction. For example, please refer to... Figure 2 The controller chip 110 integrates 16 heterogeneous controller cores.
[0023] In addition, the controller chip 110 also integrates a high-speed data interface for high-speed data interaction with external computing units. For example, the high-speed data interface can use PCIe, NVMe or higher-level communication protocols to support TB-level data throughput requirements and ensure low latency and high bandwidth requirements for data access during the training and inference of large-scale neurophysical models.
[0024] The computing unit is the hardware core responsible for executing inference tasks of large-scale neural network models. It is configured to perform computational tasks such as matrix operations, convolution operations, and activation function processing on the input data, thereby completing the inference task of the large-scale neural network model. The specific architecture of the computing unit is not limited here. Depending on actual needs, the computing unit can be implemented using dedicated architectures such as tensor cores, AI-specific accelerators, or programmable logic units, or it can be implemented using general-purpose graphics processing units (GPUs) or multi-core central processing units (CPUs) to adapt to the computational requirements of neural network models of different scales and types. The computing unit is directly connected to the controller chip 110 through a high-speed data interface, supporting end-to-end data pipeline scheduling. It can dynamically access weight parameters and feature data stored in the memory chip during model inference / training, significantly reducing memory access latency.
[0025] In this embodiment of the invention, in order to further improve storage performance, a random access memory chip 120 is also integrated inside the storage controller. The random access memory chip 120 is bonded to the controller chip 110 and uses the same address encoding rule as the controller chip 110.
[0026] For example, the random access memory (RAM) chip 120 is vertically stacked on top of the controller chip 110 using a hybrid bonding process. The RAM chip 120 and the controller chip 110 are interconnected at high density through a through-silicon via (TSV) array, achieving terabyte-level interconnect bandwidth. Each controller core of the controller chip 110 is directly connected to the RAM chip 120 through a dedicated TSV channel.
[0027] Furthermore, it should be noted that the type of random access memory chip 120 is not limited in the embodiments of the present invention, and it can be implemented using single-layer or multi-layer stacked dynamic random access memory (DRAM) or high bandwidth memory (HBM).
[0028] In this embodiment of the invention, the random access memory chip 120 can not only participate in data prefetching and temporary storage as a cache level, but can also be uniformly addressed by the controller chip 110, thereby supporting low-latency access to hot data and real-time computing collaboration.
[0029] The random access memory (RAM) chip 120 is divided into three RAM spaces: a first RAM space, a second RAM space, and a third RAM space. There are no specific restrictions on how these three RAM spaces are divided; they can be logically partitioned or physically partitioned. For example, when the RAM chip 120 has a multi-layer stacked structure, the RAM spaces can be distributed across different physical layers; while when a single-layer structure is used, the RAM spaces can be logically isolated through address mapping.
[0030] The first random access memory (RAM) space is used to store a logical address mapping table. For example, when the memory chip is a flash memory chip, the logical address mapping table is a flash translation layer (FTL) mapping table, used to achieve efficient mapping between host logical addresses and flash memory physical addresses. In specific implementations, the controller chip 110 can load the entire contents of the FTL mapping table into the first RAM when the first RAM has sufficient capacity, and when the first RAM has insufficient capacity, load only some of the hot entries from the FTL mapping table into the first RAM, while the remaining non-hot entries are stored in the flash memory chip and loaded as needed.
[0031] The second random access memory space is used to store command queues corresponding to different memory channels. Each memory channel is connected to at least one memory chip. The command queue is used to cache operation commands (such as read commands, write commands, or erase commands) issued by the controller chip to the memory chip, realizing multi-channel parallel scheduling and load balancing. For example, please refer to... Figure 3 The controller chip 110 is equipped with N storage channels, each storage channel is connected to M storage chips, and each storage channel independently selects different storage chips through chip select to achieve precise control.
[0032] The third random access memory space is used to store read data returned by the memory chip in response to operation commands, or to store write data required by the memory chip in response to operation commands. For example, when the memory controller provided by this invention is applied to the inference or training of a large-scale neural network model, the read data can be the weight parameters of the large-scale neural network model, and the write data can be intermediate data during the model inference process, gradient update data during the model training process, and so on.
[0033] Alternatively, in one embodiment, please refer to Figure 4 The controller chip 110 includes: The request scheduling unit 1110 is used to receive storage access requests from the computing unit and, according to the logical address mapping table, split the storage access request into multiple operation commands and write them into the corresponding command queue. The channel control unit 1120 is used to obtain operation commands from the command queue and send them to the memory chip for execution in parallel through the corresponding storage channel.
[0034] The above-mentioned request scheduling unit 1110 is the controller core in the controller chip 110 responsible for executing request scheduling tasks, and the channel control unit 1120 is the controller core in the controller chip 110 responsible for executing channel control tasks. The two work together to achieve efficient data read and write scheduling.
[0035] The request scheduling unit 1110 is configured to receive storage access requests from the computing unit via a high-speed data interface. These requests can include read or write requests. Write requests carry the data to be written and its corresponding logical address information, while read requests carry the logical address information of the data to be read. Furthermore, the request scheduling unit 1110 converts the logical addresses in the storage access requests into corresponding physical addresses based on a logical address mapping table stored in the first random access memory space. It then generates multiple operation commands containing the target storage channel, operation type, and physical address based on the physical address distribution and channel load. Operation commands targeting the same storage channel are written into the same command queue, thereby achieving request-level parallel processing and load balancing, effectively reducing access latency.
[0036] The channel control unit 1120 periodically polls each command queue. Once an operation command is detected, it is sent to the corresponding memory chip for execution according to the storage channel, and the execution is confirmed through an interrupt or polling mechanism.
[0037] When the storage access request is a write request, the request scheduling unit 1110 can split the write data carried by the write request into multiple write data blocks aligned with the write granularity of the storage chip, and store each of the split write data blocks in the third random access memory space. At the same time, the request scheduling unit 1110 generates write operation commands corresponding to each write data block according to the logical address mapping table, and writes these write operation commands into the target command queue after associating them with the storage address of the corresponding write data block. After the channel control unit 1120 obtains the write operation command from the command queue, it reads the corresponding write data block from the third random access memory space according to the storage address, and sends it to the corresponding storage chip through the target storage channel to execute the write, thus completing the response to the write request.
[0038] When the storage access request is a read request, the request scheduling unit 1110 converts the logical address carried by the read request into a physical address through a logical address mapping table, determines the target storage channel based on the distribution of these physical addresses, generates the corresponding read operation command, and writes each read operation command into the corresponding target command queue. After the channel control unit 1120 obtains the read operation command from the command queue, it sends it to the designated storage chip through the corresponding storage channel to execute the read, and stores the read data fragment returned by the read into the third random access memory space. After all the relevant read data fragments have been received, the random access memory chip 120 reassembles each read data fragment into complete read data in the order of request through the internal high-speed bus. Finally, the read data is returned by the request scheduling unit 1110 to the computing unit that initiated the read request through the high-speed data interface, completing the response to the read request.
[0039] Alternatively, in one embodiment, please refer to Figure 5The storage controller also includes an address mapping unit 130, which is adjacent to the random access memory chip 120. The request scheduling unit 1110 is used to send the logical address in the storage access request to the address mapping unit 130. The address mapping unit 130 is used to query the logical address mapping table according to the logical address to obtain the corresponding physical address and return the physical address to the request scheduling unit 1110. The request scheduling unit 1110 is used to generate multiple operation commands according to the physical address and write them to the corresponding command queue.
[0040] To further improve storage performance, in this embodiment of the invention, the storage controller further includes an address mapping unit 130, which is adjacent to the random access memory chip 120 and configured to perform address mapping translation operations from logical addresses to physical addresses. For example, the address mapping unit 130 can be integrated into the peripheral logic of the same random access memory chip 120, thereby reducing access latency by shortening the address mapping path. This design allows the translation from logical addresses to physical addresses to be completed at high speed close to the storage core, avoiding the performance loss caused by cross-module communication in traditional architectures, and significantly improving response efficiency, especially in high-concurrency scenarios.
[0041] Upon receiving a storage access request, the request scheduling unit 1110 sends the logical address it carries to the address mapping unit 130. The address mapping unit 130 performs a lookup operation based on the logical address mapping table stored in the first random access memory space in the random access memory chip 120 to quickly obtain the corresponding physical address and returns the physical address to the request scheduling unit 1110. The request scheduling unit 1110 generates multiple fine-grained operation commands based on the obtained physical address and writes them into the corresponding command queue.
[0042] For example, the address mapping unit 130 includes 64 independent comparators. The request scheduling unit 1110 can send 64 logical addresses in parallel to the 64 comparators of the address mapping unit 130. The random access memory chip 120 activates a row of 4096 bits of data in the stored logical address mapping table (corresponding to 64 logical address physical address mapping entries, each mapping entry occupies 64 bits). The activated 4096 bits of mapping data are simultaneously broadcast to the 64 comparators. The 64 comparators work simultaneously, completing the matching of 64 logical addresses with the mapping data within one clock cycle, and outputting the corresponding physical block address.
[0043] It should be noted that the hardware circuit implementation of the address mapping unit 130 in the embodiments of the present invention is not limited. It can be implemented by an application-specific integrated circuit, a programmable logic device or an address translation module in an existing processor architecture, or by a hardware state machine controlled by microcode, as long as it can support low-latency table lookup and matching operations between logical addresses and physical addresses.
[0044] By bringing the logical address mapping process forward and tightly coupling it to the random access memory chip 120, the address translation latency can be effectively reduced and the overall bandwidth utilization can be improved.
[0045] Optionally, in one embodiment, the request scheduling unit 1110 is further configured to, upon receiving a new storage access request from the computing unit, obtain the correlation between the new storage access request and the storage access request, and, based on the logical address mapping table and the correlation, split the new storage access request into multiple operation commands and write them into the corresponding command queue.
[0046] The correlation includes the spatial locality of data access. The request scheduling unit 1110 identifies the continuity or overlap between the new storage access request and the previous storage access request in the address space by analyzing the logical address distribution characteristics of the previous and subsequent requests.
[0047] When a new storage access request and a previous storage access request point to the same physical block or adjacent logical page, that is, when the new storage access request and the previous storage access request have spatial locality, the request scheduling unit 1110 can directly reuse the existing mapping result of the previous storage access request to generate the operation command for the new storage access request, avoiding repeated table lookups and address translations, thereby further reducing address translation overhead and reducing access latency.
[0048] Accordingly, after splitting the new storage access request into multiple operation commands, the request scheduling unit 1110 distributes the multiple operation commands to the corresponding command queues for execution. Please refer to the relevant descriptions in the above embodiments for details, which will not be repeated here.
[0049] By introducing a correlation analysis mechanism, the overhead of address mapping can be reduced and the efficiency of request response can be improved.
[0050] Alternatively, in one embodiment, please refer to Figure 6 The storage controller also includes a mode switching unit 140 for switching the operating mode of the random access memory chip 120, including a wide mode and a narrow mode. In the wide mode, the bus width of the random access memory chip 120 is consistent with the bus width of the controller chip. In the narrow mode, the bus width of the random access memory chip 120 is divided into multiple independent sub-channels.
[0051] In this embodiment of the invention, by utilizing the interconnection between the random access memory chip 120 and the controller chip 110 via a through-silicon via array, fine-grained bandwidth allocation and parallel access at the sub-channel level can be achieved.
[0052] The mode switching unit 140 can dynamically switch the working mode of the random access memory chip 120 according to the access requirements of the controller chip 110. In high-concurrency, small-granularity access scenarios, it can enable the narrow mode, divide the bus resources into multiple independent sub-channels to support multi-task parallel transmission and improve system throughput efficiency; in large-block continuous data access scenarios, it can switch to the wide mode to aggregate all bandwidth resources to achieve peak bandwidth utilization.
[0053] For example, the through-silicon via array between the controller chip 110 and the random access memory chip 120 forms a 4096-bit wide bus. When the mode switching unit 140 detects a continuous large data block access request from the controller chip 110, it switches the random access memory chip 120 to wide mode, making the entire 4096-bit bus work as a single data channel. This fully releases the bus bandwidth potential and significantly improves the efficiency of large data read and write operations. This is particularly suitable for reading and writing fixed 512-byte data blocks in large-scale neurophysical models, ensuring that the transmission of each data block can utilize the full bandwidth, reducing transmission cycles, and thus accelerating the model training and inference process. In addition, when the mode switching unit 140 detects that the controller chip 110 needs to load the logical address mapping table to the random access memory chip 120, it switches the random access memory chip 120 to narrow mode, dividing the 4096-bit bus into multiple independent sub-channels. This supports parallel small-granular access, effectively reducing the loading latency of the logical address mapping table and improving the overall response speed.
[0054] The above-mentioned mode switching unit 140 dynamically switches the working mode of the random access memory chip 120, realizing fine-grained adaptation to different access scenarios, and improving parallel processing capabilities while ensuring bandwidth utilization.
[0055] Optionally, in one embodiment, the random access memory chip 120 is a three-dimensional stacked dynamic random access memory chip.
[0056] In this embodiment of the invention, the random access memory chip 120 is implemented using a three-dimensional stacked dynamic random access memory chip. This three-dimensional stacked dynamic random access memory chip is a high-bandwidth, low-latency memory architecture. Multiple memory layers are vertically stacked using through-silicon via (TSV) technology and integrated with the controller chip 10 within the same package, achieving an exponential increase in memory bandwidth and a significant reduction in access latency. Simultaneously, multi-layer memory stacking can integrate larger capacity within a single package. Combined with the dynamic bandwidth scheduling of the mode switching unit 140, this achieves synergistic optimization of capacity and bandwidth. In narrow mode, each sub-channel can independently access different memory layers, further improving the granularity of parallel access. In wide mode, cross-layer data can be transmitted in parallel via the TSV array, fully leveraging the bandwidth advantages of the three-dimensional structure to meet the dual requirements of high bandwidth and low latency in artificial intelligence computing scenarios.
[0057] As described above, this invention provides a novel storage controller architecture, comprising: a controller chip; and a random access memory (RAM) chip, bonded to the controller chip and using the same address encoding rule as the controller chip. The RAM chip includes: a first RAM space for storing a logical address mapping table; a second RAM space for storing command queues corresponding to different storage channels, each storage channel connecting to at least one storage chip, the command queues being used to cache operation commands issued by the controller chip to the storage chips; and a third RAM space for storing read data returned by the storage chips in response to operation commands, or write data required by the storage chips in response to operation commands. Thus, by tightly integrating the RAM chip and the controller chip and sharing address encoding, zero-copy data interaction between the controller chip and the RAM chip is achieved, significantly reducing data access latency. Meanwhile, by residing the logical address mapping table in the first random access memory (RAM) space, high-speed address mapping lookup and updates are achieved; distributing the command queue in the second RAM space enables multi-channel parallel instruction scheduling, enhancing concurrent processing capabilities; and using the third RAM space as a data transfer buffer effectively coordinates the efficient flow of read and write data between the controller and the memory chip, thereby comprehensively improving overall response speed and throughput performance. This memory controller architecture significantly improves data throughput efficiency and reduces overall system power consumption through tight hardware-level coupling, making it suitable for high-concurrency, low-latency artificial intelligence computing scenarios. In typical scenarios, this memory controller architecture can support accelerated inference / training of large-scale neural network models such as the Transformer architecture, meeting the stringent high-performance storage requirements of large-scale neural network models.
[0058] In one embodiment, a storage system is also provided, please refer to Figure 7 The storage system includes a storage controller 10 and multiple storage chips 20 connected to the storage controller 10 via multiple storage channels, wherein each storage channel connects to at least one storage chip 20. The storage controller 10 can be the storage controller provided in the above embodiments, used to manage operations such as reading, writing, erasing, bad block management, and data error correction of the storage chips 20. The storage chips 20 can be implemented using non-volatile storage chips such as NAND Flash, PCRAM, RRAM, and MRAM.
[0059] It should be noted that the embodiments of the present invention do not limit the application scenarios of the above storage system. For example, the storage system is suitable for high-concurrency, low-latency artificial intelligence computing scenarios. In typical scenarios, the storage system can support the acceleration of inference / training of large-scale neural network models with architectures such as Transformer, meeting the stringent requirements of large-scale neural network models for high-performance storage.
[0060] In one embodiment, a computer device is also provided, please refer to Figure 8 The computer device includes a storage system 1 and a computing unit 2. The storage system 1 can be the storage system in the above embodiment. The computing unit 2 can be used to perform high-performance computing tasks such as artificial intelligence algorithms and big data processing. The storage system 1 provides high-bandwidth and low-latency data access support for the computing unit 2. The two work together to improve the overall computing efficiency.
[0061] Computation Unit 2 is the hardware core responsible for executing the inference task of large-scale neural network models. It is configured to perform computational tasks such as matrix operations, convolution operations, and activation function processing on the input data, thereby completing the inference task of large-scale neural network models. The specific architecture of Computation Unit 2 is not limited here. Depending on actual needs, Computation Unit 2 can be implemented using a dedicated architecture such as a tensor core, an AI-specific accelerator, or a programmable logic unit, or it can be implemented using a general-purpose graphics processing unit (GPU) or a multi-core central processing unit (CPU) to adapt to the computational requirements of neural network models of different scales and types. The Computation Unit is directly connected to Storage System 1 through a high-speed data interface, supporting end-to-end data pipeline scheduling. It can dynamically call the weight parameters and feature data stored in the storage system during model inference / training, significantly reducing memory access latency.
[0062] In practice, computer equipment can be such as servers, smart terminals, or edge computing devices.
[0063] In one embodiment, a storage access method is also provided. This storage access method is applicable to a storage controller, which includes a controller chip and a random access memory (RAM) chip. The RAM chip is bonded to the controller chip and uses the same address encoding rule as the controller chip. The RAM chip includes a first random access memory space, a second random access memory space, and a third random access memory space. Please refer to [reference needed]. Figure 9 The storage access method includes: In S310, a memory access request from the computing unit is received via the controller chip; In S320, according to the logical address mapping table stored in the first random access memory space, the controller chip splits the storage access request into multiple operation commands and writes them into the command queues corresponding to different storage channels stored in the second random access memory space. In S330, based on the third random access memory space, the controller chip sends operation commands to the connected memory chip via the memory channel for execution.
[0064] Optionally, in one embodiment, the controller chip includes a request scheduling unit and a channel control unit. The controller chip receives storage access requests from the computing unit, including: The request scheduling unit receives storage access requests from the computing unit; Based on the logical address mapping table stored in the first random access memory space, the controller chip splits the memory access request into multiple operation commands and writes them into the command queues corresponding to different memory channels stored in the second random access memory space, including: According to the logical address mapping table stored in the first random access memory space, the storage access request is split into multiple operation commands by the request scheduling unit and written into the command queues corresponding to different storage channels stored in the second random access memory space. Based on the third random access memory space, the controller chip sends operation commands to the connected memory chip via the memory channel for execution, including: The operation command is obtained from the command queue through the channel control unit and sent to the memory chip for execution in parallel through the corresponding storage channel based on the third random access memory space.
[0065] Optionally, in one embodiment, the storage controller further includes an address mapping unit adjacent to the random access memory chip. Based on the logical address mapping table stored in the first random access memory space, the address mapping unit splits the storage access request into multiple operation commands via a request scheduling unit and writes them into command queues corresponding to different storage channels stored in the second random access memory space, including: The logical address in the storage access request is sent to the address mapping unit through the request scheduling unit. The address mapping unit retrieves the corresponding physical address by querying the logical address mapping table based on the logical address, and then returns the physical address to the request scheduling unit. The request scheduling unit generates multiple operation commands based on the physical address and writes them to the corresponding command queue.
[0066] Optionally, in one embodiment, the storage access method provided by the present invention further includes: When a new storage access request is received from the computing unit through the request scheduling unit, the correlation between the new storage access request and the existing storage access request is obtained through the request scheduling unit. Based on the logical address mapping table and the association, the new storage access request is split into multiple operation commands and written to the corresponding command queue by the request scheduling unit.
[0067] Optionally, in one embodiment, the storage controller further includes a mode switching unit, and the storage access method provided by the present invention further includes: The operating mode of the random access memory chip is switched by the mode switching unit, including wide mode and narrow mode. In wide mode, the bus width of the random access memory chip is the same as that of the controller chip. In narrow mode, the bus width of the random access memory chip is divided into multiple independent sub-channels.
[0068] Optionally, in one embodiment, the random access memory chip is a three-dimensional stacked dynamic random access memory chip.
[0069] It should be noted that the specific limitations on storage access methods can be found in the limitations on storage controllers mentioned above, and will not be repeated here.
[0070] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A storage controller, characterized in that, include: Controller chip; A random access memory (RAM) chip, bonded to the controller chip and using the same address encoding rule as the controller chip, the RAM chip comprising: The first random access memory space is used to store the logical address mapping table; The second random access memory space is used to store command queues corresponding to different memory channels. Each memory channel is connected to at least one memory chip. The command queue is used to cache operation commands issued by the controller chip to the memory chip. The third random access memory space is used to store read data returned by the memory chip in response to the operation command, or to store write data that the memory chip needs to write in response to the operation command.
2. The storage controller according to claim 1, characterized in that, The controller chip includes: The request scheduling unit is used to receive storage access requests from the computing unit and, according to the logical address mapping table, split the storage access request into multiple operation commands and write them into the corresponding command queue. The channel control unit is used to obtain operation commands from the command queue and send them to the memory chip for execution in parallel through the corresponding storage channel.
3. The storage controller according to claim 2, characterized in that, The storage controller further includes an address mapping unit, which is adjacent to the random access memory chip. The request scheduling unit is used to send the logical address in the storage access request to the address mapping unit. The address mapping unit is used to query the logical address mapping table according to the logical address to obtain the corresponding physical address, and return the physical address to the request scheduling unit. The request scheduling unit is used to generate multiple operation commands based on the physical address and write them to the corresponding command queue.
4. The storage controller according to claim 2, characterized in that, The request scheduling unit is further configured to, upon receiving a new storage access request from the computing unit, obtain the correlation between the new storage access request and the existing storage access request, and, based on the logical address mapping table and the correlation, split the new storage access request into multiple operation commands and write them into the corresponding command queue.
5. The storage controller according to any one of claims 1-4, characterized in that, The storage controller further includes a mode switching unit for switching the operating mode of the random access memory chip, including a wide mode and a narrow mode. In the wide mode, the bus width of the random access memory chip is the same as the bus width of the controller chip. In the narrow mode, the bus width of the random access memory chip is divided into multiple independent sub-channels.
6. The storage controller according to any one of claims 1-4, characterized in that, The random access memory chip is a three-dimensional stacked dynamic random access memory chip.
7. A storage system, characterized in that, include: The storage controller according to any one of claims 1-6; Multiple memory chips are connected to the memory controller via multiple memory channels, wherein each memory channel is connected to at least one memory chip.
8. A computer device, characterized in that, include: Computational unit; The storage system according to claim 7.
9. A storage access method, applicable to a storage controller, characterized in that, The storage controller includes a controller chip and a random access memory (RAM) chip. The RAM chip is bonded to the controller chip and uses the same address encoding rule as the controller chip. The RAM chip includes a first RAM space, a second RAM space, and a third RAM space. The storage access method includes: The controller chip receives storage access requests from the computing unit. According to the logical address mapping table stored in the first random access memory space, the controller chip splits the storage access request into multiple operation commands and writes them into the command queues corresponding to different storage channels stored in the second random access memory space. Based on the third random access memory space, the controller chip sends the operation command to the connected memory chip via the memory channel for execution.
10. The storage access method according to claim 9, characterized in that, The controller chip further includes a mode switching unit, and the storage access method further includes: The operating mode of the random access memory chip is switched by the mode switching unit, including wide mode and narrow mode. In the wide mode, the bus width of the random access memory chip is the same as that of the controller chip. In the narrow mode, the bus width of the random access memory chip is divided into multiple independent sub-channels.