Storage controller, storage system, computer device and model reasoning method

By integrating a controller chip and a neural network processing unit into the storage controller, the computational tasks of the hybrid expert model are executed locally. This solves the technical problems that traditional storage architectures cannot meet, improves system energy efficiency, meets the high concurrency and low latency performance requirements of the hybrid expert model, achieves high concurrency and low latency response speed of the hybrid expert model, improves the high concurrency and low latency response speed of the system, and realizes the efficient execution of hybrid experts.

CN122018784APending Publication Date: 2026-05-12YEESTOR MICROELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YEESTOR MICROELECTRONICS CO LTD
Filing Date
2025-12-23
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Traditional storage architectures struggle to meet the high concurrency and low latency performance requirements of large-scale hybrid expert models during inference, especially when access patterns are highly irregular, leading to increased load pressure on storage systems.

Method used

A storage controller architecture is provided, including a controller chip and a neural network processing unit. By receiving task execution instructions and executing computational tasks of hybrid expert models, the computing tasks are executed locally at the storage end, reducing the frequency of data transmission to and from the main computing unit. Furthermore, multi-expert parallel processing and low-latency response are achieved through intelligent scheduling of computing resources.

Benefits of technology

It improves the system's energy efficiency ratio, meets the performance requirements of large-scale models, especially hybrid expert models, for high-concurrency, low-latency inference, and provides a feasible path for efficient deployment on the edge and end sides.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122018784A_ABST
    Figure CN122018784A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of artificial intelligence, and relates to a storage controller, a storage system, computer equipment and a model reasoning method. The storage controller comprises a controller chip, and the controller chip comprises a task input and output unit which is used for receiving a task execution instruction from a calculation unit; the neural network processing unit is used for executing a calculation task of at least one expert network in the hybrid expert model according to the task execution instruction to obtain a first calculation result; and the task input and output unit is also used for writing the first calculation result into a destination address indicated by the task execution instruction, and is used for the calculation unit to obtain a reasoning result in combination with a second calculation result of other expert networks in the self-execution hybrid expert model. The method can meet the performance requirements of a large-scale model, especially a hybrid expert model, for high-concurrency and low-delay reasoning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a storage controller, storage system, computer device, and model inference method. Background Technology

[0002] With the rapid development of artificial intelligence technology, the number of parameters and computational demands of large-scale models are growing exponentially. Current large-scale models generate massive random access requests during training and inference, placing extremely high demands on the bandwidth, latency, and concurrency processing capabilities of storage systems. For example, hybrid expert models, as a typical sparse activation architecture, only activate a portion of the expert network during inference, resulting in highly irregular access patterns and further exacerbating the load on storage systems. Traditional storage architectures struggle to meet the performance requirements of hybrid expert models for high-concurrency, low-latency inference. Summary of the Invention

[0003] This invention provides a storage controller, a storage system, a computer device, and a model inference method, which can meet the performance requirements of large-scale models, especially hybrid expert models, for high-concurrency, low-latency inference.

[0004] In a first aspect, embodiments of the present invention provide a storage controller, including a controller chip, the controller chip comprising: The task input / output unit is used to receive task execution instructions from the computing unit. The neural network processing unit is used to execute the computational task of at least one expert network in the hybrid expert model according to the task execution instructions, and obtain the first computational result; The task input / output unit is also used to write the calculation results to the destination address indicated by the task execution instruction, so that the calculation unit can combine the second calculation results of the other expert networks in the hybrid expert model to obtain the inference results.

[0005] Secondly, the storage system provided in the embodiments of the present invention includes: Memory chips; The storage controller provided in this embodiment of the invention.

[0006] Thirdly, the computer device provided in the embodiments of the present invention includes: Computational unit; The storage system provided in this embodiment of the invention.

[0007] Fourthly, the model reasoning method provided in the embodiments of the present invention includes: The computing unit sends task execution instructions to the storage controller. According to the task execution instructions, the computation task of at least one expert network in the hybrid expert model is executed through the storage controller to obtain the first computation result; The computational unit performs the computational tasks of the remaining expert networks in the hybrid expert model to obtain the second computational result; Based on the first and second calculation results, the reasoning result is obtained through the calculation unit.

[0008] This invention provides a novel storage controller architecture, including a controller chip comprising a task input / output unit and a neural network processing unit. The task input / output unit receives task execution instructions from a computing unit, and the neural network processing unit, based on these instructions, executes computational tasks of at least one expert network in a hybrid expert model to obtain a first computational result. The task input / output unit also writes the first computational result to a destination address indicated by the task execution instructions, allowing the computing unit to combine it with the second computational results of the remaining expert networks to obtain an inference result. This architecture transforms the storage controller from a traditional data transporter into an intelligent processing unit with computing power, enabling localized execution of computational tasks at the storage end. This significantly reduces the frequency of data transmission to and from the main computing unit, improving system energy efficiency. Combined with the dynamic activation characteristics of hybrid expert models, the storage controller can intelligently schedule computing resources according to task instructions, achieving multi-expert parallel processing and low-latency response. This meets the performance requirements of large-scale models, especially hybrid expert models, for high-concurrency, low-latency inference, providing a feasible path for the efficient deployment of hybrid expert models at the edge and endpoint. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 This is a schematic diagram of the structure of a storage controller provided in an embodiment of the present invention; Figure 2 This is another structural schematic diagram of the storage controller provided in an embodiment of the present invention; Figure 3 This is a detailed structural diagram of the neural network processing unit 120; Figure 4 This is a schematic diagram of the structure of the storage system provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of the computer device provided in an embodiment of the present invention; Figure 6This is a flowchart illustrating the model reasoning method provided in this embodiment of the invention. Detailed Implementation

[0011] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.

[0012] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0013] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0014] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."

[0015] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0016] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0017] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0018] This invention provides a storage controller, a storage system, a computer device, and a model inference method. The storage controller includes a controller chip, which comprises a task input / output unit and a neural network processing unit. The task input / output unit receives task execution instructions from an external computing unit. The neural network processing unit executes the computational tasks of at least one expert network in a hybrid expert model according to the task execution instructions to obtain a first computational result. The task input / output unit also writes the first computational result to the destination address indicated by the task execution instructions, so that the computing unit can combine the second computational results of the remaining expert networks in the hybrid expert model with its own execution to obtain an inference result and complete the final inference output.

[0019] Please refer to Figure 1 This is a schematic diagram of the structure of a storage controller disclosed in an embodiment of the present invention, as shown below. Figure 1 As shown, the storage controller may include a controller chip 100, which includes: The task input / output unit 110 is used to receive task execution instructions from the computing unit; The neural network processing unit 120 is used to execute the computational task of at least one expert network in the hybrid expert model according to the task execution instruction, and obtain the first computational result; The task input / output unit 110 is also used to write the first calculation result to the destination address indicated by the task execution instruction, and to combine the second calculation result of the other expert networks in the hybrid expert model with the calculation unit to obtain the inference result.

[0020] It should be noted that the storage controller is the core component responsible for managing the read, write, erase, bad block management, and data error correction operations of storage chips (such as NAND Flash, PCRAM, RRAM, MRAM, and other non-volatile storage chips). It implements the mapping management between logical addresses and physical addresses through a logical address mapping table and provides an efficient storage access interface to achieve efficient collaboration with external computing units (such as CPU, GPU, NPU, etc.).

[0021] To meet the stringent requirements of high-performance storage for large-scale neural network models, especially hybrid expert models, this invention provides a new storage controller architecture, and the functional modules in this architecture will be described in detail below.

[0022] The controller chip 100 is the core control unit of the entire storage controller, integrating a high-speed data interface and multiple heterogeneous controller cores. The high-speed data interface is used to realize high-speed data interaction with external computing units. For example, the high-speed data interface can use PCIe, NVMe or higher-level communication protocols to support TB-level data throughput requirements. The multiple heterogeneous controller cores are used to handle different types of storage control tasks such as data transmission, address mapping and error correction.

[0023] The computing unit is the hardware core responsible for executing inference / training tasks of large-scale neural network models. It is configured to perform computational tasks such as matrix operations, convolution operations, and activation function processing on the input data, thereby completing the inference / training tasks of large-scale neural network models. The specific architecture of the computing unit is not limited here. Depending on actual needs, the computing unit can be implemented using dedicated architectures such as tensor cores, AI-specific accelerators, or programmable logic units, or it can be implemented using general-purpose graphics processing units (GPUs) or multi-core central processing units (CPUs) to adapt to the computational requirements of neural network models of different scales and types. The computing unit is directly connected to the controller chip 100 through a high-speed data interface, supporting end-to-end data pipeline scheduling. It can dynamically call weight parameters and feature data stored in the memory chip during model inference / training, significantly reducing memory access latency. This results in low data transfer overhead and improved computational efficiency.

[0024] The task input / output unit 110, as one of the multiple controller cores integrated in the controller chip 100, is configured to interact with external computing units for instructions and data. For example, it can be used to receive task execution instructions from external computing units. Specifically, to enable the task input / output unit 110 to respond to instructions from the computing units, an AI-specific instruction set is pre-configured in the task input / output unit 110. This allows the storage controller to directly parse and execute computational tasks from the computing units, rather than simply handling data read / write requests. For example, the AI-specific instruction set may include the EXPERT_COMPUTE instruction, used to trigger the activation and computation of specific expert networks in a hybrid expert model.

[0025] The neural network processing unit 120, as one of the multiple controller cores integrated in the controller chip 100, is configured to perform computational tasks such as matrix operations, convolution operations, and activation function processing. For example, the neural network processing unit 120 can perform some of the computational tasks offloaded by external computing units, such as performing the computational tasks of at least one expert network in a hybrid expert model offloaded by a computing unit.

[0026] Understandably, with the rapid development of large-scale neural network models, hybrid expert models have become an important direction in the evolution of current artificial intelligence architectures due to their efficiency and scalability in handling complex tasks. Their core lies in dynamically allocating computing resources to the most relevant expert sub-networks, achieving "on-demand activation," thereby ensuring model capacity while avoiding the waste of computing power caused by the participation of all parameters. A typical hybrid expert model contains dozens or even hundreds of expert networks, each independently processing a specific type of input feature, dynamically scheduled by a gating mechanism. The weight parameters of all expert networks in a hybrid expert model can reach hundreds of GB or even TB levels, placing extreme demands on storage bandwidth and access latency. However, under traditional storage architectures, because the weight parameters of expert networks need to be frequently switched and called between different expert networks, frequent loading and unloading of weight parameters leads to a large amount of data migration, exacerbating storage bandwidth pressure and causing computing units to idle.

[0027] The following will use a hybrid expert model as an example to illustrate the storage controller provided by this invention.

[0028] When performing inference tasks using a hybrid expert model, the computing unit first determines a subset of expert networks that need to be activated. Then, it further identifies expert networks from this subset that need to be offloaded to the storage controller. The specific method by which the computing unit determines the expert networks to be offloaded is not limited here. For example, the computing unit can dynamically select some computational tasks to be offloaded to the storage controller based on the current load, data locality, or energy efficiency strategies. For instance, when the computing unit load is high, it may select expert networks with a large number of parameters (e.g., reaching a parameter threshold) or computationally intensive tasks to be offloaded to the storage controller; when data locality is high, it may select expert networks with a large number of parameters and a high degree of matching with the characteristics of the current task to be offloaded to the storage controller, and so on.

[0029] As described above, after determining which expert networks need to be offloaded to the storage controller, the computing unit generates corresponding task execution instructions and sends these instructions to the storage controller via a high-speed data interface. These instructions include at least the storage address of the weight parameters of the expert network to be executed, the storage address of the input feature data, the destination address for the computation result return, and the type of computation operation. Furthermore, for expert networks within the expert network subset that are not offloaded to the storage controller for execution, the computing unit retains the ability to execute their computation tasks locally, obtaining the corresponding second computation result.

[0030] On the other hand, the task input / output unit 110 in the storage controller will receive the task execution instruction from the computing unit, parse out the weight parameter storage address, input feature data storage address, calculation result return destination address and calculation operation type of the expert network to be executed, and transmit the weight parameter storage address, input feature data storage address and calculation operation type of the expert network to be executed to the neural network computing unit 120.

[0031] Accordingly, after receiving the weight parameter storage address, input feature data storage address, and calculation operation type of the expert network to be executed forwarded by the task input / output unit 110, the neural network computing unit 120 configures the operation mode of its internal computing core (such as a multiply-accumulate computing unit) according to the calculation operation type. Based on the weight parameter storage address and input feature data storage address, it reads the weight parameters and input feature data of the expert network to be executed in parallel from the corresponding storage chip through a high-concurrency storage channel, thereby using the internal computing core to complete the calculation task of the expert network to be executed and obtaining the corresponding calculation result, which is recorded as the first calculation result. Subsequently, the neural network computing unit 120 transmits the first calculation result back to the task input / output unit 110 through the high-speed bus inside the controller (such as an AXI bus).

[0032] After receiving the first calculation result, the task input / output unit 110 writes the first calculation result to the calculation result return destination address. For example, when the calculation result return destination address points to NVMe SQ / CQ, the task input / output unit 110 encapsulates the first calculation result according to the NVMe protocol format and submits it to the corresponding completion queue CQ, and triggers an interrupt to notify the computing unit to retrieve the first calculation result, or the computing unit retrieves the first calculation result after polling and detecting a CQ status update; or, for another example, when the calculation result return destination address points to the dedicated response queue of the hybrid expert model, the task input / output unit 110 encapsulates the first calculation result according to a custom semantic format and writes it to the dedicated response queue to adapt to the parsing requirements of the computing unit during the runtime of the hybrid expert model.

[0033] As described above, the storage controller, through the collaboration of the task input / output unit 110 and the neural network computing unit 120, achieves efficient offloading and execution of expert network computing tasks. The entire process relies on high-speed interfaces and internal buses to complete data flow, ensuring low-latency response. The calculation results are encapsulated and returned according to preset addresses in protocol-compatible or custom formats, realizing a closed loop of computation offloading and improving overall inference throughput. Correspondingly, after obtaining the first calculation result returned by the storage controller, the computing unit combines the second calculation results of the other expert networks in the hybrid expert model it executes (i.e., the expert networks in the expert network subset other than those offloaded to the storage controller for execution) to obtain the final inference result, completing sparse inference.

[0034] Alternatively, in one embodiment, please refer to Figure 2 The storage controller also includes an internal cache unit 200, and the controller chip 100 also includes a controller unit 130. The controller unit 130 is used to receive expert classification information from the computing unit, and load the weight parameters of the first type of expert network in the hybrid expert model from the storage chip to the cache unit of the computing unit according to the expert classification information, load the weight parameters of the second type of expert network in the hybrid expert model to the internal cache unit, and retain the weight parameters of the remaining third type of expert network in the hybrid expert model in the storage chip.

[0035] In this embodiment of the invention, the weight parameters of the hybrid expert model are persistently stored in the storage chip by default, and are dynamically loaded into different levels of cache by the memory controller as needed to adapt to the execution requirements of the computing task.

[0036] The computing unit can classify input samples according to a preset expert classification strategy, determine the category of the expert network activated for each input sample, and send the corresponding expert classification information to the storage controller to guide it in loading the weight parameters of the corresponding category. Specifically, there are three types of expert networks. The first type of expert network requires the weight parameters to reside in the cache unit of the computing unit to support high-frequency access. The second type of expert network caches the weight parameters in the cache unit 200 inside the storage controller to balance access latency and bandwidth overhead. The third type of expert network retains the weight parameters in the storage chip and loads them only on demand when activated.

[0037] It should be noted that the configuration of the above expert classification strategy in this embodiment of the invention is not specifically limited. The classification strategy can be dynamically adjusted according to the actual application scenario. For example, it can be dynamically divided based on the characteristics of the expert network, such as the calling frequency, computing density, or memory access mode. At the same time, the classification strategy can also be decided collaboratively by the computing unit and the storage controller. The distribution of weight parameters and loading priority can be optimized in real time through a runtime feedback mechanism, thereby maximizing resource utilization and inference efficiency.

[0038] For example, during inference, the controller unit 130 in the storage controller can dynamically adjust the boundary between the second and third types of expert networks based on historical access statistics, thereby preloading frequently activated expert networks from the storage chip to the internal cache unit to reduce subsequent access latency. Similarly, the computing unit can periodically evaluate the activation frequency and computational load of each expert network, upgrade high-frequency expert networks to first-type expert networks and retain them in its own cache unit to ensure the efficient execution of critical computing paths.

[0039] As shown above, after completing the expert classification of the hybrid expert model, the computing unit transmits the expert classification information to the storage controller through a high-speed data interface.

[0040] On the other hand, the controller unit 130 in the storage controller will receive expert classification information from the computing unit, and parse out the expert network category involved in the current inference task based on the expert classification information, thereby triggering the scheduling operation of the corresponding weight parameters.

[0041] Specifically, for the first type of expert network indicated by expert classification information, the controller unit 130 loads the weight parameters of the first type of expert network from the storage chip to the cache unit of the computing unit to support continuous high-frequency access; for the second type of expert network, the controller unit 130 loads its weight parameters from the storage chip to the internal cache unit 200 of the storage controller to achieve a balance between access latency and storage overhead; for the third type of expert network, its weight parameters are retained in the storage chip and are only loaded into the internal cache unit 200 of the storage controller or the cache unit of the computing unit as needed when it is activated.

[0042] Accordingly, in this embodiment, when the computing unit determines which expert network needs to be offloaded to the storage controller for execution, it can select from the second type of expert networks that have been loaded into the internal cache unit 200 of the storage controller within the expert network subset. For example, the computing unit can evaluate the memory access characteristics and computational density of each second type of expert network and prioritize the selection of expert networks that are compatible with the computing capabilities of the storage controller for offloading, so as to fully leverage the advantages of storage-computing synergy. In this way, through fine-grained collaborative scheduling between the computing unit and the storage controller, not only is hierarchical management and on-demand loading of expert network weight parameters realized, but the throughput and energy efficiency of the hybrid expert model in actual inference scenarios are also significantly improved, effectively supporting the lightweight deployment of the hybrid expert model in the storage domain.

[0043] Furthermore, it should be noted that the embodiments of the present invention do not limit the storage medium type of the internal cache unit 200, and can be implemented using single-layer or multi-layer stacked dynamic random access memory (DRAM) or high bandwidth memory (HBM) and other random access characteristic media.

[0044] Optionally, in one embodiment, the internal cache unit 200 is bonded to the controller chip 100 and uses the same address encoding rule as the controller chip 100.

[0045] In this embodiment of the invention, the internal cache unit 200 is vertically stacked on the controller chip 100 using a hybrid bonding process. The internal cache unit 200 and the controller chip 100 are interconnected at high density via a through-silicon via (TSV) array, significantly improving the data transmission bandwidth between them and enabling terabyte-level interconnect bandwidth. Each controller core of the controller chip 100 establishes a direct connection to the internal cache unit 200 through a dedicated TSV channel, ensuring conflict-free access between cores.

[0046] Furthermore, the internal cache unit 200 and the controller chip 100 are uniformly addressed in their address spaces, supporting data mapping at the page, block, or stripe granularity, thereby improving address resolution efficiency and resource scheduling flexibility. Under this architecture, the internal cache unit 200 can not only participate in data prefetching and temporary storage as a cache layer, but also be uniformly addressed by the controller chip 100, thus supporting direct access and fine-grained memory management, meeting the diverse computing tasks' requirements for low latency and high bandwidth access.

[0047] Optionally, in one embodiment, the weight parameter heat of the second type of expert network is lower than that of the weight parameter heat of the first type of expert network, but higher than that of the weight parameter heat of the third type of expert network.

[0048] The weight parameter popularity is calculated based on dynamic statistical indicators such as the frequency of expert network calls and the number of consecutive accesses during the inference process, reflecting the activity level of different experts in network execution.

[0049] In this embodiment of the invention, the computing unit monitors and periodically evaluates the popularity of the weight parameters of each expert network in the hybrid expert model, and dynamically adjusts their classification level based on popularity thresholds. These popularity thresholds include a lower threshold distinguishing between second-class and third-class expert networks, and an upper threshold distinguishing between first-class and second-class expert networks. When the popularity of a certain expert network's weight parameters consistently exceeds the upper threshold, the computing unit classifies it as a first-class expert network; when the popularity of a certain expert network's weight parameters consistently falls below the lower threshold, the computing unit classifies it as a third-class expert network; and when the popularity of a certain expert network's weight parameters lies between the upper and lower thresholds, the computing unit classifies it as a second-class expert network.

[0050] Furthermore, after each update of the weight parameters of each expert network's popularity, the computing unit synchronously refreshes its classification level determination and sends new expert classification information to the storage controller, ensuring that both parties maintain consistency in their classification status of the expert networks. Correspondingly, the controller unit 130 in the storage controller dynamically adjusts the residency strategy of each expert network's weight parameters based on the latest received classification information. For example, if an expert network moves from category 3 to category 2, the controller unit 130 loads its weight parameters from the storage chip to the storage controller's internal cache unit 200; if an expert network moves from category 2 to category 1, the controller unit 130 migrates its weight parameters from the internal cache unit 200 to the computing unit's cache unit; if an expert network moves from category 1 to category 2, the internal cache unit 200 migrates its weight parameters from the computing unit's cache unit to the storage controller's internal cache unit 200; if an expert network moves from category 2 to category 3, the controller unit 130 removes its weight parameters from the storage controller's internal cache unit 200.

[0051] Alternatively, in one embodiment, please refer to Figure 3 The neural network processing unit 120 includes a computation subunit 1210 and a cache subunit 1220. The controller unit 130 is also used to load the target weight parameters of the expert network indicated by the task execution instruction from the internal cache unit 200 or the memory chip into the cache subunit. The computation subunit is used to execute the corresponding computation task according to the target weight parameters in the cache subunit to obtain the first computation result.

[0052] In this embodiment of the invention, the neural network processing unit 120 includes a computation subunit 1210 and a cache subunit 1220. The computation subunit 1210 is the internal computation core of the neural network processing unit 120, such as a multiply-accumulate computation unit, which is responsible for performing core computation tasks such as matrix operations and vector dot products. The cache subunit 1220 serves as a temporary storage area for the neural network processing unit 120, used to cache the data required for the current computation task and the intermediate data generated during the computation process.

[0053] Accordingly, in this embodiment of the invention, after receiving the task execution instruction from the computing unit, the task input / output unit 110 in the storage controller parses out the weight parameter storage address, input feature data storage address, calculation result return destination address and calculation operation type of the expert network to be executed, and transmits the weight parameter storage address and input feature data storage address of the expert network to be executed to the controller unit 130, and transmits the calculation operation type to the computing subunit 1210.

[0054] Accordingly, after receiving the weight parameter storage address, input feature data storage address, and computation operation type of the expert network to be executed forwarded by the task input / output unit 110, the controller unit 130 determines the current location of the target weight parameter based on the weight parameter storage address. If the target weight parameter resides in the storage chip, it is loaded from the storage chip to the cache subunit. If the target weight parameter resides in the internal cache unit 200, it is migrated from the internal cache unit 200 to the cache subunit 1220. Furthermore, the controller unit 130 also loads the input feature data into the cache subunit 1220 based on the input feature data storage address, ensuring that the computation subunit 1210 can efficiently access the required data when performing computation tasks.

[0055] As described above, after loading the target weight parameters and input feature data required for the computation subunit 1210 to perform the computation task, the computation subunit 1210 performs the corresponding neural network computation operation based on the target weight parameters and input feature data loaded in the cache subunit 1220, completes the matrix multiplication, activation function processing and other processes, and obtains the first computation result.

[0056] It should be noted that the medium type of the cache subunit 1220 is not specifically limited in this embodiment of the invention. For example, it can be a high-speed storage medium such as static random access memory.

[0057] Optionally, in one embodiment, the task input / output unit 110 is used to compile multiple task execution instructions from the same computing unit into task flow instructions when it receives multiple task execution instructions from the same computing unit.

[0058] In this embodiment of the invention, in order to improve computing efficiency and resource utilization, the task input / output unit 110 aggregates multiple task execution instructions from the same computing unit and generates task flow instructions containing time-series dependencies by compiling them. This enables pipeline scheduling and parallel optimization between tasks, minimizing computational idle time and storage access conflicts while ensuring data dependency integrity, thereby improving overall computing throughput and energy efficiency.

[0059] For example, for an expert network, the task input / output unit 110 can receive three task execution instructions from the same computing unit. The first instruction requires a linear transformation of the input feature data; the second instruction requires applying an activation function to the transformed result; and the third instruction requires randomly discarding the result after the activation function. Accordingly, the task input / output unit 110 compiles these three sequentially dependent task execution instructions into a single task flow instruction, which includes the execution order of each subtask, data pointer addresses, and synchronization signal identifiers. Subsequently, the controller unit 130 schedules the data loading of the cache subunit 1220 and the computation flow of the computing subunit 1210 according to this task flow instruction, achieving cascaded execution of linear transformation, activation processing, and random discarding. This eliminates the need to write intermediate results back to the storage chip, reducing memory access overhead and improving processing efficiency.

[0060] Optionally, in one embodiment, the task input / output unit 110 is used to aggregate multiple task execution instructions from different computing units into a batch processing task instruction when receiving multiple task execution instructions from different computing units.

[0061] In this embodiment of the invention, in order to improve computing efficiency and resource utilization, the task input / output unit 110 horizontally integrates multiple task execution instructions from different computing units, merges tasks with the same computing mode into batch processing task instructions, and realizes batch allocation and parallel processing of computing resources through unified scheduling, thereby reducing instruction parsing overhead while ensuring task isolation.

[0062] For example, the task input / output unit 110 can receive linear transformation task execution instructions for different expert networks from different computing units, and then aggregate multiple independent linear transformation task execution instructions into batch matrix operation task instructions, which are uniformly allocated by the controller unit 130 to the computing subunit 1210 for parallel processing. That is, the computing subunit 1210 performs linear transformations on the input feature data of different expert networks in parallel, thereby making full use of the parallel processing capability of the computing subunit 1210.

[0063] As described above, the novel storage controller architecture provided by this invention includes a controller chip comprising a task input / output unit and a neural network processing unit. The task input / output unit receives task execution instructions from the computing unit, and the neural network processing unit, based on these instructions, executes computational tasks of at least one expert network in the hybrid expert model to obtain a first computational result. The task input / output unit also writes the first computational result to the destination address indicated by the task execution instructions, allowing the computing unit to combine it with the second computational results of the remaining expert networks to obtain the inference result. This architecture transforms the storage controller from a traditional data transporter to an intelligent processing unit with computing power, enabling localized execution of computational tasks at the storage end. This significantly reduces the frequency of data transmission to and from the main computing unit, improving system energy efficiency. Combined with the dynamic activation characteristics of the hybrid expert model, the storage controller can intelligently schedule computing resources according to task instructions, achieving multi-expert parallel processing and low-latency response. This meets the performance requirements of large-scale models, especially hybrid expert models, for high-concurrency, low-latency inference, providing a feasible path for the efficient deployment of hybrid expert models at the edge and endpoint.

[0064] In one embodiment, a memory system is also provided, please refer to Figure 4 The storage system includes a storage controller 10 and multiple storage chips 20 connected to the storage controller 10 via multiple storage channels. Each storage channel connects to at least one storage chip 20. The storage controller 10 can be the storage controller provided in the above embodiments, used to manage read / write, erase, bad block management, and data error correction operations of the storage chips 20. The storage chips 20 can be implemented using non-volatile storage chips such as NAND Flash, PCRAM, RRAM, and MRAM. For example, when the storage chip 20 uses NAND Flash, the storage channel can be, for example, a physical bus defined by the ONFI or Toggle Mode interface standard to support high-bandwidth data transmission and low-latency access.

[0065] It should be noted that the embodiments of the present invention do not limit the application scenarios of the above storage system. For example, the storage system is suitable for high-concurrency, low-latency artificial intelligence computing scenarios. In typical scenarios, the storage system can support inference / training acceleration for hybrid expert models, meeting the stringent requirements of hybrid expert models for high-performance storage.

[0066] In one embodiment, a computer device is also provided, please refer to Figure 5The computer device includes a storage system 1 and a computing unit 2. The storage system 1 can be the storage system in the above embodiment. The computing unit 2 can be used to perform computing tasks of different expert networks in the hybrid expert model. The storage system 1 provides high-bandwidth, low-latency data access support for the computing unit 2. The two work together to improve the overall computing efficiency.

[0067] Specifically, computation unit 2 is configured to perform matrix operations, convolution operations, and activation function processing on the input data to complete the computational tasks of different expert networks in the hybrid expert model. The specific architecture of computation unit 2 is not limited here; depending on actual needs, it can be implemented using a dedicated architecture such as a tensor core, an AI-specific accelerator, or a programmable logic unit, or it can be implemented using a general-purpose graphics processing unit (GPU) or a multi-core central processing unit (CPU). Computation unit 2 is directly connected to storage system 1 via a high-speed data interface (such as a data interface using PCIe, NVMe, or higher-level communication protocols), supporting end-to-end data pipeline scheduling. It can dynamically call weight parameters and feature data stored in the storage system during model inference / training, significantly reducing memory access latency.

[0068] In practice, computer equipment can be such as servers, smart terminals, or edge computing devices.

[0069] In one embodiment, a model inference method is also provided, which is applicable to the aforementioned computer device. Please refer to [reference needed]. Figure 6 The inference method of this model includes: In S310, the computing unit sends task execution instructions to the storage controller; In S320, according to the task execution instruction, the computation task of at least one expert network in the hybrid expert model is executed through the storage controller to obtain the first computation result; In S330, the computational unit performs the computational tasks of the remaining expert networks in the hybrid expert model to obtain the second computational result; In S340, the reasoning result is obtained through the calculation unit based on the first calculation result and the second calculation result.

[0070] Optionally, in one embodiment, the storage controller includes a controller chip, which includes a task input / output unit and a neural network processing unit. According to task execution instructions, the storage controller executes computational tasks of at least one expert network in the hybrid expert model to obtain a first computational result, including: The task input / output unit receives task execution instructions from the computing unit. According to the task execution instructions, the computational task of at least one expert network in the hybrid expert model is executed by the neural network processing unit to obtain the first computational result.

[0071] Optionally, in one embodiment, the storage controller further includes an internal cache unit, and the controller chip further includes a controller unit. The model inference method provided by this invention: The controller unit is used to receive expert classification information from the computing unit; Based on the expert classification information, the controller unit loads the weight parameters of the first type of expert network in the hybrid expert model from the memory chip to the cache unit of the computing unit, loads the weight parameters of the second type of expert network in the hybrid expert model to the internal cache unit, and retains the weight parameters of the remaining third type of expert network in the hybrid expert model in the memory chip.

[0072] Optionally, in one embodiment, the weight parameter heat of the second type of expert network is lower than that of the weight parameter heat of the first type of expert network, but higher than that of the weight parameter heat of the third type of expert network.

[0073] Optionally, in one embodiment, the neural network processing unit includes a computation subunit and a cache subunit, and the model inference method provided by the present invention further includes: The controller unit loads the target weight parameters of the expert network, as indicated by the task execution instructions, from the internal cache unit or the memory chip into the cache subunit. Based on the target weight parameters in the cache subunit, the corresponding calculation task is executed by the calculation subunit to obtain the first calculation result.

[0074] Optionally, in one embodiment, the internal cache unit is bonded to the controller chip and uses the same address encoding rule as the controller chip.

[0075] Optionally, in one embodiment, receiving task execution instructions from the computing unit via the task input / output unit includes: When multiple task execution instructions from the same computing unit are received through the task input / output unit, these instructions are compiled into a task flow instruction.

[0076] Optionally, in one embodiment, receiving task execution instructions from the computing unit via the task input / output unit includes: When multiple task execution instructions from different computing units are received through the task input / output unit, these instructions are aggregated into a batch processing task instruction.

[0077] It should be noted that the specific limitations on the model inference method can be found in the limitations on the storage controller mentioned above, and will not be repeated here.

[0078] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A storage controller, characterized in that, Includes a controller chip, the controller chip comprising: The task input / output unit is used to receive task execution instructions from the computing unit. A neural network processing unit is configured to execute the computational task of at least one expert network in a hybrid expert model according to the task execution instructions, and obtain a first computational result; The task input / output unit is further configured to write the first calculation result to the destination address indicated by the task execution instruction, so that the calculation unit can combine the second calculation result of the other expert networks in the hybrid expert model to obtain the inference result.

2. The storage controller according to claim 1, characterized in that, The storage controller further includes an internal cache unit, and the controller chip further includes a controller unit. The controller unit is used to receive expert classification information from the computing unit, and load the weight parameters of the first type of expert network in the hybrid expert model from the storage chip to the cache unit of the computing unit according to the expert classification information, load the weight parameters of the second type of expert network in the hybrid expert model to the internal cache unit, and retain the weight parameters of the remaining third type of expert network in the hybrid expert model in the storage chip.

3. The storage controller according to claim 2, characterized in that, The weight parameter heat of the second type of expert network is lower than that of the first type of expert network, but higher than that of the third type of expert network.

4. The storage controller according to claim 2, characterized in that, The neural network processing unit includes a computation subunit and a cache subunit. The controller unit is further configured to load the target weight parameters of the expert network indicated by the task execution instruction from the internal cache unit or the memory chip into the cache subunit. The computation subunit is configured to execute the corresponding computation task according to the target weight parameters in the cache subunit to obtain a first computation result.

5. The storage controller according to claim 2, characterized in that, The internal cache unit is bonded to the controller chip and uses the same address encoding rule as the controller chip.

6. The storage controller according to any one of claims 1-5, characterized in that, The task input / output unit is used to compile multiple task execution instructions from the same computing unit into task flow instructions when it receives multiple task execution instructions from the same computing unit.

7. The storage controller according to any one of claims 1-5, characterized in that, The task input / output unit is used to aggregate multiple task execution instructions from different computing units into a batch processing task instruction when it receives multiple task execution instructions from different computing units.

8. A storage system, characterized in that, include: Memory chips; The storage controller according to any one of claims 1-7.

9. A computer device, characterized in that, include: Computational unit; The storage system according to claim 7.

10. A model reasoning method, characterized in that, include: The computing unit sends task execution instructions to the storage controller. According to the task execution instruction, the computation task of at least one expert network in the hybrid expert model is executed through the storage controller to obtain a first computation result; The computational unit performs the computational tasks of the remaining expert networks in the hybrid expert model to obtain a second computational result. Based on the first calculation result and the second calculation result, the reasoning result is obtained through the calculation unit.