Processor system and operation method for large-scale deep learning

Through the CPU-FPGA-multi-GPU architecture and expert on-chip network E-NoC, the problems of high CPU load and communication bottlenecks in traditional processor architectures are solved, efficient task distribution and weighted aggregation computing are realized, and the processing efficiency of large-scale deep learning tasks is improved.

CN120448132AActive Publication Date: 2025-08-08SHANDONG INSPUR SCI RES INST CO LTD

Patent Information

Application Number
CN202510918820.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-08-08
Estimated Expiration
2045-07-04

AI Technical Summary

Technical Problem

When the traditional CPU + heterogeneous GPU processor architecture handles large-scale deep learning tasks, the CPU load is too high, the task distribution efficiency is low, and the communication bandwidth between the CPU and GPU expert cores is limited, resulting in waste of computing resources and performance bottlenecks.

Method used

The CPU-FPGA-multi-GPU architecture is adopted, and the host and GPU acceleration card are connected through the expert on-chip network E-NoC. Expert communication controls the splitting task of the terminal node, calculates the matching score and dispatches it to the target expert processing module, realizes weighted aggregation calculation, and optimizes the communication mechanism.

Benefits of technology

Reduce CPU burden, improve task distribution efficiency and GPU resource utilization, optimize communication performance, and improve processing efficiency and system performance of large-scale deep learning tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448132A_ABST
    Figure CN120448132A_ABST
Patent Text Reader

Abstract

The invention provides a large-scale deep learning-oriented processor system and an operation method, and belongs to the technical field of processor architecture design, and the system comprises a host, an expert network-on-chip and a plurality of GPU acceleration cards; the expert network-on-chip carries out communication between the host and each GPU acceleration card; the GPU acceleration card comprises a plurality of expert processing modules; the CPU is used for carrying out system configuration, task deployment and process monitoring, responding to a deep learning model processing request and sending the request to the expert communication control terminal node; the expert communication control terminal node splits a deep learning model processing request into a plurality of subtasks according to input data characteristics of the deep learning model, and calculates a matching score of an expert processing module; scheduling the sub-tasks to a target expert processing module, and aggregating and calculating processing results of the expert processing module; and communicating with the GPU acceleration card through the expert network-on-chip. According to the method, the CPU burden is relieved, the task distribution efficiency and the GPU utilization rate are improved, and the communication performance is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of processor architecture design, and specifically relates to a processor system and operation method for large-scale deep learning. Background Art

[0002] Processors, collectively known as CPUs, GPUs, and SoCs, are the core components of computer systems, used to interpret computer instructions and rapidly process complex data. Deep learning, a subfield of machine learning, is a core technology in artificial intelligence and plays a vital role in areas such as image recognition, natural language processing, and autonomous driving. As deep learning models continue to expand, the traditional CPU + heterogeneous GPU processor architecture faces the following challenges when handling large-scale deep learning tasks: First, excessive CPU load. In traditional architectures, the CPU is responsible for decomposing large-scale deep learning models, matching and scheduling subtasks to SM expert cores, and aggregating the results from multiple expert cores. This complex operation leads to a surge in CPU load, which in turn affects the execution efficiency of other important processes and degrades the user experience. Second, task distribution efficiency is low. The CPU's concurrent data processing capabilities are limited. Inefficient task distribution not only reduces the overall processing speed of deep learning models, but also leads to underutilization and uneven load on the GPU's multiple SM cores, exacerbating the waste of computing resources. Third, communication between the CPU and GPU expert cores relies on traditional heterogeneous transmission methods, which have limited bandwidth and cannot meet the needs of large-scale data transmission. Furthermore, the simultaneous distribution and aggregation operations further increase communication overhead, becoming a bottleneck for performance improvement.

[0003] To address the above issues, existing technologies attempt to improve performance by optimizing task scheduling algorithms or increasing communication bandwidth, but the effects are limited. Summary of the Invention

[0004] In a first aspect, an embodiment of the present application provides a processor system for large-scale deep learning, including a host, an expert on-chip network E-NoC, and several GPU accelerator cards; The host includes a CPU and an expert communication control terminal node; Expert on-chip network E-NoC, used for communication between the host and each GPU accelerator card; Each GPU accelerator card includes several expert processing modules; the expert processing modules are routing expert processing modules or shared expert processing modules; The host computer and each GPU accelerator card are connected via an expert on-chip network E-NoC; The CPU is used for system configuration, task deployment, and process monitoring, responding to deep learning model processing requests, and sending them to the expert communication control terminal node; The expert communication control endpoint is used to perform the following operations: Split deep learning model processing requests into several subtasks based on the input data characteristics of the deep learning model; Calculate the matching score of the expert processing module; Dispatching subtasks to target expert processing modules; Perform aggregate calculations on the processing results of the expert processing module; Communicate with the GPU accelerator card through the expert on-chip network E-NoC.

[0005] Furthermore, the expert communication control terminal node includes: The request receiving module receives the deep learning model processing request sent by the CPU, parses the request data and transmits it to the deep learning task splitting and comparison module; Deep learning task splitting and comparison module, used for: According to the task category in the request data, the corresponding subtask splitting rules and expert model fusion coefficients are obtained from the pre-configured task coding table; Send the subtask splitting rules and the original model data in the request data to the expert matching gate module; Send the fusion coefficients to the output weighted aggregation module; Expertly matched door modules for: According to the subtask splitting rules, the input data of the deep learning model is split into several subtasks; Calculate the matching score between each subtask and the GPU's expert processing module, and select the target expert processing module; Select the target expert processing module and generate subtask assignment instructions; The task transmission module is used to send the subtask data to the target expert processing module through the expert on-chip network E-NoC according to the subtask allocation instructions; The expert processing result receiving module is used to receive the calculation results returned by the GPU expert processing module and cache them in the output weighted aggregation module; The output weighted aggregation module is used to perform weighted aggregation on the calculation results of each cached GPU expert processing module according to the fusion coefficient to generate the final output; The response sending module is used to return the final output to the CPU.

[0006] Furthermore, the expert matching gate module includes: The subtask splitting unit is used to split the input data of the deep learning model into subtasks for parallel processing according to the feature dimension or calculation granularity based on the subtask splitting rules; Expert status reading unit, used to obtain the busy and idle status of each GPU's expert processing module in real time; The routing expert matching unit is used to perform the following operations: Calculate the matching score between each subtask and the routing expert processing module and generate a list of Top-K candidate experts; Combined with the busy and idle status of the expert processing module, a routing expert processing module with the highest matching degree and an idle status is selected from the Top-K candidate expert list as the target routing expert processing module; Shared expert matching unit, used to perform the following operations: Dynamically determine the number of shared experts to be called based on the number of deep learning model layers to which the subtask belongs; Prioritize the shared expert processing module located on the same GPU accelerator card as the target routing expert processing module as the target shared expert processing module; if there is no idle shared expert processing module on the same GPU accelerator card, call it across GPU accelerator cards through the expert on-chip network E-NoC; The expert matching and selection FSM unit is used to generate the final subtask allocation instruction according to the matching results between the routing expert processing module and the shared expert processing module.

[0007] Furthermore, the shared expert matching unit selects the shared expert processing module according to the priority level: The first priority is the idle shared expert processing module on the same GPU accelerator card as the target routing expert processing module; The second priority is the shared expert processing module of the same GPU accelerator card that can be quickly released within a set time period; The third priority is the idle shared expert processing module of different GPU accelerator cards.

[0008] Furthermore, the expert communication control terminal node adopts an FPGA core.

[0009] In a second aspect, an embodiment of the present application further provides a processor operation method for large-scale deep learning, comprising the following steps: S1. The host CPU receives the deep learning model processing request and sends the request to the expert communication control terminal node; S2. The expert communication control terminal node decomposes the deep learning model processing request into several subtasks, calculates the matching score between each subtask and the GPU's expert processing module, and selects the target expert processing module; S3. The expert communication control terminal node sends the subtask to the target expert processing module through the expert on-chip network E-NoC; S4. The target expert processing module processes the subtask and returns the processing results to the expert communication control terminal node; S5. The expert communication control terminal node performs weighted aggregation on the processing results and sends the final result back to the CPU.

[0010] Furthermore, the specific steps of step S2 are as follows: According to the task category in the request data, obtain the subtask splitting rules and fusion coefficients from the pre-configured task coding table; Split the deep learning model input data into several subtasks according to feature dimensions or calculation granularity based on the subtask splitting rules; For each subtask, do the following: Calculate the matching score with the routing expert processing module and generate a Top-K candidate list; Selecting an idle routing expert processing module with the highest matching degree from the candidate list as the target routing expert processing module; Determine the number of shared expert processing modules N based on the number of deep learning model layers L; Priority is given to the shared expert processing module on the same GPU accelerator card as the target routing expert processing module. If the resources of the same GPU accelerator card are insufficient, cross-GPU accelerator card scheduling is performed.

[0011] Furthermore, the subtask splitting rules are as follows: Perform hash calculation on the task category ID in the request data to obtain the hash address; The corresponding subtask splitting rules and fusion coefficients stored in the task coding table are located according to the hash address.

[0012] Furthermore, step S2 further includes the following steps: Read reference fusion coefficients from the preconfigured task encoding table and basic fusion coefficient ; Dynamically calculate the dynamic fusion coefficient according to the model complexity of the current subtask :

[0013] in, is the computational complexity of the current subtask; is the preset maximum load threshold; Dynamic fusion coefficient Fusion coefficient with reference Sent to the output weighted aggregation module.

[0014] Furthermore, in step S5, the formula for weighted aggregation of the processing results by the expert communication control terminal node is as follows:

[0015]

[0016] in, is the final aggregate output; Represents the subtask input features; represents the processing result of the nth shared expert processing module on the subtask input feature, represents the processing result of the nth routing expert processing module on the subtask input feature; S is the total number of shared expert processing modules, and R is the total number of routing expert processing modules; The weight of the current routing expert processing module; is the dynamic fusion coefficient, is the reference fusion coefficient.

[0017] It can be seen from the above technical solutions that this application has the following advantages: The processor system and operating method for large-scale deep learning provided in this application, through the CPU-FPGA-multi-GPU architecture, reduces the CPU burden for large-scale deep learning tasks, improves task distribution efficiency and GPU resource utilization, optimizes communication performance, improves the processing efficiency and system performance of large-scale deep learning tasks, and provides computing power support for the development and application of artificial intelligence. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for the description. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 Schematic diagram of a processor system for large-scale deep learning according to the present invention.

[0020] Figure 2 Schematic diagram of the expert communication control terminal node of the present invention.

[0021] Figure 3 Schematic diagram of the process flow of the processor operation method for large-scale deep learning of the present invention. DETAILED DESCRIPTION

[0022] In the following detailed description of a processor system for large-scale deep learning, various embodiments of the present disclosure will be described more fully. The present disclosure may have various embodiments, and adjustments and changes may be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but rather that the present disclosure should be understood to cover all adjustments, equivalents, and / or alternatives that fall within the spirit and scope of the various embodiments of the present disclosure.

[0023] For example, deep learning, a key branch of machine learning, has become a core technology in artificial intelligence, playing a key role in numerous cutting-edge fields, including image recognition, natural language processing, and autonomous driving. However, as deep learning models continue to expand, the traditional "CPU + heterogeneous GPU" processor architecture has gradually exposed a series of problems when handling large-scale deep learning tasks.

[0024] On the one hand, the CPU burden is increasing. In traditional architectures, the decomposition of deep learning models, the matching and scheduling of subtasks with SM expert cores, and the aggregation of results from multiple expert cores all fall squarely on the CPU. These complex and arduous operations increase the CPU load, disrupting the execution of other important processes and ultimately significantly reducing the user experience.

[0025] On the other hand, task distribution efficiency is far from satisfactory. Due to the limitations of the CPU's inherent concurrent data processing capabilities, task distribution efficiency is low. This not only slows down the overall processing speed of deep learning models, but also prevents the GPU's numerous SM cores from being fully utilized, resulting in an awkward load imbalance and further exacerbating the waste of computing resources.

[0026] Furthermore, communication between the CPU and GPU expert cores presents a bottleneck. Current communication methods primarily rely on traditional heterogeneous transmission, which is bandwidth-constrained and unable to meet the urgent needs of large-scale data transmission. Furthermore, distribution and aggregation operations often need to be performed simultaneously, which undoubtedly increases communication overhead and becomes a major obstacle to performance improvement.

[0027] To address the above issues, this embodiment provides a processor system for large-scale deep learning, which reduces the CPU burden, improves task distribution efficiency and GPU utilization, optimizes the communication mechanism, and improves the efficiency of large-scale deep learning task processing.

[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0029] See also Figure 1 FIG2 is a schematic diagram of a processor system for large-scale deep learning in a specific embodiment, wherein the system includes a host, an expert network-on-chip (E-NoC), and several GPU accelerator cards; The host includes a CPU and an expert communication control terminal node; Expert on-chip network E-NoC, used for communication between the host and each GPU accelerator card; Each GPU accelerator card includes several expert processing modules; the expert processing modules are routing expert processing modules or shared expert processing modules; The host computer and each GPU accelerator card are connected via an expert on-chip network E-NoC; The GPU accelerator card is the actual execution unit of deep learning tasks, and the expert processing module provides powerful parallel computing capabilities for deep learning tasks. By assigning tasks to different expert processing modules, the parallel computing advantages of the GPU can be fully utilized to achieve simultaneous processing of multiple tasks, thereby improving the processing speed and efficiency of deep learning tasks and accelerating the model training and inference process. The CPU is used for system configuration, task deployment, and process monitoring, responding to deep learning model processing requests, and sending them to the expert communication control terminal node; It should be noted that the CPU, as the core of the processor system, is responsible for system configuration, task deployment, and process monitoring. By sending deep learning model processing requests to the expert communication control terminal node, the CPU can focus on system management and coordination, avoiding excessive load caused by processing complex deep learning tasks, thereby ensuring the stable operation of the processor system and the smooth execution of other important processes. The expert communication control endpoint is used to perform the following operations: Split the deep learning model into several subtasks based on the input data characteristics of the deep learning model; Calculate the matching score of the expert processing module; Dispatching subtasks to target expert processing modules; Perform aggregate calculations on the processing results of the expert processing module; Communicate with GPU accelerator card through expert on-chip network E-NoC; It should be noted that the expert communication control terminal node loads key tasks such as splitting the deep learning model, calculating expert matching scores, scheduling subtasks, and aggregating processing results. This reduces the burden on the CPU, making the processing of deep learning tasks more professional, improving processing efficiency, and enhancing the overall performance of the processor system. The expert on-chip network E-NoC realizes fast data transmission between the host and each GPU accelerator card, breaking the communication bottleneck between the CPU and GPU in the traditional architecture, solving the problem of limited communication bandwidth, meeting the needs of large-scale deep learning tasks for large-scale data transmission, reducing communication latency, and improving the communication efficiency and overall performance of the processor system.

[0030] This embodiment provides hardware support for large-scale deep learning tasks through the collaborative work of the host, expert on-chip network and GPU accelerator card. The various components cooperate with each other to improve the efficiency of deep learning task processing.

[0031] Furthermore, as a refinement and extension of the specific implementation of the above embodiment, in order to fully illustrate the specific implementation process in this embodiment, another processor system for large-scale deep learning is provided, which includes a host, an expert on-chip network E-NoC and several GPU accelerator cards; The host includes a CPU and an expert communication control terminal node; Expert on-chip network E-NoC, used for communication between the host and each GPU accelerator card; Each GPU accelerator card includes several expert processing modules; the expert processing modules are routing expert processing modules or shared expert processing modules; It should be noted that the routing expert processing module is mainly responsible for processing certain specific and specialized features in the input; The shared expert processing module is used to capture common and global knowledge and provide basic feature extraction support for all inputs; The host computer and each GPU accelerator card are connected via an expert on-chip network E-NoC; Specifically, the host and GPU accelerator card are connected to the expert network-on-chip E-NoC via a PCIe interface equipped with RDMA; The CPU is used for system configuration, task deployment, and process monitoring, responding to deep learning model processing requests, and sending them to the expert communication control terminal node; The expert communication control endpoint is used to perform the following operations: Split deep learning model processing requests into several subtasks based on the input data characteristics of the deep learning model; Calculate the matching score of the expert processing module; Dispatching subtasks to target expert processing modules; Perform aggregate calculations on the processing results of the expert processing module; Communicate with GPU accelerator card through expert on-chip network E-NoC; like Figure 2 As shown, the expert communication control terminal node uses an FPGA core; it includes: The request receiving module receives the deep learning model processing request sent by the CPU, parses the request data and transmits it to the deep learning task splitting and comparison module; Deep learning task splitting and comparison module, used for: According to the task category in the request data, the corresponding subtask splitting rules and expert model fusion coefficients are obtained from the pre-configured task coding table; Send the sub-task splitting rule and the original model data in the request data to the expert matching gate module; Send the fusion coefficient to the output weighted aggregation module; The expert matching gate module is used for: According to the sub-task splitting rule, split the input data of the deep learning model into several sub-tasks; Calculate the matching scores of each sub-task with the expert processing modules of the GPUs, and select the target expert processing module; Select the target expert processing module and generate a sub-task allocation instruction; The expert matching gate module includes: The sub-task splitting unit is used to split the input data of the deep learning model into sub-tasks for parallel processing according to the sub-task splitting rule by feature dimension or calculation granularity; The expert status reading unit is used to obtain the busy and idle status of the expert processing modules of each GPU in real time; The routing expert matching unit is used to perform the following operations: Calculate the matching scores of each sub-task with the routing expert processing module, and generate a Top-K candidate expert list; Combined with the busy and idle status of the expert processing module, select an idle and most highly matched routing expert processing module from the Top-K candidate expert list as the target routing expert processing module; The shared expert matching unit is used to perform the following operations: Dynamically determine the number of shared experts to be called according to the number of layers of the deep learning model to which the sub-task belongs; Exemplarily, the determination rule for the number of shared experts N is: If the number of model layers L ≤ 5, then N = 1; If 5 < L ≤ 10, then N = ②; If L > 10, then N = ceil(L / 5); ceil is the ceiling function; Preferentially select the shared expert processing module located on the same GPU acceleration card as the target routing expert processing module as the target shared expert processing module; if there is no idle shared expert processing module on the same GPU acceleration card, call it across GPU acceleration cards through the expert on-chip network E-NoC; The shared expert matching unit selects the shared expert processing module according to the priority: The first priority is the idle shared expert processing module on the same GPU acceleration card as the target routing expert processing module; The second priority is the shared expert processing module on the same GPU acceleration card that can be quickly released within a set time period; The third priority is the idle shared expert processing module of different GPU accelerator cards; The expert matching and selection FSM unit is used to generate the final subtask allocation instruction according to the matching results of the routing expert processing module and the shared expert processing module; The task transmission module is used to send the subtask data to the target expert processing module through the expert on-chip network E-NoC according to the subtask allocation instructions; The expert processing result receiving module is used to receive the calculation results returned by the GPU expert processing module and cache them in the output weighted aggregation module; The output weighted aggregation module is used to perform weighted aggregation on the calculation results of each cached GPU expert processing module according to the fusion coefficient to generate the final output; The response sending module is used to return the final output to the CPU.

[0032] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0033] like Figure 3 As shown, the following is an embodiment of the processor operation method for large-scale deep learning provided by the embodiment of the present disclosure. This method and the processor operation system for large-scale deep learning in the above-mentioned embodiments belong to the same inventive concept. For details not fully described in the embodiment of the processor operation method for large-scale deep learning, please refer to the embodiment of the processor operation system for large-scale deep learning.

[0034] The method comprises the following steps: S1. The host CPU receives the deep learning model processing request and sends the request to the expert communication control terminal node; S2. The expert communication control terminal node decomposes the deep learning model processing request into several subtasks, calculates the matching score between each subtask and the GPU's expert processing module, and selects the target expert processing module; S3. The expert communication control terminal node sends the subtask to the target expert processing module through the expert on-chip network E-NoC; S4. The target expert processing module processes the subtask and returns the processing results to the expert communication control terminal node; S5. The expert communication control terminal node performs weighted aggregation on the processing results and sends the final result back to the CPU.

[0035] This embodiment ensures the execution efficiency of the entire process of deep learning tasks from request reception to result return, thereby improving user experience.

[0036] Further, as a refinement and extension of the specific implementation of the above embodiment, in order to fully illustrate the specific implementation process in this embodiment, another method for a processor to run for large-scale deep learning is provided. The method includes the following steps: S1. The CPU of the host receives a deep learning model processing request and sends the request to the expert communication control terminal node; S2. The expert communication control terminal node decomposes the deep learning model processing request into several subtasks, calculates the matching scores of each subtask with the expert processing modules of the GPUs, and selects the target expert processing module; The specific steps of step S2 are as follows: According to the task category in the request data, obtain the subtask splitting rule and the fusion coefficient from the pre-configured task encoding table; Split the input data of the deep learning model into several subtasks according to the feature dimension or the computing granularity according to the subtask splitting rule; Perform the following operations for each subtask: Calculate the matching score with the routing expert processing module and generate a Top-K candidate list; Select the idle and most highly matched routing expert processing module from the candidate list as the target routing expert processing module; Determine the number N of shared expert processing modules according to the number of layers L of the deep learning model; Specifically as follows: If L ≤ 5, then N = 1; If 5 < L ≤ 10, then N = 2; If L > 10, then N = ceil(L / 5); Preferentially select the shared expert processing modules on the same GPU acceleration card as the target routing expert processing module. If the resources of the same GPU acceleration card are insufficient, schedule across GPU acceleration cards; The subtask splitting rule is as follows: Perform a hash calculation on the task category ID in the request data to obtain a hash address; Specifically, the hash calculation uses the CRC32 algorithm to generate a 32-bit hash value; Locate the corresponding subtask splitting rule and fusion coefficient stored in the task encoding table according to the hash address; Specifically, the task encoding table uses a hash linked list structure to handle address conflicts; Step S2 further includes the following steps: Read the reference fusion coefficient from the pre-configured task encoding table and the base fusion coefficient ; Dynamically calculate the dynamic fusion coefficient according to the model complexity of the current subtask :

[0037] in, is the computational complexity of the current subtask; is the preset maximum load threshold; Dynamic fusion coefficient Fusion coefficient with reference Send to the output weighted aggregation module; It should be noted that the computational complexity of the current subtask is Obtained by counting the floating-point operations (FLOPs) of the subtask or linearly estimated based on the number of model layers involved in the subtask; S3. The expert communication control terminal node sends the subtask to the target expert processing module through the expert on-chip network E-NoC; S4. The target expert processing module processes the subtask and returns the processing results to the expert communication control terminal node; S5. The expert communication control terminal node performs weighted aggregation on the processing results and sends the final result back to the CPU; the formula for weighted aggregation of the processing results by the expert communication control terminal node in step S5 is as follows:

[0038]

[0039] in, is the final aggregate output; Represents the subtask input features; represents the processing result of the nth shared expert processing module on the subtask input feature, represents the processing result of the nth routing expert processing module on the subtask input feature; S is the total number of shared expert processing modules, and R is the total number of routing expert processing modules; The weight of the current routing expert processing module; is the dynamic fusion coefficient, is the reference fusion coefficient.

[0040] The processor operation method for large-scale deep learning provided in the embodiment of the present application can be applied to electronic devices. Those skilled in the art will understand that the electronic device structure involved in the embodiment of the present invention does not constitute a limitation on the electronic device, and the electronic device may include more or fewer components than shown, or combine certain components, or arrange components differently. In the embodiment of the present invention, the electronic device includes but is not limited to a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described and / or required herein.

[0041] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, a button, a camera, a display, and a SIM card interface, etc.

[0042] It is understood that the structures illustrated in the embodiments of the present application do not constitute specific limitations on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown, or combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0043] A processor may include one or more processing units, such as a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.

[0044] The processor can be the nerve center and command center of the electronic device. The controller can generate operation control signals based on the instruction opcode and timing signal to complete the control of instruction fetching and execution.

[0045] The processor may also include a memory for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can store instructions or data that the processor has just used or is reusing. If the processor needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces processor latency, and thus improves system efficiency.

[0046] The above-mentioned electronic device implements the processor operation method for large-scale deep learning in the present application, in which the CPU of the host receives a deep learning model processing request and sends the request to the expert communication control terminal node; the expert communication control terminal node decomposes the deep learning model processing request into several subtasks, calculates the matching score of each subtask with the expert processing module of the GPU, and selects the target expert processing module; the expert communication control terminal node sends the subtask to the target expert processing module through the expert on-chip network E-NoC; the target expert processing module processes the subtask and returns the processing result to the expert communication control terminal node; the expert communication control terminal node performs weighted aggregation on the processing result and sends the final result back to the CPU. The technical solution achieves the beneficial effect of reducing the CPU burden, improving task distribution efficiency and GPU resource utilization, optimizing the communication mechanism, and improving the efficiency of large-scale deep learning task processing through the CPU-FPGA-multi-GPU architecture.

[0047] The storage medium provided in this application stores a program product that can implement a processor operation method for large-scale deep learning.

[0048] A processor operation method for large-scale deep learning includes: the host's CPU receives a deep learning model processing request and sends the request to an expert communication control terminal node; the expert communication control terminal node decomposes the deep learning model processing request into several subtasks, calculates the matching score of each subtask with the GPU's expert processing module, and selects a target expert processing module; the expert communication control terminal node sends the subtask to the target expert processing module through the expert on-chip network E-NoC; the target expert processing module processes the subtask and returns the processing result to the expert communication control terminal node; the expert communication control terminal node performs weighted aggregation on the processing result and sends the final result back to the CPU.

[0049] In some possible implementations, the processor operation method for large-scale deep learning disclosed herein can be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps of various exemplary implementations of the present disclosure described in the above "Exemplary Method" section of this specification.

[0050] The storage medium of the present disclosure can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0051] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A processor system for large-scale deep learning, characterized in that: Includes a host, an expert on-chip network E-NoC, and several GPU accelerator cards; The host includes a CPU and an expert communication control terminal node; Expert on-chip network E-NoC, used for communication between the host and each GPU accelerator card; Each GPU accelerator card includes several expert processing modules; the expert processing modules are routing expert processing modules or shared expert processing modules; The host computer and each GPU accelerator card are connected via an expert on-chip network E-NoC; The CPU is used for system configuration, task deployment, and process monitoring, responding to deep learning model processing requests, and sending them to the expert communication control terminal node; The expert communication control endpoint is used to perform the following operations: Based on the input data characteristics of the deep learning model, the deep learning model processing request is divided into several subtasks; Calculate the matching score of the expert processing module; Dispatching subtasks to target expert processing modules; Perform aggregate calculations on the processing results of the expert processing module; Communicate with the GPU accelerator card through the expert on-chip network E-NoC.

2. The processor system for large-scale deep learning according to claim 1, characterized in that Expert communication control terminal nodes include: The request receiving module receives the deep learning model processing request sent by the CPU, parses the request data and transmits it to the deep learning task splitting and comparison module; Deep learning task splitting and comparison module, used for: According to the task category in the request data, the corresponding subtask splitting rules and expert model fusion coefficients are obtained from the pre-configured task coding table; Send the subtask splitting rules and the original model data in the request data to the expert matching gate module; Send the fusion coefficients to the output weighted aggregation module; Expertly matched door modules for: According to the subtask splitting rules, the input data of the deep learning model is split into several subtasks; Calculate the matching score between each subtask and the GPU's expert processing module, and select the target expert processing module; Select the target expert processing module and generate subtask assignment instructions; The task transmission module is used to send the subtask data to the target expert processing module through the expert on-chip network E-NoC according to the subtask allocation instructions; The expert processing result receiving module is used to receive the calculation results returned by the GPU expert processing module and cache them in the output weighted aggregation module; The output weighted aggregation module is used to perform weighted aggregation on the calculation results of each cached GPU expert processing module according to the fusion coefficient to generate the final output; The response sending module is used to return the final output to the CPU.

3. The processor system for large-scale deep learning according to claim 2, characterized in that Expert matching door modules include: A subtask splitting unit is used to split the input data of the deep learning model into subtasks for parallel processing according to feature dimensions or calculation granularity based on the subtask splitting rules; Expert status reading unit, used to obtain the busy and idle status of each GPU's expert processing module in real time; The routing expert matching unit is used to perform the following operations: Calculate the matching score between each subtask and the routing expert processing module and generate a list of Top-K candidate experts; Combined with the busy and idle status of the expert processing module, a routing expert processing module with the highest matching degree and an idle status is selected from the Top-K candidate expert list as the target routing expert processing module; Shared expert matching unit, used to perform the following operations: Dynamically determine the number of shared experts to be called based on the number of deep learning model layers to which the subtask belongs; Prioritize the shared expert processing module located on the same GPU accelerator card as the target routing expert processing module as the target shared expert processing module; if there is no idle shared expert processing module on the same GPU accelerator card, call it across GPU accelerator cards through the expert on-chip network E-NoC; The expert matching and selection FSM unit is used to generate the final subtask allocation instruction according to the matching results between the routing expert processing module and the shared expert processing module.

4. The processor system for large-scale deep learning according to claim 3, characterized in that The shared expert matching unit selects the shared expert processing module according to the priority level: The first priority is the idle shared expert processing module on the same GPU accelerator card as the target routing expert processing module; The second priority is the shared expert processing module of the same GPU accelerator card that can be quickly released within a set time period; The third priority is the idle shared expert processing module of different GPU accelerator cards.

5. The processor system for large-scale deep learning according to claim 1, characterized in that The expert communication control terminal node adopts FPGA core.

6. A processor operation method for large-scale deep learning, characterized in that: The steps include: S1. The host CPU receives the deep learning model processing request and sends the request to the expert communication control terminal node; S2. The expert communication control terminal node decomposes the deep learning model processing request into several subtasks, calculates the matching score between each subtask and the GPU's expert processing module, and selects the target expert processing module; S3. The expert communication control terminal node sends the subtask to the target expert processing module through the expert on-chip network E-NoC; S4. The target expert processing module processes the subtask and returns the processing results to the expert communication control terminal node; S5. The expert communication control terminal node performs weighted aggregation on the processing results and sends the final result back to the CPU.

7. The processor operation method for large-scale deep learning according to claim 6, characterized in that: The specific steps of step S2 are as follows: According to the task category in the request data, obtain the subtask splitting rules and fusion coefficients from the pre-configured task coding table; Split the deep learning model input data into several subtasks according to feature dimensions or calculation granularity based on the subtask splitting rules; For each subtask, do the following: Calculate the matching score with the routing expert processing module and generate a Top-K candidate list; Selecting an idle routing expert processing module with the highest matching degree from the candidate list as the target routing expert processing module; Determine the number of shared expert processing modules N based on the number of deep learning model layers L; Priority is given to the shared expert processing module on the same GPU accelerator card as the target routing expert processing module. If the resources of the same GPU accelerator card are insufficient, cross-GPU accelerator card scheduling is performed.

8. The processor operation method for large-scale deep learning according to claim 7, characterized in that: The subtask splitting rules are as follows: Perform hash calculation on the task category ID in the request data to obtain the hash address; The corresponding subtask splitting rules and fusion coefficients stored in the task coding table are located according to the hash address.

9. The processor operation method for large-scale deep learning according to claim 7, characterized in that: Step S2 also includes the following steps: Read reference fusion coefficients from the preconfigured task encoding table and basic fusion coefficient ; Dynamically calculate the dynamic fusion coefficient according to the model complexity of the current subtask : in, is the computational complexity of the current subtask; is the preset maximum load threshold; Dynamic fusion coefficient Fusion coefficient with reference Sent to the output weighted aggregation module.

10. The processor operation method for large-scale deep learning according to claim 9, characterized in that: The formula for weighted aggregation of the processing results by the expert communication control terminal node in step S5 is as follows: in, is the final aggregate output; Represents the subtask input features; represents the processing result of the nth shared expert processing module on the subtask input feature, represents the processing result of the nth routing expert processing module on the subtask input feature; S is the total number of shared expert processing modules, and R is the total number of routing expert processing modules; The weight of the current routing expert processing module; is the dynamic fusion coefficient, is the reference fusion coefficient.

Citation Information

Patent Citations

  • Scalable transfer learning with expert model

    CN115699041A

  • Hybrid expert model acceleration method and system based on locality sensitive hashing algorithm

    CN117932330A

  • Deep learning model optimization method for large-scale distributed training

    CN118535340A

  • Computing system and method for GPU (Graphics Processing Unit) computing power scheduling

    CN119645661A

  • Expert network automatic hybrid distributed algorithm based on cloud side end cross-domain data

    CN119690274A

Cited By

  • Computer system, server optimization method, electronic device and storage medium

    CN121217598A