Model reasoning acceleration method and device, storage medium and program product

By utilizing the computing resources of the main core and co-core in the substrate management controller, the model inference speed and real-time improvement are achieved, and the problem that hardware acceleration cards cannot be applied is solved.

CN120354956AActive Publication Date: 2025-07-22INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510858103.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-22
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

Due to cost, power consumption and space limitations, the substrate management controller cannot use hardware acceleration cards, resulting in slow model inference speed and insufficient real-time performance.

Method used

Using the computing resources of the main core and co-core, through the interrupt response mechanism, we determine whether the current operator is an operator to be unloaded. If so, the first calculation operation is performed on the main core. If otherwise, the second calculation operation is performed on the co-core, and the final result is generated on the main core.

Benefits of technology

Improve the model inference speed and real-time performance, make full use of main core and co-core resources, and solve the performance limitations caused by disabling external hardware acceleration cards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354956A_ABST
    Figure CN120354956A_ABST
Patent Text Reader

Abstract

The invention discloses a model reasoning acceleration method and device, a storage medium and a program product, and relates to the technical field of artificial intelligence, and the method comprises the steps: according to whether a current operator is a to-be-unloaded operator, if yes, executing all calculation operations on a calculation task of the current operator on a main core, and outputting a calculation execution result of the current operator; and if not, executing a first calculation operation on the calculation task of the current operator on the main core, outputting a first execution result of the current operator, executing a second calculation operation on the calculation task of the current operator on the auxiliary core, and outputting a second execution result of the current operator. Therefore, the technical effect of improving the reasoning speed of the model is achieved. And on the basis of an interrupt response mechanism, according to the first execution result and the second execution result, a calculation execution result of the current operator is generated on the main core, and the technical effect of improving the real-time performance of model reasoning is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, device, storage medium, and program product for accelerating model inference. Background Art

[0002] Traditional baseboard management controllers rely on low-power central processing units to run model inference. However, limited by their computing power, it is difficult to meet real-time requirements. Hardware acceleration cards are one of the mainstream technologies for accelerating artificial intelligence inference. The core idea is to offload computationally intensive tasks through dedicated hardware, such as neural network processing units, field-programmable gate arrays, and graphics processing units, thereby significantly improving performance. In high-performance scenarios such as servers and edge computing, such solutions have been widely applied. However, hardware acceleration cards have the following fatal defects in the baseboard management controller scenario: (I) Cost and hardware compatibility issues: 1. High hardware cost; 2. Interface compatibility limitations; 3. Physical space conflicts. (II) Power consumption and heat dissipation limitations: 1. Excessive power consumption; 2. Complex heat dissipation design. (III) Software ecosystem and maintenance costs: 1. Fragmented toolchains; 2. Poor model compatibility; 3. Firmware upgrade risks. Summary of the Invention

[0003] This application provides a method, device, storage medium, and program product for accelerating model inference, to at least solve the problem that in related technologies, due to cost, power consumption, space, and other limiting factors, the baseboard management controller cannot use a hardware acceleration card, making it difficult to accelerate model inference.

[0004] This application provides a method for accelerating model inference, which is applied to a baseboard management controller. The baseboard management controller includes a main core and a co-core. The method for accelerating model inference includes: In response to receiving a model inference request, load a preset model on the main core according to the model inference request, and determine whether the current operator is an operator to be offloaded; If not, perform all calculation operations on the current operator on the main core according to the input tensor of the current operator, and output the calculation execution result of the current operator; If so, perform a first calculation operation on the current operator on the main core according to the input tensor of the current operator, and output the first execution result of the current operator. At the same time, perform a second calculation operation on the current operator on the co-core, and output the second execution result of the current operator, where all calculation operations at least include the first calculation operation and the second calculation operation; Based on the interrupt response mechanism, generate the calculation execution result of the current operator on the main core according to the first execution result and the second execution result of the current operator.

[0005] The present application also provides an electronic device, including: a memory for storing a computer program; a processor for, when executing the computer program, at least implementing a model inference acceleration method including the following steps: In response to receiving a model inference request, load a preset model on the main core according to the model inference request, and determine whether the current operator is an operator to be unloaded; If not, perform all calculation operations on the current operator on the main core according to the input tensor of the current operator, and output the calculation execution result of the current operator; If so, perform a first calculation operation on the current operator on the main core according to the input tensor of the current operator, and output the first execution result of the current operator, and at the same time perform a second calculation operation on the current operator on the co-core, and output the second execution result of the current operator, where all calculation operations at least include the first calculation operation and the second calculation operation; Based on the interrupt response mechanism, generate the calculation execution result of the current operator on the main core according to the first execution result and the second execution result of the current operator.

[0006] The present application also provides a computer-readable storage medium, in which a computer program is stored, and where, when the computer program is executed by a processor, at least implementing a model inference acceleration method including the following steps: In response to receiving a model inference request, load a preset model on the main core according to the model inference request, and determine whether the current operator is an operator to be unloaded; If not, perform all calculation operations on the current operator on the main core according to the input tensor of the current operator, and output the calculation execution result of the current operator; If so, perform a first calculation operation on the current operator on the main core according to the input tensor of the current operator, and output the first execution result of the current operator, and at the same time perform a second calculation operation on the current operator on the co-core, and output the second execution result of the current operator, where all calculation operations at least include the first calculation operation and the second calculation operation; Based on the interrupt response mechanism, generate the calculation execution result of the current operator on the main core according to the first execution result and the second execution result of the current operator.

[0007] The present application also provides a computer program product, including a computer program, and where, when the computer program is executed by a processor, at least implementing a model inference acceleration method including the following steps: In response to receiving a model inference request, load a preset model on the main core according to the model inference request, and determine whether the current operator is an operator to be unloaded; If not, perform all calculation operations on the current operator on the main core according to the input tensor of the current operator, and output the calculation execution result of the current operator; If so, according to the input tensor of the current operator, perform a first calculation operation on the current operator on the main core, output the first execution result of the current operator, and at the same time perform a second calculation operation on the current operator on the co-core, output the second execution result of the current operator, where all calculation operations at least include the first calculation operation and the second calculation operation; Based on the interrupt response mechanism, generate the calculation execution result of the current operator on the main core according to the first execution result and the second execution result of the current operator.

[0008] Through the present application, according to whether the current operator is an operator to be offloaded, if so, perform all calculation operations on the calculation task of the current operator on the main core, and output the calculation execution result of the current operator; if not, perform a first calculation operation on the calculation task of the current operator on the main core, output the first execution result of the current operator, and at the same time perform a second calculation operation on the calculation task of the current operator on the co-core, output the second execution result of the current operator, which solves the problem of slow inference speed due to disabling the external hardware acceleration card and limited performance itself, and achieves the technical effect of improving the model inference speed; based on the interrupt response mechanism, generate the calculation execution result of the current operator on the main core according to the first execution result and the second execution result, and achieve the technical effect of improving the real-time performance of model inference. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1 It is a flowchart of a model inference acceleration method in an embodiment; Figure 2 It is a flowchart of the initialization of a model inference acceleration system in an embodiment; Figure 3 It is a flowchart of a model inference acceleration method on the main core in an embodiment; Figure 4 It is a flowchart of a model inference acceleration method on the co-core in an embodiment; Figure 5 It is a flowchart of aggregating the first execution result and the second execution result on the main core in an embodiment; Figure 6 It is a schematic diagram of the architecture of a model inference acceleration system in an embodiment; Figure 7 It is an internal structure diagram of an electronic device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0011] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0012] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variation thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0013] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0014] The hardware acceleration card has the following fatal defects in the baseboard management controller scenario: (1) Cost and hardware compatibility issues: 1. Excessive hardware cost: The price of a low-power neural network processing unit acceleration card far exceeds the overall hardware cost of the baseboard management controller. The proportion of the acceleration card cost therein exceeds 300%, which completely does not conform to the cost-sensitive positioning of the baseboard management controller. 2. Interface compatibility limitations: The baseboard management controller chip usually only provides low-speed interfaces, such as low-pin-count bus interfaces and enhanced serial peripheral interfaces, etc., and cannot support high-speed buses such as PCIe 3.0 / 4.0, resulting in limited bandwidth of the acceleration card. Specifically, the measured bandwidth of PCIe 2.0 x1 is only 500MB / s, which cannot meet the data throughput requirements of the neural network processing unit. 3. Physical space conflict: The common size of the baseboard management controller board is 5cm×5cm and is compact, and it cannot accommodate a standard-size acceleration card such as the M.2 2280 specification. Forced integration requires re-designing the printed circuit board layout, increasing the development cycle and risk.

[0015] (2) Power consumption and heat dissipation limitations: 1. Excessive power consumption: The power consumption of a typical neural network processing unit acceleration card is 2-4W, while the overall power consumption of the baseboard management controller usually needs to be less than 10W. The acceleration card will occupy 30%-40% of the power consumption budget therein, resulting in insufficient power supply for other key functions, such as fan control and sensor acquisition, etc. 2. Complexity of heat dissipation design: Baseboard Management Controllers (BMCs) usually rely on passive heat dissipation or radiatorless designs. However, when the accelerator card is running, it will significantly increase the board-level temperature. It has been measured that the operating temperature of the Neural Network Processing Unit (NNPU) can reach 70 °C or higher, which may trigger the over-temperature protection mechanism of the BMC and cause the system to crash.

[0016] (III) Software ecosystem and maintenance cost: 1. Fragmentation of the toolchain: Accelerator cards from different manufacturers need to be independently adapted to the software stack. For example, the NNPU requires a dedicated compiler, and the Field-Programmable Gate Array (FPGA) requires customized drivers, etc., resulting in an exponential increase in the complexity of BMC firmware development. For example, supporting two different accelerator cards simultaneously requires maintaining two code libraries, increasing the testing and maintenance costs; 2. Poor model compatibility: Accelerator cards usually only support specific operator types. For example, the NNPU is good at convolution but does not support dynamic control flow, resulting in the inability of long-tail models such as custom Long Short-Term Memory Neural Networks to run efficiently, and a large amount of manual optimization or even model reconstruction is required; 3. Firmware upgrade risk: The drivers and firmware of the hardware accelerator card need to be synchronized with the BMC firmware for upgrade. Once there is a version mismatch, such as an outdated driver version of the NNPU, it may cause the inference task to crash and affect the reliability of the server management function.

[0017] In one embodiment, as Figure 1 shown, a model inference acceleration method is provided, which is applied to a Baseboard Management Controller. The Baseboard Management Controller includes a main core and a co-core. The model inference acceleration method includes: In response to receiving a model inference request, load a preset model on the main core according to the model inference request, and determine whether the current operator is an operator to be unloaded; If not, perform all calculation operations on the current operator on the main core according to the input tensor of the current operator, and output the calculation execution result of the current operator; If so, perform a first calculation operation on the current operator on the main core according to the input tensor of the current operator, and output the first execution result of the current operator. At the same time, perform a second calculation operation on the current operator on the co-core, and output the second execution result of the current operator, where all calculation operations at least include the first calculation operation and the second calculation operation; Based on the interrupt response mechanism, generate the calculation execution result of the current operator on the main core according to the first execution result and the second execution result of the current operator.

[0018] In specific implementations, most current baseboard management controllers are built with coprocessors, i.e., the co - cores in this embodiment. However, these co - cores are usually only used for specific functions, such as running real - time operating systems and processing low - speed tasks, and thus do not participate in model inference, resulting in a situation where the main core and the co - core act independently and cannot share computing power. For example, when the main core is running a model at full load, the co - core is idle, leading to a low overall energy efficiency ratio. Therefore, in this embodiment, according to whether the current operator is an operator to be offloaded, the computing tasks of non - offloaded operators are computed by the main core of the baseboard management controller; the computing tasks of offloaded operators are partially offloaded from the main core of the baseboard management controller to the co - core of the baseboard management controller and computed in parallel by the main core and the co - core; according to the interrupt response mechanism, it is determined whether to integrate the first execution result and the second execution result on the main core of the baseboard management controller. That is, the main core of the baseboard management controller in this embodiment is responsible for processing the control flow and complex operators, and the coprocessor of the baseboard management controller focuses on rule calculation, making full use of the computing resources of the main core and the co - core in the baseboard management controller, achieving the optimal balance of computing power and energy efficiency, and thus accelerating the model inference speed of the baseboard management controller.

[0019] Specifically, according to whether the current operator is an operator to be offloaded, if so, all computing operations of the computing task of the current operator are performed on the main core, and the computing execution result of the current operator is output; if not, the first computing operation of the computing task of the current operator is performed on the main core, and the first execution result of the current operator is output. At the same time, the second computing operation of the computing task of the current operator is performed on the co - core, and the second execution result of the current operator is output, solving the problem of slow inference speed due to disabling the external hardware acceleration card and limited performance itself, achieving the technical effect of improving the model inference speed; based on the interrupt response mechanism, according to the first execution result and the second execution result, the computing execution result of the current operator is generated on the main core, achieving the technical effect of making full use of the resources of the main core and the co - core and improving the real - time performance of model inference.

[0020] As Figure 2 shown, before receiving a model inference request, the model inference acceleration method further includes: (1) System startup: After power - on, both the main core and the co - core will start the firmware system pre - burned into the storage device. After the main - core system startup is completed, the system resources will be initialized and the model inference process will be started. After the model inference process is started, a thread pool will be created and wait for the arrival of model inference tasks; (2) Hardware resource mapping: 1. Shared memory area division. When the main core starts, the memory management unit maps a fixed-size space of the configured physical address into shared memory, and configures the cache policy for both the main core side and the co-core side to ensure data consistency; 2. Interrupt controller configuration. Register the interrupt number of the co-core. For example, bind the interrupt request of the co-core to the interrupt line numbered 32, so that the main core knows which interrupt comes from the co-core, and set the interrupt priority of the co-core to be higher than IPMI (Intelligent Platform Management Interface, a standard interface for managing hardware devices in servers and other computer systems) communication but lower than temperature alarm, and enable the edge trigger mode; (3) Thread pool construction: Start the model inference process, and the process creates all working threads according to the configuration, including main core threads and co-core threads.

[0021] In addition, before determining whether the current operator is an operator to be offloaded, the model inference acceleration method further includes: Identify multiple offloadable operators on the main core for the computation graph of the preset model; If the current offloadable operator is a memory-intensive operator, then determine on the main core that the current offloadable operator is not an operator to be offloaded; If the current offloadable operator is a computation-intensive operator, then determine on the main core that the current offloadable operator is an operator to be offloaded.

[0022] In a specific implementation, in this embodiment, the main core traverses the computation graph of the preset model to identify offloadable operators; according to the operator type of the current offloadable operator, determine whether the current offloadable operator is an operator to be offloaded, and mark the input / output tensor addresses of the operator to be offloaded. Computation-intensive operators include matrix multiplication matrix operators and matrix multiplication vector operators, etc., which can be split into independent subtasks so that the main core can handle the control flow and complex operators, and the co-core focuses on regular calculations to achieve the optimal balance of computing power and energy efficiency.

[0023] In addition, according to the input tensor of the current operator, perform all calculation operations on the current operator on the main core, and output the calculation execution result of the current operator, including: Generate the calculation task of the current operator on the main core according to the current operator and the input tensor of the current operator, and write the calculation task of the current operator into the main core thread of the thread pool; Process the calculation task of the current operator in the main core thread on the main core, output the calculation execution result of the current operator, and write the calculation execution result of the current operator into the double-buffer task queue in the shared memory.

[0024] Furthermore, as Figure 3As shown, before performing the first calculation operation on the current operator on the main core according to the input tensor of the current operator and outputting the first execution result of the current operator, the model inference acceleration method further includes: On the main core, cut the calculation task of the current operator into corresponding multiple calculation subtasks, and write the multiple calculation subtasks of the current operator into a double-buffered task queue, where the amount of calculation between the multiple calculation subtasks is the same; Based on the task allocation mechanism, divide the multiple calculation subtasks of the current operator on the main core into the main core calculation subtasks and co-core calculation subtasks of the current operator; If the current calculation subtask of the current operator is a main core calculation subtask, write the current calculation subtask of the current operator into the main core thread on the main core; If the current calculation subtask of the current operator is a co-core calculation subtask, write the current calculation subtask of the current operator into the co-core thread of the thread pool on the main core; Generate a task descriptor of the current operator on the main core according to the input tensor of the current operator, and write the task descriptor of the current operator into the task descriptor queue in the shared memory.

[0025] In a specific implementation, this embodiment first updates the head pointer of the task descriptor queue, and then writes the task descriptor of the current operator into the circular buffer of the task descriptor queue. If the double-buffered task queue is full, the main core further reduces the task granularity to quickly free up queue space. For example, if each calculation subtask of the current operator includes 1024 elements, each calculation subtask is further split into new calculation subtasks including 256 elements to enter the queue faster and reduce queue backlog.

[0026] Specifically, split the calculation task of the current operator into corresponding multiple calculation subtasks, and write the multiple calculation subtasks of the current operator into a double-buffered task queue. Each calculation subtask of the current operator in the double-buffered task queue is processed by the main core thread or the co-core thread, making full use of the computing resources of the co-core in the baseboard management controller and improving the model inference speed.

[0027] Further, performing the first calculation operation on the current operator on the main core and outputting the first execution result of the current operator includes: Perform the first calculation operation on the main core calculation subtasks of the current operator in the main core thread, output the first execution result of the current operator, and write the first execution result of the current operator into the double-buffered task queue; Perform the second calculation operation on the current operator on the co-core to output the second execution result of the current operator, including: Obtain the task descriptor of the current operator in the task descriptor queue on the co-core; According to the task descriptor of the current operator, perform a second calculation operation on the co-kernel calculation subtask of the current operator in the co-kernel thread of the co-kernel, output the second execution result of the current operator, and write the second execution result of the current operator into the double-buffer task queue.

[0028] In a specific implementation, the double-buffer task queue realizes the parallelization of data production and consumption by alternately using two memory areas. The core process is that when the main core writes to the current buffer, the co-kernel processes the spare buffer; when the co-kernel completes the processing of the spare buffer, it switches to processing the current buffer, and at the same time the main core fills the spare buffer. The main core thread fills the tasks into the current buffer and then switches to the spare buffer to continue writing the next batch of tasks; the co-kernel polls the buffer switch flag bit and reads the data of the corresponding buffer according to the buffer switch flag bit to avoid read-write conflicts. The co-kernel writes the execution result into the spare buffer, updates the buffer switch flag bit, and at the same time updates the co-kernel load rate in the status register to the current load rate of the co-kernel.

[0029] Furthermore, based on the task allocation mechanism, divide multiple calculation subtasks of the current operator into co-kernel calculation subtasks of the current operator on the main core, including: Obtain the load rate of the co-kernel thread from the status register in the shared memory on the main core, and obtain the co-kernel calculation ratio of the current operator according to the load rate of the co-kernel thread; According to the calculation task of the current operator, obtain the total calculation amount of the current operator on the main core, and obtain the co-kernel calculation amount of the current operator according to the total calculation amount of the current operator and the co-kernel calculation ratio; According to multiple calculation subtasks of the current operator, obtain the average calculation amount of the current operator on the main core, and obtain the number of co-kernel subtasks of the current operator by rounding down according to the co-kernel calculation amount and the average calculation amount of the current operator; According to the number of co-kernel subtasks of the current operator, divide the corresponding number of calculation subtasks of the current operator into co-kernel calculation subtasks of the current operator on the main core.

[0030] As Figure 3 shown, in this embodiment, first initialize the dynamic scheduler and the load monitoring module, read the load rate of the co-kernel in the shared memory status register every 10 ms to dynamically adjust the task sharding strategy. If the load rate of the co-kernel is less than 30%, unload 70% of the total calculation amount of the current operator to the co-kernel; if the load rate of the co-kernel is greater than 80%, unload 20% of the total calculation amount of the current operator to the co-kernel to reduce the latency. By sharding the calculation tasks of the current operator, the computing resources of the main core and the co-kernel are fully utilized, and the model inference speed is improved.

[0031] Furthermore, based on the task allocation mechanism, divide multiple calculation subtasks of the current operator into main-core calculation subtasks of the current operator on the main core, including: Obtain the total number of subtasks of the current operator on the main core according to multiple computational subtasks of the current operator; On the main core, obtain the number of main-core subtasks of the current operator according to the total number of subtasks of the current operator and the number of co-core subtasks; According to the number of main-core subtasks of the current operator, divide the corresponding number of computational subtasks of the current operator into main-core computational subtasks of the current operator on the main core. Further, according to the input tensor of the current operator, generate a task descriptor of the current operator on the main core, including: Obtain the input tensor address, output tensor address, and operator type of the current operator on the main core, and generate a cyclic check code of the current operator according to the input tensor address, output tensor address, and operator type of the current operator; Generate a task descriptor of the current operator on the main core according to the input tensor address, output tensor address, operator type, and cyclic check code of the current operator.

[0032] Specifically, the main core generates a task descriptor of the current operator according to the input tensor address, output tensor address, operator type, and cyclic check code of the current operator, so that the co-core processes the co-core computational subtasks of the current operator in the double-buffer task queue according to the task descriptor, improving the model inference efficiency.

[0033] Further, obtain the task descriptor of the current operator in the task descriptor queue on the co-core, including: Verify the task descriptor of the current operator on the co-core according to the cyclic check code of the current operator; If the verification is successful, obtain the task descriptor of the current operator in the task descriptor queue on the co-core; If the verification fails, generate an error code and an interrupt request on the co-core; Write the error code into the error field in the status register on the co-core and send the interrupt request to the main core.

[0034] As Figure 4 shown, the co-core reads the task descriptor from the end of the circular buffer in the task descriptor queue. After verifying the cyclic check code of the successfully verified task descriptor, the operator type is parsed, ensuring the security of the task descriptor of the current operator; if the operator type is matrix operation, use single-instruction multiple-data to call the co-core optimized assembly kernel to calculate the co-core computational subtasks. If the verification fails, write the error code of the failed verification into the error_code error field, and the main core will report the error.

[0035] In addition, as Figure 5 shown, based on the interrupt response mechanism, generate the computational execution result of the current operator on the main core according to the first execution result and the second execution result of the current operator, including: In response to the main core receiving an interrupt request, compare the interrupt number of the interrupt request with the interrupt number of the co - core on the main core; If the comparison is successful, identify the error field in the status register on the main core; If the identification is successful, generate an error prompt for the current operator on the main core according to the error code in the error field; If the identification fails, obtain the second execution result of the current operator in the double - buffer task queue on the main core, and move the second execution result of the current operator into the main - core thread; Generate the calculation execution result of the current operator on the main core according to the first execution result and the second execution result of the current operator.

[0036] In a specific implementation, before the result of the co - core thread is ready, that is, before the main core receives the interrupt request from the co - core, the main - core thread enters the sleep state to reduce the idle overhead of the main core; after the result is ready, the main - core thread will be woken up for result aggregation; the main - core thread checks whether all the computational subtasks of the current operator in the main - core thread and the co - core thread are completed; if completed, transfer the calculation execution result of the current operator to the next operator; if not completed, trigger the dynamic scheduler to re - allocate tasks.

[0037] Specifically, based on the interrupt response mechanism, the main core can timely obtain the status of the co - core, such as verification errors, calculation timeouts, etc. If the co - core generates an error interrupt, the main core can ensure the continuity of the model inference service by taking over the co - core to execute calculations and other operations.

[0038] Further, comparing the interrupt number of the interrupt request with the interrupt number of the co - core on the main core includes: If the interrupt number of the interrupt request is the same as the interrupt number of the co - core, determine that the comparison is successful on the main core; If the interrupt number of the interrupt request is different from the interrupt number of the co - core, determine that the comparison fails on the main core. In addition, identifying the error field in the status register on the main core includes: Judge whether the error field includes an error code on the main core; If the error field includes an error code, determine that the identification is successful on the main core; If the error field does not include an error code, determine that the identification fails on the main core.

[0039] In a specific implementation, if the error field does not include an error code, move the second execution result of the current operator into the main - core thread on the main core, release the buffer lock; integrate the first execution result and the second execution result of the current operator in the double - buffer task queue on the main core into the calculation execution result of the current operator.

[0040] In addition, performing a second calculation operation on the co-kernel for the current operator and outputting a second execution result of the current operator further includes: If the current co-kernel sub-task calculation is completed, obtain the number of currently completed co-kernel sub-tasks on the co-kernel and compare the number of currently completed co-kernel sub-tasks with the sub-task quantity threshold; If the number of currently completed co-kernel sub-tasks is less than the sub-task quantity threshold, perform a calculation operation on the next co-kernel sub-task on the co-kernel; Record the calculation duration of the next co-kernel sub-task on the co-kernel and compare the calculation duration of the next co-kernel sub-task with the sub-task calculation duration threshold; If the calculation duration of the next co-kernel sub-task is less than the sub-task calculation duration threshold, continue to perform a calculation operation on the next co-kernel sub-task on the co-kernel until the first case or the second case is satisfied, where the first case means that the next co-kernel sub-task calculation is completed and the calculation duration of the next co-kernel sub-task is less than or equal to the sub-task calculation duration threshold, and the second case means that the next co-kernel sub-task is not calculated completed and the calculation duration of the next co-kernel sub-task is greater than the sub-task calculation duration threshold; If the next co-kernel sub-task calculation is completed and the calculation duration of the next co-kernel sub-task is less than or equal to the sub-task calculation duration threshold, compare the number of currently completed co-kernel sub-tasks with the sub-task quantity threshold; If the number of currently completed co-kernel sub-tasks is equal to the sub-task quantity threshold, generate an interrupt request on the co-kernel and write the execution result of the currently completed co-kernel sub-task into the double-buffer task queue; If the next co-kernel sub-task is not calculated completed and the calculation duration of the next co-kernel sub-task is greater than the sub-task calculation duration threshold, generate an error code and an interrupt request on the co-kernel, write the error code into the error field, and send the interrupt request to the main core.

[0041] Further, generate an error prompt for the current operator on the main core according to the error code in the error field, including: If the error code is the first code, determine that the error type is a verification error on the main core and generate an error prompt for the verification error of the current operator; If the error code is the second code, determine that the error type is a calculation timeout on the main core and generate an error prompt for the calculation timeout of the current operator.

[0042] In the specific implementation, if the subtask number threshold is 5, and the subtask calculation time threshold is 100ms, the current calculation of the third co-core subtask is completed, and the calculation operation of the fourth co-core subtask continues. The current calculation time of the fourth co-core subtask is 60ms. If the fourth co-core subtask is completed at 900ms, the calculation operation of the fifth co-core subtask will continue. If the fourth co-core subtask is still not completed at 100ms, an error code of calculation timeout is generated, and an interrupt request is sent to the main core. By setting the subtask number threshold and the subtask calculation time threshold, the main core can obtain the calculation results of the co-core in a timely manner and handle the failure of the co-core in a timely manner.

[0043] In specific implementations, examples of model reasoning acceleration methods are as follows: 1. The user triggers an inference request: a query command is sent to the main core through the interface, such as "What is the method for creating a logical disk?" 2. Model loading and computational graph analysis: The main core loads the pre-trained AI model, inputs user questions and analyzes the computational graph to identify unloadable operators such as matrix multiplication and vector operations; 3. Task slicing and dynamic unloading: The main core thread pool executes computationally intensive operators, such as activation function operators and layer normalization operators; the co-core thread detects parallel computing requirements such as weight matrix multiplication, splits the matrix into dynamic slices with a slice size of 512, which can be adjusted in real time according to the co-core load, writes the task descriptor including data address and operator type into the shared memory lock-free queue, triggers the co-core interrupt with interrupt request number 32, and notifies the co-core to obtain the task; 4. Co-core calculation: After receiving the interrupt, the co-core enters the interrupt service process and reads the task descriptor from the shared memory; the co-core uses single instruction stream multiple data stream instructions such as the dot multiplication function to perform matrix multiplication and addition, and writes the result back to the output buffer in the double-buffered task queue module; 5. Interrupt merging strategy: After processing two slices, a low-priority interrupt is triggered to notify the main core that the result is ready; 6. Main core synchronization and result aggregation: interrupt processing, the main core responds to the interrupt, reads the sharding results from the shared memory, and verifies the cyclic redundancy check; data aggregation, integrates all sharding results, and calculates the final result; 7. Graph advancement and looping: Load-aware scheduling. If the current load rate of the co-core is 78%, the shard size of the next node is dynamically adjusted to 256. Double buffer switching. The main core fills the next batch of tasks into the standby buffer, and the co-core processes the tasks in the current buffer at the same time. 8. End and continue: The graph node where the current operator is located ends, and the calculation of the next graph node continues, and so on.

[0044] It should be understood that althoughFigure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 The steps in the flowchart of Figure 5 are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 At least some of the steps in Figure 5 may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the sub-steps or stages of other steps or other steps.

[0045] In one embodiment, as shown in Figure 6 , a model inference acceleration system is provided. The model inference acceleration system is within the baseboard management controller. The model inference acceleration system includes a main core, a co-core, and a shared memory. The main core includes a model inference process. The model inference process is used to load a preset model, parse the computational graph of the preset model, and initialize a thread pool. The co-core includes a lightweight real-time operating system. The lightweight real-time operating system is used to run a task processing engine and support priority preemption. The shared memory includes a task descriptor queue, a status register, and a double-buffered task queue. The task descriptor queue is used to store metadata corresponding to model inference requests. The status register is used to provide real-time feedback on the load rate, temperature, and error code of the co-core. The double-buffered task queue is used for alternating reading and writing between the main core and the co-core.

[0046] In a specific implementation, the metadata corresponding to the model inference request in this embodiment at least includes an operator type and a data address. Specifically, the shared memory module is used to overlap computation and communication, eliminate data replication overhead, and reduce communication latency.

[0047] Furthermore, the model inference process includes a thread pool and a dynamic scheduler. The thread pool includes at least a first thread, a second thread, a third thread, and a fourth thread. The first thread and the second thread are main core threads and are used to execute compute-intensive operators. The third thread and the fourth thread are co-core threads and are used for task offloading, co-core communication, and data synchronization between the main core and the co-core. A dynamic scheduler for adjusting the size and offloading strategy of computational subtasks in real time according to the load of co - cores.

[0048] In a specific implementation, computational - intensive operators include activation - function operators and layer - normalization operators.

[0049] Furthermore, the lightweight real - time operating system includes a task - processing engine and an interrupt controller. The task - processing engine is used to parse the task - descriptor task queue and perform matrix operations on the co - core subtasks in the double - buffered task queue. The interrupt controller is used to manage task - completion notifications and exception reporting.

[0050] In a specific implementation, matrix operations include matrix - by - matrix operations and matrix - by - vector operations, etc.

[0051] Furthermore, the double - buffered task queue includes an input - output buffer and a hardware semaphore. The input - output buffer is used to provide independent production and consumption areas, realizing the physical separation of computing and communication. The hardware semaphore is used to indicate the data - readiness status and realize buffer - role switching.

[0052] In a specific implementation, the input - output buffer is used for: ① data isolation, providing independent production and consumption areas, realizing the physical separation of computing and communication. A typical implementation is two buffers: the current buffer and the standby buffer; ② parallel - processing support, allowing the producer to fill one buffer while the consumer processes the other buffer, realizing true computing - communication overlap; ③ pipeline optimization, forming a continuous pipeline of production → consumption → production → consumption to hide data - transfer latency; ④ data - consistency guarantee, ensuring that each buffer is accessed by only one role (producer or consumer) at a specific moment to avoid read - write conflicts.

[0053] The hardware semaphore is used for: ① synchronization control to coordinate the access timing of producer (computing) and consumer (communication) threads / processes, ensuring data integrity and preventing simultaneous reading and writing of the same buffer; ② status notification, where the producer notifies the consumer "data is ready" through the semaphore, and the consumer notifies the producer "the buffer is free" through the semaphore; ③ atomic - operation guarantee, providing hardware - level atomic operations to avoid the overhead of software locks and ensure the atomicity of buffer - switching operations.

[0054] Therefore, the normal working process of the double-buffered task queue is as follows: The producer obtains an empty buffer through a semaphore; the producer fills the data into the input buffer; the producer notifies the consumer through the semaphore; the consumer obtains the full buffer through the semaphore; the consumer processes the data in the output buffer; the consumer notifies the producer through the semaphore. In addition, the semaphore timeout mechanism can detect deadlocks, and the buffer status flag can detect data corruption. For the specific limitations of the model inference acceleration system, reference can be made to the limitations on the model inference acceleration method in the foregoing text, which will not be elaborated here. Each module in the above model inference acceleration system can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the electronic device in hardware form or be independent of it, or can be stored in the memory in the electronic device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0055] In one embodiment, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: In response to receiving a model inference request, load a preset model on the main core according to the model inference request, and determine whether the current operator is an operator to be unloaded; If not, perform all calculation operations on the current operator on the main core according to the input tensor of the current operator, and output the calculation execution result of the current operator; If so, perform a first calculation operation on the current operator on the main core according to the input tensor of the current operator, and output the first execution result of the current operator. At the same time, perform a second calculation operation on the current operator on the co-core, and output the second execution result of the current operator, where all calculation operations at least include the first calculation operation and the second calculation operation; Based on the interrupt response mechanism, generate the calculation execution result of the current operator on the main core according to the first execution result and the second execution result of the current operator. When the program instructions are read and executed by one or more processors, the operations corresponding to the steps in the foregoing method embodiments can also be performed. Reference can be made to the description in the foregoing text, which will not be elaborated here. Refer to Figure 7 , which exemplarily shows the architecture of the electronic device. Specifically, it may include a processor 710, a video display adapter 711, a disk drive 712, an input / output interface 713, a network interface 714, and a memory 720. The above processor 710, video display adapter 711, disk drive 712, input / output interface 713, network interface 714, and the memory 720 can be communicatively connected through a communication bus 730.

[0056] Among them, the processor 710 can be implemented in the form of a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in this application.

[0057] The memory 720 can be implemented in the form of a read-only memory (ROM), a random access memory (RAM), a static storage device, a dynamic storage device, etc. The memory 720 can store an operating system 721 for controlling the operation of the electronic device 700, and a basic input / output system (BIOS) 722 for controlling the low-level operations of the electronic device 700. In addition, a web browser 723, a data storage management 724, an icon font processing system 725, etc. can also be stored. The above-mentioned icon font processing system 725 can be the application program that specifically implements the operations of the foregoing steps in the embodiments of this application. In short, when implementing the technical solutions provided in this application through software or firmware, the relevant program codes are stored in the memory 720 and called and executed by the processor 710.

[0058] The input / output interface 713 is used to connect to the input / output module to implement information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0059] The network interface 714 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module can implement communication in a wired manner (such as USB, network cable, etc.) or in a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).

[0060] The bus 730 includes a path for transmitting information between various components of the device (such as the processor 710, the video display adapter 711, the disk drive 712, the input / output interface 713, the network interface 714, and the memory 720).

[0061] In addition, the electronic device 700 can also obtain information on specific redemption conditions from a virtual resource object redemption condition information database (not shown in the figure) for conditional judgment, etc.

[0062] It should be noted that although the above electronic device 700 only shows a processor 710, a video display adapter 711, a disk drive 712, an input / output interface 713, a network interface 714, a memory 720, a bus 730, etc., in the specific implementation process, the electronic device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the present application, and does not necessarily include all the components shown in the figure.

[0063] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing an electronic device (which can be a personal computer, a cloud server, or a network device, etc.) to execute the methods of various embodiments or some parts of the embodiments of the present application.

[0064] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: In response to receiving a model inference request, load a preset model on the main core according to the model inference request, and determine whether the current operator is an operator to be unloaded; If not, perform all calculation operations on the current operator on the main core according to the input tensor of the current operator, and output the calculation execution result of the current operator; If so, perform a first calculation operation on the current operator on the main core according to the input tensor of the current operator, and output the first execution result of the current operator. At the same time, perform a second calculation operation on the current operator on the co-core, and output the second execution result of the current operator, where all calculation operations at least include the first calculation operation and the second calculation operation; Based on the interrupt response mechanism, according to the first execution result and the second execution result of the current operator, generate the calculation execution result of the current operator on the main core. Those of ordinary skill in the art can understand that all or part of the processes in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0065] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0066] The above embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application.

[0067] In one embodiment, a computer program product is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: In response to receiving a model inference request, load a preset model on the main core according to the model inference request, and determine whether the current operator is an operator to be unloaded; If not, perform all calculation operations on the current operator on the main core according to the input tensor of the current operator, and output the calculation execution result of the current operator; If so, according to the input tensor of the current operator, perform a first calculation operation on the current operator on the main core, output the first execution result of the current operator, and at the same time perform a second calculation operation on the current operator on the co-core, output the second execution result of the current operator, where all calculation operations at least include the first calculation operation and the second calculation operation; Based on the interrupt response mechanism, generate the calculation execution result of the current operator on the main core according to the first execution result and the second execution result of the current operator. In one embodiment, a computer program product is provided, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: In response to receiving a model inference request, load a preset model on the main core according to the model inference request, and determine whether the current operator is an operator to be unloaded; If not, according to the input tensor of the current operator, perform all calculation operations on the current operator on the main core, and output the calculation execution result of the current operator; If so, according to the input tensor of the current operator, perform a first calculation operation on the current operator on the main core, output the first execution result of the current operator, and at the same time perform a second calculation operation on the current operator on the co-core, output the second execution result of the current operator, where all calculation operations at least include the first calculation operation and the second calculation operation; Based on the interrupt response mechanism, generate the calculation execution result of the current operator on the main core according to the first execution result and the second execution result of the current operator. Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer program product. When the computer program is executed, it can include the processes of the above method embodiments.

[0068] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0069] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it cannot be understood as a limitation to the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application.

Claims

1. A model inference acceleration method, which is applied to a baseboard management controller. The baseboard management controller includes a main core and a co-core, and is characterized in that, The method includes: In response to receiving a model inference request, loading a preset model on the main core according to the model inference request, and determining whether the current operator is an operator to be unloaded; If not, performing all calculation operations on the current operator on the main core according to the input tensor of the current operator, and outputting the calculation execution result of the current operator; If so, performing a first calculation operation on the current operator on the main core according to the input tensor of the current operator, and outputting a first execution result of the current operator, and at the same time performing a second calculation operation on the current operator on the co-core, and outputting a second execution result of the current operator, where the all calculation operations at least include the first calculation operation and the second calculation operation; Based on an interrupt response mechanism, generating the calculation execution result of the current operator on the main core according to the first execution result and the second execution result of the current operator.

2. The method according to claim 1, wherein Before determining whether the current operator is an operator to be unloaded, the method further includes: Identifying a calculation graph of the preset model on the main core to obtain a plurality of unloadable operators; If the current unloadable operator is a memory-intensive operator, determining that the current unloadable operator is not an operator to be unloaded on the main core; If the current unloadable operator is a compute-intensive operator, determining that the current unloadable operator is the operator to be unloaded on the main core.

3. The method according to claim 1, characterized in that The performing all calculation operations on the current operator on the main core according to the input tensor of the current operator, and outputting the calculation execution result of the current operator includes: Generating a calculation task of the current operator on the main core according to the current operator and the input tensor of the current operator, and writing the calculation task of the current operator into the main core thread of the thread pool; Processing the calculation task of the current operator in the main core thread on the main core, outputting the calculation execution result of the current operator, and writing the calculation execution result of the current operator into a double-buffer task queue in shared memory.

4. The method according to claim 3, wherein Before performing a first calculation operation on the current operator on the main core according to the input tensor of the current operator, and outputting a first execution result of the current operator, the method further includes: Cutting the calculation task of the current operator into corresponding multiple calculation subtasks on the main core, and writing the multiple calculation subtasks of the current operator into the double-buffer task queue, where the calculation amounts between the multiple calculation subtasks are the same; Based on a task allocation mechanism, dividing the multiple calculation subtasks of the current operator into a main core calculation subtask and a co-core calculation subtask of the current operator on the main core; If the current calculation subtask of the current operator is a main core calculation subtask, writing the current calculation subtask of the current operator into the main core thread on the main core; If the current calculation subtask of the current operator is a co-core calculation subtask, writing the current calculation subtask of the current operator into the co-core thread of the thread pool on the main core; Generate a task descriptor of the current operator on the main core according to the input tensor of the current operator, and write the task descriptor of the current operator into the task descriptor queue in the shared memory.

5. The method according to claim 4, characterized in that, Performing a first calculation operation on the current operator on the main core and outputting a first execution result of the current operator includes: Performing a first calculation operation on the main core calculation subtask of the current operator in the main core thread, outputting a first execution result of the current operator, and writing the first execution result of the current operator into the double-buffer task queue; Performing a second calculation operation on the current operator on the co-processor core and outputting a second execution result of the current operator includes: Obtain the task descriptor of the current operator in the task descriptor queue on the co-processor core; According to the task descriptor of the current operator, perform a second calculation operation on the co-processor core calculation subtask of the current operator in the co-processor core thread, output a second execution result of the current operator, and write the second execution result of the current operator into the double-buffer task queue.

6. The method according to claim 4, wherein Based on the task assignment mechanism, dividing multiple calculation subtasks of the current operator into co-processor core calculation subtasks of the current operator on the main core includes: Obtain the load rate of the co-processor core thread from the status register in the shared memory on the main core, and obtain the co-processor core calculation ratio of the current operator according to the load rate of the co-processor core thread; According to the calculation task of the current operator, obtain the total calculation amount of the current operator on the main core, and obtain the co-processor core calculation amount of the current operator according to the total calculation amount and co-processor core calculation ratio of the current operator; According to multiple calculation subtasks of the current operator, obtain the average calculation amount of the current operator on the main core, and obtain the number of co-processor core subtasks of the current operator by rounding down according to the co-processor core calculation amount and average calculation amount of the current operator; According to the number of co-processor core subtasks of the current operator, divide the corresponding number of calculation subtasks of the current operator into co-processor core calculation subtasks of the current operator on the main core.

7. The method according to claim 6, characterized in that Based on the task assignment mechanism, dividing multiple calculation subtasks of the current operator into main core calculation subtasks of the current operator on the main core includes: Obtain the total number of subtasks of the current operator on the main core according to multiple calculation subtasks of the current operator; Obtain the number of main core subtasks of the current operator on the main core according to the total number of subtasks and the number of co-processor core subtasks of the current operator; According to the number of main core subtasks of the current operator, divide the corresponding number of calculation subtasks of the current operator into main core calculation subtasks of the current operator on the main core.

8. The method according to claim 5, wherein Generating a task descriptor of the current operator on the main core according to the input tensor of the current operator includes: Obtain the input tensor address, output tensor address and operator type of the current operator on the main core, and generate a cyclic check code of the current operator according to the input tensor address, output tensor address and operator type of the current operator; Generate a task descriptor for the current operator on the main core according to the input tensor address, output tensor address, operator type, and cyclic redundancy check code of the current operator.

9. The method according to claim 8, characterized in that, The obtaining of the task descriptor of the current operator in the task descriptor queue on the co-core includes: Verify the task descriptor of the current operator on the co-core according to the cyclic redundancy check code of the current operator; If the verification is successful, obtain the task descriptor of the current operator in the task descriptor queue on the co-core; If the verification fails, generate an error code and an interrupt request on the co-core; Write the error code into the error field in the status register on the co-core and send the interrupt request to the main core.

10. The method according to claim 9, wherein The generating of the calculation execution result of the current operator on the main core based on the interrupt response mechanism according to the first execution result and the second execution result of the current operator includes: In response to the main core receiving the interrupt request, compare the interrupt number of the interrupt request with the interrupt number of the co-core on the main core; If the comparison is successful, identify the error field in the status register on the main core; If the identification is successful, generate an error prompt for the current operator on the main core according to the error code in the error field; If the identification fails, obtain the second execution result of the current operator in the double-buffer task queue on the main core and move the second execution result of the current operator into the main core thread; Generate the calculation execution result of the current operator on the main core according to the first execution result and the second execution result of the current operator.

11. The method according to claim 10, wherein The comparison of the interrupt number of the interrupt request with the interrupt number of the co-core on the main core includes: If the interrupt number of the interrupt request is the same as the interrupt number of the co-core, determine that the comparison is successful on the main core; If the interrupt number of the interrupt request is different from the interrupt number of the co-core, determine that the comparison fails on the main core.

12. The method according to claim 10, wherein The identification of the error field in the status register on the main core includes: Judge on the main core whether the error field includes the error code; If the error field includes the error code, determine that the identification is successful on the main core; If the error field does not include the error code, determine that the identification fails on the main core.

13. An electronic device, characterized in that, Includes: A memory for storing a computer program; A processor for implementing the steps of the model inference acceleration method according to any one of claims 1 to 12 when executing the computer program.

14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the model inference acceleration method according to any one of claims 1 to 12 when executed by a processor.

15. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the model inference acceleration method according to any one of claims 1 to 12 when executed by a processor.

Citation Information

Patent Citations

  • Operator fusion processing method and device, equipment and storage medium

    CN115203126A

  • Operator processing method, electronic equipment and storage medium

    CN115600664A

  • Operator unloading method and system of vectorization execution engine based on DPU heterogeneous architecture

    CN119781850A

  • Agent method for operating system high-frequency polling, electronic equipment and storage medium

    CN119883579A

  • Operator processing method and apparatus, and chip, computing device and storage medium

    WO2024131170A1

Cited By

  • Method for reasoning optimization of pre-training model and electronic equipment

    CN121210158A

  • Inference optimization method of pre-trained model and electronic device

    CN121210158B

  • Artificial intelligence chip

    CN121683907A