Model reasoning acceleration method and system, computer equipment and readable storage medium
By introducing virtual accelerator cards in the edge/end side inference scenario, using the CPU's memory and SIMD instruction set technology, the problem of low resource utilization caused by relying on physical accelerator cards in the existing technology is solved, and the model inference acceleration is achieved without accelerator cards is improved, and the resource utilization and overall computing efficiency of the CPU are improved.
Patent Information
- Application Number
- CN202510466856.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
In the edge/end side inference scenario, the inference calculation of neural network models is large, and traditional CPUs are limited in efficiency when dealing with large-scale parallel computing tasks. The existing technology relies on physical acceleration cards (xPUs) for model inference, and the resource utilization rate of acceleration cards is low.
By introducing a virtual acceleration card, using the CPU's memory and SIMD instruction set technology, the model file is loaded into the CPU memory space without the physical acceleration card, and the target operator matching all operators is determined from the user-state preset virtual operator library to achieve model inference acceleration.
It improves the resource utilization rate of CPU, avoids the inference acceleration in the event of limited hardware resource allocation, meets diversified business needs, and reduces the dependence on expensive and dedicated hardware.
Smart Images

Figure CN119990337A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a model reasoning acceleration method, system, computer device and readable storage medium. Background Art
[0002] In edge / end inference scenarios, the amount of inference calculations of neural network models is large, and the performance requirements of hardware devices are high. As a general-purpose processor, the traditional central processing unit (CPU) has high flexibility but limited efficiency when processing large-scale parallel computing tasks. In order to accelerate specific tasks, dedicated acceleration chips such as graphics processing units (GPUs), tensor processing units (TPUs), field-programmable gate arrays (FPGAs), etc. (collectively referred to as xPUs) have emerged. They have optimized architectures for computationally intensive tasks such as neural networks, which can greatly improve computing speed.
[0003] In the related art, the field of model reasoning acceleration directly relies on the physical accelerator card (xPU) to perform neural network reasoning tasks. The model is directly loaded into the memory space of the accelerator card, and the inference calculation is performed using the operator library built into the accelerator card. However, the model reasoning acceleration in the related art relies on the accelerator card, resulting in low resource utilization. Therefore, a model reasoning acceleration method that can improve resource utilization is needed. Summary of the invention
[0004] Based on this, it is necessary to provide a model reasoning acceleration method, system, computer device, computer-readable storage medium and computer program product that can improve resource utilization in response to the above technical problems.
[0005] In a first aspect, the present application provides a model reasoning acceleration method, comprising: Obtain a model file to be processed of the model to be processed; Parsing the model file to be processed to obtain model information of the model to be processed; If there is no physical acceleration card, the model file to be processed is loaded into the memory space of the CPU, and a target operator matching all operators in the model information is determined from a virtual operator library preset in the user state; Each of the operators is mapped to a corresponding target operator, and the reasoning of the model to be processed is completed based on the target operator.
[0006] In one embodiment, the method further comprises: Obtain the memory capacity required for executing reasoning of the model file to be processed and the current memory pool capacity; If the current memory pool capacity is smaller than the memory capacity, a dynamic expansion mechanism is triggered to determine the memory capacity in the current memory pool to be allocated to the model file to be processed; The step of parsing the model file to be processed to obtain model information of the model to be processed includes: Converting the model file to be processed into a model intermediate representation, and determining a computational graph of the model file to be processed, that is, a computational graph; According to the calculation graph, all operators of the model to be processed are obtained.
[0007] In one embodiment, mapping each of the operators to a corresponding target operator and completing the inference of the model to be processed based on the target operator includes: Mapping each of the operators to a corresponding target operator, and calling the target operator to perform calculations according to the dependency relationship between the operators in the calculation graph to obtain calculation results; Processing is performed according to the calculation results to obtain inference results, thereby completing the inference of the model to be processed.
[0008] In one embodiment, the method further comprises: If there is a physical acceleration card and the physical acceleration card supports all operators in the model information, then the model file to be processed is loaded into the video memory space of the physical acceleration card; Call the user-state application layer interface, route all the operators to the inference operator library of the physical acceleration card through the dynamic routing engine of the unified operator library, map each of the operators to the corresponding target operator in the inference operator library, and complete the inference of the model to be processed based on the inference operator library.
[0009] In one embodiment, the method further comprises: If there is a physical acceleration card and the physical acceleration card supports some of the operators, the model file to be processed is loaded into the video memory space of the physical acceleration card; Calling the user-state application layer interface, routing some of the operators to the inference operator library of the physical acceleration card through the dynamic routing engine of the unified operator library, and mapping each of the operators in some of the operators to the corresponding target operators in the inference operator library; Load the remaining operators into the memory space of the CPU, determine a virtual operator library matching all the remaining operators from the virtual accelerator card operator library preset in the user state, and map the remaining operators to the virtual operator library; Based on the inference operator library and the virtual operator library, the inference of the model to be processed is completed.
[0010] In one embodiment, the accelerating the inference of the model to be processed based on the inference operator library and the virtual operator library includes: Mapping the remaining operators to the virtual operator library, and calling corresponding operators in the virtual operator library to perform calculations according to the dependency relationships between the remaining operators to obtain a first calculation result; Sending the first calculation result to the physical acceleration card through a preset shared memory channel; Mapping some of the operators to the inference operator library, and calling corresponding operators in the inference operator library to perform calculations according to the dependency relationships between some of the operators to obtain a second calculation result; The first calculation result and the second calculation result are integrated through the physical acceleration card to complete the reasoning of the model to be processed.
[0011] In one embodiment, the method further comprises: When it is detected that the physical acceleration card is abnormal or the target operator is missing, the operator supported by the physical acceleration card or the operator corresponding to the target operator is loaded into the memory space of the CPU, and the step of determining the target operator matching all the operators in the model information from the virtual operator library preset in the user state is executed.
[0012] In a second aspect, the present application also provides a model reasoning acceleration system, the system comprising a physical hardware layer, a virtual acceleration card and an application layer, the physical hardware layer comprising a CPU, wherein: The application layer is used to initiate an inference request and determine a to-be-processed model file of the to-be-processed model carried by the inference request; The virtual acceleration card is used to obtain the model file to be processed; parse the model file to be processed to obtain model information of the model to be processed; if there is no physical acceleration card, load the model file into the memory space of the CPU, and determine the target operator that matches all operators in the model information from the virtual operator library preset in the user state; map each of the operators to the corresponding target operator, and complete the reasoning of the model to be processed based on the target operator.
[0013] In a third aspect, the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.
[0014] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of any one of the above methods when executed by a processor.
[0015] In a fifth aspect, the present application also provides a computer program product, including a computer program, which implements the steps of any of the above methods when executed by a processor.
[0016] The above-mentioned model reasoning acceleration method, system, computer device, computer-readable storage medium and computer program product, when reasoning on the model, parses the model file to be processed of the model to be processed to obtain all the operators of the model, and in the absence of a physical acceleration card, loads the model file into the memory space of the CPU, and determines the target operator that matches all the operators from the virtual operator library preset in the user state; maps each operator to its corresponding target operator, and accelerates the reasoning of the model to be processed based on the target operator, avoiding the current situation where the hardware resource configuration is limited and the reasoning acceleration cannot be achieved without an acceleration card, fully taps the potential of the CPU, and improves the resource utilization of the CPU. In addition, it meets diverse business needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 A schematic diagram of the architecture of a model reasoning acceleration system in one embodiment; Figure 2 A schematic diagram of a flow chart of a model reasoning acceleration method in one embodiment; Figure 3 A schematic diagram of a flow chart of a model reasoning acceleration method in another embodiment; Figure 4 It is a timing diagram of a model reasoning acceleration method based on a virtual accelerator card and an xPU in one embodiment; Figure 5 It is an application flow chart of a model reasoning acceleration method in one embodiment; Figure 6 A structural block diagram of a model reasoning acceleration system in one embodiment; Figure 7 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0020] In edge / end-side inference scenarios, the amount of inference calculations of neural network models is large, and the performance requirements of hardware devices are high. In order to accelerate specific tasks, dedicated acceleration chips such as GPU, TPU, FPGA, etc. (collectively referred to as xPU) came into being. However, xPU has high cost and adaptability issues. The xPU programming models and interfaces of different manufacturers are different, which makes development difficult. At the same time, in some scenarios, the hardware resource configuration is limited, and it may not be possible to equip a full range of acceleration cards, or the acceleration card functions are partially missing, making it difficult to meet diverse business needs.
[0021] Common solutions in the current model reasoning acceleration field usually directly rely on physical accelerator cards (xPUs) for neural network reasoning tasks. In the model loading stage, the model is directly loaded into the memory space of the accelerator card, and the inference calculation is performed using the operator library built into the accelerator card. When encountering operators or tasks that are not supported by the accelerator card, or when there is no physical accelerator card, a complex adaptation layer is often required to try to convert the operator or abandon the task branch, resulting in inefficient development and prone to errors. Moreover, the management of system resources is relatively extensive, and the potential for the CPU and accelerator card to work together is not fully explored. Especially in resource-constrained scenarios, resources cannot be flexibly allocated, and tasks are prone to jamming or even failure.
[0022] In view of the situation that model reasoning acceleration cannot be achieved due to the absence of a physical acceleration card, a model reasoning acceleration method is proposed. By introducing a virtual acceleration card, a virtual operator library is preset in the user state of the virtual acceleration card. When there is no physical acceleration card, the model file to be processed is loaded into the memory space of the CPU, and the target operator matching all operators is determined from the virtual operator library preset in the user state to achieve model reasoning acceleration.
[0023] The model reasoning acceleration method provided in the embodiment of the present application can be applied to Figure 1 In the system shown, the system is deployed on a terminal, and the terminal may be an edge device. Figure 1The system architecture shown is the physical hardware layer, kernel state, user state, and application layer from bottom to top. The virtual accelerator card consists of two parts: kernel state and user state. The physical hardware layer is the underlying resource managed by the virtual accelerator card, which actually performs computing tasks and may include the CPU, or the CPU and xPU, where the CPU may include at least SIMD, vector registers, etc., and the xPU may include at least GPU, TPU, FPGA, etc. It should be noted that the virtual accelerator card does not include underlying module support, and the physical hardware layer is the underlying resource managed by the virtual accelerator card system. As a software-defined middle layer, the virtual accelerator card abstracts and schedules the physical hardware (CPU, xPU) through kernel state and user state modules, but the virtual accelerator card itself does not contain physical hardware components. The relationship between the physical hardware layer and the virtual accelerator card can be compared to that between the operating system and computer hardware - the operating system manages the hardware, but the hardware itself exists independently.
[0024] The virtual accelerator card consists of two parts: user state and kernel state. The user state includes the application interface layer and the computing function layer. The application interface layer (vCardMg) serves as the interaction entrance between the user program and the virtual accelerator card system, and provides a full life cycle management interface for the virtual accelerator card. The application interface layer includes multiple sub-modules, including virtual card instance management and resource status monitoring.
[0025] Virtual card instance management is used to call the instance creation interface to trigger kernel memory pool allocation and generate a virtual card object containing a unique ID. The memory pool allocation adopts the "pre-allocation + dynamic expansion" strategy. The "pre-allocation + dynamic expansion" strategy can be understood as the system initially allocating basic memory when it is initialized. When insufficient initial memory is detected, on-demand expansion is triggered. The on-demand here can be based on a pre-set ratio or a determined expansion ratio based on the size of the model file to be processed.
[0026] Virtual card instance management can also be used to destroy instances, that is, release the associated memory pool and operator resources, and notify the kernel to recycle resources. For example, in the case of a model inference acceleration, idle memory pools and operator resources will be released, or when the inference ends, the associated memory pool and operator resources will be released. In addition, virtual card instance management can also be used for real-time status query, that is, to obtain indicators such as memory usage of each virtual card through the virtual file system, and to implement an event subscription mechanism: support asynchronous monitoring of abnormal events of acceleration cards (such as xPU overheating and insufficient memory).
[0027] The computing function layer includes a unified operator library (vOplib) and an image processing library (vImage). The unified operator library is implemented based on the SIMD instruction set and provides the operator library required for reasoning. It can also be understood as providing a unified operator interface across hardware to achieve dynamic routing and optimized execution of computing tasks. The core submodules include a virtual operator library and an operator dynamic routing engine. The virtual operator library mainly implements some commonly used operators, including the implementation of commonly used neural network operators, such as convolution and pooling. Dynamic routing engine: By maintaining the hardware support matrix to determine the execution path of each operator, the SIMD (Single Instruction Multiple Data) optimization kernel automatically selects the optimal implementation for different CPU instruction sets.
[0028] The image processing library (vImage) provides hardware-accelerated image preprocessing based on SIMD, which is compatible with the OpenCV interface specification. The core submodules of the image processing library (vImage) include image preprocessing and SIMD optimization algorithms. Image preprocessing: supports conventional image preprocessing operations such as scaling and normalization. SIMD optimization algorithm: provides hardware-accelerated image processing optimization algorithms based on SIMD, which are compatible with OpenCV.
[0029] The kernel state includes the core driver layer (vCard_Kernel) and the resource management layer, where the core driver layer is used to manage physical hardware resources and provide a securely isolated underlying operation interface. Memory pool management includes dynamic memory allocation through "pre-allocation + dynamic expansion", while establishing a secure isolation mechanism to create an independent memory namespace for each virtual card instance to prevent cross-boundary access. It should be noted that one virtual accelerator card corresponds to one virtual card instance, or multiple virtual card instances. This example uses one virtual accelerator card corresponding to one virtual card instance as an example. Providing a securely isolated underlying operation interface includes defining a hardware abstraction interface (HAL), defining a standardized xPU operation set, and shielding differences in drivers from different manufacturers.
[0030] The resource management layer includes a cross-device memory mapping module and a hardware monitor. The cross-device memory mapping module is used to create an independent memory mapping table for each virtual card instance; it receives memory allocation instructions from vCardMg (user mode) and is coordinated and executed by vCard_Kernel. Specifically, it can be manifested as: unified address space management: mapping CPU memory and xPU video memory to the same virtual address space to achieve zero-copy data transmission, and memory sharing protocol: defining cross-device memory access rules.
[0031] The hardware monitor is used to detect the use of hardware resources in real time and provide real-time data support for the dynamic routing (vOplib) and memory allocation (vCard_Kernel) of the virtual accelerator card. The specific performance includes real-time hardware status collection: reading xPU computing power utilization, temperature, power consumption and other indicators through PMU (Performance Monitoring Unit).
[0032] In the above system, by introducing a virtual accelerator card and combining it with the CPU's memory, vector registers and SIMD instruction set technology, the differences between different xPUs can be shielded, and a unified resource management and operator calling interface can be provided. Regardless of whether there is an accelerator card at the bottom layer or whether the accelerator card function is complete, the continuity of the reasoning task can be guaranteed. On the other hand, the CPU potential is deeply tapped. Through the optimized memory management module and rich operator library, the CPU is used for efficient image processing and neural network reasoning when the accelerator card resources are insufficient, thereby realizing dynamic coordination between the CPU and xPU and greatly improving the overall computing performance and resource utilization.
[0033] In one embodiment, Figure 2 As shown, a model reasoning acceleration method is provided. This embodiment uses the method applied to a terminal as an example. The terminal may be an edge device, and a Figure 1 The model reasoning acceleration system shown, it can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps: Step 202, obtaining a model file to be processed of the model to be processed.
[0034] Among them, the model to be processed can be but not limited to the image model to be processed, and the image model can be used for image classification and target object recognition. For example, automatically identify objects in photos (such as cats, dogs, cars, etc.); for example, in medical image analysis, determine whether there is abnormal tissue in the image, or identify specific people or behaviors in security monitoring scenarios. The model file to be processed can be pre-trained by the user and stored in the local disk or cloud storage, depending on the application deployment method. The model file can contain the neural network model structure and parameters, etc. The format of the model file can be but not limited to ONNX (OpenNeuralNetworkExchange), TensorFlowSavedModel / PB format and PyTorch.pth / .pt format. ONNX (OpenNeuralNetworkExchange): an open deep learning model exchange format that supports model conversion between multiple frameworks. TensorFlowSavedModel / PB format: a file format for saving TensorFlow models. PyTorch.pth / .pt format: a file format for saving PyTorch models.
[0035] It should be noted that before model inference acceleration, the system deployed on the terminal will be initialized first. System initialization includes memory subsystem initialization, operator loading and optimization, and virtual accelerator card instantiation. Memory subsystem initialization includes: working with kernel-mode driver (vCard_Kernel) through user-mode management interface (vCardMg), starting kernel-mode memory management module, i.e. memory pool management, and reserving isolated memory space for operator library, model parameters, and virtual accelerator card instance. And / or, establishing a shared memory channel between CPU and xPU through cross-device memory mapping module.
[0036] Operator loading and optimization include loading the user-mode unified operator library (vOplib), parsing the hardware support matrix, dynamically selecting the optimal operator implementation based on the current CPU instruction set, and optimizing operator parameters based on the CPU cache characteristics. Virtual accelerator card instantiation includes calling vCardMg to create a virtual card instance and establishing a communication link between the user mode and the kernel mode, where the communication link between the user mode and the kernel mode can be established through ioctl commands and shared memory.
[0037] Exemplarily, the model reasoning acceleration system deployed on the terminal is initialized, and in response to the model reasoning acceleration instruction, the model file of the model to be processed is obtained from the local disk, remote server or cloud storage. Further, after obtaining the model file, in order to ensure the reliability of the model reasoning acceleration, the accuracy of the model file can be verified, for example, the file format of the model file and the integrity of the file are verified.
[0038] Step 204: parse the model file to be processed to obtain model information of the model to be processed.
[0039] The model information includes the structural information and parameters of the model, and the structural information includes the operators of the model to be processed. Operators can be understood as basic computing units in neural networks, such as convolution, pooling, activation functions, etc. The parsing method of the model file to be processed can be parsed using existing methods, which will not be described here.
[0040] For example, the specific operation of file parsing can include identifying the format of the model file and extracting the structural information and parameters of the model. On this basis, through operator extraction, all operators in the model and their attributes are identified and extracted. The attributes may include input and output dimensions, weights, etc.
[0041] Step 206: If there is no physical acceleration card, the model file to be processed is loaded into the memory space of the CPU, and a target operator matching all operators in the model information is determined from a virtual operator library preset in the user state.
[0042] Among them, physical acceleration cards (xPU): such as GPU, TPU, FPGA and other dedicated hardware accelerators, are used to accelerate neural network reasoning tasks.
[0043] It is understandable that model reasoning acceleration in related technologies is highly dependent on a specific xPU. When the accelerator card is missing or some functions are not sound, the entire reasoning task cannot be carried out smoothly or the performance is seriously degraded. Therefore, a virtual accelerator card is introduced.
[0044] Step 208, mapping each operator to its corresponding target operator, and completing the inference of the model to be processed based on the target operator.
[0045] Exemplarily, when the model file of the model to be processed is obtained, the model file is loaded and mapped through the kernel-state memory management module and the user-state vOplib in collaboration, and the model file to be processed is parsed to obtain all operators of the model to be processed. The operator library has intelligent self-adaptation capabilities, and the runtime system will detect the hardware environment in real time. When it is detected that there is no physical acceleration card, the model file is loaded into the memory space of the CPU, and the user-state dynamic routing vOplib will quickly switch to the operator library based on the SIMD instruction set, that is, the target operator matching all operators is determined from the virtual operator library preset in the user state, and the CPU takes over the corresponding operator calculation task, and uses SIMD parallel computing optimization and other technologies based on the target operator to accelerate the calculation process. It can be understood that when model reasoning is performed based on the operator, model input data related to the model reasoning task will be input, and the model input data can be image data.
[0046] Furthermore, vector registers and SIMD instruction sets are pre-integrated in the CPU. When the model file is loaded into the CPU's memory space, the model and data will be optimized for parallel computing. The parallel computing optimization methods include but are not limited to: data layout optimization (such as allocating high-frequency access data to the CPU's L3 cache proximal memory, etc.), adaptive data blocking (i.e. dividing data blocks according to the number of SIMD channels), and instruction-level optimization (such as vectorization, instruction pipelining, etc.).
[0047] In the above embodiment, when there is a lack of dedicated hardware accelerators, the reasoning task can be completed on a general-purpose CPU by using a virtual operator library, avoiding the situation where reasoning cannot be performed due to lack of hardware, thereby improving the flexibility and compatibility of the system. In other words, regardless of whether there is a physical acceleration card, the model reasoning task can be successfully completed, which enhances the adaptability and flexibility of the system. In the absence of dedicated hardware, the existing CPU resources are fully utilized to avoid resource waste and improve overall resource utilization. In addition, the reliance on expensive dedicated hardware is reduced, the hardware procurement and maintenance costs are reduced, and more application scenarios can achieve efficient model reasoning capabilities and meet diverse business needs.
[0048] It is understandable that before the model accelerates reasoning, it is necessary to determine the capacity of the current memory pool and the memory capacity required for the model file to be processed to perform reasoning, so as to ensure that there is enough memory to complete the reasoning task. In an exemplary embodiment, the model reasoning acceleration method also includes: Obtain the memory capacity required for executing reasoning for the model file to be processed and the current memory pool capacity; if the current memory pool capacity is less than the memory capacity, trigger the dynamic expansion mechanism to determine the memory capacity in the current memory pool to be allocated to the model file to be processed.
[0049] Among them, the dynamic expansion mechanism is that the virtual card instance management calls the instance creation interface to trigger the kernel memory pool allocation. When the initial memory is insufficient, it triggers the on-demand expansion. The specific expansion method can be achieved by using the "pre-allocation + dynamic expansion" strategy for memory pool allocation, which will not be described in detail here.
[0050] Accordingly, the model file to be processed is parsed to obtain all operators of the model to be processed, including: converting the model file to be processed into an intermediate representation of the model, extracting a calculation graph according to the intermediate representation of the model; and obtaining all operators of the model to be processed according to the calculation graph.
[0051] Exemplarily, the ONNX / TensorFlow model is converted into a model intermediate representation (Unified IR), the computational graph is extracted, and all operators of the model to be processed are obtained.
[0052] In an exemplary embodiment, each operator is mapped to a corresponding target operator, and the reasoning of the model to be processed is accelerated based on the target operator, including: Map each operator to its corresponding target operator, call the target operator to perform calculations according to the dependency relationship between operators in the calculation graph, and obtain the calculation results; process the calculation results to obtain the inference results, and complete the inference of the model to be processed.
[0053] Among them, the confirmation of the target operator can be determined in the virtual operator library according to the operator type and its attributes obtained by analysis. Post-processing is performed according to the calculation results to obtain the inference results. The calculation results can be post-processed to obtain the inference results. Post-processing can be understood as a series of operations on the original output of the model to obtain results that better meet the actual application requirements. Post-processing includes result analysis and result visualization. For example, post-processing is to convert the probability value output by the model into a specific category label. For example, in the target detection task, the output result is to remove redundant bounding boxes to retain the most accurate detection results.
[0054] The inference task refers to the process of using a trained model to predict or classify new data. The inference task is based on the calculation results that have been mapped and calculated. For example, for an image classification model, the inference task is to use the model to give the category label of the image based on the input image data. The inference result refers to the final output obtained after the inference task is completed, such as the classification label, prediction value, etc.
[0055] Furthermore, during the inference calculation process, a feedback control thread can be started to monitor the memory access mode of each layer in real time. During the mapping process, a dynamic memory allocation strategy can be adopted. Specifically, by monitoring the changes in computing requirements of each layer of the model during the inference process and the fluctuations in data access frequency in real time, the feedback control mechanism is used to dynamically adjust the CPU's memory allocation ratio.
[0056] For model reasoning acceleration, there may also be physical acceleration cards, but the physical acceleration cards do not support operators or only support some operators. In addition to the above reasoning acceleration, there are also the following two situations: Case 1: If there is a physical acceleration card and the physical acceleration card supports all operators in the model information, the model file to be processed is loaded into the video memory space of the physical acceleration card; the user-state application layer interface is called, and all operators are routed to the inference operator library of the physical acceleration card through the dynamic routing engine of the unified operator library, and each operator is mapped to the corresponding target operator in the inference operator library, and the inference of the model to be processed is completed based on the inference operator library.
[0057] Exemplarily, if there is a physical acceleration card and the physical acceleration card supports all operators, the model file to be processed is loaded into the video memory space of the physical acceleration card; the application layer interface of the user state is called, and all operators are routed to the inference operator library of the physical acceleration card through the dynamic routing engine, and each operator is mapped to the corresponding target operator in the inference operator library, and the parallel computing advantage of the acceleration card is used for fast processing, and the inference of the model to be processed is accelerated based on the inference operator library. It can be understood that when executing an inference task, the execution path will be selected according to the hardware support matrix and the real-time load status. For example, the data to be inferred and the input data such as the model structure and parameters are transferred to the xPU video memory through the DMA engine, and the asynchronous computing task is started. Monitor the xPU task queue and trigger load diversion when the queuing delay exceeds the threshold (such as 5ms).
[0058] Case 2: If there is a physical acceleration card and the physical acceleration card supports some operators, the model file to be processed is loaded into the video memory space of the physical acceleration card; the application layer interface in user state is called, and some operators are routed to the inference operator library of the physical acceleration card through the dynamic routing engine, and each operator is mapped to the corresponding target operator in the inference operator library; the remaining operators are loaded into the memory space of the CPU, and the virtual operator library that matches all the remaining operators is determined from the virtual acceleration card operator library preset in the user state, and the remaining operators are mapped to the virtual operator library; based on the inference operator library and the virtual operator library, the inference of the model to be processed is completed.
[0059] Furthermore, based on the inference operator library and the virtual operator library, the inference of the model to be processed is accelerated, including: The remaining operators are mapped to a virtual operator library, and according to the dependency relationship between the remaining operators, the corresponding operators in the virtual operator library are called to perform calculations to obtain a first calculation result; the first calculation result is sent to the physical acceleration card through a preset shared memory channel, and according to the dependency relationship between some operators, the corresponding operators in the inference operator library are called to perform calculations to obtain a second calculation result; the first calculation result and the second calculation result are integrated through the physical acceleration card to complete the inference of the model to be processed.
[0060] It is understandable that in this method, the reasoning task is realized through the cooperation of the physical accelerator card and the CPU. The physical accelerator card starts the asynchronous computing task by transferring the input data to the xPU video memory through the DMA engine. The CPU adopts SIMD mode, activates the vector register management module, divides the data into blocks according to the number of SIMD channels, and uses double buffering technology to overlap data loading and calculation. Based on this synergy, the efficiency and performance of model reasoning can be effectively improved.
[0061] There may be acceleration anomalies in the process of model reasoning acceleration. In order to ensure the accuracy of reasoning acceleration, optionally, in an exemplary embodiment, the model reasoning acceleration method also includes: when an abnormality of the physical acceleration card is detected or the target operator is missing, the operator supported by the physical acceleration card or the model operator corresponding to the target operator is loaded into the memory space of the CPU, and the step of determining the target operator that matches all operators in the model information from the virtual operator library preset in the user state is executed.
[0062] For example, when an xPU operator is missing or a hardware failure is detected, the task is migrated to the virtual operator library corresponding to the CPU, the target operator is determined by mapping the operator to the virtual operator library, and the model reasoning to be processed is completed based on the target operator. Furthermore, the hardware support matrix is updated and abnormal events are recorded for reference in subsequent routing decisions. This method avoids the situation where the entire reasoning task cannot be carried out smoothly or the performance is seriously degraded due to high dependence on a specific xPU. When the accelerator card is missing or some functions are not sound,
[0063] In order to improve resource utilization, resources need to be recycled when reasoning ends. Optionally, in an exemplary embodiment, based on the kernel-mode memory management module and the user-mode vOplib, a secure erase is performed on the memory pool associated with the virtual card instance to prevent data residue, release the DMA mapping relationship, reset the mapping table, and thereby release memory resources; the vector register state of the SIMD operator is cleaned to prevent cross-task data contamination, and the compiled temporary code cache is recycled to release space, thereby resetting the operator state; optionally, a hardware-level reset is performed on the xPU to ensure that there are no residual computing tasks, and the performance monitoring unit (PMU) counter is turned off to release interrupt resources, thereby resetting the hardware context.
[0064] In one embodiment, Figure 3 As shown, a model reasoning acceleration method is provided. This embodiment takes the method applied to a terminal as an example. In this embodiment, the method includes the following steps: Step 302: Obtain the model file to be processed of the model to be processed.
[0065] Step 304: parse the model file to be processed to obtain all operators of the model to be processed.
[0066] Step 306, determine whether there is a physical acceleration card, if not, execute step 308, if yes, execute step 312.
[0067] Step 308: If there is no physical acceleration card, the model file to be processed is loaded into the memory space of the CPU, and a target operator matching all operators is determined from the virtual operator library preset in the user state.
[0068] Step 310, mapping each operator to its corresponding target operator, and accelerating the inference of the model to be processed based on the target operator.
[0069] Step 312: If there is a physical acceleration card and the physical acceleration card supports all operators, the model file to be processed is loaded into the video memory space of the physical acceleration card.
[0070] Step 314, calling the user-mode application layer interface, mapping each operator to its corresponding target operator in the inference operator library according to all operators routed to the inference operator library of the physical acceleration card through the dynamic routing engine, and accelerating the inference of the model to be processed based on the inference operator library.
[0071] Step 316: If there is a physical acceleration card and the physical acceleration card supports some operators, the model file to be processed is loaded into the video memory space of the physical acceleration card, the remaining operators are loaded into the memory space of the CPU, and a virtual operator library matching all the remaining operators is determined from the virtual acceleration card operator library preset in the user state, and the remaining operators are mapped to the virtual operator library.
[0072] Step 318, based on the inference operator library and the virtual operator library, complete the inference of the model to be processed.
[0073] It should be noted that the specific implementation of this embodiment can be achieved through the above-mentioned limited manner, which will not be elaborated here.
[0074] Based on the above model reasoning acceleration method, a timing diagram of the model reasoning acceleration method based on a virtual accelerator card and xPU is provided, such as Figure 4 As shown, it includes an application, a virtual accelerator card API, a scheduling engine, a physical accelerator card, and a CPU computing unit (i.e., a central processing unit computing unit). An inference request is initiated through the application, and the model file of the model to be processed is obtained through the virtual accelerator card API. The model file to be processed is parsed to obtain a calculation graph, and the operator hierarchical analysis is performed based on the calculation graph to allocate computing tasks. Among them, the operator hierarchical analysis includes three cases: L1 level: fully matching the native instructions of the accelerator card, L2 level: CPU-assisted calculation is required, and L3 level: completely dependent on CPU calculation. For the L1 level, if there is a physical accelerator card and all operators are supported, the model file is loaded into the video memory of the physical accelerator card for inference acceleration, and the intermediate result is returned to the scheduling engine. For the L2 level, if there is a physical accelerator card and some operators are supported, the calculation tasks are allocated through the scheduling engine, and the supported operators are allocated to the physical accelerator card to obtain the block result A, and the rest are loaded into the CPU. The reasoning tasks of the remaining operators are calculated through SIMD to obtain the block result B. Integrate block results A and block results B, return the intermediate results to the scheduling engine, and aggregate the final results through the scheduling engine and return them to the application.
[0075] For L3 level, if there is no physical acceleration card, the entire model file can be loaded into the CPU memory space, and model reasoning can be achieved through full SIMD calculation.
[0076] In the above-mentioned model reasoning acceleration method, in the case where there is no physical acceleration card, the model file is directly loaded into the CPU, and the previously integrated vector registers and SIMD instruction sets are used to optimize the model data for parallel processing to achieve model reasoning acceleration, avoiding the high dependence on a specific xPU. When the acceleration card is missing or some functions are not sound, the entire reasoning task cannot be carried out smoothly or the performance is seriously reduced; in the case where a physical acceleration card exists and only some operators are supported, the operators of the model file are split into the physical acceleration card and the CPU, and the physical acceleration card and the CPU are used for collaborative reasoning, which solves the incompatibility of operator libraries of different acceleration cards. When encountering operators that are not supported by the acceleration card, there is a lack of effective alternatives, which increases development complexity and cost. In addition, it solves the problem of single model loading and memory mapping methods in related technologies, no optimization for the characteristics of neural networks, poor memory access performance, and affecting reasoning speed.
[0077] In an exemplary embodiment, an application based on the above-mentioned model reasoning acceleration method is provided, such as Figure 5 As shown, an application flow chart of the model reasoning acceleration method is provided, which specifically includes: In response to the inference task request, the model file of the model to be processed is obtained, the model file is read, the memory capacity required for the inference execution of the model file to be processed and the current memory pool capacity are obtained, and a memory judgment is performed, that is, whether the current memory pool capacity meets the memory capacity. If it is not enough and there is currently a model inference task occupying the memory that can be released, the memory occupied by the model is released. If the current memory pool capacity is greater than the corresponding storage capacity required for the inference execution of the model file, the model file is loaded and a successful loading judgment is performed. If the loading is successful, the model file to be processed is converted into a model intermediate representation, and a calculation graph is extracted according to the model intermediate representation; according to the calculation graph, all operators of the model to be processed are obtained.
[0078] Operator hierarchical processing is performed based on the operator. The hardware environment is detected in real time through the virtual acceleration card, and the input image data is obtained. The image data is preprocessed, including image data processing such as flipping and splicing. The processed data is input into the model for inference calculation. If there is a physical acceleration card and all operators are supported, the model reasoning is completed based on the acceleration card through the acceleration card memory allocation; if there is a physical acceleration card and only some operators are supported, a hybrid computing mode is used to split the constructed calculation graph, that is, split the calculation path, and complete the model reasoning based on the CPU and xPU collaboration. If there is no physical acceleration card, a pure CPU mode is used to load the model file into the CPU memory, and the model reasoning is completed based on the virtual acceleration card through CPU memory optimization. The reasoning result is aggregated according to the calculation result of the reasoning, and the reasoning result is output. For example, the reasoning result can be an image recognition result or an image classification result.
[0079] In the above method, by detecting the physical accelerator card and its support, it is ensured that the subsequent reasoning tasks can be executed in a suitable hardware environment, avoiding task failures caused by hardware not supporting certain operators; by loading the model file into the video memory space of the physical accelerator card, the delay of data transmission is reduced, the data access speed in the reasoning process is improved, and the CPU resources are released at the same time; by calling the user-mode application layer interface, the interaction with the physical accelerator card is simplified, and the development difficulty is reduced; by utilizing the high-performance operator library of the physical accelerator card, the reasoning speed and efficiency are significantly improved, and the integrity and accuracy of the reasoning results are ensured. Overall, by dynamically coordinating the CPU and xPU, resources are reasonably allocated according to task requirements, resource idleness or overload is avoided, overall resource utilization is improved, and the problem of reasoning task interruption caused by the lack of accelerator cards or imperfect functions is effectively solved. Whether in low-cost embedded devices or complex server environments, reasoning tasks can be stably run, broadening the scope of application scenarios.
[0080] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0081] Based on the same inventive concept, the embodiment of the present application also provides a model reasoning acceleration system for implementing the model reasoning acceleration method involved above. The implementation scheme for solving the problem provided by the system is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more model reasoning acceleration system embodiments provided below can refer to the limitations of the model reasoning acceleration method above, and will not be repeated here.
[0082] In an exemplary embodiment, Figure 6 As shown, a model reasoning acceleration system is provided, the system includes a physical hardware layer, a virtual acceleration card and an application layer, the physical hardware layer includes a CPU, wherein: The application layer is used to initiate an inference request and determine the model file of the model to be processed carried by the inference request; A virtual accelerator card is used to obtain the model file to be processed; parse the model file to be processed to obtain the model information of the model to be processed; if there is no physical accelerator card, the model file to be processed is loaded into the memory space of the CPU, and the target operator matching all operators in the model information is determined from the virtual operator library preset in the user state; each operator is mapped to its corresponding target operator, and the inference of the model to be processed is completed based on the target operator.
[0083] The above-mentioned model reasoning acceleration system, when reasoning on the model, parses the model file of the model to be processed to obtain all the operators of the model. In the absence of a physical acceleration card, the model file is loaded into the memory space of the CPU, and the target operator matching all the operators is determined from the virtual operator library preset in the user state; each operator is mapped to its corresponding target operator, and the reasoning of the model to be processed is accelerated based on the target operator, avoiding the current situation where the hardware resource configuration is limited and reasoning acceleration cannot be achieved without an acceleration card, fully tapping the potential of the CPU and improving the resource utilization of the CPU. In addition, it meets diverse business needs.
[0084] Optionally, in an exemplary embodiment, the virtual accelerator card is also used to obtain the memory capacity required for executing reasoning for the model file to be processed and the current memory pool capacity; if the current memory pool capacity is less than the memory capacity, the dynamic expansion mechanism is triggered to determine the memory capacity in the current memory pool to be allocated to the model file to be processed.
[0085] Convert the model file to be processed into an intermediate representation of the model, and extract the computational graph based on the intermediate representation of the model; According to the calculation graph, all operators of the model to be processed are obtained.
[0086] Optionally, in an exemplary embodiment, the virtual accelerator card is also used to map each operator to its corresponding target operator, call the target operator to perform calculations according to the dependencies between the operators in the calculation graph, and obtain calculation results; and obtain reasoning results according to the calculation results and reasoning tasks.
[0087] Optionally, in an exemplary embodiment, the virtual accelerator card is further used to load the model file to be processed into the video memory space of the physical accelerator card if there is a physical accelerator card and the physical accelerator card supports all operators in the model information; Call the user-state application layer interface, route all operators to the inference operator library of the physical accelerator card through the dynamic routing engine of the unified operator library, map each operator to its corresponding target operator in the inference operator library, and complete the inference of the model to be processed based on the inference operator library.
[0088] Optionally, in an exemplary embodiment, the virtual acceleration card is also used to load the model file to be processed into the video memory space of the physical acceleration card if a physical acceleration card exists and the physical acceleration card supports some operators; call the application layer interface in user state, route some operators to the inference operator library of the physical acceleration card through the dynamic routing engine of the unified operator library, and map each operator in the partial operators to the corresponding target operator in the inference operator library; load the remaining operators into the memory space of the CPU, determine the virtual operator library that matches all the remaining operators from the virtual acceleration card operator library preset in the user state, and map the remaining operators to the virtual operator library; and complete the inference of the model to be processed based on the inference operator library and the virtual operator library.
[0089] Optionally, in an exemplary embodiment, the virtual acceleration card is also used to map the remaining operators to a virtual operator library, and according to the dependency relationship between the remaining operators, call the corresponding operators in the virtual operator library to perform calculations to obtain a first calculation result; send the first calculation result to the physical acceleration card through a preset shared memory channel; map some operators to an inference operator library, and according to the dependency relationship between some operators, call the corresponding operators in the inference operator library to perform calculations to obtain a second calculation result; integrate the first calculation result and the second calculation result through the physical acceleration card to complete the inference of the model to be processed.
[0090] Optionally, in an exemplary embodiment, the virtual accelerator card is also used to load the operator supported by the physical accelerator card or the model operator corresponding to the target operator into the memory space of the CPU when an abnormality of the physical accelerator card is detected or the target operator is missing.
[0091] Each module in the above model reasoning acceleration system can be implemented in whole or in part by software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.
[0092] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 7 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, a model reasoning acceleration method is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.
[0093] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0094] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0095] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0096] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0097] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0098] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.
[0099] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0100] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A model reasoning acceleration method, characterized in that: The method comprises: Obtain a model file to be processed of the model to be processed; Parsing the model file to be processed to obtain model information of the model to be processed; If there is no physical acceleration card, the model file to be processed is loaded into the memory space of the CPU, and a target operator matching all operators in the model information is determined from a virtual operator library preset in the user state; Each of the operators is mapped to a corresponding target operator, and the reasoning of the model to be processed is completed based on the target operator.
2. The method according to claim 1, characterized in that The method further comprises: Obtain the memory capacity required for executing reasoning of the model file to be processed and the current memory pool capacity; If the current memory pool capacity is smaller than the memory capacity, a dynamic expansion mechanism is triggered to determine the memory capacity in the current memory pool to be allocated to the model file to be processed; The step of parsing the model file to be processed to obtain model information of the model to be processed includes: Converting the model file to be processed into a model intermediate representation, and determining a computational graph of the model file to be processed; According to the calculation graph, all operators of the model to be processed are obtained.
3. The method according to claim 2, characterized in that Mapping each of the operators to a corresponding target operator and completing the inference of the model to be processed based on the target operator includes: Mapping each of the operators to a corresponding target operator, and calling the target operator to perform calculations according to the dependency relationship between the operators in the calculation graph to obtain calculation results; Processing is performed according to the calculation results to obtain inference results, thereby completing the inference of the model to be processed.
4. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: If there is a physical acceleration card and the physical acceleration card supports all operators in the model information, then the model file to be processed is loaded into the video memory space of the physical acceleration card; Call the user-state application layer interface, route all the operators to the inference operator library of the physical acceleration card through the dynamic routing engine of the unified operator library, map each of the operators to the corresponding target operator in the inference operator library, and complete the inference of the model to be processed based on the inference operator library.
5. The method according to claim 4, characterized in that The method further comprises: If there is a physical acceleration card and the physical acceleration card supports some of the operators, the model file to be processed is loaded into the video memory space of the physical acceleration card; Calling the user-state application layer interface, routing some of the operators to the inference operator library of the physical acceleration card through the dynamic routing engine of the unified operator library, and mapping each of the operators in some of the operators to the corresponding target operators in the inference operator library; Load the remaining operators into the memory space of the CPU, determine a virtual operator library matching all the remaining operators from the virtual accelerator card operator library preset in the user state, and map the remaining operators to the virtual operator library; Based on the inference operator library and the virtual operator library, the inference of the model to be processed is completed.
6. The method according to claim 5, characterized in that The step of accelerating the inference of the model to be processed based on the inference operator library and the virtual operator library includes: Mapping the remaining operators to the virtual operator library, and calling corresponding operators in the virtual operator library to perform calculations according to the dependency relationships between the remaining operators to obtain a first calculation result; Sending the first calculation result to the physical acceleration card through a preset shared memory channel; Mapping some of the operators to the inference operator library, and calling corresponding operators in the inference operator library to perform calculations according to the dependency relationships between some of the operators to obtain a second calculation result; The first calculation result and the second calculation result are integrated through the physical acceleration card to complete the reasoning of the model to be processed.
7. The method according to any one of claim 5 or claim 6, characterized in that: The method further comprises: When it is detected that the physical acceleration card is abnormal or the target operator is missing, the operator supported by the physical acceleration card or the operator corresponding to the target operator is loaded into the memory space of the CPU, and the step of determining the target operator matching all the operators in the model information from the virtual operator library preset in the user state is executed.
8. A model reasoning acceleration system, characterized in that: The system includes a physical hardware layer, a virtual accelerator card and an application layer, wherein the physical hardware layer includes a CPU, wherein: The application layer is used to initiate an inference request and determine a to-be-processed model file of the to-be-processed model carried by the inference request; The virtual acceleration card is used to obtain the model file to be processed; parse the model file to be processed to obtain model information of the model to be processed; if there is no physical acceleration card, load the model file to be processed into the memory space of the CPU, and determine the target operator that matches all operators in the model information from the virtual operator library preset in the user state; map each of the operators to the corresponding target operator, and complete the reasoning of the model to be processed based on the target operator.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
TensorFlow system acceleration method and device based on FPGA equipment, equipment and storage medium
CN111858036A
Multi-hardware target depth model optimization deployment architecture supporting user-defined operator
CN113934410A
CPU-GPU heterogeneous resource-oriented task scheduling method
CN114911612A
Model reasoning performance optimization method and device and related product
CN115034402A
Graph computing platform based on cloud native technology
CN117009038A
Cited By
Hierarchical hybrid expert model-based reasoning method and system, and storage medium
CN120471184A
Application dynamic compatibility method and system, electronic equipment and storage medium
CN120872402A