Method and device for deploying AI model on heterogeneous platform

By using floating-point digital signal processors to preprocess, reason and post-process image data on heterogeneous platforms, the problems of resource waste and CPU burden are solved, and more efficient hardware resource utilization and CPU burden are achieved.

CN120494008APending Publication Date: 2025-08-15BEIJING JINGWEI HIRAIN TECH CO INC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510621954.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, deep neural network models fail to make full use of computing resources when deploying to heterogeneous platforms, resulting in waste of resources and causing high computing burdens to the CPU, affecting the efficiency of handling other programs.

Method used

Image data is preprocessed, inferenced and postprocessed through floating-point digital signal processors in heterogeneous platforms, and kernel operations are written using OpenVX tools to allocate different computing tasks to each hardware to reduce CPU burden.

Benefits of technology

It realizes the full use of heterogeneous platform hardware resources, reduces the CPU computing burden, improves overall operation efficiency, and reasonably allocates CPU and other hardware tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494008A_ABST
    Figure CN120494008A_ABST
Patent Text Reader

Abstract

The invention provides a method and device for deploying an AI model on a heterogeneous platform. The method comprises the steps of obtaining a model file of the AI model to be deployed and image data of the AI model to be input; preprocessing the image data by using a first floating-point number digital signal processor in the heterogeneous platform to obtain preprocessed image data; reasoning based on a second floating-point number digital signal processor, the model file and the preprocessed image data in the heterogeneous platform to obtain a reasoning result; a third floating-point number digital signal processor in the heterogeneous platform is used for post-processing the reasoning result; and displaying the post-processed reasoning result and image data through a visual node. According to the scheme, the hardware on the heterogeneous platform is called, different AI model deployment processes are executed, and compared with the mode that the deployment processes are executed directly through a CPU, the purposes that computing resources of the hardware of the heterogeneous platform are fully utilized, and the computing burden of the CPU is reduced are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and device for deploying an AI model on a heterogeneous platform. Background Art

[0002] A heterogeneous platform refers to a system composed of different types of computing resources, such as CPUs, GPUs, and FPGAs. Deploying a deep neural network model (i.e., an AI model) onto a heterogeneous platform allows it to leverage the strengths of various hardware types to optimize the performance and efficiency of the deep neural network model.

[0003] In existing technology, the process for deploying deep neural network models on heterogeneous platforms is basically as follows: 1. Collect data and design the model structure; 2. Use a high-computing, high-memory GPU for model training; 3. Quantize the trained model using a platform-specific quantization toolchain; 4. Deploy and run the quantized model on the heterogeneous platform. Existing technology only places the AI model on the corresponding acceleration hardware for inference, while data pre-processing and post-processing of the model inference results are all handled by the CPU. Taking the deployment of the IRIS algorithm on the TDA4VM platform as an example, the current deployment solution consumes approximately 20% of CPU resources, but the two C6x (C6x floating-point digital signal processors) are completely unused.

[0004] In summary, the existing technology has two shortcomings. On the one hand, it fails to fully utilize the computing resources of heterogeneous platforms, resulting in a waste of resources; on the other hand, it imposes a high computing burden on the CPU, resulting in reduced efficiency in processing other programs. Summary of the Invention

[0005] In view of this, an embodiment of the present invention provides a method and device for deploying AI models on a heterogeneous platform, so as to fully utilize the computing resources of each hardware of the heterogeneous platform and reduce the computing burden of the CPU.

[0006] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:

[0007] A first aspect of an embodiment of the present invention discloses a method for deploying an AI model on a heterogeneous platform, which is applied to the heterogeneous platform. The method includes:

[0008] Obtaining a model file of the AI model to be deployed, and obtaining image data to be input into the AI model through a camera node in the heterogeneous platform; the AI model is pre-trained by the training platform, and the model file is obtained by converting the AI model into a format that meets quantization requirements and then performing quantization processing on the training platform;

[0009] Preprocessing the image data using a first floating-point digital signal processor in the heterogeneous platform to obtain the preprocessed image data;

[0010] Performing reasoning based on the second floating-point digital signal processor in the heterogeneous platform, the model file, and the pre-processed image data to obtain an inference result;

[0011] Post-processing the inference result using a third floating-point digital signal processor in the heterogeneous platform to obtain the inference result after the post-processing;

[0012] The inference result and the image data after the post-processing are displayed through a visualization node.

[0013] Preferably, the step of obtaining a model file of the AI model to be deployed, and obtaining image data to be input into the AI model through a camera node in the heterogeneous platform, includes:

[0014] Get the model file of the AI model to be deployed;

[0015] By executing, through the camera node in the heterogeneous platform, a process of calling the camera to obtain image data and locking the image data based on the OpenVX tool and API interface, thereby obtaining the locked image data;

[0016] The locked image data is subjected to color space conversion to obtain the image data that meets the input format requirements of the AI model.

[0017] Preferably, the preprocessing of the image data by using the first floating-point digital signal processor in the heterogeneous platform to obtain the preprocessed image data includes:

[0018] Obtaining a kernel operation corresponding to a preprocessing process written using an OpenVX tool for a first floating-point digital signal processor in the heterogeneous platform;

[0019] Utilizing a first floating-point digital signal processor in the heterogeneous platform to execute a kernel operation corresponding to a preprocessing process, and performing the normalization process on the image data;

[0020] Based on a plurality of scaling nodes connected in series that are pre-constructed using an OpenVX tool, the large-scale scaling process is performed on the image data after the normalization process to obtain the pre-processed image data.

[0021] Preferably, performing reasoning based on the second floating-point digital signal processor in the heterogeneous platform, the model file, and the pre-processed image data to obtain an inference result includes:

[0022] Obtaining a kernel operation corresponding to an inference process written using an OpenVX tool for a second floating-point digital signal processor in the heterogeneous platform;

[0023] The kernel operation is executed by using a second floating-point digital signal processor in the heterogeneous platform, and reasoning is performed based on the model file and the preprocessed image data to obtain an inference result.

[0024] Preferably, the post-processing the inference result by using the third floating-point digital signal processor in the heterogeneous platform to obtain the inference result after the post-processing includes:

[0025] Obtaining a kernel operation corresponding to a post-processing process written using an OpenVX tool for a third floating-point digital signal processor in the heterogeneous platform;

[0026] Utilizing a third floating-point digital signal processor in the heterogeneous platform to execute a kernel operation corresponding to a post-processing process, and performing splicing processing on the inference results to obtain the spliced inference results; wherein the spliced inference results are consistent with the output results of the AI model before the quantization processing in terms of data dimension;

[0027] An offset reset process is performed on the inference result after the splicing process to obtain the inference result after the post-processing.

[0028] Preferably, if the last layer of the AI model is a fully connected layer, after performing the post-processing on the inference result, the method further includes:

[0029] Using a third floating-point digital signal processor in the heterogeneous platform, performing non-maximum suppression processing on the inference result after the post-processing to obtain a target inference result with a higher confidence level;

[0030] The target inference result and the image data are displayed through a visualization node.

[0031] A second aspect of an embodiment of the present invention discloses a device for deploying an AI model on a heterogeneous platform, which is applied to a heterogeneous platform. The device includes:

[0032] An acquisition unit, configured to acquire a model file of an AI model to be deployed and image data to be input into the AI model; the AI model is pre-trained by a training platform, and the model file is obtained by the training platform converting the AI model into a format that meets quantization requirements and then performing quantization processing;

[0033] a preprocessing unit, configured to preprocess the image data using a first floating-point digital signal processor in the heterogeneous platform to obtain the preprocessed image data;

[0034] an inference unit, configured to perform inference based on the second floating-point digital signal processor in the heterogeneous platform, the model file, and the preprocessed image data to obtain an inference result;

[0035] a post-processing unit, configured to perform post-processing on the inference result using a third floating-point digital signal processor in the heterogeneous platform to obtain the inference result after the post-processing;

[0036] A display unit is used to display the inference result and the image data after the post-processing through a visualization node.

[0037] Preferably, the acquisition unit is specifically used to:

[0038] Get the model file of the AI model to be deployed;

[0039] Based on the OpenVX tool and API interface, the camera is called to obtain image data and the image data is locked;

[0040] Performing color space conversion on the locked image data to obtain the image data that meets the input format requirements of the AI model. A third aspect of an embodiment of the present invention discloses an electronic device, including: a memory and a processor;

[0041] The memory is used to store computer programs;

[0042] The processor is used to execute the computer program, specifically to implement any one of the methods disclosed in the first aspect of the embodiment of the present invention.

[0043] The fourth aspect of an embodiment of the present invention discloses a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is executed by a processor, the device where the computer-readable storage medium is located is controlled to execute any one of the methods disclosed in the first aspect of the embodiment of the present invention.

[0044] Based on the above-mentioned embodiment of the present invention, a method and device for deploying an AI model on a heterogeneous platform is provided, which is applied to a heterogeneous platform. The method includes: obtaining a model file of the AI model to be deployed and image data to be input into the AI model; the AI model is pre-trained by a training platform, and the model file is converted by the training platform into a format that meets quantization requirements and then quantized; using a first floating-point digital signal processor in the heterogeneous platform, pre-processing the image data to obtain the pre-processed image data; performing inference based on a second floating-point digital signal processor in the heterogeneous platform, the model file, and the pre-processed image data to obtain an inference result; using a third floating-point digital signal processor in the heterogeneous platform, post-processing the inference result to obtain the post-processed inference result; and displaying the post-processed inference result and the image data through a visualization node. In this solution, various hardware on the heterogeneous platform are called to execute different AI model deployment processes. Compared with the method of directly using the CPU to execute the deployment process, this method fully utilizes the computing resources of various hardware on the heterogeneous platform and reduces the computing burden of the CPU. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0046] Figure 1 This is a flowchart of a method for deploying an AI model on a heterogeneous platform disclosed in an embodiment of the present invention;

[0047] Figure 2 A hardware resource allocation diagram for a heterogeneous platform disclosed in an embodiment of the present invention;

[0048] Figure 3 This is a structural diagram of an apparatus for deploying AI models on heterogeneous platforms disclosed in an embodiment of the present invention. DETAILED DESCRIPTION

[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0050] In this application, the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0051] As can be seen from the background technology, the existing technology has two shortcomings. On the one hand, it fails to fully utilize the computing resources of heterogeneous platforms, resulting in a waste of resources; on the other hand, it imposes a high computing burden on the CPU, resulting in reduced efficiency in processing other programs.

[0052] Therefore, an embodiment of the present invention discloses a method and device for deploying AI models on a heterogeneous platform. In this solution, various hardware on the heterogeneous platform are called to execute different AI model deployment processes. Compared with the method of directly using the CPU to execute the deployment process, this method achieves the purpose of fully utilizing the computing resources of various hardware on the heterogeneous platform and reducing the computing burden of the CPU.

[0053] It should be noted that deploying an AI model refers to the process of running a trained AI model in a specific environment. In this application, the specific environment refers to a heterogeneous platform.

[0054] like Figure 1 FIG. 1 is a flowchart of a method for deploying an AI model on a heterogeneous platform disclosed in an embodiment of the present invention. The method is applied to a heterogeneous platform and mainly includes the following steps:

[0055] Step S101: Obtain the model file of the AI model to be deployed and the image data to be input into the AI model.

[0056] In step S101, the AI model is pre-trained by the training platform. After the training platform converts the AI model file into a format that meets the quantization requirements (such as ONNX format), the AI model is quantized. For AI model operators that do not support deployment on heterogeneous platforms, corresponding operator replacement and model pruning are performed.

[0057] Among them, the process of quantizing the AI model to obtain a model file includes but is not limited to using the TIDL_tools tool chain to quantize the AI model to obtain a model file.

[0058] Among them, the model files include: NET files, and IO files containing the input and output configurations of the AI model.

[0059] Quantization is the process of converting floating-point numbers to fixed-point numbers.

[0060] Operators are the basic building blocks of AI models. They define the structure and operation process of AI models, including input, output, and intermediate calculations.

[0061] It should be noted that model pruning refers to the removal of some redundant parameters of the neural network, including pruning of neurons and pruning of weights.

[0062] Exemplarily, the AI model can be a face detection model in a pre-trained IRIS algorithm.

[0063] In the specific implementation process of step S101, the model file of the AI model to be deployed is obtained; based on the OpenVX tool and API interface, the camera is called to obtain image data and the image data is locked; the locked image data is converted into a color space to obtain image data that meets the AI model input format requirements.

[0064] It is understandable that the image format required by the AI model is obtained by adding color space conversion.

[0065] Among them, the process of the camera acquiring image data can be assigned to the camera node in the heterogeneous platform for execution. For example, the camera node can use mcu2_0 in the TDA4VM platform, thereby reasonably allocating the tasks of the mpu and mcu in the CPU.

[0066] It should be noted that the OpenVX tool is an open cross-platform acceleration standard for computer vision applications, which can achieve performance and power-optimized computer vision processing, especially in embedded application cases; the camera can be a USB camera.

[0067] The API interface can be Video for Linux two (Video4Linux2), abbreviated as V4L2. V4L2 is an API interface for collecting image, video, and audio data under the Linux operating system. With the appropriate video capture device and corresponding driver, it can realize the collection of images, videos, audio, etc. It has a wide range of applications in remote conferencing, videophones, video surveillance systems, and embedded multimedia terminals.

[0068] Locking is a mutual exclusion mechanism that ensures that only one object operates on the image data in the same memory at the same time, thereby ensuring the correctness of image data reading.

[0069] Step S102: preprocessing the image data using the first floating-point digital signal processor in the heterogeneous platform to obtain preprocessed image data.

[0070] In step S102 , the heterogeneous platform may be a TDA4VM platform. In smart cockpit products, the TDA4VM platform is a commonly used heterogeneous platform.

[0071] The first floating-point digital signal processor may be a processor numbered c6x_2 on the TDA4VM platform.

[0072] The method for allocating the preprocessing process to the first floating-point digital signal processor for processing is as follows:

[0073] For the first floating-point digital signal processor, the OpenVX tool is used to customize the kernel writing corresponding to the preprocessing process, and the first floating-point digital signal processor runs the written kernel to implement the first floating-point digital signal processor to perform the preprocessing process.

[0074] Among them, the kernel writing corresponding to the preprocessing process can be completed in advance to obtain the kernel operation corresponding to the preprocessing process. Subsequently, the first floating-point digital signal processor runs the kernel operation corresponding to the preprocessing process to realize the first floating-point digital signal processor executing the preprocessing process.

[0075] In the OpenVX tool, a kernel is an implementation of a specific operation or algorithm, but it does not directly execute the operation. It is more like a template that describes the specific content of the operation.

[0076] The actual unit of computation is the node, which is bound to a specific kernel and its input and output data. A node is an instance of a kernel. Each node is created based on a kernel.

[0077] It is understandable that the preprocessing process is assigned to the first floating-point digital signal processor, thereby fully utilizing the computing resources of other hardware on the heterogeneous platform and reducing CPU occupancy.

[0078] It should be noted that preprocessing can include normalization and large-scale scaling. The specific settings for preprocessing depend on the type of AI model, and different types of AI models may have different preprocessing processes. The purpose of this solution is to offload the preprocessing process to a processor other than the CPU on heterogeneous platforms to solve the problem of excessive CPU occupancy in existing technologies.

[0079] Taking normalization and large-scale scaling as an example, the specific preprocessing process includes:

[0080] First, the image data is normalized using a first floating-point digital signal processor in the heterogeneous platform.

[0081] Then, based on multiple serially connected scaling nodes pre-built through the OpenVX tool, the normalized image data is subjected to large-scale scaling processing to obtain pre-processed image data.

[0082] For example, the image data acquired by a USB camera has a resolution of 1280*720. The AI model to be deployed is the face detection model in the IRIS algorithm, which requires an input image resolution of 128*128. Since the default scaling node of the OpenVX tool only supports 4x scaling, two scaling nodes are designed in series to achieve a large-scale scaling function exceeding 4x.

[0083] Step S103: performing inference based on the second floating-point digital signal processor in the heterogeneous platform, the model file, and the pre-processed image data to obtain an inference result.

[0084] In step S103, the second floating-point digital signal processor may be a processor numbered c7x_1 on the TDA4VM platform.

[0085] The method for allocating the inference process to the second floating-point digital signal processor for processing is as follows:

[0086] For the second floating-point digital signal processor, the OpenVX tool is used to customize the kernel writing corresponding to the reasoning process, and the second floating-point digital signal processor runs the written kernel to implement the second floating-point digital signal processor to execute the reasoning process.

[0087] Among them, the kernel writing corresponding to the reasoning process can be completed in advance to obtain the kernel operation corresponding to the reasoning process, and then the second floating-point digital signal processor runs the kernel operation corresponding to the reasoning process to realize the second floating-point digital signal processor executing the reasoning process.

[0088] It is understandable that the inference process is assigned to the second floating-point digital signal processor, thereby fully utilizing the computing resources of other hardware on the heterogeneous platform and reducing CPU usage.

[0089] It should be noted that inference refers to the process of inputting preprocessed image data into the model file to obtain the output inference result.

[0090] Step S104: using the third floating-point digital signal processor in the heterogeneous platform to post-process the inference result to obtain a post-processed inference result.

[0091] In step S104 , the third floating-point digital signal processor may be a processor numbered c6x_1 on the TDA4VM platform.

[0092] The method for allocating the inference process to the third floating-point digital signal processor is as follows:

[0093] For the third floating-point digital signal processor, the OpenVX tool is used to customize the kernel writing corresponding to the post-processing process, and the third floating-point digital signal processor runs the written kernel to implement the third floating-point digital signal processor to perform the post-processing process.

[0094] Among them, the kernel writing corresponding to the post-processing process can be completed in advance to obtain the kernel operation corresponding to the post-processing process. Subsequently, the third floating-point digital signal processor runs the kernel operation corresponding to the post-processing process to realize the third floating-point digital signal processor executing the post-processing process.

[0095] It is understandable that the post-processing process is assigned to the third floating-point digital signal processor, thereby fully utilizing the computing resources of other hardware on the heterogeneous platform and reducing CPU usage.

[0096] In the specific implementation process of step S104, first, the third floating-point digital signal processor in the heterogeneous platform is used to splice the inference results to obtain the spliced inference results; the spliced inference results are consistent with the output results of the AI model before quantization processing in terms of data dimension.

[0097] Then, the offset reset processing is performed on the inference result after the splicing processing to obtain the inference result after post-processing.

[0098] It should be noted that since the above-mentioned AI model has been quantized and pruned during the quantization process, the inference results output by the AI model need to be spliced during the post-processing process so that they are consistent with the output results of the AI model before quantization in terms of data dimension.

[0099] In addition, since the above preprocessing is performed to meet the input requirements of the AI model, after the AI model outputs the inference result, post-processing corresponding to the preprocessing is required. Taking the preprocessing including normalization processing and large-scale scaling processing as an example, the corresponding post-processing includes: offset reset processing.

[0100] Offset resetting refers to adjusting the coordinates of the AI model's inference output based on the parameters of normalization and large-scale scaling. For example, if the image is scaled twice, the coordinates output by the AI model need to be multiplied by 2.

[0101] Step S105: Display the post-processed reasoning results and image data through visualization nodes.

[0102] In the specific implementation process of step S105, rendering is performed using the OpenCV tool to display the inference results after post-processing and the image data acquired by the camera.

[0103] Among them, the display process can be assigned to the visualization node in the heterogeneous platform for execution. The visualization node actually uses mpu1_0 in the TDA4VM platform, thereby reasonably allocating the tasks of the MPU and MCU in the CPU.

[0104] In one embodiment, if the last layer of the AI model is a fully connected layer, multiple inference results will be output, and a non-maximum suppression operation needs to be performed to filter the results. Therefore, after post-processing the inference results, the method further includes:

[0105] Using the third floating-point digital signal processor in the heterogeneous platform, non-maximum suppression is performed on the post-processed inference results to obtain target inference results with higher confidence.

[0106] Use OpenCV tools to draw and display target inference results and image data.

[0107] Similarly, the presentation process can be assigned to mpu1_0 in the TDA4VM platform for execution.

[0108] For example, the last layer of the face detection model in the IRIS algorithm is a fully connected layer, which will output multiple results. It is necessary to perform a non-maximum suppression operation to filter the inference results and obtain a face frame with a higher confidence level.

[0109] Based on the method of deploying AI models on heterogeneous platforms disclosed in the above embodiment of the present invention, Figure 2 , which is a diagram of hardware resource allocation for a heterogeneous platform disclosed in an embodiment of the present invention.

[0110] Among them, the visualization node is assigned to mpu1_0 of the heterogeneous platform, the camera node is assigned to mcu2_0 of the heterogeneous platform, the post-processing node is assigned to c6x_1 of the heterogeneous platform, the pre-processing node is assigned to c6x_2 of the heterogeneous platform, and the model inference is assigned to c7x_1 of the heterogeneous platform.

[0111] This allocation method maximizes the utilization of hardware resources on heterogeneous platforms. Compared to solutions that directly place both pre- and post-processing on the CPU, the solution proposed in this invention achieves more rational hardware resource utilization, increasing the occupancy of the two C6x cores from 0% to 11% and 4%, respectively, while simultaneously reducing the CPU occupancy from 20% to 3%. This allows the CPU to perform more other tasks, improving the overall operational efficiency of the TDA4VM development board and fully demonstrating the feasibility and effectiveness of the method described in this invention. Furthermore, the tasks performed by the MPU and MCU in the CPU are rationally allocated.

[0112] Based on the method for deploying an AI model on a heterogeneous platform disclosed in the above-mentioned embodiment of the present invention, a model file of the AI model to be deployed and image data to be input into the AI model are obtained; the AI model is pre-trained by the training platform, and the model file is obtained by the training platform by converting the AI model into ONNX format and then quantizing it; the image data is pre-processed using the first floating-point digital signal processor in the heterogeneous platform to obtain pre-processed image data; reasoning is performed based on the second floating-point digital signal processor, the model file and the pre-processed image data in the heterogeneous platform to obtain an inference result; the inference result is post-processed using the third floating-point digital signal processor in the heterogeneous platform to obtain the inference result after post-processing; and the inference result and image data after post-processing are displayed through a visualization node. In this solution, various hardware on the heterogeneous platform are called to execute different AI model deployment processes. Compared with the method of directly using the CPU to execute the deployment process, the purpose of fully utilizing the computing resources of various hardware on the heterogeneous platform and reducing the computing burden of the CPU is achieved.

[0113] Corresponding to the method for deploying AI models on heterogeneous platforms disclosed in the above embodiment of the present invention, Figure 3 As shown, this is a structural diagram of a device for deploying AI models on a heterogeneous platform disclosed in an embodiment of the present invention, including: an acquisition unit 301, a preprocessing unit 302, an inference unit 303, a post-processing unit 304 and a display unit 305.

[0114] The acquisition unit 301 is used to obtain the model file of the AI model to be deployed and the image data to be input into the AI model; the AI model is pre-trained by the training platform, and the model file is obtained by the training platform by converting the AI model into ONNX format and then quantizing it.

[0115] In one embodiment, the acquiring unit 301 is specifically configured to:

[0116] Get the model file of the AI model to be deployed;

[0117] Based on OpenVX tools and API interfaces, the camera is called to obtain image data and lock the image data;

[0118] Perform color space conversion on the locked image data to obtain image data that meets the AI model input format requirements.

[0119] The preprocessing unit 302 is configured to preprocess the image data using the first floating-point digital signal processor in the heterogeneous platform to obtain preprocessed image data.

[0120] In one embodiment, the pre-processing unit 302 is specifically configured to:

[0121] Normalizing the image data using a first floating-point digital signal processor in the heterogeneous platform;

[0122] Based on multiple serially connected scaling nodes pre-built using the OpenVX tool, large-scale scaling is performed on the normalized image data to obtain pre-processed image data.

[0123] The inference unit 303 is configured to perform inference based on the second floating-point digital signal processor in the heterogeneous platform, the model file, and the preprocessed image data to obtain an inference result.

[0124] The post-processing unit 304 is configured to perform post-processing on the inference result using the third floating-point digital signal processor in the heterogeneous platform to obtain a post-processed inference result.

[0125] In one embodiment, the post-processing unit 304 is specifically configured to:

[0126] The inference results are spliced using the third floating-point digital signal processor in the heterogeneous platform to obtain a spliced inference result. The spliced inference result is consistent with the output result of the AI model before quantization in terms of data dimension.

[0127] The offset reset processing is performed on the inference result after the splicing processing to obtain the inference result after post-processing.

[0128] The display unit 305 is used to display the inference results and image data after post-processing through visualization nodes.

[0129] In one embodiment, the apparatus for deploying an AI model on a heterogeneous platform further includes:

[0130] The screening unit is used to, if the last layer of the AI model is a fully connected layer, perform non-maximum suppression on the inference results after post-processing using the third floating-point digital signal processor in the heterogeneous platform to obtain a target inference result with a higher confidence level; and display the target inference result and image data through a visualization node.

[0131] Based on the above-mentioned embodiment of the present invention, a device for deploying an AI model on a heterogeneous platform is disclosed, which obtains a model file of the AI model to be deployed and image data to be input into the AI model; the AI model is pre-trained by the training platform, and the model file is obtained by the training platform by converting the AI model into ONNX format and then quantizing it; the image data is pre-processed using the first floating-point digital signal processor in the heterogeneous platform to obtain pre-processed image data; reasoning is performed based on the second floating-point digital signal processor, the model file, and the pre-processed image data in the heterogeneous platform to obtain an inference result; the inference result is post-processed using the third floating-point digital signal processor in the heterogeneous platform to obtain the post-processed inference result; and the post-processed inference result and image data are displayed through a visualization node. In this solution, various hardware on the heterogeneous platform are called to execute different AI model deployment processes. Compared with the method of directly using the CPU to execute the deployment process, the purpose of fully utilizing the computing resources of various hardware on the heterogeneous platform and reducing the computing burden of the CPU is achieved.

[0132] An embodiment of the present invention further discloses an electronic device, comprising: a memory and a processor;

[0133] The memory is used to store computer programs; the processor is used to execute computer programs, specifically to implement a method for deploying AI models on a heterogeneous platform disclosed in the above-mentioned embodiment of the present invention.

[0134] An embodiment of the present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is executed by a processor, the device where the computer-readable storage medium is located is controlled to execute a method for deploying an AI model on a heterogeneous platform disclosed in the above embodiment of the present invention.

[0135] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0136] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0137] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for deploying AI models on heterogeneous platforms, characterized in that: Applied to heterogeneous platforms, the method includes: Obtaining a model file of the AI model to be deployed, and obtaining image data to be input into the AI model through a camera node in the heterogeneous platform; the AI model is pre-trained by the training platform, and the model file is obtained by converting the AI model into a format that meets quantization requirements and then performing quantization processing on the training platform; Preprocessing the image data using a first floating-point digital signal processor in the heterogeneous platform to obtain the preprocessed image data; Performing reasoning based on the second floating-point digital signal processor in the heterogeneous platform, the model file, and the pre-processed image data to obtain an inference result; Post-processing the inference result using a third floating-point digital signal processor in the heterogeneous platform to obtain the inference result after the post-processing; The inference result and the image data after the post-processing are displayed through a visualization node.

2. The method according to claim 1, characterized in that The obtaining of a model file of the AI model to be deployed, and obtaining image data to be input into the AI model through a camera node in the heterogeneous platform, includes: Get the model file of the AI model to be deployed; By executing, through the camera node in the heterogeneous platform, a process of calling the camera to obtain image data and locking the image data based on the OpenVX tool and API interface, thereby obtaining the locked image data; The locked image data is subjected to color space conversion to obtain the image data that meets the input format requirements of the AI model.

3. The method according to claim 1, characterized in that The method of preprocessing the image data by using the first floating-point digital signal processor in the heterogeneous platform to obtain the preprocessed image data includes: Obtaining a kernel operation corresponding to a preprocessing process written using an OpenVX tool for a first floating-point digital signal processor in the heterogeneous platform; Utilizing a first floating-point digital signal processor in the heterogeneous platform to execute a kernel operation corresponding to a preprocessing process, and performing the normalization process on the image data; Based on a plurality of scaling nodes connected in series that are pre-constructed using an OpenVX tool, the large-scale scaling process is performed on the image data after the normalization process to obtain the pre-processed image data.

4. The method according to claim 1, wherein The performing reasoning based on the second floating-point digital signal processor in the heterogeneous platform, the model file, and the preprocessed image data to obtain an inference result includes: Obtaining a kernel operation corresponding to an inference process written using an OpenVX tool for a second floating-point digital signal processor in the heterogeneous platform; The kernel operation is executed by using a second floating-point digital signal processor in the heterogeneous platform, and reasoning is performed based on the model file and the preprocessed image data to obtain an inference result.

5. The method according to claim 1, characterized in that The post-processing of the inference result by using the third floating-point digital signal processor in the heterogeneous platform to obtain the inference result after the post-processing includes: Obtaining a kernel operation corresponding to a post-processing process written using an OpenVX tool for a third floating-point digital signal processor in the heterogeneous platform; Utilizing a third floating-point digital signal processor in the heterogeneous platform to execute a kernel operation corresponding to a post-processing process, and performing splicing processing on the inference results to obtain the spliced inference results; wherein the spliced inference results are consistent with the output results of the AI model before the quantization processing in terms of data dimension; An offset reset process is performed on the inference result after the splicing process to obtain the inference result after the post-processing.

6. The method according to any one of claims 1 to 5, characterized in that: If the last layer of the AI model is a fully connected layer, then after performing the post-processing on the inference result, the method further includes: Using a third floating-point digital signal processor in the heterogeneous platform, performing non-maximum suppression processing on the inference result after the post-processing to obtain a target inference result with a higher confidence level; The target inference result and the image data are displayed through a visualization node.

7. A device for deploying AI models on heterogeneous platforms, characterized in that: Applied to heterogeneous platforms, the device includes: An acquisition unit, configured to acquire a model file of an AI model to be deployed and image data to be input into the AI model; the AI model is pre-trained by a training platform, and the model file is obtained by the training platform converting the AI model into a format that meets quantization requirements and then performing quantization processing; a preprocessing unit, configured to preprocess the image data using a first floating-point digital signal processor in the heterogeneous platform to obtain the preprocessed image data; an inference unit, configured to perform inference based on the second floating-point digital signal processor in the heterogeneous platform, the model file, and the preprocessed image data to obtain an inference result; a post-processing unit, configured to perform post-processing on the inference result using a third floating-point digital signal processor in the heterogeneous platform to obtain the inference result after the post-processing; A display unit is used to display the inference result and the image data after the post-processing through a visualization node.

8. The device according to claim 7, characterized in that The acquisition unit is specifically configured to: Get the model file of the AI model to be deployed; Based on the OpenVX tool and API interface, the camera is called to obtain image data and the image data is locked; The locked image data is subjected to color space conversion to obtain the image data that meets the input format requirements of the AI model.

9. An electronic device, characterized in that: include: memory and processor; The memory is used to store computer programs; The processor is used to execute the computer program, specifically to implement the method according to any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed by a processor, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method for accelerating multi-outlet DNN reasoning by heterogeneous processor under edge computing

    CN114662661A

  • Post-processing method, device and system under large image small target detection

    CN116152627A

  • Model testing method based on heterogeneous platform, heterogeneous chip, equipment and medium

    CN116841911A

  • Video AI reasoning optimization method and device, computer equipment and storage medium

    CN117036912A

  • Voice-driven facial animation method and device, equipment and medium

    CN117115312A