Image Inference Model Acceleration Method, System, Electronic Device and Medium
By converting RGB images to YUV420 and compressing them into JPEG bytecode, merging the ONNX model hierarchy and accelerating image inference using CUDA kernel function, the problem of slow inference detection model is solved, and the overall acceleration effect of image inference is achieved.
Patent Information
- Application Number
- CN202310480323.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-04-28
AI Technical Summary
The existing detection models are too slow to perform image inference, especially when inference is generated on GPU to generate one-dimensional sequences and two-dimensional images, which affects the efficiency of the detection model.
The RGB image is accelerated by image encoding and compressed into JPEG image bytecode, and the Huffman tree is built for decoding; the ONNX model is configured to merge the network layer and the RGB image parameter layer, and model inference is performed through the CUDA kernel function; the mask image pixel values are accessed in parallel to accelerate the conversion of inference results.
The acceleration of image encoding, decoding, model inference and inference result conversion is achieved, and the overall speed of image inference is improved, especially the processing efficiency on the GPU.
Smart Images

Figure CN116843772B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image encoding and decoding, and specifically relates to an image inference model acceleration method, system, electronic device and medium. Background Art
[0002] With the implementation of the intelligent manufacturing strategy, machine vision recognition technology plays a significant role in traditional manufacturing industries. For example, in some detection scenarios of target objects, it is necessary to train corresponding deep learning models through a large number of target object pictures, and combine the trained detection model with the image acquisition device in the detection scenario, so that the target object can be detected from the images collected by the image acquisition device through the detection model.
[0003] For traditional detection model inference for receiving and transmitting image data, traditional CPU serial encoding and decoding are used. For example, when converting a one-dimensional sequence generated by inference on a GPU into a two-dimensional image, the speed is slow, and the speed of GPU inference using full 32-bit precision is also slow. These factors will affect the speed of image inference of the detection model. Summary of the Invention
[0004] In order to solve the problem of too slow inference speed when the existing detection model performs image inference, the present invention provides an image inference model acceleration method, system, electronic device and medium.
[0005] To solve the above technical problems, the present invention provides an image inference model acceleration method, which includes image encoding acceleration, compressing the collected image and performing format conversion to convert the RGB image into a YUV420 image and compress it into JPEG image bytecode; image decoding acceleration, constructing a Huffman tree according to the JPEG image bytecode and performing Huffman tree encoding for transformation to re-decode the RGB image; model inference acceleration, configuring the ONNX model and merging the network layer of the ONNX model and the parameter layer of the RGB image respectively, and obtaining the inference result of the image through translation of the merged result; inference result conversion acceleration, performing format conversion on the inference result to obtain a single-channel mask image, creating an empty image, writing the mask pixel values in the single-channel mask image into the empty image, and obtaining the mask image through parallel access.
[0006] In a further solution of the present application, in the step of image encoding acceleration, it further includes collecting an image and compressing and sampling it to 24bit through digital signal encoding to obtain an RGB encoded image; compressing and dumping the RGB encoded image to 12bit through downsampling and color space transformation to obtain a YUV420 image; compressing a single pixel of the YUV420 image to within 8bit through the JPEG compression algorithm to obtain JPEG image bytecode.
[0007] In a further embodiment of the present application, in the step of compressing a YUV420 image by JPEG, it further includes:
[0008] The image signal is frequency-divided into multiple cosine signals through discrete cosine transform decomposition; the data order within the array of multiple cosine signals is transformed through Zigzag scanning; the chrominance quantization of the high and low frequency data within the array is performed through a JPEG quantizer; and the quantization result is further compressed through the Huffman compression algorithm.
[0009] In a further embodiment of the present application, in the step of accelerating image decoding, it further includes: reading the JPEG image byte code according to the JPEG protocol standard and adjusting the byte code order according to the byte code storage address; constructing a Huffman tree and filtering out the high-frequency components in the byte code with the adjusted order; performing data conversion on the components after high-frequency filtering through inverse quantization; and performing an inverse transformation on the data after inverse quantization conversion through inverse-order Zigzag scanning to obtain the bitmap information of the image.
[0010] In a further embodiment of the present application, in the step of accelerating model inference, it further includes configuring an ONNX model for training a deep learning model based on the h5 file format and using the Keras library; merging the network layers of the ONNX model and the parameter layers of the RGB image respectively through operator merging; translating the operators in the ONNX model into CUDA kernel functions and running the CUDA kernel functions to obtain an inference result of a one-dimensional tensor.
[0011] In a further embodiment of the present application, when merging the ONNX model and the RGB image through operator merging, multiple consecutive network layers are merged through vertical network layer merging, and tensors with the same or similar sizes in the RGB image are fused through horizontal tensor integration.
[0012] In a further embodiment of the present application, in the step of accelerating the conversion of the inference result, it further includes reading the inference result, outputting a one-dimensional floating-point number sequence from the inference result and storing it as CHW data; converting the CHW data into the RGB channel values of the corresponding pixels; merging the multi-channel tensors into a single-channel mask image through the argmax function; creating an empty image with all pixel values being 0 and writing the mask pixel values in the mask image into the empty image; accessing the mask pixel values written into the empty image through multi-core parallelism to obtain the pixel coordinates of the corresponding mask pixel values; and combining the pixel coordinates and the RGB channel values of the corresponding pixels to obtain the mask image of the corresponding pixels.
[0013] To solve the above technical problems, the present invention further provides that the image inference model acceleration system includes: an image acquisition module for acquiring an original image to be detected; a host computer configured to: compress the acquired image and convert the format to convert the RGB image into JPEG image bytecode; construct a Huffman tree based on the JPEG image bytecode and perform Huffman tree encoding transformation to re-decode the RGB image; configure the ONNX model and merge the network layer of the ONNX model and the parameter layer of the RGB image respectively, and obtain the inference result of the image by translating the merged result; perform format conversion on the inference result to obtain a single-channel mask image, create an empty image, write the mask pixel values in the single-channel mask image into the empty image, and obtain the mask image through parallel access.
[0014] To solve the above technical problems, the present invention further provides an electronic device, which includes: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the above acceleration method is executed.
[0015] To solve the above technical problems, the present invention further provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores one or more programs, and one or more of the programs can be executed by one or more processors to implement the above acceleration method.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0017] Through the above technical solution, the accelerated inference effect of the image inference model is realized by optimizing the four processes of encoding, decoding, inference, and inference result transcoding respectively;
[0018] Among them, during encoding, the image is compressed and converted to a YUV420 image to facilitate color space conversion in the GPU, improving the image encoding speed. Then, the YUV420 image is compressed into JPEG image bytecode to facilitate constructing a Huffman tree based on the JPEG image bytecode and performing reverse scanning of the Huffman encoding in the GPU to achieve accelerated image decoding;
[0019] During model inference, the ONNX model is configured and optimized to reduce the model architecture of the model and the computational amount when the model processes images, thereby achieving accelerated model inference. Finally, an empty image is created and the mask pixel values written into the empty image are accessed in parallel to obtain the mask image, thereby achieving acceleration of converting the inference result to an image. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification. Together with the following specific implementation manners, they are used to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the accompanying drawings:
[0021] Figure 1 is a flowchart of the steps of an image inference model acceleration method provided by an embodiment of the present invention;
[0022] Figure 2 is a further flowchart of the steps of step S1 in the image inference model acceleration method provided by an embodiment of the present invention;
[0023] Figure 3 is a further flowchart of the steps of step S13 in the image inference model acceleration method provided by an embodiment of the present invention;
[0024] Figure 4 is a further flowchart of the steps of step S2 in the image inference model acceleration method provided by an embodiment of the present invention;
[0025] Figure 5 is a further flowchart of the steps of step S3 in the image inference model acceleration method provided by an embodiment of the present invention;
[0026] Figure 6 is a further flowchart of the steps of step S4 in the image inference model acceleration method provided by an embodiment of the present invention; and
[0027] Figure 7 is a schematic diagram of an electronic device provided by an embodiment of the present invention. Specific Embodiment
[0028] The following details the specific implementation manners of the present invention in conjunction with the accompanying drawings. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0029] Embodiments of the present invention provide an image inference model acceleration method, aiming to achieve the speed of image inference.
[0030]
Image Inference Model Acceleration Method
[0031] As Figure 1 shown, embodiments of the present invention provide an image inference model acceleration method including the following steps
[0032] S1: Image encoding acceleration, compressing and format-converting the collected image to convert the RGB image into a YUV420 image and compressing it into JPEG image bytecode;
[0033] S2: Image decoding acceleration. Construct a Huffman tree based on the JPEG image bytecode and perform Huffman tree encoding for transformation to re-decode the RGB image;
[0034] S3: Model inference acceleration. Configure the ONNX model and merge the network layer of the ONNX model and the parameter layer of the RGB image respectively, and obtain the inference result of the image by translating the merged result;
[0035] S4: Inference result conversion acceleration. Perform format conversion on the inference result to obtain a single-channel mask image. Create an empty image, write the mask pixel values in the single-channel mask image into the empty image, and obtain the mask image through parallel access.
[0036] In step S1, the image is converted to a YUV420 image through compressive sampling and color space transformation, and the YUV420 image is compressed based on the JPEG (Joint Photographic Experts Group) image compression standard to obtain the JERG image bytecode, so as to reduce the occupied space of the image. For example, for a picture taken by an 8-megapixel camera, the file in RGB picture format occupies 24M (3264 * 2448 * 3) of space, the space occupied after compressing the picture in YUV420 format is 12M (3264 * 2448 * 1.5), and the space occupied after compressing the picture in JPEG format is about 3M. Compared with the traditional image encoding method developed based on the encapsulation of the OpenJpeg library, since CPU serial encoding is used, the image encoding speed is affected. The image encoding of the present invention is implemented based on an embedded GPU. In the GPU, the RGB color space of the image is converted into the YUV420 color space, improving the image encoding speed and realizing image encoding acceleration.
[0037] In step S2, a Huffman tree is constructed through the JPEG image bytecode and the Huffman encoding is output. Based on the data frequency in the Huffman encoding, the Huffman encoding is compressed again and re-ordered, so as to realize the decoding acceleration of the RGB image. Finally, the Huffman encoding is decompressed into a BMP (Bitmap) signal through reverse "zigzag" scanning, and the bitmap signal is converted into an RGB image through color space conversion. Compared with the traditional image decoding method developed based on the encapsulation of the OpenJpeg library, the image encoding of the present invention is implemented based on an embedded GPU. By performing reverse scanning on the Huffman encoding in the GPU, the image decoding speed is improved and image decoding acceleration is realized.
[0038] In step S3, by configuring the ONNX model (Open Neural Network Exchange) and merging multiple consecutive network layers in the ONNX model according to the granularity of the hardware, and fusing tensors with the same or similar sizes in the RGB image, the model architecture of the ONNX model and the computational amount when the model processes the image are reduced. By translating the merged result into a CUDA kernel function and running the CUDA kernel function to obtain the inference result of the image, the effect of accelerating model inference is realized. Compared with the existing method of performing full-precision inference using a model in the original ONNX format, the model inference of the present invention uses a mixed-precision inference based on the optimized ONNX format model, improving the inference speed of the model for images and realizing the acceleration of model inference.
[0039] Finally, in step S4, the single-channel mask image corresponding to the inference result is obtained by performing format conversion on the inference result. By creating an empty image and writing the mask pixel values in the single-channel mask image into the empty image, and accessing the image pixels and RGB channel values in the empty image in a multi-core parallel manner to obtain the mask image of the accessed pixels. When all the image pixels in the empty image have been accessed, the inference result can be converted into a mask image corresponding to the original image. Compared with the existing method of serially reading the inference result and writing the image values pixel by pixel, the inference result conversion of the present invention uses a method of parallelly reading the inference result and randomly writing the image pixel values, improving the speed of converting the inference result into an image and realizing the acceleration of inference result transcoding.
[0040] As Figure 2 shown, in the steps of accelerating image encoding, the following steps are further included
[0041] S11: Collect the image and compress and sample it to 24 bits through digital signal encoding to obtain an RGB encoded image;
[0042] S12: Compress and dump the RGB encoded image to 12 bits through downsampling and color space transformation to obtain a YUV420 image;
[0043] S13: Compress a single pixel of the YUV420 image to within 8 bits through the JPEG compression algorithm to obtain the JPEG image bytecode.
[0044] First, the collected image is compressed to 24 bits through digital signal encoding to obtain an RGB-encoded image. While eliminating noise and attenuation in the image channel and improving anti-interference ability, it can also achieve image compression without distortion or within a preset distortion range. Then, a downsampling operation is continued on the RGB-encoded image to filter out color signals insensitive to the human eye, and the downsampled RGB image is transferred to 12 bits through color space transformation to obtain a YUV420 image, further reducing the occupied space of the image. Finally, the YUV420 image is further compressed through the JPEG compression algorithm, and the single pixel in the YUV420 image is compressed to within 8 bits to obtain JPEG image bytecode, thereby achieving the purpose of encoding the collected original image.
[0045] In the YUV420 image, Y represents luminance, U represents hue, and V represents saturation. Its function is to describe the image color and saturation and is used to specify the color of pixels. 420 means that every four Ys share a set of UV components. Therefore, the YUV image dumped from the RGB-encoded image occupies less space. And when using YUV420 for color space transformation, human perception ability is considered, and the main signals perceived by the human eye in the image are retained to compress half of the space, accelerating transmission and processing. Therefore, it can be ensured that the dumped YUV420 image is closer to the original image perceived by the human eye, avoiding excessive image distortion and affecting subsequent inference results.
[0046] As Figure 3 shown, in the step of compressing the YUV420 image through the JPEG compression algorithm, the following steps are further included
[0047] S131: Decompose the image signal into multiple cosine signals through discrete cosine transform;
[0048] S132: Transform the data order within the array of multiple cosine signals through Zigzag scanning;
[0049] S133: Perform chrominance quantization on the high-frequency and low-frequency data within the array through a JPEG quantizer;
[0050] S134: Re-compress the quantization result through the Huffman compression algorithm.
[0051] The image is divided into N*N pixel blocks through discrete cosine transform decomposition DCT (Discrete Cosine Transform), and then discrete cosine transform is performed on each N*N pixel block one by one. That is, multiple cosine signals are frequency-divided from the YUV420 image to filter out high-frequency image information insensitive to the human eye and continue to compress space.
[0052] Since adjacent points in the multi-branch cosine signal array are also adjacent in the image, transforming the data order within the multi-branch cosine signal array through Zigzag scanning facilitates subsequent sequential processing of quantization modulation. After the DCT image transformation is completed, operations such as screening the data from high-frequency to low-frequency domain intensity signals are required. To facilitate algorithm operations, the signal is converted into a one-dimensional sequence data. When the image is written row by row from left to right into the one-dimensional data table, it will cause high-frequency and low-frequency domain signals to be mixed in different parts of the entire table, making it inconvenient to perform operations such as Huffman coding. Therefore, it is necessary to transform the data order within the signal array through Zigzag scanning.
[0053] To prevent the image from being overly compressed and affecting image quality, here, a JPEG quantizer is used to perform chrominance quantization on the high-frequency and low-frequency data within the multi-branch cosine signal array to ensure image quality. According to the principle that the human eye is more sensitive to luminance signals than chrominance signals, different quantization tables are used for the luminance component and the color difference component of the image, namely the luminance quantization table and the color difference quantization table. The elements of the quantization table are the quantization intervals. By modulating the signal with the quantization table, the floating-point high-frequency and low-frequency data are converted into integer data for convenient computer storage, and the amount of high-frequency information to be filtered and the size of the image to be retained are determined according to the compression ratio. For example, for a 3840x2160 RGB image, a compression ratio of 95% will produce an image of 2.5MB in size, while a compression ratio of 75% will produce an image of 400KB. The small changes in the compression ratio are not sensitive to the human eye to compress space.
[0054] Through the Huffman compression algorithm, different coding lengths are given to the occurrence frequencies of different characters in the quantization result, so that the average coding length of the characters is the shortest, thereby achieving the purpose of further compressing space. For example, when a computer stores data, it actually stores a bunch of 0s and 1s (binary). Each individual character requires a combination of multiple 0s and 1s to represent, which causes the computer to require a relatively large amount of space to store data. The Huffman compression algorithm uses special values to represent these combinations of 0s and 1s. For example, 0 represents A; 1111 represents B, etc. The more frequently repeated bytes, the shorter the Huffman coding length, thus achieving the purpose of compressing the stored data.
[0055] As Figure 4 shown, in the steps of accelerating image decoding, it further includes
[0056] S21: Read the JPEG image byte code according to the JPEG protocol standard and adjust the byte code order according to the byte code storage address;
[0057] S22: Construct a Huffman tree and filter out the high-frequency components in the byte code with the adjusted order;
[0058] S23: Perform data conversion on the components after high-frequency filtering through inverse quantization;
[0059] S24: Perform an inverse transform on the data after inverse quantization conversion through reverse zigzag scanning to obtain the bitmap information of the image;
[0060] S25: Re-transform the bitmap signal into an RGB image through RGB color space transformation.
[0061] Read the image width and height information of the 164th byte of the JPEG image byte header according to the JPEG protocol standard to determine the size of the image, so as to facilitate subsequent decoding of the JPEG image byte code. Obtain the image information corresponding to each pixel block according to the length of the bytes of the image information in memory, and convert the order of the bytes and the storage address order. In order to conform to the JPEG decoding algorithm habit, the big-endian mode (that is, store the high bits of the bytes starting from the low address) is used to convert the byte order and the storage address order.
[0062] In order to further compress the bytes after the big-endian conversion and improve the speed of image decoding and subsequent inference conversion, a Huffman tree is constructed and the three YUV channels of the image are separated. Each channel is decomposed into a DC component and multiple cosine components through DCT (Discrete Cosine Transform), that is, each function term of the trigonometric function after Fourier transform. Since the human eye is not sensitive to high-frequency components, the high-frequency components in the cosine can be filtered out, so as to achieve the purpose of information compression. In this stage, the YUV image is sliced into pixel blocks, and the pixel blocks are converted into components, that is, floating-point coefficients representing the frequencies of the cosine components.
[0063] When performing inverse quantization data conversion, according to the luminance and chrominance quantization tables made according to the JPEG image compression standard, the frequencies (floating-point numbers) of the cosine components obtained by the above DCT image transformation are converted into integers (that is, perform compression of the data structure), and a larger quantization interval is used to indicate the chrominance and luminance intervals concentrated in the human eye-sensitive areas, and a smaller quantization interval is used to indicate the chrominance and luminance intervals concentrated in the human eye-insensitive areas, so as to achieve a two-weight quantization compression space.
[0064] When encoding and compressing the original image, zigzag scanning is used. Therefore, when decoding the image, reverse zigzag scanning is used to perform an inverse transform on the data. For example, zigzag scanning is performed in a "zigzag" manner from the lower left to the upper right of the image, and the sequential coordinates are (0,0) → (1,0) → (0,1) → (0,2) → (1,1) → (2,0) → (3,0) →.... Reverse zigzag scanning is a positive and inverse process of sorting. After the scanning is completed, the image bitmap signal is obtained, and finally the bitmap signal is re-transformed into an RGB image through RGB color space transformation.
[0065] As Figure 5 shown, in the steps of accelerating model inference, it further includes
[0066] S31: Configure the ONNX model based on the h5 file format and using the Keras library to perform the training of the deep learning model;
[0067] S32: Merge the network layer of the ONNX model and the parameter layer of the RGB image respectively through operator merging;
[0068] S33: Translate the operators in the ONNX model into CUDA kernel functions and run the CUDA kernel functions to obtain the inference result of a one-dimensional tensor.
[0069] The h5 file is the fifth generation version of the Hierarchical Data Format (HDF5), which is a format for storing data. And for storing a large amount of data, the h5 file has great advantages, with high data storage efficiency. Use the Keras library to perform the training of the deep learning model to configure the ONNX model, which is an exchangeable data format in the field of deep learning, to prepare for subsequent compilation and inference. At the same time, it can also filter out the redundant information in the H5 format, retain the computational graph of the model, and translate the H5 file format to the ONNX format by freezing the fixed model (the "frozen" layer means that this layer does not participate in network training, that is, the parameters of this layer will not be updated. "Fixed" can also be called compilation).
[0070] When merging operators for the network layer of the ONNX model, multiple consecutive network layers are merged on the vertical network layer according to the execution granularity of the hardware. For example, the convolutional layer CONV, the normalization layer BN, and the activation layer RELU can be integrated into the CBR operation. Through horizontal tensor integration, tensors with the same or similar sizes in the RGB image are fused, that is, layers with the same or similar input tensor sizes or performing the same operation are fused into one. For example, if there are three CBR layers horizontally, they can be completely integrated into the same CBR layer. By performing vertical network layer merging and horizontal tensor integration, the model architecture of the ONNX model and the computational amount when the model processes images are reduced, the operation steps are shortened, and time is saved.
[0071] Since CUDA instructions can write programs for high-speed large-batch calculations within the GPU, translating the operators in the ONNX model into CUDA kernel functions helps to achieve the acceleration of model inference. After the translation is completed, the inference result of the image is obtained by running the CUDA kernel functions. Since the complex large-batch tensor operations in ONNX are compiled into CUDA instructions, the obtained inference result is a one-dimensional ultra-long tensor.
[0072] As Figure 6As shown, in the step of accelerating the conversion of inference results, it further includes
[0073] S41: Read the inference results, output a one-dimensional floating-point number sequence from the inference results and store it as CHW data;
[0074] S42: Convert the CHW data into the RGB channel values of the corresponding pixels;
[0075] S43: Combine the multi-channel tensors into a single-channel mask image through the argmax function;
[0076] S44: Create an empty image with all pixel values being 0 and write each mask pixel value in the mask image into the empty image;
[0077] S45: Access the mask pixel values written into the empty image in a multi-core parallel manner to obtain the pixel coordinates of the corresponding mask pixel values;
[0078] S46: Combine the pixel coordinates and the RGB channel values of the corresponding pixels to obtain the mask image of the corresponding pixels.
[0079] When converting the inference results into an image, first output a one-dimensional floating-point number sequence from the inference results and store it as CHW (Channel, Height, Width corresponding to channel, width, and height) data. Obtain the RGB channel values of the corresponding pixels by converting the CHW data. The obtained RGB channel values are multi-channel tensors. Combine the multi-channel tensors into a single-channel mask image through the argmax function. Then, create an empty image with all pixel values being 0 and write each mask pixel value in the mask image into the empty image. Thus, when accessing the mask pixel values in the empty image, the pixel coordinates of the corresponding mask pixel values can be obtained. Among them, when accessing the mask pixel values in the empty image, a multi-core parallel manner is adopted to achieve the acceleration of converting the inference results into an image. When all the mask pixel values in the mask image are written into the empty image and all the mask pixel values are accessed, the coordinates of each mask pixel value in the mask image can be obtained. Finally, combine the pixel coordinates and the RGB channel values of the corresponding pixels to obtain the mask image of the corresponding pixels.
[0080] The following is the experimental comparison data table 1 of the method of the present invention and the existing methods for image encoding, decoding, inference, and transcoding:
[0081] Table 1
[0082]
[0083] Through experiments with a large number of images, it can be seen that when encoding JPEG large images with a resolution of 3840x2160, the encoding speed can be accelerated from 150 ms to 60 ms. When decoding JPEG large images with a resolution of 9120x4060, the decoding speed can be accelerated from 360 ms to 140 ms. When inferring RGB images with a resolution of 640x352, the inference speed can be accelerated from 200 ms to 80 ms, and the splicing of the one-dimensional sequence of the inference result into a two-dimensional image can be accelerated from 400 ms to 15 ms. Overall, the two-stage model inference of about 1400 ms is accelerated to 150 ms, and the image inference speed of the model is increased by about 90%.
[0084]
Image Inference Model Acceleration System
[0085] An embodiment of the present invention further provides an image inference model acceleration system, including:
[0086] An image encoding acceleration module that compresses and converts the format of the collected image, converts the RGB image into a YUV420 image, and compresses it into JPEG image bytecode;
[0087] An image decoding acceleration module that constructs a Huffman tree based on the JPEG image bytecode and performs Huffman tree encoding for transformation to re-decode the RGB image;
[0088] A model inference acceleration module that configures the ONNX model and merges the network layer of the ONNX model and the parameter layer of the RGB image respectively, and obtains the inference result of the image through translation of the merged result;
[0089] An inference result conversion acceleration module that performs format conversion on the inference result to obtain a single-channel mask image, creates an empty image, writes the mask pixel values in the single-channel mask image into the empty image, and obtains the mask image through parallel access.
[0090] It can be understood that this image inference model acceleration system encapsulates the image inference model acceleration method in the above method embodiment in the form of compiled code in each different module to execute the corresponding functions above. This detection system needs to utilize the model in the specific steps mentioned in the above method embodiment, and the specific method will not be elaborated here.
[0091]
Electronic Device
[0092] An embodiment of the present invention further provides an electronic device 700, as Figure 7 shown, which is a schematic structural diagram of the electronic device 700 provided by the embodiment of the present invention, including:
[0093] A processor 71, a memory 72, and a bus 73; the memory 72 is used to store execution instructions, including an internal memory 721 and an external memory 722; here, the internal memory 721 is also called the main memory, which is used to temporarily store the operation data in the processor 71 and the data exchanged with the external memory 722 such as a hard disk. The processor 71 exchanges data with the external memory 722 through the internal memory 721. When the electronic device 700 runs, the processor 71 communicates with the memory 72 through the bus 73, enabling the processor 71 to execute the following instructions:
[0094] Compress and perform format conversion on the acquired image to convert the RGB image into JPEG image bytecode;
[0095] Construct a Huffman tree based on the JPEG image bytecode and perform Huffman tree encoding for transformation to re-decode the RGB image;
[0096] Configure the ONNX model and merge the network layer of the ONNX model and the parameter layer of the RGB image respectively, and obtain the inference result of the image by translating the merged result;
[0097] Perform format conversion on the inference result to obtain a single-channel mask image, create an empty image, write the mask pixel values in the single-channel mask image into the empty image, and obtain the mask image through parallel access.
[0098] Achieve the accelerated inference effect of the image inference model by optimizing the four processes of encoding, decoding, inference, and inference result transcoding respectively. Among them, during encoding, compress the image and convert the image into a YUV420 image to facilitate color space conversion in the GPU, improve the image encoding speed, and then compress the YUV420 image into JPEG image bytecode to facilitate constructing a Huffman tree based on the JPEG image bytecode and performing reverse scanning on the Huffman encoding in the GPU during decoding to achieve accelerated image decoding. During model inference, configure the ONNX model and optimize the model to reduce the model architecture of the model and the computational amount when the model processes images, thereby achieving accelerated model inference. Finally, create an empty image and obtain the mask image by accessing the mask pixel values written into the empty image in a parallel manner, thereby achieving the acceleration of the inference result to image.
[0099] It can be understood that the electronic device 700 can be integrated as an independent third-party peripheral decoder.
[0100] The embodiment of the present invention also provides a computer-readable storage medium. The computer-readable storage medium stores one or more programs, and one or more of the programs can be executed by one or more processors to implement the above acceleration method.
[0101] The preferred embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited thereto. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solution of the present invention, including combinations of various specific technical features in any suitable manner. To avoid unnecessary repetition, the present invention will not separately describe various possible combination methods. However, these simple modifications and combinations should also be regarded as the content disclosed by the present invention and fall within the protection scope of the present invention.
[0102] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. An image inference model acceleration method, characterized in that, The method includes Image encoding acceleration, compressing and format-converting the acquired image to convert the RGB image into a YUV420 image and compressing it into JPEG image bytecode; Image decoding acceleration, constructing a Huffman tree based on the JPEG image bytecode and performing reverse scanning of the Huffman encoding in the GPU to re-decode the RGB image; Model inference acceleration, configuring the ONNX model and merging the network layer of the ONNX model and the parameter layer of the RGB image respectively, and obtaining the inference result of the image by translating the merged result; Inference result conversion acceleration, performing format conversion on the inference result to obtain a single-channel mask image, creating an empty image and writing the mask pixel values in the single-channel mask image into the empty image and obtaining the mask image through parallel access; In the step of image encoding acceleration, it further includes Collecting an image and compressing and sampling it to 24bit through digital signal encoding to obtain an RGB encoded image; Compressing and dumping the RGB encoded image to 12bit through downsampling and color space transformation and obtaining a YUV420 image; Compressing a single pixel of the YUV420 image to within 8bit through the JPEG compression algorithm and obtaining JPEG image bytecode, where the image encoding is implemented based on an embedded GPU, and converting the RGB color space of the image into the YUV420 color space in the GPU; In the step of inference result conversion acceleration, it further includes Reading the inference result, outputting a one-dimensional floating-point number sequence from the inference result and storing it as CHW data; Converting the CHW data into the RGB channel values of the corresponding pixels; Merging the multi-channel tensors into a single-channel mask image through the argmax function; Creating an empty image with all pixel values being 0 and writing each mask pixel value in the mask image into the empty image; Accessing the mask pixel values written into the empty image through multi-core parallelism to obtain the pixel coordinates of the corresponding mask pixel values; Combining the pixel coordinates and the RGB channel values of the corresponding pixels to obtain the mask image of the corresponding pixels.
2. The image inference model acceleration method according to claim 1, wherein: In the step of compressing the YUV420 image by JPEG, it further includes Decomposing the image signal into multiple cosine signals by discrete cosine transform; Transforming the data order in the multi-branch cosine signal array through Zigzag scanning; Performing chrominance quantization on the high and low frequency data in the array through the JPEG quantizer; Performing re-compression on the quantization result through the Huffman compression algorithm.
3. The method for accelerating an image inference model according to claim 1, wherein: In the step of image decoding acceleration, it further includes Reading the JPEG image bytecode according to the JPEG protocol standard and adjusting the bytecode order according to the bytecode storage address; Constructing a Huffman tree and filtering out the high-frequency components in the bytecode with the adjusted order; Performing data conversion on the components after high-frequency filtering through inverse quantization; Performing inverse transformation on the data after inverse quantization conversion through reverse zigzag scanning to obtain the bitmap information of the image.
4. The method for accelerating an image inference model according to claim 1, wherein: In the step of model inference acceleration, it further includes Configuring the ONNX model based on the h5 file format and using the Keras library to perform the training of the deep learning model; Merge the network layers of the ONNX model and the parameter layers of the RGB image respectively by means of operator merging; Translate the operators in the ONNX model into CUDA kernel functions and run the CUDA kernel functions to obtain the inference result of a one-dimensional tensor.
5. The method for accelerating an image inference model according to claim 4, wherein: When merging the ONNX model and the RGB image by means of operator merging, multiple consecutive network layers are merged through vertical network layer merging, and tensors with the same or similar sizes in the RGB image are fused through horizontal tensor integration.
6. An image inference model acceleration system, characterized in that The image inference model acceleration system includes: An image encoding acceleration module that compresses and converts the format of the acquired image to convert the RGB image into a YUV420 image and compresses it into JPEG image bytecode; An image decoding acceleration module that constructs a Huffman tree based on the JPEG image bytecode and performs reverse scanning of the Huffman encoding in the GPU to re-decode the RGB image; A model inference acceleration module that configures the ONNX model and merges the network layers of the ONNX model and the parameter layers of the RGB image respectively, and obtains the inference result of the image by translating the merged result; An inference result conversion acceleration module that performs format conversion on the inference result to obtain a single-channel mask image, creates an empty image, writes the mask pixel values in the single-channel mask image into the empty image, and obtains the mask image through parallel access; In the configuration of the image encoding acceleration module, it further includes Acquire an image and compress and sample it to 24bit through digital signal encoding to obtain an RGB encoded image; Compress and dump the RGB encoded image to 12bit through downsampling and color space transformation to obtain a YUV420 image; Compress a single pixel of the YUV420 image to within 8bit through the JPEG compression algorithm to obtain JPEG image bytecode. Among them, the image encoding is implemented based on an embedded GPU, and the RGB color space of the image is converted into the YUV420 color space in the GPU; In the configuration of the inference result conversion acceleration module, it further includes Read the inference result, output a one-dimensional floating-point number sequence from the inference result and store it as CHW data; Convert the CHW data into the RGB channel values of the corresponding pixels; Merge the multi-channel tensors into a single-channel mask image through the argmax function; Create an empty image with all pixel values being 0 and write each mask pixel value in the mask image into the empty image; Access the mask pixel values written into the empty image in a multi-core parallel manner to obtain the pixel coordinates of the corresponding mask pixel values; Combine the pixel coordinates and the RGB channel values of the corresponding pixels to obtain the mask image of the corresponding pixels.
7. An electronic device, characterized in that, The electronic device includes: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the image inference model acceleration method according to any one of claims 1 to 5 is executed.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the image inference model acceleration method according to any one of claims 1-5.
Citation Information
Patent Citations
Image processing system and method
CN102238376A
TensorRT-based pedestrian re-identification method and device
CN113033337A