A neural network image preprocessing method matching hardware computing power
By performing computational analysis and image preprocessing and post-processing of neural network models, the problem of improving inference efficiency and accuracy on hardware devices with limited computing power is solved, and more efficient and accurate neural network inference is achieved.
Patent Information
- Application Number
- CN202111511504.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-06
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-12-06
AI Technical Summary
When running neural network models on hardware devices with limited computing power, how to make full use of hardware computing power to improve inference efficiency and improve the inference accuracy of neural network models on images while ensuring inference speed.
By performing computational analysis of the neural network model, the parameters of the neural network model matching the hardware computing power are determined, and image preprocessing and inference intermediate results are performed to make full use of the hardware computing power and effective information on the original image.
It improves the inference efficiency of hardware devices and the inference accuracy of neural network models, and is suitable for various neural network models such as object detection and image segmentation.
Smart Images

Figure CN114186670B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of computer vision, neural network and image processing, and in particular to a neural network image preprocessing method that matches hardware computing power. Background Art
[0002] With the popularization of artificial intelligence applications, the expansion of neural network models from the cloud to the edge of embedded hardware is not only a development trend of technology, but also an urgent need for industrial AI empowerment. With the rapid development of deep learning, neural network models have become more and more complex, and the demand for computing resources and storage space has become higher and higher, and it has become more and more difficult to run on the hardware side. In order to achieve efficient operation of neural network models on the hardware side, existing solutions include: compressing the memory space occupied by neural network models through model compression methods such as pruning, quantization, low-rank decomposition and knowledge distillation, thereby improving the inference speed of neural network models during operation; deploying neural network models on embedded hardware with stronger parallel processing capabilities, more flexible configurability and lower power consumption, such as general computing platforms such as GPUs and FPGAs and dedicated chips for specific AI scenarios.
[0003] The current neural network model technology in areas such as target detection and image segmentation is becoming more and more mature, but in actual application scenarios, there are still many problems and challenges when running neural network models on hardware devices with limited computing power. For example, under the condition of a given neural network model, how to make full use of the hardware computing power to improve the reasoning efficiency of the hardware device; how to improve the reasoning accuracy of the neural network model for images while ensuring the reasoning speed of the neural network model. In addition, in order to meet the computing power constraints, we often perform image compression operations on the original images. However, image compression operations will lose effective features on the image, which may affect the reasoning accuracy of the neural network model.
[0004] Therefore, how to use hardware computing power to improve the reasoning efficiency of hardware devices, while improving the reasoning accuracy of neural network models for images, and ensuring that target information can be quickly acquired from images through neural network models, has become a key issue in current research. Summary of the invention
[0005] In view of the above problems, the present invention provides a neural network image preprocessing method that matches hardware computing power and solves at least some of the above technical problems. Under the conditions of a given hardware device with limited computing power and a given neural network model, the neural network model parameters that match the hardware computing power are determined by performing a computational analysis on the neural network model, thereby making full use of the hardware computing power and improving the reasoning efficiency of the hardware device. Through the neural network original image preprocessing and the corresponding neural network reasoning intermediate result post-processing process, the effective information on the original image is fully utilized to improve the reasoning accuracy of the neural network model.
[0006] An embodiment of the present invention provides a neural network image preprocessing method for matching hardware computing power, comprising:
[0007] S1. Analyze the computational complexity of the neural network model and determine the optimal input resolution and batch size parameters of the neural network model that match the hardware computing power;
[0008] S2, preprocessing the original image according to the optimal input resolution and Batchsize parameter of the neural network model, so that the resolution of the preprocessed original image is consistent with the optimal input resolution of the neural network model;
[0009] S3, using the neural network model to infer the preprocessed original image to obtain inference intermediate result data;
[0010] S4. Perform post-processing operations on the intermediate inference result data to obtain an image with a final inference result.
[0011] Furthermore, the constraint condition of the hardware computing power matched by the neural network model in S1 is expressed as:
[0012] FLOPS≥Batchsize×FLOPs(H,W) (1)
[0013] Among them, FLOPS represents the total computing power of the hardware matched by the neural network; Batchsize represents the number of images inferred by the neural network model; H and W represent the height and width of the original image respectively; FLOPs(H,W) represents the computing power of the neural network model when the resolution is H*W and Batchsize is 1.
[0014] Further, in S1, the neural network model includes a convolutional layer and a fully connected layer;
[0015] The floating point operation calculation amount of the convolutional layer of the neural network model is expressed as:
[0016] (2×C i ×K 2 -1)×H o ×W o ×C o (2)
[0017] Among them, C i Represents the dimension value of the convolution layer input channel; C o Represents the dimension value of the output channel of the convolution layer; K represents the size of the convolution kernel; H o , W o Respectively represent the height and width of the output features of the convolutional layer;
[0018] The floating point operation amount of the fully connected layer of the neural network model is expressed as:
[0019] (2×I-1)×O (3)
[0020] Among them, I represents the number of input neurons in the fully connected layer, that is, the product of each dimension of the input features; O represents the number of output neurons in the fully connected layer, that is, the product of each dimension of the output features.
[0021] Further, the S2 includes:
[0022] Performing image compression processing on the original image: compressing the original image in equal proportion so that its resolution is consistent with the optimal input resolution of the neural network model, thereby obtaining a preprocessed image A;
[0023] The original image is cropped and scaled: a target area image is cropped from the original image, and the target area image is enlarged or reduced so that its resolution is consistent with the optimal input resolution of the neural network model, thereby obtaining a preprocessed image B.
[0024] Furthermore, the S3 specifically includes:
[0025] Performing a first reasoning on the preprocessed image A using the neural network model to obtain reasoning intermediate result data corresponding to the preprocessed image A; the first reasoning includes performing image feature extraction and target information prediction on the preprocessed image A;
[0026] The neural network model is used to perform a second reasoning on the preprocessed image B to obtain reasoning intermediate result data corresponding to the preprocessed image B; the second reasoning includes image feature extraction and target information prediction for the preprocessed image B.
[0027] Furthermore, the S4 specifically includes:
[0028] S41, post-processing the inference intermediate result data corresponding to the pre-processed image A to obtain the pre-processed image A with the post-processing analysis result;
[0029] S42, post-processing the inference intermediate result data corresponding to the pre-processed image B to obtain the pre-processed image B with the post-processing analysis result;
[0030] S43, integrating the pre-processed image A with the post-processing analysis result and the pre-processed image B with the post-processing analysis result to obtain an image with the final reasoning result.
[0031] Furthermore, the S41 specifically includes:
[0032] S411, performing analysis processing on the inference intermediate result data corresponding to the preprocessed image A to obtain the corresponding preprocessed image A with post-processing analysis results;
[0033] S412: Enlarge the pre-processed image A with the post-processing analysis result to make its size consistent with that of the original image.
[0034] Furthermore, the S42 specifically includes:
[0035] S421, performing analysis processing on the inference intermediate result data corresponding to the preprocessed image B to obtain the corresponding preprocessed image B with post-processing analysis results;
[0036] S422, enlarging or reducing the pre-processed image B with the post-processing analysis result so that its size is consistent with the size of the target area image;
[0037] S423 . Based on S422 , the pre-processed image B with the post-processing analysis result is integrated into the target area position of the original image.
[0038] Compared with the prior art, the neural network image preprocessing method for matching hardware computing power described in the present invention has the following beneficial effects:
[0039] 1) The present invention determines the neural network model parameters that match the hardware computing power by analyzing the computational load of the neural network model, thereby making full use of the hardware computing power and improving the reasoning efficiency of the hardware device;
[0040] 2) The present invention fully utilizes the effective information on the original image through the neural network original image preprocessing and the corresponding neural network reasoning intermediate result postprocessing process, and effectively improves the reasoning accuracy of the neural network model;
[0041] 3) The present invention is applicable to various neural network models such as target detection and image segmentation, and can be implemented on different hardware.
[0042] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings.
[0043] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0045] Figure 1 A flowchart of a neural network image preprocessing method for matching hardware computing power provided in an embodiment of the present invention.
[0046] Figure 2 The neural network raw image preprocessing process and example diagram provided for the embodiment of the present invention.
[0047] Figure 3 The present invention provides a neural network reasoning intermediate result post-processing process and example diagram.
[0048] Figure 4 The scaling and position regression process and example diagram of the inference result of the cropped and scaled image provided by the embodiment of the present invention.
[0049] Figure 5 An example diagram comparing the inference effects of the neural network model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0050] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0051] This embodiment is implemented on the FPGA hardware of Xilinx's Zynq UltraScale+MPSoC. The neural network model used to verify the design in the embodiment is a YOLOv3 target detection model that has undergone model compression. The scenario in which the embodiment is applied is an open road scenario for autonomous driving.
[0052] See also Figure 1 As shown, the embodiment of the present invention provides a neural network image preprocessing method matching hardware computing power, which specifically includes the following steps:
[0053] S1. Analyze the computational complexity of the neural network model and determine the optimal input resolution and batch size parameters of the neural network model that match the hardware computing power;
[0054] S2, preprocessing the original image according to the optimal input resolution and Batchsize parameter of the neural network model, so that the resolution of the preprocessed original image is consistent with the optimal input resolution of the neural network model;
[0055] S3, using the neural network model to infer the preprocessed original image to obtain inference intermediate result data;
[0056] S4. Perform post-processing operations on the intermediate inference result data to obtain an image with a final inference result.
[0057] The following is a detailed description of each of the above steps.
[0058] In the above step S1, computing power is a measure of hardware performance, and the amount of calculation is a measure of the complexity of the neural network model. The present invention needs to analyze the amount of calculation of the neural network model, so as to determine the optimal input resolution and batchsize parameters of the neural network model that matches the hardware computing power. Among them, the constraint condition of the hardware computing power matched by the neural network model is expressed as:
[0059] FLOPS≥Batchsize×FLOPs(H,W) (1)
[0060] Among them, FLOPS represents the total computing power of the hardware matched by the neural network, which is an indicator to measure hardware performance; Batchsize represents the number of images inferred by the neural network model; H and W represent the height and width of the original image respectively; FLOPs(H,W) represents the computing power of the neural network model when the resolution is H*W and Batchsize is 1.
[0061] The neural network model includes convolutional layers and fully connected layers;
[0062] The floating point operation calculation amount of the convolution layer of the neural network model is expressed as:
[0063] (2×C i ×K 2 -1)×H o ×W o ×C o (2)
[0064] Among them, C i Represents the dimension value of the convolution layer input channel; C o Represents the dimension value of the output channel of the convolution layer; K represents the size of the convolution kernel; H o , W o They represent the height and width of the output features of the convolutional layer respectively; i represents input; and o represents output.
[0065] The floating-point calculation amount of the fully connected layer of the neural network model is expressed as:
[0066] (2×I-1)×O (3)
[0067] Among them, I represents the number of input neurons in the fully connected layer, that is, the product of each dimension of the input features; O represents the number of output neurons in the fully connected layer, that is, the product of each dimension of the output features.
[0068] In order to further implement the above technical solution, the FPGA hardware of Xilinx's Zynq UltraScale+MPSoC selected in this embodiment can support neural network models with a computing power of less than 60BFLOPS. In this embodiment, the computational amount analysis method of the neural network model can be used to calculate that the computational amount of the YOLOv3 model compressed by the model used for verification in this embodiment is 128.54BFLOPs under the parameters of Batchsize of 1 and input resolution of 1080*1440; by analyzing the computational amount of the neural network model, the computational amount and reasoning time of the YOLOv3 model compressed by the model used for verification in this embodiment under different Batchsize parameters and input resolution parameters can be calculated, as shown in Table 1:
[0069] Table 1: The computational complexity and inference time of the YOLOv3 model of this embodiment under different parameters on the FPGA
[0070]
[0071] In Table 1, H and W represent the height and width of the original image, respectively. The YOLOv3 model can infer that H and W of the original image must be integer multiples of 32. BFLOPs is the unit of computing power. FPGA is the FPGA hardware of Xilinx's Zynq UltraScale+ MPSoC.
[0072] As can be seen from Table 1, under the condition of matching the FPGA hardware computing power, the computational amount of the YOLOv3 model corresponding to Batchsize 2 and input resolution 480*640 is 6.35 BFLOPs less than the computational amount of the YOLOv3 model corresponding to Batchsize 1 and input resolution 720*960. In other words, without increasing the amount of computation, the YOLOv3 model can be used to infer two relatively small resolution pre-processed images at the same time, and the inference time will also be less. Therefore, the input resolution of the YOLOv3 model that is finally determined to match the FPGA hardware computing power is 480*640 and the Batchsize is 2.
[0073] In the above step S2, the original image needs to be subjected to image compression preprocessing and image cropping and scaling preprocessing respectively; wherein the original image is subjected to image compression processing: the original image is proportionally compressed so that its resolution is consistent with the optimal input resolution of the neural network model, thereby obtaining a preprocessed image A (i.e., a compressed image);
[0074] The original image is cropped and scaled: the target area image is cropped from the original image, and the target area image is enlarged or reduced so that its resolution is consistent with the optimal input resolution of the neural network model, thereby obtaining the preprocessed image B (i.e., the cropped and scaled image).
[0075] In order to further implement the above technical solution, Figure 2 As shown, in this embodiment, after image compression, image cropping and image preprocessing operations of scaling, N original images with a resolution of 1080*1440 can obtain 2N images with a resolution of 480*640.
[0076] Among them, the original image is compressed:
[0077] Without changing the aspect ratio of the original image, the original image with a resolution of 1080*1440 is proportionally compressed to an image with a resolution of 480*640, which is consistent with the optimal resolution of the neural network model obtained in the above step S1, thereby obtaining a preprocessed image A (i.e., a compressed image).
[0078] Perform image cropping and scaling on the original image:
[0079] The center position and resolution of the cropped area are determined according to the area where the small targets in the original image are concentrated and the aspect ratio of the original image. If the small targets in the area need to be enlarged, the resolution value should be less than 480*640; an area is cropped from the original image according to the determined center position and resolution; the cropped area is scaled to an image with a resolution of 480*640, thereby obtaining a preprocessed image B (i.e., a cropped and scaled image).
[0080] In the above step S3, this embodiment uses the model-compressed YOLOv3 target detection model that has been successfully deployed on the FPGA hardware of Xilinx's Zynq UltraScale+MPSoC and can realize reasoning to perform reasoning on the above preprocessed image and obtain the reasoning intermediate result data. The neural network reasoning process is: the preprocessed image is subjected to the main network part of the neural network to extract the image features, and the extracted image features are predicted by the head network part of the neural network to obtain the reasoning intermediate result data containing classification and regression information. In this embodiment, the above preprocessed image A (i.e., compressed image) and preprocessed image B (i.e., cropped and scaled image) are subjected to reasoning including the image feature extraction and target information prediction process, and the reasoning intermediate result data corresponding to the preprocessed image A and the reasoning intermediate result data corresponding to the preprocessed image B containing the target detection box category and position information are obtained respectively.
[0081] In the above step S4, post-processing operations are performed on the image with the inference intermediate result, which specifically includes:
[0082] S41, post-processing the inference intermediate result data corresponding to the pre-processed image A to obtain the pre-processed image A with the post-processing analysis result;
[0083] S42, post-processing the inference intermediate result data corresponding to the pre-processed image B to obtain the pre-processed image B with the post-processing analysis result;
[0084] S43, integrating the pre-processed image A with the post-processing analysis result and the pre-processed image B with the post-processing analysis result to obtain an image with the final reasoning result.
[0085] In order to further implement the above technical solution, Figure 3 As shown, in this embodiment, 2N compressed images (i.e., preprocessed image A) and cropped scaled images (i.e., preprocessed image B) with a resolution of 480*640 are inferred by the YOLOv3 model. After corresponding post-processing operations, N target detection results on the original images with a resolution of 1080*1440 can be obtained.
[0086] Among them, S41 specifically includes:
[0087] S411, performing analysis processing on the inference intermediate result data corresponding to the preprocessed image A to obtain the corresponding preprocessed image A with post-processing analysis results;
[0088] S412, enlarging the pre-processed image A with the post-processing analysis result so that its size is consistent with that of the original image;
[0089] Specifically in this embodiment: the inference intermediate result data obtained by the YOLOv3 model inference compressed image is parsed into a target detection frame on the compressed image with a resolution of 480*640 through a non-maximum suppression operation, and the non-maximum suppression includes iteration, traversal and elimination of the detection frame; according to the compression ratio value during image preprocessing, the target detection frame is scaled to its corresponding size on the original image with a resolution of 1080*1440.
[0090] S42 specifically includes:
[0091] S421, performing analysis processing on the inference intermediate result data corresponding to the preprocessed image B to obtain the corresponding preprocessed image B with post-processing analysis results;
[0092] S422, enlarging or reducing the pre-processed image B with the post-processing analysis result so that its size is consistent with the size of the target area image;
[0093] S423 . Based on S422 , the pre-processed image B with the post-processing analysis result is integrated into the target area position of the original image.
[0094] Specifically in this embodiment: the data of the inference intermediate result obtained by the YOLOv3 model inference of the cropped and scaled image is parsed into a target detection frame on the cropped and scaled image with a resolution of 480*640 through a non-maximum suppression operation, and the non-maximum suppression includes iteration, traversal and elimination of the detection frame; the detection frame whose boundary is less than 3 pixels away from the image boundary in the parsed target detection frame is discarded, and the other detection frames are retained; according to the scaling value during image preprocessing, the target detection frame is scaled to the corresponding size on the cropped area of the original image with a resolution of 1080*1440; according to the relative position of the cropped area on the original image with a resolution of 1080*1440 during image preprocessing, the target detection frame is regressed to its corresponding position on the original image.
[0095] like Figure 4 As shown, the scaling and position regression operations of the inference results of the cropped and scaled image in this embodiment are as follows: first, the center point coordinates x, y and the length, width w, h of the detection frame are proportionally scaled, and then the center point coordinates of the detection frame are added with its offsets Δx and Δy in the image, and finally the information of the detection frame on the cropped image corresponding to the detection frame on the original image is obtained.
[0096] In the above step S43, it is necessary to integrate the result after processing in S412 and the result after processing in S423 to obtain an image with a final detection result.
[0097] In this embodiment, the following steps are specifically performed: the parsed detection frame information of the compressed image and the cropped and scaled image is merged into a set, and the redundant detection frames therein are filtered out by a non-maximum suppression operation to obtain an image with a final inference result. The non-maximum suppression includes iterating, traversing and eliminating the detection frames; specifically, Figure 5 As shown, this embodiment is based on the YOLOv3 target detection model after model compression on the FPGA hardware of Xilinx's Zynq UltraScale+ MPSoC, and compares the reasoning effect of directly compressing the original image to a resolution of 720*960 and using this method in an open road scenario of autonomous driving. Figure 5 (a) is the final inference result of directly compressing the original image to a resolution of 720*960. Figure 5 (b) is the final reasoning result of this method. Figure 5 It can be seen from the comparison that this method can detect more small targets and improve the reasoning accuracy of the neural network model.
[0098] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A neural network image preprocessing method for matching hardware computing power, characterized in that, it includes: S1. Analyze the computational complexity of the neural network model, and determine the optimal input resolution and Batchsize parameter of the neural network model that matches the hardware computing power; S2. Preprocess the original image according to the optimal input resolution and Batchsize parameter of the neural network model, so that the resolution of the preprocessed original image is consistent with the optimal input resolution of the neural network model; S3. Use the neural network model to perform inference on the preprocessed original image to obtain inference intermediate result data; S4. Perform post-processing operations on the inference intermediate result data to obtain an image with the final inference result.
2. The neural network image preprocessing method for matching hardware computing power according to claim 1, characterized in that, the constraint condition for the hardware computing power matched by the neural network model in S1 is expressed as: FLOPS≥Batchsize×FLOPs(H,W) (1) wherein, FLOPS represents the total computing power of the hardware matched by the neural network; Batchsize represents the number of images for neural network model inference; H and W respectively represent the height and width of the original image; FLOPs(H,W) represents the computational complexity of the neural network model when the resolution is H*W and Batchsize is 1.
3. The neural network image preprocessing method for matching hardware computing power according to claim 1, characterized in that: in S1, the neural network model includes a convolutional layer and a fully connected layer; the floating-point operation computational complexity of the convolutional layer of the neural network model is expressed as: (2×C i ×K 2 -1)×H o ×W o ×C o (2) Among them, C i Represents the dimension value of the convolution layer input channel; C o Represents the dimension value of the output channel of the convolution layer; K represents the size of the convolution kernel; H o , W o Respectively represent the height and width of the output features of the convolutional layer; the floating-point operation computational complexity of the fully connected layer of the neural network model is expressed as: (2×I - 1)×O (3) wherein, I represents the number of input neurons of the fully connected layer, that is, the product of each dimension of the input feature; O represents the number of output neurons of the fully connected layer, that is, the product of each dimension of the output feature.
4. The neural network image preprocessing method for matching hardware computing power according to claim 1, characterized in that: S2 includes: Perform image compression processing on the original image: perform proportional compression on the original image so that its resolution is consistent with the optimal input resolution of the neural network model, thereby obtaining preprocessing image A; Perform image cropping and scaling processing on the original image: crop the target region image from the original image, and enlarge or reduce the target region image so that its resolution is consistent with the optimal input resolution of the neural network model, thereby obtaining preprocessing image B.
5. The neural network image preprocessing method for matching hardware computing power according to claim 4, characterized in that: S3 specifically includes: Use the neural network model to perform the first inference on the preprocessing image A to obtain the inference intermediate result data corresponding to the preprocessing image A; the first inference includes image feature extraction and target information prediction on the preprocessing image A; The neural network model is used to perform a second reasoning on the preprocessed image B to obtain reasoning intermediate result data corresponding to the preprocessed image B; the second reasoning includes image feature extraction and target information prediction for the preprocessed image B.
6. A neural network image preprocessing method for matching hardware computing power as claimed in claim 5, Features: The S4 specifically includes: S41, post-processing the inference intermediate result data corresponding to the pre-processed image A to obtain the pre-processed image A with the post-processing analysis result; S42, post-processing the inference intermediate result data corresponding to the pre-processed image B to obtain the pre-processed image B with the post-processing analysis result; S43, integrating the pre-processed image A with the post-processing analysis result and the pre-processed image B with the post-processing analysis result to obtain an image with the final reasoning result.
7. A neural network image preprocessing method for matching hardware computing power as claimed in claim 6, Features: The S41 specifically includes: S411, performing analysis processing on the inference intermediate result data corresponding to the preprocessed image A to obtain the corresponding preprocessed image A with post-processing analysis results; S412: Enlarge or reduce the pre-processed image A with the post-processing analysis result to make its size consistent with that of the original image.
8. A neural network image preprocessing method for matching hardware computing power as claimed in claim 6, Features: The S42 specifically includes: S421, performing analysis processing on the inference intermediate result data corresponding to the preprocessed image B to obtain the corresponding preprocessed image B with post-processing analysis results; S422, enlarging or reducing the pre-processed image B with the post-processing analysis result so that its size is consistent with the size of the target area image; S423 . Based on S422 , the pre-processed image B with the post-processing analysis result is integrated into the target area position of the original image.
Citation Information
Patent Citations
A method for accelerating visual recognition of IOT portable device based on low-resolution compressed image
CN109086806A
Method for detecting small target of high-resolution image of any scale
CN111222474A