Embedded NPU video super-resolution reconstruction method, system and device
By deploying lightweight image super-resolution generators and discriminators on embedded NPU devices, combining the deep separation convolution and channel attention mechanisms, the problems of slow inference speed and reduced reconstruction accuracy during deep learning models deployed on embedded devices are solved, and efficient video super-resolution reconstruction is achieved.
Patent Information
- Application Number
- CN202510178410.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-18
AI Technical Summary
Existing deep learning models are constrained by computing resources and memory when deployed on embedded devices, resulting in slow inference speed and reduced rebuild accuracy.
An embedded NPU video super-resolution reconstruction method is proposed. Through a lightweight image super-resolution generator and discriminator, combined with a depth-separable convolution and channel attention mechanism, significantly reduce model parameters and calculation amount, and improve inference speed through hardware acceleration technology.
While reducing the calculation amount and parameter amount, maintaining high image quality significantly improves the inference speed and practicality of the model on embedded devices.
Smart Images

Figure CN120013764A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an embedded NPU video super-resolution reconstruction method, system and device. Background Art
[0002] Image super-resolution reconstruction technology is a processing method designed to improve image resolution and enhance image detail and quality. This technology generates a high-resolution image from a single or multiple low-resolution images by analyzing the information and structure contained in the image and learned prior knowledge. It has a wide range of applications in various fields such as surveillance video, medical imaging, satellite remote sensing, and video enhancement. Traditional image super-resolution reconstruction methods are mainly based on interpolation techniques such as nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation. Although these methods are simple to operate and only enlarge the image by increasing the pixel size, they often cannot effectively restore high-frequency details in the image, resulting in limitations in their reconstruction results.
[0003] With the rapid development of deep learning technology, significant progress has been made in the field of computer vision, particularly in deep learning-based image super-resolution reconstruction techniques. For example, the SRCNN model proposed by Dong et al. uses a simple three-layer convolutional neural network (CNN) to learn the mapping relationship between low-resolution images and high-resolution images. Subsequently, a variety of efficient super-resolution models such as RCAN, EDSR, and SRGAN have been derived. These models further enhance the network's learning ability and image reconstruction performance by increasing network depth and parameter count.
[0004] However, deploying these deep network models on embedded devices with limited computing resources and memory presents challenges. These limitations restrict the scale and complexity of deployed networks, and in practice, they can lead to issues such as slow inference speed and reduced reconstruction accuracy. Therefore, it is crucial to comprehensively balance model structure, inference speed, and reconstruction accuracy, employing lightweight and optimization strategies to adapt to the resource constraints of embedded devices. Summary of the Invention
[0005] To solve the above problems, the present invention aims to propose an embedded NPU video super-resolution reconstruction method, system and device for realizing the reconstruction of low-resolution video to high-resolution video, while reducing the amount of calculation and parameters while maintaining high image quality, and using hardware acceleration to improve the model inference speed.
[0006] To achieve the above object, the technical solution of the present invention is achieved as follows:
[0007] An embedded NPU video super-resolution reconstruction method, comprising:
[0008] S1. Data processing: Convert the original high-resolution image into a low-resolution image through a variety of combined degradation operations to simulate the image degradation process in real scenes for model training;
[0009] S2. Model Construction: A lightweight image super-resolution generator and discriminator are used to reconstruct low-resolution images into super-resolution images. The generator uses depthwise separable convolution and an effective channel attention mechanism, combined with residual modules and feature fusion technology, to significantly reduce model parameters and computational complexity while maintaining high image reconstruction quality.
[0010] S3. Model training: Use the image training dataset for training; use low-resolution images as the input of the generator network and the corresponding original high-resolution image samples as the expected output of the generator network, and use the backpropagation algorithm to train the generator network; use low-resolution images and corresponding high-resolution images as positive image pairs, and low-resolution images in the training sample dataset and the corresponding output images of the generator network as negative image pairs, and use the backpropagation algorithm to train the discriminator network;
[0011] S4, model quantization and compression: Convert the trained super-resolution model file into the intermediate ONNX format, and further convert the ONNX format model file into the RKNN format compatible with embedded devices;
[0012] S5. Deployment: Input a low-resolution video, read each frame of the video through the OpenCV library, perform necessary preprocessing on each frame, and input it into the quantized model. The model will infer the input low-resolution image and load the model weight file obtained during the training process. After processing the quantized model, it will output a super-resolution image. Finally, use the OpenCV library to reassemble the processed super-resolution image into a video, thereby completing the reconstruction of the super-resolution video.
[0013] S6. Acceleration: Use hardware acceleration to accelerate the inference process of the model deployed in step S5.
[0014] Furthermore, the data processing of step S1 includes: performing cropping, Gaussian blurring, defocusing and motion blurring simulation, and Gaussian noise addition on the original high-resolution image to generate a first reference image;
[0015] Performing JPEG compression on the first reference image and introducing image quality compression to obtain a second reference image;
[0016] The second reference image is down-sampled by bicubic interpolation and normalized to obtain a degraded low-resolution image.
[0017] Furthermore, the image reconstruction in step S2 includes: sending the low-resolution image to a shallow feature extraction module of an image super-resolution generation network to obtain a shallow feature map of the image;
[0018] The shallow feature map of the image is sent to the deep feature extraction module of the image super-resolution generation network to obtain the deep feature map of the image.
[0019] Furthermore, the deep feature extraction module includes a stack of a preset number of residual modules, each of which is composed of a cascade of 1x1 convolution, 3x3 convolution, sigmoid and an effective channel attention mechanism, and introduces a depth-wise separable convolution operation;
[0020] Sending the shallow feature map and the deep feature map of the image into a feature fusion module to obtain a fused feature map;
[0021] Sending the fused feature map to an upsampling module of an image super-resolution generation network to obtain a super-resolution image of the low-resolution image;
[0022] The super-resolution image and the high-resolution image are respectively input into an image super-resolution discrimination network to obtain a probability value of discriminating the super-resolution image as a real image;
[0023] The discriminant network is an image classification convolutional neural network that uses the LeakReLU activation function to prevent negative output necrosis;
[0024] The input image passes through the classification convolutional neural network and then passes through two linear layers to obtain the output result;
[0025] The loss functions of the generation network and the discriminant network are determined according to the probability values of the high-resolution image, the super-resolution image, and the output of the discriminator network.
[0026] Furthermore, the loss function of the discriminant network is:
[0027]
[0028] The generative network loss function consists of three parts:
[0029]
[0030] Among them L percep Using VGG loss:
[0031]
[0032] W and H represent the dimensions of the feature map in the VGG network;
[0033] i and j refer to the jth convolutional layer before the i-th max pooling layer;
[0034] The generated network adversarial loss function
[0035]
[0036] in Used to evaluate the generated super-resolution image G(x i ) and the original high-resolution image, where λ and η are weight coefficients.
[0037] Furthermore, the model training of step S3 includes: optimizing and fine-tuning the model using the prepared video training data samples;
[0038] First, obtain a batch of low-resolution image groups LR from the prepared training data Group , the batch size is set to 32, i.e., 32 low-resolution images are acquired at a time;
[0039] Secondly, the corresponding super-resolution image group SR is generated using the generative network Group ;
[0040] Next, the generated SR Group The original high-resolution image group HR in the training data Group The corresponding input is sent to the discriminant network;
[0041] At the same time, the loss function of the generator and the loss function of the discriminator are used to tune the model to complete the model training process.
[0042] Furthermore, the model quantization compression in step S4 further includes: in this conversion process, it is necessary to pre-set the model input size, mean, normalization parameters and whether to perform quantization configuration options;
[0043] Perform model quantization on the RKNN file, that is, convert floating-point numbers to fixed-point numbers so that it can run efficiently on embedded devices.
[0044] Furthermore, the acceleration of step S6 includes: utilizing a hardware image processing accelerator RGA of an embedded device to speed up image preprocessing;
[0045] Implement concurrent operations through thread pools, thereby accelerating the model's reasoning speed and improving the system's real-time performance;
[0046] Before embedded inference, RGA is used to divide large-size images into sub-blocks for super-resolution reconstruction, the sub-block splicing process is optimized, and overlapping area fusion technology is used to address the impact of splicing lines.
[0047] In order to achieve the above objectives, the present invention also discloses an embedded NPU video super-resolution reconstruction system, comprising:
[0048] Data processing module: used to convert the original high-resolution image into a low-resolution image through a variety of combined degradation operations for model training;
[0049] Model building module: Reconstruct low-resolution images into super-resolution images through a lightweight image super-resolution generator and discriminator;
[0050] Model training module: Use the image training dataset for training; use low-resolution images as the input of the generator network and the corresponding original high-resolution image samples as the expected output of the generator network, and use the back-propagation algorithm to train the generator network; use low-resolution images and corresponding high-resolution images as positive image pairs, and low-resolution images in the training sample dataset and the corresponding output images of the generator network as negative image pairs, and use the back-propagation algorithm to train the discriminator network;
[0051] Model quantization and compression module: used to convert the trained super-resolution model file into the intermediate ONNX format, and further convert the ONNX format model file into the RKNN format compatible with embedded devices;
[0052] Deployment module: Input low-resolution video, read each frame of the video through the OpenCV library, perform necessary preprocessing on each frame, and input it into the quantized model. The model will infer the input low-resolution image and load the model weight file obtained during the training process. After processing the quantized model, it will output a super-resolution image. Finally, the OpenCV library is used to reassemble the processed super-resolution image into a video, thus completing the reconstruction of the super-resolution video.
[0053] Acceleration module: Use hardware acceleration to accelerate the inference process of the model deployed in the module.
[0054] In order to achieve the above object, the present invention also discloses an embedded NPU video super-resolution reconstruction device, including a processor and a storage medium;
[0055] The storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the above method.
[0056] Beneficial Effects: This invention builds on existing image super-resolution reconstruction algorithms by replacing the original residual network with a lightweight residual structure, significantly reducing the number of model parameters and computational complexity. While ensuring image reconstruction quality, the model conversion and quantization process further compresses the model size, enhancing its portability and making it more suitable for deployment on resource-constrained embedded devices.
[0057] Furthermore, by integrating multiple hardware acceleration technologies, this invention fully exploits the potential of hardware, significantly improving the model's inference speed on embedded devices and thus enhancing its practicality. Verification on various test datasets demonstrates that the super-resolution reconstruction method of this invention not only possesses fast inference capabilities but is also highly portable, accelerating the development of deep learning super-resolution reconstruction in the embedded field. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The accompanying drawings, which constitute part of the present invention, are provided to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are provided to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:
[0059] Figure 1 This is a flowchart of the embedded NPU video super-resolution reconstruction method according to an embodiment of the present invention;
[0060] Figure 2 This is a structural diagram of an image super-resolution generator in the embedded NPU video super-resolution reconstruction method according to an embodiment of the present invention;
[0061] Figure 3 This is a flowchart of the model quantization compression in the embedded NPU video super-resolution reconstruction method according to an embodiment of the present invention;
[0062] Figure 4 This is a structural diagram of using a thread pool for model parallel operation in the embedded NPU video super-resolution reconstruction method according to an embodiment of the present invention;
[0063] Figure 5 This is a diagram of a cropping and splicing scheme in embedded reasoning in the embedded NPU video super-resolution reconstruction method according to an embodiment of the present invention;
[0064] Figure 6 This is a schematic diagram of the structure of the embedded NPU video super-resolution reconstruction system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0065] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0066] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0067] Example 1
[0068] See also Figure 1-5 : An embedded NPU video super-resolution reconstruction method, comprising:
[0069] S1. Data processing: Convert the original high-resolution image into a low-resolution image through a variety of combined degradation operations to simulate the image degradation process in real scenes for model training;
[0070] S2. Model Construction: A lightweight image super-resolution generator and discriminator are used to reconstruct low-resolution images into super-resolution images. The generator uses depthwise separable convolution and an effective channel attention mechanism, combined with residual modules and feature fusion technology, to significantly reduce model parameters and computational complexity while maintaining high image reconstruction quality.
[0071] S3. Model training: Use the image training dataset for training; use low-resolution images as the input of the generator network and the corresponding original high-resolution image samples as the expected output of the generator network, and use the backpropagation algorithm to train the generator network; use low-resolution images and corresponding high-resolution images as positive image pairs, and low-resolution images in the training sample dataset and the corresponding output images of the generator network as negative image pairs, and use the backpropagation algorithm to train the discriminator network;
[0072] S4, model quantization and compression: Convert the trained super-resolution model file into the intermediate ONNX format, and further convert the ONNX format model file into the RKNN format compatible with embedded devices;
[0073] S5. Deployment: Input a low-resolution video, read each frame of the video through the OpenCV library, perform necessary preprocessing on each frame, and input it into the quantized model. The model will infer the input low-resolution image and load the model weight file obtained during the training process. After processing the quantized model, it will output a super-resolution image. Finally, use the OpenCV library to reassemble the processed super-resolution image into a video, thereby completing the reconstruction of the super-resolution video.
[0074] S6. Acceleration: Use hardware acceleration to accelerate the inference process of the model deployed in step S5.
[0075] The combined degradation operation of this embodiment effectively simulates motion blur. The lightweight image super-resolution generator and discriminator effectively reduce model parameters and computational complexity, enhancing the network's representational capabilities while maintaining high image reconstruction quality. By utilizing image processing accelerators and thread pool technology, the inference speed on embedded devices is improved. This embodiment achieves efficient implementation of super-resolution algorithms and is suitable for real-time video processing in embedded environments.
[0076] In a specific example, the data processing in step S1 includes: performing cropping, Gaussian blurring, defocusing and motion blurring simulation, and Gaussian noise addition on the original high-resolution image to generate a first reference image;
[0077] Performing JPEG compression on the first reference image and introducing image quality compression to obtain a second reference image;
[0078] The second reference image is down-sampled by bicubic interpolation and normalized to obtain a degraded low-resolution image.
[0079] In a specific implementation, 88x88 sub-images are randomly cropped from each original high-resolution image. Gaussian blur is then applied to these sub-images, with the blur kernel sizes uniformly and randomly selected between 7x7, 9x9, …, and 21x21 to simulate the blur effect in real-world photography. Reflection padding is used to ensure spatial consistency in the blur output. Simulated defocus and motion blur effects are then added to the blurred sub-images, along with Gaussian noise, to generate the first reference image.
[0080] Secondly, the first reference image is compressed using JPEG image quality to obtain the second reference image; the JPEG quality factor is uniformly selected from [30, 95].
[0081] The second reference image is bicubicly downsampled by a factor of 2 to obtain a degraded low-resolution image.
[0082] Common interpolation functions are:
[0083] Then, data normalization is performed on the downsampled low-resolution image and the corresponding original high-resolution image, and the pixel values are scaled to the range of [0, 1] in preparation for subsequent model training.
[0084] Through the above operations, the original high-resolution image is converted into a low-resolution image with noise and blur similar to that in real life due to factors such as motion, thereby providing a dataset close to real scenes for model training.
[0085] In a specific example, the image reconstruction in step S2 includes: sending the low-resolution image to a shallow feature extraction module of an image super-resolution generation network to obtain a shallow feature map of the image;
[0086] The shallow feature map of the image is sent to the deep feature extraction module of the image super-resolution generation network to obtain the deep feature map of the image.
[0087] In the specific implementation, the image super-resolution generator is as follows Figure 2 The following parts are shown:
[0088] The input to the network is a 3-channel RGB low-resolution image. First, a 9x9 convolution layer and a PReLU activation function are used to extract the shallow feature map of the image. Then, these feature maps are sent to 10 EC-PA modules to gradually refine the deep features of the image. Subsequently, the deep feature map is processed by a 3x3 convolution layer, and the obtained feature map is spliced with the original shallow feature map to generate a fused feature map. The fused feature map then enters the PixelShuffle module for a 2x upsampling operation. Finally, the upsampled feature map is processed by a 9x9 deconvolution layer and a Tanh activation function, and the super-resolution reconstructed image is finally output;
[0089] The EC-PA module first separates the input shallow feature map xn-1 into two feature maps x'n-1 and x"n-1 through two 1x1 convolution layers respectively; x'n-1 passes through a 3x3 DSConv and a PAConv module respectively, which includes a 1x1 convolution and a Sigmoid activation function to obtain two feature maps, which are multiplied element by element and then passed through a 3x3 DSConv to obtain the feature map x'n; x"n-1 passes through a 3x3 DSConv and an ECAConv module respectively To further enhance the correlation between channels, the two feature maps obtained are concatenated in the channel dimension and fed into a 3x3 DSConv to obtain a feature map x'n; x'n and x'n are concatenated in the channel dimension and fed into a 1x1 convolution to obtain an enhanced feature map, and the number of channels is restored to the input channel number; the enhanced feature map is element-wise added to the input feature map to form the final output feature map xn; the ECAConv module implements attention calculation for each channel through global average pooling and 1D convolution, and the output feature map is element-wise multiplied with the input feature map;
[0090] The PixelShuffle upsampling module uses convolution and pixel shuffling techniques to upsample images. It receives an input of in_channels channels and generates an output feature map of in_channels*up_scale^2 channels. The 3x3 convolution kernel and 1 padding in the module ensure consistency between input and output dimensions. This is done to pre-increase the number of channels in preparation for pixel shuffling. Pixel shuffling converts the multi-channel feature map into a higher-resolution image. The up_scale parameter determines the upsampling ratio; for example, when up_scale=2, the image size is doubled. After upsampling, the feature map is activated using the PReLU function to produce the final output.
[0091] Table 1 Parameters, computational complexity, and comparison of data in the Set test dataset
[0092]
[0093] As shown in Table 1, the proposed method has lower parameter counts and computational complexity than other existing image super-resolution reconstruction methods at 2x and 4x magnifications. It also maintains high computational performance on Set5 and Set14 data.
[0094] In a specific example, the deep feature extraction module includes a stack of a preset number of residual modules, each residual module is composed of a cascade of 1x1 convolution, 3x3 convolution, sigmoid and an effective channel attention mechanism, and introduces a depth-separable convolution operation;
[0095] Sending the shallow feature map and the deep feature map of the image into a feature fusion module to obtain a fused feature map;
[0096] Sending the fused feature map to an upsampling module of an image super-resolution generation network to obtain a super-resolution image of the low-resolution image;
[0097] The super-resolution image and the high-resolution image are respectively input into an image super-resolution discrimination network to obtain a probability value of discriminating the super-resolution image as a real image;
[0098] The discriminant network is an image classification convolutional neural network that uses the LeakReLU activation function to prevent negative output necrosis;
[0099] The input image passes through the classification convolutional neural network and then passes through two linear layers to obtain the output result;
[0100] The loss functions of the generation network and the discriminant network are determined according to the probability values of the high-resolution image, the super-resolution image, and the output of the discriminator network.
[0101] In the specific implementation, the image super-resolution discriminator consists of the following parts:
[0102] The discriminator network architecture is based on a convolutional neural network and is designed to distinguish between real and generated images. The network input is a three-channel image, which is processed through a series of convolutional layers, batch normalization layers, and activation functions. The network consists of multiple convolutional blocks, each of which consists of two convolutional layers: a 3x3 convolution in the first layer and a 4x4 convolution in the second layer, followed by a batch normalization layer and a LeakyReLU activation function. Convolutional layers with a stride of 2 are used between convolutional blocks to achieve spatial downsampling, gradually reducing the size of the feature map and increasing the number of channels. Finally, a fully connected layer flattens the feature map into a one-dimensional vector, and two fully connected layers output a single real-valued value, representing the result of the image classification. The entire network design emphasizes doubling the number of channels between convolutional blocks to extract progressively higher-level features. Batch normalization and LeakyReLU activation functions are also used to stabilize the training process.
[0103] The input of the discriminator is fed into the real high-resolution image IHR and the super-resolution reconstructed image ISR and scores them. The discriminant network finally outputs a probability value, which indicates the probability that the input image is a real image.
[0104] In a specific example, the loss function of the discriminant network is:
[0105]
[0106] The generative network loss function consists of three parts:
[0107]
[0108] Among them L percep Using VGG loss:
[0109]
[0110] W and H represent the dimensions of the feature map in the VGG network;
[0111] i and j refer to the jth convolutional layer before the i-th max pooling layer;
[0112] The generated network adversarial loss function
[0113]
[0114] in Used to evaluate the generated super-resolution image G(x i ) and the original high-resolution image, where λ and η are weight coefficients.
[0115] The above content of this embodiment describes the principle of the network loss function. It can be seen that the discriminant network loss function directly calculates the loss between the original high-resolution image and the super-resolution image. The purpose is to make the super-resolved image close to the high-resolution image. The purpose of the generative network loss function is to make the results of the generative network able to deceive the discriminant network well.
[0116] In a specific example, the model training in step S3 includes: optimizing and fine-tuning the model using the prepared video training data samples;
[0117] First, obtain a batch of low-resolution image groups LR from the prepared training data Group , the batch size is set to 32, i.e., 32 low-resolution images are acquired at a time;
[0118] Secondly, the corresponding super-resolution image group SR is generated using the generative network Group;
[0119] Next, the generated SR Group The original high-resolution image group HR in the training data Group The corresponding input is sent to the discriminant network;
[0120] At the same time, the loss function of the generator and the loss function of the discriminator are used to tune the model to complete the model training process.
[0121] This process of this embodiment is intended to allow the model to gradually learn how to generate high-quality reconstructed images from low-resolution images, so as to improve its performance and adaptability.
[0122] In the specific implementation, training can be performed on an Ubuntu 16.04 system with a Tesla K80 GPU, based on PyTorch 1.11.0, using Python 3.6.8, and CUDA version 11.3.1. The training process uses the Adam optimizer with a learning rate of 10e-4 and a batch size of 32 samples.
[0123] This experiment uses the DIV2K dataset as the image training dataset and the REDS dataset as the video training dataset. The training dataset is processed by the data processing module and then fed into the generator and discriminator models for training.
[0124] After 100 rounds of iterations, the generator finally learns the sample distribution of high-resolution images during the adversarial training with the discriminator, and obtains the trained generator model file;
[0125] The model file obtained from the above training is fine-tuned for 50 rounds using the REDS video training dataset to learn how to better handle these noises, thereby improving the robustness to low-quality images and the temporal consistency features.
[0126] In a specific example, the model quantization compression in step S4 further includes: in this conversion process, it is necessary to pre-set the model input size, mean, normalization parameters, and whether to perform quantization configuration options;
[0127] Perform model quantization on the RKNN file, that is, convert floating-point numbers to fixed-point numbers so that it can run efficiently on embedded devices.
[0128] In the specific implementation, the trained super-resolution model file is converted into a model file in the RKNN format supported by the device;
[0129] The weights and activation values of the trained model are quantized and compressed. The quantization process replaces the weights and activation values of 32-bit floating-point numbers with 8-bit fixed-point numbers, and compresses the original signed continuous values into a discrete value range consisting of only 28 fixed-point numbers. The quantization process is as follows:
[0130]
[0131] Where x is a floating point number, xint is a quantized fixed point number, is a rounding operation, s is the quantization scale factor, z is the quantization zero point, and b is the quantization bit width. For example, b is 8 in the INT8 data type; clamp is a truncation operation, which is specifically defined as follows:
[0132]
[0133] Perform embedded continuous testing on the quantized network model on the corresponding test dataset to evaluate the model's accuracy, performance, and memory. If the accuracy of the inference results in the embedded system differs significantly from that in the PC, recalibrate the quantized model using the KL-Divergence quantization algorithm.
[0134] The algorithm in this embodiment updates the distribution of floating-point and fixed-point numbers by adjusting different thresholds. It then determines the maximum and minimum values of the quantization range by minimizing the similarity between the two distributions using the KL-divergence quantization algorithm. By minimizing the distribution difference between floating-point and fixed-point numbers, the KL-divergence quantization algorithm can better adapt to uneven data distributions and mitigate the impact of a few outliers. The quantized data volume is approximately 20-100 images.
[0135] In a specific example, the acceleration of step S6 includes: utilizing a hardware image processing accelerator RGA of an embedded device to speed up image preprocessing;
[0136] The thread pool enables concurrent operations, accelerating model inference speed and improving the system's real-time performance. Thread A is responsible for video reading, thread C is responsible for storing inference results, and multiple threads in thread pool B are responsible for image preprocessing, inference, and post-processing. The number of threads in the thread pool can be dynamically adjusted to accommodate different computing loads.
[0137] Before embedded inference, large-size images are divided into sub-blocks for super-resolution reconstruction using RGA. The sub-block splicing process is optimized, and overlapping region fusion technology is used to address the influence of splicing lines. Specifically, before embedded inference, large-size images are divided into sub-blocks for super-resolution reconstruction using RGA. The sub-block splicing process is optimized, and overlapping region fusion technology is used to address the influence of splicing lines. The computational cost is reduced by optimizing the sub-block processing order and splicing algorithm.
[0138] This embodiment integrates multiple hardware acceleration technologies to fully exploit the hardware's potential, significantly improving the model's inference speed on embedded devices and thus enhancing its practicality. Verification on various test datasets demonstrates that the super-resolution reconstruction method of the present invention not only possesses fast inference capabilities but is also highly portable, accelerating the development of deep learning super-resolution reconstruction in the embedded field.
[0139] In the specific implementation, the deployment module includes:
[0140] Extract image frames from the video file, load the quantized RKNN model file, and complete the model initialization;
[0141] Pre-allocate the required memory resources within the NPU and configure the model's input and output memory information. In this example, the input dimensions are set to (1, 3, 1920, 1080) and the output dimensions are set to (1, 3, 3840, 2160), corresponding to the input and output image resolutions, respectively.
[0142] Each frame is normalized and the channel order is converted to NHWC format. The processed input data is then written to the memory allocated by the NPU and the super-resolution reconstruction model is run for inference calculation.
[0143] After inference is complete, the generated data is copied from the NPU memory back to the CPU for denormalization. Finally, the reconstructed super-resolution image is synthesized into a video and exported as an output file.
[0144] The acceleration modules include:
[0145] Use a hardware image processing accelerator (RGA) to speed up image pre-processing for normalization and channel order conversion in the deployment module;
[0146] To improve reasoning efficiency, a parallel reasoning architecture based on multi-threading (thread pool) was designed;
[0147] The multi-threaded reasoning structure is shown as follows Figure 4 As shown, where:
[0148] Thread A is responsible for reading image frames from the video;
[0149] Thread C is responsible for saving the inference output results;
[0150] The number of threads in thread pool B is determined by the model's computational intensity, hardware device memory, and computing power. In this embodiment, four threads are set to implement parallel reasoning. Multiple threads in thread pool B are used for image preprocessing, model reasoning, and result post-processing.
[0151] Before inference on the embedded device, the larger-resolution image is cropped into multiple sub-blocks, and super-resolution reconstruction is performed on each sub-block separately. This strategy allows the use of a model with a smaller input resolution to reconstruct a high-resolution image, while reducing computational intensity and memory usage while ensuring the stability of the super-resolution effect.
[0152] In this embodiment, if Figure 5 , the resolution of the input image is 1920×1080. To perform 2x super resolution
[0153] For high-resolution reconstruction, the following cropping and splicing process is used:
[0154] The 1920×1080 input image is cropped into 4 sub-images, each with a resolution of 960×540;
[0155] Perform 2x super-resolution reconstruction on each sub-image to generate 4 high-resolution sub-images with a resolution of 1920×1080;
[0156] Finally, the four reconstructed sub-images are stitched into a high-resolution image with a resolution of 3840×2160;
[0157] Directly cropping the sub-images to 960×540 and reconstructing them before stitching may result in obvious stitching marks at the edges of the sub-images, affecting the visual effect.
[0158] This embodiment proposes an optimization solution that introduces overlapping sub-image areas by adjusting the cropping method to eliminate splicing traces;
[0159] Set the upper left corner of the input image as the origin and crop four sub-images with a resolution of 970×550. Each sub-image is cropped 10 pixels more in width and height to form an overlapping area. The specific cropping range is:
[0160] The first sub-image: width 0-970, height 0-550;
[0161] The second sub-image: width 950-1920, height 0-550;
[0162] The third sub-image: width 0-970, height 530-1080;
[0163] The fourth sub-image: width 950-1920, height 530-1080.
[0164] After super-resolution reconstruction of each sub-image, they are spliced in the original cropping order, and the overlapping areas are processed using weighted averaging or seamless fusion algorithms.
[0165] The cropping and splicing method of this embodiment effectively avoids the appearance of splicing marks while ensuring the integrity and quality of the final output image.
[0166] Example 2
[0167] To achieve the above purpose, see Figure 6 This embodiment also discloses an embedded NPU video super-resolution reconstruction system, including:
[0168] Data processing module: used to convert the original high-resolution image into a low-resolution image through a variety of combined degradation operations for model training;
[0169] Model building module: Reconstruct low-resolution images into super-resolution images through a lightweight image super-resolution generator and discriminator;
[0170] Model training module: Use the image training dataset for training; use low-resolution images as the input of the generator network and the corresponding original high-resolution image samples as the expected output of the generator network, and use the back-propagation algorithm to train the generator network; use low-resolution images and corresponding high-resolution images as positive image pairs, and low-resolution images in the training sample dataset and the corresponding output images of the generator network as negative image pairs, and use the back-propagation algorithm to train the discriminator network;
[0171] Model quantization and compression module: used to convert the trained super-resolution model file into the intermediate ONNX format, and further convert the ONNX format model file into the RKNN format compatible with embedded devices;
[0172] Deployment module: Input low-resolution video, read each frame of the video through the OpenCV library, perform necessary preprocessing on each frame, and input it into the quantized model. The model will infer the input low-resolution image and load the model weight file obtained during the training process. After processing the quantized model, it will output a super-resolution image. Finally, the OpenCV library is used to reassemble the processed super-resolution image into a video, thus completing the reconstruction of the super-resolution video.
[0173] Acceleration module: Use hardware acceleration to speed up the inference process of the deployed model.
[0174] Example 3
[0175] To achieve the above objectives, this embodiment also discloses an embedded NPU video super-resolution reconstruction device, including a processor and a storage medium;
[0176] The storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the above method.
[0177] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An embedded NPU video super-resolution reconstruction method, characterized in that: include: S1. Data processing: Convert the original high-resolution image into a low-resolution image through a variety of combined degradation operations to simulate the image degradation process in actual scenes for model training; S2. Model construction: Through a lightweight image super-resolution generator and discriminator, the low-resolution image is reconstructed into a super-resolution image. The generator uses deep separable convolution and an effective channel attention mechanism, combined with residual modules and feature fusion technology, to significantly reduce model parameters and calculations while maintaining high image reconstruction quality. S3, model training: use the image training dataset for training; use the low-resolution image as the input of the generator network, and the corresponding original high-resolution image sample as the expected output of the generator network, and use the back-propagation algorithm to train the generator network; use the low-resolution image and the corresponding high-resolution image as the positive image pair, and the low-resolution image in the training sample dataset and the corresponding output image of the generator network as the negative image pair, and use the back-propagation algorithm to train the discriminator network; S4, model quantization and compression: convert the trained super-resolution model file into the intermediate ONNX format, and further convert the ONNX format model file into the RKNN format compatible with embedded devices; S5, Deployment: Input low-resolution video, read each frame of the video through the opencv library, perform necessary preprocessing on each frame, and input it into the quantized model. The model will infer the input low-resolution image and load the model weight file obtained during the training process. After being processed by the quantized model, the super-resolution image is output; finally, the processed super-resolution image is reassembled into a video using the opencv library, thereby completing the reconstruction of the super-resolution video; S6. Acceleration: Use hardware acceleration to accelerate the reasoning process for the model deployed in step S5.
2. The embedded NPU video super-resolution reconstruction method according to claim 1, characterized in that: The data processing in step S1 includes: performing cropping, Gaussian blur, defocus and motion blur simulation, and Gaussian noise addition on the original high-resolution image to generate a first reference image; Performing JPEG compression on the first reference image, introducing image quality compression, and obtaining a second reference image; The second reference image is down-sampled by bicubic interpolation and normalized to obtain a degraded low-resolution image.
3. The embedded NPU video super-resolution reconstruction method according to claim 2, characterized in that: The image reconstruction in step S2 comprises: sending the low-resolution image to a shallow feature extraction module of an image super-resolution generation network to obtain a shallow feature map of the image; The shallow feature map of the image is sent to the deep feature extraction module of the image super-resolution generation network to obtain the deep feature map of the image.
4. The embedded NPU video super-resolution reconstruction method according to claim 3 is characterized in that: The deep feature extraction module includes a stack of a preset number of residual modules, each residual module is composed of a cascade of 1x1 convolution, 3x3 convolution, sigmoid and an effective channel attention mechanism, and introduces a deep separable convolution operation; Sending the shallow feature map and the deep feature map of the image to a feature fusion module to obtain a fused feature map; Sending the fused feature map to an upsampling module of an image super-resolution generation network to obtain a super-resolution image of the low-resolution image; The super-resolution image and the high-resolution image are respectively input into an image super-resolution discrimination network to obtain a probability value of discriminating the super-resolution image as a real image; The discriminant network is an image classification convolutional neural network, which uses the LeakReLU activation function to prevent negative output necrosis; The input image passes through the classification convolutional neural network and then passes through two linear layers to obtain the output result; The loss functions of the generation network and the discriminant network are determined according to the probability values of the high-resolution image, the super-resolution image and the discriminant network output.
5. The embedded NPU video super-resolution reconstruction method according to claim 4, characterized in that: The loss function of the discriminant network is: The generative network loss function consists of three parts: Where L percep Using VGG loss: W and H represent the dimensions of the feature map in the VGG network; i and j refer to the jth convolutional layer before the i-th maximum pooling layer; The generated network adversarial loss function in Used to evaluate the generated super-resolution image G(x i ) and the 1-norm distance content loss between the original high-resolution image, λ and η are weight coefficients.
6. The embedded NPU video super-resolution reconstruction method according to claim 5, characterized in that: The model training of step S3 includes: optimizing and fine-tuning the model using the prepared video training data samples; First, obtain a batch of low-resolution image groups LR from the prepared training data Group , the batch size is set to 32, i.e., 32 low-resolution images are acquired at a time; Secondly, the corresponding super-resolution image group SR is generated using the generative network Group ; Next, the generated SR Group The original high-resolution image group HR in the training data Group The corresponding input is sent to the discriminant network; At the same time, the loss function of the generator and the loss function of the discriminator are used to tune the model to complete the model training process.
7. The embedded NPU video super-resolution reconstruction method according to claim 1, characterized in that: The model quantization compression of step S4 also includes: in this conversion process, the input size, mean, normalization parameters of the model and whether to perform quantization configuration options need to be pre-set; Perform model quantization on the RKNN file, that is, convert floating-point numbers to fixed-point numbers so that it can run efficiently on embedded devices.
8. The embedded NPU video super-resolution reconstruction method according to claim 1, characterized in that: The acceleration of step S6 includes: using a hardware image processing accelerator RGA of an embedded device to speed up image preprocessing; Concurrent operations are achieved through thread pools, thereby accelerating the reasoning speed of the model and improving the real-time performance of the system; Before embedded inference, RGA is used to divide large-size images into sub-blocks for super-resolution reconstruction, the sub-block splicing process is optimized, and the overlapping area fusion technology is used to solve the influence of the splicing line.
9. An embedded NPU video super-resolution reconstruction system, characterized in that: include: Data processing module: used to convert the original high-resolution image into a low-resolution image through a variety of combined degradation operations for model training; Model building module: Reconstruct low-resolution images into super-resolution images through lightweight image super-resolution generator and discriminator; Model training module: Use image training dataset for training; use low-resolution images as input of the generator network, and the corresponding original high-resolution image samples as the expected output of the generator network, and use the back-propagation algorithm to train the generator network; use low-resolution images and corresponding high-resolution images as positive image pairs, and the low-resolution images in the training sample dataset and the corresponding output images of the generator network as negative image pairs, and use the back-propagation algorithm to train the discriminator network; Model quantization and compression module: used to convert the trained super-resolution model file into the intermediate ONNX format, and further convert the ONNX format model file into the RKNN format compatible with embedded devices; Deployment module: Input low-resolution video, read each frame of the video through the opencv library, perform necessary preprocessing on each frame, and input it into the quantized model. The model will infer the input low-resolution image and load the model weight file obtained during the training process. After being processed by the quantized model, the super-resolution image is output; finally, the processed super-resolution image is reassembled into a video using the opencv library, thereby completing the reconstruction of the super-resolution video; Acceleration module: Use hardware acceleration to accelerate the inference process for the model of the deployment module.
10. An embedded NPU video super-resolution reconstruction device, characterized in that: including processor and storage medium; The storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Single image super-resolution reconstruction method based on conditional generative adversarial network
CN110136063A
Method for generating high-quality image relative to generative adversarial super-resolution reconstruction model
CN112001847A
Video super-resolution network, and video super-resolution, encoding and decoding processing method and device
WO2023000179A1
Cited By
Large model reasoning efficiency dynamic optimization and hardware sensing compression method
CN120494006A
Picture deep learning preprocessing method based on feature compression
CN120655509A
Video processing method and device, storage medium and electronic equipment
CN120751217A
A video processing method, device, storage medium and electronic device
CN120751217B