Image processing method and device, electronic equipment, chip and medium

By combining image processing methods of convolution, pooling, deconvolution and stitching modules, the problem of excessive hardware area caused by increasing model depth is solved, and efficient image super-resolution processing and hardware landing are achieved, and image quality is kept lossless.

CN120374409APending Publication Date: 2025-07-25BEIJING X RING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410545343.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the existing image super-resolution deep learning methods, the increase in model depth leads to increased training difficulty, difficulty in landing, and the hardware area is too large, and it is impossible to take into account both image quality and inference speed.

Method used

By designing an image processing model, the convolution processing, pooling processing, deconvolution processing and splicing are combined to form a branch path to form a simpler structure, reduce parameter and computing power requirements, and combine pruning and sparse module optimization model.

Benefits of technology

It achieves more efficient image processing efficiency and smaller hardware area requirements, while maintaining almost lossless image quality, meeting the standards for hardware implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374409A_ABST
    Figure CN120374409A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method and device, electronic equipment, a chip and a medium, and relates to the field of image processing chip design. The method comprises the following steps: performing convolution processing, first processing, pooling processing and deconvolution processing on an input image to obtain a first intermediate image and a second intermediate image; splicing the first intermediate image and the second intermediate image to obtain a spliced image; performing first processing on the spliced image to obtain a third intermediate image; and performing first operation on the input image and the third intermediate image to obtain a fused image, and performing second processing on the fused image to obtain an output image. According to the method provided by the invention, a plurality of modules of convolution, pooling, deconvolution, splicing and the like are combined to form a branch path, so that the first intermediate image and the second intermediate image are respectively obtained and then are spliced and processed; compared with a continuously deepened model in the prior art, the structure is simpler, the computing power is smaller, parameters are fewer, the processing efficiency is higher, and the area is smaller after the chip is grounded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing, and in particular, to an image processing method, apparatus, electronic device, chip, and medium. Background Art

[0002] Deep learning methods for single-image super-resolution (SR) focus on improving the image quality of SR images, or consider the model inference speed and focus on the efficiency of image quality, or focus on the model inference speed, hardware performance, and image quality. The depth of the convolutional neural network is the key to image super-resolution. However, as the network deepens, the network becomes more and more difficult to train. Most models cannot be implemented on chips and will have problems such as excessive area. Summary of the Invention

[0003] The present disclosure provides an image processing method, apparatus, electronic device, chip, and medium. By designing the model structure for image processing, on the basis of the structure of the existing image processing AI model, multiple modules such as convolutional processing, pooling processing, deconvolution processing, and splicing are combined to form branch paths, so as to obtain a first intermediate image and a second intermediate image respectively, and then they are spliced and processed. Compared with the continuously deepening models in the prior art, the structure is more concise, the computing power is smaller, the number of parameters is less, the processing efficiency is higher, and the area is smaller after being implemented on the chip.

[0004] In a first aspect embodiment of the present disclosure, an image processing method is proposed. The method includes: performing convolutional processing, first processing, pooling processing, and deconvolution processing on an input image to obtain a first intermediate image and a second intermediate image; splicing the first intermediate image and the second intermediate image to obtain a spliced image; performing first processing on the spliced image to obtain a third intermediate image; performing a first operation on the input image and the third intermediate image to obtain a fused image, and performing second processing on the fused image to obtain an output image.

[0005] In some embodiments of the present disclosure, performing convolutional processing, first processing, pooling processing, and deconvolution processing on an input image to obtain a first intermediate image and a second intermediate image includes: sequentially performing first convolutional processing and first processing on the input image to obtain a first intermediate image; sequentially performing first convolutional processing, pooling processing, first processing, and deconvolution processing on the input image to obtain a second intermediate image.

[0006] In some embodiments of the present disclosure, performing first processing on the spliced image to obtain a third intermediate image includes: performing multiple second convolutional processing and activation processing on the spliced image to obtain a fourth intermediate image; performing a third convolutional processing after performing a first operation on the spliced image and the fourth intermediate image to obtain a third intermediate image.

[0007] In some embodiments of the present disclosure, the second processing includes a fourth convolution processing and an upsampling processing.

[0008] An embodiment of the second aspect of the present disclosure provides an image processing method, the method including: obtaining an input image; using a first model to process the input image to obtain an output image, wherein the first model includes: a first branch, a second branch, and a third branch, the first branch sequentially includes a first convolution module, a first processing sub-module, a splicing module, a second processing sub-module, a first addition module, and a second processing module, the second branch sequentially includes a first convolution module, a pooling module, a third processing sub-module, a deconvolution module, a splicing module, a second processing sub-module, a first addition module, and a second processing module, the third branch sequentially includes a first addition module and a second processing module; the first processing sub-module, the second processing sub-module, and the third processing sub-module are configured to perform a first processing.

[0009] In some embodiments of the present disclosure, the second processing module includes a second convolution module and an upsampling module.

[0010] In some embodiments of the present disclosure, any one of the first processing sub-module, the second processing sub-module, and the third processing sub-module includes: a fourth branch and a fifth branch, the fourth branch includes a plurality of convolution activation modules, a second addition module, and a third convolution module, the fifth branch includes a second addition module and a third convolution module, and each convolution activation module includes a fourth convolution module and an activation module.

[0011] In some embodiments of the present disclosure, the first model further includes: a pruning module and a sparsification module.

[0012] An embodiment of the third aspect of the present disclosure provides an image processing apparatus, including: a first module, a second module, a third module, and a fourth module, the first module is configured to perform a convolution processing, a first processing, a pooling processing, and a deconvolution processing on an input image to obtain a first intermediate image and a second intermediate image; the second module is configured to splice the first intermediate image and the second intermediate image to obtain a spliced image; the third module is configured to perform a first processing on the spliced image to obtain a third intermediate image; the fourth module is configured to perform a first operation on the input image and the third intermediate image to obtain a fused image, and perform a second processing on the fused image to obtain an output image.

[0013] A fourth aspect embodiment of the present disclosure provides an image processing apparatus, including an acquisition module and a processing module. The acquisition module is configured to acquire an input image; the processing module is configured to process the input image using a first model to obtain an output image, where the first model includes: a first branch, a second branch, and a third branch. The first branch sequentially includes a first convolution module, a first processing sub-module, a splicing module, a second processing sub-module, a first addition module, and a second processing module. The second branch sequentially includes a first convolution module, a pooling module, a third processing sub-module, a deconvolution module, a splicing module, a second processing sub-module, a first addition module, and a second processing module. The third branch sequentially includes a first addition module and a second processing module. The first processing sub-module, the second processing sub-module, and the third processing sub-module are configured to perform a first processing.

[0014] A fifth aspect embodiment of the present disclosure provides an electronic device, including: a processor and a memory for storing a computer program that can run on the processor. When the processor is configured to run the computer program, it executes the method described in any one of the first aspect or the second aspect embodiments of the present disclosure.

[0015] A sixth aspect embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are configured to cause a computer to execute the method described in any one of the first aspect or the second aspect embodiments of the present disclosure.

[0016] A seventh aspect embodiment of the present disclosure provides a chip, including at least one processor and a communication interface. The communication interface is configured to receive a signal input to the chip or output a signal from the chip. The processor communicates with the communication interface and implements the method described in any one of the first aspect or the second aspect through a logic circuit or by executing code instructions.

[0017] In summary, the image processing method, apparatus, electronic device, chip, and medium provided by the present disclosure include: performing convolution processing, first processing, pooling processing, and deconvolution processing on an input image to obtain a first intermediate image and a second intermediate image; splicing the first intermediate image and the second intermediate image to obtain a spliced image; performing a first processing on the spliced image to obtain a third intermediate image; performing a first operation on the input image and the third intermediate image to obtain a fused image, and performing a second processing on the fused image to obtain an output image. The image processing method provided by the present disclosure combines multiple modules such as convolution processing, pooling processing, deconvolution processing, and splicing to form branch paths to respectively obtain a first intermediate image and a second intermediate image, and then perform splicing and processing. Compared with the existing technology of continuously deepening models, it can reduce the computing power and the area required for hardening of the corresponding image processing model more, improve the efficiency of image processing, and the processed image quality is almost lossless.

[0018] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. Description of the Drawings

[0019] The drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an undue limitation on the present disclosure.

[0020] Figure 1 Schematic flowchart of an image processing method proposed in an embodiment of the present disclosure;

[0021] Figure 2 Schematic flowchart of a method for obtaining a first intermediate image and a second intermediate image proposed in an embodiment of the present disclosure;

[0022] Figure 3 Schematic flowchart of a method for obtaining a third intermediate image proposed in an embodiment of the present disclosure;

[0023] Figure 4 Schematic flowchart of an image processing method proposed in an embodiment of the present disclosure;

[0024] Figure 5 Structural diagram of a first model proposed in an embodiment of the present disclosure;

[0025] Figure 6 Structural diagrams of a first processing sub-module, a second processing sub-module, and a third processing sub-module proposed in an embodiment of the present disclosure;

[0026] Figure 7 Schematic structural diagram of an image processing apparatus proposed in an embodiment of the present disclosure;

[0027] Figure 8 Schematic structural diagram of an image processing apparatus proposed in an embodiment of the present disclosure;

[0028] Figure 9 Schematic structural diagram of an electronic device proposed in an embodiment of the present disclosure;

[0029] Figure 10 Schematic structural diagram of a chip proposed in an embodiment of the present disclosure. Detailed Embodiments

[0030] Embodiments of the present disclosure will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present disclosure and should not be construed as limiting the present disclosure.

[0031] The super-resolution algorithm of an image is to restore a given low-resolution image into a corresponding high-resolution image through a specific algorithm. By constructing various convolutional neural networks, the accuracy and efficiency of image processing are improved. The convolutional neural network obtains feature maps through convolutional operations. Units from different feature maps at each position obtain different types of features. A convolutional layer usually contains multiple feature maps with different weight vectors, enabling richer features of the image to be retained. Currently, the computing power of convolutional neural networks used for image processing is relatively high and cannot meet the standard for hardening and implementation.

[0032] In summary, to solve the technical problems in the related art, an embodiment of the present disclosure provides an image processing method. The method includes performing convolutional processing, first processing, pooling processing, and deconvolution processing on an input image to obtain a first intermediate image and a second intermediate image; splicing the first intermediate image and the second intermediate image to obtain a spliced image; performing first processing on the spliced image to obtain a third intermediate image; performing a first operation on the input image and the third intermediate image to obtain a fused image, and performing second processing on the fused image to obtain an output image. The image processing method proposed by the present disclosure can reduce the computing power of the model for implementing the above method and the area required for hardening more, improve the efficiency of image processing, and achieve almost lossless image quality after processing.

[0033] The image processing method provided by the present application will be introduced in detail below with reference to the accompanying drawings.

[0034] Figure 1 It is a schematic flowchart of an image processing method proposed by an embodiment of the present disclosure. As Figure 1 shown, the method includes the following steps. The image processing method proposed by the present disclosure corresponds to the design of the internal structure of the model.

[0035] Step 101, perform convolutional processing, first processing, pooling processing, and deconvolution processing on the input image to obtain a first intermediate image and a second intermediate image.

[0036] In some embodiments, the convolutional processing can be to process each pixel in the input image using a convolutional kernel. The size of the convolutional kernel can be customized, and the number of convolutional processing times can also be multiple. The size of the convolutional kernel for each convolutional processing can be the same or different. The present disclosure does not limit this. After the model is trained, the parameters can be changed to achieve the best training effect.

[0037] In some embodiments, the pooling process may perform an aggregation statistic on the image features that have undergone the convolution process. In other words, the pooling process may change the size of the image. For example, the length and width of the input image may each be changed to half of the original. Exemplarily, for an image with an input of 256*256, after the pooling process, an image of 128*128 can be obtained, and more image information in the original image can be acquired.

[0038] In some embodiments, the deconvolution process may denoise the image that has undergone the convolution process, that is, it can reconstruct a blurred image into a clear image. The deconvolution process can change the size of the image. For example, the length and width of the image input to the deconvolution process may each be changed to 2 times the original. Exemplarily, an image of 128*128 obtained after the pooling process is converted into an image of 256*256.

[0039] In some embodiments, the number of times of the convolution process, the first process, the pooling process, and the deconvolution process may be once or multiple times, and the present disclosure does not limit this.

[0040] Step 102: Stitch the first intermediate image and the second intermediate image to obtain a stitched image.

[0041] In the embodiments of the present disclosure, stitching the first intermediate image and the second intermediate image may be stitching the channels of the two branches together. For example, the pixel values at the same position of the first intermediate image and the second intermediate image in the same coordinate system can be merged to obtain a high-resolution stitched image.

[0042] Step 103: Perform a first process on the stitched image to obtain a third intermediate image.

[0043] In the embodiments of the present disclosure, performing a first process on the stitched image may be performing multiple convolution processes and multiple activation processes on the stitched image to perform linear processing and non-linear processing on the stitched image, removing data redundancy in the stitched image, and retaining the features of the image data. The parameters of the convolution process are variable, that is, after the model is trained multiple times, the parameters that can make the training effect of the model the best can be obtained.

[0044] Step 104: Perform a first operation on the input image and the third intermediate image to obtain a fused image, and perform a second process on the fused image to obtain an output image.

[0045] In the embodiments of the present disclosure, the first operation may be an addition operation.

[0046] Specifically, perform an addition operation on the input image and the third intermediate image to obtain a fused image.

[0047] In an embodiment of the present disclosure, fusing the input image and the third intermediate image obtained after a series of processes may be adding the gray values of the pixels at the same position of the input image and the third intermediate image or the values of each channel of the color pixels respectively.

[0048] In an embodiment of the present disclosure, the second process includes a fourth convolution process and an upsampling process.

[0049] In an embodiment of the present disclosure, the parameters of the fourth convolution process, the first convolution process, the second convolution process, and the third convolution process may be the same or different, and the present disclosure does not limit this. In other words, after the model is trained, it can be iterated multiple times to obtain the parameters that result in the best training effect.

[0050] In an embodiment of the present disclosure, the upsampling process may be to enlarge the fused image obtained after convolution processing. It may use the interpolation method, that is, on the basis of the original pixels of the fused image obtained after convolution processing, appropriate interpolation algorithms are used to insert new elements between the pixel point values. Among them, the interpolation algorithm may be any one of the nearest neighbor interpolation method, the single linear interpolation method, and the bilinear interpolation method, and the present disclosure does not limit this.

[0051] Exemplarily, for an input image of 256*256, the upsampling process may enlarge the 256*256 image to a 512*512 image.

[0052] In summary, according to the image processing method proposed by the present disclosure, by combining multiple modules such as convolution processing, pooling processing, deconvolution processing, splicing, and fusion, the computing power of the corresponding model and the area required for hardening can be reduced more, the efficiency of image processing is higher, and almost lossless image quality after processing can be achieved.

[0053] Based on Figure 1 the embodiments shown, Figure 2 a flowchart of the method for obtaining the first intermediate image and the second intermediate image is Figure 2 For Figure 1 step 101 in Figure 2 is further described as follows.

[0054] Step 201: Perform a first convolution process and a first process on the input image in sequence to obtain a first intermediate image.

[0055] In an embodiment of the present disclosure, the first process may be multiple convolution processes and multiple activation processes, that is, performing linear transformation of multiple convolutions and non-linear transformation of activation functions on the input image obtained after convolution processing, so that the obtained first intermediate image can remove data redundancy in the input image and retain the features of the input image data.

[0056] In an embodiment of the present disclosure, the parameters of the first convolution process and the parameters of the multiple convolution processes in the first process may be the same or different. In other words, after the corresponding model is trained, the parameters of each convolution process are obtained through multiple iterations, so that the final output image has a better effect.

[0057] Step 202: Perform a first convolution process, a pooling process, a first process, and a deconvolution process on the input image in sequence to obtain a second intermediate image.

[0058] In an embodiment of the present disclosure, after performing the first convolution process on the input image, a pooling process is then performed to aggregate and statistically analyze the image features after the convolution process, and the length and width of the image are each changed to half of the original. For example, if the input image is an image of 256*256, it remains an image of 256*256 after the first convolution process, and an image of 128*128 can be obtained after the pooling process, that is, more image information in the original image can be obtained.

[0059] In an embodiment of the present disclosure, the 128*128 image obtained after the pooling process and the first process can be reconstructed into a clear image through the deconvolution process, that is, the size of the image is changed. For example, the 128*128 image is converted into a 256*256 image to achieve image reconstruction.

[0060] In the above embodiment, the pooling process and the deconvolution process cooperate to form a downsampling branch, which can increase the receptive field of the corresponding image processing model without increasing the total number of network layers.

[0061] Based on Figure 1-2 the embodiment shown, Figure 3 is a flowchart of a method for obtaining a third intermediate image. Figure 3 For Figure 1 step 103 in, Figure 3 a further description is as follows:

[0062] Step 301: Perform multiple convolution-activation processes on the spliced image to obtain a fourth intermediate image.

[0063] In an embodiment of the present disclosure, the convolution-activation process includes a second convolution process and an activation process. In other words, Figure 1 , Figure 2 the first process in can be multiple convolution-activation processes.

[0064] In an embodiment of the present disclosure, performing multiple convolution-activation processes on the spliced image may involve performing multiple second convolution processes and activation processes on the spliced image. Among them, the parameters of the second convolution process may be the same as or different from those of the first convolution process. The activation process may be to append an activation function to the image that has undergone the second convolution process, mapping the linear transformation to a non-linear space to achieve image segmentation. The parameters of the second convolution process may be parameters obtained through multiple iterations of model training to achieve better training effects.

[0065] In some embodiments, the activation function in the activation process may be any one of the ReLU (Rectified Linear Unit) activation function, Leaky ReLU, PReLU, and ELU, and the present disclosure is not limited thereto.

[0066] Exemplarily, a 3*3 convolution kernel is used in the second convolution process, and the ReLU activation function is used in the activation process. The formula for the ReLU activation function is: f(x) = max(0, x), where x is the input value and f(x) is the output value. When x > 0, f(x) = x; when x <= 0, f(x) = 0.

[0067] Step 302: After performing a first operation on the spliced image and the fourth intermediate image, perform a third convolution process to obtain a third intermediate image.

[0068] In an embodiment of the present disclosure, the first operation may be an addition operation.

[0069] Specifically, after performing an addition operation on the spliced image and the fourth intermediate image, perform a third convolution process to obtain a third intermediate image.

[0070] In an embodiment of the present disclosure, the parameters of the third convolution process may be the same as or different from those of the first convolution process and the second convolution process. The parameters of the third convolution process may be parameters obtained after multiple iterations during the model training process.

[0071] The process of the first process in the above embodiments is the same as that of Figure 1 、 Figure 2 The first process in Figure 1 、 Figure 2 may refer to the steps in the embodiment shown in Figure 3

[0072] In summary, the image processing method proposed in the present disclosure combines convolution processing, activation processing, deconvolution processing, pooling processing, activation processing, and upsampling processing in an image processing model to form branch paths, thereby obtaining a first intermediate image and a second intermediate image, and then performing splicing and processing. Compared with continuously deepening models, it can make the image processing more efficient, the performance of image super-resolution processing better, and the computing power and area of the image processing model corresponding to the above image processing method decrease more, enabling it to meet the standards for hardening and implementation.

[0073] Figure 4 This is the method flow chart of the image processing method proposed in the embodiment of the present disclosure. As Figure 4 shown, it includes the following steps:

[0074] Step 401, obtain an input image.

[0075] Exemplarily, obtain an input image of 256*256.

[0076] Step 402, use a first model to process the input image to obtain an output image.

[0077] In the embodiment of the present disclosure, the first model includes a first branch, a second branch, and a third branch. The first branch sequentially includes a first convolution module, a first processing sub-module, a splicing module, a second processing sub-module, a first addition module, and a second processing module. The second branch sequentially includes a first convolution module, a pooling module, a third processing sub-module, a deconvolution module, a splicing module, a second processing sub-module, a first addition module, and a second processing module. The third branch sequentially includes a first addition module and a second processing module. The first processing sub-module, the second processing sub-module, and the third processing sub-module are used to perform the first processing.

[0078] In the embodiment of the present disclosure, the pooling module is used to perform pooling processing, that is, aggregate statistics on the image and change the size of the image. Exemplarily, the length and width of the image input to the pooling module can be changed to half of the original, that is, for an input image of 256*256, an image of 128*128 can be obtained after pooling processing, and more image information in the original image can be obtained.

[0079] In the embodiment of the present disclosure, the deconvolution module is used to perform deconvolution processing, that is, denoise the image and change the size of the image. The length and width of the image input to the deconvolution module are changed to 2 times the original. Exemplarily, an image of 128*128 is converted to an image of 256*256 to achieve image magnification.

[0080] In an embodiment of the present disclosure, the first convolution module is used to perform first convolution processing. The parameters of the first convolution processing can be preset. For example, the first convolution module can perform first convolution processing with a 5*5 convolution kernel. During the training process of the first model, through multiple iterations, the parameters of the convolution processing that result in the best training effect are obtained.

[0081] In an embodiment of the present disclosure, the second processing module includes a second convolution module and an upsampling module. Among them, the second convolution module is used to perform second convolution processing. The parameters of the second convolution processing can be the same as or different from the parameters of the first convolution processing. For example, the second convolution module can perform second convolution processing with a 5*5 convolution kernel. During the training process of the first model, through multiple iterations, the parameters of the convolution processing that result in the best training effect are obtained.

[0082] In an embodiment of the present disclosure, the upsampling module is used to perform upsampling processing, that is, to enlarge the image. It can use the interpolation method. Based on the original pixels, appropriate interpolation algorithms are used to insert new elements between the pixel values. The interpolation algorithm can be any one of the nearest neighbor interpolation method, the bilinear interpolation method, and the bicubic interpolation method. The present disclosure is not limited thereto.

[0083] In an embodiment of the present disclosure, any one of the first processing sub-module, the second processing sub-module, and the third processing sub-module includes: a fourth branch and a fifth branch. The fourth branch includes multiple convolution activation modules, a second addition module, and a third convolution module. The fifth branch includes a second addition module and a third convolution module. Each convolution activation module includes a fourth convolution module and an activation module.

[0084] In an embodiment of the present disclosure, the third convolution module is used to perform third convolution processing, and the fourth convolution module is used to perform fourth convolution processing. The parameters of the third convolution processing and the fourth convolution processing can be the same as or different from the parameters of the first convolution processing and the second convolution processing. The present disclosure is not limited thereto. For example, the convolution kernels in the third convolution module and the fourth convolution module can be 3*3 matrices.

[0085] In an embodiment of the present disclosure, the parameters of the convolution processing in the first convolution module, the second convolution module, the third convolution module, and the fourth convolution module can be the parameter values that result in the best training effect obtained through multiple iterations during the training process of the first model.

[0086] In an embodiment of the present disclosure, the structures of the first processing sub-module, the second processing sub-module, and the third processing sub-module are the same, and they are all modules used to perform first processing.

[0087] For example, the structure of the first model is as Figure 5As shown, where CONV represents a module for performing convolution processing, RLFB represents a module for performing first processing, cat represents a module for performing concatenation processing, the plus sign represents a module for performing addition processing, and CONV and PIXELSHUFFLE together constitute a second processing module for performing second processing. MaxPooling represents a module for performing pooling processing, and DeConv represents a module for performing deconvolution processing.

[0088] Exemplarily, the structural diagrams of the first processing sub-module, the second processing sub-module, and the third processing sub-module are as Figure 6 shown. Among them, CONV represents a module for performing convolution processing, RELU represents a module for performing activation processing, the plus sign represents a module for performing addition processing, and each combination of CONV and RELU can be represented as a convolution activation module. The module for performing first processing can include multiple convolution activation modules, for example, three. In an embodiment of the present disclosure, the first model further includes a pruning module and a sparsification module.

[0089] Among them, the pruning module is used to perform pruning processing on the trained model. In other words, the pruning processing can optimize the model, remove redundant connections and neurons in the model, thereby reducing the amount of computation and memory occupancy, improving the inference speed of the model. At the same time, the pruning processing can also prevent overfitting and improve the generalization ability of the model.

[0090] Among them, the sparsification module is used to perform sparsification processing on the model after pruning processing. In other words, the sparsification processing can optimize the model after pruning processing, reduce the storage requirements and amount of computation of the model, improve the efficiency and performance of the model, reduce the use of computing resources and memory, and enable the model after sparsification processing to be applied in mobile devices or embedded systems.

[0091] In the above embodiment, by adding a pruning module and a sparsification module, the computing power and area of the first model can both be reduced to meet the standard of hardening and implementation.

[0092] In summary, the image processing method proposed in the present disclosure designs the model according to the structure of the first model proposed in the present disclosure, making the first model more convenient for quantization after training to meet the standard of hardening and implementation. At the same time, with the cooperation of pruning and sparsification processing, the computing power and area of the model are reduced more, and the loss of image quality is minimized in image processing.

[0093] Figure 7 It is a schematic structural diagram of an image processing apparatus 700 according to an embodiment of the present disclosure. As Figure 7 shown, the apparatus includes: a first module 710, a second module 720, a third module 730, and a fourth module 740.

[0094] The first module 710 is configured to perform convolution processing, first processing, pooling processing, and deconvolution processing on the input image to obtain a first intermediate image and a second intermediate image;

[0095] The second module 720 is configured to splice the first intermediate image and the second intermediate image to obtain a spliced image;

[0096] The third module 730 is configured to perform first processing on the spliced image to obtain a third intermediate image;

[0097] The fourth module 740 is configured to perform a first operation on the input image and the third intermediate image to obtain a fused image, and perform second processing on the fused image to obtain an output image.

[0098] In some embodiments, the first module is further configured to sequentially perform first convolution processing and first processing on the input image to obtain a first intermediate image; sequentially perform first convolution processing, pooling processing, first processing, and deconvolution processing on the input image to obtain a second intermediate image.

[0099] In some embodiments, the third module is further configured to perform multiple second convolution processing and activation processing on the spliced image to obtain a fourth intermediate image; perform a first operation on the spliced image and the fourth intermediate image, and then perform third convolution processing to obtain a third intermediate image.

[0100] In some embodiments, the second processing includes fourth convolution processing and upsampling processing.

[0101] In summary, the image processing device proposed in the present disclosure obtains a first intermediate image and a second intermediate image by performing convolution processing, first processing, pooling processing, and deconvolution processing on the input image; splices the first intermediate image and the second intermediate image to obtain a spliced image; performs first processing on the spliced image to obtain a third intermediate image; performs a first operation on the input image and the third intermediate image to obtain a fused image, and performs second processing on the fused image to obtain an output image. The image processing device proposed in the present disclosure combines multiple modules such as convolution processing, pooling processing, deconvolution processing, splicing, and fusion, so that the designed model has less computing power and fewer parameters compared to the original model, can improve the efficiency and performance of image processing, achieve almost lossless image quality after processing, and the corresponding model has less computing power and area, meeting the standard for hardening and implementation.

[0102] Figure 8 It is a schematic structural diagram of an image processing device 800 according to an embodiment of the present disclosure. As Figure 8 shown, the device includes: an acquisition module 810 and a processing module 820.

[0103] The acquisition module 810 is configured to acquire an input image.

[0104] The processing module 820 is used to process the input image using the first model to obtain an output image. The first model includes: a first branch, a second branch, and a third branch. The first branch sequentially includes a first convolution module, a first processing sub-module, a splicing module, a second processing sub-module, a first addition module, and a second processing module. The second branch sequentially includes a first convolution module, a pooling module, a third processing sub-module, a deconvolution module, a splicing module, a second processing sub-module, a first addition module, and a second processing module. The third branch sequentially includes a first addition module and a second processing module. The first processing sub-module, the second processing sub-module, and the third processing sub-module are used to perform the first processing.

[0105] In some embodiments, the second processing module includes a second convolution module and an upsampling module.

[0106] In some embodiments, any one of the first processing sub-module, the second processing sub-module, and the third processing sub-module includes: a fourth branch and a fifth branch. The fourth branch includes a plurality of convolution activation modules, a second addition module, and a third convolution module. The fifth branch includes a second addition module and a third convolution module. Each convolution activation module includes a fourth convolution module and an activation module.

[0107] In some embodiments, the first model further includes: a pruning module and a sparsification module.

[0108] In summary, the image processing apparatus proposed in the present disclosure designs the model according to the structure of the first model proposed in the present disclosure, making the first model more convenient for quantization after training, meeting the standard for hardening and implementation. At the same time, with the cooperation of pruning and sparse processing, the computing power and area of the model are reduced more, and the loss of image quality can be minimized in image processing.

[0109] Figure 9 It is a schematic structural diagram of an electronic device 900 for implementing the above image processing method shown according to an exemplary embodiment.

[0110] Refer to Figure 9 , the electronic device 900 may include one or more of the following components: a processing component 902, a memory 904, a power component 906, a multimedia component 908, an audio component 910, an input / output (I / O) interface 912, a sensor component 914, and a communication component 916.

[0111] The processing component 902 generally controls the overall operation of the electronic device 900, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 902 may include one or more processors 920 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 902 may include one or more modules to facilitate the interaction between the processing component 902 and other components. For example, the processing component 902 may include a multimedia module to facilitate the interaction between the multimedia component 908 and the processing component 902.

[0112] The memory 904 is configured to store various types of data to support the operation of the electronic device 900. Examples of such data include instructions for any application or method operating on the electronic device 900, contact data, phone book data, messages, pictures, videos, etc. The memory 904 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks, or optical disks.

[0113] The power component 906 provides power to various components of the electronic device 900. The power component 906 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 900.

[0114] The multimedia component 908 includes a screen that provides an output interface between the electronic device 900 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 908 includes a front camera and / or a rear camera. When the electronic device 900 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0115] The audio component 910 is configured to output and / or input audio signals. For example, the audio component 910 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 900 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 904 or transmitted via the communication component 916. In some embodiments, the audio component 910 further includes a speaker for outputting audio signals.

[0116] The I / O interface 912 provides an interface between the processing component 902 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.

[0117] The sensor component 914 includes one or more sensors for providing an assessment of the status of various aspects of the electronic device 900. For example, the sensor component 914 can detect the on / off state of the electronic device 900, the relative positioning of components, such as the display and keypad of the electronic device 900. The sensor component 914 can also detect a change in the position of the electronic device 900 or a component of the electronic device 900, the presence or absence of user contact with the electronic device 900, the orientation or acceleration / deceleration of the electronic device 900, and a change in the temperature of the electronic device 900. The sensor component 914 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 914 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 914 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0118] The communication component 916 is configured to facilitate communication between the electronic device 900 and other devices in a wired or wireless manner. The electronic device 900 can access a wireless network based on communication standards, such as WiFi, 2G or 3G, 4G LTE, 5G NR (New Radio), or a combination thereof. In an exemplary embodiment, the communication component 916 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 916 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0119] In an exemplary embodiment, the electronic device 900 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0120] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 904 including instructions, and the above instructions can be executed by a processor 920 of the electronic device 900 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0121] An embodiment of the present disclosure also proposes a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the image processing method described in the above embodiments of the present disclosure.

[0122] An embodiment of the present disclosure also proposes a computer program product, including a computer program, and the computer program is used to execute the image processing method described in the above embodiments of the present disclosure when being executed by a processor.

[0123] An embodiment of the present disclosure also proposes a chip, including at least one processor and a communication interface, where the communication interface is used to receive a signal input to the chip or output a signal from the chip, and the processor communicates with the communication interface and implements the image processing method described in the above embodiments of the present disclosure through logic circuits or by executing code instructions.

[0124] Figure 10 FIG. 1000 is a schematic structural diagram of a chip for implementing the above image processing method according to an exemplary embodiment. Referring to Figure 10 , the chip 1000 includes at least one communication interface 1001 and a processor 1002. The communication interface 1001 is used to receive a signal input to the chip 1000 or output a signal from the above chip 1000, and the processor 1002 communicates with the communication interface 1001 and implements the image processing method described in the above embodiments of the present disclosure through logic circuits or by executing code instructions.

[0125] It should be noted that the terms "first", "second", etc. in the description of the present disclosure, the claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0126] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples" or "some examples", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0127] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment or part of code including one or more executable instructions for implementing a specific logical function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations, where the functions can be executed in a manner that is not shown or discussed in sequence, including in a substantially simultaneous manner or in a reverse order according to the functions involved, which should be understood by those skilled in the technical field to which the embodiments of the present disclosure belong.

[0128] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processing module, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (control method), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0129] It should be understood that various parts of the embodiments of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0130] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0131] In addition, each functional unit in various embodiments of the present disclosure may be integrated into one processing module, may exist separately as individual physical units, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, or the like.

[0132] Although the embodiments of the present disclosure have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.

Claims

1. An image processing method, characterized in that, Comprising: Performing convolution processing, first processing, pooling processing, and deconvolution processing on the input image to obtain a first intermediate image and a second intermediate image; Splicing the first intermediate image and the second intermediate image to obtain a spliced image; Performing the first processing on the spliced image to obtain a third intermediate image; Performing a first operation on the input image and the third intermediate image to obtain a fused image, and performing a second processing on the fused image to obtain an output image.

2. The method according to claim 1, wherein The performing convolution processing, first processing, pooling processing, and deconvolution processing on the input image to obtain a first intermediate image and a second intermediate image comprises: Sequentially performing a first convolution processing and the first processing on the input image to obtain a first intermediate image; Sequentially performing the first convolution processing, the pooling processing, the first processing, and the deconvolution processing on the input image to obtain a second intermediate image.

3. The method according to claim 1, characterized in that, The performing the first processing on the spliced image to obtain a third intermediate image comprises: Performing multiple convolution-activation processings on the spliced image to obtain a fourth intermediate image, where the convolution-activation processing comprises a second convolution processing and an activation processing; Performing the first operation on the spliced image and the fourth intermediate image, and then performing a third convolution processing to obtain a third intermediate image.

4. The method according to claim 1, characterized in that, The second processing comprises a fourth convolution processing and an upsampling processing.

5. An image processing method, characterized in that, Comprising: Obtaining an input image; Using a first model to process the input image to obtain an output image, wherein the first model comprises: a first branch, a second branch, and a third branch. The first branch sequentially comprises a first convolution module, a first processing sub-module, a splicing module, a second processing sub-module, a first addition module, and a second processing module. The second branch sequentially comprises the first convolution module, a pooling module, a third processing sub-module, a deconvolution module, the splicing module, the second processing sub-module, the first addition module, and the second processing module. The third branch sequentially comprises the first addition module and the second processing module. The first processing sub-module, the second processing sub-module, and the third processing sub-module are used to perform a first processing.

6. The method according to claim 5, wherein The second processing module comprises a second convolution module and an upsampling module.

7. The method according to claim 5, characterized in that Any one of the first processing sub-module, the second processing sub-module, and the third processing sub-module comprises: a fourth branch and a fifth branch. The fourth branch comprises a plurality of convolution activation modules, a second addition module, and a third convolution module. The fifth branch comprises the second addition module and the third convolution module. Each convolution activation module comprises a fourth convolution module and an activation module.

8. The method according to claim 5, characterized in that, The first model further comprises: a pruning module and a sparsification module.

9. An image processing apparatus, characterized in that, Comprising a first module, a second module, a third module, and a fourth module, wherein the first module is used to perform convolution processing, first processing, pooling processing, and deconvolution processing on the input image to obtain a first intermediate image and a second intermediate image; The second module is used to splice the first intermediate image and the second intermediate image to obtain a spliced image; The third module is configured to perform the first processing on the spliced image to obtain a third intermediate image; The fourth module is configured to perform a first operation on the input image and the third intermediate image to obtain a fused image, and perform a second processing on the fused image to obtain an output image.

10. An image processing apparatus, characterized in that, It includes an acquisition module and a processing module, The acquisition module is configured to acquire an input image; The processing module is configured to use a first model to process the input image to obtain an output image, wherein, the first model includes: a first branch, a second branch, and a third branch. The first branch sequentially includes a first convolution module, a first processing sub-module, a splicing module, a second processing sub-module, a first addition module, and a second processing module. The second branch sequentially includes the first convolution module, a pooling module, a third processing sub-module, a deconvolution module, the splicing module, the second processing sub-module, the first addition module, and the second processing module. The third branch sequentially includes the first addition module and the second processing module. The first processing sub-module, the second processing sub-module, and the third processing sub-module are configured to perform a first processing.

11. An electronic device, characterized in that, It includes: a processor and a memory for storing a computer program that can run on the processor, wherein, when the processor is configured to run the computer program, it executes the method according to any one of claims 1-5 or 6-8.

12. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-5 or 6-8.

13. A chip, characterized in that, It includes at least one processor and a communication interface. The communication interface is configured to receive a signal input to the chip or output a signal from the chip. The processor communicates with the communication interface and implements the method according to any one of claims 1-5 or 6-8 through logic circuits or by executing code instructions.