Training method of image processing model, and image processing method and device

By adding multiple attention units and Sigmoid functions to the image processing model, the model structure is simplified, and efficient image super-resolution reconstruction is realized, solving the problems of model complexity and high computing cost in the existing technology, and improving image processing speed and image quality.

CN120375142APending Publication Date: 2025-07-25BEIJING X RING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410432860.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-10
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing image super-resolution method based on convolutional neural networks relies on a large number of parameters and complex network structures, resulting in problems such as large computing power, high power consumption and high latency.

Method used

Multiple attention units are added to the image processing model, stacked in series and generated attention weights using Sigmoid function, simplifying the model structure, and super-resolution reconstruction of images through a fast parameterless attention network.

Benefits of technology

It reduces the complexity of the model and the calculation cost, improves the running speed and image quality of the image processing model, and solves problems such as large computing power, high power consumption and high latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375142A_ABST
    Figure CN120375142A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing model training method, an image processing method and an image processing device. Wherein the image processing model comprises a first convolution unit, a plurality of attention units, a second convolution unit, a Concat unit, a third convolution unit and an up-sampling unit, the plurality of attention units are stacked in series, and the first convolution unit, the plurality of attention units, the second convolution unit, the Concat unit, the third convolution unit and the up-sampling unit are sequentially connected in series. A Sigmoid function is adopted in each attention unit to generate an attention weight, the training method comprises the steps that a training set image and a corresponding label image are acquired, and the resolution ratio of the training set image is smaller than that of the label image; performing super-resolution reconstruction on the training set image based on the initial image processing model to obtain a super-resolution image; generating a loss value according to the super-resolution image and the label image; and training an image processing model according to the loss value. According to the embodiment of the invention, the complexity and calculation cost of the model can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of computer vision and image processing, and artificial intelligence fields such as deep learning, and particularly relates to a method for training an image processing model, an image processing method, and an apparatus. Background Art

[0002] In the fields of computer vision and image processing, especially in the field of image super-resolution (ISR), significant development has been achieved in recent years. The image super-resolution technology aims to recover high-resolution images from low-resolution images, which is very important for improving image quality and details, especially in applications such as high-definition display and advanced image editing.

[0003] In related technologies, the super-resolution method based on a convolutional neural network (CNN) has become a research hotspot. By training a deep network to learn the mapping relationship from low resolution to high resolution, it can effectively recover image details and improve the quality of super-resolution. However, this method usually relies on a large number of parameters and complex network structures to learn the mapping from low-resolution to high-resolution images, resulting in problems such as high model computing power, high power consumption, and high latency. Summary of the Invention

[0004] The present disclosure provides a method for training an image processing model, an image processing method, and an apparatus.

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a method for training an image processing model. The image processing model includes a first convolutional unit, a plurality of attention units, a second convolutional unit, a connection Concat unit, a third convolutional unit, and an upsampling unit. The plurality of attention units are stacked in series. The first convolutional unit, the plurality of attention units, the second convolutional unit, the Concat unit, the third convolutional unit, and the upsampling unit are connected in series in sequence. An Sigmoid function is used in each of the attention units to generate attention weights. The training method includes:

[0006] Obtain a training set image and a label image corresponding to the training set image, where the resolution of the training set image is less than the resolution of the label image;

[0007] Perform super-resolution reconstruction on the training set image based on an initial image processing model to obtain a super-resolved image;

[0008] Generate a loss value according to the super-resolved image and the label image;

[0009] Train the image processing model according to the loss value.

[0010] According to a second aspect of the embodiments of the present disclosure, there is provided an image processing method, including:

[0011] Obtain an image to be processed;

[0012] Input the image to be processed into a preset image processing model to obtain a target image corresponding to the image to be processed, where the resolution of the target image is greater than the resolution of the image to be processed; wherein, the image processing model is trained by the method described above.

[0013] According to a third aspect of the embodiments of the present disclosure, there is provided a training device for an image processing model. The image processing model includes a first convolutional unit, a plurality of attention units, a second convolutional unit, a connection Concat unit, a third convolutional unit, and an upsampling unit. The plurality of attention units are stacked in series. The first convolutional unit, the plurality of attention units, the second convolutional unit, the Concat unit, the third convolutional unit, and the upsampling unit are connected in series in sequence. An Sigmoid function is used in each of the attention units to generate attention weights. The training device includes:

[0014] An acquisition module, configured to acquire a training set image and a label image corresponding to the training set image, where the resolution of the training set image is less than the resolution of the label image;

[0015] A super-resolution reconstruction module, configured to perform super-resolution reconstruction on the training set image based on an initial image processing model to obtain a super-resolved image;

[0016] A generation module, configured to generate a loss value according to the super-resolved image and the label image;

[0017] A training module, configured to train the image processing model according to the loss value.

[0018] According to a fourth aspect of the embodiments of the present disclosure, there is provided an image processing device, including:

[0019] An acquisition module, configured to acquire an image to be processed;

[0020] A prediction module, configured to input the image to be processed into a preset image processing model to obtain a target image corresponding to the image to be processed, where the resolution of the target image is greater than the resolution of the image to be processed; wherein, the image processing model is trained by the method described above.

[0021] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including: one or more processors; wherein, the electronic device is configured to execute the training method of the image processing model described in the foregoing first aspect, or execute the image processing method described in the second aspect.

[0022] According to a sixth aspect of the embodiments of the present disclosure, there is provided a storage medium storing instructions, which, when running on an electronic device, cause the electronic device to execute the training method of the image processing model described in the first aspect, or execute the image processing method described in the second aspect.

[0023] According to a seventh aspect of the embodiments of the present disclosure, there is provided a computer program product including a computer program, which, when executed by a processor, implements the steps of the method described in the first aspect, or implements the steps of the method described in the foregoing second aspect.

[0024] According to an eighth aspect of the embodiments of the present disclosure, there is provided a chip, including: one or more processors; wherein, the processor is configured to call instructions to cause the chip to execute the training method of the image processing model described in the foregoing first aspect, or execute the image processing method described in the foregoing second aspect.

[0025] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:

[0026] Improve the model structure of the image processing model, add multiple attention units to the model, and the multiple attention units are stacked in series. The Sigmoid function is used for calculation in each attention unit, so that the attention network does not need to set parameters, simplifies the model structure, reduces the complexity and calculation cost of the model, greatly improves the running speed of the image processing model, and solves the problems of large computing power, high power consumption and high latency of the image processing model in the related art. Since the attention mechanism is added, the model can still effectively focus on the important features of the image, and the image quality effect in the super-resolution task can be improved.

[0027] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.

[0029] Figure 1A It is a schematic diagram of the network structure of the image processing model provided by the embodiments of the present disclosure.

[0030] Figure 1BSchematic diagram of the attention unit provided by an embodiment of the present disclosure.

[0031] Figure 2 Flowchart of a method for training an image processing model provided by an embodiment of the present disclosure.

[0032] Figure 3 Flowchart of a method for training an image processing model provided by an embodiment of the present disclosure.

[0033] Figure 4 Example diagram of reparameterization processing for a convolutional layer provided by an embodiment of the present disclosure.

[0034] Figure 5 Flowchart of an image processing method provided by an embodiment of the present disclosure.

[0035] Figure 6 Block diagram of a training device for an image processing model provided by an embodiment of the present disclosure.

[0036] Figure 7 Block diagram of an image processing device provided by an embodiment of the present disclosure.

[0037] Figure 8 Block diagram of an electronic device 800 shown according to an exemplary embodiment.

[0038] Figure 9 Schematic diagram of the structure of a chip 900 proposed by an embodiment of the present disclosure. Detailed implementation

[0039] Here, exemplary embodiments will be described in detail, and their examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0040] Image super-resolution technology aims to recover high-resolution images from low-resolution images, which is very important for improving image quality and details, especially in applications such as high-definition display and advanced image editing.

[0041] In the related art, the super-resolution method based on Convolutional Neural Network (CNN) has become a research hotspot. By training a deep network to learn the mapping relationship from low resolution to high resolution, it can effectively restore image details and improve the quality of super-resolution. However, this method usually relies on a large number of parameters and complex network structures to learn the mapping from low-resolution to high-resolution images, resulting in problems such as high model computing power, high power consumption, and high latency.

[0042] Based on this, the present disclosure provides a method for training an image processing model, an image processing method, and an apparatus. Among them, the image processing model in the present disclosure involves a fast parameter-free attention network, which can be used to improve the resolution of images and make the images clearer. The technical solutions provided by the embodiments of the present disclosure can be applied to scenarios that require high-resolution images, such as satellite image analysis, medical imaging, security monitoring, high-definition video production, etc.

[0043] Figure 1A It is a schematic diagram of the network structure of the image processing model provided by the embodiments of the present disclosure. As Figure 1A shown, the image processing model may include a first convolutional unit 101, an attention unit 102, a second convolutional unit 103, a Concat (connection) unit 104, a third convolutional unit 105, and an upsampling unit 106. Among them, there may be multiple attention units 102, Figure 1A The number and form of the attention units shown are only for illustration and do not constitute a limitation on the embodiments of the present disclosure. In practical applications, there may be two or more attention units, or there may be three or more attention units. Figure 1A The attention units shown take six attention units 102 as an example.

[0044] As Figure 1A shown, multiple attention units 102 are stacked in series. The first convolutional unit 101, multiple attention units 102, the second convolutional unit 103, the Concat unit 104, the third convolutional unit 105, and the upsampling unit 106 are connected in series in sequence. Each attention unit 102 uses a sigmoid function to generate attention weights. Exemplarily, the first convolutional unit 101, the second convolutional unit 103, and the third convolutional unit 105 may respectively select a convolutional kernel with a size of 3×3. Exemplarily, the attention unit 102 is a basic building block of the image processing model, and can perform sigmoid calculation on features and use them as attention maps.

[0045] As Figure 1BAs shown, each attention unit 102 may include a first convolutional layer 1021, a first activation function layer 1022, a second convolutional layer 1023, a second activation function layer 1024, a third convolutional layer 1025, and a third activation function layer 1026. The first convolutional layer 1021, the first activation function layer 1022, the second convolutional layer 1023, the second activation function layer 1024, and the third convolutional layer 1025 are connected in series in sequence, and the third activation function layer 1026 uses the Sigmoid function. For each attention unit 102, the input feature image is processed by the attention unit 102 to obtain a first feature image, a second feature image, and an attention weight after being processed by the attention unit 102; wherein, the first feature image output by the previous attention unit 102 serves as the input of the next attention unit 102. Exemplarily, both the first activation function layer 1022 and the second activation function layer 1024 may be SiLU (Sigmoid Gated Linear Unit) functions.

[0046] Exemplarily, if the internal receptive field size of an attention unit 102 is 7×7, the more attention units 102 are stacked, the larger the long-range attention receptive field (such as 2n + 1). Exemplarily, the first convolutional layer 1021, the second convolutional layer 1023, and the third convolutional layer 1025 may respectively select a convolutional kernel with a size of 3×3. It should be noted that the attention unit 102 uses the Sigmoid function to generate the attention weight, making the attention unit 102 a parameter-free attention network, that is, the image processing model adopts a fast parameter-free attention mechanism, making the attention unit 102 a lighter attention network. It does not require additional parameters to learn the attention weight, but determines it through a fixed calculation rule. Such a mechanism can reduce the complexity and computational cost of the model while still being able to effectively focus on important specific parts of the image.

[0047] The embodiment of the present disclosure also provides a training method for the image processing model described in the above embodiment. Figure 2 is a flowchart of the training method for the image processing model provided by the embodiment of the present disclosure. It should be noted that for the description of the network structure of this image processing model, reference can be made to the relevant descriptions of the image processing model in the above Figure 1A and Figure 1B shown embodiments, which will not be elaborated here.

[0048] As Figure 2 shown, the training method may include but is not limited to the following steps.

[0049] In step 201, a training set image and a label image corresponding to the training set image are obtained.

[0050] Exemplarily, training set images and label images corresponding to the training set images can be obtained from the publicly available DIV2K dataset. In an embodiment of the present disclosure, the resolution of the training set images is less than the resolution of the label images. That is to say, the label images can be high-resolution images corresponding to the training set images, which is convenient for training an image processing model using the training set images and their corresponding label images, and learning the mapping relationship from low resolution to high resolution by training the image processing model, so that the image processing model has super-resolution capabilities.

[0051] In step 202, super-resolution reconstruction is performed on the training set images based on the initial image processing model to obtain super-resolved images.

[0052] In some embodiments, the training set images are input into the initial image processing model. An initial feature image is obtained through a first convolutional unit, and the initial feature image is processed through a plurality of attention units and a second convolutional unit to obtain a first shallow feature image, a second shallow feature image, a first deep feature image, and a second deep feature image. The first shallow feature image, the second shallow feature image, the first deep feature image, and the second deep feature image are concatenated through a Concat unit to obtain a fused feature image, and the fused feature image is sequentially processed through a third convolutional unit and an upsampling unit, so that the super-resolved images output by the image processing model can be obtained.

[0053] In step 203, a loss value is generated according to the super-resolved images and the label images.

[0054] In some embodiments, a preset loss function can be used to calculate the loss value according to the super-resolved images and the label images. Exemplarily, the loss function can be an L1 norm loss function, and the loss function can be expressed as follows:

[0055] Loss = |y - gt|(1)

[0056] Where Loss is the loss function, y is the super-resolved image output by the image processing model, gt is the label image of the input image of the image processing model, that is, the high-resolution image corresponding to the input image, and || is the absolute value function. It should be noted that the loss function can be in other forms, and the present disclosure does not make any limitations in this regard and will not elaborate further.

[0057] In step 204, the image processing model is trained according to the loss value.

[0058] In some embodiments, the parameters of the image processing model can be adjusted according to the loss value and the backpropagation gradient to obtain the trained image processing model. Exemplarily, the gradient required by the backpropagation algorithm is calculated based on the loss value, and the gradient is used for backpropagation to update the parameters of the image processing model until the loss value is less than or equal to the preset value or the number of model iterations reaches the preset number, and the iterative training is stopped, thereby obtaining the trained image processing model.

[0059] Through the embodiments of the present disclosure, the model structure of the image processing model is improved, and a plurality of attention units are added to the model. The plurality of attention units are stacked in series, and the Sigmoid function is used for calculation in each attention unit, so that the attention network does not need to set parameters, simplifies the model structure, thereby reducing the complexity and computational cost of the model, greatly improving the running speed of the image processing model, and thus solving the problems of large computing power, high power consumption, and high latency of the image processing model in the related art. Since the attention mechanism is added, the model can still effectively focus on the important features of the image, and the image quality effect in the super-resolution task can be improved.

[0060] Figure 3 is a flowchart of the training method of the image processing model provided by the embodiments of the present disclosure. It should be noted that for the description of the network structure of the image processing model, reference can be made to the relevant descriptions of the image processing model in the embodiments described above Figure 1A and Figure 1B shown, and details are not described herein again. As Figure 3 shown, the training method may include but is not limited to the following steps.

[0061] In step 301, a training set image and a label image corresponding to the training set image are obtained.

[0062] For the optional implementation manner of step 301, reference can be made to Figure 2 the optional implementation manner of step 201 in Figure 2 and other related parts in the embodiments involved, and details are not described herein again.

[0063] In step 302, the training set image is input into the initial image processing model.

[0064] In step 303, the training set image is subjected to a convolution operation through the first convolution unit to obtain an initial feature image of the training set image on the first convolution unit.

[0065] In step 304, the initial feature image is processed through a plurality of attention units and the second convolution unit to obtain a first shallow feature image, a second shallow feature image, a first deep feature image, and a second deep feature image.

[0066] In some embodiments, the first feature image output by the first attention unit among the multiple attention units can be multiplied pointwise with the attention weights output by the second attention unit among the multiple attention units to obtain a first shallow feature image; the initial feature image is multiplied pointwise with the attention weights output by the third attention unit among the multiple attention units to obtain a second shallow feature image; the second feature image output by the second attention unit is determined as the first deep feature image, and the feature image obtained after being processed by the multiple attention units and the second convolutional unit in sequence is determined as the second deep feature image. Among them, the first attention unit is the first attention unit among the multiple attention units, the second attention unit is the penultimate attention unit among the multiple attention units, and the third attention unit is the last attention unit among the multiple attention units.

[0067] In an alternative implementation, the initial feature image is sequentially processed by each of the multiple attention units to obtain a first feature image after being processed by the multiple attention units. The first feature image output by the first attention unit among the multiple attention units is multiplied pointwise with the attention weights output by the second attention unit among the multiple attention units to obtain a first shallow feature image. The initial feature image is multiplied pointwise with the attention weights output by the third attention unit among the multiple attention units to obtain a second shallow feature image. The second feature image output by the second attention unit is determined as the first deep feature image. The first feature image after being processed by the multiple attention units is subjected to a convolution operation through the second convolutional unit to obtain a second deep feature image.

[0068] Exemplarily, Figure 1AThe network structure of the illustrated image processing model. Among the multiple attention units 102, there are 6 attention units 102. Taking the series stacking of 6 attention units 102 as an example, the initial feature image output by the first convolutional unit is used as the input of the attention network composed of these 6 series-stacked attention units 102. That is, the initial feature image is sequentially processed through each of these 6 attention units to obtain the first feature image after being processed by these 6 attention units. Here, the first feature image is the feature image output by the 6th attention unit. Multiply the first feature image output by the first attention unit among these 6 attention units with the attention weight output by the penultimate attention unit (i.e., the 5th attention unit) among these 6 attention units to obtain the first shallow feature image. Multiply the initial feature image with the attention weight output by the last attention unit (i.e., the 6th attention unit) among these 6 attention units to obtain the second shallow feature image. Determine the second feature image output by the 5th attention unit as the first deep feature image. Perform a convolution operation on the first feature image (i.e., the first feature image output by the 6th attention unit) after being processed by 6 attention units 102 through the second convolutional unit to obtain the second deep feature image.

[0069] In some embodiments, for each attention unit, the input feature image can be processed through this attention unit to obtain the first feature image, the second feature image, and the attention weight after being processed by this attention unit. Among them, the first feature image output by the previous attention unit serves as the input of the next attention unit.

[0070] In a possible implementation manner, taking one attention unit as an example, the optional implementation manners for processing the input feature image through this attention unit to obtain the first feature image, the second feature image, and the attention weight after being processed by this attention unit include: sequentially processing the input feature image through the first convolutional layer, the first activation function layer, the second convolutional layer, the second activation function layer, and the third convolutional layer to obtain an intermediate feature image; performing a fusion process on the intermediate feature image and the input feature image to obtain a fusion feature; processing the intermediate feature image through the third activation function layer to obtain the attention weight after being processed by the attention unit; performing a dot product operation on the fusion feature and the attention weight after being processed by the attention unit to obtain the first feature image after being processed by the attention unit; performing a convolution operation on the input feature image through the first convolutional layer to obtain the second feature image after being processed by the attention unit.

[0071] Exemplarily, Figure 1BTaking the network structure of the attention unit shown as an example, the feature image input to the attention unit can be processed successively through a first convolutional layer, a first activation function layer, a second convolutional layer, a second activation function layer, and a third convolutional layer to obtain an intermediate feature image. The intermediate feature image and the feature image input to the attention unit are fused to obtain a fused feature. The intermediate feature image is processed through a third activation function layer to obtain an attention weight after being processed by the attention unit. The fused feature is multiplied pointwise with the attention weight after being processed by the attention unit to obtain a first feature image after being processed by the attention unit. The feature image input to the attention unit is subjected to a convolution operation through a first convolutional layer to obtain a second feature image after being processed by the attention unit.

[0072] Exemplarily, the internal receptive field size of an attention unit 102 is 7×7. Then, the more attention units 102 are stacked, the larger the long-range attention receptive field is (such as 2n + 1). In addition, the embodiments of the present disclosure implement long-range attention feature fusion through point multiplication, which can improve the running speed of the model.

[0073] In step 305, the first shallow feature image, the second shallow feature image, the first deep feature image, and the second deep feature image are concatenated through a Concat unit to obtain a fused feature image.

[0074] Exemplarily, the first shallow feature image, the second shallow feature image, the first deep feature image, and the second deep feature image are concatenated along the channel dimension through a Concat unit to obtain a concatenated fused feature image. It should be noted that the first shallow feature image, the second shallow feature image, the first deep feature image, and the second deep feature image can also be concatenated according to other dimensions. The present disclosure does not make any limitations in this regard and will not elaborate further.

[0075] In step 306, after the fused feature image is subjected to a convolution operation through a third convolutional unit, it is subjected to an upsampling operation through an upsampling unit to obtain a super-resolved image.

[0076] Exemplarily, after the fused feature image is subjected to a convolution operation through a third convolutional unit, the feature image output by the third convolutional unit is used as the input of the upsampling unit, and the input is subjected to an upsampling operation through the upsampling unit to obtain a super-resolved image.

[0077] In step 307, a loss value is generated based on the super-resolved image and the label image.

[0078] For the optional implementation of step 307, reference can be made to Figure 2 the optional implementation of step 203 in Figure 2For other related parts in the embodiments involved, they will not be elaborated here.

[0079] In step 308, the image processing model is trained according to the loss value.

[0080] For the optional implementation of step 308, reference can be made to Figure 2 the optional implementation of step 204 in Figure 2 For other related parts in the embodiments involved, they will not be elaborated here.

[0081] Optionally, in some embodiments, for the training stage of the image processing model, the reparameterization technique is adopted to expand each convolution in at least part of the convolutions in the image processing model into multiple convolution structures; after the image processing model is trained, the multiple convolution structures expanded for each convolution are merged to obtain at least part of the trained convolutions.

[0082] Exemplarily, in the model training stage, the reparameterization technique is adopted to perform reparameterization processing on the convolutional layers in the image processing model. Exemplarily, all convolutional layers in the image processing model are expanded, which can be more easily fitted and improve the training effect; after training, each expanded convolutional layer is merged into one convolutional layer. For example, as Figure 4 shown, taking the convolutional kernel size of a convolutional layer as 3×3 as an example, in the model training stage, the reparameterization technique is adopted to perform reparameterization processing on the convolutional kernel, that is, after the reparameterization processing, this convolutional layer becomes composed of 4 convolutions, and after training, these 4 convolutions are merged into one convolutional layer. Thus, in the model training stage, by adopting the reparameterization technique to perform reparameterization processing on the convolutional layer, the training effect can be improved; after the model is trained, the expanded convolutional layers are merged into the corresponding convolutional layers, which can simplify the model structure while ensuring the model effect.

[0083] By implementing the embodiments of the present disclosure, by adopting the fast parameter-free attention mechanism, the model structure can be simplified. In addition, by implementing long-range attention feature fusion through dot multiplication, the model running speed can be improved. In the model training stage, the reparameterization technique is adopted to perform reparameterization processing on the convolutional layers in the model, that is: all convolutional layers are expanded, which can be more easily fitted and improve the training effect; after training, each expanded convolutional layer is merged into one convolutional layer, which can simplify the model structure while ensuring the model effect.

[0084] The embodiments of the present disclosure also provide an image processing method. Figure 5 It is a flowchart of the image processing method provided by the embodiments of the present disclosure. As Figure 5 shown, the image processing method may include but is not limited to the following steps.

[0085] In step 501, an image to be processed is obtained.

[0086] In step 502, the image to be processed is input into a preset image processing model to obtain a target image corresponding to the image to be processed.

[0087] In some embodiments, the resolution of the target image is greater than that of the image to be processed. Among them, in some embodiments, the image processing model may be an image super-resolution model, and the image processing model may be trained by the training method shown in any of the foregoing embodiments.

[0088] Figure 6 It is a block diagram of a training device for an image processing model provided by an embodiment of the present disclosure. It should be noted that for the description of the network structure of the image processing model, reference may be made to the relevant description of the image processing model in the foregoing Figure 1A and Figure 1B shown embodiments, which will not be elaborated here. As Figure 6 shown, the training device may include: an acquisition module 601, a super-resolution reconstruction module 602, a generation module 603, and a training module 604.

[0089] Among them, the acquisition module 601 is used to acquire a training set image and a label image corresponding to the training set image, and the resolution of the training set image is less than that of the label image.

[0090] The super-resolution reconstruction module 602 is used to perform super-resolution reconstruction on the training set image based on an initial image processing model to obtain a super-resolved image.

[0091] The generation module 603 is used to generate a loss value according to the super-resolved image and the label image.

[0092] The training module 604 is used to train the image processing model according to the loss value.

[0093] In some embodiments, the super-resolution reconstruction module 602 is specifically used for: inputting the training set image into the initial image processing model; performing a convolution operation on the training set image through a first convolution unit to obtain an initial feature image of the training set image on the first convolution unit; processing the initial feature image through a plurality of attention units and a second convolution unit to obtain a first shallow feature image, a second shallow feature image, a first deep feature image, and a second deep feature image; splicing the first shallow feature image, the second shallow feature image, the first deep feature image, and the second deep feature image through a Concat unit to obtain a fused feature image; performing a convolution operation on the fused feature image through a third convolution unit and then performing an upsampling operation through an upsampling unit to obtain a super-resolved image.

[0094] In some embodiments, the super-resolution reconstruction module 602 is specifically configured to: process the initial feature image through each of a plurality of attention units in sequence to obtain a first feature image after being processed by the plurality of attention units; perform a dot product of the first feature image output by the first attention unit among the plurality of attention units and the attention weight output by the second attention unit among the plurality of attention units to obtain a first shallow feature image; wherein, the first attention unit is the first attention unit among the plurality of attention units, and the second attention unit is the second-to-last attention unit among the plurality of attention units; perform a dot product of the initial feature image and the attention weight output by the third attention unit among the plurality of attention units to obtain a second shallow feature image; wherein, the third attention unit is the last attention unit among the plurality of attention units; determine the second feature image output by the second attention unit as the first deep feature image; perform a convolution operation on the first feature image after being processed by the plurality of attention units through the second convolution unit to obtain a second deep feature image.

[0095] In some embodiments, each attention unit includes a first convolutional layer, a first activation function layer, a second convolutional layer, a second activation function layer, a third convolutional layer, and a third activation function layer. The first convolutional layer, the first activation function layer, the second convolutional layer, the second activation function layer, and the third convolutional layer are connected in series in sequence, and the third activation function layer uses a Sigmoid function; wherein, for each attention unit, the input feature image is processed through the attention unit to obtain a first feature image, a second feature image, and an attention weight after being processed by the attention unit; wherein, the first feature image output by the previous attention unit serves as the input of the next attention unit.

[0096] In some embodiments, optional implementation manners of processing the input feature image through the attention unit to obtain a first feature image, a second feature image, and an attention weight after being processed by the attention unit include: processing the input feature image through the first convolutional layer, the first activation function layer, the second convolutional layer, the second activation function layer, and the third convolutional layer in sequence to obtain an intermediate feature image; performing a fusion process on the intermediate feature image and the input feature image to obtain a fused feature; processing the intermediate feature image through the third activation function layer to obtain the attention weight after being processed by the attention unit; performing a dot product operation on the fused feature and the attention weight after being processed by the attention unit to obtain the first feature image after being processed by the attention unit; performing a convolution operation on the input feature image through the first convolutional layer to obtain the second feature image after being processed by the attention unit.

[0097] In some embodiments, the super-resolution reconstruction module 602 is further configured to: during the training phase of the image processing model, adopt the reparameterization technique to expand each convolution in at least part of the convolutions in the image processing model into a plurality of convolution structures; after the image processing model is trained, merge the plurality of convolution structures expanded from each convolution to obtain at least part of the trained convolutions.

[0098] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0099] Figure 7 It is a block diagram of the image processing device provided by an embodiment of the present disclosure. As Figure 7 shown, the image processing device may include: an acquisition module 701, a prediction module 702. Among them, the acquisition module 701 is configured to acquire an image to be processed; the prediction module 702 is configured to input the image to be processed into a preset image processing model to obtain a target image corresponding to the image to be processed, and the resolution of the target image is greater than the resolution of the image to be processed; wherein, the image processing model is trained by the training method of any of the foregoing embodiments.

[0100] As Figure 8 shown, it is a block diagram of an electronic device 800 according to an embodiment of the present disclosure. The electronic device 800 is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device 800 may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present application described and / or claimed herein.

[0101] As Figure 8As shown, the electronic device 800 includes: one or more processors 801, a memory 802, and interfaces for connecting the various components, including a high-speed interface and a low-speed interface. The various components are interconnected using different buses and can be mounted on a common motherboard or otherwise mounted as required. The processor can process instructions executed within the electronic device 800, including instructions stored in the memory or on the memory for graphical information to display a GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, if necessary, multiple processors and / or multiple buses can be used with multiple memories and multiple memories. Similarly, multiple electronic devices 800 can be connected, with each device providing part of the necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Figure 8 Here, a processor 801 is taken as an example.

[0102] The memory 802 is the non-transitory computer-readable storage medium provided by the present disclosure. Among them, the memory stores instructions executable by at least one processor, so that the at least one processor executes the training method and / or the image processing method of the image processing model provided by the present disclosure. The non-transitory computer-readable storage medium of the present disclosure stores computer instructions, and the computer instructions are used to cause a computer to execute the training method and / or the image processing method of the image processing model provided by the present disclosure.

[0103] As a non-transitory computer-readable storage medium, the memory 802 can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as program instructions / modules corresponding to the training method and / or the image processing method of the image processing model in the embodiments of the present disclosure. The processor 801 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 802, that is, implements the training method and / or the image processing method in the above method embodiments.

[0104] The memory 802 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and applications required for at least one function; the data storage area can store data created according to the use of the electronic device 800, etc. In addition, the memory 802 can include a high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 802 optionally includes a memory remotely provided relative to the processor 801, and these remote memories can be connected to the electronic device 800 through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0105] The electronic device 800 may further include: an input device 803 and an output device 804. The processor 801, the memory 802, the input device 803, and the output device 804 may be connected via a bus or other means. Figure 8 Here, the example of connection via a bus is taken.

[0106] The input device 803 can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the electronic device 800, such as input devices like touchscreens, keypads, mice, trackpads, touchpads, pointing sticks, one or more mouse buttons, trackballs, joysticks, etc. The output device 804 may include a display device, an auxiliary lighting device (e.g., an LED), and a haptic feedback device (e.g., a vibration motor), etc. The display device may include, but is not limited to, a liquid crystal display (LCD), a light-emitting diode (LED) display, and a plasma display. In some embodiments, the display device may be a touchscreen.

[0107] Various embodiments of the systems and techniques described herein may be implemented in digital electronic circuit systems, integrated circuit systems, dedicated ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0108] These computing programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor, and these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., a disk, an optical disc, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.

[0109] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0110] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), the Internet, and blockchain network.

[0111] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with blockchain.

[0112] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 802 including instructions, which can be executed by a processor 801 of an electronic device 800 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0113] Figure 9It is a schematic structural diagram of a chip 900 proposed by an embodiment of the present disclosure. For the case where the electronic device can be a chip or a chip system, reference can be made to Figure 9 the schematic structural diagram of the chip 900 shown, but not limited thereto. As Figure 9 shown, the chip 900 may include one or more processors 901. The chip 900 is used to execute any of the above methods.

[0114] In some embodiments, the chip 900 further includes one or more interface circuits 902. Optionally, terms such as interface circuit, interface, and transceiver pin can be replaced with each other. In some embodiments, the chip 900 further includes one or more memories 903 for storing data. Optionally, all or part of the memories 903 may be outside the chip 900. Optionally, the interface circuit 902 is connected to the memory 903, and the interface circuit 902 can be used to receive data from the memory 903 or other devices, and the interface circuit 902 can be used to send data to the memory 903 or other devices. For example, the interface circuit 902 can read the data stored in the memory 903 and send the data to the processor 901.

[0115] In some embodiments, the interface circuit 902 executes at least one of the communication steps such as sending and / or receiving in the above method. The interface circuit 902 executing the communication steps such as sending and / or receiving in the above method means, for example, that the interface circuit 902 executes data interaction between the processor 901, the chip 900, the memory 903, or the transceiver device. In some embodiments, the processor 901 executes at least one of the other steps.

[0116] In various embodiments such as virtual devices, physical devices, and chips, the various modules and / or devices described can be combined or separated arbitrarily according to the situation. Optionally, some or all of the steps can also be executed collaboratively by multiple modules and / or devices, which is not limited herein.

[0117] The present disclosure also proposes a program product. When the above program product is executed by an electronic device 800, the electronic device 800 is caused to execute any of the above methods. Optionally, the above program product is a computer program product.

[0118] The present disclosure also proposes a computer program. When it runs on a computer, the computer is caused to execute any of the above methods.

[0119] Other embodiments of the present invention will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention following the general principles of the invention and including known common general knowledge or conventional technical means in the technical field not disclosed in this disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the invention are pointed out by the following claims.

[0120] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A training method for an image processing model, characterized in that The training method includes: Obtaining a training set of images and the label images corresponding to the training set of images, wherein the resolution of the training set of images is less than the resolution of the label images; Performing super-resolution reconstruction on the training set of images based on an initial image processing model to obtain a super-resolved image; Generating a loss value according to the super-resolved image and the label image; Training the image processing model according to the loss value.

2. The method according to claim 1, characterized in that, The image processing model includes a first convolutional unit, a plurality of attention units, a second convolutional unit, a connection Concat unit, a third convolutional unit, and an upsampling unit. The plurality of attention units are stacked in series. The first convolutional unit, the plurality of attention units, the second convolutional unit, the Concat unit, the third convolutional unit, and the upsampling unit are connected in series in sequence. In each of the attention units, a sigmoid function is used to generate an attention weight. The performing super-resolution reconstruction on the training set of images based on the initial image processing model to obtain a super-resolved image includes: Inputting the training set of images into the initial image processing model; Performing a convolution operation on the training set of images through the first convolutional unit to obtain an initial feature image of the training set of images on the first convolutional unit; Processing the initial feature image through the plurality of attention units and the second convolutional unit to obtain a first shallow feature image, a second shallow feature image, a first deep feature image, and a second deep feature image; Splicing the first shallow feature image, the second shallow feature image, the first deep feature image, and the second deep feature image through the Concat unit to obtain a fused feature image; Performing a convolution operation on the fused feature image through the third convolutional unit and then performing an upsampling operation through the upsampling unit to obtain the super-resolved image.

3. The method according to claim 2, wherein The processing the initial feature image through the plurality of attention units and the second convolutional unit to obtain a first shallow feature image, a second shallow feature image, a first deep feature image, and a second deep feature image includes: Sequentially processing the initial feature image through each of the plurality of attention units to obtain a first feature image after being processed by the plurality of attention units; Performing a dot product on the first feature image output by the first attention unit among the plurality of attention units and the attention weight output by the second attention unit among the plurality of attention units to obtain the first shallow feature image; wherein, the first attention unit is the first attention unit among the plurality of attention units, and the second attention unit is the second-to-last attention unit among the plurality of attention units; Performing a dot product on the initial feature image and the attention weight output by the third attention unit among the plurality of attention units to obtain the second shallow feature image; wherein, the third attention unit is the last attention unit among the plurality of attention units; Determine the second feature image output by the second attention unit as the first deep feature image; Perform a convolution operation on the first feature image processed by the multiple attention units through the second convolution unit to obtain the second deep feature image.

4. The method according to claim 2 or 3, characterized in that, Each of the attention units includes a first convolutional layer, a first activation function layer, a second convolutional layer, a second activation function layer, a third convolutional layer, and a third activation function layer. The first convolutional layer, the first activation function layer, the second convolutional layer, the second activation function layer, and the third convolutional layer are connected in series in sequence, and the third activation function layer uses the Sigmoid function; wherein, For each of the attention units, process the input feature image through the attention unit to obtain the first feature image, the second feature image, and the attention weight processed by the attention unit; wherein, the first feature image output by the previous attention unit serves as the input of the next attention unit.

5. The method according to claim 4, wherein The processing of the input feature image through the attention unit to obtain the first feature image, the second feature image, and the attention weight processed by the attention unit includes: Process the input feature image sequentially through the first convolutional layer, the first activation function layer, the second convolutional layer, the second activation function layer, and the third convolutional layer to obtain an intermediate feature image; Perform a fusion process on the intermediate feature image and the input feature image to obtain a fused feature; Process the intermediate feature image through the third activation function layer to obtain the attention weight processed by the attention unit; Perform a dot product operation on the fused feature and the attention weight processed by the attention unit to obtain the first feature image processed by the attention unit; Perform a convolution operation on the input feature image through the first convolutional layer to obtain the second feature image processed by the attention unit.

6. The method according to claim 1, wherein The method further includes: For the training stage of the image processing model, adopt the reparameterization technique to expand each convolution in at least part of the convolutions in the image processing model into multiple convolution structures; After the image processing model is trained, merge the multiple convolution structures expanded by each convolution to obtain the at least part of the convolutions that have been trained.

7. An image processing method, characterized in that Includes: Obtain an image to be processed; Input the image to be processed into a preset image processing model to obtain a target image corresponding to the image to be processed, and the resolution of the target image is greater than the resolution of the image to be processed; wherein, the image processing model is trained by the method according to any one of claims 1 to 6.

8. A training device for an image processing model, characterized in that, The image processing model includes a first convolutional unit, a plurality of attention units, a second convolutional unit, a Concat unit, a third convolutional unit, and an upsampling unit. The plurality of attention units are stacked in series. The first convolutional unit, the plurality of attention units, the second convolutional unit, the Concat unit, the third convolutional unit, and the upsampling unit are connected in series in sequence. In each of the attention units, a sigmoid function is used to generate attention weights. The training device includes: An acquisition module, configured to acquire training set images and label images corresponding to the training set images, wherein the resolution of the training set images is less than the resolution of the label images; A super-resolution reconstruction module, configured to perform super-resolution reconstruction on the training set images based on an initial image processing model to obtain super-resolved images; A generation module, configured to generate a loss value according to the super-resolved images and the label images; A training module, configured to train the image processing model according to the loss value.

9. An image processing apparatus, characterized in that, Including: An acquisition module, configured to acquire an image to be processed; A prediction module, configured to input the image to be processed into a preset image processing model to obtain a target image corresponding to the image to be processed, wherein the resolution of the target image is greater than the resolution of the image to be processed; wherein, the image processing model is trained by the method according to any one of claims 1 to 6.

10. An electronic device, characterized in that, Including: One or more processors; Wherein, the electronic device is configured to execute the training method of the image processing model according to any one of claims 1 to 6, or execute the image processing method according to claim 7.

11. A chip, characterized in that, Including: One or more processors; Wherein, the processor is configured to call instructions to cause the chip to execute the training method of the image processing model according to any one of claims 1 to 6, or execute the image processing method according to claim 7.

12. A storage medium storing instructions, characterized in that, When the instructions run on the electronic device, the electronic device is caused to execute the training method of the image processing model according to any one of claims 1 to 6, or execute the image processing method according to claim 7.

13. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 6, or implements the steps of the method according to claim 7.