Image processing method, electronic equipment and computer readable storage medium
By using convolutional kinetics of different shapes to perform feature extraction on low-resolution images, the problem of poor image details and edge recovery quality in the prior art is solved, and clear recovery of high-resolution images is achieved.
Patent Information
- Application Number
- CN202510486541.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
AI Technical Summary
In existing methods of reconstructing low-resolution images into high-resolution images, the quality of details and edge recovery is poor, resulting in the restoration of details and edges remains blurry.
Convolution kernels of different shapes (different sizes) are used to extract the image features of low-resolution images, and high-resolution images are generated using the first convolution kernel and the second convolution kernel to experience more context information through convolution kernels of different shapes, enhancing the details and edge recovery quality of the image.
Improves the detail and edge recovery quality of high-resolution images, enhances image recovery performance, and achieves clearer image recovery.
Smart Images

Figure CN120335688A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and particularly relates to an image processing method, an electronic device, and a computer-readable storage medium. Background Art
[0002] Image resolution is an important indicator to measure image quality, which describes the amount of information stored in the image and the clarity of the image. Specifically, image resolution refers to the number of pixels displayed per unit length (usually in inches) in the image, and the unit is PPI (Pixels Per Inch). The higher the resolution, the more pixels the image contains, the more details the image contains, and the higher the image quality.
[0003] Higher-resolution images can retain finer details, such as tiny lesions in medical images and facial features in video surveillance. Reconstructing a low-resolution image into a high-resolution image can restore the details lost in the low-resolution image, such as restoring tiny lesions in medical images and restoring facial features in video surveillance.
[0004] However, in the existing methods for reconstructing a lower-resolution image into a higher-resolution image, there is a problem of poor restoration quality of details and edges. For example, the restored details and edges are still relatively blurred. Summary of the Invention
[0005] In view of this, embodiments of this application provide an image processing method, an electronic device, and a computer-readable storage medium, which can improve the restoration quality of details and edges of a higher-resolution image and improve the performance of image restoration when reconstructing a low-resolution image into a high-resolution image.
[0006] In a first aspect, an embodiment of this application provides an image processing method, including: obtaining a first image to be processed; where the first image has a first resolution; determining a first image feature of the first image; where the first image feature is used to characterize the feature of the image content in the first image; extracting features from the first image feature by using a first convolution kernel and a second convolution kernel to obtain a feature extraction result; where the first convolution kernel includes a K*K convolution kernel, and the second convolution kernel includes an M*N convolution kernel, and M and N are not equal; generating a second image based on the feature extraction result; where the second image has a second resolution, and the second resolution is greater than the first resolution.
[0007] Second aspect, an embodiment of the present application provides a model training method, including: obtaining training samples; wherein, the training samples include sample images and target images; inputting the sample images into an image resolution enhancement model to obtain predicted images; determining a loss value based on the difference between the predicted images and the target images; training the image resolution enhancement model based on the loss value to obtain a trained image resolution enhancement model.
[0008] Third aspect, an embodiment of the present application provides an image processing apparatus, including: an obtaining module, configured to obtain a first image to be processed; wherein, the first image has a first resolution. A determining module, configured to determine a first image feature of the first image; wherein, the first image feature is used to characterize the feature of the image content in the first image. A feature extraction module, configured to perform feature extraction on the first image feature by using a first convolution kernel and a second convolution kernel to obtain a feature extraction result. An image generation module, configured to generate a second image based on the feature extraction result; wherein, the second image has a second resolution, and the second resolution is greater than the first resolution.
[0009] Fourth aspect, an embodiment of the present application provides a model training apparatus, including: an obtaining module, configured to obtain training samples; wherein, the training samples include sample images and target images. A prediction module, configured to input the sample images into an image resolution enhancement model to obtain predicted images. A determining module, configured to determine a loss value based on the difference between the predicted images and the target images. A training module, configured to train the image resolution enhancement model based on the loss value to obtain a trained image resolution enhancement model.
[0010] Fifth aspect, an embodiment of the present application provides an electronic device, including: a processor; a memory for storing processor-executable instructions, wherein the processor is configured to execute the image processing method in the first aspect or the model training method in the second aspect above.
[0011] Sixth aspect, an embodiment of the present application provides a computer-readable storage medium, where the storage medium stores a computer program, and the computer program is used to execute the image processing method in the first aspect or the model training method in the second aspect above.
[0012] Seventh aspect, an embodiment of the present application provides a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor of a computer device, the computer device is enabled to execute the image processing method in the first aspect or the model training method in the second aspect above.
[0013] The embodiments of the present application provide an image processing method, an electronic device, and a computer-readable storage medium. By using convolution kernels of different shapes (different sizes) to extract image features from a lower-resolution image, more image features representing context information can be obtained. These more context-information-representing image features are used to enhance the details and edges of the image, improve the restoration quality of the details and edges of a higher-resolution image, and enhance the performance of image restoration. Description of the Drawings
[0014] Figure 1 Shown is a schematic diagram of the change of a mobile phone interface provided by an exemplary embodiment of the present application.
[0015] Figure 2 Shown is a schematic flowchart of an image processing method provided by an exemplary embodiment of the present application.
[0016] Figure 3 Shown is a schematic diagram of using an image resolution enhancement model to enhance the image resolution provided by an exemplary embodiment of the present application.
[0017] Figure 4 Shown is a schematic diagram of using an image resolution enhancement model to enhance the image resolution provided by another exemplary embodiment of the present application.
[0018] Figure 5 Shown is a schematic diagram of a normal convolution kernel and a convolution kernel with a dilation rate of 2 provided by an exemplary embodiment of the present application.
[0019] Figure 6 Shown is provided by an exemplary embodiment of the present application regarding Figure 4 a schematic diagram of the deep feature extraction sub-module in the middle and deep feature extraction module.
[0020] Figure 7 Shown is provided by an exemplary embodiment of the present application regarding Figure 4 a schematic diagram of the structure of the middle and deep feature extraction module.
[0021] Figure 8 Shown is a schematic flowchart of an image processing method provided by another exemplary embodiment of the present application.
[0022] Figure 9 Shown is a schematic flowchart of a model training method for an image resolution enhancement model provided by an exemplary embodiment of the present application.
[0023] Figure 10 Shown is a schematic diagram of the structure of an image processing device provided by an exemplary embodiment of the present application.
[0024] Figure 11 Shown is a schematic diagram of the structure of a model training device provided by an exemplary embodiment of the present application.
[0025] Figure 12 The block diagram of an electronic device 1100 for executing an image processing method or a model training method provided by an exemplary embodiment of the present application is shown. Detailed implementation manners
[0026] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0027] Application overview
[0028] As described in the background art above, in the existing methods for reconstructing a lower-resolution image into a higher-resolution image, there are problems with poor restoration quality of details and edges. For example, the restored details and edges are still relatively blurred.
[0029] To solve the above technical problems, an embodiment of the present application provides an image processing method, which includes: extracting image features of a lower-resolution image by using a first convolution kernel and a second convolution kernel to obtain a feature extraction result, and converting the lower-resolution image into a higher-resolution image based on the feature extraction result. Wherein, the first convolution kernel includes a K*K convolution kernel, that is, the first convolution kernel is a square convolution kernel, and the second convolution kernel includes an M*N convolution kernel, where M and N are not equal, that is, the second convolution kernel is a rectangular convolution kernel.
[0030] It can be understood that convolution kernels of different shapes (different sizes) correspond to different ranges of receptive fields and can sense more context information. Here, the context information refers to the correlation relationship between the value of a pixel or the pixel values of a region and the values of surrounding pixels or regions in an image or image features. In the embodiments of the present application, by extracting image features of a lower-resolution image through convolution kernels of different shapes (different sizes), more image features representing context information can be obtained. The more image features representing context information are used to enhance the details and edges of the image, improve the restoration quality of the details and edges of the higher-resolution image, and improve the performance of image restoration.
[0031] Exemplary scenario
[0032] The image processing method provided by the embodiments of the present application can be applied to an electronic device, which can be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, and a personal digital assistant (PDA), an augmented reality (AR) / virtual reality (VR) device, a medical imaging device, a monitoring device, a satellite imaging device, and other electronic devices with image processing capabilities. The embodiments of the present application do not impose special restrictions on the specific form of the electronic device.
[0033] Taking the electronic device as a mobile phone as an example, an exemplary description is given for the scenario of adjusting the image quality of pictures in the gallery (also known as the album). Figure 1 Shown is a schematic diagram of the change of the mobile phone interface provided by an exemplary embodiment of the present application. As Figure 1 shown in (a), the mobile phone 100 displays the interface 101, and the interface 101 includes a gallery icon. In response to the user clicking on the gallery icon, the mobile phone 100 displays the gallery interface 102 as shown in Figure 1 shown in (b). The gallery interface 102 includes multiple pictures.
[0034] If the user wants to adjust the image quality of a certain image in the gallery, the mobile phone 100 can be used to perform corresponding editing operations on the image. Specifically, for example, as Figure 1 shown in (b), the gallery interface 102 includes a thumbnail A1, and the content of the thumbnail A1 is the wings of a butterfly. In response to the user's click operation on the thumbnail A1, the mobile phone 100 displays an enlarged view of the thumbnail A1. For example, as Figure 1 shown in (c), the interface 103 includes an enlarged view of the thumbnail A1: image A2. The interface 103 includes editing controls. In response to the user's click operation on the editing controls, the mobile phone 100 displays the interface 104 as shown in Figure 1 shown in (d). The interface 104 includes controls corresponding to various editing function options, such as one-key beautification, cropping and rotation, filters, enhancement, graffiti, and other controls. In response to the user's click operation on the enhancement control 1041, the mobile phone 100 displays as Figure 1The interface 105 shown in (e) of [the figure]. The interface 105 includes various function options corresponding to enhancement functions, such as controls for sharpness, fading, noise, etc. Among them, the sharpness control 1051 is used to adjust the image resolution to adjust the sharpness of the image. The mobile phone 100 responds to the user's operation of clicking the sharpness control 1051 and displays the sharpness value scale control 1052. The mobile phone 100 responds to the user's sliding of the sharpness value scale control 1052 and adjusts the sharpness of the image A2. For example, as shown in (f) of [the figure], it displays the interface 106, which includes the image A3. The image A3 is a picture after changing the image A2 to the corresponding sharpness value, and the resolution of the image A3 is higher than that of the image A2. Figure 1 As shown in (f) of [the figure], it displays the interface 106, which includes the image A3. The image A3 is a picture after changing the image A2 to the corresponding sharpness value, and the resolution of the image A3 is higher than that of the image A2.
[0035] During the process of the mobile phone 100 changing the image A2 to the image A3, specifically, the mobile phone 100 can use the first convolution kernel and the second convolution kernel to extract the image features of the lower-resolution image A2 to obtain the feature extraction result. Then, based on this feature extraction result, it changes the lower-resolution image A2 to a higher-resolution image A3.
[0036] In addition to the above scenarios, improving the image resolution can also be applied to other scenarios. In another example, it can improve the image resolution in the medical image enhancement scenario. For example, the medical imaging device can perform 4-fold super-resolution processing on the pathological section image. 4-fold super-resolution means 4 times super-resolution (Super-Resolution, SR), which refers to magnifying the resolution of the lower-resolution image by 4 times.
[0037] In another example, it can improve the image resolution in the video surveillance enhancement scenario. For example, the surveillance device can real-time improve the resolution of the video from 1080P (1920 * 1080 pixels) to 4K (3840 * 2160 pixels).
[0038] In another example, it can improve the image resolution in the satellite image processing scenario. For example, it can improve the low-resolution multispectral image (Multispectral Images, MSI) to a high resolution while retaining the spectral information. Among them, the multispectral image usually contains spectral information in multiple bands, and these bands cover multiple spectral ranges from visible light to near-infrared.
[0039] In one example, it can use the image resolution improvement model to improve the image resolution. Specifically, the image resolution improvement model can use the first convolution kernel and the second convolution kernel to extract the image features of the lower-resolution image A2 to obtain the feature extraction result. Then, based on this feature extraction result, it changes the lower-resolution image A2 to a higher-resolution image A3.
[0040] The image processing method can be executed by an electronic device. Further, an image resolution enhancement model can be deployed on the electronic device to execute the image processing method provided in the embodiments of the present application. Further, the electronic device can be communicatively connected to other electronic devices (such as a server), and the other electronic devices are used to execute the image processing method provided in the embodiments of the present application to obtain an image with enhanced resolution, and transmit the image with enhanced resolution, such as a higher-resolution image A3, to the electronic device.
[0041] It should be understood that the above application scenario examples are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited thereto. On the contrary, the embodiments of the present application can be applied to any scenario where they may be applicable.
[0042] Exemplary method
[0043] Figure 2 The flowchart of the image processing method provided in an exemplary embodiment of the present application is shown. Figure 2 The method can be executed by an electronic device, such as Figure 1 the mobile phone 100 in Figure 2 As shown, the image processing method may include the following.
[0044] 210: Obtain a first image to be processed; wherein, the first image has a first resolution.
[0045] In some cases, the electronic device can be triggered to obtain the first image to be processed under the user's operation for image processing of the first image. For example, as Figure 1 shown, the mobile phone 100 responds to the user's sliding of the clarity value scale control 1052 to adjust the clarity of image A2.
[0046] In some cases, it may be that when the first image is collected, the first image to be processed is obtained for image processing of the first image. For example, after a medical imaging device takes a medical image, the medical image can be directly subjected to image processing without any corresponding triggering operation by the user.
[0047] The first image can be an image to be processed in various scenarios. For example, the first image can be a photo taken by a mobile phone, a picture in the mobile phone gallery, a medical image taken by a medical imaging device, each frame of a surveillance video taken by a surveillance device, a satellite image taken by a satellite imaging device, etc., but is not limited thereto.
[0048] 220: Determine the first image feature of the first image; wherein, the first image feature is used to characterize the feature of the image content in the first image.
[0049] An image feature is a tensor that is used to represent the features of the image content in an image. A tensor is a multi-dimensional array. For example, an image feature is a tensor that includes multiple feature maps, where each feature map can be regarded as a "channel" representing the feature representation of the image at different levels. An image feature is the result of feature extraction from an image. The feature extraction method can be a feature extraction method based on deep learning. For example, the feature extraction method can be a feature extraction method based on a Convolutional Neural Network (CNN).
[0050] 230: Feature extraction is performed on the first image feature using a first convolutional kernel and a second convolutional kernel to obtain a feature extraction result.
[0051] The size of the convolutional kernel determines the pixel range covered each time the feature extraction (such as a convolution operation) slides on the image or feature map. Different-shaped (different-sized) convolutional kernels correspond to different ranges of receptive fields, which can sense more context information. More image features representing context information are used to enhance the details and edges of the image, improve the recovery quality of the details and edges of the higher-resolution image, and improve the performance of image restoration.
[0052] The size of the convolutional kernel refers to the width and height of the convolutional kernel, usually represented as size * size. For example, the size of the first convolutional kernel can be represented as K * K, that is, the first convolutional kernel includes a K * K convolutional kernel, and the first convolutional kernel is a square convolutional kernel. The size of the second convolutional kernel can be represented as M * N, that is, the second convolutional kernel includes an M * N convolutional kernel, where M and N are not equal, and M, N, and K are integers greater than 0.
[0053] In one example, the electronic device can perform feature extraction on a part of the first image feature using the first convolutional kernel to obtain a feature extraction result, and perform feature extraction on another part of the first image feature using the second convolutional kernel to obtain another feature extraction result, so as to facilitate the subsequent generation of a second image based on the two feature extraction results.
[0054] 240: Generate a second image based on the feature extraction result; wherein, the second image has a second resolution, and the second resolution is greater than the first resolution.
[0055] The feature extraction result includes more image features representing context information. More image features representing context information are used to enhance the details and edges of the image, and can transform the first image with a lower resolution into a second image with a higher resolution.
[0056] In one example, the second image can be generated through an image resolution enhancement model. For example, Figure 3The following is a schematic diagram of improving the image resolution using an image resolution improvement model provided by an exemplary embodiment of the present application. As Figure 3 shown, by inputting a lower-resolution image, i.e., the first image, into the image resolution improvement model, a higher-resolution image, i.e., the second image, can be output. Among them, the image resolution improvement model can execute the above steps 210 to 240. Further, in one example, Figure 4 The following is a schematic diagram of improving the image resolution using an image resolution improvement model provided by another exemplary embodiment of the present application. As Figure 4 shown, the image resolution improvement model includes a shallow feature extraction module, a deep feature extraction module, a feature fusion module, and a reconstruction module. The deep feature extraction module is used to execute the above steps 220 to 230, and the reconstruction module is used to generate the second image. The shallow feature extraction module and the feature fusion module will be specifically introduced below.
[0057] The embodiment of the present application provides an image processing method. By using convolution kernels of different shapes (different sizes) to extract image features of a lower-resolution image, more image features representing context information can be obtained. The more image features representing context information are used to enhance the details and edges of the image, improve the recovery quality of the details and edges of the higher-resolution image, and improve the performance of image restoration.
[0058] According to an embodiment of the present application, the first image feature includes a first partial image feature and a second partial image feature. Using a first convolution kernel and a second convolution kernel to extract features from the first image feature, the feature extraction result includes: using the first convolution kernel to extract features from the first partial image feature of the first image feature to obtain a first feature extraction result, and using the second convolution kernel to extract features from the second partial image feature of the first image feature to obtain a second feature extraction result. Generating the second image based on the feature extraction result includes: generating the second image based on the first feature extraction result and the second feature extraction result.
[0059] Compared with using the first convolution kernel and the second convolution kernel to extract features from the first image feature respectively, dividing the image feature into two parts and processing these two parts of image features respectively can achieve a lower computational complexity and at the same time obtain a better image recovery quality.
[0060] In one example, the first image feature is divided into channels to obtain a first channel division feature and a second channel division feature. Among them, the first channel division feature is the first partial image feature, and the second channel division feature is the second partial image feature.
[0061] Further, in one example, the first image feature can be divided into channels according to a bisecting ratio to obtain a first partial image feature and a second partial image feature, that is, the first image feature is equally divided into the first partial image feature and the second partial image feature according to channels. Further, in one example, the first image feature can be randomly divided into the first partial image feature and the second partial image feature, or the first image feature can be equally divided into the first partial image feature and the second partial image feature according to a set equal division rule.
[0062] For example, if the first image feature includes 10 feature maps, the 10 feature maps are equally divided into two parts, one part includes 5 feature maps, and the first convolutional kernel is used to extract features from the 5 feature maps to obtain a first feature extraction result, and the other part also includes 5 feature maps, and the second convolutional kernel is used to extract features from the 5 feature maps to obtain a second feature extraction result.
[0063] In another example, the first image feature can be divided into channels according to a ratio to obtain a first partial image feature and a second partial image feature. For example, the ratio of the first partial image feature to the second partial image feature is 6:4, that is, the first partial image feature is 60% of the first image feature, and the second partial image feature is the remaining 40% of the first image feature. Further, for example, if the first image feature includes 10 feature maps, the first image is divided into two parts according to a ratio of 6:4, one part includes 6 feature maps, and the first convolutional kernel is used to extract features from the 6 feature maps to obtain a first feature extraction result, and the other part includes 4 feature maps, and the second convolutional kernel is used to extract features from the 4 feature maps to obtain a second feature extraction result.
[0064] In this embodiment, instead of separately extracting features of the first image feature using the first convolutional kernel and the second convolutional kernel, the first image feature is divided into two partial image features, and the first convolutional kernel and the second convolutional kernel are used to extract features of the two partial image features respectively. In this way, the computational amount can be reduced, the efficiency of feature extraction can be ensured to a certain extent, and while ensuring the efficiency of feature extraction, more context information can also be extracted, the details and edge restoration quality of higher-resolution images can be improved, and the performance of image restoration can be improved.
[0065] According to an embodiment of the present application, M is equal to 1, N is equal to K, and the second convolutional kernel includes a 1*K convolutional kernel; M is equal to K, N is equal to 1, and the second convolutional kernel further includes a K*1 convolutional kernel. Using the second convolutional kernel to extract features from the second partial image feature to obtain a second feature extraction result includes: performing a convolutional operation on the second partial image feature using the 1*K convolutional kernel to obtain a first convolutional result, performing a convolutional operation on the first convolutional result using the K*1 convolutional kernel to obtain a second convolutional result, and obtaining the second feature extraction result based on the second convolutional result.
[0066] The image feature can be a tensor including multiple feature maps. A feature map is a two-dimensional array, including feature values in the height direction and feature values in the width direction. The feature values in the height direction can also be called the vertical feature values, and the feature values in the width direction can also be called the horizontal feature values.
[0067] In this embodiment, a 1*K convolutional kernel is used to perform a convolution operation on the second part of the image features to obtain a first convolution result, which can expand the receptive field in one direction, that is, the width direction; a K*1 convolutional kernel is used to perform a convolution operation on the first convolution result to obtain a second convolution result, which can expand the receptive field in the other direction, that is, the height direction. The combination of the two can obtain more context information in the height and width directions, improve the restoration quality of the details and edges of the higher-resolution image, and improve the performance of image restoration.
[0068] According to an embodiment of the present application, M is equal to 1, N is equal to K, and the second convolutional kernel further includes a 1*K dilated convolutional kernel. M is equal to K, N is equal to 1, and the second convolutional kernel further includes a K*1 dilated convolutional kernel. The dilation rate of the 1*K dilated convolutional kernel and the K*1 dilated convolutional kernel is a first preset value. Obtaining a second feature extraction result based on the second convolution result includes: performing dilated convolution on the second convolution result based on the 1*K dilated convolutional kernel and the first preset value to obtain a first dilated convolution result, and performing dilated convolution on the first dilated convolution result based on the K*1 dilated convolutional kernel and the first preset value to obtain a second dilated convolution result, and the second dilated convolution result is the aforementioned second feature extraction result.
[0069] Dilated Convolution is a special convolution operation. It expands the receptive field of the convolutional kernel by inserting holes (or intervals) in the convolutional kernel. If the dilation rate is 1, the elements in the convolutional kernel are continuous, and this kind of convolutional kernel is an ordinary convolutional kernel. If the dilation rate is greater than 1, that is, the first preset value is greater than 1, there are intervals between the elements in the convolutional kernel, which is a non-ordinary convolutional kernel, and this convolutional kernel can be called a dilated convolutional kernel. Figure 5 The following is a schematic diagram of an ordinary convolutional kernel and a convolutional kernel with a dilation rate of 2 provided by an exemplary embodiment of the present application. As Figure 5 shown, in the current perspective, from left to right is the horizontal direction, X represents the horizontal direction, from bottom to top is the vertical direction, Y represents the vertical direction, and d represents the dilation rate. As Figure 5 shown in (a) of [], the dilation rate d of this 3×3 convolutional kernel is equal to 1, and the elements in the convolutional kernel are continuous. As Figure 5 shown in (b) of [], the dilation rate d is equal to 2, and there is 1 interval between the elements in the convolutional kernel. This kind of convolutional kernel is larger than Figure 5Compared with the convolution kernel shown in (a) therein, when sliding on the image or feature map each time, it covers a larger range of pixels, can sense more context information, and more image features representing context information are used to enhance the details and edges of the image, improving the restoration quality of the details and edges of the higher-resolution image and the performance of image restoration.
[0070] For example, Figure 6 The following shows a schematic diagram of the deep feature extraction sub-module in the deep feature extraction module in Figure 4 an exemplary embodiment of the present application. As shown in Figure 6 , the deep feature extraction sub-module includes an image feature extraction module 1, and the image feature extraction module 1 includes a "1*K, DW-conv" module, a "K*1, DW-conv" module, a "1*K, DW-D-conv" module, and a "K*1, DW-D-conv" module. Among them, the "1*K, DW-conv" module is used to perform a convolution operation on the second part of the image features using a 1*K convolution kernel to obtain a first convolution result; the "K*1, DW-conv" module is used to perform a convolution operation on the first convolution result using a K*1 convolution kernel to obtain a second convolution result; the "1*K, DW-D-conv" module is used to perform a dilated convolution on the second convolution result based on a 1*K convolution kernel and a first preset value to obtain a first dilated convolution result; the "K*1, DW-D-conv" module is used to perform a dilated convolution on the first dilated convolution result based on a K*1 convolution kernel and a first preset value to obtain a second dilated convolution result. Among them, the image feature extraction module 1 can also be called a multi-shape multi-scale large kernel attention module (MM-LKA).
[0071] In the embodiment of the present application, the convolution kernel of the dilated convolution (dilated convolution kernel) covers a larger range of pixels when sliding on the image or feature map each time, can sense more context information, and more image features representing context information are used to enhance the details and edges of the image, improving the restoration quality of the details and edges of the higher-resolution image and the performance of image restoration.
[0072] According to an embodiment of the present application, the first convolution kernel includes a K*K convolution kernel, and feature extraction is performed on the first part of the image features using the first convolution kernel to obtain a first feature extraction result, including: performing a convolution operation on the first part of the image features using a K*K convolution kernel to obtain a third convolution result, and obtaining the first feature extraction result based on the third convolution result.
[0073] In the embodiment of the present application, the first convolution kernel uniformly captures edge information in all directions through a symmetric receptive field, effectively reconstructs edge details, improves the restoration quality of details and edges of an image with a relatively high resolution, and improves the performance of image restoration.
[0074] According to an embodiment of the present application, the first convolution kernel further includes a K×K dilated convolution kernel, and the dilation rate of the K×K dilated convolution kernel is a second preset value. Obtaining a second feature extraction result based on the third convolution result includes: performing dilated convolution on the third convolution result based on the K×K dilated convolution kernel and the second preset value to obtain a third dilated convolution result, and the third dilated convolution result is the aforementioned first feature extraction result.
[0075] As described above, if the dilation rate is greater than 1, there are intervals between the elements in the convolution kernel, which is a non-ordinary convolution kernel. Therefore, the second preset value is a value greater than 1, and the first preset value and the second preset value may be equal or not equal.
[0076] For example, please continue to refer to Figure 6 , the image feature extraction module 1 further includes a "K×K, DW-conv" module and a "K×K, DW-D-conv". Among them, the "K×K, DW-conv" module is used to perform a convolution operation on the first part of the image features using a K×K convolution kernel to obtain a third convolution result; the "K×K, DW-D-conv" is used to perform dilated convolution on the third convolution result based on the K×K convolution kernel and the second preset value to obtain a third dilated convolution result.
[0077] In the embodiment of the present application, the convolution kernel of dilated convolution (dilated convolution kernel) covers a larger pixel range each time it slides on an image or a feature map, can sense more context information, and more image features representing context information are used to enhance the details and edges of the image, improve the restoration quality of details and edges of an image with a relatively high resolution, and improve the performance of image restoration.
[0078] According to an embodiment of the present application, generating a second image based on the first feature extraction result and the second feature extraction result includes: splicing the second dilated convolution result and the third dilated convolution result to obtain a first splicing result; performing a convolution operation on the first splicing result to obtain a fourth convolution result; performing a dot product of the fourth convolution result and the first image features to obtain a first dot product result; generating a second image based on the first dot product result.
[0079] For example, please continue to refer to Figure 6, the image feature extraction module 1 further includes a "Concat" module and a "Conv-1" module. Among them, the "Concat" module is used to splice the second dilated convolution result and the third dilated convolution result to obtain a first splicing result; the "Conv-1" module is used to perform a convolution operation on the first splicing result to obtain a fourth convolution result; the deep feature extraction sub-module is further used to perform a dot product of the fourth convolution result and the first image feature to obtain a first dot product result. Then, the image resolution enhancement model is used to generate a second image based on the first dot product result.
[0080] In one example, a convolution kernel with a size of 1*1 can be used to perform a convolution operation on the first splicing result to obtain a fourth convolution result.
[0081] In the embodiment of the present application, after multiple convolution operations, the image features usually lose some detail information of the original image. By performing a dot product of the fourth convolution result and the first image feature, more detail information represented by the first image feature can be retained, which can better enhance the details and edges of the image, improve the recovery quality of the details and edges of the higher-resolution image, and improve the performance of image restoration.
[0082] According to an embodiment of the present application, before determining the first image feature of the first image, it further includes: determining the second image feature of the first image, and determining the first image feature of the first image, including: performing a star convolution operation on the second image feature to obtain the first image feature.
[0083] In one example, shallow feature extraction is performed on the first image to obtain a second image feature. For example, please continue to refer to Figure 4 , the second image feature can be obtained by performing shallow feature extraction on the first image through a shallow feature extraction module.
[0084] The second image feature is also used to characterize the features of the image content in the first image. Feature extraction refers to extracting features from an image that can characterize the image content. These features can be edges, textures, colors, shapes, etc. in the image, and useful information for subsequent enhancing the resolution of the image is extracted. Shallow feature extraction can refer to an operation processed by one convolutional layer.
[0085] The implicit high-dimensional non-linear feature space does not simply increase the number of channels of the image features, that is, the channel dimension of the first image feature can be equal to the channel dimension of the second image feature. The core of the star convolution operation lies in mapping the input features to the implicit high-dimensional non-linear feature space through a pixel-by-pixel dot product operation, thereby enhancing the representation ability of the image features. In the embodiment of the present application, by enhancing the representation ability of the image features, the recovery quality of the details and edges of the higher-resolution image is improved, and the performance of image restoration is improved.
[0086] According to an embodiment of the present application, performing a star convolution operation on the second image feature to obtain the first image feature includes: performing a convolution operation on the second image feature to obtain a fifth convolution result; performing multiple star convolution operations on the second image feature, using the output result of each star convolution operation as the input of the subsequent convolution operation until the output result of the last time among the multiple times is obtained; performing a convolution operation on the output result of each star convolution operation to obtain a sixth convolution result; and obtaining the first image feature based on the fifth convolution result, the sixth convolution result corresponding to the output result of each star convolution operation, and the output result of the last time among the multiple times.
[0087] For example, please continue to refer to Figure 6 , the deep feature extraction sub-module further includes an image feature extraction module 2, and the image feature extraction module 2 includes a plurality of "StarConv" modules and a plurality of "Conv-1" modules; wherein, the plurality of "StarConv" modules are connected end to end in sequence, and are used to perform multiple star convolution operations on the second image feature, using the output result of each star convolution operation as the input of the subsequent convolution operation until the output result of the last time among the multiple times is obtained; each "Conv-1" module is respectively connected to the output of a "StarConv" module, and is used to perform a convolution operation on the output result of each star convolution operation to obtain a sixth convolution result. The image feature extraction module 2 can also be called a star distillation module (Star distillation module, SDM).
[0088] In one example, a convolution kernel with a size of 1*1 can be used to perform a convolution operation on the output result of each star convolution operation to obtain a sixth convolution result.
[0089] In one example, the fifth convolution result, the sixth convolution result corresponding to the output result of each star convolution operation, and the output result of the last time among the multiple times are concatenated (concatenate, concat) to obtain the first image feature. Concatenation means connecting two or more tensors end to end in a specific dimension to generate a larger tensor.
[0090] In another example, the fifth convolution result, the sixth convolution result corresponding to the output result of each star convolution operation, and the output result of the last time among the multiple times are concatenated to obtain a concatenation result, and then feature extraction is performed on the concatenation result to obtain the first image feature.
[0091] For example, please continue to refer to Figure 6, the deep feature extraction sub-module further includes: a "Concat" module and a "Conv-1" module. Among them, the "Concat" module is used to splice the fifth convolution result, the sixth convolution result corresponding to the output result of each star convolution operation, and the output result of the last time among multiple times to obtain a splicing result, and the "Conv-1" module is used to extract features from the splicing result to obtain the first image feature.
[0092] In the embodiment of the present application, the image features extracted by different convolutional layers can represent different degrees of detail information and semantic information. The fifth convolution result, the sixth convolution result corresponding to the output result of each star convolution operation, and the output result of the last time among multiple times can simultaneously extract different degrees of detail information and context information, which can better enhance the details and edges of the image, improve the restoration quality of the details and edges of the higher-resolution image, and improve the performance of image restoration.
[0093] According to an embodiment of the present application, generating a second image based on the feature extraction result includes: obtaining a third image feature based on the second image feature and the feature extraction result; generating a second image based on the third image feature.
[0094] In the embodiment of the present application, obtaining a third image feature based on the second image feature and the feature extraction result can extract more context information in the second image feature, which can better enhance the details and edges of the image, improve the restoration quality of the details and edges of the higher-resolution image, and improve the performance of image restoration.
[0095] In an example, for example, please continue to refer to Figure 6 , the feature extraction result includes the second atrous convolution result and the third atrous convolution result. According to the above technical solution, the first dot product result is obtained, and the first dot product result is subjected to a convolution operation to obtain the seventh convolution result; the seventh convolution result is subjected to a pixel normalization (Pixel Norm) operation to obtain a pixel normalization result; the pixel normalization result is dot-added to the second image feature to obtain a first dot-added result (which can also be called the third image feature), and a second image is generated based on the first dot-added result.
[0096] According to an embodiment of the present application, generating a second image based on the third image feature includes: updating the third image feature to the second image feature, and repeating the step of obtaining the third image feature based on the second image feature and the feature extraction result; generating a second image based on the third image feature obtained by repeatedly executing the step of obtaining the third image feature based on the second image feature and the feature extraction result.
[0097] Repeat the step of obtaining the third image feature based on the second image feature and the feature extraction result, that is, obtain the third image feature based on the second image feature and the feature extraction result multiple times; wherein, the third image feature obtained each time based on the second image feature and the feature extraction result is used as the second image feature for the next time.
[0098] For example, Figure 7 Shown is a schematic structural diagram of the intermediate and deep feature extraction module provided by an exemplary embodiment of the present application. As Figure 4 shown, the intermediate and deep feature extraction module includes multiple intermediate and deep feature extraction sub-modules. The third image feature obtained by each intermediate and deep feature extraction sub-module based on the second image feature and the feature extraction result is used as the second image feature for the next intermediate and deep feature extraction sub-module, and so on. The image resolution enhancement model generates a second image based on the third image feature output by each intermediate and deep feature extraction sub-module. Further, in one example, the structure of the intermediate and deep feature extraction sub-module can refer to the description of Figure 7 above, which will not be elaborated here. Figure 6
[0099] In one example, continue to refer to Figure 4 As Figure 4 shown, the feature fusion module is used to splice the third image features output by each intermediate and deep feature extraction sub-module, and the reconstruction module generates a second image based on the splicing result.
[0100] In the embodiments of the present application, the image features extracted by different convolutional layers can represent different levels of detail information and context information. The third image feature obtained each time based on the second image feature and the feature extraction result can retain different levels of detail information and context information at the same time, which can better enhance the details and edges of the image, improve the restoration quality of the details and edges of the higher-resolution image, and improve the performance of image restoration.
[0101] According to an embodiment of the present application, an image resolution enhancement model is used to enhance the resolution of an image. Specifically, the image resolution enhancement model obtains a first image to be processed; wherein, the first image has a first resolution; the image resolution enhancement model determines a first image feature of the first image; wherein, the first image feature is used to represent the feature of the image content in the first image; the image resolution enhancement model uses a first convolution kernel and a second convolution kernel to perform feature extraction on the first image feature to obtain a feature extraction result; wherein, the first convolution kernel includes a K*K convolution kernel, and the second convolution kernel includes an M*N convolution kernel, and M and N are not equal; the image resolution enhancement model generates a second image based on the feature extraction result; wherein, the second image has a second resolution, and the second resolution is greater than the first resolution.
[0102] Figure 8 The following is a schematic flowchart of an image processing method provided by another exemplary embodiment of the present application. Figure 8 The embodiment is Figure 2 an example of the embodiment. To avoid repetition, for the same parts, reference may be made to the description in the above embodiment, and details will not be elaborated here. As Figure 8 shown, the image processing method may include the following steps.
[0103] 810: Obtain a first image to be processed; wherein, the first image has a first resolution.
[0104] 820: Determine a second image feature of the first image.
[0105] 830: Perform a star convolution operation on the second image feature to obtain a first image feature.
[0106] 840: Use a first convolution kernel to extract features from a first part of the image features of the first image feature to obtain a first feature extraction result, and use a second convolution kernel to extract features from a second part of the image features to obtain a second feature extraction result.
[0107] 850: Generate a second image based on the first feature extraction result and the second feature extraction result; wherein, the second image has a second resolution, and the second resolution is greater than the first resolution.
[0108] To improve the image resolution, an image resolution improvement model can be used. The inference process of the image resolution improvement model was introduced above. Next, the training process of the image resolution improvement model will be introduced.
[0109] Figure 9 The following is a schematic flowchart of a model training method for an image resolution improvement model provided by an exemplary embodiment of the present application. Figure 9 The method can be executed by a training device, and the training device can be a computer, a server, etc. As Figure 9 shown, the model training method may include the following steps.
[0110] 910: Obtain training samples; wherein, the training samples include sample images and target images.
[0111] Specifically, in the training stage, the target image can be the true label of the image obtained by the image resolution improvement model. The sample image can be used as the input sample of the model resolution improvement model.
[0112] 920: Input the sample image into the image resolution improvement model to obtain a predicted image.
[0113] In one example, a sample image can be rotated to different angles, such as 90 degrees, 180 degrees, 270 degrees, and input into an image resolution enhancement model. The sample images rotated to various angles correspond to a target image, which can enable the image resolution enhancement model to learn more diverse features, so that when facing images at different angles, it can more accurately enhance the resolution of the images.
[0114] 930: Determine a loss value based on the difference between the predicted image and the target image.
[0115] In one example, determine a loss value based on the difference between the predicted image and the target image and a preset loss function.
[0116] 940: Train the image resolution enhancement model based on the loss value to obtain a trained image resolution enhancement model.
[0117] In one example, use the Adaptive Moment Estimation (Adam) algorithm to adjust the parameters of the image resolution enhancement model to obtain a trained image resolution enhancement model.
[0118] The embodiments of the present application provide a model training method. By inputting a sample image into an image resolution enhancement model to obtain a predicted image, a trained image resolution enhancement model can be obtained based on the difference between the predicted image and the target image. The image resolution enhancement model obtained by this model training method can improve the resolution of the input image.
[0119] Exemplary device
[0120] Figure 10 The following shows a schematic structural diagram of an image processing device provided by an exemplary embodiment of the present application. As Figure 10 shown, the image processing device 1000 includes: an acquisition module 1010, a determination module 1020, a feature extraction module 1030, and an image generation module 1040.
[0121] The acquisition module 1010 is configured to acquire a first image to be processed; wherein, the first image has a first resolution. The determination module 1020 is configured to determine a first image feature of the first image; wherein, the first image feature is used to characterize the feature of the image content in the first image. The feature extraction module 1030 is configured to perform feature extraction on the first image feature by using a first convolution kernel and a second convolution kernel to obtain a feature extraction result. The image generation module 1040 is configured to generate a second image based on the feature extraction result; wherein, the second image has a second resolution, and the second resolution is greater than the first resolution.
[0122] An embodiment of the present application provides an image processing device. By using convolution kernels of different shapes (different sizes) to extract image features of a lower-resolution image, more image features representing context information can be obtained. The more image features representing context information are used to enhance the details and edges of the image, improve the restoration quality of the details and edges of a higher-resolution image, and improve the performance of image restoration.
[0123] According to an embodiment of the present application, the feature extraction module 1030 is configured to use a first convolution kernel to extract a first part of the image features of the first image feature to obtain a first feature extraction result, and use a second convolution kernel to extract a second part of the image features of the first image feature to obtain a second feature extraction result. The image generation module 1040 is configured to generate a second image based on the first feature extraction result and the second feature extraction result.
[0124] According to an embodiment of the present application, M is equal to 1, N is equal to K, the second convolution kernel includes a 1*K convolution kernel, M is equal to K, N is equal to 1, the second convolution kernel further includes a K*1 convolution kernel. The feature extraction module 1030 is configured to perform a convolution operation on the second part of the image features using the 1*K convolution kernel to obtain a first convolution result, perform a convolution operation on the first convolution result using the K*1 convolution kernel to obtain a second convolution result, and obtain the second feature extraction result based on the second convolution result.
[0125] According to an embodiment of the present application, M is equal to 1, N is equal to K, the second convolution kernel further includes a 1*K dilated convolution kernel, M is equal to K, N is equal to 1, the second convolution kernel further includes a K*1 dilated convolution kernel. The dilation rate of the 1*K dilated convolution kernel and the K*1 dilated convolution kernel is a first preset value. The feature extraction module 1030 is configured to perform dilated convolution on the second convolution result based on the 1*K dilated convolution kernel and the first preset value to obtain a first dilated convolution result, perform dilated convolution on the first dilated convolution result based on the K*1 dilated convolution kernel and the first preset value to obtain a second dilated convolution result, and the second dilated convolution result is the aforementioned second feature extraction result.
[0126] According to an embodiment of the present application, the first convolution kernel includes a K*K convolution kernel. The feature extraction module 1030 is configured to include: performing a convolution operation on the first part of the image features using the K*K convolution kernel to obtain a third convolution result, and obtaining the first feature extraction result based on the third convolution result.
[0127] According to an embodiment of the present application, the feature extraction module 1030 is configured to splice the second dilated convolution result and the third dilated convolution result to obtain a first splicing result; perform a convolution operation on the first splicing result to obtain a fourth convolution result; and perform a dot product of the fourth convolution result and the first image feature to obtain a first dot product result.
[0128] According to an embodiment of the present application, the determination module 1020 is configured to determine a second image feature of the first image before determining the first image feature of the first image. Determining the first image feature of the first image includes: performing a star convolution operation on the second image feature to obtain the first image feature.
[0129] According to an embodiment of the present application, the feature extraction module 1030 is configured to perform a convolution operation on the second image feature to obtain a fifth convolution result; perform multiple star convolution operations on the second image feature, and use the output result of each star convolution operation as the input of the subsequent convolution operation until the output result of the last time among the multiple times is obtained; perform a convolution operation on the output result of each star convolution operation to obtain a sixth convolution result; and obtain the first image feature based on the fifth convolution result, the sixth convolution result corresponding to the output result of each star convolution operation, and the output result of the last time among the multiple times.
[0130] According to an embodiment of the present application, generating a second image based on the feature extraction result includes: obtaining a third image feature based on the second image feature and the feature extraction result; and generating the second image based on the third image feature.
[0131] According to an embodiment of the present application, the image generation module 1040 is configured to update the third image feature to the second image feature, and repeat the step of obtaining the third image feature based on the second image feature and the feature extraction result; and generate the second image based on the third image feature obtained by repeatedly performing the step of obtaining the third image feature based on the second image feature and the feature extraction result.
[0132] According to an embodiment of the present application, to improve the resolution of an image using an image resolution improvement model. Specifically, the acquisition module 1010 is configured to obtain a first image to be processed using the image resolution improvement model; wherein, the first image has a first resolution; the determination module 1020 is configured to determine the first image feature of the first image using the image resolution improvement model; wherein, the first image feature is used to characterize the feature of the image content in the first image; the feature extraction module 1030 is configured to perform feature extraction on the first image feature using the image resolution improvement model with a first convolution kernel and a second convolution kernel to obtain a feature extraction result; wherein, the first convolution kernel includes a K*K convolution kernel, and the second convolution kernel includes an M*N convolution kernel, and M and N are not equal; the image generation module 1040 is configured to generate a second image based on the feature extraction result using the image resolution improvement model; wherein, the second image has a second resolution, and the second resolution is greater than the first resolution.
[0133] It should be understood that the operations and functions of the acquisition module 1010, the determination module 1020, the feature extraction module 1030, and the image generation module 1040 in the above embodiments can be referred to the above Figure 2 or Figure 8The descriptions in the image processing method provided in the embodiments are not repeated here to avoid redundancy.
[0134] Figure 11 The following is a schematic structural diagram of a model training device provided by an exemplary embodiment of the present application. As Figure 11 shown, the model training device 1100 includes: an acquisition module 1110, a prediction module 1120, a determination module 1130, and a training module 1140.
[0135] The acquisition module 1110 is configured to acquire training samples; wherein, the training samples include sample images and target images. The prediction module 1120 is configured to input the sample images into an image resolution enhancement model to obtain predicted images. The determination module 1130 is configured to determine a loss value based on the difference between the predicted images and the target images. The training module 1140 is configured to train the image resolution enhancement model based on the loss value to obtain a trained image resolution enhancement model.
[0136] It should be understood that the operations and functions of the acquisition module 1010, the first determination module 1020, the second determination module 1030, and the training module 1040 in the above embodiments may refer to the descriptions in the model training method provided in the above Figure 9 embodiments. To avoid redundancy, they are not repeated here.
[0137] An embodiment of the present application provides a model training device. By inputting sample images into an image resolution enhancement model, predicted images are obtained, so that a trained image resolution enhancement model can be obtained based on the difference between the predicted images and the target images. The image resolution enhancement model obtained by this model training method can improve the resolution of the input images.
[0138] Figure 12 The following is a block diagram of an electronic device 1100 for executing an image processing method or a model training method provided by an exemplary embodiment of the present application. The electronic device 1100 may specifically be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), an augmented reality (AR) / virtual reality (VR) device, a medical imaging device, a monitoring device, a satellite imaging device, the aforementioned training device, or other devices.
[0139] Refer to Figure 12, the electronic device 1200 includes a processing component 1210, which further includes one or more processors, and memory resources represented by a memory 1220 for storing instructions executable by the processing component 1210, such as application programs. The application programs stored in the memory 1220 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1210 is configured to execute instructions to perform the above-mentioned image processing method or model training method.
[0140] The electronic device 1200 may further include a power component configured to perform power management of the electronic device 1200, a wired or wireless network interface configured to connect the electronic device 1200 to a network, and an input / output (I / O) interface. The electronic device 1200 may be operated based on an operating system stored in the memory 1220, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.
[0141] The embodiments of the present application further provide a non-transitory computer-readable storage medium. When the instructions in the storage medium are executed by the processor of the above-mentioned electronic device 1200, the electronic device 1200 is enabled to execute an image processing method or a model training method.
[0142] The embodiments of the present application further provide a computer program product. The computer program product includes a computer program. When the computer program is executed by the processor of a computer device, the computer device is enabled to execute the image processing method or the model training method provided in any of the above embodiments.
[0143] All of the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present application, which will not be elaborated herein one by one.
[0144] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0145] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated herein.
[0146] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.
[0147] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0148] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0149] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical discs that can store program verification codes.
[0150] It should be noted that in the description of this application, the terms "first", "second", "third", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "multiple" is two or more.
[0151] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0152] The above are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, etc. made within the spirit and principles of this application shall be included within the protection scope of this application.
Claims
1. An image processing method, characterized in that, Including: Obtain a first image to be processed; wherein, the first image has a first resolution; Determine a first image feature of the first image; wherein, the first image feature is used to characterize the feature of the image content in the first image; Use a first convolution kernel and a second convolution kernel to perform feature extraction on the first image feature to obtain a feature extraction result; wherein, the first convolution kernel includes a K*K convolution kernel, and the second convolution kernel includes an M*N convolution kernel, and M and N are not equal; Generate a second image based on the feature extraction result; wherein, the second image has a second resolution, and the second resolution is greater than the first resolution.
2. The image processing method according to claim 1, wherein The first image feature includes a first partial image feature and a second partial image feature. The using the first convolution kernel and the second convolution kernel to perform feature extraction on the first image feature to obtain a feature extraction result includes: Use the first convolution kernel to perform feature extraction on the first partial image feature of the first image feature to obtain a first feature extraction result; Use the second convolution kernel to perform feature extraction on the second partial image feature of the first image feature to obtain a second feature extraction result; The generating a second image based on the feature extraction result includes: generating a second image based on the first feature extraction result and the second feature extraction result.
3. The image processing method according to claim 2, wherein M is equal to 1, N is equal to K, the second convolution kernel includes a 1*K convolution kernel, M is equal to K, N is equal to 1, the second convolution kernel further includes a K*1 convolution kernel. The using the second convolution kernel to perform feature extraction on the second partial image feature of the first image feature to obtain a second feature extraction result includes: Use the 1*K convolution kernel to perform a convolution operation on the second partial image feature to obtain a first convolution result; Use the K*1 convolution kernel to perform a convolution operation on the first convolution result to obtain a second convolution result; Obtain a second feature extraction result based on the second convolution result.
4. The image processing method according to claim 3, wherein M is equal to 1, N is equal to K, the second convolution kernel further includes a 1*K dilated convolution kernel, M is equal to K, N is equal to 1, the second convolution kernel further includes a K*1 dilated convolution kernel. The dilation rate of the 1*K dilated convolution kernel and the K*1 dilated convolution kernel is a first preset value. The obtaining a second feature extraction result based on the second convolution result includes: Perform dilated convolution on the second convolution result based on the 1*K dilated convolution kernel and the first preset value to obtain a first dilated convolution result; Perform dilated convolution on the first dilated convolution result based on the K*1 dilated convolution kernel and the first preset value to obtain a second dilated convolution result; wherein, the second dilated convolution result is the second feature extraction result.
5. The image processing method according to any one of claims 2 to 4, characterized in that, The first convolution kernel includes a K*K convolution kernel. The using the first convolution kernel to perform feature extraction on the first partial image feature of the first image feature to obtain a first feature extraction result includes: Use the K*K convolution kernel to perform a convolution operation on the first partial image feature to obtain a third convolution result; Obtain a first feature extraction result based on the third convolution result.
6. The image processing method according to claim 5, wherein The first convolutional kernel further includes a K×K dilated convolutional kernel, and the dilation rate of the K×K dilated convolutional kernel is a second preset value. Obtaining a first feature extraction result based on the third convolutional result includes: Performing dilated convolution on the third convolutional result based on the K×K dilated convolutional kernel and the second preset value to obtain a third dilated convolutional result; wherein, the third dilated convolutional result is the first feature extraction result.
7. The image processing method according to claim 6, wherein The generating a second image based on the first feature extraction result and the second feature extraction result includes: Concatenating the second dilated convolutional result and the third dilated convolutional result to obtain a first concatenation result; Performing a convolutional operation on the first concatenation result to obtain a fourth convolutional result; Performing a dot product of the fourth convolutional result and the first image feature to obtain a first dot product result; Generating a second image based on the first dot product result.
8. The image processing method according to claim 1, wherein Before determining the first image feature of the first image, it further includes: determining a second image feature of the first image; The determining the first image feature of the first image includes: Performing a star convolutional operation on the second image feature to obtain a first image feature.
9. The image processing method according to claim 8, wherein The performing a star convolutional operation on the second image feature to obtain a first image feature includes: Performing a convolutional operation on the second image feature to obtain a fifth convolutional result; Performing a plurality of star convolutional operations on the second image feature, using the output result of each star convolutional operation as the input of the subsequent convolutional operation until obtaining the output result of the last time among the plurality of times; Performing a convolutional operation on the output result of each star convolutional operation to obtain a sixth convolutional result; Obtaining a first image feature based on the fifth convolutional result, the sixth convolutional result corresponding to the output result of each star convolutional operation, and the output result of the last time among the plurality of times.
10. The image processing method according to claim 8, wherein The generating a second image based on the feature extraction result includes: Obtaining a third image feature based on the second image feature and the feature extraction result; Generating the second image based on the third image feature.
11. The image processing method according to claim 10, characterized in that The generating the second image based on the third image feature includes: Updating the third image feature to a second image feature, and repeatedly executing the step of obtaining a third image feature based on the second image feature and the feature extraction result; Generating the second image based on the third image feature obtained by repeatedly executing the step of obtaining a third image feature based on the second image feature and the feature extraction result.
12. The image processing method according to any one of claims 1 to 4, characterized in that, It includes: An image resolution enhancement model obtains a first image to be processed; wherein, the first image has a first resolution; The image resolution enhancement model determines a first image feature of the first image; wherein, the first image feature is used to characterize the feature of the image content in the first image; The image resolution enhancement model performs feature extraction on the first image feature by using a first convolutional kernel and a second convolutional kernel to obtain a feature extraction result; wherein, the first convolutional kernel includes a K×K convolutional kernel, the second convolutional kernel includes an M×N convolutional kernel, and M and N are not equal; The image resolution improvement model generates a second image based on the feature extraction result; wherein, the second image has a second resolution, and the second resolution is greater than the first resolution.
13. An electronic device, characterized in that, It includes: A processor; A memory for storing instructions executable by the processor, wherein, the processor is used to execute the image processing method described in any one of claims 1 to 12 above.
14. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is used to execute the image processing method described in any one of claims 1 to 12 above.