Image Processing Method, Electronic Device, and Computer Program Product
Through the multi-feature extraction neural network, high-frequency feature prediction processing is performed on the image to be processed and fused, solving the problem of low high-frequency image accuracy in the prior art and achieving higher image processing accuracy.
Patent Information
- Application Number
- CN202111101722.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-18
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-09-18
AI Technical Summary
The existing super-resolution image processing methods are difficult to effectively improve the accuracy of high-frequency images, and the generated textures are uncontrollable, resulting in low accuracy of the output results.
Through a multi-feature extraction neural network, high-frequency feature prediction processing is performed on the image to be processed, multiple output images with different high-frequency features are obtained, and these output images are fused to generate high-frequency images with a resolution higher than those to be processed.
By fusion of multiple output images with different high-frequency characteristics, the accuracy of the high-frequency image can be effectively improved, reflecting the characteristics of different aspects of the low-frequency image, thereby improving the accuracy of image processing.
Smart Images

Figure CN114022354B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an image processing method, electronic equipment and computer program product. Background Art
[0002] Super-resolution (SR) is a visual task to restore a high-resolution image given a low-resolution image. In recent years, CNN super-resolution has gradually replaced traditional methods such as back-projection, sparse representation and dictionary learning, becoming the mainstream super-resolution method.
[0003] Because the content in a low-resolution image is not very recognizable, it may deal with many or even infinite high-resolution forms. Therefore, super-resolution is ill-posed. The traditional single-branch method directly learns low-resolution images and easily produces blurry outputs. Some GAN-based methods add GAN loss to the single-branch network to produce high-frequency textures similar to high-resolution images (HR), but in essence they can only obtain the output of a single high-frequency feature, and the generated texture is uncontrollable, resulting in low accuracy of the output results. Summary of the invention
[0004] In view of this, an object of the present invention is to provide an image processing method, an electronic device and a computer program product to improve the accuracy of high-frequency images.
[0005] In a first aspect, an embodiment of the present invention provides an image processing method, the method comprising: obtaining an image to be processed; performing high-frequency feature prediction processing on the image to be processed through a multi-feature extraction neural network to obtain multiple output images with different high-frequency features; wherein the similarity between any two output images is less than a first similarity; fusing the multiple output images to obtain a high-frequency image corresponding to the image to be processed; wherein the resolution of the high-frequency image is higher than that of the image to be processed.
[0006] Further, the multi-feature extraction neural network described above includes multiple layers, and each layer includes multiple feature extraction modules. The feature extraction modules in the previous layer are connected to at least one feature extraction module in the next layer. The interconnected feature extraction modules located in different layers constitute a channel of the multi-feature extraction neural network. The step of performing high-frequency feature prediction processing on the image to be processed through the multi-feature extraction neural network to obtain multiple output images with different high-frequency features includes: performing first feature extraction on the image to be processed through the multi-feature extraction neural network to obtain initial features; performing second feature extraction on the initial features through the current channel among the multiple channels of the multi-feature extraction neural network to obtain the output features corresponding to this channel; wherein, the resolution of the output features corresponding to any one channel is higher than the resolution of the initial features; performing upsampling processing on the output features of each channel respectively to obtain multiple output images.
[0007] Further, the step of performing second feature extraction on the initial features through the current channel among the multiple channels of the multi-feature extraction neural network to obtain the output features corresponding to this channel includes: performing second feature extraction on the input features through the current feature extraction module to obtain the initial output features corresponding to this feature extraction module; wherein, if the current feature extraction module is the first feature extraction module in the channel, the input features are the initial features, otherwise, the input features are the output features of the previous feature extraction module; combining the initial output features and the input features corresponding to the current feature extraction module as the output features of the current feature extraction module; determining whether the current feature extraction module is the last feature extraction module in the current channel; if not, continue to perform the second feature extraction operation; if so, fusing the output features of the current feature extraction module with the input features of the first feature extraction module in each layer of the current channel, and determining the fusion result as the output features corresponding to the current channel.
[0008] Further, the step of fusing the multiple output images to obtain the high-frequency image corresponding to the image to be processed includes: determining the weights of each output image through the fusion network; performing weighted summation on each output image according to the corresponding weights to obtain the high-frequency image.
[0009] Further, the above multi-feature extraction neural network is obtained through the following training: obtaining a sample image and an initial neural network; wherein, the sample image includes a training image and its corresponding label image, and the resolution of the label image is higher than that of the training image; performing third feature extraction on the sample image to obtain training features corresponding to the training image; performing fourth feature extraction on the training features through the initial neural network to obtain multiple predicted features; the resolution of the predicted features is higher than that of the training features; determining a predicted image corresponding to the predicted features according to the predicted features; determining a total loss value according to the predicted image and its corresponding label image; adjusting the parameters of the multi-feature extraction neural network according to the total loss value until the first training stop condition is satisfied.
[0010] Further, the above multi-feature extraction neural network includes multiple layers, and each layer includes multiple feature extraction modules. The feature extraction modules of the previous layer are connected to at least one feature extraction module of the next layer. The mutually connected feature extraction modules located in different layers constitute a channel of the multi-feature extraction neural network. The step of performing fourth feature extraction on the training features through the multi-feature extraction neural network to obtain multiple predicted features includes: performing fourth feature extraction on the input features through the current feature extraction module to obtain the initial output features corresponding to the feature extraction module; wherein, if the current feature extraction module is the first feature extraction module in the channel, the input features are the training features, otherwise, the input features are the output features of the previous feature extraction module; combining the initial output features with the input features corresponding to the current feature extraction module as the output features of the current feature extraction module; determining whether the current feature extraction module is the last feature extraction module in the current channel; if not, continue to perform the fourth feature extraction operation; if so, fusing the output features of the current feature extraction module with the input features of the first feature extraction module of each layer in the current channel, and determining the fusion result as the predicted features corresponding to the current channel.
[0011] Further, the step of determining the total loss value according to the predicted image and its corresponding label image includes: obtaining a first loss value corresponding to the predicted image according to each predicted image and the label image; obtaining a second loss value corresponding to the two predicted images according to every two predicted images and the label image; determining the total loss value according to the first loss value and the second loss value.
[0012] Further, the step of obtaining the second loss value corresponding to the two predicted images according to each two predicted images and the label image includes: determining a first positive sample according to the label image, determining a first negative sample according to the first predicted image among the two predicted images, and determining a first anchor sample according to the second predicted image among the two predicted images; calculating a triplet loss value through the first positive sample, the first negative sample, and the first anchor sample to obtain a first triplet loss value; determining a second positive sample according to the label image, determining a second anchor sample according to the first predicted image, and determining a second negative sample according to the second predicted image; calculating a triplet loss value through the second positive sample, the second negative sample, and the second anchor sample to obtain a second triplet loss value; adding the first triplet loss value and the second triplet loss value to obtain the second loss value.
[0013] Further, the above method further includes: extracting a first gray value of the first predicted image and a second gray value of the second predicted image; determining a first residual according to the first gray value and the label image; determining a second residual according to the second gray value and the label image; the step of determining a first positive sample according to the label image, determining a first negative sample according to the first predicted image among the two predicted images, and determining a first anchor sample according to the second predicted image among the two predicted images includes: determining the first residual as the first negative sample; determining the second residual as the first anchor sample; the step of determining a second positive sample according to the label image, determining a second anchor sample according to the first predicted image, and determining a second negative sample according to the second predicted image includes: determining the first residual as the second anchor sample; determining the second residual as the second negative sample.
[0014] Further, the above multi-feature extraction neural network has a tree-shaped topology of M*N, where M represents the number of layers of the multi-feature extraction neural network, and N represents the number of initial modules included in the first layer of the multi-feature extraction neural network. Each initial module in the first layer and the intermediate layer of the multi-feature extraction neural network is connected to N initial modules in the next layer.
[0015] Further, the above fusion network is trained through the following method; obtaining an initial fusion network; determining a predicted fusion image according to each predicted image and its corresponding initial weight; determining a third loss value according to the predicted fusion image and the label image; adjusting the parameters of the initial fusion network according to the third loss value until the second training stop condition is satisfied.
[0016] In a second aspect, an embodiment of the present invention further provides an electronic device, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the image processing method in the first aspect above.
[0017] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions, which, when called and executed by a processor, cause the processor to implement the image processing method in the first aspect above.
[0018] In a fourth aspect, an embodiment of the present invention further provides a computer program product including a computer program, which, when executed by a processor, implements the image processing method in the first aspect above.
[0019] For the image processing method provided by the embodiment of the present invention, first, an image to be processed is obtained, and the image to be processed is subjected to high-frequency feature prediction processing through a multi-feature extraction neural network to obtain a plurality of output images with different high-frequency features, and then the plurality of output images are fused to obtain a high-frequency image corresponding to the image to be processed. Since the plurality of output images with different high-frequency features can reflect different aspects of the features of the low-frequency image, the high-frequency image obtained based on the plurality of output images with different high-frequency features can effectively improve the accuracy of the high-frequency image.
[0020] Other features and advantages of the present disclosure will be described in the following description, or some features and advantages can be inferred from the description or determined without doubt, or can be learned by implementing the above technologies of the present disclosure.
[0021] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, details are described as follows. Description of the Drawings
[0022] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a schematic structural diagram of an electronic system provided by an embodiment of the present invention;
[0024] Figure 2 It is a flowchart of an image processing method provided by an embodiment of the present invention;
[0025] Figure 3 It is a schematic structural diagram of a multi-feature extraction neural network provided by an embodiment of the present invention;
[0026] Figure 4Schematic diagram of the structure and connection relationship of one channel in a multi - feature extraction neural network provided by an embodiment of the present invention;
[0027] Figure 5 Schematic diagram of the comparison between a low - frequency image and a predicted high - frequency image obtained through the multi - feature extraction neural network provided by an embodiment of the present invention;
[0028] Figure 6 Schematic diagram of a fusion network fusing four output images provided by an embodiment of the present invention;
[0029] Figure 7 Schematic diagram of a high - frequency image obtained through an image processing method provided by an embodiment of the present invention;
[0030] Figure 8 Flowchart of a first neural network training method provided by an embodiment of the present invention;
[0031] Figure 9 Schematic diagram of the calculation process of a second loss value provided by an embodiment of the present invention;
[0032] Figure 10 Schematic diagram of the comparison between high - frequency images obtained through different image processing methods;
[0033] Figure 11 Schematic diagram of the comparison of indexes of different image processing methods;
[0034] Figure 12 Schematic diagram of an image processing device provided by an embodiment of the present invention;
[0035] Figure 13 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0037] In recent years, important progress has been made in the research of technologies such as computer vision, deep learning, machine learning, image processing, and image recognition based on artificial intelligence. Artificial Intelligence (AI) is an emerging science and technology that studies and develops theories, methods, technologies, and application systems for simulating and extending human intelligence. The discipline of artificial intelligence is a comprehensive discipline that involves many technical categories such as chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, and neural networks. As an important branch of artificial intelligence, computer vision specifically enables machines to recognize the world. Computer vision technologies usually include face recognition, live detection, fingerprint recognition and anti-counterfeiting verification, biometric recognition, face detection, pedestrian detection, object detection, pedestrian recognition, image processing, image recognition, image semantic understanding, image retrieval, character recognition, video processing, video content recognition, behavior recognition, 3D reconstruction, virtual reality, augmented reality, Simultaneous Localization and Mapping (SLAM), computational photography, robot navigation and positioning, and other technologies. With the research and progress of artificial intelligence technology, this technology has been applied in many fields, such as security, urban management, traffic management, building management, park management, face access, face attendance, logistics management, warehouse management, robots, intelligent marketing, computational photography, mobile phone imaging, cloud services, smart home, wearable devices, driverless, autonomous driving, intelligent healthcare, face payment, face unlocking, fingerprint unlocking, human identity verification, smart screen, smart TV, cameras, mobile Internet, webcasting, beauty, makeup, medical beauty, intelligent temperature measurement, and other fields.
[0038] Current image processing methods for low-frequency images often perform high-frequency prediction based on a single input image, resulting in a low accuracy of the high-frequency image. Based on this, the embodiments of the present invention provide an image processing method, an electronic device, and a computer program product to improve the accuracy of the high-frequency image.
[0039] Refer to Figure 1 the structural schematic diagram of the electronic system 100 shown. This electronic system can be used to implement the image processing method and device of the embodiments of the present invention.
[0040] As Figure 1 shown in the structural schematic diagram of an electronic system, the electronic system 100 includes one or more processing devices 102, one or more storage devices 104, an input device 106, an output device 108, and one or more image acquisition devices 110. These components are interconnected through a bus system 112 and / or other forms of connection mechanisms (not shown). It should be noted that Figure 1 the components and structure of the electronic system 100 shown are only exemplary and not restrictive. According to needs, the electronic system can also have other components and structures.
[0041] The processing device 102 can be a server, a smart terminal, or a device including a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities. It can process the data of other components in the electronic system 100 and can also control other components in the electronic system 100 to execute the image processing function.
[0042] The storage device 104 can include one or more computer program products. The computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage medium. The processing device 102 can run the program instructions to implement the client functions (implemented by the processing device) in the embodiments of the present invention below and / or other desired functions. Various application programs and various data can also be stored in the computer-readable storage medium, such as various data used and / or generated by the application programs, etc.
[0043] The input device 106 can be a device used by the user to input instructions and can include one or more of a keyboard, a mouse, a microphone, and a touch screen, etc.
[0044] The output device 108 can output various information (such as images or sounds) to the outside (for example, to the user) and can include one or more of a display, a speaker, etc.
[0045] The image acquisition device 110 can acquire the image to be processed and store the image to be processed in the storage device 104 for use by other components.
[0046] Exemplarily, the devices for implementing the image processing method, electronic device, and computer program product according to the embodiments of the present invention can be integrally arranged or separately arranged. For example, the processing device 102, the storage device 104, the input device 106, and the output device 108 can be integrally arranged, while the image acquisition device 110 is arranged at a specified position where an image can be acquired. When the devices in the above electronic system are integrally arranged, the electronic system can be implemented as a smart terminal such as a camera, a smart phone, a tablet computer, a computer, a vehicle-mounted terminal, etc.
[0047] Figure 2 It is a flowchart of an image processing method provided for the embodiments of the present invention. Refer to Figure 2 and the method includes the following steps:
[0048] S202: Acquire the image to be processed;
[0049] The image to be processed is a low-frequency image obtained by a photographing device. Specifically, it is an image with a lower resolution compared to a high-frequency image. For example, it can be a low-resolution image taken by a photographing device such as a mobile phone. By processing the low-frequency image, a high-frequency image can be obtained.
[0050] S204: Perform high-frequency feature prediction processing on the image to be processed through a multi-feature extraction neural network to obtain multiple output images with different high-frequency features.
[0051] The multi-feature extraction neural network can be a single-input and multiple-output network structure. For example, it can be a tree network structure, which can include two layers or more layers. Multiple branches of the last layer output multiple output images. For a multi-layer tree structure, each layer has multiple branches, and each branch of the current layer is connected to multiple branches of the next layer to obtain more different branch outputs.
[0052] Among them, the similarity between any two output images is less than a first similarity. The first similarity can be a preset similarity value, which can be set according to experience. The greater the similarity between any two output images, the more similar the two images are. On the contrary, the smaller the similarity, the greater the difference between the two images. In the embodiments of the present invention, the similarity between every two output images is less than the first similarity, indicating that the difference between every two images is relatively large. That is, based on a low-frequency image, multiple output images with relatively large differences are obtained. These output images with relatively large differences can better reflect different high-frequency features of the low-frequency image from multiple aspects.
[0053] The multi-feature extraction neural network is obtained by training on a sample image pair including multiple low-frequency images and high-frequency images. The training process of the multi-feature extraction neural network will be elaborated in detail below and will not be repeated here.
[0054] S206: Fuse the multiple output images to obtain the high-frequency image corresponding to the image to be processed.
[0055] The multiple output images can be upsampled once or multiple times to be converted into images with the same dimension, and then the multiple images with the same dimension are fused to obtain the corresponding high-frequency image. Among them, the resolution of the high-frequency image is higher than that of the low-frequency image.
[0056] The above image processing method provided by the embodiments of the present invention first obtains the image to be processed, performs high-frequency feature prediction processing on the image to be processed through a multi-feature extraction neural network to obtain multiple output images with different high-frequency features, and then fuses the multiple output images to obtain the high-frequency image corresponding to the image to be processed. Since the multiple output images with different high-frequency features can reflect the features of different aspects of the low-frequency image, therefore, the high-frequency image obtained based on the multiple output images with different high-frequency features can effectively improve the accuracy of the high-frequency image.
[0057] In some possible implementation manners, the multi-feature extraction neural network includes multiple layers, and each layer includes multiple feature extraction modules. The feature extraction modules in the previous layer are connected to at least one feature extraction module in the next layer, and the mutually connected feature extraction modules located in different layers constitute a channel of the multi-feature extraction neural network.
[0058] In one example, the multi-feature extraction neural network can be an M*N tree topology structure, where M represents the number of layers of the multi-feature extraction neural network, and N represents the number of initial modules included in the first layer of the multi-feature extraction neural network. Each feature extraction module in the first layer and the intermediate layer of the multi-feature extraction neural network is connected to N feature extraction modules in the next layer.
[0059] For example Figure 3 as shown Figure 3 is a schematic structural diagram of a multi-feature extraction neural network provided by the embodiments of the present invention. Figure 3 The multi-feature extraction neural network shown in is a 2*2 full binary tree structure. The multi-feature extraction neural network includes two network layers. The first layer includes two nodes (i.e., feature extraction modules), and the second layer includes four nodes. Each node in the first layer is respectively connected to two nodes in the second layer. Each node adopts an RG (Residual Group) module commonly used in the super-resolution task. 2 RGs are stacked to form a branch module. Each RG is composed of 4 RCAB (Residual Channel Attention Block) modules. Four output features can be output through each node, and the four output features are connected to the upsampling module to obtain four output images. The upsampling module can specifically be a pixel shuffle module.
[0060] Among them, RCAB represents a commonly used super-resolution network module. The input of RCAB is a feature. The input feature passes through two branches. The first branch is directly connected to the module result through a residual connection. The second branch is further divided into two sub-branches. The first sub-branch is a convolutional layer (for example, it can be two convolutions), and the other sub-branch is a channel attention layer. It should be noted that the above-mentioned number of network layers, the number of RG modules and RCAB modules are only exemplary, and specifically, it can be other combinations of numbers.
[0061] Based on this, the step of performing high-frequency feature prediction processing on the image to be processed through the multi-feature extraction neural network in the above step S204 to obtain multiple output images with different high-frequency features can specifically include:
[0062] (1) Perform first feature extraction on the image to be processed through the multi-feature extraction neural network to obtain an initial feature;
[0063] After obtaining the low-frequency image, the multi-feature extraction neural network can be used to perform feature extraction on the low-frequency image to obtain the initial feature corresponding to the low-frequency image. Specifically, the feature extraction module in the multi-feature extraction neural network can be used to extract the initial feature.
[0064] (2) Perform second feature extraction on the initial feature through the current channel in multiple channels of the multi-feature extraction neural network to obtain the output feature corresponding to this channel;
[0065] Among them, the resolution of the output feature corresponding to any one channel is higher than the resolution of the initial feature;
[0066] (3) Perform upsampling processing on the output feature of each channel respectively to obtain multiple output images.
[0067] The purpose of the upsampling processing is to obtain multiple output images of the same size. The output feature of each channel can be subjected to one upsampling operation or multiple upsampling operations.
[0068] In order to improve the similarity between the output feature obtained by each channel in the above multi-feature extraction neural network and the initial feature and avoid the loss of features during the transmission in each layer, in each channel of the multi-feature extraction neural network, the following second feature extraction operation is performed on the initial feature:
[0069] (1) Perform second feature extraction on the input feature through the current feature extraction module to obtain the initial output feature corresponding to this feature extraction module; among them, if the current feature extraction module is the first feature extraction module in the channel, the input feature is the initial feature, otherwise, the input feature is the output feature of the previous feature extraction module;
[0070] (2) Combine the initial output feature and the input feature corresponding to the current feature extraction module as the output feature of the current feature extraction module;
[0071] (3) Determine whether the current feature extraction module is the last feature extraction module in the current channel; if not, continue to perform the second feature extraction operation;
[0072] (4) If so, fuse the output feature of the current feature extraction module with the input features of the first feature extraction module in each layer of the current channel, and determine the fusion result as the output feature corresponding to the current channel.
[0073] Figure 4 For Figure 3 a schematic diagram of the structure and connection relationship of one channel in the multi-feature extraction neural network shown in Figure 4 as shown, the left box is the first layer in the current channel, and the right box is the second layer in the current channel. In each layer, there are two feature extraction modules, that is Figure 3 the two RCAB modules in Figure 4 constitute a feature extraction module. The following takes Figure 4 as an example to illustrate the process of obtaining the output feature of the current channel by performing the second feature extraction operation on the initial feature: For ease of explanation, the four feature extraction modules in the current channel are sequentially named M1 - M4. At M1, since M1 is the first feature extraction module in the current channel, the input feature of M1 is the initial feature. Perform the second feature extraction on the initial feature through M1 to obtain the initial output feature m1 out-tmp of M1. Add m11 to the input feature of M1 (i.e., the initial feature) to obtain the output feature m1 out of M1. At this time, determine whether M1 is the last feature extraction module in the current channel. After judgment, M1 is not the last feature extraction module. Therefore, use m1 out as the input feature of the next feature extraction module M2, and repeat the above second feature extraction operation. The operation processes of M2, M3, and M4 are the same as that of M1 above and will not be elaborated here. When the output feature m4 out is obtained through M4 out , after judgment, M4 is the last feature extraction module in the current channel. Therefore, fuse m4 out with the input feature of the first feature extraction module in the first layer (i.e., the initial feature) and the input feature of the first feature extraction module in the second layer (i.e., the input feature m2 out ) of M3 to obtain the output feature of the current channel.
[0074] The output image obtained by the multi-feature extraction neural network provided in the above embodiments of the present invention can reflect the features of the low-frequency image to be processed from multiple perspectives and obtain a high-frequency image with a relatively high resolution. For example, Figure 5 is a schematic comparison diagram of the low-frequency image and the predicted high-frequency image of the multi-feature extraction neural network, Figure 5 in which the large left image is the low-frequency image, LR in the right image is the low-frequency feature to be predicted, the four images of Branch1 - Branch4 are the four output images of the multi-feature extraction neural network, and HR is the finally obtained high-frequency image.
[0075] In some possible implementation manners, the step of fusing the multiple output images to obtain the high-frequency image corresponding to the low-frequency image may specifically be to determine the weight of each output image through a fusion network; perform weighted summation on each output image according to the corresponding weight to obtain the high-frequency image.
[0076] Specifically, the fusion network can be designed as two convolutional layers with a convolutional kernel size of 3*3 and a softmax layer. In this fusion network, the multiple output images obtained by the multi-feature extraction neural network are convolved twice together to convert the multiple output images into the same dimension, and then passed through the softmax layer to determine the fusion weight of each output image. According to the fusion weight map, the four output images are weighted pixel by pixel respectively. Further, the weighted images are added pixel by pixel to generate the final fusion output image, that is, the high-frequency image.
[0077] Figure 6 is a schematic diagram of the fusion of four output images by the fusion network of the embodiments of the present invention. For example, Figure 6 as shown, the large left image is the low-frequency image, and the four small right images are the four output images output by the multi-feature extraction neural network. The thicker the box area in the image, the higher the corresponding weight during fusion, that is, the clearer part has a larger weight value, and the more blurred part has a smaller weight value. In this way, the obtained fusion image takes more into account the clear parts in each output image, ensuring the accuracy of the fusion image.
[0078] Figure 7 is a schematic diagram of the high-frequency image obtained by the image processing method provided in the embodiments of the present invention. For example, Figure 7 as shown, the left side of the straight line is the original low-frequency image, and the right side is the high-frequency image obtained by the above method. It can be seen that the image processing method provided in the embodiments of the present invention can effectively obtain a high-frequency image with a relatively high resolution.
[0079] The method for processing low-frequency images provided by the embodiments of the present invention obtains a high-frequency image by predicting high-frequency features of the low-frequency image based on a first neural network. The accuracy of the high-frequency image depends to a large extent on the stability of the multi-feature extraction neural network and the accuracy of the output results. Therefore, on the basis of the above method, the embodiments of the present invention also provide a training method for the multi-feature extraction neural network. Refer to Figure 8 the flowchart of the training method for the multi-feature extraction neural network shown in
[0080] S802: Obtain a sample image and an initial neural network;
[0081] Among them, the sample image includes a training image and its corresponding label image, and the resolution of the label image is higher than that of the training image;
[0082] S804: Perform third feature extraction on the sample image to obtain the training features corresponding to the training image;
[0083] S806: Perform fourth feature extraction on the training features through the initial neural network to obtain multiple predicted features;
[0084] Among them, the resolution of the predicted features is higher than that of the training features.
[0085] S808: Determine the predicted image corresponding to the predicted features according to the predicted features;
[0086] Specifically, it may be to perform upsampling processing on the multiple predicted features respectively to obtain the predicted images corresponding to the multiple predicted features.
[0087] S810: Determine the total loss value according to the predicted image and its corresponding label image.
[0088] S812: Adjust the parameters of the multi-feature extraction neural network according to the total loss value until the first training stop condition is satisfied.
[0089] In some possible implementation manners, the above multi-feature extraction neural network includes multiple layers, and each layer includes multiple feature extraction modules. The feature extraction modules in the previous layer are connected to at least one of the feature extraction modules in the next layer. The interconnected feature extraction modules located in different layers form a channel of the multi-feature extraction neural network;
[0090] Based on the structure of the above multi-feature extraction neural network, the step of performing fourth feature extraction on the training features through the multi-feature extraction neural network in S806 to obtain multiple predicted features may specifically be:
[0091] (1) The input features are subjected to fourth feature extraction by the current feature extraction module to obtain the initial output features corresponding to this feature extraction module; wherein, if the current feature extraction module is the first feature extraction module in the said channel, the input features are the training features, otherwise, the input features are the output features of the previous feature extraction module;
[0092] (2) The initial output features are combined with the input features corresponding to the current feature extraction module as the output features of the current feature extraction module;
[0093] (3) Determine whether the current feature extraction module is the last feature extraction module in the current channel; if not, continue to perform the fourth feature extraction operation;
[0094] (4) If so, fuse the output features of the current feature extraction module with the input features of the first feature extraction module in each layer of the current channel, and determine the fused result as the predicted features corresponding to the current channel.
[0095] In some possible implementation manners, the total loss value in step S810 above can be determined according to the following method:
[0096] Based on each predicted image and the label image, obtain the first loss value corresponding to this predicted image; based on every two predicted images and the label image, obtain the second loss value corresponding to these two predicted images; determine the total loss value according to the first loss value and the second loss value.
[0097] Specifically, the total loss value of the multi-feature extraction neural network is determined by the first loss value and the second loss value. The first loss value can specifically be the L2 loss (mean squared error) loss value, and the second loss value can specifically be the Triplet loss (triplet loss) loss value. Calculate the L2 loss value for each output image and the label image respectively, and then sum up these L2 loss values to form the overall L2 loss. Calculate the Triplet loss value pairwise between each output image.
[0098] In some possible implementation manners, the calculation process of the above second loss value is as follows:
[0099] (1) Determine the first positive sample according to the label image, determine the first negative sample according to the first predicted image among the two predicted images, and determine the first anchor sample according to the second predicted image among the two predicted images;
[0100] (2) Calculate the triplet loss value through the first positive sample, the first negative sample, and the first anchor sample to obtain the first triplet loss value;
[0101] (3) Determine the second positive sample according to the labeled image, determine the second anchor sample according to the first predicted image, and determine the second negative sample according to the second predicted image;
[0102] (4) Calculate the triplet loss value through the second positive sample, the second negative sample, and the second anchor sample to obtain the second triplet loss value;
[0103] (5) Add the first triplet loss value and the second triplet loss value to obtain the second loss value.
[0104] Specifically, the above second loss value is the loss value obtained by the triplet loss function. The input of the triplet loss function tripletLoss includes positive samples, negative samples, and anchor samples. In order to achieve the effect that the difference between any two images is large, while the difference between these two images and the labeled image is small, the triplet loss value can be calculated between any two output images. Specifically, calculate the triplet loss value twice for any two output images respectively, and then combine these two triplet loss values as the triplet loss value of these two images. For example, for Image 1 and Image 2, first use the labeled image as the positive sample, Image 1 as the negative sample, and Image 2 as the anchor sample to calculate the triplet loss value to obtain loss1. Then use the labeled image as the positive sample, Image 1 as the anchor sample, and Image 2 as the negative sample to calculate the triplet loss value once to obtain loss2. Then loss1 + loss2 is the second loss value corresponding to Image 1 and Image 2.
[0105] Because the super-resolution task is mainly to restore high-frequency textures, if the triplet loss is directly used to calculate the loss value between output images, other dissimilation may occur in addition to textures, such as color, brightness, etc. Therefore, based on the above method, the present invention further improves the triplet loss, determines the triplet loss value based on the grayscale information of the output images. Based on this, the determination process of the above second loss value further includes:
[0106] Extract the first grayscale value of the first predicted image and the second grayscale value of the second predicted image; determine the first residual according to the first grayscale value and the labeled image; determine the second residual according to the second grayscale value and the labeled image;
[0107] Based on this, the step of determining the first positive sample according to the labeled image, determining the first negative sample according to the first predicted image among the two predicted images, and determining the first anchor sample according to the second predicted image among the two predicted images includes:
[0108] Determine the first residual as the first negative sample; determine the second residual as the first anchor sample;
[0109] Further, the steps of determining the second positive sample according to the labeled image, determining the second anchor sample according to the first predicted image, and determining the second negative sample according to the second predicted image include:
[0110] Determine the first residual as the second anchor sample; determine the second residual as the second negative sample.
[0111] Figure 9 The figure is a schematic diagram of the calculation process of a second loss value provided by an embodiment of the present invention. As Figure 9 shown, represents the i-th output image, represents the j-th output image, where i is not equal to j. I HR represents the labeled image, that is, the high-frequency image. First, the three images are grayscaled to obtain the gray Y channel of the images (the Y channel in YUV is the gray channel), that is, Y HR and Gray-scaling the features can prevent hue alienation. Then, the gray channels are normalized through the Norm layer to highlight texture alienation and prevent brightness alienation.
[0112] After that, the values of the gray channels of image i and image j are respectively subtracted from the labeled image to obtain residuals, that is, and And calculate the triplet loss value according to the residuals. The triplet loss value obtained based on the residuals can further highlight texture alienation.
[0113] In some possible implementation manners, the above multi-feature extraction neural network in the embodiment of the present invention is an M*N tree topology structure, where M represents the number of layers of the multi-feature extraction neural network, and N represents the number of initial modules included in the first layer of the multi-feature extraction neural network. Each initial module in the first layer and the middle layer of the multi-feature extraction neural network is connected to N initial modules in the next layer.
[0114] In some possible implementation manners, the above fusion network can be trained by the following method:
[0115] (1) Obtain an initial fusion network; (2) Determine a predicted fusion image according to each predicted image and its corresponding initial weight; (3) Determine a third loss value according to the predicted fusion image and the labeled image; (4) Adjust the parameters of the initial fusion network according to the third loss value until the second training stop condition is satisfied.
[0116] The above-mentioned initial weights can be calculated through the softmax layer in the fusion network. Each output image is weighted and summed according to the initial weights to obtain a fused image. The loss value is calculated between the fused image and the label image to obtain a third loss value. Any general loss function can be used in the calculation process of the loss value, and the embodiments of the present invention do not limit this. The parameters of the initial fusion network are continuously adjusted according to the third loss function until the second training stop condition is met, and the training of the initial fusion network is completed. Among them, the second training stop condition can be the number of training times or the third loss value is less than a preset value.
[0117] For ease of understanding, the following specifically describes the training process of the multi-feature extraction neural network and the fusion network provided by the present invention in combination with a specific application scenario. The training process specifically includes the following steps:
[0118] (1) Obtain a low-frequency image Fig1 and a high-frequency image Fig2, and perform feature extraction on the low-frequency image to obtain initial features.
[0119] (2) Set the initial structure and parameters of the multi-feature extraction neural network;
[0120] Among them, the multi-feature extraction neural network includes two layers, and each layer includes 2 RG modules, that is, 2 RG modules stacked to form a branch. Each RG is composed of 4 RCAB modules. The upsampling module uses a pixel shuffle module.
[0121] (3) Input the initial image into the multi-feature extraction neural network. In the first RG module of the first layer, the input of the first RCAB module is the initial feature, and the output is feature 1. The input of the second RCAB module is the combined feature 1 obtained by combining the initial feature and feature 1, and the output is feature 2.
[0122] (4) Enter the first RCAB module of the first RG module in the second layer, and the input is the combined feature 2 obtained by combining the combined feature 1 and feature 2. The input determination methods of other RCAB modules are similar to those of the above modules.
[0123] (5) Continue to calculate the output values of other modules until the outputs of all modules in the last layer are obtained, that is, four output images S1, S2, S3, and S4.
[0124] (6) Calculate the triplet loss values between S1 - S4 pairwise to obtain six loss values Tloss1 - Tloss6.
[0125] (7) Calculate the loss values loss1 - loss4 between S1 - S4 and the high-frequency image Fig2 respectively.
[0126] (8) Add up all of loss1 - loss4 and Tloss1 - Tloss6 to obtain the total loss value.
[0127] (9) Determine whether the total loss value is less than the preset loss value. If so, the multi - feature extraction neural network training is completed; otherwise, repeat the above steps (3) - (8).
[0128] (10) Process the initial features through the trained multi - feature extraction neural network to obtain four high - frequency images S11, S21, S31, and S41.
[0129] (11) Set the initial structure and parameters of the fusion network;
[0130] The fusion network includes two convolutional layers with a convolutional kernel size of 3x3 and a softmax layer.
[0131] (12) Input the four images S11 - S41 into the fusion network to obtain the fused image M1;
[0132] (13) Calculate the loss value between the fused image M1 and the high - frequency image Fig2, and determine whether the loss value is less than the preset fusion loss value. If so, the training is completed; otherwise, repeat the above steps (10) - (11).
[0133] To further verify the beneficial technical effects of the image processing method provided by the embodiments of the present invention, the technical effects of the present invention are verified from three aspects as follows.
[0134] First, compare the method provided by the embodiments of the present invention with various methods in the prior art in terms of experimental results. The low - frequency image processing methods in the prior art include the Bicubic upsampling method, the single - branch methods SRResNet, EDSR, and RCAN with direct input and output, the method ESRGAN using GAN, the multi - output method SRFlow, and the methods LP - KPN and CDC based on the divide - and - conquer idea. Using PSNR (Peak Signal to Noise Ratio) and SSIM (Structural SIMilarity) as comparison metrics, the comparison results are shown in Table 1:
[0135] Table 1
[0136]
[0137] Among them, the RealSR dataset is obtained by zooming with two types of single-lens reflex cameras, Canon and Nikon. There are 595 scenes collected in the RealSR dataset. It is divided into three super-resolution magnification factors (x2, x3, x4). The total number of training data is 1,265, and 100 test data for each magnification factor are randomly selected.
[0138] To verify the universality of the model, the present invention is trained on RealSR and the effect is verified on DRealSR. The DRealSR dataset is another real dataset collected using multiple single-lens reflex cameras.
[0139] D2CRealSR is a dataset with an x8 magnification factor collected by a Sony single-lens reflex camera, and the effect comparison is carried out on this dataset. In Table 1, Ours is the experimental result obtained by the method provided in the embodiment of the present invention. It can be seen from Table 1 that the image processing method provided in the embodiment of the present invention has achieved good objective indicators (PSNR, SSIM) in all experiments.
[0140] Second, the high-frequency images obtained by the method provided in the embodiment of the present invention and the methods in the prior art are compared in terms of the visualization accuracy of the images. The comparison results are as Figure 10 shown. Figure 10 The figure respectively shows the comparison diagrams of the high-frequency images obtained by high-frequency prediction for a part of the low-frequency image of Canon_045 (x4). Among them, Our represents the high-frequency image obtained by the image processing method provided in the embodiment of the present invention. Compared with other prior arts, the picture clarity of the high-frequency image corresponding to our is significantly higher.
[0141] Third, Figure 11 Regarding the number of network parameters used by each method and the image effect when obtaining high-frequency images, the black circles in the figure represent the number of network parameters and indicators used by the method provided in the embodiment of the present invention, and the circles with black outer frames represent the number of network parameters and indicators of other methods. The horizontal axis of the figure represents the size of the number of parameters, and the vertical axis represents the PSNR index at the x4 magnification factor of the RealSR dataset. For the convenience of viewing, the larger the circle, the larger the number of parameters. It can be seen that the methods provided in the embodiments of the present invention are all located in the upper left corner, and the circles are the smallest, indicating that our method has a small number of parameters and at the same time has a better super-resolution effect.
[0142] Based on the above method embodiments, the embodiments of the present invention also provide an image processing device. Refer to Figure 12 shown, the device includes:
[0143] An acquisition module 1202, configured to acquire an image to be processed;
[0144] A processing module 1204 is configured to perform high-frequency feature prediction processing on an image to be processed through a multi-feature extraction neural network, so as to obtain a plurality of output images with different high-frequency features; wherein, the similarity between any two output images is less than a first similarity.
[0145] A fusion module 1206 is configured to fuse the plurality of output images to obtain a high-frequency image corresponding to the image to be processed; wherein, the resolution of the high-frequency image is higher than that of the image to be processed.
[0146] The above image processing device provided by the embodiments of the present invention first obtains an image to be processed, performs high-frequency feature prediction processing on the image to be processed through a multi-feature extraction neural network to obtain a plurality of output images with different high-frequency features, and then fuses the plurality of output images to obtain a high-frequency image corresponding to the image to be processed. Since the plurality of output images with different high-frequency features can reflect the features of different aspects of the low-frequency image, the accuracy of the high-frequency image can be effectively improved based on the high-frequency image obtained from the plurality of output images with different high-frequency features.
[0147] The above multi-feature extraction neural network includes multiple layers, and each layer includes a plurality of feature extraction modules. The feature extraction modules of the previous layer are connected to at least one feature extraction module of the next layer. The mutually connected feature extraction modules located in different layers constitute a channel of the multi-feature extraction neural network. The process of performing high-frequency feature prediction processing on the image to be processed through the multi-feature extraction neural network to obtain a plurality of output images with different high-frequency features includes: performing first feature extraction on the image to be processed through the multi-feature extraction neural network to obtain initial features; performing second feature extraction on the initial features through the current channel in the multiple channels of the multi-feature extraction neural network to obtain output features corresponding to the channel; wherein, the resolution of the output features corresponding to any one channel is higher than the resolution of the initial features; and respectively performing upsampling processing on the output features of each channel to obtain a plurality of output images.
[0148] The process of performing second feature extraction on the initial features through the current channel among multiple channels of the multi-feature extraction neural network to obtain the output features corresponding to this channel includes: performing second feature extraction on the input features through the current feature extraction module to obtain the initial output features corresponding to this feature extraction module; where if the current feature extraction module is the first feature extraction module in the channel, the input features are the initial features, otherwise, the input features are the output features of the previous feature extraction module; combining the initial output features and the input features corresponding to the current feature extraction module as the output features of the current feature extraction module; determining whether the current feature extraction module is the last feature extraction module in the current channel; if not, continue to perform the second feature extraction operation; if so, fuse the output features of the current feature extraction module with the input features of the first feature extraction module in each layer of the current channel, and determine the fused result as the output features corresponding to the current channel.
[0149] The steps of fusing multiple output images to obtain the high-frequency image corresponding to the image to be processed include: determining the weights of each output image through the fusion network; performing weighted summation on each output image according to the corresponding weights to obtain the high-frequency image.
[0150] The above multi-feature extraction neural network is trained in the following manner: obtaining sample images and an initial neural network; where the sample images include training images and their corresponding label images, and the resolution of the label images is higher than that of the training images; performing third feature extraction on the sample images to obtain the training features corresponding to the training images; performing fourth feature extraction on the training features through the initial neural network to obtain multiple prediction features; the resolution of the prediction features is higher than that of the training features; determining the prediction images corresponding to the prediction features according to the prediction features; determining the total loss value according to the prediction images and their corresponding label images; adjusting the parameters of the multi-feature extraction neural network according to the total loss value until the first training stop condition is met.
[0151] The above multi-feature extraction neural network includes multiple layers, and each layer includes multiple feature extraction modules. The feature extraction modules in the previous layer are connected to at least one feature extraction module in the next layer. The interconnected feature extraction modules located in different layers form a channel of the multi-feature extraction neural network. The process of performing the fourth feature extraction on the training features through the multi-feature extraction neural network to obtain multiple predicted features includes: performing the fourth feature extraction on the input features through the current feature extraction module to obtain the initial output features corresponding to the feature extraction module. Wherein, if the current feature extraction module is the first feature extraction module in the channel, the input features are the training features; otherwise, the input features are the output features of the previous feature extraction module. Combining the initial output features with the input features corresponding to the current feature extraction module as the output features of the current feature extraction module. Determining whether the current feature extraction module is the last feature extraction module in the current channel. If not, continue to perform the fourth feature extraction operation. If so, fusing the output features of the current feature extraction module with the input features of the first feature extraction module in each layer of the current channel, and determining the fusion result as the predicted features corresponding to the current channel.
[0152] The process of determining the total loss value based on the predicted image and its corresponding labeled image includes: obtaining the first loss value corresponding to the predicted image according to each predicted image and the labeled image; obtaining the second loss value corresponding to the two predicted images according to every two predicted images and the labeled image; determining the total loss value according to the first loss value and the second loss value.
[0153] The process of obtaining the second loss value corresponding to the two predicted images according to every two predicted images and the labeled image includes: determining the first positive sample according to the labeled image, determining the first negative sample according to the first predicted image among the two predicted images, and determining the first anchor sample according to the second predicted image among the two predicted images; calculating the triplet loss value through the first positive sample, the first negative sample, and the first anchor sample to obtain the first triplet loss value; determining the second positive sample according to the labeled image, determining the second anchor sample according to the first predicted image, and determining the second negative sample according to the second predicted image; calculating the triplet loss value through the second positive sample, the second negative sample, and the second anchor sample to obtain the second triplet loss value; adding the first triplet loss value and the second triplet loss value to obtain the second loss value.
[0154] The above device further includes: a grayscale value extraction module, configured to extract a first grayscale value of the first predicted image and a second grayscale value of the second predicted image; a first residual determination module, configured to determine a first residual according to the first grayscale value and the label image; a second residual determination module, configured to determine a second residual according to the second grayscale value and the label image; the process of determining the first positive sample according to the label image, determining the first negative sample according to the first predicted image among the two predicted images, and determining the first anchor sample according to the second predicted image among the two predicted images includes: determining the first residual as the first negative sample; determining the second residual as the first anchor sample; the process of determining the second positive sample according to the label image, determining the second anchor sample according to the first predicted image, and determining the second negative sample according to the second predicted image includes: determining the first residual as the second anchor sample; determining the second residual as the second negative sample.
[0155] The above multi-feature extraction neural network has a tree-shaped topology of M*N, where M represents the number of layers of the multi-feature extraction neural network, and N represents the number of initial modules included in the first layer of the multi-feature extraction neural network. Each initial module in the first layer and the intermediate layer of the multi-feature extraction neural network is connected to N initial modules in the next layer.
[0156] The above fusion network is obtained by training through the following method; obtaining an initial fusion network; determining a predicted fusion image according to each predicted image and its corresponding initial weight; determining a third loss value according to the predicted fusion image and the label image; adjusting the parameters of the initial fusion network according to the third loss value until the second training stop condition is satisfied.
[0157] The image processing device provided by the embodiments of the present invention has the same implementation principle and the same technical effects as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the embodiments of the above device, reference may be made to the corresponding content in the foregoing image processing method embodiments.
[0158] The embodiments of the present invention also provide an electronic device, as Figure 13 shown, is a schematic structural diagram of the electronic device. Among them, the electronic device includes a processor 1301 and a memory 1302. The memory 1302 stores computer-executable instructions that can be executed by the processor 1301, and the processor 1301 executes the computer-executable instructions to implement the above image processing method.
[0159] In Figure 13 the illustrated embodiment, the electronic device further includes a bus 1303 and a communication interface 1304. Among them, the processor 1301, the communication interface 1304, and the memory 1302 are connected through the bus 1303.
[0160] Among them, the memory 1302 may include high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk memory. The communication connection between this system network element and at least one other network element is achieved through at least one communication interface 1304 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 1303 can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The bus 1303 can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 8 only a two-way arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0161] The processor 1301 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 1301 or by instructions in the form of software. The above-mentioned processor 1301 can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor 1301 reads the information in the memory and combines its hardware to complete the steps of the image processing method in the foregoing embodiments.
[0162] An embodiment of the present invention also provides a computer-readable storage medium storing computer-executable instructions, which, when called and executed by a processor, cause the processor to implement the above image processing method. For specific implementation, reference can be made to the foregoing method embodiment and will not be elaborated herein.
[0163] A computer program product of the image processing method, apparatus, and electronic device provided by an embodiment of the present invention includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the method described in the foregoing method embodiment. For specific implementation, reference can be made to the method embodiment and will not be elaborated herein.
[0164] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0165] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program code.
[0166] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and should not be construed as indicating or implying relative importance.
[0167] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any technician familiar with the technical field of the present invention can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. An image processing method, characterized in that, the method includes: obtaining an image to be processed; performing high-frequency feature prediction processing on the image to be processed through a multi-feature extraction neural network to obtain multiple output images with different high-frequency features; wherein, the similarity between any two of the output images is less than a first similarity; fusing the multiple output images to obtain a high-frequency image corresponding to the image to be processed; wherein, the resolution of the high-frequency image is higher than that of the image to be processed; wherein, the multi-feature extraction neural network includes multiple layers, each layer includes multiple feature extraction modules, the feature extraction modules of the previous layer are connected to at least one of the feature extraction modules of the next layer, and the interconnected feature extraction modules located in different layers constitute a channel of the multi-feature extraction neural network; the step of performing high-frequency feature prediction processing on the image to be processed through a multi-feature extraction neural network to obtain multiple output images with different high-frequency features includes: performing first feature extraction on the image to be processed through the multi-feature extraction neural network to obtain initial features; performing second feature extraction on the initial features through the current channel among the multiple channels of the multi-feature extraction neural network to obtain output features corresponding to the channel; wherein, the resolution of the output features corresponding to any one of the channels is higher than the resolution of the initial features; performing upsampling processing on the output features of each channel respectively to obtain multiple output images; wherein, the step of performing second feature extraction on the initial features through the current channel among the multiple channels of the multi-feature extraction neural network to obtain output features corresponding to the channel includes: performing second feature extraction on the input features through the current feature extraction module to obtain initial output features corresponding to the feature extraction module; wherein, if the current feature extraction module is the first feature extraction module in the channel, the input feature is the initial feature, otherwise, the input feature is the output feature of the previous feature extraction module; combining the initial output features and the input features corresponding to the current feature extraction module as the output features of the current feature extraction module; judging whether the current feature extraction module is the last feature extraction module in the current channel; if not, continue to perform the second feature extraction operation; if so, fusing the output features of the current feature extraction module with the input features of the first feature extraction module of each layer in the current channel, and determining the fusion result as the output features corresponding to the current channel.
2. The method according to claim 1, characterized in that, the step of fusing the multiple output images to obtain a high-frequency image corresponding to the image to be processed includes: determining the weight of each output image through a fusion network; performing weighted summation on each output image according to the corresponding weight to obtain a high-frequency image.
3. The method according to claim 1, characterized in that, the multi-feature extraction neural network is obtained through the following training method: Obtain a sample image and an initial neural network; wherein, the sample image includes a training image and its corresponding label image, and the resolution of the label image is higher than that of the training image; Perform third feature extraction on the sample image to obtain training features corresponding to the training image; Perform fourth feature extraction on the training features through the initial neural network to obtain multiple prediction features; the resolution of the prediction features is higher than that of the training features; Determine a prediction image corresponding to the prediction features according to the prediction features; Determine a total loss value according to the prediction image and its corresponding label image; Adjust the parameters of the multi-feature extraction neural network according to the total loss value until the first training stop condition is met.
4. The method according to claim 3, wherein, The multi-feature extraction neural network includes multiple layers, and each layer includes multiple feature extraction modules. The feature extraction modules in the previous layer are connected to at least one of the feature extraction modules in the next layer. The interconnected feature extraction modules located in different layers form a channel of the multi-feature extraction neural network; The step of performing fourth feature extraction on the training features through the multi-feature extraction neural network to obtain multiple prediction features includes: Perform fourth feature extraction on the input features through the current feature extraction module to obtain initial output features corresponding to the feature extraction module; wherein, if the current feature extraction module is the first feature extraction module in the channel, the input features are the training features, otherwise, the input features are the output features of the previous feature extraction module; Combine the initial output features with the input features corresponding to the current feature extraction module as the output features of the current feature extraction module; Determine whether the current feature extraction module is the last feature extraction module in the current channel; if not, continue to perform the fourth feature extraction operation; If so, fuse the output features of the current feature extraction module with the input features of the first feature extraction module in each layer of the current channel, and determine the fusion result as the prediction features corresponding to the current channel.
5. The method according to claim 3, wherein, The step of determining a total loss value according to the prediction image and its corresponding label image includes: Obtain a first loss value corresponding to the prediction image according to each prediction image and the label image; Obtain a second loss value corresponding to the two prediction images according to every two prediction images and the label image; Determine a total loss value according to the first loss value and the second loss value.
6. The method according to claim 5, wherein, The step of obtaining a second loss value corresponding to the two prediction images according to every two prediction images and the label image includes: Determine a first positive sample according to the label image, determine a first negative sample according to the first prediction image among the two prediction images, and determine a first anchor sample according to the second prediction image among the two prediction images; Calculate the triplet loss value through the first positive sample, the first negative sample, and the first anchor sample to obtain the first triplet loss value; Determine the second positive sample according to the label image, determine the second anchor sample according to the first predicted image, and determine the second negative sample according to the second predicted image; Calculate the triplet loss value through the second positive sample, the second negative sample, and the second anchor sample to obtain the second triplet loss value; Add the first triplet loss value and the second triplet loss value to obtain the second loss value.
7. The method according to claim 6, wherein, the method further includes: Extract the first grayscale value of the first predicted image and the second grayscale value of the second predicted image; Determine the first residual according to the first grayscale value and the label image; Determine the second residual according to the second grayscale value and the label image; The step of determining the first positive sample according to the label image, the first negative sample according to the first predicted image among the two predicted images, and the first anchor sample according to the second predicted image among the two predicted images includes: Determine the first residual as the first negative sample; Determine the second residual as the first anchor sample; The step of determining the second positive sample according to the label image, the second anchor sample according to the first predicted image, and the second negative sample according to the second predicted image includes: Determine the first residual as the second anchor sample; Determine the second residual as the second negative sample.
8. The method according to any one of claims 1-7, wherein, The multi-feature extraction neural network has a tree topology of M*N, M represents the number of layers of the multi-feature extraction neural network, N represents the number of initial modules included in the first layer of the multi-feature extraction neural network, and each of the initial modules in the first layer and the intermediate layer of the multi-feature extraction neural network is connected to N initial modules in the next layer.
9. The method according to claim 3, wherein, The fusion network is obtained by training through the following method; Obtain an initial fusion network; Determine a predicted fusion image according to each predicted image and its corresponding initial weight; Determine a third loss value according to the predicted fusion image and the label image; Adjust the parameters of the initial fusion network according to the third loss value until the second training stop condition is satisfied.
10. An electronic device, wherein, It includes a processor and a memory, the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method according to any one of claims 1 to 9.
11. A computer-readable storage medium, wherein, The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the method according to any one of claims 1 to 9.
12. A computer program product, including a computer program, wherein, When the computer program is executed by a processor, it implements the method according to any one of claims 1-9.
Citation Information
Patent Citations
Visual image enhancement generation method, system and device and storage medium
CN113066013A