Low-resolution image processing method, electronic device and computer program product
By aligning and fusion processing of feature maps for multi-frame low-resolution images, the problem of unstable alignment effect under the influence of noise in the prior art is solved, and the stability and accuracy of high-resolution images are improved.
Patent Information
- Application Number
- CN202111111165.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-18
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-09-18
AI Technical Summary
The prior art fails to effectively consider the influence of noise in the super-resolution processing of multi-frame low-resolution images, resulting in unstable alignment effect and low super-segment image accuracy.
By obtaining the reference feature map and the initial feature map set of the video frame sequence to be processed, the reference feature map is used to align the initial feature map set, generate the alignment feature map set, and fuse the alignment feature map set and the reference feature map to finally obtain a high-resolution image through reconstruction.
Effectively remove noise in the initial feature map, improve image alignment, and thus improve the stability and accuracy of high-resolution images.
Smart Images

Figure CN114202457B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a low-resolution image processing method, electronic equipment and computer program product. Background Art
[0002] Super-resolution is a widely studied problem. Its task is to generate a corresponding high-resolution image given a low-resolution image. According to the form of the low-resolution image input, super-resolution can be divided into single-frame image super-resolution and multi-frame image super-resolution.
[0003] Multi-frame super-resolution is the reconstruction of the original high-definition image using multiple low-resolution images. Low-resolution images are usually taken by handheld smartphones in multi-frame mode. In multi-frame mode, different image frames will change due to the shaking of the mobile phone camera or changes in light. Therefore, the quality of multi-frame low-resolution images is often low and has a lot of noise. In addition, the jitter between multi-frame images is large. Due to the poor image quality, it is difficult to align them using traditional methods.
[0004] In the prior art, deformable convolution is used to implicitly estimate the motion field of video frames, that is, a multi-level resolution deformable convolution module is used to solve the problem of aligning multiple frames of images. However, the influence of noise is not considered in the alignment process, and the alignment effect is very unstable, resulting in low accuracy of the final super-resolution result. Summary of the invention
[0005] In view of this, an object of the present invention is to provide a low-resolution image processing method, electronic device and computer program product to improve the accuracy of a high-resolution image obtained through multiple frames of low-resolution images.
[0006] In a first aspect, an embodiment of the present invention provides a method for processing a low-resolution image, the method comprising: obtaining a reference feature map and an initial feature map set corresponding to a video frame sequence to be processed; wherein the reference feature map is a feature map corresponding to a reference video frame in the video frame sequence to be processed, and each initial feature map in the initial feature map set is a feature map corresponding to a video frame other than the reference video frame in the video frame sequence to be processed; aligning each initial feature map in the initial feature map set according to the reference feature map to obtain an aligned feature map set corresponding to the initial feature map set; wherein each aligned feature map in the aligned feature map set corresponds to an initial feature map, and the similarity between the aligned feature map and the reference feature map is greater than the similarity between the initial feature map and the reference feature map; fusing the aligned feature map set and the reference feature map to obtain a fused feature map; reconstructing the fused feature map to obtain a target image; wherein the resolution of the target image is higher than the resolution of all video frames in the video frame to be processed.
[0007] Furthermore, the step of aligning each initial feature map in the initial feature map set according to the reference feature map to obtain an aligned feature map set corresponding to the initial feature map set includes: performing downsampling and upsampling on the reference feature map, respectively, to obtain a first upsampling feature corresponding to the reference feature map; performing downsampling and upsampling on the current initial feature map, respectively, to obtain a second upsampling feature corresponding to the initial feature map; determining the aligned feature map corresponding to the initial feature map according to the current initial feature map, the first upsampling feature and the second upsampling feature; and counting the aligned feature maps corresponding to all the initial feature maps to generate an aligned feature map set.
[0008] Further, the above-mentioned step of downsampling and upsampling the reference feature map respectively to obtain the first upsampling feature corresponding to the reference feature map includes: downsampling the reference feature map a first preset number of times to obtain a first preset number of first intermediate features; upsampling the first intermediate feature with the smallest size a first preset number of times to obtain the first upsampling feature; the above-mentioned step of downsampling and upsampling the initial feature map in the initial feature map set respectively to obtain the second upsampling feature corresponding to the initial feature map includes: downsampling the initial feature map in the initial feature map set a first preset number of times to obtain a first preset number of second intermediate features; upsampling the second intermediate feature with the smallest size a first preset number of times to obtain the second upsampling feature.
[0009] Furthermore, the step of performing upsampling for a first preset number of times on the second intermediate feature with the smallest size to obtain the second upsampled feature includes: combining the output feature of the last upsampling and the second intermediate feature with the same size as the output feature to obtain a first combined feature; wherein the input of the first upsampling is the second intermediate feature with the smallest size; the first combined feature corresponding to the last upsampling is a combination of the output feature of the last upsampling and the initial feature map; determining whether the first preset number of upsampling has been performed, and if so, determining the first combined feature as the second upsampled feature; otherwise, using the first combined feature as the input of the current upsampling and continuing to perform upsampling operations on the first combined feature.
[0010] Furthermore, the step of determining the alignment feature map corresponding to the initial feature map based on the initial feature map, the first up-sampled feature and the second up-sampled feature includes: determining the offset corresponding to the second intermediate feature based on each second intermediate feature and the first intermediate feature having the same size as the first intermediate feature; determining the initial alignment feature corresponding to the second intermediate feature based on the second intermediate feature and the offset corresponding to the second intermediate feature; and determining the alignment feature map based on the initial alignment feature.
[0011] Furthermore, the above-mentioned step of determining the initial alignment feature corresponding to the second intermediate feature based on the second intermediate feature and the offset corresponding to the second intermediate feature includes: performing a first convolution operation on the second intermediate feature and the offset corresponding to the second intermediate feature to obtain a first convolution feature corresponding to the second intermediate feature; and determining the initial alignment feature corresponding to the second intermediate feature based on the first convolution feature.
[0012] Furthermore, the above-mentioned step of determining the alignment feature map based on the initial alignment feature includes: performing the following upsampling operation on the initial alignment feature for a first preset number of times: combining the output of the previous upsampling and the initial alignment feature of the same size as the output to obtain a second combined feature; wherein the input of the first upsampling is the initial alignment feature of the smallest size; determining whether the upsampling operation has been performed for the first preset number of times, and if so, determining the second combined feature as the alignment feature map; otherwise, continuing the upsampling operation using the second combined feature as the input of the current upsampling.
[0013] Furthermore, the above-mentioned step of fusing the aligned feature atlas set and the reference feature map to obtain a fused feature map includes: determining a fusion weight of the aligned feature map based on the reference feature map and the aligned feature map; and performing weighted summation of the reference feature map and the aligned feature map based on the fusion weight to obtain a fused feature map.
[0014] Furthermore, the above-mentioned step of determining the fusion weight of the aligned feature map based on the reference feature map and the aligned feature map includes: determining the global weight corresponding to each aligned feature map based on the reference feature map; normalizing each global weight to obtain the fusion weight corresponding to the aligned feature map.
[0015] Furthermore, the above-mentioned step of reconstructing the fused feature map to obtain the target image includes: performing the following second convolution operation on the fused feature map at least twice to obtain the second convolution feature: obtaining the current third combined feature; wherein, the current third combined feature corresponding to the first second convolution operation is the fused feature map, and the current third combined feature corresponding to the non-first second convolution operation is determined by combining the combination of the second convolution features obtained by all the executed second convolution operations with the fused feature map; performing the second convolution operation on the third combined feature to obtain the second convolution feature; continuing to perform the second convolution operation until the number of second convolution operations reaches a second preset number, and determining the second convolution feature of the most recent second convolution operation as the second convolution feature corresponding to the fused feature map; and determining the target image according to the second convolution feature corresponding to the fused feature map.
[0016] In a second aspect, an embodiment of the present invention further provides an electronic device, including a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the low-resolution image processing method of the first aspect mentioned above.
[0017] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the low-resolution image processing method of the first aspect mentioned above.
[0018] In a fourth aspect, an embodiment of the present invention further provides a computer program product, including a computer program, which, when executed by a processor, implements the low-resolution image processing method of the first aspect. The low-resolution image processing method, electronic device, and computer program product provided by the embodiment of the present invention process the reference frame in the video frame sequence to be processed to obtain a reference feature map corresponding to the video frame sequence to be processed, process the video frames other than the reference frame in the video frame sequence to be processed to obtain an initial feature map set corresponding to the video frame sequence to be processed, align each of the initial feature maps in the initial feature map set according to the reference feature map to obtain an aligned feature map set corresponding to the initial feature map set, fuse the aligned feature map set and the reference feature map to obtain a fused feature map, and reconstruct the fused feature map to obtain a target image. The present invention aligns the initial feature map with the reference feature map, and performs image fusion and reconstruction based on the aligned feature map after alignment, effectively removes noise in the initial feature map, improves the alignment effect, and thereby improves the stability and image accuracy of the target image determined according to the reference feature map and the aligned feature map.
[0019] Other features and advantages of the present disclosure will be set forth in the following description, or some features and advantages may be inferred or unambiguously determined from the description, or may be learned by implementing the above-mentioned technology of the present disclosure.
[0020] In order to make the above-mentioned objectives, features and advantages of the present disclosure more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0022] Figure 1 A schematic diagram of the structure of an electronic system provided by an embodiment of the present invention;
[0023] Figure 2 A flowchart of a low-resolution image processing method provided by an embodiment of the present invention;
[0024] Figure 3 A flowchart of another low-resolution image processing method provided by an embodiment of the present invention;
[0025] Figure 4 A schematic diagram of the structure of an attention network provided by an embodiment of the present invention;
[0026] Figure 5 A schematic diagram of a reconstruction network structure provided by an embodiment of the present invention;
[0027] Figure 6 A schematic diagram of a process of a low-resolution image processing method provided by an embodiment of the present invention;
[0028] Figure 7 A schematic diagram of a low-resolution image processing device provided by an embodiment of the present invention;
[0029] Figure 8 A visual comparison diagram of experimental results obtained by the low-resolution image processing method provided by an embodiment of the present invention and a processing method in the prior art;
[0030] Fig. 9 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0032] In recent years, research on computer vision, deep learning, machine learning, image processing, image recognition and other technologies based on artificial intelligence has made important progress. Artificial Intelligence (AI) is an emerging science and technology that studies and develops theories, methods, technologies and application systems for simulating and extending human intelligence. Artificial intelligence is a comprehensive discipline involving many types of technologies such as chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, neural networks, etc. Computer vision, as an important branch of artificial intelligence, specifically allows machines to recognize the world. Computer vision technology usually includes face recognition, liveness detection, fingerprint recognition and anti-counterfeiting verification, biometric recognition, face detection, pedestrian detection, target detection, pedestrian recognition, image processing, image recognition, image semantic understanding, image retrieval, text recognition, video processing, video content recognition, behavior recognition, 3D reconstruction, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), computational photography, robot navigation and positioning and other technologies. With the research and advancement of artificial intelligence technology, this technology has been applied in many fields, such as security, urban management, traffic management, building management, park management, facial access, facial attendance, logistics management, warehouse management, robots, intelligent marketing, computational photography, mobile phone imaging, cloud services, smart homes, wearable devices, unmanned driving, automatic driving, smart medical care, facial payment, facial unlocking, fingerprint unlocking, identity verification, smart screens, smart TVs, cameras, mobile Internet, live streaming, beauty, makeup, medical beauty, and smart temperature measurement.
[0033] The current super-resolution image processing method does not consider the influence of image noise on the alignment effect and image accuracy during the process. Based on this, the embodiments of the present invention provide a low-resolution image processing method, device and electronic device to improve the image accuracy of a high-resolution image obtained by using multiple frames of low-resolution images.
[0034] Reference Figure 1 The electronic system 100 is a schematic structural diagram of the electronic system 100. The electronic system can be used to implement the low-resolution image processing method and device of the embodiment of the present invention.
[0035] like Figure 1 The electronic system 100 includes one or more processing devices 102, one or more storage devices 104, an input device 106, an output device 108, and one or more image acquisition devices 110. These components are interconnected via a bus system 112 and / or other forms of connection mechanisms (not shown). It should be noted that Figure 1 The components and structure of the electronic system 100 shown are merely exemplary and not limiting. The electronic system may also have other components and structures as required.
[0036] The processing device 102 can be a server, an intelligent terminal, or a device including a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities. It can process data of other components in the electronic system 100 and can also control other components in the electronic system 100 to perform low-resolution image processing functions.
[0037] The storage device 104 may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processing device 102 may run the program instructions to implement the client functions and / or other desired functions in the embodiments of the present invention (implemented by the processing device) described below. Various applications and various data, such as various data used and / or generated by the application, may also be stored in the computer-readable storage medium.
[0038] The input device 106 may be a device used by a user to input instructions, and may include one or more of a keyboard, a mouse, a microphone, a touch screen, and the like.
[0039] The output device 108 may output various information (eg, images or sounds) to the outside (eg, a user), and may include one or more of a display, a speaker, and the like.
[0040] The image acquisition device 110 may acquire a video frame sequence to be processed, and store the video frame sequence in the storage device 104 for use by other components.
[0041] Exemplarily, the various devices in the method, device and electronic device for implementing the low-resolution image processing according to the embodiment of the present invention can be integrated or dispersed, such as integrating the processing device 102, the storage device 104, the input device 106 and the output device 108 into one body, and setting the image acquisition device 110 at a designated position where the image can be acquired. When the various devices in the above electronic system are integrated, the electronic system can be implemented as a smart terminal such as a camera, a smart phone, a tablet computer, a computer, a vehicle-mounted terminal, etc.
[0042] Figure 2 A flowchart of a low-resolution image processing method provided by an embodiment of the present invention is shown in FIG. Figure 2 , the method comprises the following steps:
[0043] S202: Obtain a reference feature map and an initial feature map set corresponding to the video frame sequence to be processed;
[0044] The sequence of video frames to be processed is a plurality of low-resolution images acquired by a camera device, for example, a plurality of photos taken continuously by a mobile phone of a target object. Since the multiple frames of images may be offset due to the shaking of the camera device, the embodiment of the present invention first determines a reference video frame in the sequence of video frames to be processed. The reference video frame may be a frame of image with the highest resolution, a frame of image selected randomly, or the first frame of image. After the reference video frame is determined, all other frame images must be aligned to the reference frame to achieve alignment of multiple frame images. Feature extraction is performed on the reference video frame to obtain a reference feature map.
[0045] Feature extraction is performed on other video frames in the video frame sequence to obtain the corresponding initial feature map of each video frame. The initial feature maps corresponding to all other frame images except the reference frame constitute the initial feature map set corresponding to the video frame sequence to be processed.
[0046] It should be noted that the reference video frame may be determined by taking the first video frame obtained by shooting as the reference video frame, or taking the middle video frame as the reference video frame, which is not limited in the embodiment of the present invention.
[0047] Specifically, a neural network may be used to acquire the reference feature map and the initial feature map set, for example, a residual network may be used. The present invention does not limit the specific method of feature extraction.
[0048] S204: performing alignment processing on each initial feature map in the initial feature map set according to the reference feature map to obtain an aligned feature map set corresponding to the initial feature map set;
[0049] In this step, each initial feature map in the initial feature map set is aligned to obtain an aligned feature map corresponding to the initial feature map. All aligned feature maps constitute an aligned feature map set. The purpose of alignment is to reduce the difference between the initial feature map and the reference feature map. Therefore, the similarity between the aligned feature map and the reference feature map is greater than the similarity between the initial feature map and the reference feature map.
[0050] The specific alignment method will be described in detail below and will not be repeated here.
[0051] S206: Fusing the aligned feature atlas and the reference feature map to obtain a fused feature map;
[0052] S208: Reconstruct the fused feature map to obtain a target image.
[0053] The above-mentioned low-resolution image processing method provided by an embodiment of the present invention first obtains a reference feature map and an initial feature map set corresponding to a video frame sequence to be processed, aligns each of the initial feature maps in the initial feature map set according to the reference feature map to obtain an aligned feature map set corresponding to the initial feature map set, then fuses the aligned feature map set and the reference feature map to obtain a fused feature map, and finally reconstructs the fused feature map to obtain a target image. The initial feature map is aligned with the reference feature map, and image fusion and reconstruction are performed based on the aligned aligned feature map, which effectively removes noise in the initial feature map and improves the effect of image alignment, thereby improving the stability and image accuracy of the target image determined according to the reference feature map and the aligned feature map.
[0054] In some possible implementations, each alignment feature in the above alignment feature atlas can be obtained by the following method:
[0055] (1) performing downsampling and upsampling on the reference feature map to obtain a first upsampling feature corresponding to the reference feature map;
[0056] The reference feature map is downsampled and upsampled the same number of times, and the output feature of the last upsampling operation is the first upsampled feature. For example, the reference feature map can be downsampled twice, and then upsampled twice based on the features obtained by the second downsampling, and finally the first upsampled feature is obtained. Since the size of the feature map obtained by the downsampling operation gradually decreases, and the size of the feature map obtained by the upsampling operation gradually increases, the process of downsampling and upsampling can be called a feature pyramid, and features of different scales can be extracted through the feature pyramid. The step size of each sampling in the feature pyramid is the same, for example, a convolution operation with a step size of 2.
[0057] (2) performing downsampling and upsampling on the current initial feature map to obtain a second upsampling feature corresponding to the initial feature map;
[0058] Similar to the processing of the reference feature map, each initial feature map in the initial feature map set is subjected to downsampling and upsampling respectively, and the output feature of the last upsampling process is the second upsampled feature corresponding to the initial feature map. In order to achieve a better effect of aligning the initial feature map with the reference feature map, the number of upsampling in the initial feature map is the same as the number of upsampling in the reference feature map, the number of downsampling in the initial feature map is the same as the number of downsampling in the reference feature map, and the step size of each sampling operation in the initial feature map is the same as the step size of each sampling operation in the reference feature map.
[0059] (3) determining an alignment feature map corresponding to the initial feature map according to the current initial feature map, the first up-sampled feature, and the second up-sampled feature;
[0060] The aligned feature map is a feature map obtained by transforming the initial feature map, and the obtained feature map has a smaller deviation from the reference feature map, that is, the aligned feature map is more similar to the reference feature map. Therefore, the similarity between the aligned feature map and the reference feature map is greater than the similarity between the initial feature map and the reference feature map.
[0061] After obtaining the aligned feature map corresponding to each initial feature map, the aligned feature maps corresponding to all the initial feature maps are counted to generate an aligned feature map set.
[0062] In some possible implementations, the method for determining the first upsampling feature may specifically be:
[0063] The reference feature map is downsampled a first preset number of times to obtain a first preset number of first intermediate features; and the first intermediate feature with the smallest size is upsampled a first preset number of times to obtain a first upsampled feature.
[0064] For the sake of ease of description, N1 is used to represent the first preset number of times. For example, N1 can be set to 2. Then the above process can be specifically: down-sampling the reference feature map f0 for the first time to obtain the first intermediate feature f01, down-sampling f01 for the second time to obtain the first intermediate feature f02, that is, obtaining two first intermediate features f01 and f02.
[0065] Furthermore, f02 is upsampled for the first time to obtain the second feature f03, and feature f03 is upsampled for the second time to obtain feature f04, which is the first upsampled feature.
[0066] In some possible implementations, the process of determining the second upsampling feature may specifically be:
[0067] (1) downsampling the initial feature graph in the initial feature graph set a first preset number of times to obtain a first preset number of second intermediate features;
[0068] (2) upsampling the second intermediate feature with the smallest size for a first preset number of times to obtain a second upsampled feature; (3) combining the output feature of the last upsampling and the second intermediate feature with the same size as the output feature to obtain a first combined feature;
[0069] In order to avoid losing the features of the initial feature map during the upsampling process, before each upsampling, the output feature of the last upsampling and the second intermediate feature of the same size as the output feature are combined to obtain a first combined feature; wherein the input of the first upsampling is the second intermediate feature with the smallest size; the first combined feature corresponding to the last upsampling is the combination of the output feature of the last upsampling and the initial feature map; (4) determining whether the first preset number of upsamplings has been performed, and if so, determining the first combined feature as the second upsampling feature; (5) otherwise, using the first combined feature as the input of the current upsampling, and continuing to perform upsampling operations on the first combined feature.
[0070] Continuing with the previous example, the initial feature map f1 is downsampled for the first time to obtain the second intermediate feature f11, and f11 is downsampled for the second time to obtain the second intermediate feature f12. The second intermediate feature f12 with the smallest size is upsampled for the first time to obtain the output feature f13. The feature f11 is the same as f13 in the second intermediate features {f11, f12}, so f11 is combined with f13 to obtain the first combined feature f14, and f14 is upsampled for the second time to obtain the output feature f15. Since N1 upsamplings are performed, the second upsampling is the last upsampling. f15 is combined with the initial feature map f1 to obtain the first combined feature f16, and f16 is the second upsampled feature.
[0071] It should be noted that the first upsampled feature can be determined in the same way for the reference feature map:
[0072] (1) downsampling the reference feature map a first preset number of times to obtain a first preset number of intermediate features;
[0073] (2) combining the output feature of the last upsampling and the intermediate feature of the same size as the output feature to obtain a combined feature; wherein the input of the first upsampling is the intermediate feature with the smallest size; and the combined feature corresponding to the last upsampling is the combination of the output feature of the last upsampling and the reference feature map;
[0074] (3) determining whether upsampling has been performed for a first preset number of times, and if so, determining the combined feature as a first upsampling feature;
[0075] (4) Otherwise, the combined feature is used as the input of the current upsampling, and the upsampling operation is continued on the combined feature.
[0076] After obtaining the first intermediate features corresponding to the reference feature maps of different sizes and the second intermediate features corresponding to the initial feature maps, the alignment feature map of each initial feature map can be determined based on these intermediate features. Specifically:
[0077] (1) determining an offset corresponding to each second intermediate feature according to each second intermediate feature and a first intermediate feature having the same size as the first intermediate feature;
[0078] Since the reference feature map and the initial feature map are upsampled and downsampled the same number of times and have the same step size, each second intermediate feature has a first intermediate feature of the same size. Based on this, each second intermediate feature and the first intermediate feature of the same size are processed to obtain the offset corresponding to the second feature, that is, each second feature has a corresponding offset.
[0079] (2) performing a first convolution operation on the second intermediate feature and the offset corresponding to the second intermediate feature to obtain a first convolution feature corresponding to the second intermediate feature;
[0080] A first convolution operation is performed on each second intermediate feature and its corresponding offset to obtain a first convolution feature corresponding to the second intermediate feature.
[0081] In some possible implementations, the first convolution operation may be a deformable convolution operation with a non-fixed convolution kernel shape. The multi-layer cascaded deformable convolution used in the embodiment of the present invention can more accurately align image features from multiple scales.
[0082] (3) Determine an initial alignment feature map corresponding to the second intermediate feature based on the first convolution feature.
[0083] It can be understood that each second intermediate feature corresponds to an initial alignment feature map, that is, multiple initial alignment feature maps can be obtained.
[0084] (4) Performing the following upsampling operation on the initial alignment features for a first preset number of times:
[0085] Starting the upsampling operation from the smallest initial alignment feature and fusing the output of the previous upsampling at the end of each upsampling process can further eliminate the noise in the initial feature map and improve the stability and accuracy of the final output alignment feature.
[0086] Specifically, the alignment feature map can be determined in the following manner:
[0087] 4-1: Combine the output of the last upsampling and the initial alignment feature of the same size as the output to obtain a second combined feature; wherein the input of the first upsampling is the initial alignment feature of the smallest size;
[0088] 4-2: Determine whether a first preset number of upsampling operations are performed, and if so, determine the second combined feature as an alignment feature;
[0089] 4-3: Otherwise, the second combined feature is used as the input of the current upsampling to continue the upsampling operation.
[0090] The above-mentioned embodiment of the present invention obtains features of different scales by downsampling twice, and upsampling starts from the minimum scale and adds it to the features of the previous layer to obtain multi-scale features with rich information. Due to the interpolation operation in the up- and down-sampling process, the first up-sampled features and the second up-sampled features finally obtained are denoised, and finally the features of different frames are aligned by features of different scales, thereby achieving the effect of aggregating information, and the obtained aggregated information effectively removes the noise of the image in the video frame image sequence.
[0091] After the initial feature map is processed to obtain multiple aligned feature maps, the aligned feature map needs to be fused with the reference feature to obtain a final high-resolution image. In order to enhance the fusion effect, the embodiment of the present invention further provides another low-resolution image processing method based on the above method, see Figure 3 , the method specifically comprises:
[0092] S302: Obtain a reference feature map and an initial feature map set corresponding to the video frame sequence to be processed;
[0093] S304: performing alignment processing on each initial feature map in the initial feature map set according to the reference feature map to obtain an aligned feature map set corresponding to the initial feature map set;
[0094] The above process S302-S304 may refer to steps S202-S204 in the embodiment of the present invention, which will not be described in detail here.
[0095] S306: Determine a fusion weight of the aligned feature map according to the reference feature map and the aligned feature map;
[0096] Specifically, the fusion weight can be determined as follows:
[0097] (1) Determine the global weight corresponding to each aligned feature map according to the reference feature map;
[0098] (2) Normalize each global weight to obtain the fusion weight corresponding to the alignment feature.
[0099] Among them, the normalization process can be completed by using a normalization algorithm or a normalized neural network, for example, normalization is performed through an attention mechanism.
[0100] S308: performing weighted summation on the reference feature map and the alignment feature map according to the fusion weight to obtain a fused feature map;
[0101] Specifically, the attention network can be used to obtain the fusion feature map, such as Figure 4As shown in the schematic diagram of the attention network structure, the reference feature map and the initial feature map are first reduced in dimension through 1x1 convolution, and the reference feature map and the initial feature map are inner-producted to obtain the global weight between each initial feature map and the reference feature map. All global weights are converted into normalized weights through softmax operation, and then the normalized weights are multiplied by the corresponding reference feature map and initial feature map, and the feature dimension is restored through 1x1 convolution, and the corresponding initial feature map is added in the form of residual to obtain the final output fused feature map.
[0102] The method provided by the embodiment of the present invention uses non-local relationship information to fuse features after aligning multiple frame features. Since each frame has a spatial relationship with the reference frame, and the image itself has large jitter and errors may occur in the alignment stage, this spatial relationship is a non-local response. Based on this, the cross-frame attention network used in the present invention extracts non-local information between each frame and the reference frame to enhance the fusion result, which can further eliminate the noise of the low-resolution image and improve the fusion effect.
[0103] S310: Reconstruct the fused feature map to obtain a target image.
[0104] After obtaining the fused feature map, the fused feature map is further reconstructed to obtain a high-resolution target image. The target image can be determined as follows:
[0105] (1) Perform the following second convolution operation on the fused features at least twice to obtain the second convolution feature:
[0106] Obtain the current third combined feature; wherein the current third combined feature corresponding to the first second convolution operation is a fusion feature, and the current third combined feature corresponding to the non-first second convolution operation is a combination of second convolution features obtained by performing the second convolution operation; perform the second convolution operation on the third combined feature to obtain a second convolution feature; continue to perform the second convolution operation until the number of second convolution operations reaches a second preset number, and determine the second convolution feature of the most recent second convolution operation as the second convolution feature corresponding to the fusion feature;
[0107] (2) Determine the target image based on the second convolution feature corresponding to the fused feature map.
[0108] Specifically, a reconstruction network can be used to reconstruct the fused feature map, such as Figure 5 The reconstruction network structure diagram shown in Figure 1, where LRCG stands for long range information aggregation group, in which the input of each module is the input of all previous inputs aggregated together and then reduced in dimension through 1x1 convolution. WARB stands for wide activation residual network, in which the input is first increased in dimension through 1x1 convolution, and then reduced in dimension after passing through the activation function.
[0109] The entire convolution module includes multiple LRCG sub-modules, each LRCG sub-module contains 2 WARB sub-modules. The fused feature is input into the first WARG of the first LRCG to obtain feature 1, and then feature 1 is connected with the input fused feature map to obtain the input of the second WARB and perform convolution operation again. The second WARB outputs feature 2, and feature 2 is connected with feature 1 and the fused feature to obtain the output of the first LRCG, and so on, until the final output, that is, the high-resolution image, is obtained.
[0110] The convolution method of long-term information in the embodiment of the present invention avoids the loss of information during the reconstruction process and utilizes multi-layer semantic features to further improve the accuracy of high-resolution images.
[0111] For ease of understanding, the following Figure 6 This paper introduces the application scenario flow of a low-resolution image processing method. Figure 6 The process diagram of the low-resolution image processing method provided by the present invention is as follows: the method takes the video frame to be processed as an example, which includes 3 images, and the feature maps corresponding to the three images are fig1, fig2 and fig3, respectively, where fig1 is a reference feature map. Figure 6 As shown, the method includes:
[0112] Step 1: Downsample fig1 twice to get fig11 and fig12, and downsample fig2 twice to get fig21 and fig22.
[0113] Step 2: Upsample fig12 to get fig13, and upsample fig22 to get fig23.
[0114] Step 3: Combine fig13 and fig11 to get fig14, and combine fig23 and fig21 to get fig24.
[0115] Step 4: Continue upsampling fig14 to obtain fig15, and continue upsampling fig24 to obtain fig25.
[0116] Step 5: Combine fig15 and fig1 to get fig16, and combine fig25 and fig2 to get fig26. Fig26 is the second upsampling feature, and the upsampling operation is completed.
[0117] Step 6: For the smallest downsampled features fig12 and fig22, calculate the deviation between the two to obtain D1, perform deformable convolution (DCN) operation on D1 and fig22 to obtain the aligned feature map A1.
[0118] Step 7: Perform the same operation as step 6 on fig14 and fig24 to obtain the aligned feature map A2. Similarly, perform the same operation as step 6 on fig16 and fig26 to obtain the aligned feature map A3.
[0119] Step 8: Upsample A1 to obtain A11.
[0120] Step 9: Combine A11 and A2 to obtain A12, and continue to upsample A12 to obtain A13.
[0121] Step 10: Combine A13 with A3 to obtain the final aligned feature map Aout1.
[0122] Step 11: For fig3, use the same operation as fig2 to obtain the alignment feature map Aout2 corresponding to fig3.
[0123] Step 12: Fuse fig1, Aout1, and Aout2 to obtain the fused feature map Xout.
[0124] Specifically, convolution operations and inner product operations can be performed on fig1, Aout1 and Aout2 to obtain the corresponding global weights, the global weights can be softmaxed to obtain normalized weights, and the weighted sum of fig1, Aout1 and Aout2 can be performed to obtain the fused feature Xout.
[0125] Step 13: Reconstruct the long-term information of the fused feature map Xout to obtain a high-resolution image.
[0126] Based on the above method embodiment, the present invention also provides a low-resolution image processing device, see Figure 7 As shown, the device comprises:
[0127] The acquisition module 702 is used to acquire a reference feature map and an initial feature map set corresponding to the video frame sequence to be processed; wherein the reference feature map is a feature map corresponding to a reference video frame in the video frame sequence to be processed, and each initial feature map in the initial feature map set is a feature map corresponding to a video frame other than the reference video frame in the video frame sequence to be processed;
[0128] An alignment module 704 is used to align each initial feature map in the initial feature map set according to the reference feature map to obtain an aligned feature map set corresponding to the initial feature map set; wherein each aligned feature map in the aligned feature map set corresponds to an initial feature map, and the similarity between the aligned feature map and the reference feature map is greater than the similarity between the initial feature map and the reference feature map;
[0129] A fusion module 706 is used to fuse the aligned feature atlas and the reference feature map to obtain a fused feature map;
[0130] The reconstruction module 708 is used to reconstruct the fused feature map to obtain a target image; wherein the resolution of the target image is higher than the resolution of all video frames in the video frame to be processed.
[0131] The above-mentioned low-resolution image processing device provided by an embodiment of the present invention first obtains a reference feature map and an initial feature map set corresponding to a sequence of video frames to be processed, performs alignment processing on each of the initial feature maps in the initial feature map set according to the reference feature map, obtains an aligned feature map set corresponding to the initial feature map set, then fuses the aligned feature map set and the reference feature map to obtain a fused feature map, and finally reconstructs the fused feature map to obtain a target image. The present invention performs alignment processing on the initial feature map through the reference feature map, and performs image fusion and reconstruction based on the aligned feature map after alignment, effectively removes noise in the initial feature map, improves the alignment effect, and further improves the stability and image accuracy of the target image determined according to the reference feature map and the aligned feature map.
[0132] The above process of aligning each initial feature map in the initial feature map set according to the reference feature map to obtain the aligned feature map set corresponding to the initial feature map set includes: performing downsampling and upsampling on the reference feature map respectively to obtain the first upsampling feature corresponding to the reference feature map; performing downsampling and upsampling on the current initial feature map respectively to obtain the second upsampling feature corresponding to the initial feature map; determining the aligned feature map corresponding to the initial feature map according to the current initial feature map, the first upsampling feature and the second upsampling feature; and counting the aligned feature maps corresponding to all the initial feature maps to generate an aligned feature map set.
[0133] The above-mentioned process of downsampling and upsampling the reference feature map respectively to obtain the first upsampling feature corresponding to the reference feature map includes: downsampling the reference feature map a first preset number of times to obtain a first preset number of first intermediate features; upsampling the first intermediate feature with the smallest size a first preset number of times to obtain the first upsampling feature; the above-mentioned process of downsampling and upsampling the initial feature map in the initial feature map set respectively to obtain the second upsampling feature corresponding to the initial feature map includes: downsampling the initial feature map in the initial feature map set a first preset number of times to obtain a first preset number of second intermediate features; upsampling the second intermediate feature with the smallest size a first preset number of times to obtain the second upsampling feature.
[0134] The above-mentioned process of upsampling the second intermediate feature with the smallest size for a first preset number of times to obtain the second upsampled feature includes: combining the output feature of the last upsampling and the second intermediate feature with the same size as the output feature to obtain a first combined feature; wherein the input of the first upsampling is the second intermediate feature with the smallest size; the first combined feature corresponding to the last upsampling is a combination of the output feature of the last upsampling and the initial feature map; determining whether the first preset number of upsampling has been performed, and if so, determining the first combined feature as the second upsampled feature; otherwise, using the first combined feature as the input of the current upsampling and continuing to perform upsampling operations on the first combined feature.
[0135] The above process of determining the alignment feature map corresponding to the initial feature map based on the initial feature map, the first up-sampled feature and the second up-sampled feature includes: determining the offset corresponding to the second intermediate feature based on each second intermediate feature and the first intermediate feature of the same size as the first intermediate feature; determining the initial alignment feature corresponding to the second intermediate feature based on the second intermediate feature and the offset corresponding to the second intermediate feature; and determining the alignment feature map based on the initial alignment feature.
[0136] The above process of determining the initial alignment feature corresponding to the second intermediate feature based on the second intermediate feature and the offset corresponding to the second intermediate feature includes: performing a first convolution operation on the second intermediate feature and the offset corresponding to the second intermediate feature to obtain a first convolution feature corresponding to the second intermediate feature; and determining the initial alignment feature corresponding to the second intermediate feature based on the first convolution feature.
[0137] The above process of determining the alignment feature map based on the initial alignment feature includes: performing the following upsampling operation on the initial alignment feature for a first preset number of times: combining the output of the previous upsampling and the initial alignment feature of the same size as the output to obtain a second combined feature; wherein the input of the first upsampling is the initial alignment feature with the smallest size; determining whether the upsampling operation has been performed for the first preset number of times, and if so, determining the second combined feature as the alignment feature map; otherwise, continuing the upsampling operation using the second combined feature as the input of the current upsampling.
[0138] The above-mentioned process of fusing the aligned feature atlas set and the reference feature map to obtain the fused feature map includes: determining the fusion weight of the aligned feature map according to the reference feature map and the aligned feature map; and performing weighted summation of the reference feature map and the aligned feature map according to the fusion weight to obtain the fused feature map.
[0139] The above process of determining the fusion weight of the aligned feature map based on the reference feature map and the aligned feature map includes: determining the global weight corresponding to each aligned feature map based on the reference feature map; normalizing each global weight to obtain the fusion weight corresponding to the aligned feature map.
[0140] The above-mentioned process of reconstructing the fused feature map to obtain the target image includes: performing the following second convolution operation on the fused feature map at least twice to obtain the second convolution feature: obtaining the current third combined feature; wherein, the current third combined feature corresponding to the first second convolution operation is the fused feature map, and the current third combined feature corresponding to the non-first second convolution operation is determined by combining the combination of the second convolution features obtained by all the executed second convolution operations with the fused feature map; performing the second convolution operation on the third combined feature to obtain the second convolution feature; continuing to perform the second convolution operation until the number of second convolution operations reaches a second preset number, and determining the second convolution feature of the most recent second convolution operation as the second convolution feature corresponding to the fused feature map; and determining the target image according to the second convolution feature corresponding to the fused feature map.
[0141] The low-resolution image processing device provided in the embodiment of the present invention has the same implementation principle and technical effects as those of the aforementioned method embodiment. For the sake of brief description, for parts not mentioned in the embodiment of the above-mentioned device, reference may be made to the corresponding contents in the aforementioned low-resolution image processing method embodiment.
[0142] In order to further verify the beneficial technical effects of the low-resolution image processing method provided by the embodiment of the present invention, the present invention conducts experiments on synthetic RAW domain images and real RAW domain image data, visually comparing the Figure 8 As shown, Figure 8 The left side is the original image, the right side is the low-resolution image, and the four advanced models used for comparison are EDSR, RRDB, WDSR and RCAN, as well as the EDVR model in video super-resolution. HR is the actual high-resolution image. Figure 8 It can be seen that the high-resolution image obtained by the method provided in the embodiment of the present invention is closest to the real HR image and has high image accuracy.
[0143] The following Table 1 is the test results obtained by testing various models in the prior art and the method provided by the embodiment of the present invention, wherein PSNR, SSIM and LPIPS are all indicators for evaluating the quality of the result image. The larger the PSNR and SSIM values, the higher the resolution of the result image, and the smaller the LPIPS value, the higher the resolution of the result image.
[0144] Table 1
[0145]
[0146] The embodiment of the present invention further provides an electronic device, such as Fig. 9 As shown, it is a schematic diagram of the structure of the electronic device, wherein the electronic device includes a processor 901 and a memory 902, the memory 902 stores computer executable instructions that can be executed by the processor 901, and the processor 901 executes the computer executable instructions to implement the above-mentioned low-resolution image processing method.
[0147] exist Fig. 9 In the illustrated embodiment, the electronic device further includes a bus 903 and a communication interface 904 , wherein the processor 901 , the communication interface 904 and the memory 902 are connected via the bus 903 .
[0148] Among them, the memory 902 may include a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 904 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 903 can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 903 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig. 9 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0149] The processor 901 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit in the processor 901 or the instruction in the form of software. The above processor 901 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiment of the present invention can be directly embodied as a hardware decoding processor to execute, or the hardware and software modules in the decoding processor can be executed. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor 901 reads the information in the memory and completes the steps of the low-resolution image processing method of the above-mentioned embodiment in combination with its hardware.
[0150] The embodiment of the present invention further provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions prompt the processor to implement the above-mentioned low-resolution image processing method. The specific implementation can refer to the aforementioned method embodiment, which will not be repeated here. The low-resolution image processing method and the computer program product of the electronic device provided by the embodiment of the present invention include a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the method described in the aforementioned method embodiment. The specific implementation can refer to the method embodiment, which will not be repeated here.
[0151] Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0152] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0153] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.
[0154] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for processing a low-resolution image, characterized in that: The method comprises: Obtaining a reference feature map and an initial feature map set corresponding to a sequence of video frames to be processed; wherein the reference feature map is a feature map corresponding to a reference video frame in the sequence of video frames to be processed, and each initial feature map in the initial feature map set is a feature map corresponding to a video frame in the sequence of video frames to be processed except the reference video frame; Performing alignment processing on each of the initial feature maps in the initial feature map set according to the reference feature map to obtain an aligned feature map set corresponding to the initial feature map set; wherein each aligned feature map in the aligned feature map set corresponds to one of the initial feature maps, and the similarity between the aligned feature map and the reference feature map is greater than the similarity between the initial feature map and the reference feature map; Fusing the aligned feature atlas and the reference feature map to obtain a fused feature map; Reconstructing the fused feature map to obtain a target image; wherein the resolution of the target image is higher than the resolution of all video frames in the video frame to be processed; The step of reconstructing the fused feature map to obtain a target image includes: performing the following second convolution operation on the fused feature map at least twice to obtain a second convolution feature: obtaining a current third combined feature; wherein the current third combined feature corresponding to the first second convolution operation is the fused feature map, and the current third combined feature corresponding to the non-first second convolution operation is determined by combining a combination of second convolution features obtained by all executed second convolution operations with the fused feature map; performing a second convolution operation on the third combined feature to obtain a second convolution feature; continuing to perform the second convolution operation until the number of second convolution operations reaches a second preset number, and determining the second convolution feature of the most recent second convolution operation as the second convolution feature corresponding to the fused feature map; and determining the target image according to the second convolution feature corresponding to the fused feature map.
2. The method according to claim 1, characterized in that The step of performing alignment processing on each of the initial feature maps in the initial feature map set according to the reference feature map to obtain an aligned feature map set corresponding to the initial feature map set comprises: Performing downsampling and upsampling on the reference feature map respectively to obtain a first upsampling feature corresponding to the reference feature map; Performing downsampling and upsampling processing on the current initial feature map respectively to obtain a second upsampling feature corresponding to the initial feature map; Determine an alignment feature map corresponding to the initial feature map according to the current initial feature map, the first up-sampled feature, and the second up-sampled feature; The aligned feature maps corresponding to all the initial feature maps are counted to generate an aligned feature map set.
3. The method according to claim 2, characterized in that The step of performing downsampling processing and upsampling processing on the reference feature map respectively to obtain a first upsampling feature corresponding to the reference feature map includes: Downsampling the reference feature map a first preset number of times to obtain a first preset number of first intermediate features; Upsampling the first intermediate feature with the smallest size by the first preset number of times to obtain a first upsampled feature; The step of performing downsampling processing and upsampling processing on the initial feature graph in the initial feature graph set to obtain a second upsampling feature corresponding to the initial feature graph comprises: Downsampling the initial feature graph in the initial feature graph set a first preset number of times to obtain a first preset number of second intermediate features; The second intermediate feature with the smallest size is upsampled the first preset number of times to obtain a second upsampled feature.
4. The method according to claim 3, characterized in that The step of performing upsampling of the second intermediate feature with the smallest size by the first preset number of times to obtain a second upsampled feature comprises: The output feature of the last upsampling and the second intermediate feature of the same size as the output feature are combined to obtain a first combined feature; wherein the input of the first upsampling is the second intermediate feature with the smallest size; the first combined feature corresponding to the last upsampling is a combination of the output feature of the last upsampling and the initial feature map; Determining whether the first preset number of upsamplings has been performed, and if so, determining the first combined feature as a second upsampling feature; Otherwise, the first combined feature is used as the input of the current upsampling, and the upsampling operation is continued on the first combined feature.
5. The method according to claim 3, characterized in that: The step of determining an alignment feature map corresponding to the initial feature map according to the initial feature map, the first up-sampled feature, and the second up-sampled feature comprises: Determine an offset corresponding to each of the second intermediate features according to each of the second intermediate features and the first intermediate feature having the same size as the first intermediate feature; Determine an initial alignment feature corresponding to the second intermediate feature according to the second intermediate feature and the offset corresponding to the second intermediate feature; An alignment feature map is determined according to the initial alignment feature.
6. The method according to claim 5, characterized in that The step of determining an initial alignment feature corresponding to the second intermediate feature according to the second intermediate feature and the offset corresponding to the second intermediate feature comprises: Performing a first convolution operation on the second intermediate feature and the offset corresponding to the second intermediate feature to obtain a first convolution feature corresponding to the second intermediate feature; An initial alignment feature corresponding to the second intermediate feature is determined according to the first convolution feature.
7. The method according to claim 5, characterized in that The step of determining an alignment feature map according to the initial alignment feature comprises: The initial alignment feature is subjected to the following upsampling operation for the first preset number of times: The output of the last upsampling and the initial alignment feature of the same size as the output are combined to obtain a second combined feature; wherein the input of the first upsampling is the initial alignment feature of the smallest size; Determining whether the first preset number of upsampling operations are performed, and if so, determining the second combined feature as an alignment feature map; Otherwise, the upsampling operation is continued by taking the second combined feature as the input of the current upsampling.
8. The method according to any one of claims 1 to 7, characterized in that The step of fusing the alignment feature atlas and the reference feature map to obtain a fused feature map comprises: The fusion weight of the aligned feature map is determined according to the reference feature map and the aligned feature map; and the reference feature map and the aligned feature map are weightedly summed according to the fusion weight to obtain a fused feature map.
9. The method according to claim 8, characterized in that The step of determining a fusion weight of the aligned feature map according to the reference feature map and the aligned feature map comprises: Determine a global weight corresponding to each of the aligned feature maps according to the reference feature map; Each of the global weights is normalized to obtain a fusion weight corresponding to the aligned feature map.
10. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the method according to any one of claims 1 to 9.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the method according to any one of claims 1 to 9.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and storage medium
CN110070511A
Video super-resolution processing method and device, and storage medium
CN112700392A
Image beautifying processing method and device, storage medium and electronic equipment
CN113077397A