Image enhancement processing method and storage medium
Patent Information
- Application Number
- CN202210913936.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2042-07-29
AI Technical Summary
而通常用户所观看视频的视频质量参差不齐,如部分视频的视频分辨率较低,影响用户观看效果,基于此,如何提高视频的视频分辨率成为了当前的研究热点
[0017] This application embodiment can acquire an image to be enhanced and a reference image of the image to be enhanced. The reference image can be obtained based on the previous frame of the image to be enhanced in the target video. Then, the image to be enhanced and the reference image can be fused to obtain a fused image. Further, feature extraction can be performed on the image to be enhanced to obtain the target image features of the image to be enhanced, and feature extraction can be performed on the fused image to obtain the enhanced image features of the image to be enhanced. Thus, the enhanced image of the image to be enhanced can be reconstructed based on the target image features and the enhanced image features. The image resolution of the enhanced image is greater than that of the image to be enhanced. In this way, the image resolution can be improved, and thus the video resolution can also be improved.
Smart Images

Figure CN117541486B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular to an image enhancement processing method and storage medium. Background Technology
[0002] With the rapid development of multimedia technology, a massive amount of video data has emerged, making video playback a common operation in users' daily lives. However, the quality of videos watched by users often varies greatly; for example, some videos have low resolution, affecting the viewing experience. Therefore, improving video resolution has become a current research hotspot. Summary of the Invention
[0003] This application provides an image enhancement processing method and storage medium that can improve image resolution and, consequently, video resolution.
[0004] In a first aspect, embodiments of this application provide an image enhancement processing method, including:
[0005] Obtain the image to be enhanced and a reference image of the image to be enhanced; the reference image is obtained based on the previous frame of the image to be enhanced in the target video;
[0006] The fusion unit is used to fuse the image to be enhanced and the reference image to obtain a fused image;
[0007] Feature extraction is performed on the image to be enhanced to obtain the target image features of the image to be enhanced, and feature extraction is performed on the fused image to obtain the enhanced image features of the image to be enhanced;
[0008] An enhanced image is reconstructed from the target image features and the enhanced image features; the image resolution of the enhanced image is greater than that of the image to be enhanced.
[0009] Secondly, embodiments of this application provide an image enhancement processing apparatus, comprising:
[0010] An acquisition unit is used to acquire the image to be enhanced and a reference image of the image to be enhanced; the reference image is obtained based on the previous frame image of the image to be enhanced in the target video;
[0011] The fusion unit is used to fuse the image to be enhanced and the reference image to obtain a fused image;
[0012] The extraction unit is used to extract features from the image to be enhanced to obtain the target image features of the image to be enhanced, and to extract features from the fused image to obtain the enhanced image features of the image to be enhanced;
[0013] The reconstruction unit is used to reconstruct an enhanced image of the image to be enhanced based on the target image features and the enhanced image features; the image resolution of the enhanced image is greater than the image resolution of the image to be enhanced.
[0014] Thirdly, embodiments of this application provide a computer device, the computer device including: a processor and a memory, the processor being configured to execute the method described in the first aspect above.
[0015] Fourthly, embodiments of this application also provide a computer-readable storage medium storing program instructions that, when executed, implement the method described in the first aspect above.
[0016] Fifthly, embodiments of this application also provide a computer program product or computer program, which includes computer instructions that, when executed by a processor, implement the method described in the first aspect above.
[0017] This application embodiment can acquire an image to be enhanced and a reference image of the image to be enhanced. The reference image can be obtained based on the previous frame of the image to be enhanced in the target video. Then, the image to be enhanced and the reference image can be fused to obtain a fused image. Further, feature extraction can be performed on the image to be enhanced to obtain the target image features of the image to be enhanced, and feature extraction can be performed on the fused image to obtain the enhanced image features of the image to be enhanced. Thus, the enhanced image of the image to be enhanced can be reconstructed based on the target image features and the enhanced image features. The image resolution of the enhanced image is greater than that of the image to be enhanced. In this way, the image resolution can be improved, and thus the video resolution can also be improved. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of the architecture of an image enhancement processing system provided in an embodiment of this application;
[0020] Figure 2 This is a schematic flowchart of an image enhancement processing method provided in an embodiment of this application;
[0021] Figure 3aThis is a schematic diagram of a structure for acquiring a target video provided in an embodiment of this application;
[0022] Figure 3b This is a schematic diagram of the structure of a target model provided in an embodiment of this application;
[0023] Figure 3c This is a schematic diagram of the structure of a first feature extraction module provided in an embodiment of this application;
[0024] Figure 3d This is a schematic diagram of the structure of a second feature extraction module provided in an embodiment of this application;
[0025] Figure 3e This is a schematic diagram of the structure of a downsampling module provided in an embodiment of this application;
[0026] Figure 3f This is a schematic diagram of another downsampling module provided in an embodiment of this application;
[0027] Figure 3g This is a schematic diagram of the structure of a residual module provided in an embodiment of this application;
[0028] Figure 4a This is a schematic diagram of the structure of a downsampling submodule provided in an embodiment of this application;
[0029] Figure 4b This is a schematic diagram of another downsampling submodule provided in an embodiment of this application;
[0030] Figure 4c This is a schematic diagram of the structure of an upsampling module provided in an embodiment of this application;
[0031] Figure 4d This is a schematic diagram of the structure of a conversion module provided in an embodiment of this application;
[0032] Figure 4e This is a schematic diagram of another target model provided in an embodiment of this application;
[0033] Figure 4f This is a schematic diagram showing a comparison between an image to be enhanced and an enhanced image, provided in an embodiment of this application.
[0034] Figure 5 This is a schematic flowchart of another image enhancement processing method provided in an embodiment of this application;
[0035] Figure 6 This is a schematic diagram of the structure of an initial model provided in an embodiment of this application;
[0036] Figure 7 This is a schematic diagram of the structure of an image enhancement processing device provided in an embodiment of this application;
[0037] Figure 8 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0038] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0039] With the continuous development of internet technology, artificial intelligence (AI) technology has also seen significant advancements. AI technology refers to the theories, methods, techniques, and application systems that utilize digital computers or computers-controlled machines to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science; it primarily aims to understand the essence of intelligence and produce new intelligent machines that can react in a manner similar to human intelligence, enabling these machines to possess multiple functions such as perception, reasoning, and decision-making. Accordingly, AI technology is a multidisciplinary field, mainly encompassing computer vision (CV), speech processing, natural language processing, and machine learning (ML) / deep learning.
[0040] Computer vision (CV) is a science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in recognizing and measuring targets, and then performs image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), autonomous driving, intelligent transportation, and other technologies, as well as common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0041] Machine learning is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of AI and the fundamental approach to enabling computer devices to possess intelligence. Deep learning, on the other hand, is a technique that utilizes deep neural network systems for machine learning. Machine learning / deep learning typically includes artificial neural networks, reinforcement learning (RL), supervised learning, and unsupervised learning. Supervised learning refers to training models using training samples with known categories (labeled categories), while unsupervised learning refers to training models using training samples with unknown categories (unlabeled categories).
[0042] Based on the aforementioned AI technologies, including computer vision and machine learning / deep learning, this application proposes an image enhancement processing scheme to enhance a target video and obtain the corresponding enhanced video. The enhanced video has a higher resolution than the target video. It is understood that enhancing the target video requires sequentially performing image enhancement processing on all image frames in the target video to obtain enhanced images for each frame. The enhanced image corresponding to each frame can then be used to construct the enhanced video corresponding to the target video. Considering that the image enhancement process is consistent for each frame in the target video, this application mainly uses the image enhancement processing of one frame in the target video as an example. This frame can be referred to as the image to be enhanced.
[0043] The general principle of the image enhancement processing scheme in this application embodiment is as follows: An image to be enhanced and a reference image of the image to be enhanced can be obtained, and the enhanced image corresponding to the image to be enhanced can be determined based on these two frames. The reference image can be obtained based on the previous frame of the image to be enhanced in the target video, and the image resolution of the enhanced image is greater than that of the image to be enhanced. Optionally, the image to be enhanced and the reference image can be fused first to obtain a fused image; then, the enhanced image corresponding to the image to be enhanced can be determined using the fused image and the image to be enhanced. In one embodiment, feature extraction can be performed on the image to be enhanced to obtain the target image features of the image to be enhanced, and feature extraction can be performed on the fused image to obtain the enhanced image features of the image to be enhanced; after obtaining these two features, the enhanced image of the image to be enhanced can be reconstructed based on these two features. By implementing the above scheme, image enhancement processing can be achieved to improve the image resolution of the image, and video enhancement processing can also be achieved to improve the video resolution of the video. During the image enhancement process, the processing effect of the previous frame (such as a reference image) can be used to assist the image enhancement processing of the current frame (such as the image to be enhanced), which can effectively ensure the continuity between image frames, thereby improving the image enhancement effect of the image to be enhanced, and also improving the video enhancement effect of the video.
[0044] In practical implementation, the image enhancement processing scheme mentioned above can be executed by a computer device, which can be a terminal or a server. The terminal mentioned here can include, but is not limited to, smartphones, tablets, laptops, desktop computers, smart TVs, etc. Various clients (apps) can run on the terminal, such as multimedia playback clients, social media clients, browser clients, news feed clients, educational clients, and so on. The server mentioned here can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, etc.
[0045] For example, when the computer device is a server, if any object requires video enhancement, the video to be enhanced (i.e., the target video mentioned above) can be uploaded to the server through any terminal. The server then uses the image enhancement processing scheme to perform image enhancement processing on each frame of the target video, thereby obtaining the enhanced video. Figure 1 As shown.
[0046] Based on the above description of image enhancement processing schemes, this application proposes an image enhancement processing method, which is mainly described using a computer device as the execution subject; please refer to... Figure 2 The image enhancement processing method may include the following steps S201-S204:
[0047] S201, Obtain the image to be enhanced and a reference image of the image to be enhanced.
[0048] The image to be enhanced can be any frame from a target video. The target video can be a video to be enhanced, which may be distorted due to compression or other factors. In this case, the video to be enhanced needs to be decompressed to achieve video clarity. The target video can be of any video format and size, without restriction. For example, the video format can be mp4, avi, etc., and the video size can be 512×384, 640×352, etc. The images in the target video (such as the image to be enhanced, reference images, etc.) can also be images of any format, such as png, bmp, jpg, etc.
[0049] The reference image can be obtained from the previous frame of the image to be enhanced in the target video. If the image to be enhanced is the first frame of the target video, then the previous frame can refer to the image to be enhanced itself; that is, the reference image can be the image to be enhanced. If the image to be enhanced is not the first frame of the target video (such as the second frame, third frame, etc.), then the reference image can be the previous frame; or, the reference image is obtained by enhancing the previous frame, and the resolution of the image to be enhanced is smaller than the resolution of the reference image. Considering that if the reference image is directly the previous frame of the image to be enhanced, it may contain a lot of noise and little detail, making it unsuitable for denoising and detail generation in the image to be enhanced, it is preferable to use the image obtained by enhancing the previous frame as the reference image. In this case, the reference image has less noise and richer detail, which is more conducive to denoising and detail generation in the image to be enhanced.
[0050] The specific implementation of the reference image obtained by performing image enhancement processing on the previous frame image mentioned above is similar to the implementation process of obtaining the corresponding enhanced image from the image to be enhanced. The specific description of performing image enhancement processing on the previous frame image is not provided here.
[0051] In one implementation, the target video mentioned above can be obtained when it is determined that there is a need for video enhancement.
[0052] Optionally, the existence of a video enhancement need can be determined when the computer device receives a video enhancement request. For example, an object (which could be any user) can send a video enhancement request for a target video to the computer device, causing the computer device to receive the request. Upon receiving the request, the computer device determines that a video enhancement need exists for the target video. In one embodiment, when an object needs to enhance a target video, it can perform relevant operations through a user interface on a terminal to send a video enhancement request for the target video to the computer device. See, for example... Figure 3a As shown: The terminal used by the object can display a user interface on the terminal screen. This user interface can include at least a data setting area marked 301 and a confirmation control marked 302. If the object wants to enhance a target video, it can input relevant information about the target video in the data setting area 301. This information can be the target video itself or the target video's storage area. Then, it can perform a trigger operation (such as a click or press) on the confirmation control 302, thereby triggering the terminal used by the object to obtain the target video based on the relevant information in the data setting area 301. After the terminal obtains the target video, it can send a video enhancement request for the target video to the computer device. This video enhancement request can carry the target video.
[0053] Optionally, video enhancement requirements can also be triggered by a scheduled enhancement task. For example, an enhancement scheduled task can be set up, specifying the trigger conditions for enhancing a target video, which may be pre-stored in a storage area. For instance, the trigger condition could be that the current time reaches a preset enhancement time; or, if a new video is stored in the storage area, the corresponding target video could be that new video; and so on.
[0054] S202, the image to be enhanced and the reference image are fused to obtain a fused image.
[0055] The fusion process can refer to concatting two frames of images in the channel direction. Specifically, step S202 can be implemented by concatting the image to be enhanced and the reference image in the channel direction. The data obtained by concatenation is the fused image.
[0056] In one implementation, the fusion process can be achieved by calling a fusion module in the target model. For example, the image to be enhanced and a reference image can be input into the fusion module, and the output of the fusion module is the fused image. The specific structure of the target model can be found in the following description.
[0057] S203, perform feature extraction on the image to be enhanced to obtain the target image features of the image to be enhanced, and perform feature extraction on the fused image to obtain the enhanced image features of the image to be enhanced.
[0058] In one implementation, the target image features and the enhanced image features can be obtained by calling a target model. This target model can include two branches: the first branch can be used to extract features from the image to be enhanced, obtaining the target image features; the second branch can be used to extract features from the fused image, obtaining the enhanced image features. The first branch can include a first feature extraction module, and the second branch can include the fusion module mentioned in step S202 and the second feature extraction module. For example, the target model can be found in [reference needed]. Figure 3b As shown.
[0059] Based on this, it can be seen that the target image features can be obtained by calling the first feature extraction module in the target model. For example, the image to be enhanced can be input into this first feature extraction module, and the output of this first feature extraction module is the target image features of the image to be enhanced. The first feature extraction module can simply extract features from the image to be enhanced without changing the number of channels of the image to be enhanced, such as 3 channels. Optionally, the first feature extraction module can be composed of a convolutional layer (conv) + an activation layer (such as the ReLU activation function), for example, see [link to documentation]. Figure 3c As shown. Similarly, enhanced image features can be obtained by calling the second feature extraction module in the target model. For example, the fused image can be input into the second feature extraction module, and the output of the second feature extraction module is the enhanced image feature of the image to be enhanced.
[0060] In one implementation, the specific implementation of feature extraction from the fused image to obtain the enhanced image features of the image to be enhanced may include the following steps s11-s13:
[0061] s11, perform feature extraction on the fused image to obtain the first image feature of the fused image.
[0062] Feature extraction here can refer to preliminary feature extraction and dimensionality upscaling of the fused image. This dimensionality upscaling can refer to increasing the number of channels. For example, if the number of channels in the fused image is 3, then the number of channels corresponding to the first image feature can be 6, or other values. Feature extraction here can also have inter-frame offset ablation function to eliminate the differences between the image information corresponding to the image to be enhanced and the reference image in the fused image.
[0063] s12, the first image features are downsampled to obtain the downsampled features of the fused image.
[0064] Downsampling can modify the size of the first image features, such as by changing their resolution and number of channels, to obtain downsampled features for the fused image. For example, pooling, bicubic interpolation, and space-to-depth downsampling can be used. Considering that space-to-depth downsampling not only performs downsampling but also preserves the original details in the image features, it can be prioritized. For instance, if the downsampling is a 2x downsampling, the resolution of the first image features can be reduced to half, but the number of channels can be increased fourfold, resulting in downsampled features. This is equivalent to reconstructing the first image features without changing their corresponding feature values, and preserving the details in the first image features during downsampling to prevent loss of details in the image to be enhanced due to downsampling.
[0065] s13 performs upsampling on the downsampled features to obtain the enhanced image features of the image to be enhanced.
[0066] Upsampling can involve restoring the resolution and number of channels of downsampled features to obtain enhanced image features of the image to be enhanced. These enhanced image features have the same resolution and number of channels as the target image features to ensure they are in the same dimension for subsequent image reconstruction. For example, the downsampling method corresponding to this upsampling process could be depthtospace, which corresponds to spacetodepth.
[0067] By downsampling to change the image size during image enhancement processing and then upsampling to restore the size, the scale diversity of the receptive field can be effectively improved, thereby enhancing and ensuring the subjective effect (such as viewing effect) of the subsequent enhanced image.
[0068] In this context, in one embodiment, step s11 can be implemented by a feature extraction submodule, step s12 by a downsampling module, and step s13 by an upsampling module; that is, the second feature extraction module may include a feature extraction submodule, a downsampling module, and an upsampling module. For example, the composition of the second feature extraction module may be as follows: Figure 3d As shown in the figure. The feature extraction submodule can have the same module structure as the first feature extraction module.
[0069] In one implementation, step s12 can be further implemented as follows: First, feature extraction can be performed on the first image features to obtain the second image features of the fused image; then, downsampling processing can be performed on the second image features to obtain the initial downsampled features of the fused image; finally, feature extraction can be performed on the initial downsampled features to obtain the downsampled features of the fused image. Based on this, it can be seen that in this case, the downsampling module can include: a forward feature extraction module, a downsampling submodule (or a down-sample block), and a backward feature extraction module. For example, the composition of the downsampling module can be as follows: Figure 3e As shown.
[0070] The forward feature extraction module can include n residual blocks (Res blocks), and the backward feature extraction module can also include n residual blocks (or Res blocks). For example, the module structure of the downsampling module can also be as follows: Figure 3f As shown, the module structure of any residual module can be as follows: Figure 3g As shown, the downsampling module is a symmetrical structure consisting of n Resblocks and a downsampling sub-module (or a down-sample block). This type of downsampling module can also be called a Dual ResNet. The first n Resblocks (i.e., the forward feature extraction module) perform deep feature extraction on the first image features. The middle down-sample block downsamples the second image features (e.g., downsampling by a factor of two). The n Resblocks connected to the down-sample block then perform deep semantic analysis on the previously extracted features (i.e., the initial downsampled features), laying the groundwork for subsequent reconstruction and restoration.
[0071] In one implementation, the specific implementation of determining the initial downsampling features in step s12 above may further include: First, the resolution and number of channels of the second image features can be transformed using a target downsampling method to obtain the transformed second image features. Then, the transformed second image features can be used to perform feature difference ablation processing on the image differences between the image to be enhanced and the reference image to obtain the initial downsampling features of the fused image. For example, if the image to be enhanced and the reference image have spatial displacement differences, such as a difference of 4 pixels, after downsampling processing (e.g., 2x downsampling), this 4-pixel difference can be reduced to half of its original value, i.e., a difference of 2 pixels. Then, feature difference ablation processing can further reduce this 2-pixel difference. In this case, the downsampling submodule may include: a target downsampling submodule and a feature ablation module. For example, the composition of the downsampling submodule may be as follows: Figure 4a As shown.
[0072] The target downsampling submodule employs a space-to-depth downsampling method. This method reduces the resolution of the second image feature while increasing the number of channels, essentially reshaping the data of the second image feature. Specifically, it reconstructs the data in the height and width dimensions of the second image feature, expanding the data in the channel dimension without altering the feature values themselves. For example, if the downsampling factor is 2, this method can reduce the resolution of the second image feature to half its original value while increasing the number of channels fourfold. For instance, consider a feature with dimensions W×H×C, where W, H, and C represent width, height, and number of channels, respectively. If the second image feature is 100×100×3, the downsampled feature will have dimensions of 50×50×12. In short, space-to-depth moves spatial data (W and H dimensions) to the depth dimension (channel dimension). Next, a feature ablation module (which may include convolutional layers and activation layers) is connected to perform feature ablation on the transformed second image features output by the target downsampling submodule. In this case, the downsampling submodule can also be as follows: Figure 4b As shown.
[0073] In summary, the space-to-depth downsampling method utilizes all data from the second image features during downsampling without losing any data, thus preserving image details and preventing loss of detail due to downsampling. It can also alter noise distribution, transforming dense noise into a more dispersed one. Simply put, each channel can be considered an image, increasing the number of channels after space-to-depth (e.g., if space-to-depth uses a 2x downsampling, the number of channels can be quadrupled), resulting in a more dispersed noise distribution. Space-to-depth also reduces resolution, effectively decreasing computation time and improving the reconstruction speed of the enhanced image. Furthermore, the subsequent dimensionality reduction using convolutional layers after downsampling also helps eliminate noise.
[0074] In one implementation, step s13 can be further implemented as follows: First, the number of channels in the downsampled features can be transformed, for example, the number of channels in the downsampled features can be reduced to 12 channels, thus obtaining the transformed downsampled features; then, the resolution and number of channels of the transformed downsampled features can be restored using a target upsampling method to obtain the enhanced image features of the image to be enhanced; the resolution and number of channels of the enhanced image features are consistent with those of the target image features to ensure that they are in the same dimension as the target image features and can be used for subsequent image reconstruction. The target upsampling method can be depthtospace, which transforms channel data into spatial data. For example, if the size of the transformed downsampled features is 50×50×12, and the size of the enhanced image features obtained after upsampling is 100×100×3, it can be seen that there are 12 channels in the transformed downsampled features. Therefore, during upsampling, the channel data corresponding to these 12 channels can be converted into data in the spatial dimensions (W and H dimensions) of the enhanced image features. It can also increase the resolution (e.g., if the upsampling is 2x upsampling, the resolution can be increased by 2x), restoring the converted downsampled features to the same resolution and number of channels as the target image features (or the image to be enhanced in the input target model).
[0075] In this case, the upsampling module may include a conversion module and an upsampling submodule. For example, the composition of the upsampling module may be as follows: Figure 4c As shown. The transformation module can adopt a conv+relu+conv structure, such as... Figure 4d As shown in the figure. Furthermore, based on the above description, the overall structure of the target model can be seen as follows: Figure 4e As shown, and through Figure 4eThe model structure diagram also shows that the target model is a network structure with a symmetrical structure. The network structure is simple and the target model can have high overall computational efficiency, thereby improving the efficiency of image enhancement processing and thus increasing the generation speed of enhanced images (or enhanced videos).
[0076] S204. Image reconstruction is performed using the features of the target image and the features of the enhanced image to obtain the enhanced image of the image to be enhanced.
[0077] In this context, the image resolution of the enhanced image is greater than that of the image to be enhanced, and both the enhanced image and the image to be processed have the same image content. For example, see... Figure 4f As shown, Figure 4f The image marked 41 is the image to be enhanced, which can be a frame from a video that has been compressed and distorted; the image marked 42 is the enhanced image. It can be seen that compared to the image to be enhanced 41, the enhanced image 42 has significantly improved clarity and richer details and textures. Therefore, the distortion caused by video compression can be effectively removed, resulting in a video with higher resolution.
[0078] Image reconstruction can refer to adding features from two images. Therefore, step S204 can be implemented by adding the target image features and the enhanced image features, and using the result as the enhanced image of the image to be enhanced. In one implementation, step S204 can be implemented using the reconstruction module in the target model, for example... Figure 3b As shown, the target image features and the enhanced image features can be input into the reconstruction module, and the output of the reconstruction module is the enhanced image of the image to be enhanced.
[0079] In this embodiment, image enhancement processing can be implemented to improve the image resolution of an image, and video enhancement processing can also be implemented to improve the video resolution of a video. Furthermore, automated image enhancement processing can be performed using a target model, thereby improving the automation and intelligence of enhancement, and increasing image enhancement efficiency, which in turn improves video enhancement efficiency. During image enhancement processing, the processing effect of the previous frame (such as a reference image) can be used to assist the image enhancement processing of the current frame (such as the image to be enhanced), effectively ensuring the continuity between image frames, thus improving the image enhancement effect of the image to be enhanced, and also improving the video enhancement effect of the video. Simultaneously, by utilizing adjacent frames, temporal information can be used to ensure the continuity between frames, and detailed information between frames can be extracted, which helps to remove noise and complement detailed information in the current frame, thereby achieving decompression distortion of the video and improving the video resolution.
[0080] Please see Figure 5 This is a schematic flowchart illustrating another image enhancement processing method provided in this application embodiment. This application embodiment is mainly described using a computer device as the execution subject; please refer to... Figure 5 The image enhancement processing method may include the following steps S501-S503:
[0081] S501, Obtain at least one training sample for training the initial model and the label image of each training sample.
[0082] A training sample can include a sample image and a reference sample image. The reference sample image is the previous frame of the sample image in the sample video. The label image is an enhanced version of the sample image, meaning that the image resolution of the label image is greater than that of the sample image, and the label image and the sample image have the same image content. Alternatively, the sample image is a blurry frame, while the label image is a high-resolution version of the sample image.
[0083] In one implementation, two consecutive sample images included in the training samples can be obtained from a single sample video. The acquisition of a training sample is described below. Specifically, a sample video can be acquired, and consecutive image frames can be extracted from any sample video to obtain multiple consecutive image frames. Any two consecutive image frames can then be used as a training sample. For example, after extracting consecutive image frames from a sample video, images 1, 2, 3, 4, and 5 can be obtained. Images 1 and 2 can form one training sample, images 2 and 3 can form another, images 3 and 4 can form yet another, and images 4 and 5 can also form a training sample. Thus, one or more training samples can be obtained using a single sample video.
[0084] In another implementation, the initial model can be trained using training samples from different scenarios to improve the generalization ability of the trained target model. This ensures that the trained target model achieves good performance, and the same model can be used for image or video enhancement processing in different scenarios, without needing to change the model based on scene variations. The scene type classification is not specifically limited; for example, scene types can include classroom scenes, playground scenes, cafeteria scenes, dormitory scenes, etc.
[0085] In this context, step S501 can be implemented as follows: Videos of different scene types can be acquired, and these acquired videos can be used as sample videos. Specifically, training samples for a given scene type can be obtained from sample videos of one scene type; similarly, training samples for different scene types can be obtained from sample videos of different scene types. For any given scene type of sample video, continuous image frames can be extracted to obtain multiple consecutive images, which can then be used to obtain training samples for that scene type.
[0086] In some cases, a video may contain multiple scene types. For example, a video may contain both a playground scene and a cafeteria scene. In this case, scene type detection can be performed on the image frames in the video to distinguish the scene type of each frame. In this scenario, step S501 can be implemented as follows: First, a sample video can be acquired, and consecutive image frames can be extracted from the sample video to obtain multiple consecutive images. After obtaining multiple consecutive images, scene type detection can be performed on each frame. Following the principle that images of the same scene type are grouped into the same group, the multiple frames are grouped according to the detection results to obtain one or more image groups. Each image group corresponds to one scene type. For any image group, two consecutive frames in that image group can be used as a training sample, thus obtaining multiple training samples.
[0087] Optionally, scene type detection can be based on image content. In one embodiment, for any two frames (a first image and a second image), if the shared image content between the first image and the second image exceeds a preset percentage, the first and second images can be classified as belonging to the same scene type. If the shared image content between the first image and the second image does not exceed the preset percentage, the first and second images can be classified as belonging to different scene types. The preset percentage can be pre-set, such as 80%, 70%, etc., and its specific value is not limited. Based on this, it can be seen that scene type detection can determine the scene type of each frame in the sample video. It should also be noted that this scene type detection does not require explicit knowledge of the specific scene type of each image; it only needs to determine whether any two frames belong to the same scene type.
[0088] In another embodiment, multiple reference images can be pre-acquired, with each reference image corresponding to a scene type. Then, for any frame in the sample video, the image content of that frame is matched with each of the reference images. If the percentage of identical image content between the image content of that frame and that of a certain reference image exceeds a preset percentage, then the scene type corresponding to that reference image can be determined as the scene type of that frame. In this way, the scene type of each frame in the sample video can be determined.
[0089] For example, for a sample video, after extracting consecutive image frames, N frames can be obtained. After scene type detection, the scene type of all consecutive images from frame 1 to frame n1 is type 1, the scene type of all consecutive images from frame n2 to frame n3 is type 2, and the scene type of all consecutive images from frame n4 to frame N is type 3. Frames n1 and n2 are consecutive, and frames n3 and n4 are consecutive. Based on these results, all images from frame 1 to frame n1 can be grouped into one image group, all images from frame n2 to frame n3 can be grouped into one image group, and all images from frame n4 to frame N can be grouped into one image group. Then, for any image group, two consecutive frames within that group can be used as training samples, thus obtaining corresponding training samples for each scene type.
[0090] It should be noted that the sample videos obtained above can be either first-resolution videos or second-resolution videos. First-resolution videos can refer to videos with a resolution higher than a first preset resolution, while second-resolution videos can refer to videos with a resolution lower than a second preset resolution. The first preset resolution can be greater than or equal to the second preset resolution. In other words, first-resolution videos are for videos with higher resolutions, while second-resolution videos are for videos with lower resolutions.
[0091] Optionally, if the sample video is a first-resolution video, for the training samples determined from this sample video, since the sample video is a first-resolution video, the image resolution of the images in the sample video is relatively high, while the image resolution of the images included in the training samples needs to be lower. Therefore, after extracting two consecutive frames, these two frames can be degraded (or blurred) to reduce their resolution. For example, image blurring can be done by adding Gaussian noise, performing Gaussian blur, or adding decompression noise. Then, the two frames after data augmentation can be used as a training sample. The first frame of these two consecutive frames can be used as the sample reference image, and the second frame can be used as the sample image. As for the label image of the sample image, it can be the image of the sample image before data augmentation, that is, the image extracted from the sample video can be directly used as the label image. It is known that the sample image is the degraded version of the label image.
[0092] Optionally, if the sample video is a second-resolution video, since the image resolution of the images in the sample video is low, and the image resolution of the images included in the training samples also needs to be low, then any two consecutive frames in the target video can be used as a training sample. The first frame of these two consecutive frames can be used as the sample reference image, and the second frame can be used as the sample image. As for the label image of the sample image, it can be the image after image enhancement processing of the sample image. The image enhancement processing here can be the image enhancement processing in the traditional technique, as long as the image resolution of the label image is greater than that of the sample image.
[0093] In this embodiment of the application, the training batch size (patch size) when training the initial model can be set to 64, 96, 128, etc., and the specific training batch size can be determined according to the video memory size of the computer device; for example, if the video memory is large, the corresponding training batch size can be set to a larger value, and if the video memory is small, the corresponding training batch size can be set to a smaller value.
[0094] S502: For any training sample, call the initial model to perform image enhancement processing on any training sample to obtain the predicted image of any training sample.
[0095] The initial model can be found in, for example, as follows: Figure 3b or Figure 4eAs shown, the process of the initial model enhancing the images of the training samples and obtaining the predicted images during the training phase is similar to the process of using the target model to enhance the images of the images to be enhanced and obtaining the enhanced images during the inference phase (i.e., steps S201-S204). The specific implementation process of step S502 can be referred to the description of steps S201-S204, and will not be repeated here.
[0096] S503: Based on the predicted image of any training sample and the label image of any training sample, train the initial model to obtain the target model.
[0097] In one implementation, the model result of the initial model in the embodiments of this application can be as follows: Figure 3b In this case, for any training sample, the model loss value of the initial model can be calculated based on the difference between the predicted image and the label image of any training sample. The initial model can then be trained based on this model loss value. For example, the model parameters in the initial model can be updated in the direction of reducing the model loss value, thereby obtaining the trained initial model (i.e., the target model). When calculating the model loss value of the initial model, the loss function corresponding to that initial model can be the L2 loss function, meaning the L2 loss function can be used to calculate the model loss value. In this implementation, the computer device can calculate the model loss value using the following formula (1):
[0098] L = ||I t -I s || (1)
[0099] Where L represents the model loss value calculated for any training sample, and I t I represents the label image of any training sample. s This represents the predicted image for any training sample.
[0100] In another implementation, the initial model in this embodiment may include a generation module and a discrimination module. The generation module can be used to generate corresponding predicted images based on training samples. The network structure of the generation module can be as follows: Figure 3b As shown; the discrimination module can be used to determine whether the predicted image generated by the generation module is real. For example, the initial model can be as follows: Figure 6As shown. In this case, under the initial model structure, the predicted image of any training sample can be obtained by image processing of any training sample using the generation module in the initial model. Based on this, the specific implementation of step S502 can be: for any training sample, call the generation module of the initial model to perform image processing on any training sample, thereby obtaining the predicted image of any training sample. In this case, the specific implementation of step S503 can include: inputting the predicted image and the label image of any training sample into the discrimination module in the initial model to obtain the discrimination result of the predicted image of any training sample; then, the initial model can be trained based on the predicted image, the label image, and the discrimination result to obtain the trained initial model; furthermore, the generation module in the trained initial model can be determined as the target model.
[0101] Specifically, the aforementioned method of training the initial model based on the predicted image, label image, and discrimination result to obtain the trained initial model can be implemented as follows: Calculate the model loss value of the initial model based on the predicted image, label image, and discrimination result; then train the initial model based on the model loss value to obtain the trained initial model. Optionally, a first loss value can be calculated based on the predicted image and label image, and a second loss value can be calculated based on the discrimination result; after obtaining these two loss values, the model loss value of the initial model can be calculated based on these two loss values.
[0102] Optionally, the sum of the first and second loss values can be used as the model loss value; in this implementation, the computer device can calculate the model loss value using the following formula (2):
[0103] L = L1 + L GAN (2)
[0104] Here, L1 represents the first loss value, and the loss function corresponding to L1 can be the L2 loss function. L1 can be used to represent the difference between the predicted image and the sample image. For example, L1 can specifically be L1 = ||I t -I s ||;L GAN L represents the second loss value. GAN The corresponding loss function can be an adversarial loss function, for example, L. GAN Specifically, it could be L GAN =-∑logD(I t ,I s ), where D(·) represents the effect of the discrimination module, the purpose of which is to determine whether the predicted image generated by the generation module is real.
[0105] Optionally, the first loss value and the second loss value can be weighted and summed to obtain the model loss value; in this implementation, the computer device can calculate the model loss value using the following formula (3):
[0106] L = k1 × L1 + k2 × L GAN (3)
[0107] Where k1 is the weight coefficient of L1, and k2 is the weight coefficient of L... GAN The weighting coefficients.
[0108] As can be seen, in training the initial model, this embodiment can decompose the video into continuous images and input two consecutive distorted images (such as a reference sample image and a sample image) into the initial model. One branch can fuse the two images along the channel direction to obtain a fused image. Then, inter-frame offset ablation is performed on the fused image, and the features from the inter-frame offset ablation are input into a dual residual network for feature extraction to obtain the features corresponding to the fused image. Another branch can directly extract features from the next frame (i.e., the sample image) of the two consecutive distorted images to obtain the corresponding features. Furthermore, a predicted image can be reconstructed based on the features from these two branches, and L2 loss and adversarial loss are calculated with the labeled image. The calculated loss values are then used to train the initial model. This embodiment can utilize an adversarial network (a combination of a generator and a discriminator module) to improve the initial model's ability to generate details in the sample image, thereby improving the enhancement effect of the predicted image corresponding to the sample image, and ultimately resulting in a better model performance.
[0109] The above method embodiments are illustrative examples of the methods in this application. The descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. For example, after training the target model, the image to be enhanced can be obtained, and the enhancement processing of the image to be enhanced can be performed based on the target model to obtain the enhanced image of the image to be enhanced. This will not be elaborated here.
[0110] In this embodiment, the processing effect of the previous frame (such as a reference image) can be used to assist the image enhancement processing of the current frame (such as the image to be enhanced), effectively ensuring the continuity between image frames and thus improving the image enhancement effect of the image to be enhanced, as well as the video enhancement effect of the video. Simultaneously, image information from adjacent frames can be effectively utilized, aiding in noise removal and the complementarity of detail information in the current frame. Furthermore, the training sample data corresponding to the initial target model during training can be data from multiple scenes, thereby effectively improving the generalization ability of the target model. When switching scenes, the target model can be used for adaptive image enhancement processing without requiring changes to the target model due to scene changes.
[0111] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of an image enhancement processing device provided in an embodiment of this application. Specifically, the device is disposed in a computer device, and the device includes: an acquisition unit 701, a fusion unit 702, an extraction unit 703, and a reconstruction unit 704;
[0112] The acquisition unit 701 is used to acquire the image to be enhanced and a reference image of the image to be enhanced; the reference image is obtained based on the previous frame image of the image to be enhanced in the target video;
[0113] The fusion unit 702 is used to fuse the image to be enhanced and the reference image to obtain a fused image;
[0114] The extraction unit 703 is used to extract features from the image to be enhanced to obtain the target image features of the image to be enhanced, and to extract features from the fused image to obtain the enhanced image features of the image to be enhanced;
[0115] The reconstruction unit 704 is used to reconstruct an enhanced image of the image to be enhanced based on the target image features and the enhanced image features; the image resolution of the enhanced image is greater than the image resolution of the image to be enhanced.
[0116] Furthermore, when the extraction unit 703 extracts features from the fused image to obtain the enhanced image features of the image to be enhanced, it is specifically used for:
[0117] Feature extraction is performed on the fused image to obtain the first image feature of the fused image;
[0118] The first image features are downsampled to obtain the downsampled features of the fused image;
[0119] The downsampled features are upsampled to obtain the enhanced image features of the image to be enhanced.
[0120] Furthermore, when the extraction unit 703 performs downsampling processing on the first image features to obtain the downsampled features of the fused image, it is specifically used for:
[0121] The first image features are extracted to obtain the second image features of the fused image;
[0122] The second image features are downsampled to obtain the initial downsampled features of the fused image;
[0123] Feature extraction is performed on the initial downsampled features to obtain the downsampled features of the fused image.
[0124] Further, when the extraction unit 703 performs downsampling processing on the second image features to obtain the initial downsampling features of the fused image, it is specifically used for:
[0125] The resolution and number of channels of the second image feature are transformed by using a target downsampling method to obtain the transformed second image feature;
[0126] The image differences between the image to be enhanced and the reference image are ablated using the transformed second image features to obtain the initial downsampling features of the fused image.
[0127] Furthermore, when the extraction unit 703 performs upsampling processing on the downsampled features to obtain the enhanced image features of the image to be enhanced, it is specifically used for:
[0128] The downsampled features are converted to a higher channel count to obtain the converted downsampled features.
[0129] The resolution and number of channels of the transformed downsampled features are restored by using a target upsampling method to obtain the enhanced image features of the image to be enhanced; the enhanced image features have the same resolution and number of channels as the target image features.
[0130] Furthermore, if the image to be enhanced is the first frame image in the target video, then the reference image is the image to be enhanced;
[0131] If the image to be enhanced is not the first frame in the target video, then the reference image is obtained by performing image enhancement processing on the previous frame, and the image resolution of the image to be enhanced is smaller than the image resolution of the reference image.
[0132] Furthermore, the device also includes a training unit 705, which is further configured to:
[0133] Obtain at least one training sample for training an initial model and label images for each training sample; a training sample includes a sample image and a reference sample image, wherein the reference sample image is the previous frame of the sample image in the sample video; the label image is an enhanced image of the sample image, wherein the image resolution of the label image is greater than that of the sample image;
[0134] For any training sample, the initial model is invoked to perform image enhancement processing on the training sample to obtain a predicted image of the training sample.
[0135] Based on the predicted image of any training sample and the label image of any training sample, the initial model is trained to obtain the target model.
[0136] Further, the initial model includes a generation module and a discrimination module. The predicted image of any training sample is obtained by performing image enhancement processing on the training sample using the generation module in the initial model. When the training unit 705 trains the initial model based on the predicted image and the label image of any training sample to obtain the target model, it is specifically used for:
[0137] The predicted image and the label image of any training sample are input into the discrimination module to obtain the discrimination result of the predicted image of any training sample.
[0138] Based on the predicted image, the labeled image, and the discrimination result, the initial model is trained to obtain the trained initial model;
[0139] The generating module in the trained initial model is determined as the target model.
[0140] Furthermore, when the training unit 705 acquires at least one training sample for training the initial model, it specifically performs the following:
[0141] Acquire a sample video and extract consecutive image frames from the sample video to obtain multiple consecutive images;
[0142] Scene type detection is performed on each frame of image, and the frames of image are grouped according to the principle that the same scene type is divided into the same group. Based on the detection results of each frame of image, the frames of image are grouped to obtain one or more image groups; each image group corresponds to one scene type.
[0143] For any image group, any two consecutive frames in the image group are used as training samples to obtain at least one training sample.
[0144] It is understood that the division of units in this embodiment is illustrative and merely a logical functional division; in actual implementation, there may be other division methods. The functional units in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units can be implemented in hardware or as software functional units.
[0145] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Specifically, the computer device includes: a memory 801 and a processor 802.
[0146] In one embodiment, the computer device may further include a data interface 803, through which the memory 801, processor 802, and data interface 803 can exchange data.
[0147] The memory 801 may include volatile memory; the memory 801 may also include non-volatile memory; the memory 801 may also include a combination of the above types of memory. The processor 802 may be a central processing unit (CPU). The processor 802 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), or any combination thereof.
[0148] The memory 801 is used to store programs, and the processor 802 can call the programs stored in the memory 801 to perform the following steps:
[0149] Obtain the image to be enhanced and a reference image of the image to be enhanced; the reference image is obtained based on the previous frame of the image to be enhanced in the target video;
[0150] The image to be enhanced and the reference image are fused together to obtain a fused image;
[0151] Feature extraction is performed on the image to be enhanced to obtain the target image features of the image to be enhanced, and feature extraction is performed on the fused image to obtain the enhanced image features of the image to be enhanced;
[0152] An enhanced image is reconstructed from the target image features and the enhanced image features; the image resolution of the enhanced image is greater than that of the image to be enhanced.
[0153] Furthermore, when the processor 802 extracts features from the fused image to obtain the enhanced image features of the image to be enhanced, it is specifically used for:
[0154] Feature extraction is performed on the fused image to obtain the first image feature of the fused image;
[0155] The first image features are downsampled to obtain the downsampled features of the fused image;
[0156] The downsampled features are upsampled to obtain the enhanced image features of the image to be enhanced.
[0157] Furthermore, when the processor 802 performs downsampling processing on the first image features to obtain the downsampled features of the fused image, it is specifically used for:
[0158] The first image features are extracted to obtain the second image features of the fused image;
[0159] The second image features are downsampled to obtain the initial downsampled features of the fused image;
[0160] Feature extraction is performed on the initial downsampled features to obtain the downsampled features of the fused image.
[0161] Further, when the processor 802 performs downsampling processing on the second image features to obtain the initial downsampling features of the fused image, it is specifically used for:
[0162] The resolution and number of channels of the second image feature are transformed by using a target downsampling method to obtain the transformed second image feature;
[0163] The image differences between the image to be enhanced and the reference image are ablated using the transformed second image features to obtain the initial downsampling features of the fused image.
[0164] Furthermore, when the processor 802 performs upsampling processing on the downsampled features to obtain the enhanced image features of the image to be enhanced, it is specifically used for:
[0165] The downsampled features are converted to a higher channel count to obtain the converted downsampled features.
[0166] The resolution and number of channels of the transformed downsampled features are restored by using a target upsampling method to obtain the enhanced image features of the image to be enhanced; the enhanced image features have the same resolution and number of channels as the target image features.
[0167] Furthermore, if the image to be enhanced is the first frame image in the target video, then the reference image is the image to be enhanced;
[0168] If the image to be enhanced is not the first frame in the target video, then the reference image is obtained by performing image enhancement processing on the previous frame, and the image resolution of the image to be enhanced is smaller than the image resolution of the reference image.
[0169] Furthermore, the processor 802 is also used for:
[0170] Obtain at least one training sample for training an initial model and label images for each training sample; a training sample includes a sample image and a reference sample image, wherein the reference sample image is the previous frame of the sample image in the sample video; the label image is an enhanced image of the sample image, wherein the image resolution of the label image is greater than that of the sample image;
[0171] For any training sample, the initial model is invoked to perform image enhancement processing on the training sample to obtain a predicted image of the training sample.
[0172] Based on the predicted image of any training sample and the label image of any training sample, the initial model is trained to obtain the target model.
[0173] Further, the initial model includes a generation module and a discrimination module. The predicted image of any training sample is obtained by performing image enhancement processing on the training sample using the generation module in the initial model. When the processor 802 trains the initial model based on the predicted image and the label image of any training sample to obtain the target model, it is specifically used for:
[0174] The predicted image and the label image of any training sample are input into the discrimination module to obtain the discrimination result of the predicted image of any training sample.
[0175] Based on the predicted image, the labeled image, and the discrimination result, the initial model is trained to obtain the trained initial model;
[0176] The generating module in the trained initial model is determined as the target model.
[0177] Furthermore, when the processor 802 acquires at least one training sample for training the initial model, it specifically performs the following:
[0178] Acquire a sample video and extract consecutive image frames from the sample video to obtain multiple consecutive images;
[0179] Scene type detection is performed on each frame of image, and the frames of image are grouped according to the principle that the same scene type is divided into the same group. Based on the detection results of each frame of image, the frames of image are grouped to obtain one or more image groups; each image group corresponds to one scene type.
[0180] For any image group, any two consecutive frames in the image group are used as training samples to obtain at least one training sample.
[0181] Embodiments of this application also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements this application. Figure 2 or Figure 5 The method described in the corresponding embodiment can also be implemented. Figure 7 The apparatus described in the embodiments corresponding to this application will not be repeated here.
[0182] The computer-readable storage medium can be an internal storage unit of the device described in any of the foregoing embodiments, such as the device's hard drive or memory. The computer-readable storage medium can also be an external storage device of the device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the device. Further, the computer-readable storage medium can include both internal and external storage units of the device. The computer-readable storage medium is used to store the computer program and other programs and data required by the terminal device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0183] Embodiments of this application also provide a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various embodiments described above.
[0184] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0185] The above-disclosed embodiments are merely some of the embodiments of this application, and should not be construed as limiting the scope of this application. Those skilled in the art can understand that implementing all or part of the above embodiments and making equivalent changes in accordance with the claims of this application still fall within the scope of the invention.
Claims
1. An image enhancement processing method characterized by, The method includes: Obtain the image to be enhanced and a reference image of the image to be enhanced; the reference image is obtained by performing image enhancement processing on the previous frame image of the image to be enhanced in the target video; The image to be enhanced and the reference image are fused to obtain a fused image. The fusion process includes stitching the image to be enhanced and the reference image together in the channel direction. Feature extraction is performed on the image to be enhanced to obtain the target image features of the image to be enhanced, and feature extraction is performed on the fused image to obtain the first image features of the fused image; Forward feature extraction is performed on the first image features to obtain the second image features of the fused image; The resolution and number of channels of the second image feature are transformed by a target downsampling method to obtain the transformed second image feature. The target downsampling method includes space-to-depth based downsampling. The image differences between the image to be enhanced and the reference image are ablated using the transformed second image features to obtain the initial downsampling features of the fused image; Backward feature extraction is performed on the initial downsampled features to obtain the downsampled features of the fused image; The downsampled features are upsampled to obtain the enhanced image features of the image to be enhanced. The enhanced image features have the same resolution and number of channels as the target image features. An enhanced image is reconstructed from the target image features and the enhanced image features; the image resolution of the enhanced image is greater than that of the image to be enhanced.
2. The method of claim 1, wherein, The upsampling process of the downsampled features to obtain the enhanced image features of the image to be enhanced includes: The downsampled features are converted to a higher channel count to obtain the converted downsampled features. The resolution and number of channels of the transformed downsampled features are restored by using a target upsampling method to obtain the enhanced image features of the image to be enhanced.
3. The method according to claim 1, characterized in that, If the image to be enhanced is the first frame image in the target video, then the reference image is the image to be enhanced; If the image to be enhanced is not the first frame in the target video, then the reference image is obtained by performing image enhancement processing on the previous frame, and the image resolution of the image to be enhanced is smaller than the image resolution of the reference image.
4. The method according to any one of claims 1 to 3, characterized in that, The target image features, the enhanced image features, and the enhanced image are obtained by calling the target model; the method further includes: Obtain at least one training sample for training an initial model and label images for each training sample; a training sample includes a sample image and a reference sample image, wherein the reference sample image is the previous frame of the sample image in the sample video; the label image is an enhanced image of the sample image, wherein the image resolution of the label image is greater than that of the sample image; For any training sample, the initial model is invoked to perform image enhancement processing on the training sample to obtain a predicted image of the training sample. Based on the predicted image of any training sample and the label image of any training sample, the initial model is trained to obtain the target model.
5. The method of claim 4, wherein, The initial model includes a generation module and a discrimination module. The predicted image of any training sample is obtained by performing image enhancement processing on the training sample using the generation module in the initial model. The process of training the initial model based on the predicted image and the label image of any training sample to obtain the target model includes: The predicted image and the label image of any training sample are input into the discrimination module to obtain the discrimination result of the predicted image of any training sample. Based on the predicted image, the labeled image, and the discrimination result, the initial model is trained to obtain the trained initial model; The generating module in the trained initial model is determined as the target model.
6. The method of claim 4, wherein, Obtaining at least one training sample for training the initial model includes: Acquire a sample video and extract consecutive image frames from the sample video to obtain multiple consecutive images; Scene type detection is performed on each frame of image, and the frames of image are grouped according to the principle that the same scene type is divided into the same group. Based on the detection results of each frame of image, the frames of image are grouped to obtain one or more image groups; each image group corresponds to one scene type. For any image group, any two consecutive frames in the image group are used as training samples to obtain at least one training sample.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program instructions that, when executed, are used to implement the method as described in any one of claims 1-6.
8. A computer device, comprising: The computer device includes a processor and a memory, the processor being configured to perform the method as described in any one of claims 1-6.
9. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed by a processor, implement the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Image processing method and device, computer equipment and storage medium
CN111476719A
Ultra resolution implementation method and device for image frames
WO2022007895A1