Image processing method and device, electronic equipment and storage medium

CN120765482BActive Publication Date: 2026-09-29BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510874449.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2026-09-29
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

相关技术中,通常采用对图像超分辨率(Super-Resolution,SR)技术以对图像分辨率进行优化得到超分图像;但在超分辨率处理过程中,可能导致得到的超分图像中的目标物出现结构畸变、过锐白边、纹理不真实等问题,故亟待提出一种图像处理方法以提升图像显示质量

Benefits of technology

[0009]本公开实施例提供的图像处理方法,通过确定待处理的目标超分图像中需要优化的、表征目标超分图像中目标物对应的待优化特征所在的目标区域,基于待优化特征在多个网络模型中确定目标网络模型以使得该目标网络模型对该待优化特征进行优化处理得到目标超分图像对应的超分前的目标原始图像中目标区域的第一特征图,将第一特征图中的第一特征与目标超分图像中对应目标区域的第二特征进行融合处理,得到目标区域的第二特征图,基于第二特征图对所述目标超分图像进行重建,得到重建后的目标超分图像;通过使用目标网络模型对目标超分图像对应的目标原始图像中需要优化的目标区域的特征进行优化处理,并将优化处理后得到的特征与对应目标超分图像中相应区域的特征进行融合以实现对目标超分图像中相应区域的特征进行优化重建,提升了目标超分图像的显示质量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765482B_ABST
    Figure CN120765482B_ABST
Patent Text Reader

Abstract

The application discloses an image processing method and device, electronic equipment and storage medium, and relates to the technical field of computer processing. The method comprises the following steps: determining a target region needing optimization in a target super-resolution image to be processed, wherein the target super-resolution image is used to represent an image of a target original image after super-resolution processing; determining a target network model based on a to-be-optimized feature in multiple network models, wherein the network model is configured to perform optimization processing of a corresponding type feature on an input image, and the type features corresponding to all network models are not all the same; inputting the target original image into the target network model to obtain a first feature map of the target region in the target original image; performing fusion processing on a first feature in the first feature map and a second feature of a corresponding target region in the target super-resolution image to obtain a second feature map of the target region; and reconstructing the target super-resolution image based on the second feature map to obtain a reconstructed target super-resolution image, thereby improving the display quality of the target super-resolution image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer processing technology, specifically to an image processing method, apparatus, electronic device, storage medium, and program product. Background Technology

[0002] With the development of multimedia technology, the visual quality of images, as carriers of information dissemination, is constantly being improved. In related technologies, super-resolution (SR) techniques are commonly used to optimize image resolution and obtain super-resolution images. However, during super-resolution processing, problems such as structural distortion, overly sharp white edges, and unrealistic textures may occur in the resulting super-resolution images. Therefore, there is an urgent need to propose an image processing method to improve image display quality. Summary of the Invention

[0003] In view of this, the present disclosure provides an image processing method, apparatus, electronic device, storage medium, and program product to improve image display quality.

[0004] In a first aspect, this disclosure provides an image processing method, the method comprising: determining a target region to be optimized in a target super-resolution image to be processed, the target region representing the region where the feature to be optimized corresponding to the target object in the target super-resolution image is located, the target super-resolution image representing the image after super-resolution processing of the original target image; determining a target network model among multiple network models based on the feature to be optimized, the multiple network models being configured to perform corresponding type feature optimization processing on the input image, the type features corresponding to the multiple network models being not all the same; inputting the original target image into the target network model to obtain a first feature map of the target region in the original target image; fusing the first feature in the first feature map with the second feature of the corresponding target region in the target super-resolution image to obtain a second feature map of the target region; and reconstructing the target super-resolution image based on the second feature map to obtain a reconstructed target super-resolution image.

[0005] Secondly, this disclosure provides an image processing apparatus, comprising: a first determining module, configured to determine a target region to be optimized in a target super-resolution image to be processed, wherein the target region represents the region where the feature to be optimized corresponding to the target object in the target super-resolution image is located, and the target super-resolution image represents the image after super-resolution processing of the original target image; a second determining module, configured to determine a target network model among multiple network models based on the feature to be optimized, wherein the multiple network models are respectively configured to perform corresponding type feature optimization processing on the input image, and the type features corresponding to the multiple network models are not all the same; an input module, configured to input the original target image into the target network model to obtain a first feature map of the target region in the original target image; a fusion module, configured to fuse the first feature in the first feature map with the second feature of the corresponding target region in the target super-resolution image to obtain a second feature map of the target region; and a reconstruction module, configured to reconstruct the target super-resolution image based on the second feature map to obtain a reconstructed target super-resolution image.

[0006] Thirdly, this disclosure provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the image processing method described in the first aspect or any corresponding embodiment thereof.

[0007] Fourthly, this disclosure provides a computer-readable storage medium storing computer instructions for causing a computer to perform the image processing method described in the first aspect or any corresponding embodiment thereof.

[0008] Fifthly, this disclosure provides a computer program product, including computer instructions for causing a computer to perform the image processing method described in the first aspect or any corresponding embodiment thereof.

[0009] The image processing method provided in this disclosure determines the target region in the target super-resolution image that needs optimization and represents the feature corresponding to the target object in the target super-resolution image. Based on the feature to be optimized, a target network model is determined among multiple network models so that the target network model optimizes the feature to be optimized to obtain a first feature map of the target region in the original target image before super-resolution. The first feature in the first feature map is fused with the second feature of the corresponding target region in the target super-resolution image to obtain a second feature map of the target region. The target super-resolution image is reconstructed based on the second feature map to obtain a reconstructed target super-resolution image. By using a target network model to optimize the features of the target region in the original target image corresponding to the target super-resolution image, and fusing the optimized features with the features of the corresponding region in the target super-resolution image, the method achieves optimized reconstruction of the features of the corresponding region in the target super-resolution image, thereby improving the display quality of the target super-resolution image. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the specific embodiments or related technologies of this disclosure, the accompanying drawings used in the description of the specific embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a schematic diagram illustrating an application scenario of the image processing method according to an embodiment of the present disclosure;

[0012] Figure 2 This is a schematic flowchart of an image processing method according to an embodiment of the present disclosure;

[0013] Figure 3A This is a schematic diagram corresponding to the image processing method according to the embodiments of this disclosure;

[0014] Figure 3B This is a schematic diagram corresponding to the image processing method according to the embodiments of this disclosure;

[0015] Figure 4 This is a structural block diagram of an image processing apparatus according to an embodiment of the present disclosure;

[0016] Figure 5 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0018] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0019] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0020] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0021] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0022] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0023] With the development of multimedia technology, the visual quality of images, as carriers of information dissemination, is constantly being improved. Super-resolution (SR) technology is commonly used to optimize image resolution and obtain super-resolution images, specifically as follows: Figure 1 As shown, the image to be super-resolution processed is sent to the server 20 through the terminal device 10, so that the computing power of the server 20 can be combined to perform super-resolution and other processing; however, during the super-resolution processing, the target objects in the obtained super-resolution image may have problems such as structural distortion, overly sharp white edges, and unrealistic textures.

[0024] Specifically, in related technologies, super-resolution (SR) processing of images is usually achieved using network models built on convolutional neural networks (CNNs) or generative adversarial networks (GANs). However, due to structural defects in the network models themselves, in low-resolution and blurry scenes, a large amount of image detail information is lost at low resolution. The convolutional kernels in CNN and GAN models are unable to accurately capture the complete structural features of objects. In the reconstruction process, errors occur in piecing together or distorting the shape of objects, causing buildings, people, objects, etc. in the image to exhibit geometric deformations that do not conform to realistic logic, resulting in structural distortion problems that affect visual recognition and understanding. In addition, in order to compensate for the lack of clarity caused by low resolution, related models tend to over-enhance edge information, which makes the originally smooth transition of object edges become abnormally sharp, appearing as abrupt bright white lines, i.e., "white edges," in the image. These "white borders" not only detract from the aesthetics of the image but may also mislead visual attention, obscuring the true content the image is meant to convey or causing artifacts. In high-definition scenes, while the adversarial game mechanism between the discriminator and generator in GAN networks aims to drive the generator to produce more realistic high-resolution images, in practice, the generator sometimes fabricates textures that do not exist in the original scene in order to "deceive" the discriminator and gain approval. These false textures may seem to increase the detail of the image, but in reality, they contradict the physical characteristics of the actual shooting scene, destroying the semantic authenticity of the image. For example, generating vegetation textures that do not conform to the natural growth pattern in natural landscape images, or having strange patterns on building surfaces that do not conform to the material texture. There is also insufficient fidelity in preserving subtle but crucial details in the original high-definition material. At the same time, in pursuing higher resolution metrics, CNN and GAN may easily ignore subtle changes in light and shadow, color transitions, and other details in the original high-definition scene, smoothing them out or incorrectly modifying them. This causes the super-resolution image to lose the delicate texture and artistic atmosphere contained in the original image, making it difficult to meet the professional needs of fields with demanding image quality requirements, such as film and television production and digital preservation of artworks.

[0025] According to an embodiment of this disclosure, an image processing method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0026] This embodiment provides an image processing method that can be used on any electronic device capable of executing the method, such as a server. Figure 2 This is a flowchart of an image processing method according to an embodiment of the present disclosure, such as... Figure 2 As shown, the process includes the following steps:

[0027] Step S101: Determine the target region that needs to be optimized in the target super-resolution image to be processed. The target region represents the region where the feature to be optimized is located corresponding to the target object in the target super-resolution image. The target super-resolution image is used to represent the image after super-resolution processing of the original target image.

[0028] For example, the target original image may be a low-resolution image caused by the performance limitations of the image acquisition device or the bandwidth transmission and storage device performance limitations; the target super-resolution image may be an image obtained after super-resolution processing of the low-resolution target original image; the target object to be optimized features in the target super-resolution image may include the target object’s structural features (such as contour morphology features, etc.), edge features, and texture features, etc., which are target object-related display features that affect the display quality of the target object. This application embodiment does not limit this.

[0029] One way to determine the features to be optimized corresponding to the target object in the super-resolution image is to compare the target super-resolution image with the original target image before super-resolution, and then use the features that do not meet the requirements as the features to be optimized. For example, when comparing contour morphology features, such as... Figure 3A and Figure 3B As shown, Figure 3A The original image 1 contains object A, whose outline is rectangular. The "dashed" area of ​​the rectangle indicates that the resolution of this part of the outline is low. Figure 3B The target super-resolution image 2 corresponds to the original target image 1. Target object B in the super-resolution image 2 is obtained after super-resolution processing of target object A. Super-resolution processing causes distortion of the object's low-resolution contour regions. The rate of change (e.g., curvature) of the contour features of the target object before and after super-resolution can be detected. If the rate of change exceeds a preset threshold, the contour feature is considered a feature to be optimized. The region containing this feature can then be selected. For edge features, the edge sharpness of the target object before and after super-resolution processing can be compared to determine if it is a feature to be optimized. Similarly, for texture features, the clarity and accuracy of the texture features before and after super-resolution processing can be compared to determine if they are a feature to be optimized. The region containing this feature can then be selected to obtain the target region for optimization.

[0030] Step S102: Determine the target network model among multiple network models based on the features to be optimized. The multiple network models are configured to perform corresponding type feature optimization processing on the input image, and the type features corresponding to the multiple network models are not all the same.

[0031] For example, in this embodiment, multiple network models can be pre-configured to optimize features of different types of target objects. The network models of each type are pre-trained in supervised mode using a large number of image training samples containing different target objects, enabling them to optimize the features of target objects in the target region of the input image. The multiple network models in this embodiment may include network models for structural feature optimization, edge feature optimization, and texture feature optimization. The types of features corresponding to the multiple network models configured in this embodiment are not all the same. Multiple network models can be configured for optimizing the same type of feature. The same type of feature to be optimized can be input into multiple network models that can be used to optimize that type of feature, and the multiple optimization results can be combined to obtain the final feature optimization result. Alternatively, the multiple network models configured for optimizing the same type of feature may correspond to different types of target objects. For example, for structural feature optimization, a network model for optimizing the structural features of architectural targets and a network model for optimizing the structural features of artworks can be configured. The appropriate network model for structural feature optimization can be selected based on the type of target object. Based on the features to be optimized of the target object, the network model used for corresponding feature optimization is determined from among the multiple network models as the target network model.

[0032] Step S103: Input the original target image into the target network model to obtain the first feature map of the target region in the original target image.

[0033] For example, the target original image corresponding to the target super-resolution image is input into the target network model, so that the target network model optimizes the target original image with corresponding features. Since the target super-resolution image has display quality problems, the target original image corresponding to the target super-resolution image is used as the optimization basis, which makes it easier for the target network model to perform accurate optimization. In this embodiment of the application, the target original image is converted to a target format, such as PNG format, before being input into the target network model, and noise reduction and normalization processing are performed to improve the quality of the target original image.

[0034] The optimization results of this target network model can yield the first feature map corresponding to the same target region. Specifically, this network model for structural feature optimization can determine the optimization direction and effect of the target object during super-resolution processing based on the type of target object contained in the identified original target image, ensuring the fusion effect when fused with the target super-resolution image. The specific optimization direction and effect can be obtained through supervised training during the model training phase. Based on the feature map output by the target network model, the first feature map corresponding to the target region is obtained by combining the spatial location information of the target region. Taking this target network model as an example, the original target image is input into this network model for structural feature optimization, enabling the network model to perform structural optimization on the target objects in the original target image. Specifically, it can optimize the structural resolution of the corresponding structural features to be optimized to improve the clarity of the corresponding structural features, and then construct the first feature map based on the pixel features corresponding to the target region.

[0035] Step S104: The first feature in the first feature map is fused with the second feature of the corresponding target region in the target super-resolution image to obtain the second feature map of the target region.

[0036] For example, the first feature in the first feature map corresponding to the target region in the original target image is fused with the second feature of the corresponding target region in the super-resolution target image. Specifically, the features at the corresponding level can be fused according to their pixel spatial locations at the pixel level or sub-pixel level. The fusion process can be to superimpose the first feature and the second feature, and obtain the second feature map based on the superimposed features.

[0037] Step S105: Reconstruct the target super-resolution image based on the second feature map to obtain the reconstructed target super-resolution image. The second feature map is obtained by fusing the first feature with the second feature of the corresponding target region in the target super-resolution image. The features of the second feature map replace the features of the corresponding target region in the target super-resolution image, thereby realizing the "texturing" reconstruction process of the target region corresponding to the target super-resolution image and improving the display quality of the target super-resolution image.

[0038] The image processing method provided in this embodiment determines the target region in the target super-resolution image that needs optimization and represents the feature corresponding to the target object in the target super-resolution image. Based on the feature to be optimized, a target network model is determined among multiple network models so that the target network model optimizes the feature to be optimized to obtain a first feature map of the target region in the original target image before super-resolution. The first feature in the first feature map is fused with the second feature of the corresponding target region in the target super-resolution image to obtain a second feature map of the target region. The target super-resolution image is reconstructed based on the second feature map to obtain the reconstructed target super-resolution image. By using the target network model to optimize the features of the target region in the original target image corresponding to the target super-resolution image, and fusing the optimized features with the features of the corresponding region in the target super-resolution image, the method achieves optimized reconstruction of the features of the corresponding region in the target super-resolution image, thereby improving the display quality of the target super-resolution image.

[0039] In some optional implementations, determining the features to be optimized includes: inputting the target super-resolution image into an image quality detection model to obtain quality detection results of multiple types of features corresponding to the target object in the target super-resolution image; and taking the types of features whose quality detection results do not meet the preset quality conditions as features to be optimized.

[0040] For example, the image quality detection model can be pre-trained using multiple before-and-after super-resolution image comparison data. Through supervised training, the image quality detection model can accurately distinguish whether the corresponding target object in the super-resolution image has display quality problems. Features whose quality detection results do not meet preset quality conditions are taken as features to be optimized. When the detected features to be optimized include multiple types, such as the structural features, edge features, and texture features of the target object in the target super-resolution image all needing optimization, the target region containing the corresponding structural feature to be optimized is marked in the original target image, and the original target image is input into the target network model for structural feature optimization; similarly, the target region containing the corresponding edge feature to be optimized is marked in the original target image, and the original target image is input into the target network model for edge feature optimization; and the target region containing the corresponding texture feature to be optimized is marked in the original target image, and the original target image is input into the target network model for texture feature optimization.

[0041] In some optional implementations, step S104 includes: determining the first weight data corresponding to the first feature and the second weight data corresponding to the second feature; and performing a fusion process on the first feature and the second feature based on the first weight data and the second weight data.

[0042] For example, the first weight data corresponding to the first feature and the second weight data corresponding to the second feature can be pre-configured. For instance, the first weight data and the second weight data can each be configured to 0.5. That is, when fusing the first feature and the second feature, by assigning weights of 0.5 respectively, the first feature and the second feature each contribute half of the fusion result. Alternatively, the first weight data corresponding to the first feature and the second weight data corresponding to the second feature can also be determined based on the feature quality of the corresponding target object's feature to be optimized in the target super-resolution image. For example, multiple levels of quality problem severity can be pre-configured. When the structural features of the target object in the target super-resolution image need optimization, but its structural feature quality is... If the problem is a minor quality issue, a weight of 0.7 can be assigned, and the weight of the first feature corresponding to the original target image can be set to 0.3. This means that the second feature can be prioritized during feature fusion, and the first feature is used to assist in optimization. Similarly, if the edge features of the target object in the super-resolution image need to be optimized, but the edge feature quality problem is a major quality issue, a weight of 0.3 can be assigned to the second weight data, and the weight of the first feature corresponding to the original target image can be set to 0.7. This means that the first feature can be prioritized during feature fusion, and the second feature is used to assist in optimization. The specific fusion method is shown in the following formula (1):

[0043] F = w1·F1 + w2·F2(1)

[0044] Where F represents the fused features, w1+w2=1, and w1 and w2 are the weights of the first feature F1 and the second feature F2, respectively.

[0045] In some optional implementations, determining the first weight data corresponding to the first feature and the second weight data corresponding to the second feature includes: obtaining the target quality index data corresponding to the feature to be optimized; and determining the first weight data corresponding to the first feature and the second weight data corresponding to the second feature based on the target quality index data corresponding to the feature to be optimized and the corresponding weight allocation strategy.

[0046] For example, the target quality index data corresponding to the feature to be optimized can be determined by an image quality detection model. Specifically, when the image quality detection model determines that the quality detection result of any type of feature in the target super-resolution image does not meet the preset quality conditions, it can determine that the feature of that type is the feature to be optimized and output the corresponding target quality index data of the feature to be optimized. For example, when the feature to be optimized is the structural feature of the target object, the target quality index data such as the sharpness and structural deformation degree corresponding to the structural feature that does not meet the quality requirements can be output. The target quality index data is matched with a pre-built weight allocation strategy. The weight allocation strategy pre-configures the second weight data corresponding to the target quality index data of different types of features. Then, according to the determined target quality index data and the weight allocation strategy, the second weight data corresponding to the second feature is matched. According to the correlation between the second weight data and the first weight data, the first weight data of the first feature can be determined.

[0047] In this embodiment of the application, the weight data corresponding to different target quality index data in the weight allocation strategy can be fixed, or it can be adaptively adjusted according to the user's requirements for the image representation style. The specific adjustment method is not limited. Those skilled in the art can make adaptive configurations in advance according to the weight adaptation rules under different styles. For example, the second weight data matched according to the target quality index data corresponding to the current target super-resolution image is 0.6 and the first weight data is 0.4. When the user inputs style features (such as artistic style), the second weight data and the first weight data are further adjusted based on the pre-matched style adjustment strategy (such as adjusting the second weight data to 0.65 and the first weight data to 0.35) to ensure the display quality of the generated target super-resolution image while meeting the user's requirements for the display style.

[0048] In some optional implementations, the target super-resolution image to be processed is the target super-resolution image corresponding to any video frame in the video to be processed.

[0049] For example, the video to be processed may contain video frames whose resolution does not meet the requirements due to limitations in the parameters of the acquisition device itself, or limitations in bandwidth, storage space, etc. Specifically, the video to be processed is segmented into frames, and super-resolution processing is performed on the video frames whose resolution does not meet the requirements. The super-resolution processing results can be detected using an image quality detection model, and the video frames whose quality does not meet the requirements are taken as the target super-resolution image. When the quality of the super-resolution processing results corresponding to multiple video frames does not meet the requirements, the target super-resolution image corresponding to each video frame is reconstructed using the target super-resolution image reconstruction process described in the above embodiment.

[0050] In some alternative implementations, the multiple network models include a first network model for structural feature optimization, a second network model for edge feature optimization, and a third network model for texture feature optimization, as well as a fourth network model for maintaining consistency of motion features between frames.

[0051] For example, in the embodiments of this application, the first network model for structural feature optimization can be a Structure Expert Network (SEN) for structural recovery, which uses a CNN with multi-scale dilated convolutions to capture structural information at different scales; the second network model for edge feature optimization can be an Edge Expert Network (EEN) for edge optimization, which introduces a GAN based on perceptual loss to make edge enhancement more in line with the characteristics of human vision; the third network model for texture feature optimization can be a Texture Expert Network (TEN) for texture generation, which combines semantic segmentation to assist in texture generation to ensure the rationality of the texture; and the fourth network model for maintaining the consistency of motion features between frames can be a Motion Expert Network (MEN) for ensuring the consistency between frames, which uses a combination of optical flow estimation and recurrent neural networks to track the motion of objects between frames to maintain the coherence of the video sequence.

[0052] For the video to be processed, super-resolution processing of a single frame may lead to inconsistencies in the position and shape of objects between adjacent frames, resulting in visual flickering or misalignment. Therefore, for the target super-resolution image as a video frame, the image quality detection model can detect the position or shape features of the target object in the current target super-resolution image based on the position or shape of the target object in the adjacent frames corresponding to the target super-resolution image. This determines whether the target object's position in the current target super-resolution image is within the target location region or whether its shape features have changed in a way that does not meet the requirements. When it is determined that the target super-resolution image as a video frame has corresponding quality problems, a fourth network model used to maintain the consistency of motion features between frames can be used in conjunction with optical flow estimation information in the video containing the target super-resolution image to optimize the position or shape features of the target object in the original target image corresponding to the target super-resolution image. Specifically, optical flow estimation methods such as PWC-Net can be used to obtain pixel-level motion vectors (u,v) to characterize the pixel correspondence between frames. This allows the trained fourth network model, which is used to maintain the consistency of motion features between frames, to optimize the position or shape features of the target object in the original image based on these pixel-level motion vectors (u,v).

[0053] In some optional implementations, step S105 includes: when the resolution of the second feature map does not meet the preset requirements, performing resolution optimization processing on the second feature map to obtain a resolution-optimized second feature map; and reconstructing the target super-resolution image based on the resolution-optimized second feature map to obtain a reconstructed target super-resolution image.

[0054] For example, in order to save computing power, when optimizing the target object in the original target image in the embodiments of this application, the resolution of the optimized second feature map may not meet the preset requirements for image display. For example, the required resolution is 1080P, but in order to save computing power, it is first optimized to 720P. At this time, the resolution of the second feature map can be optimized by deconvolution layer or upsampling module to achieve 1080P. Based on the second feature map after resolution optimization, the target super-resolution image is reconstructed to obtain the reconstructed target super-resolution image.

[0055] In some optional implementations, the method further includes performing visual optimization processing on the reconstructed target super-resolution image when the visual effect data of the reconstructed target super-resolution image does not meet the usage requirements.

[0056] For example, the visual effect data may include color data or contrast data of the target super-resolution image, etc., and this application embodiment does not limit this. The visual effect data of the reconstructed target super-resolution image can be detected by an image quality detection model. When the detected visual effect data does not meet the usage requirements, a pre-trained visual processing model can be used for optimization. Furthermore, adaptive visual optimization can be performed by combining the image style corresponding to the target super-resolution image.

[0057] As a specific implementation of this application, for videos originating from older equipment, with low resolution and blurry, shaky images, the front-end preprocessing unit first breaks the video down frame by frame, converts it into a suitable processing format (such as PNG format), and filters out most noise interference to stabilize the data baseline. Next, structural experts in the multi-expert network model cluster use a multi-scale feature capture mechanism to outline the contours of people and vehicles in the image, repairing structural deformations caused by low resolution; edge experts follow perception optimization principles to strengthen object edges and eliminate over-sharpness and white edges; texture experts combine scene semantic information to reasonably "complete" realistic textures for background areas including roads, building walls, etc.; the motion expert network model uses optical flow tracking and motion compensation technology to effectively counteract image shakiness and ensure smooth transitions between frames. The fusion scheduling center dynamically optimizes the fusion weights of each expert's output based on the real-time changes in display quality and motion amplitude of each frame; finally, the back-end reconstruction and optimization unit transforms the fusion results into clear and stable high-resolution video, providing reliable and accurate visual data for security monitoring; or, for a high-definition movie clip with artistic value but limited resolution, the goal is to improve resolution while preserving the original artistic texture. The preprocessing stage, while maintaining a standardized format, prioritizes the preservation of original color information. Structural experts reconstruct the scene architecture and character postures from the film; edge experts optimize the edges of facial features and clothing with a cinematic aesthetic, creating a soft visual experience; texture experts reference the style and cultural background of other films of the same period, adding realistic textures that fit the context; and motion experts simulate the unique camera movements of filmmaking to ensure the video is smooth and rhythmic after super-resolution. Through comprehensive processing, the film achieves a significant increase in resolution while fully meeting the needs of film restoration.

[0058] The image processing method provided in this application integrates multiple expert networks with specific expertise, such as Structural Expert Network (SEN), Edge Expert Network (EEN), Texture Expert Network (TEN), and Motion Expert Network (MEN), allowing each expert network to focus on solving the corresponding feature dimension in the video super-resolution process. In low-resolution, blurry scenes, SEN employs multi-scale dilated convolution, which can capture structural cues at different scales in blurry images with severe information loss, accurately restore the original geometry of objects, effectively suppress structural distortion, and ensure the accuracy of the basic framework of the image. EEN introduces a GAN based on perceptual loss to conform to the characteristics of human visual perception. The system optimizes edges using a natural approach, avoiding harsh over-sharpening. While improving clarity, it eliminates white edges, resulting in natural and smooth edge transitions that align with visual aesthetics. MEN combines optical flow estimation with recurrent neural networks to closely track object motion trajectories, overcoming frame-to-frame continuity disruptions caused by low resolution and image jitter, significantly reducing flickering and ensuring stable and smooth video playback. In high-definition scenes, TEN combines semantic segmentation to assist in texture generation. Instead of blindly pursuing rich detail, it relies on the semantic information of the image to ensure that every generated texture is realistic and reasonable, eliminating false and abrupt textures, maintaining the realism and semantic logic of the image, and comprehensively improving video quality.

[0059] This embodiment also provides an image processing apparatus for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0060] This embodiment provides an image processing device, such as... Figure 4 As shown, it includes:

[0061] The first determining module 201 is used to determine the target region that needs to be optimized in the target super-resolution image to be processed. The target region represents the region where the feature to be optimized is located corresponding to the target object in the target super-resolution image. The target super-resolution image is used to represent the image after super-resolution processing of the original target image.

[0062] The second determining module 202 is used to determine the target network model among multiple network models based on the features to be optimized. The multiple network models are configured to perform corresponding type feature optimization processing on the input image, and the type features corresponding to the multiple network models are not all the same.

[0063] Input module 203 is used to input the original target image into the target network model to obtain the first feature map of the target region in the original target image;

[0064] The fusion module 204 is used to fuse the first feature in the first feature map with the second feature of the corresponding target region in the target super-resolution image to obtain the second feature map of the target region.

[0065] The reconstruction module 205 is used to reconstruct the target super-resolution image based on the second feature map to obtain the reconstructed target super-resolution image.

[0066] The image processing apparatus provided in this embodiment determines the target region in the target super-resolution image to be optimized, which represents the feature to be optimized corresponding to the target object in the target super-resolution image. Based on the feature to be optimized, a target network model is determined among multiple network models so that the target network model optimizes the feature to be optimized to obtain a first feature map of the target region in the original target image before super-resolution. The first feature in the first feature map is fused with the second feature of the corresponding target region in the target super-resolution image to obtain a second feature map of the target region. Based on the second feature map, the target super-resolution image is reconstructed to obtain a reconstructed target super-resolution image. By using the target network model to optimize the features of the target region in the original target image corresponding to the target super-resolution image, and fusing the optimized features with the features of the corresponding region in the target super-resolution image, the apparatus achieves optimized reconstruction of the features of the corresponding region in the target super-resolution image, thereby improving the display quality of the target super-resolution image.

[0067] In some optional embodiments, the apparatus further includes: a third determining module, used to input the target super-resolution image into the image quality detection model to obtain the quality detection results of multiple types of features corresponding to the target object in the target super-resolution image; and to take the types of features whose quality detection results do not meet the preset quality conditions as features to be optimized.

[0068] In some optional implementations, the fusion module 204 includes: a determining submodule, used to determine the first weight data corresponding to the first feature and the second weight data corresponding to the second feature; and a fusion submodule, used to perform fusion processing on the first feature and the second feature based on the first weight data and the second weight data.

[0069] In some optional implementations, the determining submodule includes: an acquisition unit, used to acquire target quality index data corresponding to the feature to be optimized; and a determining unit, used to determine first weight data corresponding to the first feature and second weight data corresponding to the second feature based on the target quality index data corresponding to the feature to be optimized and the corresponding weight allocation strategy.

[0070] In some optional implementations, the target super-resolution image to be processed is the target super-resolution image corresponding to any video frame in the video to be processed.

[0071] In some alternative implementations, the multiple network models include: a first network model for structural feature optimization; a second network model for edge feature optimization; a third network model for texture feature optimization; and a fourth network model for maintaining consistency of motion features between frames.

[0072] In some optional implementations, the reconstruction module 205 includes: an optimization submodule, used to perform resolution optimization processing on the second feature map when the resolution of the second feature map does not meet the preset requirements, to obtain a resolution-optimized second feature map; and a reconstruction submodule, used to reconstruct the target super-resolution image based on the resolution-optimized second feature map, to obtain a reconstructed target super-resolution image.

[0073] In some alternative implementations, the apparatus further includes an optimization module for performing visual optimization processing on the reconstructed target super-resolution image when the visual effect data of the reconstructed target super-resolution image does not meet the usage requirements.

[0074] The image processing apparatus provided in this disclosure can execute the image processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method. Further functional descriptions of the various modules and units described above are the same as in the corresponding embodiments described above, and will not be repeated here.

[0075] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure.

[0076] The following is a detailed reference. Figure 5 The diagram illustrates a structural schematic suitable for implementing an electronic device according to embodiments of the present disclosure. The electronic device may include a processor (e.g., a central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from memory 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of the electronic device. The processor 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0077] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; memory devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown, and more or fewer devices may be implemented or have instead.

[0078] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a memory 508, or installed from a ROM 502. When the computer program is executed by the processor 501, it performs the functions defined in the image processing method of embodiments of this disclosure.

[0079] Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0080] This disclosure also provides a computer-readable storage medium in which the methods described in this disclosure can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium may also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the image processing method shown in the above embodiments is implemented.

[0081] A portion of this disclosure can be applied to computer program products, such as computer program instructions, which, when executed by a computer, can invoke or provide methods and / or technical solutions according to this disclosure through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, and installation package files. Accordingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions; the computer compiling the instructions and then executing the corresponding compiled program; the computer reading and executing the instructions; or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0082] Although embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. An image processing method, characterized in that, The method includes: Determine the target region that needs to be optimized in the target super-resolution image to be processed. The target region represents the region where the feature to be optimized is located corresponding to the target object in the target super-resolution image. The target super-resolution image is used to represent the image after super-resolution processing of the original target image. Based on the features to be optimized, a target network model is determined among multiple network models. The multiple network models are configured to perform corresponding type feature optimization processing on the input image, and the type features corresponding to the multiple network models are not all the same. The original target image is input into the target network model to obtain the first feature map of the target region in the original target image; The first feature in the first feature map is fused with the second feature of the corresponding target region in the target super-resolution image to obtain the second feature map of the target region. The target super-resolution image is reconstructed based on the second feature map to obtain the reconstructed target super-resolution image.

2. The method according to claim 1, characterized in that, Determining the feature to be optimized includes: The target super-resolution image is input into the image quality detection model to obtain the quality detection results of multiple types of features corresponding to the target object in the target super-resolution image; The type of feature whose quality test results do not meet the preset quality conditions is taken as the feature to be optimized.

3. The method according to claim 1, characterized in that, The step of fusing the first feature in the first feature map with the second feature of the corresponding target region in the target super-resolution image includes: Determine the first weight data corresponding to the first feature and the second weight data corresponding to the second feature; The first feature and the second feature are fused based on the first weight data and the second weight data.

4. The method according to claim 3, characterized in that, Determining the first weight data corresponding to the first feature and the second weight data corresponding to the second feature includes: Obtain the target quality index data corresponding to the feature to be optimized; Based on the target quality index data corresponding to the feature to be optimized and the corresponding weight allocation strategy, the first weight data corresponding to the first feature and the second weight data corresponding to the second feature are determined.

5. The method according to claim 1, characterized in that, The target super-resolution image to be processed is the target super-resolution image corresponding to any video frame in the video to be processed.

6. The method according to claim 5, characterized in that, The plurality of network models include: The first network model used for structural feature optimization; A second network model for edge feature optimization; A third network model for texture feature optimization; and, A fourth network model is used to maintain consistency of motion features between frames.

7. The method according to claim 1, characterized in that, The step of reconstructing the target super-resolution image based on the second feature map to obtain the reconstructed target super-resolution image includes: When the resolution of the second feature map does not meet the preset requirements, the resolution of the second feature map is optimized to obtain a resolution-optimized second feature map. Based on the second feature map after resolution optimization, the target super-resolution image is reconstructed to obtain the reconstructed target super-resolution image.

8. The method according to claim 1 or 6, characterized in that, The method further includes: If the visual effect data of the reconstructed target super-resolution image does not meet the usage requirements, visual optimization processing is performed on the reconstructed target super-resolution image.

9. An image processing apparatus, characterized in that, The device includes: The first determining module is used to determine the target region that needs to be optimized in the target super-resolution image to be processed. The target region represents the region where the feature to be optimized is located corresponding to the target object in the target super-resolution image. The target super-resolution image is used to represent the image after super-resolution processing of the original target image. The second determining module is used to determine a target network model among multiple network models based on the features to be optimized. The multiple network models are respectively configured to perform corresponding type feature optimization processing on the input image, and the type features corresponding to the multiple network models are not all the same. The input module is used to input the original target image into the target network model to obtain a first feature map of the target region in the original target image; The fusion module is used to fuse the first feature in the first feature map with the second feature of the corresponding target region in the target super-resolution image to obtain the second feature map of the target region. The reconstruction module is used to reconstruct the target super-resolution image based on the second feature map to obtain the reconstructed target super-resolution image.

10. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the image processing method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the image processing method according to any one of claims 1 to 8.

12. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the image processing method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Image super-resolution enhancement method and system

    CN115409704A

  • Training method of image enhancement model, and image enhancement method and device

    CN118154445A