Image processing method and device, electronic equipment and storage medium

By using multiple network models to optimize and fuse the features of the target area during image super-resolution processing, the image display quality problem is solved, structural distortion and texture authenticity are improved, and high-quality display of film, television and art works is achieved.

CN120765482APending Publication Date: 2025-10-10BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510874449.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing image super-resolution technology can easily lead to problems such as target object structural distortion, overly sharp white edges, and unrealistic textures during the processing process, making it difficult to meet the professional image quality requirements of film and television production and artistic works.

Method used

By determining the areas that need to be optimized in the target super-resolution image, multiple network models are used to optimize the structure, edge and texture features respectively, and then combined with the feature map of the target original image for fusion reconstruction to improve the image display quality.

Benefits of technology

It effectively suppresses structural distortion, eliminates white edges, generates realistic textures, and enhances image visual recognition and understanding, meeting the high-quality display requirements of film, television, and art works.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765482A_ABST
    Figure CN120765482A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and device, electronic equipment and a storage medium, and relates to the technical field of computer processing.The method comprises the steps that a target area needing to be optimized in a to-be-processed target super-resolution image is determined, and the target super-resolution image is used for representing an image obtained after super-resolution processing of a target original image; determining a target network model in a plurality of network models based on the to-be-optimized features, the network models being configured to perform optimization processing of corresponding type features on the input image, the type features corresponding to all the network models being not completely the same; inputting a target original image into the target network model to obtain a first feature map of a target area in the target original image; performing fusion processing on a first feature in the first feature map and a second feature corresponding to the target region in the target super-division image to obtain a second feature map of the target region; and reconstructing the target super-resolution image based on the second feature map to obtain the reconstructed target super-resolution image, thereby improving the display quality of the target super-resolution image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer processing technology, and in particular to an image processing method, device, electronic device, storage medium, and program product. Background Art

[0002] With the development of multimedia technology, the expectation for the visual quality of images, as a carrier of information dissemination, continues to rise. In related technologies, super-resolution (SR) technology is often used to optimize image resolution to produce super-resolution images. However, the super-resolution process can cause structural distortion, overly sharp white edges, and unrealistic textures in the resulting super-resolution images. Therefore, it is urgent to propose an image processing method to improve image display quality. Summary of the Invention

[0003] In view of this, the present disclosure provides an image processing method, apparatus, electronic device, storage medium, and program product to improve image display quality.

[0004] In a first aspect, the present disclosure provides an image processing method, comprising: determining a target area to be optimized in a target super-resolution image to be processed, the target area representing an area where features to be optimized corresponding to a target object in the target super-resolution image are located, and the target super-resolution image is used to represent the image of the target original image after super-resolution processing; determining a target network model from a plurality of network models based on the features to be optimized, the plurality of network models being respectively configured to perform optimization processing on corresponding types of features of an input image, and the types of features corresponding to the plurality of network models are not all the same; inputting the target original image into the target network model to obtain a first feature map of the target area in the target original image; fusing a first feature in the first feature map with a second feature corresponding to the target area in the target super-resolution image to obtain a second feature map of the target area; and reconstructing the target super-resolution image based on the second feature map to obtain a reconstructed target super-resolution image.

[0005] In a second aspect, the present disclosure provides an image processing device, which includes: a first determination module for determining a target area to be optimized in a target super-resolution image to be processed, the target area representing the area where the features to be optimized corresponding to the target object in the target super-resolution image are located, and the target super-resolution image is used to represent the image of the target original image after super-resolution processing; a second determination module for determining a target network model from a plurality of network models based on the features to be optimized, the plurality of network models being respectively configured to perform optimization processing on the input image for features of corresponding types, and the types of features corresponding to the plurality of network models are not all the same; an input module for inputting the target original image into the target network model to obtain a first feature map of the target area in the target original image; a fusion module for fusing the first feature in the first feature map with the second feature corresponding to the target area in the target super-resolution image to obtain a second feature map of the target area; a reconstruction module for reconstructing the target super-resolution image based on the second feature map to obtain a reconstructed target super-resolution image.

[0006] In a third aspect, the present disclosure provides an electronic device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the image processing method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0007] In a fourth aspect, the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the image processing method of the first aspect or any corresponding embodiment thereof.

[0008] In a fifth aspect, the present disclosure provides a computer program product, comprising computer instructions for causing a computer to execute the image processing method of the first aspect or any corresponding embodiment thereof.

[0009] The image processing method provided by the embodiment of the present disclosure determines a target area in a target super-resolution image to be processed, where a feature to be optimized corresponding to a target object in the target super-resolution image is located; determines a target network model from multiple network models based on the feature to be optimized so that the target network model optimizes the feature to be optimized to obtain a first feature map of the target area in the target original image before super-resolution corresponding to the target super-resolution image; fuses a first feature in the first feature map with a second feature of the corresponding target area in the target super-resolution image to obtain a second feature map of the target area; reconstructs the target super-resolution image based on the second feature map to obtain a reconstructed target super-resolution image; optimizes the features of the target area to be optimized in the target original image corresponding to the target super-resolution image by using the target network model; and fuses the features obtained after the optimization with the features of the corresponding area in the corresponding target super-resolution image to optimize and reconstruct the features of the corresponding area in the target super-resolution image, thereby improving the display quality of the target super-resolution image. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the related technologies, the following briefly introduces the drawings required for use in the specific embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0011] Figure 1 is a schematic diagram of an application scenario of the image processing method according to an embodiment of the present disclosure;

[0012] Figure 2 is a flowchart of an image processing method according to an embodiment of the present disclosure;

[0013] Figure 3A is a schematic diagram corresponding to an image processing method according to an embodiment of the present disclosure;

[0014] Figure 3B is a schematic diagram corresponding to an image processing method according to an embodiment of the present disclosure;

[0015] Figure 4 is a structural block diagram of an image processing apparatus according to an embodiment of the present disclosure;

[0016] Figure 5 Schematic diagram of the hardware structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0017] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present disclosure.

[0018] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0019] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0020] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0021] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0022] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0023] With the development of multimedia technology, the expectation of visual quality of images as a carrier of information dissemination continues to increase. Super-resolution (SR) technology is usually used to optimize the image resolution to obtain super-resolution images. Specifically, Figure 1 As shown, the image to be super-resolution processed is sent to the server 20 through the terminal device 10, so that super-resolution processing can be performed in combination with the computing power of the server 20; however, during the super-resolution processing, the target object in the obtained super-resolution image may have structural distortion, overly sharp white edges, unrealistic texture and other problems.

[0024] Specifically, in the relevant technologies, network models built based on convolutional neural networks (CNN) or generative adversarial networks (GAN) are usually used to achieve super-resolution (SR) processing of image resolution. However, due to the structural defects of the network models themselves, in low-resolution and blurred scenes, since a large amount of image detail information is lost at low resolution, the convolution kernels in the CNN and GAN models are difficult to accurately capture the complete structural features of the object. In the reconstruction process, the shapes of objects are incorrectly pieced together or distorted, causing the buildings, people, and objects in the picture to present geometric deformations that are inconsistent with realistic logic, resulting in structural distortion problems, affecting visual recognition and understanding; and in order to compensate for the lack of clarity caused by low resolution, the relevant models tend to over-enhance edge information, which makes the edges of objects that originally had smooth transitions become abnormally sharp, which appear as abrupt bright white lines, namely "white edges" on the image. These "white edges" not only destroy the aesthetic quality of the image but can also mislead visual attention, obscuring the image's intended content or causing artifacts. In high-definition scenarios, the adversarial game between the discriminator and the generator in the GAN network is designed to force the generator to produce more realistic high-resolution images. However, in practice, to "trick" the discriminator and gain recognition, the generator sometimes fabricates textures that don't exist in the original scene. These false textures appear to add detail, but in reality contradict the physical characteristics of the real scene, undermining the semantic authenticity of the image. For example, vegetation textures that don't conform to natural growth patterns can be generated in natural scenery, or strange patterns on building surfaces that don't match the material texture. Furthermore, the fidelity of subtle but critical details in the original high-definition footage is insufficient. Furthermore, in their pursuit of high-resolution metrics, CNNs and GANs tend to overlook subtle lighting and color transitions in the original HD scene, smoothing or incorrectly modifying them. This causes the super-resolution image to lose the delicate texture and artistic atmosphere of the original image. This makes it difficult to meet the professional needs of demanding image quality in fields such as film and television production and digital preservation of artworks.

[0025] According to an embodiment of the present disclosure, an image processing method embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0026] In this embodiment, an image processing method is provided, which can be used on any electronic device capable of executing the method, such as a server. Figure 2 is a flowchart of an image processing method according to an embodiment of the present disclosure, such as Figure 2 As shown, the process includes the following steps:

[0027] Step S101, determining a target region to be optimized in a target super-resolution image to be processed, wherein the target region represents a region where features to be optimized corresponding to a target object in the target super-resolution image are located, and the target super-resolution image is used to represent an image after the target original image is super-resolved.

[0028] Exemplarily, the target original image may be a low-resolution image caused by performance limitations of the image acquisition device or bandwidth transmission and storage device performance limitations; the target super-resolution image may be an image obtained after super-resolution processing of the low-resolution target original image; the features to be optimized of the target object in the target super-resolution image may include structural features of the target object (such as contour morphological features, etc.), edge features, and texture features, etc., which are target-related display features that affect the display quality of the target object. The embodiments of the present application are not limited to this.

[0029] The method for determining the features to be optimized corresponding to the target object in the target super-resolved image can be to compare the target super-resolved image with the target original image before super-resolved features of the corresponding type, and use the features whose comparison results do not meet the requirements as the features to be optimized; taking the comparison of contour morphological features as an example, Figure 3A and Figure 3B As shown, Figure 3A The original target image 1 contains the target object A, the outline of which is a rectangle. The "dashed line" area of ​​the rectangle indicates that the resolution of the outline of this part is low; Figure 3B The target super-resolution image 2 corresponds to the target original image 1. The target object B in the target super-resolution image 2 is obtained by super-resolution processing of the target object A. The super-resolution processing technology causes the object contour to be distorted in the contour area with lower resolution in the target object A. The change rate (such as curvature) of the contour morphological features of the target object before and after super-resolution can be detected. When the change rate is greater than a preset threshold, the contour morphological feature is characterized as a feature to be optimized. The area where the feature to be optimized is located can be selected to obtain the area where the feature to be optimized is located. For edge features, the edge sharpness of the target object in the image before and after super-resolution processing can be compared to determine whether the edge feature of the target object is a feature to be optimized. Similarly, for texture features, the clarity and accuracy of the texture features of the target object in the image before and after super-resolution can be compared to determine whether the texture feature is a feature to be optimized, and the area where the feature to be optimized is located can be selected to obtain the target area to be optimized.

[0030] Step S102: determining a target network model from a plurality of network models based on the features to be optimized, wherein the plurality of network models are respectively configured to optimize corresponding types of features of the input image, and the types of features corresponding to the plurality of network models are not all the same.

[0031] For example, in embodiments of the present application, multiple network models can be pre-configured to optimize different types of features of a target object. The network models of the corresponding types are pre-trained using a large number of image training samples containing different targets, so that the network models of the corresponding types can optimize the features of the target object in the target area contained in the input image. The multiple network models in embodiments of the present application may include a network model for optimizing structural features, a network model for optimizing edge features, and a network model for optimizing texture features. The types of features corresponding to the multiple network models configured in embodiments of the present application are not all the same. For optimizing the same type of features, multiple network models can be configured. The same type of features to be optimized can be input into multiple network models that can be used to optimize that type of features, and the multiple optimization results can be combined to obtain the final feature optimization result. Alternatively, the multiple networks configured to optimize the same type of features can correspond to different types of targets. For example, for optimizing structural features, a network model can be configured for optimizing the structural features of architectural targets and a network model for optimizing the structural features of artworks. The corresponding network model for optimizing structural features can be selected based on the type of target object. Based on the target object's features to be optimized, the network model for optimizing the corresponding features is determined from the multiple network models as the target network model.

[0032] Step S103: input the target original image into the target network model to obtain a first feature map of the target area in the target original image.

[0033] Exemplarily, the target original image corresponding to the target super-resolution image is input into the target network model, so that the target network model optimizes the corresponding type of features of the target original image. Since the target super-resolution image has display quality problems, the target original image corresponding to the target super-resolution image is used as the optimization basis to facilitate accurate optimization of the target network model. In an embodiment of the present application, the target original image is converted to a target format before being input into the target network model, such as into PNG format, and denoising and normalization are performed to improve the quality of the target original image.

[0034] The first feature map corresponding to the target region can be obtained from the optimization result of the target network model. Specifically, the network model for structure feature optimization can determine the optimization direction and optimization effect of the target object when performing super-resolution processing based on the identified target object type contained in the target original image, so that the fusion effect can be ensured when the target super-resolution image is fused subsequently. The optimization direction and optimization effect can be obtained through supervised training in the model training stage. Based on the feature map output by the target network model, the first feature map corresponding to the target region is obtained in combination with the spatial position information of the target region. Taking the target network model as an example of the network model for structure feature optimization, the target original image is input into the network model for structure feature optimization, so that the network model for structure feature optimization can optimize the structure of the target object in the target original image. Specifically, the structure resolution of the corresponding structure feature to be optimized can be optimized to improve the clarity of the corresponding structure feature, and then the first feature map can be constructed according to the pixel features corresponding to the target region.

[0035] In step S104, the first feature in the first feature map is fused with the second feature of the corresponding target region in the target super-resolution image to obtain a second feature map of the target region.

[0036] Exemplarily, the first feature in the first feature map corresponding to the target region in the target original image is fused with the second feature of the corresponding target region in the target super-resolution image. Specifically, the features at the corresponding level can be fused according to the pixel space position at the pixel level or sub-pixel level. The fusion processing mode can be superimposing the first feature and the second feature, and obtaining the second feature map according to the superimposed features.

[0037] In step S105, the target super-resolution image is reconstructed based on the second feature map to obtain a reconstructed target super-resolution image. The second feature map obtained by fusing the first feature with the second feature of the corresponding target region in the target super-resolution image replaces the features of the corresponding target region in the target super-resolution image, realizing the "mapping" reconstruction processing of the target region corresponding to the target super-resolution image, and improving the display quality of the target super-resolution image.

[0038] The image processing method provided in the embodiment determines a target region in a target super-resolution image to be processed, which needs to be optimized and in which a to-be-optimized feature corresponding to a target object in the target super-resolution image is located, determines a target network model based on the to-be-optimized feature in multiple network models, so that the target network model optimizes and processes the to-be-optimized feature to obtain a first feature map of the target region in a target original image corresponding to the target super-resolution image before super-resolution, fuses a first feature in the first feature map with a second feature of a corresponding target region in the target super-resolution image, obtains a second feature map of the target region, and reconstructs the target super-resolution image based on the second feature map to obtain a reconstructed target super-resolution image. The features of the target region in the target original image corresponding to the target super-resolution image are optimized and processed by using the target network model, and the features obtained after the optimization processing are fused with the features of the corresponding region in the corresponding target super-resolution image to realize the optimization and reconstruction of the features of the corresponding region in the target super-resolution image, thereby improving the display quality of the target super-resolution image.

[0039] In some optional embodiments, the to-be-optimized feature is determined by inputting the target super-resolution image into an image quality detection model to obtain quality detection results of multiple types of features corresponding to the target object in the target super-resolution image, and taking a type of feature whose quality detection result does not satisfy a preset quality condition as the to-be-optimized feature.

[0040] For example, the image quality detection model can be trained in advance by using a plurality of image pairs before and after super-resolution, so that the image quality detection model can accurately distinguish whether the corresponding target object in the super-resolution image has a display quality problem through supervised training. The corresponding type of feature whose quality detection result does not satisfy the preset quality condition is taken as the to-be-optimized feature. When the detected to-be-optimized feature includes multiple types, such as the structure feature, the edge feature and the texture feature of the target object in the target super-resolution image, which all need to be optimized, the target region in which the corresponding to-be-optimized structure feature is located is marked in the target original image, and the target original image is input into a target network model for structure feature optimization. Similarly, the target region in which the corresponding to-be-optimized edge feature is located is marked in the target original image, and the target original image is input into a target network model for edge feature optimization. In addition, the target region in which the corresponding to-be-optimized texture feature is located is marked in the target original image, and the target original image is input into a target network model for texture feature optimization.

[0041] In some optional embodiments, the step S104 includes: determining first weight data corresponding to the first feature and second weight data corresponding to the second feature; and fusing the first feature and the second feature according to the first weight data and the second weight data.

[0042] Exemplarily, the first weight data corresponding to the first feature and the second weight data corresponding to the second feature can be pre-configured, such as configuring the first weight data and the second weight data to be 0.5 respectively, that is, when the first feature and the second feature are fused, by assigning a weight of 0.5 respectively, the first feature and the second feature each contribute half to the fusion result when fused; or the first weight data corresponding to the first feature and the second weight data corresponding to the second feature can also be determined according to the feature quality of the target object to be optimized in the target super-resolution image, such as pre-configuring multiple levels of quality problem degrees, when it is detected that the structural features of the target object in the target super-resolution image need to be optimized, but the quality of its structural features is problematic. If the quality problem is of a relatively mild degree, a weight of 0.7 can be assigned to the second weight data, and the weight of the first feature corresponding to the corresponding target original image can be set to 0.3, which means that the second feature can be prioritized as the basis during feature fusion, and the first feature is used for auxiliary optimization; similarly, when the edge features of the target object in the target super-resolution image need to be optimized, but the quality problem of the edge features is of a relatively severe degree, a weight of 0.3 can be assigned to the second weight data, and the weight of the first feature corresponding to the corresponding target original image can be set to 0.7, which means that the first feature can be prioritized as the basis during feature fusion, and the second feature is used for auxiliary optimization; the specific fusion method is shown in the following formula (1):

[0043] F=w1·F1+w2·F2(1)

[0044] Among them, F is the fused feature, w1+w2=1, w1 and w2 are the weights of the first feature F1 and the second feature F2 respectively.

[0045] In some optional embodiments, determining the first weight data corresponding to the first feature and the second weight data corresponding to the second feature includes: obtaining target quality indicator data corresponding to the feature to be optimized; determining the first weight data corresponding to the first feature and the second weight data corresponding to the second feature based on the target quality indicator data corresponding to the feature to be optimized and the corresponding weight allocation strategy.

[0046] Exemplarily, the target quality index data corresponding to the feature to be optimized can be determined by an image quality detection model. Specifically, when the image quality detection model determines that the quality detection result of any type of feature in the target super-resolution image does not meet the preset quality conditions, it can determine that the type of feature is the feature to be optimized and can also output the corresponding target quality index data of the corresponding feature to be optimized; for example, when the feature to be optimized is the structural feature of the target object, the target quality index data such as clarity and degree of structural deformation corresponding to the structural feature that does not meet the quality requirements can be output accordingly, and the target quality index data is matched with a pre-constructed weight allocation strategy, in which the second weight data corresponding to the target quality index data of different types of features is pre-configured. The second weight data corresponding to the corresponding second feature is matched according to the determined target quality index data and the weight allocation strategy; according to the association between the second weight data and the first weight data, the first weight data of the first feature can be determined.

[0047] In an embodiment of the present application, the weight data corresponding to different target quality indicator data in the weight allocation strategy can be fixed, or the weight data corresponding to different target quality indicator data can be adaptively adjusted according to the user's requirements for the image representation style received. The specific adjustment method is not limited. Those skilled in the art can make adaptive configurations in advance according to the weight adaptation rules under different styles. For example, the second weight data matched according to the target quality indicator data corresponding to the current target super-resolution image is 0.6 and the first weight data is 0.4. When the style features input by the user (such as artistic style) are received, the second weight data and the first weight data are further adjusted based on the pre-matched style adjustment strategy (such as adjusting the second weight data to 0.65 and the first weight data to 0.35) to ensure the display quality of the generated target super-resolution image while meeting the user's requirements for the display style.

[0048] In some optional implementations, the target super-resolution image to be processed is a target super-resolution image corresponding to any video frame in the video to be processed.

[0049] Exemplarily, the video to be processed may contain video frames whose resolution does not meet the requirements due to the parameter limitations of the acquisition device itself or the limitations of bandwidth, storage space, etc. Specifically, the video to be processed is frame-divided and the video frames whose resolution does not meet the requirements are super-resolution processed. The super-resolution processing results can be detected using an image quality detection model, and the video frames whose quality does not meet the requirements are used as the target super-resolution images. When the quality of the super-resolution processing results corresponding to multiple video frames does not meet the requirements, the target super-resolution image corresponding to each video frame is reconstructed using the target super-resolution image in the above embodiment.

[0050] In some optional embodiments, the multiple network models include, in addition to the first network model for structural feature optimization, the second network model for edge feature optimization, and the third network model for texture feature optimization, a fourth network model for maintaining consistency of motion features between frames.

[0051] Exemplarily, the first network model for structural feature optimization in the embodiment of the present application can be a Structure Expert Network (SEN) for structural recovery, which uses a CNN with multi-scale dilated convolution to capture structural information at different scales; the second network model for edge feature optimization can be an Edge Expert Network (EEN) for edge optimization, which introduces a GAN based on perceptual loss to make edge enhancement more in line with the visual characteristics of the human eye; the third network model for texture feature optimization can be a Texture Expert Network (TEN) for texture generation, which combines semantic segmentation to assist in generating texture to ensure the rationality of the texture; the fourth network model for maintaining the consistency of motion features between frames can be a Motion Expert Network (MEN) to ensure consistency between frames, which uses a combination of optical flow estimation and recurrent neural networks to track the motion of objects between frames to maintain the coherence of the video sequence.

[0052] For a video to be processed, when a single frame is super-resolved, the position and shape of the target object between adjacent frames may become inconsistent, resulting in visual flickering or misalignment. Therefore, for the target super-resolved image as a video frame, the image quality detection model can detect the position or shape features of the target object in the current target super-resolved image based on the position or shape of the target object in the adjacent frames corresponding to the target super-resolved image, and determine whether the position of the target object in the current target super-resolved image is within the target position area or whether the shape features have changed in a manner that does not meet the requirements. When it is determined that the target super-resolved image as a video frame has corresponding quality problems, the fourth network model for maintaining consistency of motion features between frames can be used in combination with the optical flow estimation information in the video where the target super-resolved image is located to optimize the position or shape features of the target object in the target original image corresponding to the target super-resolved image. Specifically, an optical flow estimation method such as an optical flow network (PWC-Net) can be used to obtain a pixel-level motion vector (u, v) for characterizing the inter-frame pixel correspondence relationship, so that the fourth network model trained to maintain the consistency of inter-frame motion features can optimize the position or morphological features of the target object in the target original image based on the pixel-level motion vector (u, v).

[0053] In some optional embodiments, step S105 includes: when the resolution of the second feature map does not meet the preset requirements, performing resolution optimization processing on the second feature map to obtain a second feature map after resolution optimization processing; based on the second feature map after resolution optimization processing, reconstructing the target super-resolution image to obtain a reconstructed target super-resolution image.

[0054] For example, in order to save computing power, when the embodiment of the present application optimizes the target object in the target original image, the resolution of the optimized second feature map may not meet the preset requirements for image display. For example, the required resolution is 1080P, but in order to save computing power, it is first optimized to 720P. At this time, the second feature map can be optimized through a deconvolution layer or an upsampling module to reach 1080P. Based on the second feature map after resolution optimization, the target super-resolution image is reconstructed to obtain a reconstructed target super-resolution image.

[0055] In some optional implementations, the method further includes: when the visual effect data of the reconstructed target super-resolution image does not meet the usage requirements, performing visual optimization processing on the reconstructed target super-resolution image.

[0056] For example, the visual effect data may include color data or contrast data of the target super-resolution image, etc., which is not limited in the embodiments of the present application. The visual effect data of the reconstructed target super-resolution image can be tested using an image quality detection model. If the detected visual effect data does not meet the usage requirements, a pre-trained visual processing model can be used for optimization processing. Furthermore, adaptive visual optimization can be performed in combination with the image style corresponding to the target super-resolution image.

[0057] As one specific embodiment of the present application, for the video from old devices, low resolution and blurred picture, first, the front-end preprocessing unit can disassemble the video by frame, convert it into a suitable processing format (such as PNG format), and filter out most of the noise interference at the same time, stabilize the data baseline. Then, the structural expert in the multi-expert network model cluster uses the multi-scale feature capture mechanism to outline the contours of people and vehicles in the picture, and repair the structural deformation caused by low resolution; the edge expert follows the perception optimization criterion to strengthen the object edges and eliminate the phenomenon of over-sharpness and white edges; the texture expert combines the scene semantic information to reasonably "complete" the real texture for the background area containing roads, building walls, etc. The motion expert network model uses the optical flow tracking and motion compensation technology to effectively offset the picture shaking and ensure smooth transition between frames. The fusion scheduling center dynamically optimizes the fusion weight of each expert output according to the real-time display quality, motion amplitude and other indicators of each frame; finally, the back-end reconstruction optimization unit converts the fusion result into a clear and stable high-resolution video, providing reliable and accurate visual materials for security monitoring; or for a high-definition movie segment with artistic value but limited resolution, it is expected to improve the resolution while retaining the original artistic quality. The preprocessing link focuses on protecting the original color information on the basis of regularizing the format. The structural expert restores the scene architecture and character posture in the film; the edge expert optimizes the edges of the character's face and clothes with film aesthetics as the guide, creating a soft visual experience; the texture expert adds real textures that fit the situation by referring to the style and cultural background of the same period movie; the motion expert simulates the unique camera effect of film shooting to ensure that the super-resolution video is smooth and rhythmic. After all-round processing, the film meets the needs of film restoration while its resolution is greatly improved.

[0058] The image processing method provided in the embodiment of the present application integrates multiple expert networks with specific expertise, such as structure expert network (SEN), edge expert network (EEN), texture expert network (TEN) and motion expert network (MEN), so that each expert network focuses on solving the corresponding problem feature dimension in the video super-resolution process; in low-resolution and blurred scenes, SEN adopts multi-scale dilated convolution to capture structural clues of different scales in blurred images with severe information loss, accurately restore the original geometric shape of the object, effectively suppress structural distortion, and ensure that the basic framework of the image is accurate; EEN introduces GAN based on perceptual loss to conform to the characteristics of human visual perception. It optimizes edges in a targeted way to avoid harsh over-sharpening, eliminates white edges while improving clarity, and makes edge transitions natural and smooth, in line with visual aesthetics; MEN combines optical flow estimation with recurrent neural networks to closely track the movement trajectory of objects, overcome the destruction of inter-frame continuity caused by factors such as low resolution and image jitter, greatly reduce flickering, and ensure the stability and smoothness of dynamic video playback; in high-definition scenes, TEN combines semantic segmentation to assist in texture generation. It does not blindly pursue rich details, but ensures that every generated texture is real and reasonable based on the semantic information of the image, eliminating the appearance of false and abrupt textures, maintaining the realism and semantic logic of the picture, and comprehensively improving the video quality.

[0059] This embodiment also provides an image processing device for implementing the above-mentioned embodiments and preferred implementations. Details already described will not be repeated. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0060] This embodiment provides an image processing device, such as Figure 4 Shown, including:

[0061] The first determining module 201 is used to determine a target region to be optimized in a target super-resolved image to be processed, where the target region represents a region where features to be optimized corresponding to the target object in the target super-resolved image are located, and the target super-resolved image represents an image after the target original image has been super-resolved;

[0062] A second determining module 202 is configured to determine a target network model from a plurality of network models based on the features to be optimized, wherein the plurality of network models are respectively configured to optimize corresponding types of features of the input image, and the types of features corresponding to the plurality of network models are not all the same;

[0063] An input module 203 is configured to input a target original image into a target network model to obtain a first feature map of a target region in the target original image;

[0064] A fusion module 204 is configured to fuse the first feature in the first feature map with the second feature corresponding to the target area in the target super-resolved image to obtain a second feature map of the target area;

[0065] The reconstruction module 205 is configured to reconstruct the target super-resolution image based on the second feature map to obtain a reconstructed target super-resolution image.

[0066] The image processing device provided by this embodiment determines a target area in a target super-resolution image to be processed, where a feature to be optimized corresponding to a target object in the target super-resolution image is located, determines a target network model from multiple network models based on the feature to be optimized so that the target network model optimizes the feature to be optimized to obtain a first feature map of the target area in the target original image before super-resolution corresponding to the target super-resolution image, fuses a first feature in the first feature map with a second feature of the corresponding target area in the target super-resolution image to obtain a second feature map of the target area, reconstructs the target super-resolution image based on the second feature map to obtain a reconstructed target super-resolution image; optimizes the features of the target area to be optimized in the target original image corresponding to the target super-resolution image by using the target network model, and fuses the features obtained after the optimization with the features of the corresponding area in the corresponding target super-resolution image to optimize and reconstruct the features of the corresponding area in the target super-resolution image, thereby improving the display quality of the target super-resolution image.

[0067] In some optional embodiments, the device also includes: a third determination module, which is used to input the target super-resolution image into an image quality detection model to obtain quality detection results of multiple type features corresponding to the target object in the target super-resolution image; and use the type features whose quality detection results do not meet the preset quality conditions as features to be optimized.

[0068] In some optional embodiments, the fusion module 204 includes: a determination submodule for determining the first weight data corresponding to the first feature and the second weight data corresponding to the second feature; and a fusion submodule for performing fusion processing on the first feature and the second feature based on the first weight data and the second weight data.

[0069] In some optional embodiments, the determination submodule includes: an acquisition unit for acquiring target quality indicator data corresponding to the feature to be optimized; a determination unit for determining first weight data corresponding to the first feature and second weight data corresponding to the second feature based on the target quality indicator data corresponding to the feature to be optimized and the corresponding weight allocation strategy.

[0070] In some optional implementations, the target super-resolution image to be processed is a target super-resolution image corresponding to any video frame in the video to be processed.

[0071] In some optional embodiments, the multiple network models include: a first network model for structural feature optimization; a second network model for edge feature optimization; a third network model for texture feature optimization; and a fourth network model for maintaining consistency of motion features between frames.

[0072] In some optional embodiments, the reconstruction module 205 includes: an optimization submodule, which is used to perform resolution optimization processing on the second feature map when the resolution of the second feature map does not meet the preset requirements, so as to obtain a second feature map after resolution optimization processing; and a reconstruction submodule, which is used to reconstruct the target super-resolution image based on the second feature map after resolution optimization processing, so as to obtain a reconstructed target super-resolution image.

[0073] In some optional embodiments, the apparatus further includes: an optimization module configured to perform visual optimization processing on the reconstructed target super-resolution image when the visual effect data of the reconstructed target super-resolution image does not meet usage requirements.

[0074] The image processing device provided in the embodiments of the present disclosure can execute the image processing method provided in any embodiment of the present disclosure, and has the functional modules and beneficial effects corresponding to the execution method. The further functional description of each of the above modules and units is the same as that of the corresponding embodiment above, and will not be repeated here.

[0075] Figure 5 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure.

[0076] The following specific reference Figure 5 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure. The electronic device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a memory 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the electronic device are also stored in the RAM 503. The processor 501, ROM 502, and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0077] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 5 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all of the devices shown, and more or fewer devices may be implemented or possessed instead.

[0078] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 509, or installed from the memory 508, or installed from the ROM 502. When the computer program is executed by the processor 501, the above-mentioned functions defined in the image processing method of the embodiment of the present disclosure are performed.

[0079] Figure 5 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0080] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the image processing method shown in the above embodiment is implemented.

[0081] Part of the present disclosure can be applied as a computer program product, for example, computer program instructions, when executed by a computer, through the operation of the computer, the method and / or technical solutions according to the present disclosure can be invoked or provided. Those skilled in the art should understand that the form of computer program instructions in computer readable medium includes but is not limited to source file, executable file, installation package file and the like, and accordingly, the way of computer program instructions executed by computer includes but is not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Here, the computer readable medium can be any available computer readable storage medium or communication medium accessible to the computer.

[0082] Although the embodiments of the present disclosure are described in conjunction with the drawings, various modifications and changes can be made by those skilled in the art without departing from the spirit and scope of the present disclosure, and such modifications and changes fall within the scope defined by the appended claims.

Claims

1. An image processing method, characterized in that: The method comprises: Determining a target region to be optimized in a target super-resolved image to be processed, wherein the target region represents a region where features to be optimized corresponding to a target object in the target super-resolved image are located, and the target super-resolved image is used to represent an image of the target original image after super-resolved processing; Determining a target network model from a plurality of network models based on the features to be optimized, wherein the plurality of network models are respectively configured to perform optimization processing on corresponding types of features of the input image, and the types of features corresponding to the plurality of network models are not all the same; Inputting the target original image into the target network model to obtain a first feature map of the target area in the target original image; Fusing the first feature in the first feature map with the second feature corresponding to the target area in the target super-resolved image to obtain a second feature map of the target area; The target super-resolution image is reconstructed based on the second feature map to obtain a reconstructed target super-resolution image.

2. The method according to claim 1, characterized in that Determining the feature to be optimized includes: Inputting the target super-resolution image into an image quality detection model to obtain quality detection results of multiple types of features corresponding to the target object in the target super-resolution image; The type feature whose quality detection result does not meet the preset quality condition is used as the feature to be optimized.

3. The method according to claim 1, characterized in that The fusing the first feature in the first feature map with the second feature of the corresponding target area in the target super-resolved image includes: Determining first weight data corresponding to the first feature and second weight data corresponding to the second feature; The first feature and the second feature are fused according to the first weight data and the second weight data.

4. The method according to claim 3, characterized in that The determining first weight data corresponding to the first feature and second weight data corresponding to the second feature includes: Obtaining target quality indicator data corresponding to the feature to be optimized; According to the target quality indicator data corresponding to the feature to be optimized and the corresponding weight allocation strategy, first weight data corresponding to the first feature and second weight data corresponding to the second feature are determined.

5. The method according to claim 1, wherein The target super-resolution image to be processed is a target super-resolution image corresponding to any video frame in the video to be processed.

6. The method according to claim 5, characterized in that The multiple network models include: A first network model for structural feature optimization; A second network model for edge feature optimization; A third network model for texture feature optimization; and The fourth network model is used to maintain the consistency of motion features between frames.

7. The method according to claim 1, characterized in that The reconstructing the target super-resolution image based on the second feature map to obtain a reconstructed target super-resolution image includes: When the resolution of the second feature map does not meet the preset requirement, performing resolution optimization processing on the second feature map to obtain a second feature map after resolution optimization processing; Based on the second feature map after the resolution optimization processing, the target super-resolution image is reconstructed to obtain a reconstructed target super-resolution image.

8. The method according to claim 1 or 6, characterized in that The method further comprises: When the visual effect data of the reconstructed target super-resolution image does not meet the use requirements, a visual optimization process is performed on the reconstructed target super-resolution image.

9. An image processing device, characterized in that: The device comprises: A first determination module is configured to determine a target region to be optimized in a target super-resolved image to be processed, wherein the target region represents a region in the target super-resolved image where features to be optimized corresponding to the target object are located, and the target super-resolved image represents an image of the target original image after super-resolved processing; a second determining module, configured to determine a target network model from a plurality of network models based on the features to be optimized, wherein the plurality of network models are respectively configured to perform optimization processing on corresponding types of features of the input image, and the types of features corresponding to the plurality of network models are not all the same; An input module, configured to input the target original image into the target network model to obtain a first feature map of the target area in the target original image; A fusion module, configured to fuse the first feature in the first feature map with the second feature corresponding to the target area in the target super-resolved image to obtain a second feature map of the target area; A reconstruction module is used to reconstruct the target super-resolution image based on the second feature map to obtain a reconstructed target super-resolution image.

10. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the image processing method according to any one of claims 1 to 8 by executing the computer instructions.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the image processing method according to any one of claims 1 to 8.

12. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the image processing method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Image super-resolution enhancement method and system

    CN115409704A

  • Training method of image enhancement model, and image enhancement method and device

    CN118154445A

  • Image processing method, apparatus, electronic device and computer readable storage medium

    US20210125313A1