Image Processing Method, Apparatus, Electronic Device, and Storage Medium
Through the trained image processing model, feature extraction and style integration of images to be processed is solved, and the problem of unsatisfactory image quality improvement processing efficiency in the prior art is achieved, and the effect of efficiently improving image quality and maintaining the real style is achieved.
Patent Information
- Application Number
- CN202210161789.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-02-22
AI Technical Summary
When the prior art improves image quality, the processing efficiency is not ideal and it is difficult to meet the practical application needs.
By calling the trained image processing model, the image to be processed is downsampled feature extraction, the image feature map and style feature map are determined, and the upsampled feature extraction is performed based on the style feature map to obtain the target image.
It improves the efficiency of image quality improvement, makes the target image more in line with the style of the original image, has more realistic visual effects, and meets practical application needs.
Smart Images

Figure CN115311152B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to technical fields such as artificial intelligence, cloud technology, and audio and video. Specifically, the present application relates to an image processing method, apparatus, electronic device, and storage medium. Background Art
[0002] With the rapid development of science and technology and the improvement of people's living needs, images or videos (which are essentially images) have appeared in various life scenarios. For example, more and more young people like to record or share their lives through images (photos) or videos. However, during the shooting or network transmission of a portrait, the image may be unclear and of low quality due to various reasons such as inaccurate focusing, too much noise, or too high a picture compression ratio. Especially for a human image (such as the user's face image), it will affect the user's experience. How to improve the image quality through image processing technology has become one of the important issues for researchers to study.
[0003] Currently, although there are already various different ways to improve image quality in the related art, the processing efficiency is not ideal and needs to be improved. Summary of the Invention
[0004] Embodiments of the present application provide an image processing method, apparatus, electronic device, and storage medium, which can solve the problem that the processing efficiency is not ideal when improving the image quality in the prior art. The technical solutions are as follows:
[0005] According to one aspect of the embodiments of the present application, an image processing method is provided, and the method includes:
[0006] Obtain a to-be-processed image of a target object;
[0007] Perform the following processing on the to-be-processed image by invoking a trained image processing model to obtain a target image corresponding to the to-be-processed image, where the image quality of the target image is higher than that of the to-be-processed image:
[0008] Perform downsampling feature extraction on the to-be-processed image to obtain an image feature map corresponding to the to-be-processed image;
[0009] Determine a style feature map corresponding to the to-be-processed image according to the image feature map;
[0010] Perform upsampling feature extraction on the image feature map based on the style feature map to obtain the target image.
[0011] Optionally, performing downsampling feature extraction on the to-be-processed image to obtain an image feature map corresponding to the to-be-processed image includes:
[0012] Perform downsampling feature extraction on the image to be processed to obtain at least one scale of image feature maps corresponding to the image to be processed;
[0013] Determine the style feature map corresponding to the image to be processed according to the image feature maps, including:
[0014] For each scale, determine the style feature map corresponding to this scale according to the image feature map of this scale;
[0015] Perform upsampling feature extraction on the image feature maps based on the style feature maps to obtain the target image, including:
[0016] Perform at least one upsampling feature extraction on the target feature map through at least one upsampling feature extraction module to obtain the target image. Wherein, in each upsampling feature extraction process, the style feature map of the corresponding scale is fused into the feature map of the corresponding scale obtained by each upsampling feature extraction. The target feature map is the image feature map of the smallest scale among at least one scale of image feature maps, and each upsampling feature extraction module corresponds to a style feature map of one scale.
[0017] Optionally, at least one scale of image feature maps includes at least two scales of image feature maps, and at least one upsampling feature extraction module includes at least two cascaded upsampling feature extraction modules;
[0018] Perform at least one upsampling feature extraction on the target feature map through at least one upsampling feature extraction module to obtain the target image, including:
[0019] For each upsampling feature extraction module, use the style feature map corresponding to this upsampling feature extraction module as an adjustment weight to adjust the network parameters of this upsampling feature extraction module. Perform upsampling processing on the input feature map corresponding to this upsampling feature extraction module through the adjusted network parameters to obtain the output feature map of this upsampling feature extraction module, and use the last output feature map as the target image;
[0020] Wherein, the input of the first upsampling feature extraction module is the target feature map, and the input of the module other than the first upsampling feature extraction module is the output feature map of the previous upsampling feature extraction module of this upsampling feature extraction module.
[0021] Optionally, the image processing model includes an image encoding network and an image generation network. The image encoding network includes a feature extraction module and a feature mapping module. By calling the trained image processing model, perform the following processing on the image to be processed to obtain the target image corresponding to the image to be processed, including:
[0022] Input the image to be processed into the feature extraction module, and perform downsampling feature extraction on the image to be processed through the feature extraction module to obtain at least one scale of image feature maps corresponding to the image to be processed;
[0023] Input the image feature maps of each scale into the feature mapping module respectively to obtain the style feature maps corresponding to each scale;
[0024] Based on the style feature maps corresponding to each scale and the target feature map, perform at least one upsampling feature extraction on the target feature map through the image generation network to obtain the target image corresponding to the image to be processed.
[0025] Optionally, the image processing model is trained in the following way:
[0026] Obtain the first training data and the initial neural network model. The training data includes multiple image pairs, each image pair includes a first sample image and a second sample image corresponding to the same image content, and the image quality of the second sample image in each image pair is higher than that of the first sample image. The initial neural network model includes an initial image processing model and an initial discriminator;
[0027] Train the initial neural network model based on the first training data until the value of the total loss function corresponding to the model meets the training end condition, and obtain the trained neural network model. Take the trained image processing model as the image processing model;
[0028] Among them, the input of the initial image processing model is each first sample image, and the output is the predicted target image corresponding to each first sample image. The input of the initial discriminator is the predicted target image corresponding to each first sample image and the second sample image, and the output is the discrimination result of the predicted target image corresponding to each first sample image and the second sample image;
[0029] The total loss function includes a first loss function and a second loss function. The value of the first loss function characterizes the difference between the predicted target image corresponding to each first sample image and the second sample image, and the value of the second loss function characterizes the accuracy of the discrimination result of the predicted target image corresponding to each first sample image and the second sample image.
[0030] Optionally, the image generation network of the initial image processing model and the initial discriminator are pre-trained using the second training data. When training the initial neural network model based on the first training data, the learning rate of the image encoding network of the initial image processing model is greater than the learning rates of the image generation network of the initial image processing model and the initial discriminator.
[0031] Optionally, obtaining the first training data includes:
[0032] Obtain each second sample image, perform image degradation processing on each second sample image respectively, and obtain each processed image;
[0033] Use the processed image corresponding to each second sample image as the first sample image corresponding to the second sample image, and obtain first training data based on each second sample image and the corresponding processed image.
[0034] Optionally, the image degradation processing includes at least one of adding blur processing, reducing image resolution processing, adding noise processing, or image format compression processing; the first training data includes images processed by at least two image degradation processing methods.
[0035] According to another aspect of the embodiments of the present application, an image processing device is provided, and the device includes:
[0036] An image acquisition module, configured to acquire a to-be-processed image of a target object;
[0037] An image processing module, configured to perform downsampling feature extraction on the to-be-processed image by calling a trained image processing model, obtain an image feature map corresponding to the to-be-processed image, and determine a style feature map corresponding to the to-be-processed image according to the image feature map; perform upsampling feature extraction on the image feature map based on the style feature map to obtain a target image; wherein, the image quality of the target image is higher than the image quality of the to-be-processed image.
[0038] Optionally, when the image processing module performs downsampling feature extraction on the to-be-processed image to obtain an image feature map corresponding to the to-be-processed image, it is specifically configured to:
[0039] Perform downsampling feature extraction on the to-be-processed image to obtain at least one scale of image feature map corresponding to the to-be-processed image;
[0040] When the image processing module determines a style feature map corresponding to the to-be-processed image according to the image feature map, it is specifically configured to:
[0041] For each scale, determine the style feature map corresponding to the scale according to the image feature map corresponding to the scale;
[0042] When the image processing module performs upsampling feature extraction on the image feature map based on the style feature map to obtain a target image, it is specifically configured to:
[0043] Performing at least one upsampling feature extraction on the target feature map through at least one upsampling feature extraction module to obtain the target image, where in each upsampling feature extraction process, the style feature map of the corresponding scale is fused into the feature map of the corresponding scale obtained by each upsampling feature extraction, the target feature map is the image feature map of the smallest scale among the image feature maps of at least one scale, and each upsampling feature extraction module corresponds to a style feature map of one scale.
[0044] Optionally, the image feature maps of at least one scale include image feature maps of at least two scales, and the at least one upsampling feature extraction module includes at least two cascaded upsampling feature extraction modules;
[0045] When the image processing module performs at least one upsampling feature extraction on the target feature map through at least one upsampling feature extraction module to obtain the target image, it is specifically used for:
[0046] For each upsampling feature extraction module, using the style feature map corresponding to this upsampling feature extraction module as an adjustment weight to adjust the network parameters of this upsampling feature extraction module, performing upsampling processing on the input feature map corresponding to this upsampling feature extraction module through the adjusted network parameters to obtain the output feature map of this upsampling feature extraction module, and using the last output feature map as the target image;
[0047] Among them, the input of the first upsampling feature extraction module is the target feature map, and the input of the modules other than the first upsampling feature extraction module is the output feature map of the previous upsampling feature extraction module of this upsampling feature extraction module.
[0048] Optionally, the image processing model includes an image encoding network and an image generation network. The image encoding network includes a feature extraction module and a feature mapping module. When the image processing module performs the following processing on the image to be processed by calling the trained image processing model to obtain the target image corresponding to the image to be processed, it is specifically used for:
[0049] Inputting the image to be processed into the feature extraction module, and performing downsampling feature extraction on the image to be processed through the feature extraction module to obtain at least one scale of image feature maps corresponding to the image to be processed;
[0050] Inputting the image feature maps of each scale into the feature mapping module respectively to obtain the style feature maps corresponding to each scale;
[0051] Based on the style feature maps corresponding to each scale and the target feature map, performing at least one upsampling feature extraction on the target feature map through the image generation network to obtain the target image corresponding to the image to be processed.
[0052] Optionally, the image processing model is obtained by training the model training module in the following manner:
[0053] Obtain first training data and an initial neural network model. The training data includes multiple image pairs, each image pair including a first sample image and a second sample image corresponding to the same image content. The image quality of the second sample image in each image pair is higher than that of the first sample image. The initial neural network model includes an initial image processing model and an initial discriminator;
[0054] Train the initial neural network model based on the first training data until the value of the total loss function corresponding to the model meets the training end condition, obtaining the trained neural network model, and use the trained image processing model as the image processing model;
[0055] Wherein, the input of the initial image processing model is each first sample image, and the output is the predicted target image corresponding to each first sample image. The input of the initial discriminator is the predicted target image and the second sample image corresponding to each first sample image, and the output is the discrimination result of the predicted target image and the second sample image corresponding to each first sample image;
[0056] The total loss function includes a first loss function and a second loss function. The value of the first loss function characterizes the difference between the predicted target image and the second sample image corresponding to each first sample image, and the value of the second loss function characterizes the accuracy of the discrimination result of the predicted target image and the second sample image corresponding to each first sample image.
[0057] Optionally, the image generation network and the initial discriminator of the initial image processing model are pre-trained using second training data. When training the initial neural network model based on the first training data, the learning rate of the image encoding network of the initial image processing model is greater than the learning rates of the image generation network and the initial discriminator of the initial image processing model.
[0058] Optionally, when the model training module obtains the first training data, it specifically is used for:
[0059] Obtain each second sample image, perform image degradation processing on each second sample image respectively to obtain each processed image;
[0060] Use the processed image corresponding to each second sample image as the first sample image corresponding to the second sample image, and based on each second sample image and the corresponding processed image, obtain the first training data.
[0061] Optionally, the image degradation processing includes at least one of adding blur processing, reducing image resolution processing, adding noise processing, or image format compression processing; the first training data includes images processed by at least two image degradation processing methods.
[0062] According to another aspect of the embodiments of the present application, an electronic device is provided. The electronic device includes: a memory, a processor, and a computer program stored on the memory. It is characterized in that the processor executes the computer program to implement any one of the steps in the above image processing method.
[0063] According to still another aspect of the embodiments of the present application, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, any one of the steps in the above image processing method is implemented.
[0064] According to one aspect of the embodiments of the present application, a computer program product is provided. When the computer program is executed by a processor, any one of the steps in the above image processing method is implemented.
[0065] The beneficial effects of the solution provided by the embodiments of the present application are as follows:
[0066] In the embodiments of the present application, when processing a to-be-processed image based on an image processing model, after obtaining the image feature map of the to-be-processed image, a style feature map corresponding to the to-be-processed image is obtained based on the image feature map. When obtaining a high-quality target image by performing upsampling processing on the image feature map, the style feature map is incorporated into the upsampling process, so that the obtained target image can better conform to the original style of the to-be-processed image. While the quality of the obtained target image is improved, it still has a more realistic style feature, thus better meeting the actual application requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the accompanying drawings required for description in the embodiments of the present application.
[0068] Figure 1 It is a schematic diagram of the system architecture of an image processing system applicable to the image processing method provided by the embodiments of the present application;
[0069] Figure 2a It is a schematic diagram of a to-be-processed image provided by the embodiments of the present application;
[0070] Figure 2b It is a schematic diagram of a target image provided by the embodiments of the present application;
[0071] Figure 3 It is a schematic diagram of the flow of an image processing method provided by the embodiments of the present application;
[0072] Figure 4 It is a schematic diagram of a GAN network provided by the embodiments of the present application;
[0073] Figure 5 Schematic diagram of a StyleGAN v2 network provided by an embodiment of the present application;
[0074] Figure 6 Schematic structural diagram of an image processing model provided by an embodiment of the present application;
[0075] Figure 7 Schematic structural diagram of an image processing apparatus provided by an embodiment of the present application;
[0076] Figure 8 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0077] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the implementation manners described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.
[0078] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components and / or their combinations supported by the art of the present technology. It should be understood that when we say that an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by the term, for example, "A and / or B" or "at least one of A or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".
[0079] To make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0080] Optionally, the image processing method provided in the embodiments of the present application can be implemented based on artificial intelligence (AI) technology. For example, after obtaining the image to be processed, the image to be processed can be processed based on a trained image processing model to obtain a target image with higher quality. Among them, artificial intelligence technology is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0081] Among them, the method provided in the embodiments of the present application can be specifically implemented based on computer vision technology (CV) in artificial intelligence technology. Among them, computer vision is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition, detection, and measurement, and further perform graphics processing to make the computer-processed image more suitable for human eyes to observe or be transmitted to an instrument for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0082] The image processing model involved in the embodiments of the present application can be obtained through machine learning training. Among them, machine learning (ML) is the study of how computers simulate or implement human learning behaviors to acquire new knowledge or skills, and reorganize the existing knowledge structure to continuously improve its own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0083] Optionally, the data processing involved in the embodiments of the present application can be implemented based on cloud technology. For example, for the data calculation involved in the training process of the model, cloud computing can be used. Cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" seem to be infinitely expandable to users, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage.
[0084] Next, through the description of several exemplary embodiments, the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described. It should be noted that the following embodiments can refer to, draw on, or combine with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.
[0085] It should be noted that the image to be processed in the embodiments of the present application can be any image that needs to be processed. That is to say, the specific form of the target object in the image to be processed is not limited in the embodiments of the present application. Optionally, the image to be processed can be an item image, a person image, an image of a certain part of a person, such as a face image, or a landscape image.
[0086] In an alternative embodiment of the present application, for relevant data such as user information (such as the user's face image), when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. That is to say, if the embodiments of the present application involve data related to users, these data need to be obtained under the condition of user authorization and consent and compliance with the relevant laws, regulations, and standards of the country and region.
[0087] The image processing method provided by this application can be executed by any electronic device, including but not limited to user terminals or servers. Optionally, the image processing method provided by the embodiments of this application can be implemented as an independent application program (such as image processing software) or as a functional module / plug-in of an existing application program. By applying the image processing method provided by the embodiments of this application to the application program, when the user needs to edit and process an image, the user terminal can implement the repair of the image to be processed by executing the computer program corresponding to the functional module / plug-in to obtain the target image. Alternatively, after the user terminal obtains the image to be processed selected by the user through the client of the application program, it can send the image to the server of the application program. The server can implement the repair of the image to be processed by executing the computer program corresponding to the functional module / plug-in and return the repaired target image to the user terminal.
[0088] In the embodiments of this application, the specific type of the above user terminal is not limited in the embodiments of this application, and can include but not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, wearable electronic devices, AR / VR devices, etc. The above server can be a cloud server or a physical server, and can be a single server or a server cluster.
[0089] The following describes the image processing method provided by this application in combination with the accompanying drawings and optional implementation manners in the embodiments of this application.
[0090] Figure 1 FIG. is a schematic diagram of the system architecture of an image processing system applicable to the image processing method provided by the embodiments of this application. As Figure 1 shown in the figure, the system may include a terminal device 10 (i.e., a user terminal) and a server 20. The terminal device 10 and the server 20 perform data transmission through a network. A trained image processing model can be deployed in the server 20. After the terminal device 10 obtains the image to be processed, it sends the image to be processed to the server 20. Among them, the image to be processed is an image with low image quality, such as the blurred image shown in Figure 2a . Further, after the server 20 receives the image to be processed, it can call the trained image processing model to process the image to be processed to obtain the target image corresponding to the image to be processed. At this time, the obtained target image (such as Figure 2b shown) has improved image quality compared to the image to be processed, is more in line with the original style of the image to be processed, and has more realistic style characteristics. Further, the server 20 returns the target image to the terminal device 10, and the terminal device 10 displays the target image to the user.
[0091] The following is based on Figure 1Taking the server as the execution entity of the method, an optional implementation manner of the image processing method provided in this application will be described for the image processing system shown in the following. As Figure 3 shown, an image processing method is provided in an embodiment of this application, and the method may include:
[0092] Step S201, obtaining a to-be-processed image of a target object.
[0093] Wherein, the target object may refer to any object such as a person, an animal, a plant, a building, etc., and the embodiment of this application does not limit the type of the target object. The to-be-processed image may be any image that needs to improve the quality, may be an image obtained in real time by an image acquisition device, or may be a frame image extracted from a video, and the embodiment of this application does not limit it. Optionally, the to-be-processed image may be a face image. For the convenience of description, in some example descriptions below, the face image will be used as an example for illustration.
[0094] Step S202, processing the to-be-processed image through a trained image processing model for steps S301 to S303 to obtain a target image corresponding to the to-be-processed image, wherein the image quality of the target image is higher than that of the to-be-processed image.
[0095] Wherein, the image quality may be characterized by image metrics. For example, it may be characterized by one or more metrics such as the resolution of the image or the clarity of the image. Optionally, the image quality of the target image being higher than that of the to-be-processed image may mean that at least one image metric of the target image is better than that of the to-be-processed image. For example, it may be that the clarity of the target image is higher than the image quality of the to-be-processed image.
[0096] Steps S301 to S303 of processing the to-be-processed image through a trained image processing model will be described below.
[0097] Step S301, performing downsampling feature extraction on the to-be-processed image to obtain an image feature map corresponding to the to-be-processed image.
[0098] Step S302, determining a style feature map corresponding to the to-be-processed image according to the image feature map;
[0099] Optionally, perform downsampling feature extraction on the obtained image to be processed. At this time, an image feature map corresponding to the image to be processed can be obtained, and based on the obtained image feature map, a style feature map corresponding to the image to be processed can be determined. Among them, the style in the embodiments of the present application can be understood as the factors or information that affect the visual effect of the image to be processed. For example, when taking an image of a person, if different shooting parameters are used by the shooting device or the shooting is performed in different shooting environments, the clarity of some of the obtained images may be higher and that of some may be lower. At this time, the image feature maps of the images are different, and further, the obtained image feature maps can be extracted into different style feature maps to reflect the degree of blurriness of the image. Among them, the image to be processed can be a face image. When the image to be processed is a face image as an example, the style of the image to be processed can be understood as the information that can affect the visual effects such as the actual color of the face in the image. Correspondingly, the style feature map can be understood as the feature vector of the style influencing factors obtained based on the image feature map of the image to be processed, that is, the mathematical representation of the influencing factors.
[0100] Step S303, perform upsampling feature extraction on the image feature map based on the style feature map to obtain the target image.
[0101] Optionally, after obtaining the image feature map corresponding to the image to be processed, upsampling feature extraction can be performed on the image feature map. In order to make the finally obtained target image more conform to the original style of the image to be processed and have more realistic style characteristics, at this time, upsampling feature extraction can be performed on the image feature map according to the determined style feature map to obtain the target image. The size of the target image can be the same as or different from the size of the image to be processed, which is not limited in the embodiments of the present application.
[0102] In the embodiments of the present application, when processing the image to be processed based on the image processing model, after obtaining the image feature map of the image to be processed, the style feature map corresponding to the image to be processed will be obtained based on this image feature map. When obtaining a high-quality target image by performing upsampling processing on the image feature map, the style feature map will be incorporated into the upsampling process, so that the obtained target image can more conform to the original style of the image to be processed. While the quality of the obtained target image is improved, it still has more realistic style characteristics, so as to better meet the actual application requirements.
[0103] A possible implementation manner is provided in the embodiments of the present application. Performing downsampling feature extraction on the image to be processed to obtain the image feature map corresponding to the image to be processed includes:
[0104] Perform downsampling feature extraction on the image to be processed to obtain at least one scale of image feature map corresponding to the image to be processed;
[0105] Determine the style feature map corresponding to the image to be processed according to the image feature map, including:
[0106] For each scale, determine the style feature map corresponding to this scale according to the image feature map of the corresponding scale;
[0107] Perform upsampling feature extraction on the image feature map based on the style feature map to obtain the target image, including:
[0108] Perform at least one upsampling feature extraction on the target feature map through at least one upsampling feature extraction module to obtain the target image. Wherein, in each upsampling feature extraction process, fuse the style feature map of the corresponding scale into the feature map of the corresponding scale obtained by each upsampling feature extraction. The target feature map is the image feature map of the smallest scale among at least one scale of image feature maps, and each upsampling feature extraction module corresponds to a style feature map of one scale.
[0109] Optionally, when performing downsampling feature extraction on the image to be processed, at least one downsampling feature extraction can be performed on the image to be processed. At this time, at least one scale of image feature maps corresponding to the image to be processed will be obtained. When obtaining the image feature map corresponding to one scale, the style feature map corresponding to this scale can be determined according to the image feature map corresponding to this scale. For example, when obtaining the image feature map corresponding to scale A, the style feature map corresponding to scale A can be determined according to the image feature map corresponding to scale A.
[0110] Furthermore, the image feature map of the smallest scale among at least one scale of image feature maps obtained during downsampling feature extraction can be used as the target feature map, and at least one upsampling feature extraction is performed on the target feature map through at least one upsampling feature extraction module. Each upsampling feature extraction module corresponds to a style feature map of one scale. The corresponding scale can refer to the scale of the input feature map of this upsampling feature extraction module, or the scale of the output feature map of this upsampling feature extraction module. The embodiments of the present application do not limit this. For example, assume that an upsampling feature extraction module can extract an image feature map of 4 (size of the feature map) * 4 to an image feature map of 8 * 8. At this time, the scale corresponding to this upsampling feature extraction module can be 4 * 4 or 8 * 8. However, it should be noted that the meaning of the scale corresponding to each upsampling feature extraction module should be the same. For example, when the scale corresponding to the upsampling feature extraction module refers to the scale of the input feature map of this upsampling feature extraction module, the scales corresponding to all upsampling feature extraction modules should also be the scale of the input feature map.
[0111] Further, in the process of performing at least one upsampling feature extraction on the target feature map through at least one upsampling feature extraction module, in each upsampling feature extraction process, the style feature map of the corresponding scale can be fused into the feature map of the corresponding scale obtained by each upsampling feature extraction. For example, assuming that when performing one upsampling feature extraction through an upsampling feature module, the scale corresponding to this upsampling feature module is A, at this time, the style feature map corresponding to scale A can be fused into the image feature map obtained through this upsampling feature module.
[0112] In an embodiment of the present application, a possible implementation manner is provided. The image feature maps of at least one scale include image feature maps of at least two scales, and at least one upsampling feature extraction module includes at least two cascaded upsampling feature extraction modules;
[0113] Performing at least one upsampling feature extraction on the target feature map through at least one upsampling feature extraction module to obtain a target image, including:
[0114] For each upsampling feature extraction module, using the style feature map corresponding to this upsampling feature extraction module as an adjustment weight to adjust the network parameters of this upsampling feature extraction module, performing upsampling processing on the input feature map corresponding to this upsampling feature extraction module through the adjusted network parameters to obtain the output feature map of this upsampling feature extraction module, and using the last output feature map as the target image;
[0115] Among them, the input of the first upsampling feature extraction module is the target feature map, and the input of the modules other than the first upsampling feature extraction module is the output feature map of the previous upsampling feature extraction module of this upsampling feature extraction module.
[0116] Optionally, in the embodiments of the present application, the image to be processed can be subjected to downsampling feature extraction at least twice. At this time, at least two scales of image feature maps can be obtained, and at least two cascaded upsampling feature extraction modules can be included for upsampling feature extraction. Further, when performing at least one upsampling feature extraction on the target feature map through at least one upsampling feature extraction module to obtain the target image, the target feature map can be input into the first upsampling feature extraction module. For this first upsampling feature extraction module, the style feature map corresponding to this module can be used as the adjustment weight to adjust the network parameters of this upsampling feature extraction module. Then, the input target feature map is subjected to upsampling processing through the adjusted network parameters to obtain the output feature map of this module, and the output feature map of this module is input into the second upsampling feature extraction module. At the same time, based on the same principle, the style feature map corresponding to the second upsampling feature extraction module is used as the adjustment weight to adjust the network parameters of the second upsampling feature extraction module. Then, the input output feature map is subjected to upsampling processing through the adjusted network parameters to obtain the output feature map of this module, until the output feature map of the last upsampling feature extraction module is obtained, and this output feature map is used as the target feature map.
[0117] In one example, it is assumed that the image to be processed is subjected to downsampling feature extraction twice, and the 16*16 image feature map is downsampled to the 8*8 image feature map and the 4*4 image feature map respectively, and a total of 2 cascaded upsampling feature extraction modules are included, and the scale corresponding to each sampling feature extraction module is the size of the input feature map; further, the style feature map corresponding to the 8*8 scale can be determined based on the 8*8 image feature map, and the style feature map corresponding to the 4*4 scale can be determined based on the 4*4 image feature map. Correspondingly, the 4*4 image feature map can be used as the target feature map and input into the first upsampling feature extraction module, and the style feature map corresponding to the 4*4 scale is used as the adjustment weight to adjust the network parameters of this upsampling feature extraction module. The input 4*4 image feature map is subjected to upsampling processing through the adjusted network parameters to obtain the 8*8 image feature map, and then the 8*8 image feature map is input into the second upsampling feature extraction module, and the style feature map corresponding to the 8*8 scale is used as the adjustment weight to adjust the network parameters of this upsampling feature extraction module. The 8*8 image feature map is subjected to upsampling processing through the adjusted network parameters to obtain the 16*16 image feature map, and this 16*16 image feature map is used as the target image.
[0118] As an alternative embodiment, in the embodiments of the present application, after performing downsampling feature extraction on the image to be processed at least twice to obtain image feature maps of at least two scales, the target feature map can be used as the initial feature map to be adjusted, and then the style feature map corresponding to the scale of the initial feature map to be adjusted can be used as the adjustment weight to adjust the initial feature map to be adjusted, obtaining the adjusted feature map; further, upsampling feature extraction can be performed on the adjusted feature map to obtain a new feature map. At this time, the new feature map to be processed can be used as the new feature map to be processed, and the style feature map corresponding to the scale of the new feature map to be processed can be used as the adjustment weight to adjust the new feature map to be processed, obtaining the adjusted feature map, and then upsampling feature extraction is performed on the adjusted feature map again to obtain the extracted feature map. Then, the extracted feature map can be used as the new feature map to be processed, and the process of using the style feature maps corresponding to each scale as the adjustment weights to adjust the feature maps of the corresponding scales obtained during the upsampling feature extraction process to obtain the target image, and based on performing upsampling feature extraction on the adjusted feature map and using the extracted feature map as the new feature map to be processed is repeated until the number of times of upsampling feature extraction is equal to the number of style feature maps, and the feature map obtained by the last upsampling feature extraction is used as the target image.
[0119] In an example, it is assumed that downsampling feature extraction is performed on the image to be processed 3 times, obtaining feature maps of three scales, namely the image feature map of 16*16 (the size of the feature map), the image feature map of 8*8, and the image feature map of 4*4; further, the style feature map corresponding to 16*16 can be determined based on the 16*16 feature map, the style feature map corresponding to 8*8 can be determined based on the 8*8 feature map, and the style feature map corresponding to 4*4 can be determined based on the 4*4 feature map. Correspondingly, the image feature map of 4*4 can be used as the target feature map, and then the style feature map corresponding to 4*4 can be used as the adjustment weight to adjust the image feature map of 4*4, obtaining the adjusted feature map, and upsampling feature extraction is performed on the adjusted feature map to obtain the image feature map of 8*8; further, the style feature map corresponding to 8*8 can be used as the adjustment weight to adjust the image feature map of 8*8, obtaining the adjusted feature map, and upsampling feature extraction is performed on the adjusted feature map to obtain the image feature map of 16*16. The style feature map corresponding to 16*16 is used as the adjustment weight to adjust the image feature map of 16*16, obtaining the adjusted feature map, and upsampling feature extraction is performed on the adjusted feature map to obtain the image feature map of 32*32. At this time, the image feature map of 32*32 is the target image.
[0120] In an embodiment of the present application, a possible implementation is provided. The image processing model includes an image encoding network and an image generation network. The image encoding network includes a feature extraction module and a feature mapping module. By calling the trained image processing model, the following processing is performed on the image to be processed to obtain the target image corresponding to the image to be processed, including:
[0121] Input the image to be processed into the feature extraction module, and perform downsampling feature extraction on the image to be processed through the feature extraction module to obtain at least one scale of image feature maps corresponding to the image to be processed;
[0122] Input the image feature maps of each scale into the feature mapping module respectively to obtain the style feature maps corresponding to each scale;
[0123] Based on the style feature maps corresponding to each scale and the target feature map, perform at least one upsampling feature extraction on the target feature map through the image generation network to obtain the target image corresponding to the image to be processed.
[0124] Optionally, the method provided in the embodiment of the present application can be implemented by using an image processing model. The image processing model includes an image encoding network and an image generation network, and the image encoding network can include a feature extraction module and a feature mapping module. The image generation network includes at least one upsampling feature extraction module for performing at least one upsampling feature extraction on the input feature map. Correspondingly, when obtaining the target image corresponding to the image to be processed based on this image processing model, the image to be processed can be input into the feature extraction module for at least one downsampling feature extraction to obtain at least one scale of image feature maps corresponding to the image to be processed, and then the image feature maps of each scale can be input into the feature mapping module respectively to obtain the style feature maps corresponding to each scale. At this time, the style feature maps corresponding to each scale and the target feature map can be input into the image generation network. The at least one upsampling feature extraction module included in the image generation network can perform at least one upsampling feature extraction on the target feature map. When performing upsampling feature extraction at each scale, the style feature map corresponding to the corresponding scale is used to adjust the network weights (such as convolutional network weights) during upsampling feature extraction. At this time, the target feature map is obtained as the target image after at least one upsampling feature extraction. Correspondingly, when the style feature maps are different, the image generation network can enhance the image quality of the input image to different degrees according to different style feature maps.
[0125] In an embodiment of the present application, a possible implementation is provided. The image processing model is trained in the following manner:
[0126] Obtain the first training data and the initial neural network model. The training data includes multiple image pairs, each image pair including a first sample image and a second sample image corresponding to the same image content. The image quality of the second sample image in each image pair is higher than that of the first sample image. The initial neural network model includes an initial image processing model and an initial discriminator;
[0127] Train the initial neural network model based on the first training data until the value of the total loss function corresponding to the model meets the training end condition, obtaining the trained neural network model, and use the trained image processing model as the image processing model;
[0128] Among them, the input of the initial image processing model is each first sample image, and the output is the predicted target image corresponding to each first sample image. The input of the initial discriminator is the predicted target image corresponding to each first sample image and the second sample image, and the output is the discrimination result of the predicted target image corresponding to each first sample image and the second sample image;
[0129] The total loss function includes a first loss function and a second loss function. The value of the first loss function characterizes the difference between the predicted target image corresponding to each first sample image and the second sample image, and the value of the second loss function characterizes the accuracy of the discrimination result of the predicted target image corresponding to each first sample image and the second sample image.
[0130] Optionally, the above-mentioned image processing model can be obtained by training the initial neural network model based on the acquired first training data. Among them, the first training data includes multiple image pairs, each image pair including a first sample image and a second sample image corresponding to the first sample image. The image content of the first sample image and the second sample image is the same, and the image quality of the second sample image is higher than that of the first sample image, while the initial neural network model includes an initial image processing model and an initial discriminator.
[0131] Further, when training the initial neural network model, each first sample image can be input into the initial neural network model to obtain the predicted target image corresponding to each first sample image. At this time, the first loss function can be determined based on the predicted target image and the first sample image corresponding to each first sample image. The value of the first loss function characterizes the difference between the predicted target image and the first sample image corresponding to each first sample image. When the value of the first loss function is smaller, it indicates that the difference between the sample image and the corresponding predicted target image is smaller. Then, the predicted target image corresponding to each first sample image and the second sample image can be respectively input into the initial discriminator to obtain the discrimination results of each predicted target image and the second sample image, and the second loss function can be determined based on the discrimination results of the predicted target image corresponding to each first sample image and the second sample image. The value of the second loss function characterizes the accuracy of the discrimination results of the predicted target image corresponding to each first sample image and the second sample image. The smaller the value of the second loss function, the more accurate the discrimination results of the predicted target image and the second sample image. The total loss function corresponding to the initial neural network model includes the first loss function and the second loss function. At this time, it can be determined whether the value of the total loss function satisfies the training end condition based on the values of the first loss function and the second loss function. If not, the network parameters of the initial image processing model can be adjusted to obtain an adjusted image processing model, and then the adjusted image processing model can be continuously trained based on the first training data until the value of the corresponding total loss function satisfies the training condition. At this time, the discriminator obtained can determine the judged target object as true, and the generated target image can deceive the discriminator. Among them, the training condition can be set according to actual needs, and this application embodiment does not limit this, for example, it can be that the value of the total loss function is less than a set threshold. In the embodiment of this application, the value of the first loss function is the generation loss, and the value of the second loss function can be understood as the recognition loss. The training objective of the initial neural network model is to make the generation loss and the recognition loss as small as possible, that is, in the embodiment of this application, the training of the model is constrained by the loss functions from two different dimensions, so as to better improve the performance of the image processing model.
[0132] Optionally, in the embodiment of this application, the total loss function can be expressed as:
[0133] L = L GAN + L LPIPS
[0134] Wherein, L represents the total loss function, L LPIPS represents the first loss function, and L GAN represents the second loss function. Regarding the specific loss function adopted for the first loss function, this application does not limit it. Optionally, the first loss function can be expressed using the L1 loss.
[0135] Optionally, L corresponding to an image pair LPIPS can be expressed as:
[0136] L LPIPS = |G(input) - GT| 1
[0137] where G(input) represents the predicted target image corresponding to the first sample image, GT represents the second sample image corresponding to the first sample image, and |G(input) - GT| 1 represents the L1 loss between G(input) and GT.
[0138] Optionally, L GAN can be expressed as:
[0139]
[0140] where D(GT) represents the discrimination result of the second sample image, G(input) represents the discrimination result of the first sample image, E[log(D(GT))] represents the expectation of the discrimination result of the second sample image, E[log(1 - D(G(input)))] represents the expectation of the discrimination result of the first sample image, and the loss corresponding to the image pair is the sum of the differences in pixel values of all corresponding pixels in these two images.
[0141] In an embodiment of the present application, a possible implementation manner is provided. The image generation network and the initial discriminator of the initial image processing model are pre-trained using the second training data. When training the initial neural network model based on the first training data, the learning rate of the image encoding network of the initial image processing model is greater than the learning rates of the image generation network and the initial discriminator of the initial image processing model.
[0142] Optionally, in practical applications, the structures of the various parts of the image processing model are not limited in this embodiment of the present application and can be selected according to actual application requirements. Optionally, as long as it can generate a high-definition portrait and put it into the portrait restoration framework in a similar encoding manner, it can be used as the image generation network of the initial image processing model described above. For example, Figure 4As shown, the GAN (Generative Adversarial Networks) network can generate highly realistic and high-definition images. It can map a simple distribution (such as Gaussian noise) to a complex distribution similar to the real image distribution (i.e., the generated image distribution). At this time, the generator network in the GAN network can be inserted into the initial image processing model as the generator network, and the discriminator of the GAN network can be used as the initial discriminator. In recent years, the image quality generated by the GAN network has been continuously improved. For example, PGAN (Progressive GAN (Progressive Generative Adversarial Networks, an improved GAN network)), BigGAN (a type of GAN network), StyleGAN (a type of GAN network), and StyleGAN v2 (Style GAN version 2, an improved StyleGAN v1 network) can all generate highly realistic images. At this time, the generator networks of PGAN, ProgressiveGAN, BigGAN, StyleGAN v1, and StyleGAN v2 can all be used as the image generator network of the above-mentioned initial image processing model.
[0143] In one example, assume that the generator network of the StyleGAN v2 network is used as the image generator network of the initial image processing model. At this time, the training data for training the StyleGAN v2 network is the second training data. Among them, when training the StyleGAN v2 network, the schematic diagram of the network structure to be trained can be as Figure 5 shown. This network structure can include the mapping network and generator network of the StyleGAN v2 network, as well as the discriminator that needs to be pre-trained. In the StyleGAN v2 network, the random noise z can be mapped to the latent variable w (a vector) through the mapping network. The initial input of the generator network is a 4*4 feature map. After passing through each upsampling module, the resolution becomes 2 times. After passing through a total of 8 upsampling modules, the resolution is finally increased to 1024*1024. In each upsampling module, the latent variable w modulates the input feature map through the AdaIN (Adaptive Instance Normalization) method, thereby affecting the appearance of the final generated image. Among them, the process of generating the final image from the noise z can be expressed as:
[0144] w = Mapping(z)
[0145] Image = Generator(w)
[0146] Among them, Mapping(z) represents mapping the random noise z to the latent variable w, and Image represents the generated image.
[0147] After training based on the second training data, the image generation network and discriminator that meet the training end condition can be used as the image generation network and initial discriminator of the initial neural network model, and the initial neural network model can be trained based on the first training data. Among them, since the image generation network and the initial discriminator have already been pre-trained, in order to improve the training efficiency, different learning rates can be set for the image encoding network of the initial image processing model, the image generation network of the initial image processing model, and the initial discriminator for training. For example, the learning rate of the image encoding network of the initial image processing model can be set to be greater than the learning rates of the image generation network and the initial discriminator of the initial image processing model (for example, the ratio of the learning rate of the image encoding network of the initial image processing model, the learning rate of the image generation network of the initial image processing model, and the learning rate of the initial discriminator is set to 100:10:1), so as to reduce the amount of data processing during training and improve the training efficiency.
[0148] In an embodiment of the present application, a possible implementation manner is provided. Obtaining the first training data includes:
[0149] Obtaining each second sample image, and respectively performing image degradation processing on each second sample image to obtain each processed image;
[0150] Taking the processed image corresponding to each second sample image as the first sample image corresponding to the second sample image, and based on each second sample image and the corresponding processed image, obtaining the first training data.
[0151] Furthermore, each second sample image can be obtained, and image degradation processing can be respectively performed on each second sample image to obtain each processed image. At this time, the processed image corresponding to each second sample image can be used as the first sample image corresponding to the second sample image, so as to obtain the first training data. It can be understood that the first sample images in the first training data include images obtained by using at least two image degradation processing methods, and the images obtained by using at least two image degradation processing methods can refer to images obtained by performing different image degradation processing methods on one second sample image, or can be images obtained by performing different image degradation processing methods on different second sample images.
[0152] In an embodiment of the present application, a possible implementation manner is provided. The image degradation processing includes at least one of adding blur processing, reducing image resolution processing, adding noise processing, or image format compression processing; the first training data includes images obtained by using at least two image degradation processing methods.
[0153] Optionally, the image degradation process refers to the process of reducing the image quality, which may specifically include at least one of adding a blur process, reducing the image resolution process, adding a noise process, or an image format compression process. Among them, adding a blur process includes randomly adding Gaussian blur, motion blur, etc. Among them, the standard deviation of Gaussian blur is randomly selected within a certain range, and the motion blur includes 38 kinds of custom blurs; reducing the image resolution process refers to reducing the image resolution, and the image quality can be reduced by means of downsampling. For which specific downsampling method or methods to sample, this application does not limit. For example, biliner (bilinear) downsampling, bicubic (bicubic downsampling), area (area downsampling), etc. can be randomly selected to reduce the image resolution; adding a noise process includes Gaussian noise, Poisson noise, etc., and the noise intensity can be randomly selected within a certain range; the image format compression process may refer to JPEG (an image file format) compression, which simulates the decrease in image quality during the process of saving a picture, and the compression ratio can be randomly selected between 5% and 50%.
[0154] Optionally, the method provided in the embodiments of the present application can be applied to application scenarios such as old photo restoration and picture portrait clarification. For better understanding, the following will be described in detail in combination with specific application scenarios. In this example, the image to be processed is a face image to be restored (i.e., a low-resolution image with a size of 32*32), and the restored face image (i.e., a high-resolution image with a size of 32*32) can be obtained based on the image processing model. Among them, as Figure 6 shown, the image processing model includes an image encoding network and an image generation network. The image encoding network includes a feature extraction module and a feature mapping module. The feature extraction module is composed of a cascade of multiple downsampling modules (illustrated by three in this example), and the feature mapping module refers to a fully connected layer. The image generation network is composed of a cascade of multiple upsampling modules (illustrated by three in this example).
[0155] Further, the low-resolution image can be input into the feature extraction module. The three downsampling feature extraction modules in the feature extraction module perform three downsamplings on the low-resolution image to obtain three scales of image feature maps, namely, the 16*16 image feature, the 8*8 image feature map, and the 4*4 image feature map. Then, the 16*16 image feature, the 8*8 image feature map, and the 4*4 image feature map are respectively input into the fully connected layer to obtain the style feature maps corresponding to each scale. Further, the 4*4 image feature map (i.e., the target feature map) is input into the first upsampling feature extraction module in the image generation network, and based on the style feature map corresponding to the 4*4 scale as the adjustment weight, the network parameters of this upsampling feature extraction module are adjusted. The input 4*4 image feature map is upsampled through the adjusted network parameters to obtain an 8*8 image feature map. Then, the 8*8 image feature map is input into the second upsampling feature extraction module, and based on the style feature map corresponding to the 8*8 scale as the adjustment weight, the network parameters of this upsampling feature extraction module are adjusted. The 8*8 image feature map is upsampled through the adjusted network parameters to obtain a 16*16 image feature map, and is input into the third upsampling feature extraction module. Based on the style feature map corresponding to the 16*16 scale as the adjustment weight, the network parameters of this upsampling feature extraction module are adjusted. The 16*16 image feature map is upsampled through the adjusted network parameters to obtain a 32*32 image feature map. At this time, this 32*32 image feature map is the restored face image (i.e., the high-resolution image).
[0156] An embodiment of the present application provides a device, as Figure 7 shown. The image processing device 70 may include: an image acquisition module 701 and an image processing module 702. Among them,
[0157] The image acquisition module 701 is configured to acquire a to-be-processed image of a target object;
[0158] The image processing module 702 is configured to perform downsampling feature extraction on the to-be-processed image by calling a trained image processing model to obtain an image feature map corresponding to the to-be-processed image, and determine a style feature map corresponding to the to-be-processed image according to the image feature map; perform upsampling feature extraction on the image feature map based on the style feature map to obtain a target image; wherein, the image quality of the target image is higher than that of the to-be-processed image.
[0159] Optionally, when the image processing module performs downsampling feature extraction on the to-be-processed image to obtain an image feature map corresponding to the to-be-processed image, it is specifically configured to:
[0160] Perform downsampling feature extraction on the to-be-processed image to obtain at least one scale of image feature map corresponding to the to-be-processed image;
[0161] When the image processing module determines the style feature map corresponding to the image to be processed according to the image feature map, it is specifically used for:
[0162] For each scale, determine the style feature map corresponding to that scale according to the image feature map of the corresponding scale;
[0163] When the image processing module performs upsampling feature extraction on the image feature map based on the style feature map to obtain the target image, it is specifically used for:
[0164] Perform at least one upsampling feature extraction on the target feature map through at least one upsampling feature extraction module to obtain the target image. Among them, in each upsampling feature extraction process, fuse the style feature map of the corresponding scale into the feature map of the corresponding scale obtained by each upsampling feature extraction. The target feature map is the image feature map of the smallest scale among at least one scale of image feature maps, and each upsampling feature extraction module corresponds to a style feature map of one scale.
[0165] Optionally, the at least one scale of image feature maps includes at least two scales of image feature maps, and the at least one upsampling feature extraction module includes at least two cascaded upsampling feature extraction modules;
[0166] When the image processing module performs at least one upsampling feature extraction on the target feature map through at least one upsampling feature extraction module to obtain the target image, it is specifically used for:
[0167] For each upsampling feature extraction module, use the style feature map corresponding to the upsampling feature extraction module as the adjustment weight to adjust the network parameters of the upsampling feature extraction module. Perform upsampling processing on the input feature map corresponding to the upsampling feature extraction module through the adjusted network parameters to obtain the output feature map of the upsampling feature extraction module, and use the last output feature map as the target image;
[0168] Among them, the input of the first upsampling feature extraction module is the target feature map, and the input of the modules other than the first upsampling feature extraction module is the output feature map of the previous upsampling feature extraction module of the upsampling feature extraction module.
[0169] Optionally, the image processing model includes an image encoding network and an image generation network. The image encoding network includes a feature extraction module and a feature mapping module. When the image processing module performs the following processing on the image to be processed by calling the trained image processing model to obtain the target image corresponding to the image to be processed, it is specifically used for:
[0170] Input the image to be processed into the feature extraction module, and perform downsampling feature extraction on the image to be processed through the feature extraction module to obtain at least one scale of image feature maps corresponding to the image to be processed;
[0171] Input the image feature maps of each scale into the feature mapping module respectively to obtain the style feature maps corresponding to each scale;
[0172] Based on the style feature maps corresponding to each scale and the target feature map, perform at least one upsampling feature extraction on the target feature map through the image generation network to obtain the target image corresponding to the image to be processed.
[0173] Optionally, the image processing model is obtained by the model training module through the following method:
[0174] Obtain the first training data and the initial neural network model. The training data includes multiple image pairs, each image pair includes a first sample image and a second sample image corresponding to the same image content, and the image quality of the second sample image in each image pair is higher than that of the first sample image. The initial neural network model includes an initial image processing model and an initial discriminator;
[0175] Train the initial neural network model based on the first training data until the value of the total loss function corresponding to the model meets the training end condition, and obtain the trained neural network model. Use the trained image processing model as the image processing model;
[0176] Among them, the input of the initial image processing model is each first sample image, and the output is the predicted target image corresponding to each first sample image. The input of the initial discriminator is the predicted target image corresponding to each first sample image and the second sample image, and the output is the discrimination result of the predicted target image corresponding to each first sample image and the second sample image;
[0177] The total loss function includes a first loss function and a second loss function. The value of the first loss function characterizes the difference between the predicted target image corresponding to each first sample image and the second sample image, and the value of the second loss function characterizes the accuracy of the discrimination result of the predicted target image corresponding to each first sample image and the second sample image.
[0178] Optionally, the image generation network and the initial discriminator of the initial image processing model are pre-trained using the second training data. When training the initial neural network model based on the first training data, the learning rate of the image encoding network of the initial image processing model is greater than the learning rates of the image generation network and the initial discriminator of the initial image processing model.
[0179] Optionally, when the model training module obtains the first training data, it is specifically used for:
[0180] Obtain each second sample image, perform image degradation processing on each second sample image respectively to obtain each processed image;
[0181] Use the processed image corresponding to each second sample image as the first sample image corresponding to the second sample image, and obtain first training data based on each second sample image and the corresponding processed image.
[0182] Optionally, the image degradation processing includes at least one of adding blurring processing, reducing image resolution processing, adding noise processing, or image format compression processing; the first training data includes images processed by at least two image degradation processing methods.
[0183] The device according to the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device according to the embodiments of the present application correspond to the steps in the methods according to the embodiments of the present application. For the detailed function description of each module of the device, reference can specifically be made to the description in the corresponding method shown above, and details are not described herein again.
[0184] An electronic device is provided in an embodiment of the present application, including a memory, a processor, and a computer program stored on the memory, and the processor executes the above computer program to implement the steps of the image processing method.
[0185] In an optional embodiment, an electronic device is provided, as Figure 8 shown Figure 8 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between this electronic device and other electronic devices, such as sending and / or receiving data, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiment of the present application.
[0186] The processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0187] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 only a thick line is shown herein, but it does not mean that there is only one bus or one type of bus.
[0188] The memory 4003 may be a ROM (Read Only Memory) or other type of static storage device that can store static information and instructions, a RAM (Random Access Memory) or other type of dynamic storage device that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.
[0189] The memory 4003 is used to store the computer program for implementing the embodiments of the present application, and is controlled by the processor 4001 to execute. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0190] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.
[0191] The embodiments of the present application further provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.
[0192] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the description and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than that shown in the drawings or described in words.
[0193] It should be understood that although the flowcharts in the embodiments of the present application indicate various operation steps by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated in this application, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.
[0194] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, other similar implementation means based on the technical idea of the present application also belong to the protection scope of the embodiments of the present application.
Claims
1. An image processing method, characterized in that, comprising: obtaining a to-be-processed image of a target object; performing the following processing on the to-be-processed image by invoking a trained image processing model to obtain a target image corresponding to the to-be-processed image, wherein the image quality of the target image is higher than that of the to-be-processed image: performing downsampling feature extraction on the to-be-processed image to obtain an image feature map corresponding to the to-be-processed image; determining a style feature map corresponding to the to-be-processed image according to the image feature map; wherein the style feature map characterizes factors or information affecting the visual effect of the to-be-processed image; based on the style feature map, performing upsampling feature extraction on the image feature map through an upsampling feature extraction module of the image processing model to obtain the target image, wherein the upsampling feature extraction comprises the following steps: taking the style feature map as an adjustment weight of the upsampling feature map extraction module, adjusting the network parameters of the upsampling feature extraction module, and performing upsampling processing on the image feature map corresponding to the to-be-processed image through the adjusted network parameters to obtain an output feature map of the upsampling feature extraction module.
2. The method according to claim 1, characterized in that, the performing downsampling feature extraction on the to-be-processed image to obtain an image feature map corresponding to the to-be-processed image comprises: performing downsampling feature extraction on the to-be-processed image to obtain at least one scale of image feature maps corresponding to the to-be-processed image; the determining a style feature map corresponding to the to-be-processed image according to the image feature map comprises: for each scale, determining a style feature map corresponding to the scale according to the image feature map of the corresponding scale; the performing upsampling feature extraction on the image feature map based on the style feature map to obtain the target image comprises: performing at least one upsampling feature extraction on a target feature map through at least one upsampling feature extraction module to obtain the target image, the target feature map being the image feature map of the smallest scale among the at least one scale of image feature maps, and each upsampling feature extraction module corresponding to a style feature map of one scale.
3. The method according to claim 2, characterized in that, the at least one scale of image feature maps comprises at least two scales of image feature maps, and the at least one upsampling feature extraction module comprises at least two cascaded upsampling feature extraction modules; the performing at least one upsampling feature extraction on a target feature map through at least one upsampling feature extraction module to obtain the target image comprises: for each upsampling feature extraction module, taking the style feature map corresponding to the upsampling feature extraction module as an adjustment weight, adjusting the network parameters of the upsampling feature extraction module, performing upsampling processing on the input feature map corresponding to the upsampling feature extraction module through the adjusted network parameters to obtain an output feature map of the upsampling feature extraction module, and taking the last output feature map as the target image; Among them, the input of the first upsampling feature extraction module is the target feature map, and the input of the modules other than the first upsampling feature extraction module is the output feature map of the previous upsampling feature extraction module of this upsampling feature extraction module.
4. The method according to claim 2 or 3, wherein, the image processing model includes an image encoding network and an image generation network, the image encoding network includes a feature extraction module and a feature mapping module, and the following processing is performed on the image to be processed by calling the trained image processing model to obtain the target image corresponding to the image to be processed, including: inputting the image to be processed into the feature extraction module, and performing downsampling feature extraction on the image to be processed through the feature extraction module to obtain at least one scale of image feature maps corresponding to the image to be processed; inputting the image feature maps of each scale into the feature mapping module respectively to obtain the style feature maps corresponding to each scale; based on the style feature maps corresponding to each scale and the target feature map, performing at least one upsampling feature extraction on the target feature map through the image generation network to obtain the target image corresponding to the image to be processed.
5. The method according to claim 4, wherein, the image processing model is obtained by training in the following manner: acquiring first training data and an initial neural network model, the first training data includes a plurality of image pairs, each image pair includes a first sample image and a second sample image corresponding to the same image content, and the image quality of the second sample image in each image pair is higher than that of the first sample image, and the initial neural network model includes an initial image processing model and an initial discriminator; training the initial neural network model based on the first training data until the value of the total loss function corresponding to the model meets the training end condition, obtaining the trained neural network model, and using the trained image processing model as the image processing model; wherein, the input of the initial image processing model is each of the first sample images, the output is the predicted target image corresponding to each of the first sample images, the input of the initial discriminator is the predicted target image corresponding to each of the first sample images and the second sample image, and the output is the discrimination result of the predicted target image corresponding to each of the first sample images and the second sample image; the total loss function includes a first loss function and a second loss function, the value of the first loss function characterizes the difference between the predicted target image corresponding to each of the first sample images and the second sample image, and the value of the second loss function characterizes the accuracy of the discrimination result of the predicted target image corresponding to each of the first sample images and the second sample image.
6. The method according to claim 5, wherein, The image generation network of the initial image processing model and the initial discriminator are pre-trained using second training data. When training the initial neural network model based on the first training data, the learning rate of the image encoding network of the initial image processing model is greater than the learning rates of the image generation network of the initial image processing model and the initial discriminator.
7. The method according to claim 5, wherein, the obtaining of the first training data includes: obtaining each second sample image, and performing image degradation processing on each of the second sample images to obtain each processed image; using the processed image corresponding to each second sample image as the first sample image corresponding to the second sample image, and based on each of the second sample images and the corresponding processed images, obtaining the first training data.
8. The method according to claim 7, wherein, the image degradation processing includes at least one of adding blurring processing, reducing image resolution processing, adding noise processing, or image format compression processing; and the first training data includes images processed by at least two image degradation processing methods.
9. An image processing apparatus, wherein, it includes: an image acquisition module, configured to acquire a to-be-processed image of a target object; an image processing module, configured to perform downsampling feature extraction on the to-be-processed image by invoking a trained image processing model to obtain an image feature map corresponding to the to-be-processed image, and determine a style feature map corresponding to the to-be-processed image according to the image feature map; wherein, the style feature map characterizes factors or information in the to-be-processed image that affect the visual effect of the image; based on the style feature map, perform upsampling feature extraction on the image feature map by the upsampling feature extraction module of the image processing model to obtain a target image; wherein, the image quality of the target image is higher than the image quality of the to-be-processed image, and the upsampling feature extraction includes the following steps: using the style feature map as the adjustment weight of the upsampling feature extraction module, adjusting the network parameters of the upsampling feature extraction module, and performing upsampling processing on the image feature map corresponding to the to-be-processed image by the adjusted network parameters to obtain the output feature map of the upsampling feature extraction module.
10. The apparatus according to claim 9, wherein, when the image processing module is configured to perform downsampling feature extraction on the to-be-processed image to obtain an image feature map corresponding to the to-be-processed image, it is specifically configured to: perform downsampling feature extraction on the to-be-processed image to obtain at least one scale of image feature map corresponding to the to-be-processed image; when the image processing module is configured to determine a style feature map corresponding to the to-be-processed image according to the image feature map, it is specifically configured to; for each scale, determine the style feature map corresponding to the scale according to the image feature map corresponding to the scale when the image processing module is configured to perform upsampling feature extraction on the image feature map based on the style feature map to obtain the target image, it is specifically configured to: Performing at least one upsampling feature extraction on the target feature map through at least one upsampling feature extraction module to obtain the target image, where the target feature map is the image feature map with the smallest scale among the image feature maps of the at least one scale, and each upsampling feature extraction module corresponds to a style feature map of one scale.
11. The apparatus according to claim 10, wherein, the image feature maps of the at least one scale include image feature maps of at least two scales, and the at least one upsampling feature extraction module includes at least two cascaded upsampling feature extraction modules; when the image processing module is used to perform at least one upsampling feature extraction on the target feature map through at least one upsampling feature extraction module to obtain the target image, it is specifically configured to: For each upsampling feature extraction module, using the style feature map corresponding to this upsampling feature extraction module as an adjustment weight to adjust the network parameters of this upsampling feature extraction module, performing upsampling processing on the input feature map corresponding to this upsampling feature extraction module through the adjusted network parameters to obtain the output feature map of this upsampling feature extraction module, and using the last output feature map as the target image; wherein, the input of the first upsampling feature extraction module is the target feature map, and the input of the modules other than the first upsampling feature extraction module is the output feature map of the previous upsampling feature extraction module of this upsampling feature extraction module.
12. The apparatus according to claim 10 or 11, wherein, the image processing model includes an image encoding network and an image generation network, the image encoding network includes a feature extraction module and a feature mapping module, and when the image processing module is used to perform the following processing on the to-be-processed image by calling the trained image processing model to obtain the target image corresponding to the to-be-processed image, it is specifically configured to: Inputting the to-be-processed image into the feature extraction module, and performing downsampling feature extraction on the to-be-processed image through the feature extraction module to obtain at least one scale of image feature maps corresponding to the to-be-processed image; Inputting the image feature maps of each scale into the feature mapping module respectively to obtain the style feature maps corresponding to each scale; Based on the style feature maps corresponding to each scale and the target feature map, performing at least one upsampling feature extraction on the target feature map through the image generation network to obtain the target image corresponding to the to-be-processed image.
13. The apparatus according to claim 12, wherein, the image processing model is obtained by the model training module through the following method: Obtaining first training data and an initial neural network model, the first training data includes a plurality of image pairs, each image pair includes a first sample image and a second sample image corresponding to the same image content, the image quality of the second sample image in each image pair is higher than that of the first sample image, and the initial neural network model includes an initial image processing model and an initial discriminator; Training the initial neural network model based on the first training data until the value of the total loss function corresponding to the model meets the training end condition, obtaining the trained neural network model, and using the trained image processing model as the image processing model; Wherein, the input of the initial image processing model is each of the first sample images, and the output is the predicted target image corresponding to each of the first sample images. The input of the initial discriminator is the predicted target image corresponding to each of the first sample images and the second sample images, and the output is the discrimination result of the predicted target image corresponding to each of the first sample images and the second sample images; The total loss function includes a first loss function and a second loss function. The value of the first loss function characterizes the difference between the predicted target image corresponding to each of the first sample images and the second sample images, and the value of the second loss function characterizes the accuracy of the discrimination result of the predicted target image corresponding to each of the first sample images and the second sample images.
14. The apparatus according to claim 13, wherein, The image generation network of the initial image processing model and the initial discriminator are pre-trained using the second training data. When training the initial neural network model based on the first training data, the learning rate of the image encoding network of the initial image processing model is greater than the learning rates of the image generation network of the initial image processing model and the initial discriminator.
15. The apparatus according to claim 13, wherein, When the model training module obtains the first training data, it is specifically configured to: Obtain each second sample image, perform image degradation processing on each of the second sample images respectively to obtain each processed image; Use the processed image corresponding to each second sample image as the first sample image corresponding to the second sample image, and based on each of the second sample images and the corresponding processed images, obtain the first training data.
16. The apparatus according to claim 15, wherein, The image degradation processing includes at least one of adding blurring processing, reducing image resolution processing, adding noise processing, or image format compression processing; the first training data includes images processed by at least two image degradation processing methods.
17. An electronic device, including a memory, a processor, and a computer program stored on the memory, wherein, The processor executes the computer program to implement the steps of the method according to any one of claims 1-8.
18. A computer-readable storage medium, on which a computer program is stored, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-8.
19. A computer program product, including a computer program, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-8.
Citation Information
Patent Citations
Method and device for repairing target face image
CN113592724A