Image processing method, apparatus and system
By employing a two-stage video enhancement scheme using generative adversarial networks, video frames containing noise, blur, and compression impairment are first mapped to clean, low-resolution video frames, and then to high-resolution video frames. This solves the problem of insufficient robustness and stability of existing video enhancement algorithms in real-world scenarios, achieving a higher level of visual perception.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-28
- Publication Date
- 2026-03-31
AI Technical Summary
Existing video enhancement algorithms have poor robustness and stability in real-world scenarios, especially when dealing with noise, blur, and compression impairment.
A two-stage video enhancement scheme based on generative adversarial networks is adopted. The video frames containing noise, blur and compression impairment are first mapped into clean low-resolution video frames through convolutional neural networks, and then mapped into high-resolution video frames. The image processing is performed using a processing model composed of the first and second generative networks.
This improves the robustness and stability of video enhancement algorithms in real-world scenarios, enhances visual perception, and solves the problem of poor robustness and stability of existing image processing methods in real-world scenarios.
Smart Images

Figure CN114004750B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and more specifically, to an image processing method, apparatus, and system. Background Technology
[0002] Currently, in order to achieve the restoration of old films, 4K remastering, and video quality enhancement on online video platforms, video enhancement algorithms based on convolutional neural networks can be used. Without additional input or interaction, these algorithms can intelligently improve the resolution of videos, remove blur, noise, and compression damage, and significantly improve the quality and viewing experience of videos.
[0003] However, existing video enhancement algorithms are mainly based on image enhancement algorithms, which perform poorly in real-world scenarios, mainly due to issues such as poor robustness and stability, and blurred details.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] This application provides an image processing method, apparatus, and system to at least address the technical problem of poor robustness and stability of image processing methods in real-world scenarios in related technologies.
[0006] According to one aspect of the embodiments of this application, an image processing method is provided, comprising: receiving a first image; processing the first image using a processing model to obtain a second image, wherein the resolution of the second image satisfies a first preset condition with the resolution of the first image, the processing model is used to input the first image into a first generator network to obtain a third image, and input the third image into a second generator network to obtain a second image, wherein the resolution of the third image satisfies a second preset condition with the resolution of the first image; and displaying the second image.
[0007] According to another aspect of the embodiments of this application, an image processing method is also provided, including: acquiring a first image; processing the first image using a processing model to obtain a second image, wherein the resolution of the second image and the resolution of the first image satisfy a first preset condition, and the processing model is used to input the first image into a first generator network to obtain a third image, and input the third image into a second generator network to obtain a second image, wherein the resolution of the third image and the resolution of the first image satisfy a second preset condition.
[0008] According to another aspect of the embodiments of this application, an image processing method is also provided, comprising: receiving a model training request; obtaining training samples and an initial model corresponding to the model training request, wherein the training samples include: a first image sample and a second image sample, the resolution of the second image sample and the resolution of the first image sample satisfying a first preset condition; training the initial model using the training samples to obtain a processing model, wherein the processing model is used to input the first image sample into a first generator network to obtain a third image sample, and input the third image sample into a second generator network to obtain a fourth image sample, the resolution of the third image sample and the resolution of the first image sample satisfying a second preset condition; and outputting the processing model.
[0009] According to another aspect of the embodiments of this application, an image processing method is also provided, comprising: acquiring training samples and an initial model, wherein the training samples include: a first image sample and a second image sample, the resolution of the second image sample and the resolution of the first image sample satisfying a first preset condition; training the initial model using the training samples to obtain a processing model, wherein the processing model is used to input the first image sample into a first generator network to obtain a third image sample, and input the third image sample into a second generator network to obtain a fourth image sample, the resolution of the third image sample and the resolution of the first image sample satisfying a second preset condition.
[0010] According to another aspect of the embodiments of this application, an image processing apparatus is also provided, comprising: a receiving module for receiving a first image; a processing module for processing the first image using a processing model to obtain a second image, wherein the resolution of the second image satisfies a first preset condition with the resolution of the first image, the processing model is used to input the first image into a first generating network to obtain a third image, and input the third image into a second generating network to obtain a second image, wherein the resolution of the third image satisfies a second preset condition with the resolution of the first image; and a display module for displaying the second image.
[0011] According to another aspect of the embodiments of this application, an image processing apparatus is also provided, including: an acquisition module for acquiring a first image; and a processing module for processing the first image using a processing model to obtain a second image, wherein the resolution of the second image and the resolution of the first image satisfy a first preset condition, and the processing model is used to input the first image into a first generation network to obtain a third image, and input the third image into a second generation network to obtain a second image, wherein the resolution of the third image and the resolution of the first image satisfy a second preset condition.
[0012] According to another aspect of the embodiments of this application, an image processing apparatus is also provided, comprising: a receiving module for receiving a model training request; an acquiring module for acquiring training samples and an initial model corresponding to the model training request, wherein the training samples include: a first image sample and a second image sample, the resolution of the second image sample and the resolution of the first image sample satisfying a first preset condition; a training module for training the initial model using the training samples to obtain a processing model, wherein the processing model is used to input the first image sample into a first generator network to obtain a third image sample, and input the third image sample into a second generator network to obtain a fourth image sample, the resolution of the third image sample and the resolution of the first image sample satisfying a second preset condition; and an output module for outputting the processing model.
[0013] According to another aspect of the embodiments of this application, an image processing apparatus is also provided, comprising: an acquisition module for acquiring training samples and an initial model, wherein the training samples include: a first image sample and a second image sample, the resolution of the second image sample and the resolution of the first image sample satisfying a first preset condition; and a training module for training the initial model using the training samples to obtain a processing model, wherein the processing model is used to input the first image sample into a first generator network to obtain a third image sample, and input the third image sample into a second generator network to obtain a fourth image sample, the resolution of the third image sample and the resolution of the first image sample satisfying a second preset condition.
[0014] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored program, wherein, when the program is running, it controls the device where the computer-readable storage medium is located to perform the above-described image processing method.
[0015] According to another aspect of the embodiments of this application, a computer terminal is also provided, including: a memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program executes the above-described image processing method when it runs.
[0016] According to another aspect of the embodiments of this application, an image processing system is also provided, including: a processor; and a memory connected to the processor, configured to provide the processor with instructions for processing the following processing steps: receiving a first image; processing the first image using a processing model to obtain a second image, wherein the resolution of the second image and the resolution of the first image satisfy a first preset condition, the processing model being configured to input the first image into a first generation network to obtain a third image, and input the third image into a second generation network to obtain a second image, wherein the resolution of the third image and the resolution of the first image satisfy a second preset condition; and displaying the second image.
[0017] In this embodiment, the purpose of video quality enhancement is achieved by using a processing model composed of a first generator network and a second generator network to process the image. Since the processing model includes a first generator network and a second generator network, and the input and output images of the first generator network have the same resolution, while the input and output images of the second generator network have different resolutions, the robustness and stability of the video enhancement algorithm in real-world scenarios are increased, and the visual perception effect is improved. This solves the technical problem of poor robustness and stability of image processing methods in real-world scenarios in related technologies. Attached Figure Description
[0018] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0019] Figure 1 This is a hardware structure block diagram of a computer terminal for implementing an image processing method according to an embodiment of this application;
[0020] Figure 2 This is a flowchart of a first image processing algorithm according to an embodiment of this application;
[0021] Figure 3 This is a schematic diagram of an optional model training stage according to an embodiment of this application;
[0022] Figure 4 This is a schematic diagram of an optional model inference stage according to an embodiment of this application;
[0023] Figure 5 This is a flowchart of a second image processing algorithm according to an embodiment of this application;
[0024] Figure 6 This is a flowchart of a third image processing algorithm according to an embodiment of this application;
[0025] Figure 7 This is a flowchart of the fourth image processing algorithm according to an embodiment of this application;
[0026] Figure 8 This is a schematic diagram of a first image processing apparatus according to an embodiment of this application;
[0027] Figure 9 This is a schematic diagram of a second image processing apparatus according to an embodiment of this application;
[0028] Figure 10 This is a schematic diagram of a third image processing apparatus according to an embodiment of this application;
[0029] Figure 11 This is a schematic diagram of a fourth image processing apparatus according to an embodiment of this application; and
[0030] Figure 12 This is a structural block diagram of a computer terminal according to an embodiment of this application. Detailed Implementation
[0031] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0033] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:
[0034] Video quality enhancement: This refers to the task of repairing low-quality videos and improving their overall image quality. Specifically, it includes removing blur, noise, and compression artifacts, as well as increasing video resolution.
[0035] Generative Adversarial Networks (GANs) can refer to an unsupervised learning method that learns by having two neural networks play against each other.
[0036] Current image super-resolution algorithms are mainly trained on simulated data pairs. The low-quality images in the simulated data pairs are mainly obtained by bicubic downsampling of high-quality images. However, the models trained on the simulated data pairs do not perform well in real-world scenarios.
[0037] Currently, image super-resolution algorithms based on convolutional neural networks are mainly divided into two categories. The first category is image super-resolution algorithms oriented towards objective numerical indicators, such as Image Super-Resolution Using Very Deep Residual Channel Attention Networks (RCAN), whose goal is to obtain better objective indicators, such as Peak Signal-to-Noise Ratio (PSNR). The second category is image super-resolution algorithms oriented towards human perception, such as Enhanced Super-Resolution Generative Adversarial Networks (ESRGAN), whose goal is to obtain better visual effects.
[0038] The first type of image super-resolution algorithm is usually trained using L1 loss (minimum absolute value loss). This algorithm can achieve certain results, but it suffers from problems such as blurred details. The second type of image super-resolution algorithm is usually trained by adding GAN loss. This algorithm can generate more details, but it suffers from problems such as poor stability.
[0039] To address the aforementioned technical issues, this application provides a two-stage video enhancement scheme based on generative adversarial networks. Video frames containing noise, blur, and compression impairment are first mapped into clean, low-resolution video frames through a convolutional neural network 1, and then mapped into high-resolution video frames through a convolutional neural network 2. This not only increases the robustness and stability of the video enhancement algorithm in real-world scenarios but also achieves higher visual perception effects.
[0040] Example 1
[0041] According to an embodiment of this application, an image processing method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0042] The methods and embodiments provided in this application can be executed on mobile terminals, computer terminals, or similar computing devices. Figure 1 A hardware block diagram of a computer terminal (or mobile device) for implementing an image processing method is shown. Figure 1As shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0043] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuitry are generally referred to herein as "data processing circuitry". This data processing circuitry may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). This data processing circuitry serves as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0044] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the image processing method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the above-mentioned image processing method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0045] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0046] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0047] It should be noted here that, in some optional embodiments, the above... Figure 1 The computer device (or mobile device) shown may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be noted that... Figure 1 This is only one instance of a specific particular instance and is intended to illustrate the types of components that may exist in the aforementioned computer device (or mobile device).
[0048] It should be noted here that, in some embodiments, the above... Figure 1 The computer device (or mobile device) shown has a touch display (also referred to as a "touchscreen" or "touch display screen"). In some embodiments, the above... Figure 1 The computer device (or mobile device) shown has a graphical user interface (GUI), through which users can interact with the GUI by touching and / or gesturing on a touch-sensitive surface. The human-computer interaction functions here may include the following interactions: creating web pages, drawing, word processing, creating electronic documents, playing games, video conferencing, instant messaging, sending and receiving emails, call interface, playing digital video, playing digital music, and / or web browsing, etc. Executable instructions for performing the above human-computer interaction functions are configured / stored in one or more processor-executable computer program products or readable storage media.
[0049] Under the aforementioned operating environment, this application provides the following: Figure 2 The image processing algorithm shown. Figure 2 This is a flowchart of a first image processing algorithm according to an embodiment of this application. For example... Figure 2 As shown, the method may include the following steps:
[0050] Step S202: Receive the first image.
[0051] The first image in the above steps can be a low-resolution video frame that requires video quality enhancement, containing noise, blur, or compression impairment, but is not limited to these. This video frame can be a video frame directly uploaded by the user, or a video frame extracted from a user-uploaded video. For example, taking an online video platform as an example, the first image can be a video frame from a user-uploaded live video.
[0052] In one optional embodiment, the image processing algorithm provided in this application can be executed by the client. The user can directly capture images or videos in real time using a shooting device, which then transmits the images or videos to the client, allowing the client to receive the first image. Alternatively, the user can directly select a pre-stored image or video by operating controls in the interactive interface provided by the client, allowing the client to receive the first image based on the user's selection.
[0053] In another alternative embodiment, the image processing algorithm provided in this application can be executed by a server. The user can directly capture images or videos in real time using a shooting device and send them to the server via the internet, thus allowing the server to receive the first image. Alternatively, the user can directly operate controls in the interactive interface provided by the server to select pre-stored images or videos, allowing the server to receive the first image based on the user's selection.
[0054] It should be noted that the aforementioned shooting devices can be cameras, camcorders, smartphones, tablets, laptops, etc., but are not limited to these.
[0055] Step S204: The first image is processed using a processing model to obtain a second image, wherein the resolution of the second image and the resolution of the first image satisfy a first preset condition. The processing model is used to input the first image into a first generator network to obtain a third image, and input the third image into a second generator network to obtain a second image, wherein the resolution of the third image and the resolution of the first image satisfy a second preset condition.
[0056] The processing model in the above steps can be a two-stage video enhancement model based on generative adversarial networks. That is, the processing model contains two deep convolutional neural networks, which serve as two different generative networks. Since the image processing tasks of the two generative networks are different, the network structures of the two generative networks are different. This application does not limit the specific structure of the generative network. Any network structure that can complete the corresponding image processing task can be applied to the embodiments of this application.
[0057] The first preset condition in the above steps can refer to the condition that the resolution of the second image differs significantly from the resolution of the first image. Optionally, the condition that the resolution of the second image satisfies the first preset condition includes: the difference between the resolution of the second image and the resolution of the first image is greater than a first preset value. The second preset condition can refer to the condition that the resolution of the third image differs slightly from the resolution of the first image. Optionally, the condition that the resolution of the third image satisfies the second preset condition includes: the difference between the resolution of the third image and the resolution of the first image is less than a second preset value. The first and second preset values can be set according to actual needs, and they can also be the same. A resolution difference greater than the first preset value indicates a significant difference in resolution between the two images, meaning the two images have different resolutions; a resolution difference less than the second preset value indicates a small difference in resolution between the two images, meaning the two images have the same resolution.
[0058] The second image in the above steps can be an image obtained by repairing the first image and improving its image quality. In this embodiment, the example is taken where the resolution of the first image is lower than that of the second image. The first image can refer to a low-quality, low-resolution video frame, the third image can refer to a high-quality, low-resolution video frame, and the second image can refer to a high-quality, high-resolution video frame.
[0059] It should be noted that when the image processing algorithm provided in this application is executed by the client, the processing model needs to be stored on the client; when the image processing algorithm provided in this application is executed by the server, the processing model needs to be stored on the server. However, the client's storage space and computing resources are limited. In order to reduce the client's resource consumption and improve processing efficiency and effect, this embodiment of the application uses the execution of the above-mentioned image processing algorithm by the server as an example for explanation.
[0060] Step S206: Display the second image.
[0061] In one optional embodiment, to facilitate user confirmation of whether the second image meets their needs, the client or server, after processing the first image using a processing model to obtain the second image, can directly display the second image in the user's interactive interface, allowing the user to view the video enhancement effect. If the second image does not meet the user's needs, the user can provide feedback to the client or server, which will then update the network weights of the processing model to improve the video enhancement effect.
[0062] Based on the technical solution provided in the above embodiments of this application, after receiving the first image, the first image can be processed using a processing model composed of a first generator network and a second generator network, and the processed second image can be displayed. It is readily apparent that, since the processing model includes both a first generator network and a second generator network, and the input and output images of the first generator network have the same resolution, while the input and output images of the second generator network have different resolutions, the purpose of video quality enhancement is achieved. This increases the robustness and stability of the video enhancement algorithm in real-world scenarios, improves the visual perception effect, and thus solves the technical problem of poor robustness and stability of image processing methods in real-world scenarios in related technologies.
[0063] In the above embodiments of this application, the method further includes: acquiring a second image sample; preprocessing the second image sample to obtain a first image sample, wherein the preprocessing includes at least one of the following: downsampling, adding blur information, adding noise information, and image compression; and training a processing model using the first image sample and the second image sample.
[0064] The second image sample in the above steps can be a high-quality, high-resolution video frame captured from the network.
[0065] In one optional embodiment, to ensure that the trained processing model can better reflect the degradation in real-world scenarios, low-quality, low-resolution video frames can be obtained through artificial degradation methods. In this embodiment, the artificial degradation methods include downsampling, blurring, adding noise, and compression, but are not limited to these, and are not simply bicubic downsampling. Through these methods, pairs of real high-quality, high-resolution video frames and real low-quality, low-resolution video frame data can be obtained, and these video frame data pairs can then be used as training samples for model training.
[0066] It should be noted that after obtaining low-quality, low-resolution video frames through manual degradation, all video frames can be further cleaned to obtain the final video frame data pairs used for model training.
[0067] In the above embodiments of this application, training the processing model using the first image sample and the second image sample includes: inputting the first image sample into a first generator network to obtain a third image sample; inputting the third image sample into a second generator network to obtain a fourth image sample; inputting the second image sample and the fourth image sample into a discriminator network in the processing model to obtain a discrimination result; obtaining a target loss value of the processing model based on the second image sample, the third image sample, the fourth image sample, and the discrimination result; and updating the network parameters of the processing model based on the target loss value.
[0068] The purpose of the first and second generator networks in the above steps is to generate fake high-quality, high-resolution video frames, while the task of the discrimination network is to distinguish between real and fake high-quality, high-resolution video frames, that is, to distinguish whether the input video frame is a real video frame or a video frame generated by the second generator network.
[0069] The fourth image sample in the above steps can be a high-quality, high-resolution video frame generated by the second generation network. The higher the processing accuracy of the second generation network, the higher the similarity between the fourth image sample and the second image sample, and the easier it is to fool the discriminator network.
[0070] The target loss value in the above steps can be the total loss value of the entire processing model, and can be determined by the following loss values: the loss value used to constrain the generated fourth and third image samples to be similar to the second image sample in terms of content, the loss value used to constrain the fourth and second image samples to be similar in terms of perception, and the adversarial loss value.
[0071] It should be noted that the discriminant network is only used during model training. This application does not limit the specific structure of the discriminant network; any network structure capable of performing the corresponding discrimination task can be applied to the embodiments of this application. To implement the training process of the first and second generator networks, the discriminant network can be trained first, and then the adversarial loss value can be determined based on the trained discriminant network, thereby achieving the purpose of training the first and second generator networks.
[0072] In the above embodiments of this application, obtaining the target loss value of the processing model based on the second image sample, the third image sample, the fourth image sample, and the discrimination result includes: inputting the second image sample and the fourth image sample into the minimum absolute value loss function to obtain a first loss value; performing downsampling processing on the second image sample, and inputting the processed image sample and the third image sample into the minimum absolute value loss function to obtain a second loss value; inputting the second image sample and the fourth image sample into the perceptual loss function to obtain a third loss value; obtaining a fourth loss value based on the discrimination result; and obtaining a weighted sum of the first loss value, the second loss value, the third loss value, and the fourth loss value to obtain the target loss value.
[0073] The processed image sample in the above steps can be an image sample obtained by bicubic downsampling the second image sample.
[0074] In an optional embodiment, for the second and fourth image samples, an L1 loss can be calculated to obtain a first loss value, denoted as L. pix1 For the processed image samples and the third image sample, the L1 loss can also be calculated to obtain the second loss value, denoted as L.pix2 For the second and fourth image samples, a perceptual loss can also be calculated to obtain a third loss value, denoted as L. perceptual For the discrimination result, the adversarial loss can be calculated to obtain the fourth loss value, denoted as L. GAN The final target loss value is L = aL pix1 +bL pix2 +cL perceptual +dL GAN , where a, b, c, and d are the weights of the four loss values, respectively.
[0075] It should be noted that for the four loss values mentioned above, L pix1 and L pix2 The loss value L is used to constrain the content similarity between the generated fourth and third image samples and the second image sample. perceptual The loss value is used to constrain the fourth image sample to be perceptually similar to the second image sample.
[0076] It should also be noted that the processed image samples and the third image samples can be input into another discriminant network, and a new adversarial loss can be calculated based on the output of the discriminant network. In this case, the processing model can include two generator networks and two discriminant networks.
[0077] The following is combined Figure 3 and Figure 4 A preferred embodiment of this application will be described in detail, wherein, Figure 3 This shows the model training phase. Figure 4 The model inference stage is shown.
[0078] like Figure 3 As shown, the processing flow during the model training phase is as follows: Collect real high-quality video frames and manually downgrade them to obtain real low-quality video frames; clean the data to obtain pairs of real high-quality and real low-quality video frame data; (e.g., ...) Figure 3 As shown by the solid line, the real low-quality video frame LQ is passed through convolutional neural network 1 (generator 1) to obtain a clean intermediate result IR. The intermediate result IR is then passed through convolutional neural network 2 (generator 2) to generate a high-quality video frame HR. Figure 3 As shown by the dashed line, the L1 loss is calculated for the generated high-quality video frame HR and the real high-quality video frame GT, denoted as L. pix1 The L1 loss is calculated for the intermediate result IR and the high-quality video frames GT after bicubic downsampling, denoted as L. pix2The perceptual loss, denoted as L_perceptual, is calculated for the generated high-quality video frame HR and the real high-quality video frame GT. The generated high-quality video frame HR and the real high-quality video frame GT are then input into convolutional neural network 3 (discriminator 1) to calculate the adversarial loss L. GAN Thus, the final loss L = aL can be obtained. pix1 +bL pix2 +cL perceptual +dL GAN This completes the model training process.
[0079] like Figure 4 As shown, the processing flow in the model inference stage is as follows: Real low-quality video frames are input into Convolutional Neural Network 1 (Generator 1) to obtain intermediate results, and then the intermediate results are input into Convolutional Neural Network 2 (Generator 2) to generate high-quality video frames. By repeatedly performing the above operations, low-quality video can be converted into high-quality video.
[0080] The above scheme first maps noisy, blurred, and compressed video frames into clean, low-resolution video frames using Convolutional Neural Network 1, and then maps them into high-quality, high-resolution video frames using Convolutional Neural Network 2. By incorporating various degradation methods from real-world scenes, such as noise, blur, and compression, the robustness of the video enhancement algorithm is greatly enhanced, and the problem of poor performance of models trained on bicubic downsampling data in real-world scenes is addressed to some extent. By introducing Generative Adversarial Networks, the detail blurring problem of current image super-resolution algorithms guided by objective numerical indicators is solved. By mapping noisy video frames into clean video frames, the instability of current human perception-oriented image super-resolution algorithms in noisy scenes is addressed.
[0081] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0082] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0083] Example 2
[0084] According to an embodiment of this application, an image processing method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0085] Figure 5 This is a flowchart of a second image processing algorithm according to an embodiment of this application. For example... Figure 5 As shown, the method may include the following steps:
[0086] Step S502: Obtain the first image.
[0087] The first image in the above steps can be a low-resolution video frame that requires video quality enhancement, containing noise, blur, or compression impairment, but is not limited to these. This video frame can be a video frame directly uploaded by the user, or a video frame extracted from a user-uploaded video. For example, taking an online video platform as an example, the first image can be a video frame from a user-uploaded live video.
[0088] Step S504: The first image is processed using a processing model to obtain a second image, wherein the resolution of the second image and the resolution of the first image satisfy a first preset condition. The processing model is used to input the first image into a first generation network to obtain a third image, and input the third image into a second generation network to obtain a second image, wherein the resolution of the third image and the resolution of the first image satisfy a second preset condition.
[0089] The processing model in the above steps can be a two-stage video enhancement model based on generative adversarial networks. That is, the processing model contains two deep convolutional neural networks, which serve as two different generative networks. Since the image processing tasks of the two generative networks are different, the network structures of the two generative networks are different. This application does not limit the specific structure of the generative network. Any network structure that can complete the corresponding image processing task can be applied to the embodiments of this application.
[0090] The first preset condition in the above steps can refer to the condition that the resolution of the second image differs significantly from the resolution of the first image. Optionally, the condition that the resolution of the second image satisfies the first preset condition includes: the difference between the resolution of the second image and the resolution of the first image is greater than a first preset value. The second preset condition can refer to the condition that the resolution of the third image differs slightly from the resolution of the first image. Optionally, the condition that the resolution of the third image satisfies the second preset condition includes: the difference between the resolution of the third image and the resolution of the first image is less than a second preset value. The first and second preset values can be set according to actual needs, and they can also be the same. A resolution difference greater than the first preset value indicates a significant difference in resolution between the two images, meaning the two images have different resolutions; a resolution difference less than the second preset value indicates a small difference in resolution between the two images, meaning the two images have the same resolution.
[0091] The second image in the above steps can be an image obtained by repairing the first image and improving its image quality. In this embodiment, the example is taken where the resolution of the first image is lower than that of the second image. The first image can refer to a low-quality, low-resolution video frame, the third image can refer to a high-quality, low-resolution video frame, and the second image can refer to a high-quality, high-resolution video frame.
[0092] In the above embodiments of this application, the method further includes: acquiring a second image sample; preprocessing the second image sample to obtain a first image sample, wherein the preprocessing includes at least one of the following: downsampling, adding blur information, adding noise information, and image compression; and training a processing model using the first image sample and the second image sample.
[0093] The second image sample in the above steps can be a high-quality, high-resolution video frame captured from the network.
[0094] In the above embodiments of this application, training the processing model using the first image sample and the second image sample includes: inputting the first image sample into a first generator network to obtain a third image sample; inputting the third image sample into a second generator network to obtain a fourth image sample; inputting the second image sample and the fourth image sample into a discriminator network in the processing model to obtain a discrimination result; obtaining a target loss value of the processing model based on the second image sample, the third image sample, the fourth image sample, and the discrimination result; and updating the network parameters of the processing model based on the target loss value.
[0095] The purpose of the first and second generator networks in the above steps is to generate fake high-quality, high-resolution video frames, while the task of the discrimination network is to distinguish between real and fake high-quality, high-resolution video frames, that is, to distinguish whether the input video frame is a real video frame or a video frame generated by the second generator network.
[0096] The fourth image sample in the above steps can be a high-quality, high-resolution video frame generated by the second generation network. The higher the processing accuracy of the second generation network, the higher the similarity between the fourth image sample and the second image sample, and the easier it is to fool the discriminator network.
[0097] The target loss value in the above steps can be the total loss value of the entire processing model, and can be determined by the following loss values: the loss value used to constrain the generated fourth and third image samples to be similar to the second image sample in terms of content, the loss value used to constrain the fourth and second image samples to be similar in terms of perception, and the adversarial loss value.
[0098] In the above embodiments of this application, obtaining the target loss value of the processing model based on the second image sample, the third image sample, the fourth image sample, and the discrimination result includes: inputting the second image sample and the fourth image sample into the minimum absolute value loss function to obtain a first loss value; performing downsampling processing on the second image sample, and inputting the processed image sample and the third image sample into the minimum absolute value loss function to obtain a second loss value; inputting the second image sample and the fourth image sample into the perceptual loss function to obtain a third loss value; obtaining a fourth loss value based on the discrimination result; and obtaining a weighted sum of the first loss value, the second loss value, the third loss value, and the fourth loss value to obtain the target loss value.
[0099] The processed image sample in the above steps can be an image sample obtained by bicubic downsampling the second image sample.
[0100] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0101] Example 3
[0102] According to an embodiment of this application, an image processing method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0103] Figure 6 This is a flowchart of a third image processing algorithm according to an embodiment of this application. For example... Figure 6 As shown, the method may include the following steps:
[0104] Step S602: Receive model training request.
[0105] To train a model with high processing accuracy, a large number of training samples are often required for multiple training iterations, resulting in a significant amount of data and computational overhead. To reduce resource consumption on user devices (such as smartphones, tablets, laptops, PDAs, and personal computers), model training can be performed on a server, with only the trained model deployed on the user device for ease of use. Furthermore, to significantly reduce the computational burden on user devices, the trained model can be directly deployed on the server. The user device connects to the server via a specific interface, sending the data to be processed. The server then uses the deployed model to process the data and returns the results to the user device.
[0106] The model training request in the above steps can be generated according to the user's model usage needs. The request can carry the data type to be processed and the expected processing result. For example, in the medical field, the data type to be processed may be low-quality video frames, and the expected processing result may be high-quality video frames.
[0107] Since the server can provide model training services to different users with varying model needs, specific model structures, and training samples, in one optional embodiment, an interactive interface can be provided on the user device. The user inputs a model training request in an input area, and the user device then sends the request to the server via the network. For greater targeting, the server can offer different model training schemes based on the user's type, allowing the user to select one in the input area. The user device can then generate a model training request based on the user's selection and send it to the server via the network.
[0108] Step S604: Obtain the training samples and initial model corresponding to the model training request. The training samples include: a first image sample and a second image sample. The resolution of the second image and the resolution of the first image satisfy a first preset condition.
[0109] The first preset condition in the above steps can refer to the condition that the resolution of the second image differs significantly from the resolution of the first image. Optionally, the condition that the resolution of the second image satisfies the first preset condition includes: the difference between the resolution of the second image and the resolution of the first image is greater than a first preset value. The second preset condition can refer to the condition that the resolution of the third image differs slightly from the resolution of the first image. The first preset value can be set according to actual needs. A resolution difference greater than the first preset value indicates a significant difference in resolution between the two images, and they can be considered to have different resolutions.
[0110] The training samples in the above steps can be low-quality, low-resolution video frames and high-quality, high-resolution video frame data pairs. In this embodiment, the resolution of the first image sample is less than that of the second image sample. The first image sample can be a low-quality, low-resolution video frame containing noise, blur, or compression impairment, but is not limited to these. The second image sample can be a high-quality, high-resolution video frame.
[0111] Step S606: Train the initial model using training samples to obtain a processing model. The processing model is used to input the first image sample into the first generator network to obtain the third image sample, and input the third image sample into the second generator network to obtain the fourth image sample. The resolution of the third image sample and the resolution of the first image sample satisfy the second preset condition.
[0112] The processing model in the above steps can be a two-stage video enhancement model based on generative adversarial networks. That is, the processing model contains two deep convolutional neural networks, which serve as two different generative networks. Since the image processing tasks of the two generative networks are different, the network structures of the two generative networks are different. This application does not limit the specific structure of the generative network. Any network structure that can complete the corresponding image processing task can be applied to the embodiments of this application.
[0113] The second preset condition in the above steps can refer to the condition that the resolution of the third image is relatively close to the resolution of the first image. Optionally, the condition that the resolution of the third image and the resolution of the first image satisfy the second preset condition includes: the difference between the resolution of the third image and the resolution of the first image is less than a second preset value. The second preset value can be set according to actual needs, and the first and second preset values can also be the same. A resolution difference less than the second preset value indicates that the resolution difference between the two images is small, and the two images can be considered to have the same resolution.
[0114] Step S608: Output the processing model.
[0115] In one alternative embodiment, if the processing model needs to be deployed on the client, the server can transmit the processing model to the client via the network; if the processing model needs to be deployed on the server, the processing model can be directly put online, so that users can use the online processing model to enhance video quality.
[0116] In the above embodiments of this application, obtaining the training samples corresponding to the model training request includes: obtaining a second image sample; preprocessing the second image sample to obtain a first image sample, wherein the preprocessing includes at least one of the following: downsampling, adding blur information, adding noise information, and image compression.
[0117] The second image sample in the above steps can be a high-quality, high-resolution video frame captured from the network.
[0118] In the above embodiments of this application, training the initial model using training samples includes: inputting the second image sample and the fourth image sample into the discriminant network in the processing model to obtain a discrimination result; obtaining the target loss value of the processing model based on the second image sample, the third image sample, the fourth image sample and the discrimination result; and updating the network parameters of the processing model based on the target loss value.
[0119] The purpose of the first and second generator networks in the above steps is to generate fake high-quality, high-resolution video frames, while the task of the discrimination network is to distinguish between real and fake high-quality, high-resolution video frames, that is, to distinguish whether the input video frame is a real video frame or a video frame generated by the second generator network.
[0120] The fourth image sample in the above steps can be a high-quality, high-resolution video frame generated by the second generation network. The higher the processing accuracy of the second generation network, the higher the similarity between the fourth image sample and the second image sample, and the easier it is to fool the discriminator network.
[0121] The target loss value in the above steps can be the total loss value of the entire processing model, and can be determined by the following loss values: the loss value used to constrain the generated fourth and third image samples to be similar to the second image sample in terms of content, the loss value used to constrain the fourth and second image samples to be similar in terms of perception, and the adversarial loss value.
[0122] In the above embodiments of this application, obtaining the target loss value of the processing model based on the second image sample, the third image sample, the fourth image sample, and the discrimination result includes: inputting the second image sample and the fourth image sample into the minimum absolute value loss function to obtain a first loss value; performing downsampling processing on the second image sample, and inputting the processed image sample and the third image sample into the minimum absolute value loss function to obtain a second loss value; inputting the second image sample and the fourth image sample into the perceptual loss function to obtain a third loss value; obtaining a fourth loss value based on the discrimination result; and obtaining a weighted sum of the first loss value, the second loss value, the third loss value, and the fourth loss value to obtain the target loss value.
[0123] The processed image sample in the above steps can be an image sample obtained by bicubic downsampling the second image sample.
[0124] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0125] Example 4
[0126] According to an embodiment of this application, an image processing method is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0127] Figure 7 This is a flowchart of a fourth image processing algorithm according to an embodiment of this application. For example... Figure 7 As shown, the method may include the following steps:
[0128] Step S702: Obtain training samples and an initial model, wherein the training samples include: a first image sample and a second image sample, and the resolution of the second image satisfies a first preset condition with the resolution of the first image.
[0129] The training samples in the above steps can be low-quality, low-resolution video frames and high-quality, high-resolution video frame data pairs. In this embodiment, the resolution of the first image sample is less than that of the second image sample. The first image sample can be a low-quality, low-resolution video frame containing noise, blur, or compression impairment, but is not limited to these. The second image sample can be a high-quality, high-resolution video frame.
[0130] The first preset condition in the above steps can refer to the condition that the resolution of the second image differs significantly from the resolution of the first image. Optionally, the condition that the resolution of the second image satisfies the first preset condition includes: the difference between the resolution of the second image and the resolution of the first image is greater than a first preset value. The second preset condition can refer to the condition that the resolution of the third image differs slightly from the resolution of the first image. The first preset value can be set according to actual needs. A resolution difference greater than the first preset value indicates a significant difference in resolution between the two images, and they can be considered to have different resolutions.
[0131] Step S704: Train the initial model using training samples to obtain a processing model. The processing model is used to input the first image sample into the first generator network to obtain the third image sample, and input the third image sample into the second generator network to obtain the fourth image sample. The resolution of the third image sample and the resolution of the first image sample satisfy the second preset condition.
[0132] The processing model in the above steps can be a two-stage video enhancement model based on generative adversarial networks. That is, the processing model contains two deep convolutional neural networks, which serve as two different generative networks. Since the image processing tasks of the two generative networks are different, the network structures of the two generative networks are different. This application does not limit the specific structure of the generative network. Any network structure that can complete the corresponding image processing task can be applied to the embodiments of this application.
[0133] The second preset condition in the above steps can refer to the condition that the resolution of the third image is relatively close to the resolution of the first image. Optionally, the condition that the resolution of the third image and the resolution of the first image satisfy the second preset condition includes: the difference between the resolution of the third image and the resolution of the first image is less than a second preset value. The second preset value can be set according to actual needs, and the first and second preset values can also be the same. A resolution difference less than the second preset value indicates that the resolution difference between the two images is small, and the two images can be considered to have the same resolution.
[0134] In the above embodiments of this application, obtaining training samples includes: obtaining a second image sample; preprocessing the second image sample to obtain a first image sample, wherein the preprocessing includes at least one of the following: downsampling, adding blur information, adding noise information, and image compression.
[0135] The second image sample in the above steps can be a high-quality, high-resolution video frame captured from the network.
[0136] In the above embodiments of this application, training the initial model using training samples includes: inputting the second image sample and the fourth image sample into the discriminant network in the processing model to obtain a discrimination result; obtaining the target loss value of the processing model based on the second image sample, the third image sample, the fourth image sample and the discrimination result; and updating the network parameters of the processing model based on the target loss value.
[0137] The purpose of the first and second generator networks in the above steps is to generate fake high-quality, high-resolution video frames, while the task of the discrimination network is to distinguish between real and fake high-quality, high-resolution video frames, that is, to distinguish whether the input video frame is a real video frame or a video frame generated by the second generator network.
[0138] The fourth image sample in the above steps can be a high-quality, high-resolution video frame generated by the second generation network. The higher the processing accuracy of the second generation network, the higher the similarity between the fourth image sample and the second image sample, and the easier it is to fool the discriminator network.
[0139] The target loss value in the above steps can be the total loss value of the entire processing model, and can be determined by the following loss values: the loss value used to constrain the generated fourth and third image samples to be similar to the second image sample in terms of content, the loss value used to constrain the fourth and second image samples to be similar in terms of perception, and the adversarial loss value.
[0140] In the above embodiments of this application, obtaining the target loss value of the processing model based on the second image sample, the third image sample, the fourth image sample, and the discrimination result includes: inputting the second image sample and the fourth image sample into the minimum absolute value loss function to obtain a first loss value; performing downsampling processing on the second image sample, and inputting the processed image sample and the third image sample into the minimum absolute value loss function to obtain a second loss value; inputting the second image sample and the fourth image sample into the perceptual loss function to obtain a third loss value; obtaining a fourth loss value based on the discrimination result; and obtaining a weighted sum of the first loss value, the second loss value, the third loss value, and the fourth loss value to obtain the target loss value.
[0141] The processed image sample in the above steps can be an image sample obtained by bicubic downsampling the second image sample.
[0142] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0143] Example 5
[0144] According to embodiments of this application, an image processing apparatus for implementing the above-described image processing method is also provided, such as... Figure 8 As shown, the device 800 includes: a receiving module 802, a processing module 804, and a display module 806.
[0145] The receiving module 802 is used to receive a first image; the processing module 804 is used to process the first image using a processing model to obtain a second image, wherein the resolution of the second image and the resolution of the first image satisfy a first preset condition; the processing model is used to input the first image into a first generating network to obtain a third image, and input the third image into a second generating network to obtain a second image, wherein the resolution of the third image and the resolution of the first image satisfy a second preset condition; and the display module 806 is used to display the second image.
[0146] It should be noted that the receiving module 802, processing module 804, and display module 806 mentioned above correspond to steps S202 to S206 in Embodiment 1. The three modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0147] In the above embodiments of this application, the device further includes an acquisition module and a training module.
[0148] The acquisition module is used to acquire a second image sample; the processing module is also used to preprocess the second image sample to obtain a first image sample, wherein the preprocessing includes at least one of the following: downsampling, adding blur information, adding noise information, and image compression; the training module is used to train the processing model using the first image sample and the second image sample.
[0149] In the above embodiments of this application, the training module includes: a first input unit, a second input unit, a third input unit, a processing unit, and an update unit.
[0150] The first input unit is used to input the first image sample into the first generator network to obtain the third image sample; the second input unit is used to input the third image sample into the second generator network to obtain the fourth image sample; the third input unit is used to input the second image sample and the fourth image sample into the discriminator network in the processing model to obtain the discrimination result; the processing unit is used to obtain the target loss value of the processing model based on the second image sample, the third image sample, the fourth image sample and the discrimination result; and the update unit is used to update the network parameters of the processing model based on the target loss value.
[0151] In the above embodiments of this application, the processing unit includes: a first input subunit, a second input subunit, a third input subunit, a processing subunit, and an acquisition subunit.
[0152] The first input subunit is used to input the second image sample and the fourth image sample into the minimum absolute value loss function to obtain the first loss value; the second input subunit is used to downsample the second image sample and input the processed image sample and the third image sample into the minimum absolute value loss function to obtain the second loss value; the third input subunit is used to input the second image sample and the fourth image sample into the perceptual loss function to obtain the third loss value; the processing subunit is used to obtain the fourth loss value based on the discrimination result; and the acquisition subunit is used to obtain the weighted sum of the first loss value, the second loss value, the third loss value and the fourth loss value to obtain the target loss value.
[0153] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0154] Example 6
[0155] According to embodiments of this application, an image processing apparatus for implementing the above-described image processing method is also provided, such as... Figure 9 As shown, the device 900 includes an acquisition module 902 and a processing module 904.
[0156] The acquisition module 902 is used to acquire a first image; the processing module 904 is used to process the first image using a processing model to obtain a second image, wherein the resolution of the second image and the resolution of the first image satisfy a first preset condition; the processing model is used to input the first image into a first generation network to obtain a third image, and input the third image into a second generation network to obtain a second image, wherein the resolution of the third image and the resolution of the first image satisfy a second preset condition.
[0157] It should be noted that the receiving module 902 and processing module 904 mentioned above correspond to steps S502 to S504 in Embodiment 2. The two modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0158] In the above embodiments of this application, the device further includes a training module.
[0159] The acquisition module is further used to acquire a second image sample; the processing module is further used to preprocess the second image sample to obtain a first image sample, wherein the preprocessing includes at least one of the following: downsampling, adding blur information, adding noise information, and image compression; the training module is used to train the processing model using the first image sample and the second image sample.
[0160] In the above embodiments of this application, the training module includes: a first input unit, a second input unit, a third input unit, a processing unit, and an update unit.
[0161] The first input unit is used to input the first image sample into the first generator network to obtain the third image sample; the second input unit is used to input the third image sample into the second generator network to obtain the fourth image sample; the third input unit is used to input the second image sample and the fourth image sample into the discriminator network in the processing model to obtain the discrimination result; the processing unit is used to obtain the target loss value of the processing model based on the second image sample, the third image sample, the fourth image sample and the discrimination result; and the update unit is used to update the network parameters of the processing model based on the target loss value.
[0162] In the above embodiments of this application, the processing unit includes: a first input subunit, a second input subunit, a third input subunit, a processing subunit, and an acquisition subunit.
[0163] The first input subunit is used to input the second image sample and the fourth image sample into the minimum absolute value loss function to obtain the first loss value; the second input subunit is used to downsample the second image sample and input the processed image sample and the third image sample into the minimum absolute value loss function to obtain the second loss value; the third input subunit is used to input the second image sample and the fourth image sample into the perceptual loss function to obtain the third loss value; the processing subunit is used to obtain the fourth loss value based on the discrimination result; and the acquisition subunit is used to obtain the weighted sum of the first loss value, the second loss value, the third loss value and the fourth loss value to obtain the target loss value.
[0164] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0165] Example 7
[0166] According to embodiments of this application, an image processing apparatus for implementing the above-described image processing method is also provided, such as... Figure 10 As shown, the device 1000 includes: a receiving module 1002, an acquisition module 1004, a training module 1006, and an output module 1008.
[0167] The receiving module 1002 is used to receive a model training request; the acquiring module 1004 is used to acquire the training samples and the initial model corresponding to the model training request, wherein the training samples include: a first image sample and a second image sample, and the resolution of the second image sample and the resolution of the first image sample satisfy a first preset condition; the training module 1006 is used to train the initial model using the training samples to obtain a processing model, wherein the processing model is used to input the first image sample into a first generator network to obtain a third image sample, and input the third image sample into a second generator network to obtain a fourth image sample, and the resolution of the third image sample and the resolution of the first image sample satisfy a second preset condition; the output module 1008 is used to output the processing model.
[0168] It should be noted that the receiving module 1002, acquiring module 1004, training module 1006, and output module 1008 mentioned above correspond to steps S602 to S608 in Embodiment 3. The four modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0169] In the above embodiments of this application, the acquisition module includes an acquisition unit and a processing unit.
[0170] The acquisition unit is used to acquire a second image sample; the processing unit is used to preprocess the second image sample to obtain a first image sample, wherein the preprocessing includes at least one of the following: downsampling, adding blur information, adding noise information, and image compression.
[0171] In the above embodiments of this application, the training module includes: an input unit, a processing unit, and an update unit.
[0172] The input unit is used to input the second image sample and the fourth image sample into the discrimination network in the processing model to obtain the discrimination result; the processing unit is used to obtain the target loss value of the processing model based on the second image sample, the third image sample, the fourth image sample and the discrimination result; and the update unit is used to update the network parameters of the processing model based on the target loss value.
[0173] In the above embodiments of this application, the processing unit includes: a first input subunit, a second input subunit, a third input subunit, a processing subunit, and an acquisition subunit.
[0174] The first input subunit is used to input the second image sample and the fourth image sample into the minimum absolute value loss function to obtain the first loss value; the second input subunit is used to downsample the second image sample and input the processed image sample and the third image sample into the minimum absolute value loss function to obtain the second loss value; the third input subunit is used to input the second image sample and the fourth image sample into the perceptual loss function to obtain the third loss value; the processing subunit is used to obtain the fourth loss value based on the discrimination result; and the acquisition subunit is used to obtain the weighted sum of the first loss value, the second loss value, the third loss value and the fourth loss value to obtain the target loss value.
[0175] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0176] Example 8
[0177] According to embodiments of this application, an image processing apparatus for implementing the above-described image processing method is also provided, such as... Figure 11 As shown, the device 1100 includes an acquisition module 1102 and a training module 1104.
[0178] The acquisition module 1102 is used to acquire training samples and an initial model. The training samples include a first image sample and a second image sample. The resolution of the second image sample and the resolution of the first image sample satisfy a first preset condition. The training module 1104 is used to train the initial model using the training samples to obtain a processing model. The processing model is used to input the first image sample into a first generator network to obtain a third image sample, and input the third image sample into a second generator network to obtain a fourth image sample. The resolution of the third image sample and the resolution of the first image sample satisfy a second preset condition.
[0179] It should be noted that the acquisition module 1102 and training module 1104 mentioned above correspond to steps S702 to S704 in Embodiment 4. The two modules and their corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0180] In the above embodiments of this application, the acquisition module includes an acquisition unit and a processing unit.
[0181] The acquisition unit is used to acquire a second image sample; the processing unit is used to preprocess the second image sample to obtain a first image sample, wherein the preprocessing includes at least one of the following: downsampling, adding blur information, adding noise information, and image compression.
[0182] In the above embodiments of this application, the training module includes: an input unit, a processing unit, and an update unit.
[0183] The input unit is used to input the second image sample and the fourth image sample into the discrimination network in the processing model to obtain the discrimination result; the processing unit is used to obtain the target loss value of the processing model based on the second image sample, the third image sample, the fourth image sample and the discrimination result; and the update unit is used to update the network parameters of the processing model based on the target loss value.
[0184] In the above embodiments of this application, the processing unit includes: a first input subunit, a second input subunit, a third input subunit, a processing subunit, and an acquisition subunit.
[0185] The first input subunit is used to input the second image sample and the fourth image sample into the minimum absolute value loss function to obtain the first loss value; the second input subunit is used to downsample the second image sample and input the processed image sample and the third image sample into the minimum absolute value loss function to obtain the second loss value; the third input subunit is used to input the second image sample and the fourth image sample into the perceptual loss function to obtain the third loss value; the processing subunit is used to obtain the fourth loss value based on the discrimination result; and the acquisition subunit is used to obtain the weighted sum of the first loss value, the second loss value, the third loss value and the fourth loss value to obtain the target loss value.
[0186] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0187] Example 9
[0188] According to an embodiment of this application, an image processing system is also provided, comprising:
[0189] Processor. And
[0190] The memory, connected to the processor, is used to provide the processor with instructions to perform the following processing steps: receiving a first image; processing the first image using a processing model to obtain a second image, wherein the resolution of the second image and the resolution of the first image satisfy a first preset condition; the processing model is used to input the first image into a first generation network to obtain a third image, and input the third image into a second generation network to obtain a second image, wherein the resolution of the third image and the resolution of the first image satisfy a second preset condition; and displaying the second image.
[0191] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0192] Example 10
[0193] Embodiments of this application may provide a computer terminal, which may be any computer terminal device in a group of computer terminals. Optionally, in this embodiment, the aforementioned computer terminal may also be replaced by a mobile terminal or other terminal device.
[0194] Optionally, in this embodiment, the computer terminal may be located in at least one of a plurality of network devices in a computer network.
[0195] In this embodiment, the computer terminal described above can execute the program code for the following steps in the image processing method: receiving a first image; processing the first image using a processing model to obtain a second image, wherein the resolution of the second image and the resolution of the first image satisfy a first preset condition, the processing model is used to input the first image into a first generation network to obtain a third image, and input the third image into a second generation network to obtain a second image, wherein the resolution of the third image and the resolution of the first image satisfy a second preset condition; and displaying the second image.
[0196] Optionally, Figure 12 This is a structural block diagram of a computer terminal according to an embodiment of this application. Figure 12 As shown, the computer terminal A may include one or more (only one is shown in the figure) processors 1202 and memory 1204.
[0197] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the image processing method and apparatus in this application embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned image processing method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0198] The processor can call the information and application program stored in the memory through the transmission device to perform the following steps: receiving a first image; processing the first image using a processing model to obtain a second image, wherein the resolution of the second image and the resolution of the first image satisfy a first preset condition; the processing model is used to input the first image into a first generation network to obtain a third image, and input the third image into a second generation network to obtain a second image, wherein the resolution of the third image and the resolution of the first image satisfy a second preset condition; and displaying the second image.
[0199] Optionally, the processor may also execute program code for the following steps: acquiring a second image sample; preprocessing the second image sample to obtain a first image sample, wherein the preprocessing includes at least one of the following: downsampling, adding blur information, adding noise information, and image compression; and training the processing model using the first image sample and the second image sample.
[0200] Optionally, the processor may also execute program code for the following steps: inputting a first image sample into a first generator network to obtain a third image sample; inputting the third image sample into a second generator network to obtain a fourth image sample; inputting the second and fourth image samples into a discriminator network in the processing model to obtain a discrimination result; obtaining a target loss value for the processing model based on the second, third, and fourth image samples and the discrimination result; and updating the network parameters of the processing model based on the target loss value.
[0201] Optionally, the processor may also execute program code for the following steps: inputting the second image sample and the fourth image sample into the minimum absolute value loss function to obtain a first loss value; performing downsampling processing on the second image sample, and inputting the processed image sample and the third image sample into the minimum absolute value loss function to obtain a second loss value; inputting the second image sample and the fourth image sample into the perceptual loss function to obtain a third loss value; obtaining a fourth loss value based on the discrimination result; and obtaining a weighted sum of the first loss value, the second loss value, the third loss value, and the fourth loss value to obtain the target loss value.
[0202] The processor can call the information and application program stored in the memory through the transmission device to perform the following steps: acquire a first image; process the first image using a processing model to obtain a second image, wherein the resolution of the second image and the resolution of the first image satisfy a first preset condition; the processing model is used to input the first image into a first generation network to obtain a third image, and input the third image into a second generation network to obtain a second image, wherein the resolution of the third image and the resolution of the first image satisfy a second preset condition.
[0203] The processor can access information and applications stored in the memory via a transmission device to perform the following steps: receiving a model training request; obtaining training samples and an initial model corresponding to the model training request, wherein the training samples include: a first image sample and a second image sample, the resolution of the second image sample and the resolution of the first image sample satisfying a first preset condition; training the initial model using the training samples to obtain a processing model, wherein the processing model is used to input the first image sample into a first generator network to obtain a third image sample, and input the third image sample into a second generator network to obtain a fourth image sample, the resolution of the third image sample and the resolution of the first image sample satisfying a second preset condition; and outputting the processing model.
[0204] The processor can access information and applications stored in the memory via a transmission device to perform the following steps: acquiring training samples and an initial model, wherein the training samples include: a first image sample and a second image sample, the resolution of the second image sample and the resolution of the first image sample satisfying a first preset condition; training the initial model using the training samples to obtain a processing model, wherein the processing model is used to input the first image sample into a first generator network to obtain a third image sample, and input the third image sample into a second generator network to obtain a fourth image sample, the resolution of the third image sample and the resolution of the first image sample satisfying a second preset condition.
[0205] This application provides an image processing solution. By utilizing a processing model composed of a first generator network and a second generator network to process images, the goal of video quality enhancement is achieved. Since the processing model includes both a first generator network and a second generator network, and the input and output images of the first generator network have the same resolution while the input and output images of the second generator network have different resolutions, the robustness and stability of the video enhancement algorithm in real-world scenarios are increased, improving the visual perception effect. This solves the technical problem of poor robustness and stability of image processing methods in real-world scenarios in related technologies.
[0206] Those skilled in the art will understand that Figure 12 The structure shown is for illustrative purposes only. The computer terminal can also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, and other terminal devices. Figure 12 This does not limit the structure of the aforementioned electronic device. For example, computer terminal A may also include components that are more... Figure 12 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 12 The different configurations shown.
[0207] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0208] Example 11
[0209] Embodiments of this application also provide a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the image processing method provided in the above embodiments.
[0210] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0211] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: receiving a first image; processing the first image using a processing model to obtain a second image, wherein the resolution of the second image satisfies a first preset condition with the resolution of the first image, the processing model is used to input the first image into a first generation network to obtain a third image, and input the third image into a second generation network to obtain a second image, wherein the resolution of the third image satisfies a second preset condition with the resolution of the first image; and displaying the second image.
[0212] Optionally, the storage medium is further configured to store program code for performing the following steps: acquiring a second image sample; preprocessing the second image sample to obtain a first image sample, wherein the preprocessing includes at least one of the following: downsampling, adding blur information, adding noise information, and image compression; and training a processing model using the first image sample and the second image sample.
[0213] Optionally, the storage medium is further configured to store program code for performing the following steps: inputting a first image sample into a first generator network to obtain a third image sample; inputting the third image sample into a second generator network to obtain a fourth image sample; inputting the second and fourth image samples into a discriminator network in the processing model to obtain a discrimination result; obtaining a target loss value for the processing model based on the second, third, and fourth image samples and the discrimination result; and updating the network parameters of the processing model based on the target loss value.
[0214] Optionally, the storage medium is further configured to store program code for performing the following steps: inputting the second image sample and the fourth image sample into the minimum absolute value loss function to obtain a first loss value; downsampling the second image sample, and inputting the processed image sample and the third image sample into the minimum absolute value loss function to obtain a second loss value; inputting the second image sample and the fourth image sample into the perceptual loss function to obtain a third loss value; obtaining a fourth loss value based on the discrimination result; and obtaining a weighted sum of the first loss value, the second loss value, the third loss value, and the fourth loss value to obtain a target loss value.
[0215] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: acquiring a first image; processing the first image using a processing model to obtain a second image, wherein the resolution of the second image satisfies a first preset condition with the resolution of the first image, and the processing model is used to input the first image into a first generation network to obtain a third image, and input the third image into a second generation network to obtain a second image, wherein the resolution of the third image satisfies a second preset condition with the resolution of the first image.
[0216] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: receiving a model training request; obtaining training samples and an initial model corresponding to the model training request, wherein the training samples include: a first image sample and a second image sample, the resolution of the second image sample and the resolution of the first image sample satisfying a first preset condition; training the initial model using the training samples to obtain a processing model, wherein the processing model is used to input the first image sample into a first generator network to obtain a third image sample, and input the third image sample into a second generator network to obtain a fourth image sample, the resolution of the third image sample and the resolution of the first image sample satisfying a second preset condition; and outputting the processing model.
[0217] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining training samples and an initial model, wherein the training samples include: a first image sample and a second image sample, the resolution of the second image sample and the resolution of the first image sample satisfying a first preset condition; training the initial model using the training samples to obtain a processing model, wherein the processing model is used to input the first image sample into a first generator network to obtain a third image sample, and input the third image sample into a second generator network to obtain a fourth image sample, the resolution of the third image sample and the resolution of the first image sample satisfying a second preset condition.
[0218] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0219] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0220] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0221] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0222] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0223] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0224] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. An image processing method comprising: receiving a first image; processing the first image using a processing model to obtain a second image, wherein a resolution of the second image and a resolution of the first image satisfy a first preset condition, the processing model is configured to input the first image into a first generative network to obtain a third image, and input the third image into a second generative network to obtain the second image, a resolution of the third image and the resolution of the first image satisfy a second preset condition, the first preset condition is configured to indicate that the resolution of the second image is different from the resolution of the first image, and a quality of the second image is higher than the first image, and the second preset condition is configured to indicate that the resolution of the third image is the same as the resolution of the first image, and the quality of the third image is higher than the first image; displaying the second image; the method further comprises: during training of the processing model, updating network parameters in the processing model based on a target loss value, the target loss value is determined by a content loss value, a perception loss value and an adversarial loss value, wherein the content loss value is configured to indicate a loss value for constraining a fourth image sample and a third image sample to be similar in content to a second image sample, the perception loss value is configured to indicate a loss value for constraining the fourth image sample and the second image sample to be similar in perception, and the third image sample is obtained by inputting a preprocessed second image sample into the first generative network, and the fourth image sample is obtained by inputting the third image sample into the second generative network.
2. The method of claim 1, wherein, the method further comprises: obtaining the second image sample; preprocessing the second image sample to obtain a first image sample, wherein the preprocessing comprises at least one of the following: down-sampling, increasing blur information, increasing noise information and image compression; training the processing model using the first image sample and the second image sample.
3. The method of claim 2, wherein, training the processing model using the first image sample and the second image sample comprises: inputting the first image sample into the first generative network to obtain the third image sample; inputting the second image sample and the fourth image sample into a discriminative network in the processing model to obtain a discriminative result; obtaining a target loss value of the processing model based on the second image sample, the third image sample, the fourth image sample and the discriminative result; updating network parameters of the processing model based on the target loss value.
4. The method of claim 3, wherein, obtaining a target loss value of the processing model based on the second image sample, the third image sample, the fourth image sample and the discriminative result comprises: inputting the second image sample and the fourth image sample into a least absolute value loss function to obtain a first loss value; performing down-sampling processing on the second image sample, and inputting the processed image sample and the third image sample into a least absolute value loss function to obtain a second loss value; inputting the second image sample and the fourth image sample into a perception loss function to obtain a third loss value; obtaining a fourth loss value based on the discrimination result; obtaining a weighted sum of the first loss value, the second loss value, the third loss value and the fourth loss value to obtain the target loss value.
5. The method of claim 1, wherein, The first preset condition that the resolution of the second image and the resolution of the first image satisfy includes that a difference between the resolution of the second image and the resolution of the first image is greater than a first preset value.
6. The method of claim 1, wherein, The second preset condition that the resolution of the third image and the resolution of the first image satisfy includes that a difference between the resolution of the third image and the resolution of the first image is less than a second preset value.
7. An image processing method, comprising: obtaining a first image; processing the first image by using a processing model to obtain a second image, wherein a resolution of the second image and a resolution of the first image satisfy a first preset condition, the processing model is used for inputting the first image into a first generation network to obtain a third image, and inputting the third image into a second generation network to obtain the second image, a resolution of the third image and a resolution of the first image satisfy a second preset condition, the first preset condition is used for indicating that the resolution of the second image is different from the resolution of the first image, and a quality of the second image is higher than that of the first image, and the second preset condition is used for indicating that the resolution of the third image is the same as the resolution of the first image, and a quality of the third image is higher than that of the first image; The method further comprises: in the process of training the processing model, updating network parameters in the processing model based on a target loss value, the target loss value being determined by a content loss value, a perception loss value and an adversarial loss value, wherein the content loss value is used to indicate a loss value for constraining a fourth image sample and a third image sample generated to be similar in content to a second image sample, the perception loss value is used to indicate a loss value for constraining the fourth image sample and the second image sample to be similar in perception, the third image sample being obtained by inputting a preprocessed second image sample into the first generation network, and the fourth image sample being obtained by inputting the third image sample into the second generation network.
8. The method of claim 7, wherein, The method further comprises: obtaining the second image sample; preprocessing the second image sample to obtain a first image sample, wherein the preprocessing includes at least one of the following: downsampling, adding blur information, adding noise information and image compression; training the processing model by using the first image sample and the second image sample.
9. The method of claim 8, wherein, Training the processing model by using the first image sample and the second image sample includes: inputting the first image sample into the first generation network to obtain the third image sample; inputting the second image sample and the fourth image sample into a discrimination network in the processing model to obtain a discrimination result; obtaining a target loss value of the processing model based on the second image sample, the third image sample, the fourth image sample and the discrimination result; updating network parameters of the processing model based on the target loss value.
10. An image processing method, comprising: receiving a model training request; obtaining training samples and an initial model corresponding to the model training request, wherein the training samples comprise a first image sample and a second image sample, and a resolution of the second image sample satisfies a first preset condition with respect to a resolution of the first image sample; training the initial model by using the training samples to obtain a processing model, wherein the processing model is configured to input the first image sample into a first generation network to obtain a third image sample, and input the third image sample into a second generation network to obtain a fourth image sample, a resolution of the third image sample satisfies a second preset condition with respect to the resolution of the first image sample, the first preset condition is configured to indicate that a resolution of the second image is different from a resolution of the first image, and a quality of the second image is higher than that of the first image, and the second preset condition is configured to indicate that the resolution of the third image is the same as the resolution of the first image, and a quality of the third image is higher than that of the first image; outputting the processing model; the method further comprises: updating network parameters in the processing model based on a target loss value during training of the processing model, the target loss value is determined by a content loss value, a perception loss value and an adversarial loss value, wherein the content loss value is configured to indicate a loss value for constraining the generated fourth image sample and the third image sample to be similar in content to the second image sample, the perception loss value is configured to indicate a loss value for constraining the fourth image sample and the second image sample to be similar in perception, the third image sample is obtained by inputting a preprocessed second image sample into the first generation network, and the fourth image sample is obtained by inputting the third image sample into the second generation network.
11. The method of claim 10, wherein, obtaining the training samples corresponding to the model training request comprises: obtaining the second image sample; preprocessing the second image sample to obtain the first image sample, wherein the preprocessing comprises at least one of downsampling, adding blur information, adding noise information and image compression.
12. The method of claim 10, wherein, training the initial model by using the training samples comprises: inputting the second image sample and the fourth image sample into a discrimination network in the processing model to obtain a discrimination result; obtaining a target loss value of the processing model based on the second image sample, the third image sample, the fourth image sample and the discrimination result; updating network parameters of the processing model based on the target loss value.
13. An image processing method, comprising: obtaining training samples and an initial model, wherein the training samples comprise a first image sample and a second image sample, and a resolution of the second image sample satisfies a first preset condition with respect to a resolution of the first image sample; The initial model is trained by using the training sample, to obtain a processing model, wherein the processing model is used to input the first image sample into a first generation network to obtain a third image sample, and input the third image sample into a second generation network to obtain a fourth image sample, the resolution of the third image sample meets a second preset condition with the resolution of the first image sample, the first preset condition is used to represent that the resolution of the second image is different from the resolution of the first image, the quality of the second image is higher than that of the first image, and the second preset condition is used to represent that the resolution of the third image is the same as that of the first image, and the quality of the third image is higher than that of the first image. The method further includes: during the training of the processing model, updating network parameters in the processing model based on a target loss value, the target loss value being determined by a content loss value, a perception loss value and an adversarial loss value, wherein the content loss value is used to represent a loss value for constraining the generated fourth image sample and the third image sample to be similar in content to the second image sample, the perception loss value is used to represent a loss value for constraining the fourth image sample and the second image sample to be similar in perception, the third image sample is obtained by inputting the preprocessed second image sample into the first generation network, and the fourth image sample is obtained by inputting the third image sample into the second generation network.
14. The method of claim 13, wherein, The training sample is obtained by: obtaining the second image sample; preprocessing the second image sample to obtain a first image sample, wherein the preprocessing includes at least one of the following: down-sampling, adding blur information, adding noise information and image compression.
15. The method of claim 13, wherein, The initial model is trained by using the training sample, to obtain a processing model, wherein the processing model is used to input the first image sample into a first generation network to obtain a third image sample, and input the third image sample into a second generation network to obtain a fourth image sample, the resolution of the third image sample meets a second preset condition with the resolution of the first image sample, the first preset condition is used to represent that the resolution of the second image is different from the resolution of the first image, the quality of the second image is higher than that of the first image, and the second preset condition is used to represent that the resolution of the third image is the same as that of the first image, and the quality of the third image is higher than that of the first image.
16. An image processing apparatus, comprising: a receiving module configured to receive a first image; a processing module configured to process the first image by using a processing model to obtain a second image, wherein the resolution of the second image meets a first preset condition with the resolution of the first image, the processing model is used to input the first image into a first generation network to obtain a third image, and input the third image into a second generation network to obtain the second image, the resolution of the third image meets a second preset condition with the resolution of the first image, the first preset condition is used to represent that the resolution of the second image is different from the resolution of the first image, the quality of the second image is higher than that of the first image, and the second preset condition is used to represent that the resolution of the third image is the same as that of the first image, and the quality of the third image is higher than that of the first image. a display module configured to display the second image; the device is further configured to update network parameters in the processing model based on a target loss value during training of the processing model, the target loss value being determined based on a content loss value, a perception loss value, and an adversarial loss value, wherein the content loss value is used to represent a loss value that constrains a fourth image sample and a third image sample generated to be similar in content to a second image sample, the perception loss value is used to represent a loss value that constrains the fourth image sample and the second image sample to be similar in perception, the third image sample is obtained by inputting a preprocessed second image sample into the first generation network, and the fourth image sample is obtained by inputting the third image sample into the second generation network.
17. An image processing device, comprising: an acquisition module configured to acquire a first image; a processing module configured to process the first image using a processing model to obtain a second image, wherein a resolution of the second image and a resolution of the first image satisfy a first preset condition, the processing model is configured to input the first image into a first generation network to obtain a third image, and input the third image into a second generation network to obtain the second image, a resolution of the third image and a resolution of the first image satisfy a second preset condition, the first preset condition is used to represent that the resolution of the second image is different from the resolution of the first image, and a quality of the second image is higher than a quality of the first image, and the second preset condition is used to represent that the resolution of the third image is the same as the resolution of the first image, and the quality of the third image is higher than the quality of the first image; the device is further configured to update network parameters in the processing model based on a target loss value during training of the processing model, the target loss value being determined based on a content loss value, a perception loss value, and an adversarial loss value, wherein the content loss value is used to represent a loss value that constrains a fourth image sample and a third image sample generated to be similar in content to a second image sample, the perception loss value is used to represent a loss value that constrains the fourth image sample and the second image sample to be similar in perception, the third image sample is obtained by inputting a preprocessed second image sample into the first generation network, and the fourth image sample is obtained by inputting the third image sample into the second generation network.
18. An image processing device, comprising: a receiving module configured to receive a model training request; an acquisition module configured to acquire a training sample and an initial model corresponding to the model training request, wherein the training sample comprises a first image sample and a second image sample, and a resolution of the second image sample and a resolution of the first image sample satisfy a first preset condition. a training module, configured to train the initial model by using the training sample to obtain a processing model, wherein the processing model is configured to input the first image sample into a first generation network to obtain a third image sample, and input the third image sample into a second generation network to obtain a fourth image sample, the resolution of the third image sample and the resolution of the first image sample satisfy a second preset condition, the first preset condition is used to indicate that the resolution of the second image is different from the resolution of the first image, the quality of the second image is higher than that of the first image, and the second preset condition is used to indicate that the resolution of the third image is the same as that of the first image, and the quality of the third image is higher than that of the first image; an output module, configured to output the processing model; in the process of training the processing model, the network parameters in the processing model are updated based on a target loss value, the target loss value is determined by a content loss value, a perception loss value and an adversarial loss value, wherein the content loss value is used to indicate a loss value for constraining the generated fourth image sample and the third image sample to be similar in content to the second image sample, the perception loss value is used to indicate a loss value for constraining the fourth image sample and the second image sample to be similar in perception, the third image sample is obtained by inputting the preprocessed second image sample into the first generation network, and the fourth image sample is obtained by inputting the third image sample into the second generation network.
19. An image processing apparatus, comprising: an acquisition module, configured to acquire a training sample and an initial model, wherein the training sample comprises a first image sample and a second image sample, and the resolution of the second image sample and the resolution of the first image sample satisfy a first preset condition; a training module, configured to train the initial model by using the training sample to obtain a processing model, wherein the processing model is configured to input the first image sample into a first generation network to obtain a third image sample, and input the third image sample into a second generation network to obtain a fourth image sample, the resolution of the third image sample and the resolution of the first image sample satisfy a second preset condition, the first preset condition is used to indicate that the resolution of the second image is different from the resolution of the first image, the quality of the second image is higher than that of the first image, and the second preset condition is used to indicate that the resolution of the third image is the same as that of the first image, and the quality of the third image is higher than that of the first image; The device is further configured to update network parameters in the processing model based on a target loss value during training of the processing model, the target loss value being determined based on a content loss value, a perception loss value, and an adversarial loss value, wherein the content loss value is indicative of a loss value constraining a fourth image sample and a third image sample generated to be similar in content to a second image sample, the perception loss value is indicative of a loss value constraining the fourth image sample and the second image sample to be similar in perception, the third image sample being obtained by inputting a preprocessed second image sample into the first generative network, and the fourth image sample being obtained by inputting the third image sample into the second generative network.
20. A computer readable storage medium, the computer readable storage medium comprising a stored program, wherein, The program is configured to control a device in which the computer readable storage medium is located to perform the image processing method of any one of claims 1-15 when the program is running.
21. A computer terminal comprising a memory and a processor, the processor being arranged to run a program stored in the memory, wherein, The program is configured to perform the image processing method of any one of claims 1-15 when the program is running.
22. An image processing system comprising: a processor; and a memory connected to the processor, configured to provide the processor with instructions to process the following processing steps: receiving a first image; processing the first image using a processing model to obtain a second image, wherein a resolution of the second image and a resolution of the first image satisfy a first preset condition, the processing model is configured to input the first image into a first generative network to obtain a third image, and input the third image into a second generative network to obtain the second image, a resolution of the third image and a resolution of the first image satisfy a second preset condition, the first preset condition is configured to indicate that the resolution of the second image is different from the resolution of the first image, and a quality of the second image is higher than that of the first image, the second preset condition is configured to indicate that the resolution of the third image is the same as the resolution of the first image, and a quality of the third image is higher than that of the first image; displaying the second image; and updating network parameters in the processing model based on a target loss value during training of the processing model, the target loss value being determined based on a content loss value, a perception loss value, and an adversarial loss value, wherein the content loss value is indicative of a loss value constraining a fourth image sample and a third image sample generated to be similar in content to a second image sample, the perception loss value is indicative of a loss value constraining the fourth image sample and the second image sample to be similar in perception, the third image sample being obtained by inputting a preprocessed second image sample into the first generative network, and the fourth image sample being obtained by inputting the third image sample into the second generative network.
Citation Information
Patent Citations
Method of model training, image transmitting, and image processing, and related device and equipment
CN109525859A
Super-resolution reconstruction method, device and equipment apparatus and storage medium
CN111340711A