Image processing method, system, storage medium and terminal device

By using a hybrid degradation model, combining fixed and learnable degradation operators, and expanding the degradation processing space, the problem that training sample pairs in existing technologies cannot simulate actual degradation is solved, and more accurate super-resolution image processing is achieved.

CN115115510BActive Publication Date: 2025-10-03TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210614074.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-10-03
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

In existing super-resolution image acquisition methods, the degradation processing of training sample pairs cannot fully simulate the actual degradation process, resulting in poor neural network performance.

Method used

A hybrid degradation model is adopted, which combines a first degradation submodule with a fixed degradation operator and a second degradation submodule with a learnable degradation operator to expand the degradation processing space and train a super-resolution model.

Benefits of technology

The accuracy and effect of the super-resolution model are improved, which can better cover the real degradation and generate more accurate high-definition images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115115510B_ABST
    Figure CN115115510B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention discloses an image processing method, system, storage medium and terminal device, which are applied to the field of information processing technology based on artificial intelligence. In the process of training a super-resolution model, the image processing system can first determine a first super-resolution model, and determine a first training sample pair based on the first super-resolution model to train a hybrid degradation model, and then determine a second training sample pair based on the hybrid degradation model to adjust the first super-resolution model to obtain a second super-resolution model. The hybrid degradation model used in this process includes a first degradation submodule with a fixed degradation operator and a second degradation submodule with a learnable degradation operator, which can synthesize real degradation that cannot be simulated by the traditional fixed degradation operator in the first degradation submodule, so that the degradation processing space of the overall hybrid degradation model is greatly expanded, which can cover more real degradations, thereby making the adjusted second super-resolution model more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information processing technology based on artificial intelligence, and in particular to an image processing method, system, storage medium and terminal device. Background Art

[0002] In the real world, low-resolution images are affected by various degradation processes (such as blur, noise, and compression). Moreover, most of these degradations are complex and unknown, which increases the complexity and difficulty of image processing when processing low-resolution images to obtain corresponding super-resolution images.

[0003] An existing super-resolution image acquisition method mainly uses a trained neural network to obtain a super-resolution image of any image. The training process of the neural network needs to be based on training sample pairs, where the training sample pairs include a high-definition image and its corresponding low-definition image. In order to obtain training samples, traditional degradation methods are generally used to degrade any high-definition image to obtain a low-definition image to form a training sample pair.

[0004] In this process, the degradation processing used when obtaining the training sample pairs cannot fully simulate the actual degradation process, so the effect of the neural network obtained based on the training sample pairs is not very good. Summary of the Invention

[0005] The embodiments of the present invention provide an image processing method, system, storage medium and terminal device, which improve the accuracy of a super-resolution model.

[0006] An embodiment of the present invention provides an image processing method, including:

[0007] Determining a first super-resolution model, where the first super-resolution model is used to obtain a high-resolution image corresponding to any low-resolution image;

[0008] Acquire a plurality of high-definition sample images according to the first super-resolution model and a plurality of low-definition sample images;

[0009] Determining a first training sample pair according to the low-definition sample image and the high-definition sample image, and training a hybrid degradation model according to the first training sample pair, wherein the hybrid degradation model includes a first degradation submodule having a fixed degradation operator and a second degradation submodule having a learnable degradation operator;

[0010] determining a second training sample pair according to the hybrid degradation model;

[0011] The first super-resolution model is adjusted according to the second training sample pair to obtain a second super-resolution model.

[0012] Another embodiment of the present invention provides an image processing system, including:

[0013] a model determining unit, configured to determine a first super-resolution model, wherein the first super-resolution model is used to obtain a high-resolution image corresponding to any low-resolution image;

[0014] a high-definition acquisition unit, configured to acquire a plurality of high-definition sample images according to the first super-resolution model and a plurality of low-definition sample images;

[0015] a degradation training unit, configured to determine a first training sample pair based on the low-definition sample image and the high-definition sample image, and train a hybrid degradation model based on the first training sample pair, wherein the hybrid degradation model includes a first degradation submodule having a fixed degradation operator and a second degradation submodule having a learnable degradation operator;

[0016] A training pair unit, configured to determine a second training sample pair according to the mixed degradation model;

[0017] A model adjustment unit is used to adjust the first super-resolution model according to the second training sample pair to obtain a second super-resolution model.

[0018] Another aspect of the embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a plurality of computer programs, wherein the computer programs are suitable for being loaded by a processor and executing the image processing method as described in the first aspect of the embodiment of the present invention.

[0019] Another aspect of the present invention provides a terminal device, including a processor and a memory;

[0020] The memory is used to store multiple computer programs, and the computer programs are used to be loaded by the processor and execute the image processing method as described in one aspect of an embodiment of the present invention; the processor is used to implement each computer program in the multiple computer programs.

[0021] As can be seen, in the method of this embodiment, during the super-resolution model training process, the image processing system can first determine a first super-resolution model, and then determine a first training sample pair based on the first super-resolution model to train a hybrid degradation model. Then, based on the hybrid degradation model, a second training sample pair is determined to adjust the first super-resolution model to obtain a second super-resolution model. The hybrid degradation model used in this process includes a first degradation submodule with a fixed degradation operator and a second degradation submodule with a learnable degradation operator. This allows the hybrid degradation model to synthesize real-world degradations that cannot be simulated by the traditional fixed degradation operator in the first degradation submodule. This significantly expands the degradation processing space of the overall hybrid degradation model, allowing it to cover more real-world degradations, thereby making the adjusted second super-resolution model more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0023] Figure 1 is a schematic diagram of an image processing method provided by an embodiment of the present invention;

[0024] Figure 2 is a flowchart of an image processing method provided in one embodiment of the present invention;

[0025] Figure 3 is a flowchart of training a first super-resolution model in one embodiment of the present invention;

[0026] Figure 4a is a schematic diagram of a circulation model in one embodiment of the present invention;

[0027] Figure 4b Schematic diagram of a cyclic model acquiring a high-definition image corresponding to a video frame at any moment in one embodiment of the present invention;

[0028] Figure 5 is a schematic diagram of a mixed degradation model in one embodiment of the present invention;

[0029] Figure 6 This is a flowchart of a super-resolution model training method provided by an application embodiment of the present invention;

[0030] Figure 7 is a schematic diagram of a distributed system to which an image processing method in another application embodiment of the present invention is applied;

[0031] Figure 8 is a schematic diagram of a block structure in another application embodiment of the present invention;

[0032] Figure 9 is a schematic diagram of the logical structure of an image processing system provided by an embodiment of the present invention;

[0033] Figure 10 This is a schematic diagram of the logical structure of a terminal device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0035] The terms "first", "second", "third", "fourth", etc. (if any) in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the invention described herein can, for example, be implemented in orders other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or apparatus that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or apparatus.

[0036] The embodiment of the present invention provides an image processing method, which mainly obtains a super-resolution model by training an image processing system. The super-resolution model is used to obtain a high-resolution image corresponding to a low-resolution image. Specifically, Figure 1 As shown in the figure, the image processing system can train the super-resolution model through the following steps:

[0037] Determine a first super-resolution model, where the first super-resolution model is used to obtain a high-resolution image corresponding to any low-resolution image; obtain multiple high-definition sample images based on the first super-resolution model and the multiple low-definition sample images; determine a first training sample pair based on the low-definition sample images and the high-definition sample images, and train a hybrid degradation model based on the first training sample pair, where the hybrid degradation model includes a first degradation submodule having a fixed degradation operator and a second degradation submodule having a learnable degradation operator; determine a second training sample pair based on the hybrid degradation model; and adjust the first super-resolution model based on the second training sample pair to obtain a second super-resolution model.

[0038] The hybrid degradation model, the first super-resolution model, and the second super-resolution model are all machine learning models based on artificial intelligence. Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI is the study of the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0039] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0040] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning through demonstration.

[0041] In this way, the hybrid degradation model used in the above-mentioned training process to obtain the second super-resolution model includes: a first degradation submodule with a fixed degradation operator and a second degradation submodule with a learnable degradation operator. This allows the hybrid degradation model to synthesize real degradations that cannot be simulated by the traditional fixed degradation operator in the first degradation submodule. This greatly expands the degradation processing space of the overall hybrid degradation model, can cover more real degradations, and thus makes the adjusted second super-resolution model more accurate.

[0042] An embodiment of the present invention provides an image processing method, which is mainly a method executed by an image processing system. The flow chart is as follows: Figure 2 Shown, including:

[0043] Step 101: Determine a first super-resolution model, where the first super-resolution model is used to obtain a high-resolution image corresponding to any low-resolution image.

[0044] Specifically, if Figure 3 As shown, the image processing system can train the first super-resolution model through the following steps:

[0045] A. Determine a third training sample pair according to the explicit degradation model. The third training sample pair includes multiple sample pairs. Each sample pair includes a low-definition image and its corresponding high-definition image.

[0046] The explicit degradation model primarily consists of traditional fixed degradation operators, such as blur, noise, downsampling, and compression. These fixed degradation operators lack learning capabilities. The explicit degradation model can be used to calculate fixed degradation operators for any HD image to obtain a corresponding low-resolution image, thus forming a sample pair with the low-resolution image and its corresponding HD image.

[0047] B. Training the first super-resolution model based on the third training sample. Specifically, the first super-resolution model can be trained by the following steps:

[0048] B1. Determine the initial super-resolution model.

[0049] It is understood that when determining the initial super-resolution model, the image processing system also determines the initial values ​​of the parameters of the multi-layer structure and each layer of the super-resolution model. The parameters of the initial super-resolution model refer to fixed parameters used in the calculation of each layer of the super-resolution model that do not require constant re-assignment, such as parameter scale, number of network layers, and user vector length.

[0050] Specifically, if Figure 4a As shown, the super-resolution initial model may include the following structure: a loop module 10, which is used to calculate the LR of the video frame at any moment. t , the previous moment video frame LR t-1 , the next moment video frame LR t+1 , the previous moment video frame LR obtained by the loop module 10 t-1 Corresponding high-definition image SR t-1 The loop module 10 obtains the video frame LR at the previous moment t-1 Supports high-definition image SR t-1 Hidden state information at any moment, get the video frame LR t Corresponding high-definition image SR tIn a specific implementation, the loop module 10 may include: a feature extraction submodule 110, multiple levels of residual submodules (including a first-level residual submodule 111 and a second-level residual submodule 112), and an output submodule 113, wherein:

[0051] The feature extraction submodule 110 is used to extract feature information of a video frame at any moment, a video frame at the previous moment, a video frame at the next moment, a high-definition image corresponding to the video frame at the previous moment obtained by the loop module 10, and hidden state information when the loop module 10 obtains the high-definition image corresponding to the video frame at the previous moment.

[0052] The residual submodule 111 of the first level among the multiple levels is used to transform the feature information extracted by the feature extraction submodule 110, and the residual submodule 112 of the second level among the multiple levels is used to transform the feature information obtained by the residual submodule 111 of the first level or another residual submodule 112 of the second level.

[0053] The output submodule 113 is configured to output a high-definition image corresponding to a video frame at any moment based on feature information obtained by transformation of the residual submodules at multiple levels.

[0054] It should be noted that, in this embodiment, the residual submodule (ResBlk) of each level may include multiple residual submodules connected in series, so that the residual submodule 112 of the second level will further transform the feature information obtained by transforming a residual submodule connected in series in other levels.

[0055] B2. Obtain high-definition images corresponding to each low-definition image in the third training sample pair through the super-resolution initial model.

[0056] It should be noted that, in this embodiment, each low-resolution image in the third training sample pair obtained above can be a plurality of video frames at consecutive moments contained in a video file. In this way, when obtaining a high-definition image through the super-resolution initial model determined in step B1 above, as shown in FIG. Figure 4a and Figure 4b Shown, including:

[0057] The feature extraction submodule 110 in the loop module 10 first extracts the low-definition image included in the third training sample pair, and the video frame LR at any moment t , the previous moment video frame LR t-1 , the next moment video frame LR t+1 , the previous moment video frame LR obtained before the loop module 10 t-1 Corresponding high-definition image SR t-1 The loop module 10 obtains the video frame LR at the previous moment t-1 Supports high-definition image SRt-1 The characteristic information of the hidden state information at that time.

[0058] After the feature extraction submodule 110 is transformed by the first-level multiple residual submodules 111 and the second-level multiple residual submodules 112, the output submodule 113 integrates the feature information obtained by the transformation of the residual submodules at each level, and then outputs the video frame LR at any moment. t Corresponding high-definition image SR t .

[0059] It should be noted that if Figure 4b The three loop modules 10 shown do not mean that the aforementioned super-resolution initial model actually includes three loop modules 10 , but rather the loop modules 10 involved in the process of acquiring the high-definition image corresponding to the video frame at each moment.

[0060] B3. Adjust parameter values ​​in the initial super-resolution model based on the high-definition image obtained by the initial super-resolution model and the corresponding high-definition image in the third training sample pair to obtain a first super-resolution model.

[0061] Specifically, the image processing system will first calculate the loss function related to the super-resolution initial model based on the high-definition image obtained by the super-resolution initial model in the above step B2 and the corresponding high-definition image in the third training sample pair. The loss function is used to indicate the error between the high-definition image obtained by the super-resolution initial model and the actual high-definition image of each low-definition image (obtained in the third training sample pair), such as the cross-entropy loss function.

[0062] The training process of the first super-resolution model is to minimize the value of the above-mentioned error. This training process continuously optimizes the parameter values ​​of the parameters in the super-resolution initial model determined in step B1 above through a series of mathematical optimization methods such as backpropagation derivation and gradient descent, and minimizes the calculated value of the above-mentioned loss function. Specifically, when the calculated value of the loss function is large, such as greater than a preset value, it is necessary to change the parameter value, such as reducing the weight value of a certain neuron connection, so that the function value of the loss function calculated according to the adjusted parameter value is reduced.

[0063] It should be noted that the above steps B2 to B3 are an adjustment of the parameter values ​​in the super-resolution initial model based on the high-definition image obtained by the super-resolution initial model. In actual applications, it is necessary to continuously loop through the above steps B2 to B3 until the adjustment of the parameter values ​​meets certain stop conditions.

[0064] Therefore, after executing steps B2 to B3 of the above embodiment, the image processing system needs to determine whether the current parameter adjustment satisfies a preset stopping condition. If so, the process ends and the parameter values ​​adjusted in step B3 are used as the parameter values ​​of the trained first super-resolution model. If not, the system returns to executing steps B2 to B3 for the initial super-resolution model after adjusting the parameter values. The preset stopping condition includes, but is not limited to, any of the following conditions: the difference between the currently adjusted parameter value and the last adjusted parameter value is less than a threshold, i.e., the adjusted parameter value has reached convergence; and the number of parameter adjustments equals a preset number of times.

[0065] Step 102: obtaining high-definition sample images corresponding to a plurality of low-definition sample images according to a first super-resolution model.

[0066] Specifically, in the process of obtaining high-definition sample images, for any low-definition sample image, multiple scaled low-definition sample images corresponding to any low-definition sample image can be obtained according to multiple scaling factors, and a pre-selected set of high-definition sample images corresponding to the multiple scaled low-definition sample images can be obtained according to the first super-resolution model, and then a high-definition sample image can be selected from the pre-selected set.

[0067] For example, a low-definition sample image is scaled using scaling factors of 1, 0.9, 0.7, 0.5, and 0.3, respectively, to obtain multiple scaled low-definition sample images. These multiple scaled low-definition sample images are then input into a first super-resolution model, and a preselected set of high-definition sample images is output. As the resolution of the input image to the first super-resolution model decreases, artifacts in the output high-definition sample images gradually decrease. However, if the low-definition sample image is scaled significantly (for example, with a scaling factor of 0.3), details or information in the output high-definition sample image may be lost. Therefore, a high-definition sample image outputted for the scaled low-definition sample image corresponding to a scaling factor of 0.5 may be selected, and the scaled low-definition sample image and its corresponding high-definition sample image may be used to form a first training sample pair. This allows for a good balance between artifact elimination and detail loss.

[0068] Step 103: determine a first training sample pair based on the low-definition sample image and the high-definition sample image, and train a hybrid degradation model based on the first training sample pair, wherein the hybrid degradation model includes a first degradation submodule having a fixed degradation operator and a second degradation submodule having a learnable degradation operator.

[0069] Specifically, when determining the first training sample pair, a high-definition sample image selected in the above step and the corresponding scaled low-definition sample image can be used to form any first training sample pair.

[0070] In this embodiment, the first degradation submodule with a fixed degradation operator can be specifically described in the above-mentioned explicit degradation model and will not be repeated here. The second degradation submodule with a learnable degradation operator is mainly composed of some learnable degradation operators. These degradation operators have learning capabilities. The second degradation submodule can perform calculations based on the learned degradation operators for any high-definition image to obtain a corresponding low-definition image.

[0071] The specific method for training the hybrid degradation model may be similar to the method for training the first super-resolution model, which will not be described in detail here. The difference is that the first training sample pair is used in training the hybrid degradation model.

[0072] Specific as Figure 5 As shown, the hybrid degradation model in this embodiment may include multiple degradation units connected in series. Each degradation unit may include a first degradation submodule 20 and a second degradation submodule 21. The second degradation submodule 21 may specifically be a micro-neural network with two to three convolutional layers. It can capture the main features of the feature information obtained by degradation in the first degradation submodule 20 and then integrate these main features into the next first degradation submodule 20. Because the degradation operator in the second degradation submodule 21 is learnable and can synthesize real degradation that cannot be simulated by the traditional degradation operator in the first degradation submodule 20, the degradation processing space of the overall hybrid degradation model is greatly expanded, and can cover more real degradation.

[0073] Step 104: Determine a second training sample pair according to the mixed degradation model.

[0074] Specifically, the hybrid degradation model can perform calculations based on a fixed degradation operator and a learned degradation operator for any high-definition image to obtain a corresponding low-definition image, thereby forming a second training sample pair with the low-definition image and its corresponding high-definition image.

[0075] Step 105: Adjust the first super-resolution model according to the second training sample pair to obtain a second super-resolution model.

[0076] The specific method used in adjusting the first super-resolution model can be similar to steps B2 to B3 in the above-mentioned method for training the first super-resolution model, and will not be repeated here. The difference is that the second training sample pair is used in adjusting the first super-resolution model.

[0077] As can be seen, in the method of this embodiment, during the super-resolution model training process, the image processing system can first determine a first super-resolution model, and then determine a first training sample pair based on the first super-resolution model to train a hybrid degradation model. Then, based on the hybrid degradation model, a second training sample pair is determined to adjust the first super-resolution model to obtain a second super-resolution model. The hybrid degradation model used in this process includes a first degradation submodule with a fixed degradation operator and a second degradation submodule with a learnable degradation operator. This allows the hybrid degradation model to synthesize real-world degradations that cannot be simulated by the traditional fixed degradation operator in the first degradation submodule. This significantly expands the degradation processing space of the overall hybrid degradation model, allowing it to cover more real-world degradations, thereby making the adjusted second super-resolution model more accurate.

[0078] The following uses a specific application example to illustrate the image processing method of an embodiment of the present invention. The method of this embodiment is mainly applied to the scenario of processing comic video frames. The image processing method of this embodiment may include the following two parts:

[0079] (1) Figure 6 As shown, the super-resolution model can be trained through the following steps to obtain the second super-resolution model:

[0080] Step 201: Determine an explicit degradation model, where the explicit degradation model has a traditional fixed degradation operator, and determine a third training sample pair through the explicit degradation model.

[0081] Specifically, the explicit degradation model can perform degradation processing on each HD video frame in the animation HD video based on a fixed degradation operator to obtain the corresponding low-definition image, and each HD video frame and its corresponding low-definition image are used as a third training sample pair.

[0082] Step 202: Determine the super-resolution initial model. The structure of the super-resolution initial model can be as described above. Figure 4a As shown, no further description is given here.

[0083] Step 203: Based on the third training sample pair and the aforementioned initial super-resolution model, a first super-resolution model is trained.

[0084] Step 204: Determine a first training sample pair according to the first super-resolution model.

[0085] Specifically, the low-definition video frames in the low-definition animation video can be scaled based on multiple scaling factors, and then the multiple scaled low-definition sample images obtained are input into the first super-resolution model to obtain a pre-selected set of multiple high-definition sample images, and high-definition sample images with better effects are selected from the pre-selected set, and then the selected high-definition sample images and their corresponding scaled low-definition sample images are used to form a first training sample pair.

[0086] Step 205 : training a hybrid degradation model based on the first training sample pair, where the hybrid degradation model includes a first degradation submodule having a fixed degradation operator and a second degradation submodule having a learnable degradation operator.

[0087] Step 206: Determine a second training sample pair according to the mixed degradation model.

[0088] Specifically, the hybrid degradation model can perform degradation processing on each high-definition video frame in the animation high-definition video based on a fixed degradation operator and a learnable degradation operator to obtain the corresponding low-definition image, and each high-definition video frame and its corresponding low-definition image are used as a second training sample pair.

[0089] Step 207 : Adjust the first super-resolution model according to the second training sample pair to obtain a second super-resolution model, and then preset the first super-resolution model into the image processing system.

[0090] (2) Process any low-definition video to obtain a high-definition video

[0091] Specifically, the image processing system can obtain low-definition video frames at each moment in any low-definition video, call the second super-resolution model preset in the system, and the second super-resolution model can perform calculations based on the low-definition video frames at each moment to obtain high-definition images at each moment.

[0092] In actual application, when the second super-resolution model is trained using the method in this embodiment to obtain a high-definition image, and the low-definition image is processed using existing methods (such as Real-ESRGAN, RealBasicVSR) to obtain a high-definition image, for example, the high-definition image obtained by using the second super-resolution model in this embodiment has a better effect and does not have defects such as ghosting.

[0093] In addition, when two second super-resolution models are trained using the method in this embodiment (super-resolution model 1 and super-resolution model 2 are trained based on the second training sample pairs obtained from the hybrid degradation model with 1 and 3 learnable degradation operators), the low-definition images are processed to obtain high-definition images, and the low-definition images are processed using existing methods (such as Real-ESRGAN and RealBasicVSR) to obtain high-definition images, and the parameters involved in this process are calculated, specifically including: learned parameters (params), runtime, image quality evaluation (Natural image quality evaluator, NIQE) and MANIQA of the obtained high-definition images, as shown in Table 1 below. It can be seen that the time required to obtain high-definition images using the embodiment of the present invention is greatly reduced, and the quality of the obtained high-definition images is high:

[0094]

[0095] The models reported in Table 1 are all trained or fine-tuned on the animation dataset.

[0096] It can be seen that the image processing method in the embodiment of the present invention can achieve the following effects:

[0097] (1) The second training sample pairs generated by the hybrid degradation method are more consistent with the distribution of low-definition animation videos in the real world, making the second super-resolution model trained based on the second training sample pairs more accurate. The hybrid degradation method can be composed of traditional fixed degradation operators (such as blur, noise, downsampling operators) and degradation operators that can be learned from the real world.

[0098] (2) An effective and efficient multi-scale network structure is adopted as the super-resolution model, and the efficiency of the unidirectional recurrent network and the effectiveness of the sliding window are utilized to achieve better super-resolution effects on real-world animation videos.

[0099] The following uses another specific application example to illustrate the image processing method of the present invention. The image processing system in the embodiment of the present invention is mainly a distributed system 100, which may include a client 300 and multiple nodes 200 (any form of computing device connected to the network, such as a server, a user terminal), and the client 300 and the node 200 are connected through network communication.

[0100] Taking the distributed system as the blockchain system as an example, see Figure 7This is a schematic diagram of an optional architecture for a distributed system 100 provided in an embodiment of the present invention, applied to a blockchain system. The system consists of multiple nodes 200 (any type of computing device connected to a network, such as a server or user terminal) and clients 300. The nodes form a peer-to-peer (P2P) network. The P2P protocol is an application layer protocol that runs on top of the Transmission Control Protocol (TCP). In a distributed system, any machine, such as a server or terminal, can join and become a node. Nodes include hardware, middleware, operating system, and application layers.

[0101] See also Figure 7 The functions of each node in the blockchain system shown include:

[0102] 1) Routing: A basic function of a node, used to support communication between nodes.

[0103] In addition to the routing function, nodes can also have the following functions:

[0104] 2) Applications, deployed in the blockchain, implement specific services based on actual business needs, record data related to the implementation of functions to form record data, carry digital signatures in the record data to indicate the source of the task data, and send the record data to other nodes in the blockchain system for other nodes to add the record data to a temporary block when they successfully verify the source and integrity of the record data.

[0105] For example, the services implemented by the application include: code that implements image processing functions, which mainly include:

[0106] Determine a first super-resolution model, where the first super-resolution model is used to obtain a high-resolution image corresponding to any low-resolution image; obtain multiple high-definition sample images based on the first super-resolution model and the multiple low-definition sample images; determine a first training sample pair based on the low-definition sample images and the high-definition sample images, and train a hybrid degradation model based on the first training sample pair, where the hybrid degradation model includes a first degradation submodule having a fixed degradation operator and a second degradation submodule having a learnable degradation operator; determine a second training sample pair based on the hybrid degradation model; and adjust the first super-resolution model based on the second training sample pair to obtain a second super-resolution model.

[0107] 3) Blockchain, including a series of blocks that are connected to each other in the order of their generation. Once a new block is added to the blockchain, it will not be removed. The block records the record data submitted by the nodes in the blockchain system.

[0108] See also Figure 8 This is an optional schematic diagram of the block structure provided by an embodiment of the present invention. Each block includes the hash value of the transaction records stored in the block (the hash value of the current block) and the hash value of the previous block. The blocks are connected by hash values ​​to form a blockchain. In addition, the block may also include information such as the timestamp when the block was generated. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each data block contains relevant information used to verify the validity of the information (anti-counterfeiting) and generate the next block.

[0109] The embodiment of the present invention further provides an image processing system, the structural diagram of which is shown in FIG. Figure 9 Specifically, it may include:

[0110] The model determination unit 30 is configured to determine a first super-resolution model, where the first super-resolution model is used to obtain a high-resolution image corresponding to any low-resolution image.

[0111] Specifically, the model determination unit 30 is specifically used to determine a third training sample pair based on the explicit degradation model; the third training sample pair includes multiple sample pairs, each sample pair includes a low-definition image and its corresponding high-definition image; and the first super-resolution model is trained based on the third training sample pair.

[0112] Among them, when training the first super-resolution model according to the third training sample pair, the model determination unit 30 is specifically used to determine the super-resolution initial model; obtain the high-definition images corresponding to each low-definition image in the third training sample pair through the super-resolution initial model; and adjust the parameter values ​​in the super-resolution initial model according to the high-definition images obtained by the super-resolution initial model and the corresponding high-definition images in the third training sample pair to obtain the first super-resolution model.

[0113] Among them, when determining the super-resolution initial model, the model determination unit 30 is specifically used to determine that the super-resolution initial model includes the following structure: a loop module, wherein; the loop module is used to obtain the high-definition image corresponding to the video frame at any moment based on the video frame at any moment, the video frame at the previous moment, the video frame at the next moment, the high-definition image corresponding to the video frame at the previous moment obtained by the loop module, and the hidden state information of the loop module when obtaining the high-definition image corresponding to the video frame at the previous moment. The loop module includes: a feature extraction submodule, a plurality of levels of residual submodules and an output submodule; the feature extraction submodule is used to extract feature information of a video frame at any moment, a video frame at a previous moment, a video frame at a next moment, a high-definition image corresponding to the video frame at the previous moment obtained by the loop module, and hidden state information of the loop module when obtaining the high-definition image corresponding to the video frame at the previous moment; the residual submodule of the first level among the multiple levels is used to transform the feature information extracted by the feature extraction submodule, and the residual submodule of the second level among the multiple levels is used to transform the feature information obtained by the residual submodule of the first level or another second level; the output submodule is used to output the high-definition image corresponding to the video frame at any moment based on the feature information transformed by the residual submodules of the multiple levels.

[0114] Furthermore, the model determination unit 30 is further configured to stop adjusting the parameter value when the number of adjustments to the parameter value is equal to a preset number, or when the difference between the currently adjusted fixed parameter value and the last adjusted parameter value is less than a threshold.

[0115] The high-definition acquisition unit 31 is configured to acquire a plurality of high-definition sample images according to the first super-resolution model determined by the model determination unit 30 and a plurality of low-definition sample images.

[0116] The high-definition acquisition unit 31 is specifically used to obtain multiple scaled low-definition sample images corresponding to any low-definition sample image according to multiple scaling factors; obtain a pre-selected set of high-definition sample images corresponding to the multiple scaled low-definition sample images according to the first super-resolution model; and select a high-definition sample image from the pre-selected set.

[0117] The degradation training unit 32 is used to determine a first training sample pair based on the low-definition sample image and the high-definition sample image acquired by the high-definition acquisition unit 31, and to train a hybrid degradation model based on the first training sample pair, wherein the hybrid degradation model includes a first degradation submodule having a fixed degradation operator and a second degradation submodule having a learnable degradation operator.

[0118] When determining the first training sample pair, the degradation training unit 32 may use the selected high-definition sample image and the corresponding scaled low-definition sample image to form a first training sample pair.

[0119] The training pair unit 33 is configured to determine a second training sample pair according to the mixed degradation model obtained by the degradation training unit 32 .

[0120] The model adjustment unit 34 is configured to adjust the first super-resolution model according to the second training sample pair determined by the training pair unit 33 to obtain a second super-resolution model.

[0121] As can be seen, during the super-resolution model training process of the image processing system of this embodiment, the model determination unit 30 can first determine a first super-resolution model, and the degradation training unit 32 can determine a first training sample pair based on the first super-resolution model to train a hybrid degradation model. The training pair unit 33 then determines a second training sample pair based on the hybrid degradation model, and the model adjustment unit 34 adjusts the first super-resolution model based on the second training sample pair to obtain a second super-resolution model. The hybrid degradation model used in this process includes a first degradation submodule with a fixed degradation operator and a second degradation submodule with a learnable degradation operator. This allows the hybrid degradation model to synthesize real degradations that cannot be simulated by the traditional fixed degradation operator in the first degradation submodule. This greatly expands the degradation processing space of the overall hybrid degradation model, allowing it to cover more real degradations, thereby making the adjusted second super-resolution model more accurate.

[0122] The embodiment of the present invention further provides a terminal device, the structural diagram of which is shown in FIG. Figure 10 As shown, the terminal device may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPUs) 40 (for example, one or more processors) and a memory 41, and one or more storage media 42 (for example, one or more mass storage devices) storing application programs 421 or data 422. Memory 41 and storage medium 42 may be temporary storage or permanent storage. The program stored in the storage medium 42 may include one or more modules (not shown), each module may include a series of instruction operations in the terminal device. Furthermore, the CPU 40 may be configured to communicate with the storage medium 42 to execute a series of instruction operations in the storage medium 42 on the terminal device.

[0123] Specifically, the application 421 stored in the storage medium 42 includes an image processing application, and the application may include the model determination unit 30, high-definition acquisition unit 31, degradation training unit 32, training pair unit 33, and model adjustment unit 34 of the aforementioned image processing system, which are not described in detail here. Furthermore, the central processing unit 40 may be configured to communicate with the storage medium 42 and execute a series of operations corresponding to the image processing application stored in the storage medium 42 on the terminal device.

[0124] The terminal device may also include one or more power supplies 43, one or more wired or wireless network interfaces 44, one or more input and output interfaces 45, and / or one or more operating systems 423, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0125] The steps performed by the image processing system in the above method embodiment can be based on the Figure 10 The structure of the terminal device shown.

[0126] Furthermore, another aspect of an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a plurality of computer programs, wherein the computer programs are suitable for being loaded by a processor and executing the image processing method executed by the above-mentioned image processing system.

[0127] Another aspect of the present invention provides a terminal device, including a processor and a memory;

[0128] The memory is used to store multiple computer programs, and the computer programs are used to be loaded by the processor and executed by the image processing method executed by the above-mentioned image processing system; the processor is used to implement each computer program in the multiple computer programs.

[0129] In addition, according to one aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described above.

[0130] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.

[0131] The above is a detailed introduction to an image processing method, system, storage medium, and terminal device provided by an embodiment of the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.

Claims

1. An image processing method, characterized in that: include: Determining a first super-resolution model, where the first super-resolution model is used to obtain a high-resolution image corresponding to any low-resolution image; Acquire multiple high-definition sample images based on the first super-resolution model and multiple low-definition sample images, wherein the first super-resolution model is obtained by training an initial super-resolution model, and the initial super-resolution model includes a loop module, wherein the loop module is used to acquire the high-definition image corresponding to the video frame at any moment based on a video frame at any moment, a video frame at a previous moment, a video frame at a next moment, a high-definition image corresponding to the video frame at the previous moment obtained by the loop module, and hidden state information of the loop module when acquiring the high-definition image corresponding to the video frame at the previous moment, wherein the video frame at any moment and the high-definition image corresponding to the video frame at any moment are the same video frame in content but have different resolutions; Determining a first training sample pair according to the low-definition sample image and the high-definition sample image, and training a hybrid degradation model according to the first training sample pair, wherein the hybrid degradation model includes a first degradation submodule having a fixed degradation operator and a second degradation submodule having a learnable degradation operator; determining a second training sample pair according to the hybrid degradation model; The first super-resolution model is adjusted according to the second training sample pair to obtain a second super-resolution model.

2. The method according to claim 1, wherein Determining the first super-resolution model specifically includes: Determining a third training sample pair according to the explicit degradation model; the third training sample pair includes a plurality of sample pairs, each sample pair including a low-definition image and its corresponding high-definition image; The first super-resolution model is trained based on the third training sample pair.

3. The method according to claim 2, wherein The step of training the first super-resolution model according to the third training sample pair specifically includes: Determine the initial super-resolution model; Obtaining high-definition images corresponding to the low-definition images in the third training sample pair through the super-resolution initial model; According to the high-definition image obtained by the super-resolution initial model and the corresponding high-definition image in the third training sample pair, the parameter values ​​in the super-resolution initial model are adjusted to obtain the first super-resolution model.

4. The method according to claim 1, wherein The loop module includes: a feature extraction submodule, a plurality of levels of residual submodules and an output submodule; The feature extraction submodule is used to extract feature information of a video frame at any moment, a video frame at a previous moment, a video frame at a next moment, a high-definition image corresponding to the video frame at the previous moment obtained by the loop module, and hidden state information when the loop module obtains the high-definition image corresponding to the video frame at the previous moment; The residual submodule of the first level among the multiple levels is used to transform the feature information extracted by the feature extraction submodule, and the residual submodule of the second level among the multiple levels is used to transform the feature information obtained by the residual submodule of the first level or another second level; The output submodule is used to output the high-definition image corresponding to the video frame at any moment based on the feature information obtained by the transformation of the residual submodules of the multiple levels.

5. The method according to claim 3, wherein Also includes: When the number of times the parameter value is adjusted is equal to a preset number, or when the difference between the currently adjusted fixed parameter value and the last adjusted parameter value is less than a threshold, the adjustment of the parameter value is stopped.

6. The method according to any one of claims 1 to 5, characterized in that The step of obtaining a plurality of high-definition sample images according to the first super-resolution model and the plurality of low-definition sample images specifically includes: Acquire multiple scaled low-definition sample images corresponding to any low-definition sample image according to multiple scaling factors; Acquire a preselected set of high-definition sample images corresponding to the plurality of scaled low-definition sample images according to the first super-resolution model; A high-definition sample image is selected from the pre-selected set.

7. The method according to claim 6, wherein Determining the first training sample pair based on the low-definition sample image and the high-definition sample image specifically includes: using the selected high-definition sample image and the corresponding scaled low-definition sample image to form a first training sample pair.

8. An image processing system, characterized in that: include: a model determination unit, configured to determine a first super-resolution model, wherein the first super-resolution model is configured to obtain a high-resolution image corresponding to any low-resolution image, the first super-resolution model being trained by an initial super-resolution model, the initial super-resolution model comprising a loop module, the loop module being configured to obtain a high-definition image corresponding to the video frame at any moment based on a video frame at any moment, a video frame at a previous moment, a video frame at a next moment, a high-definition image corresponding to the video frame at the previous moment obtained by the loop module, and hidden state information of the loop module when obtaining the high-definition image corresponding to the video frame at the previous moment, wherein the video frame at any moment and the high-definition image corresponding to the video frame at any moment are the same video frame in content but have different resolutions; a high-definition acquisition unit, configured to acquire a plurality of high-definition sample images according to the first super-resolution model and a plurality of low-definition sample images; a degradation training unit, configured to determine a first training sample pair based on the low-definition sample image and the high-definition sample image, and train a hybrid degradation model based on the first training sample pair, wherein the hybrid degradation model includes a first degradation submodule having a fixed degradation operator and a second degradation submodule having a learnable degradation operator; A training pair unit, configured to determine a second training sample pair according to the mixed degradation model; A model adjustment unit is used to adjust the first super-resolution model according to the second training sample pair to obtain a second super-resolution model.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of computer programs, and the computer programs are suitable for being loaded by a processor and executing the image processing method according to any one of claims 1 to 7.

10. A terminal device, characterized in that: including processor and memory; The memory is used to store a plurality of computer programs, and the computer programs are used to be loaded by the processor and execute the image processing method according to any one of claims 1 to 7; the processor is used to implement each computer program in the plurality of computer programs.

11. A computer program product, comprising computer instructions stored in a computer-readable storage medium; a processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to perform the steps of the image processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image processing apparatus and image processing method

    CN110088799A

  • Video super-resolution method based on convolutional neural network and mixed resolution

    CN110120011A

  • Method and device for acquiring blind super-resolution image, and storage medium

    CN112927137A

  • Online updated image blind super-resolution reconstruction method and device

    CN113487476A