Image quality recovery method and device, storage medium and electronic equipment
Through the stream matching model, the image to be processed is fused with the random noise image, the noise speed is predicted to be denoised, and the image degradation process is reversed, which solves the problem of poor image recovery effect in the prior art and achieves efficient image quality recovery.
Patent Information
- Application Number
- CN202510348872.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, image quality recovery methods such as generating adversarial models and diffusion models have poor results in the denoising process, lack effective guidance or introduction of uncertain factors, resulting in low image recovery effect.
Using the stream matching model, by fusing the to-processed image with the random noise image, injecting the stream matching model as a generation condition, predicting the noise addition speed and denoising, reversing the image degradation process, and recovering high-quality images.
Effectively denoising low-quality images and restoring high-quality images, solving the problem of poor image recovery effect in the prior art and achieving efficient image quality recovery.
Smart Images

Figure CN120374434A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular, to a method, device, storage medium, and electronic device for image quality restoration. Background Art
[0002] In today's era of information explosion, images have become an important carrier for information transmission. However, due to limitations in shooting equipment, transmission processes, or storage conditions, image quality often faces degradation problems.
[0003] Image quality degradation is ultimately due to the noise added to the original image, and there are various types of noise, such as rain, fog, night, motion blur, white noise, etc. Therefore, image quality restoration is essentially a process of denoising the image.
[0004] Then, how to effectively denoise low-quality images to restore them to high-quality images is an urgent problem to be solved. Summary of the Invention
[0005] Embodiments of this specification provide a method, device, storage medium, and electronic device for image quality restoration to partially solve the problems existing in the above-mentioned prior art.
[0006] Embodiments of this specification adopt the following technical solutions:
[0007] A method for image quality restoration provided by this specification, the method includes:
[0008] Obtain an image to be processed and a random noise image;
[0009] Fuse the image to be processed and the random noise image to obtain a fused image;
[0010] Input a standard noise image into a pre-trained flow matching model, and inject the fused image as a generation condition into the flow matching model, so that the flow matching model predicts the noise addition speed corresponding to the standard noise image based on the generation condition, and denoises the standard noise image according to the noise addition speed to obtain a restored image.
[0011] A device for image quality restoration provided by this specification, the device includes:
[0012] An acquisition module, configured to obtain an image to be processed and a random noise image;
[0013] A fusion module, configured to fuse the image to be processed and the random noise image to obtain a fused image;
[0014] A denoising module is configured to input a standard noise image into a pre-trained flow matching model, inject the fused image as a generation condition into the flow matching model, enable the flow matching model to predict the noise addition speed corresponding to the standard noise image based on the generation condition, and denoise the standard noise image according to the noise addition speed to obtain a restored image.
[0015] A computer-readable storage medium provided in this specification stores a computer program, and when the computer program is executed by a processor, the method for image quality restoration described above is implemented.
[0016] An electronic device provided in this specification includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method for image quality restoration described above is implemented.
[0017] At least one of the technical solutions adopted in the embodiments of this specification can achieve the following beneficial effects:
[0018] An embodiment of this specification discloses a method for image quality restoration. The method injects a fused image obtained by fusing a to-be-processed image and a random noise image as a generation condition into a flow matching model, enables the flow matching model to predict the noise addition speed of adding noise to a standard noise image under the guidance of the fused image, and denoises the standard noise image according to the predicted noise addition speed to obtain a restored image, which can effectively denoise a low-quality to-be-processed image and obtain a high-quality image corresponding to the to-be-processed image. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of this specification, form a part of this specification, and the illustrative embodiments and descriptions thereof of this specification are used to explain this specification and do not constitute an improper limitation to this specification. In the drawings:
[0020] Figure 1 is a flowchart of the method for image quality restoration provided by an embodiment of this specification;
[0021] Figure 2 is a schematic diagram of the method for training a flow matching model provided by an embodiment of this specification;
[0022] Figure 3 is a schematic diagram of a device for image quality restoration provided by an embodiment of this specification;
[0023] Figure 4 is a schematic diagram of the structure of an electronic device provided by an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] In the prior art, the denoising of low-quality images is often performed through generative adversarial models or diffusion models to restore high-quality images. However, the generative adversarial model essentially directly models the process of restoring from a low-quality image to a high-quality image onto an end-to-end network, lacking effective guidance in the middle, so the effect of image restoration is poor. Although the diffusion model can use the low-quality image as a generation condition as guidance, uncertain factors are introduced in the denoising process, so the effect of image restoration is not high either.
[0025] In the embodiments of this specification, the degradation process of the image from high quality to low quality is modeled as a number of noise addition processes with a constant noise addition speed. In this way, the degradation process is a reversible process, and the inverse process of the degradation process is the restoration process of restoring the image from low quality to high quality. Therefore, the embodiments of this specification can use a flow matching model to predict the noise addition speed in the degradation process of the image from high quality to low quality, and then denoise the standard noise image according to the noise addition speed to obtain the restored image.
[0026] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this specification.
[0027] The following details the technical solutions provided by each embodiment of this specification in conjunction with the drawings.
[0028] Figure 1 The following is a flowchart of the method for image quality restoration provided by the embodiments of this specification, including the following steps:
[0029] S100: Obtain the image to be processed and a random noise image.
[0030] In the embodiments of this specification, the image to be processed is the low-quality image that needs to perform image quality restoration. The device that executes Figure 1 the method shown to perform image quality restoration on the image to be processed can be any electronic device capable of running a flow matching model, such as a server, etc. The following will be described by taking a server as an example.
[0031] The server needs to first obtain the low-quality image that needs to perform image quality restoration as the image to be processed.
[0032] Since the subsequent flow matching model is used to restore the image quality, in the embodiments of this specification, it is necessary to determine the deterministic degradation path from a high-quality image to a low-quality image, and then reverse this deterministic degradation path to restore the image quality of the low-quality image. That is to say, the degradation path that needs to be determined in the embodiments of this specification must meet two conditions: one is reversibility, and the other is determinacy.
[0033] For reversibility, the process of degrading a high-quality image to a low-quality image is represented as a stochastic process z t , where t is the number of degradation times. That is to say, z0 represents the state where the original high-quality image in the degradation process has not degraded at all. After t times of degradation, this process is represented as z t . Then, it can be expressed by an ordinary differential equation as v is the velocity field, representing the degradation speed, that is, the noise distribution added to the high-quality image at the t-th degradation.
[0034] Although the process represented is a reversible process, however, it cannot be directly applied to the degradation process of high-quality images because the degradation of images is often not reversible. The reason is that for a low-quality image, or for any intermediate image in the process of degrading a high-quality image to a low-quality image, it may be degraded from different high-quality images through different degradation processes. This makes the degradation process from a high-quality image to a low-quality image uncertain, which is also the reason why uncertainty factors are introduced when using diffusion models for denoising in the prior art.
[0035] Obviously, the more severely an image is degraded and the lower its quality, the more corresponding high-quality images it has (the high-quality image corresponding to a low-quality image means that this low-quality image may be degraded from this high-quality image through a certain degradation process). That is to say, the greater its uncertainty, and vice versa, the smaller its uncertainty. When the high-quality image has not been degraded, it has determinacy, and as the degradation process continues, this uncertainty increases continuously, but the information in the image is not lost. From the perspective of mutual information, in the entire degradation process, no matter how much a high-quality image degrades into a low-quality image, the mutual information between this low-quality image and the original high-quality image is constant. Then, since the degradation process from a high-quality image to a low-quality image should originally be deterministic, but the represented degradation process is uncertain, and because no matter what kind of degradation occurs, the mutual information is not lost, therefore, an auxiliary variable y t is introduced in the embodiments of this specification to represent this uncertainty, that is, z t = [x t ; y t , x tis the image after the high-quality image x0 has been degraded t times.
[0036] z t = [x t ; y t indicates that the high-quality image x0 should have degenerated into x according to a certain degradation process t , but due to the existence of the auxiliary variable y t , the originally deterministic degradation process becomes an uncertain degradation process At t = 0, y t = 0, at this time z t is deterministic, and as t increases, the uncertainty of z t becomes greater and greater, or in other words, the entropy is greater.
[0037] Therefore, for determinism, this embodiment of the specification introduces a random noise image, and the introduced random noise image can increase the entropy of the degradation process z t when the number of degradation times t increases, and maintain determinism when the number of degradation times t is 0.
[0038] Among them, the random noise image can be a noise image with any random distribution, such as a Gaussian noise image.
[0039] S102: Fuse the image to be processed and the random noise image to obtain a fused image.
[0040] After the server obtains the low-quality image to be processed and introduces the random noise image, it can fuse the image to be processed and the random noise image to obtain a fused image.
[0041] Specifically, the image to be processed can be encoded first to obtain the image features of the image to be processed, and then the image features of the image to be processed and the random noise image are fused to obtain a fused image. When encoding the image to be processed, specifically, the image to be processed can be encoded by VAE, that is, the image to be processed is input into a variational autoencoder (VAE) to obtain the image features of the image to be processed output by the VAE. When fusing, the image features of the image to be processed and the random noise image can be directly spliced to obtain a fused image.
[0042] On the one hand, the fused image can represent the degraded image x t (that is, the current image to be processed), and on the other hand, it can also represent the uncertainty in the degradation process when the original high-quality image x0 degrades to the current x t through the random noise image.
[0043] S104: Input the standard noise image into a pre-trained flow matching model, and inject the fusion image as a generation condition into the flow matching model, so that the flow matching model predicts the noise addition speed corresponding to the standard noise image based on the generation condition, and denoises the standard noise image according to the noise addition speed to obtain a restored image.
[0044] Since the fusion image obtained through step S102 can represent both the degraded image x t , and can also represent the uncertainty during the degradation process when the original high-quality image x0 degrades to the current x t , therefore, the server can inject the fusion image as a generation condition into the flow matching model.
[0045] In essence, the flow matching model is also a diffusion model. Therefore, its denoising process also starts from a standard noise image and the standard noise image needs to be input into the pre-trained flow matching model. Among them, the standard noise image can be a noise image with any distribution, such as a white noise image.
[0046] Different from the diffusion model in the prior art, for one-time noise addition, the diffusion model in the prior art pays more attention to what kind of image an image is before noise addition and what kind of image it is after noise addition, while the flow matching model pays more attention to what kind of noise distribution the noise added in this noise addition is. Therefore, after injecting the above fusion image as a generation condition into the flow matching model, the flow matching model can know that: an unknown high-quality image, after going through the degradation process with the uncertainty described by the above random noise image, forms the current low-quality image x t . Thus, the flow matching model can use the above information obtained as a prior, and based on this prior, predict what kind of noise distribution was added to the original high-quality image during the degradation process assuming there is no such uncertainty. The predicted noise distribution is the noise addition speed corresponding to the standard noise image.
[0047] After predicting the above noise addition speed, since the degradation process is reversible, the flow matching model can directly reverse the above degradation process into a denoising process, that is, denoise the standard noise image according to the noise addition speed to complete a denoising process.
[0048] In the embodiments of this specification, the degradation process of degrading a high-quality image to a low-quality image is regarded as a degradation process that has been degraded T times, where T is the maximum number of noise addition times. Therefore, when denoising a standard noise image, step S104 also needs to be iteratively executed, and the above denoising process is executed T times. That is, after denoising the standard noise image once using step S104 to obtain a restored image, this restored image is re-input into the flow matching model as the standard noise image, so that the flow matching model continues to denoise the re-determined standard noise image based on the generation conditions until the number of denoising times reaches the maximum number of noise addition times T.
[0049] Through the above method, the image quality restoration process by the flow matching model is regarded as the inverse process of the image degradation process, and the uncertainty in the degradation process of the random noise image quantization coding is introduced. The flow matching model uses the degraded low-quality image and the uncertainty of the degradation process to the low-quality image as the generation conditions to denoise the standard noise image, and can obtain the high-quality image corresponding to the low-quality image, realizing the restoration of the image quality of the low-quality image.
[0050] Furthermore, the above Figure 1 The training method of the flow matching model can be as Figure 2 shown. Figure 2 It is a schematic diagram of the method for training the flow matching model provided by the embodiments of this specification, including the following steps:
[0051] S200: Obtain a sample high-quality image, the sample low-quality image corresponding to the sample high-quality image, and a random noise image.
[0052] In the embodiments of this specification, the server can pre-train the flow matching model using the Figure 2 shown method, and then perform image quality restoration using the Figure 1 shown method.
[0053] When training the flow matching model, the server can first obtain a sample high-quality image and the sample low-quality image corresponding to the sample high-quality image, that is, the sample low-quality image degraded from the sample high-quality image, and also obtain a random noise image used to encode the uncertainty in the degradation process.
[0054] S202: Fuse the sample high-quality image, the sample low-quality image, and the random noise image to obtain a sample fused image.
[0055] Since in the embodiments of this specification, it is necessary to encode and represent the uncertainty during the degradation process through the introduced random noise image. When degrading to the sample low-quality image, the entropy is the largest, indicating the greatest uncertainty. When not degraded, it is 0, indicating that the degradation process is deterministic at this time. Therefore, the sample high-quality image, the sample low-quality image, and the random noise image can be fused according to the current noise addition times t. The larger t is, the more degradation times the sample high-quality image x0 has experienced, the less component of the sample high-quality image in the sample fusion image, and the greater the components of the sample low-quality image and the random noise image.
[0056] Specifically, when fusing the sample high-quality image, the sample low-quality image, and the random noise image, the current noise addition times t can be randomly determined according to the preset maximum noise addition times T, where 1 ≤ t ≤ T. According to the current noise addition times t, the fusion weights corresponding to the sample high-quality image, the sample low-quality image, and the random noise image are determined respectively. Then, the sample high-quality image, the sample low-quality image, and the random noise image are weighted and fused according to the respective fusion weights. Among them, the fusion weight corresponding to the sample high-quality image is negatively correlated with the current noise addition times t, and the fusion weights corresponding to the sample low-quality image and the random noise image are positively correlated with the current noise addition times t.
[0057] Furthermore, when performing weighted fusion on the sample high-quality image, the sample low-quality image, and the random noise image, the sample high-quality image and the sample low-quality image can be encoded first to obtain the image features of the sample high-quality image and the image features of the sample low-quality image, and then the image features of the sample high-quality image, the image features of the sample low-quality image, and the random noise image are weighted and fused. When encoding the sample high-quality image and the sample low-quality image respectively, VAE can also be used for encoding.
[0058] If in step S102, when the server fuses the image to be processed and the random noise image, the method of directly splicing the image features of the image to be processed and the random noise image is used for fusion, then in step S202, the server can also directly splice the image features of the sample low-quality image and the random noise image. At this time, the sample low-quality image and the random noise image jointly correspond to the same fusion weight. Then the sample fusion image can be 1 / t × x0 + (1 - 1 / t) × [e; x T , where x T is the image feature of the sample low-quality image corresponding to the sample high-quality image x0, T is the maximum noise addition times, e is the random noise image, [e; x Trepresents the image features of the spliced sample low-quality image and the random noise image. It can be seen that as t increases, the component of the sample high-quality image in the sample fusion image decreases, the component of the sample low-quality image increases, and the component of the random noise image representing uncertainty also increases.
[0059] S204: Input the standard noise image into the flow matching model to be trained, and inject the sample fusion image as a sample generation condition into the flow matching model to be trained, so that the flow matching model to be trained predicts the noise addition speed corresponding to the standard noise image based on the sample generation condition as the predicted speed.
[0060] After obtaining the above sample fusion image through step S202, the sample fusion image can be injected into the flow matching model to be trained as a generation condition. Among them, the flow matching model can be a flow matching model based on U-Net. Then the server injects the sample fusion image into the U-Net in the flow matching model, so that the U-Net predicts the degradation process z under the guidance of this sample fusion condition t Degraded from z0 for T (the maximum number of degradations) times to z T When it is, the noise distribution added to the sample high-quality image at the t-th degradation is the noise addition speed corresponding to the standard noise image, that is, the predicted speed.
[0061] S206: Determine the labeled noise addition speed according to the sample high-quality image and the sample low-quality image, and determine the loss value according to the difference between the predicted speed and the labeled noise addition speed.
[0062] In the embodiment of this specification, when determining the labeled noise addition speed, the total noise distribution included in the sample low-quality image relative to the sample high-quality image can be determined according to the sample high-quality image and the sample low-quality image, and then the entire degradation process z can be determined according to the maximum number of noise additions T and the total noise distribution t Degraded from z0 for T times to z T When it is, the noise distribution added to the sample high-quality image on average each time of degradation is used as the labeled noise addition speed.
[0063] It should be noted that the process of predicting the noise distribution added to the sample image at the t-th degradation through the flow matching model to be trained in step S204 of this specification and the process of determining the labeled noise addition speed in step S206 do not have a sequential execution order and can be executed synchronously.
[0064] After predicting the noise distribution (i.e., the predicted speed) added to the sample image at the t-th degradation and determining the labeled noise addition speed, the loss value can be determined according to the difference between the predicted speed and the labeled noise addition speed. Among them, the loss value is positively correlated with the difference. The greater the difference between the predicted speed and the labeled noise addition speed, the greater the loss value, and vice versa, the smaller the loss value.
[0065] Specifically, in step S202, when fusing the sample high-quality image, the sample low-quality image, and the random noise image, multiple current noise addition times t can be randomly determined according to the preset maximum noise addition times T. For each current noise addition time t, a sample fused image is obtained respectively. And through step S204, for each current noise addition time t, the predicted speed is predicted by the flow matching model to be trained. Since the predicted speed predicted by the flow matching model under ideal conditions for different current noise addition times t should be equal to the labeled noise addition speed, when determining the loss value according to the difference between the predicted speed and the labeled noise addition speed, the mean square error loss value between the predicted speed predicted by the flow matching model to be trained for each current noise addition time t and the labeled noise addition speed can be determined.
[0066] S208: Taking reducing the loss value as the training objective, adjust the model parameters of the flow matching model to be trained.
[0067] Taking minimizing the above mean square error loss as the training objective and adjusting the model parameters of the flow matching model can enable the flow matching model to learn how to predict any low-quality image x under the premise of eliminating the uncertainty represented by the random noise image as a prior according to the uncertainty in the degradation process. t The noise distribution added at the t-th degradation, and try to keep x t The noise addition speed at each degradation in the degradation process that has been experienced is uniform.
[0068] Through Figure 2 After training the flow matching model by the method shown, the trained flow matching model can be applied to Figure 1 The image quality restoration method shown to realize the image quality restoration of the low-quality image to be processed.
[0069] In addition, as Figure 1 shown in the image quality restoration method, when denoising through the flow matching model, the guidance for the flow matching model mainly has two aspects. On the one hand, it is the degraded image x t (i.e., the current image to be processed), and on the other hand, it is the uncertainty in the degradation process when the high-quality image x0 represented by the random noise image degrades to the current x t . In the embodiments of this specification, more guidance, that is, more generation conditions, can also be provided for the flow matching model to further improve the effect of the flow matching model for image quality restoration.
[0070] Specifically, in addition to Figure 1In addition to injecting the shown fused image as a generation condition into the flow matching model, the type information of the noise contained in the image to be processed can also be injected into the flow matching model as one of the generation conditions. The types of noise contained in the image to be processed may include, but are not limited to, one or a combination of several of rain, fog, night, motion blur, and white noise (this is because an image often contains more than one type of noise). Correspondingly, when training the flow matching model by the training method shown in Figure 2 the type of noise contained in the sample low-quality image should also be injected into the flow matching model to be trained as one of the generation conditions. In order to improve the image restoration quality, the type information of the noise that needs to be injected as one of the generation conditions can be the visual coding features of this type of noise. That is, when performing the image quality restoration method shown in Figure 1 first, the type of noise contained in the image to be processed can be determined, then the noise of this type can be visually coded to obtain the visual coding features of this type of noise, and the visual coding features can be injected into the flow matching model as one of the generation conditions; correspondingly, when performing the method for training the flow matching model as shown in Figure 2 it is also necessary to first determine the type of noise contained in the sample low-quality image, then visually code the noise of this type to obtain the visual coding features of this type of noise, and inject the visual coding features into the flow matching model to be trained as one of the generation conditions.
[0071] The above is a method for image quality restoration provided by the embodiments of this specification. Based on the same idea, this specification also provides corresponding devices, storage media, and electronic devices.
[0072] Figure 3 The following is a schematic diagram of a device for image quality restoration provided by the embodiments of this specification. The device includes:
[0073] An acquisition module 301, configured to acquire an image to be processed and a random noise image;
[0074] A fusion module 302, configured to fuse the image to be processed and the random noise image to obtain a fused image;
[0075] A denoising module 303, configured to input a standard noise image into a pre-trained flow matching model, and inject the fused image into the flow matching model as a generation condition, so that the flow matching model predicts the noise addition speed corresponding to the standard noise image based on the generation condition, and denoise the standard noise image according to the noise addition speed to obtain a restored image.
[0076] Optionally, the fusion module 302 is specifically configured to encode the image to be processed to obtain the image features of the image to be processed; fuse the image features of the image to be processed and the random noise image to obtain a fused image.
[0077] Optionally, the fusion module 302 is specifically configured to splice the image features of the image to be processed and the random noise image.
[0078] Optionally, the apparatus further includes:
[0079] A training module 304, configured to obtain a sample high-quality image, a sample low-quality image corresponding to the sample high-quality image, and a random noise image; fuse the sample high-quality image, the sample low-quality image, and the random noise image to obtain a sample fusion image; input a standard noise image into a flow matching model to be trained, and inject the sample fusion image as a sample generation condition into the flow matching model to be trained, so that the flow matching model to be trained predicts a noise addition speed corresponding to the standard noise image based on the sample generation condition as a predicted speed; determine a labeled noise addition speed according to the sample high-quality image and the sample low-quality image, and determine a loss value according to the difference between the predicted speed and the labeled noise addition speed; wherein, the loss value is positively correlated with the difference; and adjust model parameters of the flow matching model to be trained with the goal of reducing the loss value.
[0080] Optionally, the training module 304 is specifically configured to determine a current noise addition number t; determine respective fusion weights corresponding to the sample high-quality image, the sample low-quality image, and the random noise image according to the current noise addition number t; wherein, the fusion weight corresponding to the sample high-quality image is negatively correlated with the current noise addition number t, and the fusion weights corresponding to the sample low-quality image and the random noise image are positively correlated with the current noise addition number t; and perform weighted fusion on the sample high-quality image, the sample low-quality image, and the random noise image according to the respective fusion weights corresponding to the sample high-quality image, the sample low-quality image, and the random noise image.
[0081] Optionally, the noise addition speed predicted by the flow matching model to be trained based on the sample generation condition is: the noise distribution added when performing the t-th noise addition on the sample high-quality image;
[0082] The training module 304 is specifically configured to determine a total noise distribution included in the sample low-quality image according to the sample high-quality image and the sample low-quality image; and determine, according to a preset maximum noise addition number and the total noise distribution, a noise distribution added on average each time when performing noise addition on the sample high-quality image as the labeled noise addition speed.
[0083] Optionally, the training module 304 is specifically configured to randomly determine a plurality of current noise addition numbers t according to a preset maximum noise addition number;
[0084] Determine the mean square error loss value between each predicted speed and the labeled noise addition speed according to the predicted speeds respectively obtained for each current noise addition count t based on the labeled noise addition speed and the flow matching model to be trained.
[0085] This specification also provides a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, it can be used to execute the above-provided method for image quality restoration.
[0086] Based on Figure 1 the method for image quality restoration shown, embodiments of this specification also provide Figure 4 the structural schematic diagram of the electronic device shown. As Figure 4 , at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above method for image quality restoration.
[0087] The above are only embodiments of this specification and are not used to limit this specification. For those skilled in the art, various changes and modifications can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.
Claims
1. A method for image quality restoration, the method comprising: Obtaining an image to be processed and a random noise image; Fusing the image to be processed and the random noise image to obtain a fused image; Inputting a standard noise image into a pre-trained flow matching model, and injecting the fused image as a generation condition into the flow matching model, so that the flow matching model predicts the noise addition speed corresponding to the standard noise image based on the generation condition, and denoises the standard noise image according to the noise addition speed to obtain a restored image.
2. The method according to claim 1, wherein fusing the image to be processed and the random noise image to obtain a fused image specifically comprises: Encoding the image to be processed to obtain the image features of the image to be processed; Fusing the image features of the image to be processed and the random noise image to obtain a fused image.
3. The method according to claim 2, wherein fusing the image features of the image to be processed and the random noise image specifically comprises: Stitching the image features of the image to be processed and the random noise image.
4. The method according to claim 1, wherein pre-training the flow matching model specifically comprises: Obtaining a sample high-quality image, a sample low-quality image corresponding to the sample high-quality image, and a random noise image; Fusing the sample high-quality image, the sample low-quality image, and the random noise image to obtain a sample fused image; Inputting a standard noise image into the flow matching model to be trained, and injecting the sample fused image as a sample generation condition into the flow matching model to be trained, so that the flow matching model to be trained predicts the noise addition speed corresponding to the standard noise image based on the sample generation condition, as the predicted speed; Determining the labeled noise addition speed according to the sample high-quality image and the sample low-quality image, and determining a loss value according to the difference between the predicted speed and the labeled noise addition speed; wherein, the loss value is positively correlated with the difference; Taking reducing the loss value as a training objective, and adjusting the model parameters of the flow matching model to be trained.
5. The method according to claim 4, wherein fusing the sample high-quality image, the sample low-quality image, and the random noise image specifically comprises: Determining the current noise addition times t; According to the current noise addition times t, determining the fusion weights corresponding to the sample high-quality image, the sample low-quality image, and the random noise image respectively; wherein, the fusion weight corresponding to the sample high-quality image is negatively correlated with the current noise addition times t, and the fusion weights corresponding to the sample low-quality image and the random noise image are positively correlated with the current noise addition times t; Performing weighted fusion on the sample high-quality image, the sample low-quality image, and the random noise image according to the fusion weights corresponding to the sample high-quality image, the sample low-quality image, and the random noise image respectively.
6. The method according to claim 5, wherein the noise addition speed predicted by the flow matching model to be trained based on the sample generation condition is: the noise distribution added when performing the t-th noise addition on the sample high-quality image; Determine the labeled noise addition speed according to the high-quality sample image and the low-quality sample image, specifically including: Determine the total noise distribution included in the low-quality sample image according to the high-quality sample image and the low-quality sample image; Determine the noise distribution added each time when adding noise to the high-quality sample image on average according to the preset maximum number of noise addition times and the total noise distribution, as the labeled noise addition speed.
7. The method according to claim 6, determining the current number of noise addition times t, specifically including: Randomly determine multiple current numbers of noise addition times t according to the preset maximum number of noise addition times; Determine the loss value according to the difference between the predicted speed and the labeled noise addition speed, specifically including: Determine the mean square error loss value between each predicted speed and the labeled noise addition speed according to the labeled noise addition speed and the predicted speeds respectively predicted by the flow matching model to be trained for each current number of noise addition times t.
8. An image quality restoration device, the device includes: An acquisition module, configured to acquire an image to be processed and a random noise image; A fusion module, configured to fuse the image to be processed and the random noise image to obtain a fused image; A denoising module, configured to input a standard noise image into a pre-trained flow matching model, and inject the fused image as a generation condition into the flow matching model, so that the flow matching model predicts the noise addition speed corresponding to the standard noise image based on the generation condition, and denoise the standard noise image according to the noise addition speed to obtain a restored image.
9. A computer-readable storage medium, the storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1-7 above is implemented.
10. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method described in any one of claims 1-7 above is implemented.
Citation Information
Cited By
Speech enhancement method and device based on stream matching, equipment and medium
CN120913576A