Image processing method and device

By repairing old film video frames based on the diffusion model and constructing training data pairs, the problem of difficult to truly simulate old film quality degradation in the prior art is solved, and a more efficient old film quality repair effect is achieved.

CN120017908APending Publication Date: 2025-05-16SHANGHAI HODE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510127944.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing old film repair methods generate training data by artificially degrading quality, making it difficult to truly simulate the deterioration of the picture quality of old films, resulting in unsatisfactory repair results.

Method used

By obtaining the first video frame of the low-quality video, repairing it based on the diffusion model, and obtaining the second video frame; then constructing training data pairs based on the first video frame and the second video frame, training the target repair model to obtain the trained target repair model; input the video to be repaired into the model for repair, and obtaining the repaired video.

Benefits of technology

Through the diffusion model, high-quality video frames corresponding to low-quality video frames are constructed, and training data pairs that are closer to reality are constructed, which improves the image quality repair effect of the target repair model, so that it can more effectively restore real low-quality images to high-definition images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017908A_ABST
    Figure CN120017908A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image processing method. The image processing method comprises the following steps: acquiring a first video frame of a low-quality video; repairing the first video frame based on a diffusion model to obtain a second video frame; constructing a training data pair based on the first video frame and the second video frame; training a target repair model based on the training data to obtain a trained target repair model; and inputting a to-be-repaired video into the trained target repairing model for repairing to obtain a repaired video. According to the technical scheme of the embodiment of the invention, the restoration model can learn the capability of restoring a real low-quality image to a high-definition image, and the restoration effect of the restoration model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to an image processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Art

[0002] Old film restoration refers to the use of algorithms to repair image quality problems in old and low-quality video resources, such as old TV shows, old movies, and old cartoons, in order to improve the viewing picture quality.

[0003] At present, the method of restoring old films is mainly to artificially degrade high-definition images to generate training data pairs of low-quality and high-definition images, and then use these data to train the model so that the model has the ability to restore low-quality images to high-definition images.

[0004] However, methods based on artificial degradation often fail to truly simulate the image quality degradation of old films, resulting in less than ideal restoration effects of old films.

[0005] It should be noted that the above content is not necessarily prior art, nor is it intended to limit the scope of patent protection of this application. Summary of the invention

[0006] The embodiments of the present application provide an image processing method, apparatus, computer device, computer-readable storage medium, and computer program product to solve or alleviate one or more of the technical problems raised above.

[0007] One aspect of an embodiment of the present application provides an image processing method, the method comprising: Get the first video frame of the low-quality video; Repairing the first video frame based on a diffusion model to obtain a second video frame; constructing a training data pair based on the first video frame and the second video frame; Training a target repair model based on the training data to obtain a trained target repair model; The video to be repaired is input into the trained target repair model for repair, to obtain a repaired video.

[0008] Optionally, the repairing the first video frame based on the diffusion model to obtain the second video frame includes: The first video frame is divided into N frames according to the length of the first video frame to obtain a plurality of groups of first video frame sequences, where N is a natural number greater than 1; Inputting each group of the first video frame sequences into the diffusion model for restoration, respectively, to obtain a second video frame sequence corresponding to each group of the first video frame sequences; Correspondingly, constructing a training data pair based on the first video frame and the second video frame includes: Pairing each group of the first video frame sequence with the corresponding second video frame sequence to form preliminary data pairs; A training data pair is constructed based on the preliminary data pair.

[0009] Optionally, constructing a training data pair based on the preliminary data pair comprises: Evaluating the preliminary data pairs, and screening the preliminary data pairs to obtain remaining data pairs based on the evaluation results; The training data pairs are obtained based on the remaining data pairs.

[0010] Optionally, the evaluating the preliminary data pairs and selecting the remaining data pairs from the preliminary data pairs based on the evaluation result includes: evaluating a large multimodal model employing image quality perception based on said preliminary data; Remaining data pairs are obtained based on the evaluation results and the preliminary data pairs.

[0011] Optionally, the multimodal large model is evaluated based on information of image modality and information of text modality, the information of image modality includes information of the first video frame sequence and the second video frame sequence, and the information of text modality is guiding information for evaluating, screening or rejecting input images.

[0012] Optionally, obtaining remaining data pairs based on the evaluation results and the preliminary data pairs includes: Based on the evaluation results, preliminary screening data pairs are obtained from the preliminary data pairs; receiving subjective evaluation information of the preliminary screening data; The preliminary screened data pairs are screened a second time based on the subjective evaluation information to obtain the remaining data pairs.

[0013] Another aspect of an embodiment of the present application provides an image processing device, the device comprising: An acquisition module, used for acquiring a first video frame of a low-quality video; A first restoration module, configured to restore the first video frame based on a diffusion model to obtain a second video frame; A construction module, configured to construct a training data pair based on the first video frame and the second video frame; A training module, used for training a target repair model based on the training data to obtain a trained target repair model; The second restoration module is used to input the video to be restored into the trained target restoration model for restoration, so as to obtain a restored video.

[0014] Another aspect of an embodiment of the present application provides a computer device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.

[0015] Another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method described above is implemented.

[0016] Another aspect of an embodiment of the present application provides a computer program product, including a computer program, which implements the method described above when executed by a processor.

[0017] The above technical solution adopted in the embodiment of the present application may have the following advantages: By obtaining a first video frame of a low-quality video, the first video frame is repaired based on a diffusion model to obtain a second video frame; training data is constructed based on the first video frame and the second video frame, and a target repair model is trained based on the training data to obtain a trained target repair model; the video to be repaired is input into the trained target repair model for repair to obtain a repaired video, and a high-quality video frame corresponding to the low-quality video frame can be constructed through the diffusion model, and a training data pair that is closer to reality is constructed to train the target repair model, so that the target repair model can learn the ability to restore real low-quality images to high-definition images, thereby effectively improving the image quality repair effect of the target repair model. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings exemplarily illustrate the embodiments and constitute a part of the specification, and together with the text description of the specification, are used to explain the exemplary implementation of the embodiments. The embodiments shown are for illustrative purposes only and do not limit the scope of the claims. In all drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0019] Figure 1 The principle diagram of the old film restoration technology in the related art is schematically shown; Figure 2 The schematic diagram shows a principle diagram of an image processing method according to the first embodiment of the present application; Figure 3 The flowchart of the image processing method according to the first embodiment of the present application is schematically shown; Figure 4 Schematically shows Figure 3 Flow chart of sub-steps of step S102; Figure 5 Schematically shows Figure 3 Flow chart of sub-steps of step S104; Figure 6 Schematically shows Figure 5 Flow chart of sub-steps of step S206; Figure 7 Schematically shows Figure 6 Flow chart of sub-steps of step S300; Figure 8 Schematically shows Figure 7 Flow chart of sub-steps of step S402; Fig. 9 The following is a schematic diagram showing an example flow chart of an image processing method according to the first embodiment of the present application; Fig.10 A block diagram schematically shows an image processing device according to the second embodiment of the present application; and Fig.11 The hardware architecture diagram of the computer device according to the third embodiment of the present application is schematically shown. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0021] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0022] In the description of the present application, it should be understood that the numerical labels before the steps do not indicate the order in which the steps are executed, but are only used to facilitate the description of the present application and to distinguish each step, and therefore should not be understood as a limitation on the present application.

[0023] First, the following terms are explained: Old film restoration: Old film restoration refers to the use of intelligent technology or manual means to restore old and poor-quality film or TV series sources, making their quality clearer than the original, thereby improving the viewing experience.

[0024] Diffusion Model: It is a generative model that generates new data samples by gradually adding noise to the data and then learning how to recover the original data from the noise. This process is a bit like blurring a painting and then gradually restoring its details.

[0025] Image Quality-Aware Multimodal Large Model: is an AI model that can understand and evaluate image quality. It combines information from multiple modalities such as images and text, and can quantitatively and qualitatively evaluate image clarity, contrast, color saturation, noise level, etc., and output human-understandable descriptions.

[0026] Degradation: It is an important step in generating training data. By degrading high-quality images to low-quality images, input-output pairs can be constructed for supervised training.

[0027] Peak signal-to-noise ratio (PSNR): is a commonly used image quality evaluation indicator. It is mainly used to measure the similarity between two images and is often used to evaluate the performance of algorithms such as image compression, image restoration, and image denoising.

[0028] Structural Similarity Index (SSIM): is an indicator to measure the similarity between two images. Compared with traditional indicators such as mean square error (MSE), SSIM is more in line with the perceptual characteristics of the human visual system and can better reflect the degree of image distortion.

[0029] Secondly, in order to facilitate those skilled in the art to understand the technical solutions provided in the embodiments of the present application, the relevant technologies are described below: Artificial intelligence-based old film restoration technology has become one of the important technologies widely used by major video platforms. Its purpose is to give classic old films new vitality and enhance the user experience of watching old films. The common restoration method is to comprehensively optimize the image quality of the video through deep learning models. For example, for blurred images, texture details can be enhanced to make the image sharper; for blurred or distorted face areas, facial features can be accurately restored; random noise in the video can be effectively removed through noise reduction technology; and the natural tones of the picture can be restored through color correction technology, and even realistic colors can be added to black and white videos, and so on.

[0030] Mainstream old film restoration techniques are usually based on supervised learning frameworks, such as Figure 1 As shown in the figure, by artificially degrading the existing high-definition images, a training pair of low-quality images and high-definition images is generated; then a model is trained using these data so that the model has the ability to restore low-quality images to high-definition images. However, artificial degradation methods usually rely on specific algorithms, such as downsampling, adding Gaussian noise, JPEG compression or blur processing. Although the degraded images generated by these algorithms have obvious quality differences from high-definition images, they are far less complex than the image quality degradation in actual old films. Image quality problems in old films are often caused by a combination of multiple factors, including physical losses caused by age (such as scratches, fading, mildew, etc.), technical limitations of camera equipment (such as low resolution, motion blur), degradation of the storage process (such as tape noise, film grain), and distortion of the digital transcoding process. These problems are often difficult to accurately simulate with a single or simple degradation method.

[0031] To this end, the present application embodiment provides an image processing technical solution. In this technical solution, if Figure 2 As shown in the figure, low-quality images in old films are used as input, and the diffusion model's powerful ability to restore details is used to complete the missing image information in the low-quality images. In combination with subjective and objective evaluations, high-quality images that meet human perception are selected to construct training data pairs that are closer to reality. Since the distribution of low-quality images is beyond the range of human subjective perception, the accuracy and rationality of the estimation of low-quality images are difficult to evaluate. Therefore, the method of simulating low-quality images in related technologies, the restoration model is likely to not learn the ability to restore real low-quality images to real high-definition images. On the contrary, the accuracy and rationality of high-quality images are easier to evaluate and are more in line with the subjective perception of the human eye. Therefore, by constructing high-quality images from low-quality images and screening the constructed high-quality images to construct training data pairs, training data that is closer to the distribution of real data can be constructed, thereby effectively improving the model's image quality restoration ability. See below for details.

[0032] The technical solutions of the present application are described below through multiple embodiments. It should be noted that these embodiments can be implemented in a variety of different forms and should not be interpreted as being limited to the embodiments described here.

[0033] Embodiment 1 Figure 3 The flowchart of the image processing method according to the first embodiment of the present application is schematically shown.

[0034] like Figure 3 As shown, the image processing method may include steps S100 to S108, wherein: Step S100: acquiring a first video frame of a low-quality video.

[0035] Step S102: repair the first video frame based on a diffusion model to obtain a second video frame.

[0036] Step S104: construct a training data pair based on the first video frame and the second video frame.

[0037] Step S106: training a target repair model based on the training data to obtain a trained target repair model.

[0038] Step S108, inputting the video to be repaired into the trained target repair model for repair, to obtain a repaired video.

[0039] The image processing method provided in this embodiment obtains a first video frame of a low-quality video, and repairs the first video frame based on a diffusion model to obtain a second video frame; constructs training data based on the first video frame and the second video frame, and trains a target repair model based on the training data to obtain a trained target repair model; inputs the video to be repaired into the trained target repair model for repair to obtain a repaired video, and a high-quality video frame corresponding to the low-quality video frame can be constructed through the diffusion model, and a training data pair that is closer to reality is constructed to train the target repair model, so that the target repair model can learn the ability to restore real low-quality images to high-definition images, thereby effectively improving the image quality repair effect of the target repair model.

[0040] The following combination Figure 3 , each step in steps S100~S108 and other optional steps are explained in detail.

[0041] Step S100 , get the first video frame of the low-quality video.

[0042] Low-quality videos can be videos with lower picture quality, specifically old film videos, such as old TV series, old movies, old cartoons and old documentaries, etc. Of course, they can also be other videos with lower picture quality, such as videos shot with old equipment.

[0043] Specifically, each video frame of the low-quality video may be extracted to obtain the first video frame of the low-quality video. There may be multiple low-quality videos, that is, the first video frame may be extracted from multiple low-quality videos.

[0044] Step S102 , repairing the first video frame based on the diffusion model to obtain a second video frame.

[0045] Specifically, the first video frame may be input into the diffusion model for restoration, and the restored video frame may be used as the second video frame. When the first video frame is input into the diffusion model, one or more first video frames may be input into the diffusion model.

[0046] The diffusion model can be pre-trained, specifically, a large number of low-quality video frames can be input into the diffusion model for pre-training, and then the output of the diffusion model is evaluated to update the parameters of the diffusion model. After the output of the diffusion model reaches a certain accuracy, a pre-trained diffusion model is obtained. Then, the first video frame is input into the pre-trained diffusion model for restoration to obtain a second video frame, so that the restoration quality of the second video frame is better.

[0047] Step S104 , constructing a training data pair based on the first video frame and the second video frame.

[0048] Specifically, each first video frame and its corresponding second video frame may be used to construct a set of training data pairs, or two or more first video frames and their corresponding second video frames may be used to construct a set of training data. For example, the first video frames include video frame a, video frame b, video frame c, etc., and the corresponding second video frames after restoration are video frame A, video frame B, video frame C, etc., then video frame a and video frame A may be used as a set of training data pairs, or (video frame a and video frame b) and (video frame A and video frame B) may be used as a set of training data pairs.

[0049] In an optional embodiment, if Figure 4 As shown, step S102 (i.e., repairing the first video frame based on the diffusion model to obtain the second video frame) may include: Step S200: dividing the first video frame into N frames of length to obtain a plurality of first video frame sequences, where N is a natural number greater than 1.

[0050] Step S202: input each group of the first video frame sequences into the diffusion model for restoration, so as to obtain a second video frame sequence corresponding to each group of the first video frame sequences.

[0051] Correspondingly, if Figure 5 As shown, step S104 may include: Step S204 , pairing each group of first video frame sequences with the corresponding second video frame sequences to form preliminary data pairs.

[0052] Step S206: constructing a training data pair based on the preliminary data pair.

[0053] For example, assuming that there are 12 first video frames and N is 3, the first video frames are divided according to the length of N frames, and 4 groups of first video frame sequences can be obtained, namely, the 1st to 3rd frames are a group, the 4th to 6th frames are a group, the 7th to 9th frames are a group, and the 10th to 12th frames are a group.

[0054] After the first video frame sequence is obtained by segmentation, each group of the first video frame sequence can be input into the diffusion model for repair, and the second video frame sequence corresponding to each group of the first video frame sequence can be obtained. Then, the first video frame sequence can be paired with the corresponding second video frame sequence to form a preliminary data pair, and then the training data pair can be constructed based on the preliminary data pair. In the case where the diffusion model repair effect is good, the training data pair can be constructed based on the preliminary data pair by directly using the preliminary data pair as the training data pair.

[0055] In this embodiment, the first video frame is divided into N frames to obtain several groups of first video frame sequences, each group of first video frame sequences is input into the diffusion model for repair, and the second video frame sequence corresponding to each group of first video frame sequences is obtained. Each group of first video frame sequences is paired with the corresponding second video frame sequence to form a preliminary data pair, and a training data pair is constructed based on the preliminary data pair. Since multiple video frames contain timing information, compared with single-frame processing, inputting the video frame sequence into the diffusion model for repair can obtain better repair quality.

[0056] In an optional embodiment, if Figure 6 As shown, in step S206, constructing a training data pair based on the preliminary data pair may include: Step S300 , evaluating the preliminary data pairs, and screening the preliminary data pairs to obtain remaining data pairs based on the evaluation results.

[0057] Step S302: Obtain the training data pairs based on the remaining data pairs.

[0058] When evaluating the preliminary data, subjective and / or objective methods may be used, such as asking experts to conduct subjective evaluation, or using some objective indicators for objective evaluation, such as using indicators such as PSNR or SSIM to objectively evaluate the preliminary data. When multiple evaluations are included, different weights may be assigned to each evaluation, and then the multiple evaluations are weighted and summed according to the weights to obtain a comprehensive evaluation result, and then the preliminary data pairs are screened according to the comprehensive evaluation results to obtain the remaining data pairs.

[0059] After screening, the remaining data pairs can be directly used as training data pairs, or a part of the remaining data pairs can be selected as training data pairs according to certain rules.

[0060] In this embodiment, preliminary data pairs are evaluated, and remaining data pairs are screened from the preliminary data pairs based on the evaluation results, and training data pairs are obtained based on the remaining data pairs. Since the second video frame obtained by the diffusion model may contain some video frames that do not conform to the human eye perception, the preliminary data pairs are evaluated and screened to construct training data pairs. This can make the constructed training data pairs more reasonable and accurate, thereby improving the training effect of the model and improving the accuracy of model repair.

[0061] In an optional embodiment, if Figure 7 As shown, in step S300, the preliminary data pairs are evaluated, and the remaining data pairs are screened from the preliminary data pairs based on the evaluation results, which may include: Step S400 , evaluating a large multimodal model using image quality perception based on preliminary data.

[0062] Step S402, obtaining remaining data pairs based on the evaluation results and the preliminary data pairs.

[0063] Specifically, data related to the preliminary data pair can be obtained, such as image data, audio data, and text data, wherein the audio data can be the corresponding audio of the preliminary data pair, which can be converted into text content through speech recognition, and the text data can be the corresponding subtitles, titles, fragment information, and other related text content of the preliminary data pair; then, the relevant data is input into the multimodal large model of image quality perception, and the multimodal large model of image quality perception evaluates the preliminary data pair according to the relevant data to obtain the evaluation result; then, according to the evaluation result, some unreasonable preliminary data are eliminated to obtain the remaining data pairs. For example, the multimodal large model of image quality perception can output an evaluation score for all preliminary data pairs according to the relevant data, and then the preliminary data pairs with too low scores are eliminated according to the evaluation score to obtain the remaining data pairs.

[0064] Since the multimodal large model of image quality perception can accurately evaluate the image quality based on information from multiple modalities, evaluating the preliminary data pairs using the multimodal large model of image quality perception can improve the accuracy of the evaluation while also improving the efficiency of the evaluation.

[0065] In an optional embodiment, the multimodal large model is evaluated based on information of image modality and information of text modality, wherein the information of image modality includes information of a first video frame sequence and a second video frame sequence, and the information of text modality is guiding information for evaluating, screening or rejecting input images.

[0066] Among them, the information of text modality can be, for example, "Does the restored image in the image pair meet the high image quality requirements?", "Does the restored image in the image pair have illogical textures?", "Does the restored image in the image pair have objects that cannot be clearly defined?" and other guiding information for evaluation, screening or elimination. Therefore, the multimodal large model of image quality perception determines whether the restored image (second video frame) in the preliminary data pair is reasonable or whether there are related problems based on the guidance of these text modal information and the comparison between the first video frame sequence and the second video frame sequence, thereby obtaining an evaluation result. The information of text modality can be implemented in the form of a text template, and the relevant text template can be pre-configured and then selected according to actual needs.

[0067] In this embodiment, by using the information of the first video frame sequence and the second video frame sequence as the image modality information of the multimodal large model of image quality perception, and using the guidance information for evaluating, screening or rejecting the input image as the text modality information of the multimodal large model, the multimodal large model can accurately obtain the evaluation result according to the guidance of the text modality information and the comparison of the video frames before and after restoration, thereby improving the accuracy of the multimodal large model evaluation.

[0068] In an optional embodiment, if Figure 8 As shown, in step S402, obtaining the remaining data pairs based on the evaluation results and the preliminary data pairs may include: Step S500: Screening the preliminary data pairs based on the evaluation results to obtain preliminary screening data pairs.

[0069] Step S502: receiving subjective evaluation information on the preliminary screening data.

[0070] Step S504: performing a secondary screening on the preliminary screened data pairs based on the subjective evaluation information to obtain the remaining data pairs.

[0071] Subjective evaluation information can be input by personnel with professional knowledge, for example, it can be a mark for scoring or eliminating preliminary screening data.

[0072] Specifically, after obtaining the evaluation results of the multimodal large model of image quality perception, the preliminary data is screened according to the evaluation results to obtain preliminary screened data pairs; then a group of personnel with professional knowledge of image evaluation can be requested to conduct subjective evaluation of the preliminary screened data pairs, and obtain the subjective evaluation information of these professional knowledge personnel on the preliminary screened data pairs; then, the preliminary screened data pairs are screened again according to the subjective evaluation information, and the preliminary screened data pairs with higher subjective evaluation information are retained as the remaining data pairs.

[0073] In this embodiment, preliminary screened data pairs are obtained by screening from preliminary data pairs based on evaluation results, subjective evaluation information on the preliminary screened data is received, and the preliminary screened data pairs are secondary screened based on the subjective evaluation information to obtain remaining data pairs. Training data pairs can be constructed in combination with subjective and objective evaluations, thereby effectively improving the quality of training data pairs and improving the effects of model training and repair.

[0074] Step S106 , training the target repair model based on the training data to obtain a trained target repair model.

[0075] The target restoration model can be a deep learning-based model, such as a generative adversarial network (GAN) or a convolutional neural network.

[0076] Specifically, each set of training data pairs can be input into the target repair model for training, and after the training reaches a preset accuracy, a trained target repair model is obtained.

[0077] Step S108 , input the video to be repaired into the trained target repair model for repair, and obtain the repaired video.

[0078] The video to be restored may be a video of lower image quality, such as an old film video.

[0079] Specifically, when the video quality needs to be repaired, the video frames of the video to be repaired can be extracted and input into the trained target repair model for repair to obtain the repaired video. Since the target repair model has learned how to repair from real low-quality video frames to reasonable high-quality video frames, by inputting low-quality video frames into the trained target repair model, high-quality and reasonable video frames can be obtained, and the repair effect is relatively ideal.

[0080] In order to make this application easier to understand, the following Fig. 9 Provide an example application. Figure 7 As shown, it can roughly include the following main processes: 1. Divide the old video frames into segments to obtain multiple video frame sequences (segments); 2. Input these video frame sequences into the diffusion model for repair to obtain the repaired video frame sequences; 3. Pair the video frame sequences before and after restoration to form data pairs; 4. Inputting the data pair into a large multimodal model of image quality perception, and simultaneously inputting text modality information for guidance into the large multimodal model of image quality perception; 5. Determine whether the test passes. If so, proceed to step 6, otherwise proceed to step 8; 6. Conduct subjective human eye evaluation on the data pairs that pass the test; 7. Determine whether the test is passed. If so, proceed to step 9; otherwise, proceed to step 8. 8. Discard the image pair; 9. Get the training data pairs for training.

[0081] In this exemplary application, the video frames of the old film are repaired by the diffusion model. The powerful ability of the diffusion model to restore details can be used to complete the missing image information in the low-quality video frames of the old film, and a set of candidate high-quality images can be obtained; then, the low-quality video frames are paired with the high-quality video frames and input into the multimodal large model of image quality perception. At the same time, the text modality information is input into the multimodal large model. The multimodal large model can be used to accurately evaluate the image quality based on the image modality information and the text modality information, and a more accurate evaluation result can be obtained; finally, combined with the subjective evaluation of the human eye, the quality of the constructed training data set can be further improved by combining subjective and objective evaluations, thereby improving the training effect of the repair model, so that the repair model can learn to learn from real low-quality images to corresponding high-quality and reasonable high-quality images, and ultimately improve the repair effect of the repair model.

[0082] Embodiment 2 Fig.10 The block diagram of the image processing device according to the second embodiment of the present application is schematically shown. The device can be divided into one or more program modules, one or more program modules are stored in a storage medium, and are executed by one or more processors to complete the embodiment of the present application. The program module referred to in the embodiment of the present application refers to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. Fig.10 As shown, the apparatus 500 may include an acquisition module 510, a first repair module 520, a construction module 530, a training module 540, and a second repair module 550, wherein: An acquisition module 510 is used to acquire a first video frame of a low-quality video; A first restoration module 520, configured to restore the first video frame based on a diffusion model to obtain a second video frame; A construction module 530, configured to construct a training data pair based on the first video frame and the second video frame; A training module 540 is used to train a target repair model based on the training data to obtain a trained target repair model; The second restoration module 550 is used to input the video to be restored into the trained target restoration model for restoration, so as to obtain a restored video.

[0083] In an optional embodiment, the first repair module 520 is further configured to: The first video frame is divided into N frames according to the length of the first video frame to obtain a plurality of groups of first video frame sequences, where N is a natural number greater than 1; Inputting each group of the first video frame sequences into the diffusion model for restoration, respectively, to obtain a second video frame sequence corresponding to each group of the first video frame sequences; Correspondingly, the construction module 530 is also used for: Pairing each group of the first video frame sequence with the corresponding second video frame sequence to form preliminary data pairs; A training data pair is constructed based on the preliminary data pair.

[0084] In an optional embodiment, the construction module 530 is further configured to: Evaluating the preliminary data pairs, and screening the preliminary data pairs to obtain remaining data pairs based on the evaluation results; The training data pairs are obtained based on the remaining data pairs.

[0085] In an optional embodiment, the construction module 530 is further configured to: evaluating a large multimodal model employing image quality perception based on said preliminary data; Remaining data pairs are obtained based on the evaluation results and the preliminary data pairs.

[0086] In an optional embodiment, the multimodal large model is evaluated based on information of image modality and information of text modality, the information of image modality includes information of the first video frame sequence and the second video frame sequence, and the information of text modality is guiding information for evaluating, screening or rejecting input images.

[0087] In an optional embodiment, the construction module 530 is further configured to: Based on the evaluation results, preliminary screening data pairs are obtained from the preliminary data pairs; receiving subjective evaluation information of the preliminary screening data; The preliminary screened data pairs are screened a second time based on the subjective evaluation information to obtain the remaining data pairs.

[0088] Embodiment 3 Fig.11The schematic diagram of the hardware architecture of a computer device 10000 suitable for implementing the image processing method according to the third embodiment of the present application is schematically shown. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle terminal, a game console, a virtual device, a workbench, a digital assistant, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server, or a server cluster composed of multiple servers), etc. Fig.11 As shown, the computer device 10000 includes but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 10010 can be an internal storage module of the computer device 10000, such as a hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 can also be an external storage device of the computer device 10000, such as a plug-in hard disk equipped on the computer device 10000, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the memory 10010 can also include both the internal storage module of the computer device 10000 and its external storage device. In this embodiment, the memory 10010 is generally used to store an operating system and various application software installed in the computer device 10000, such as program codes of image processing methods, etc. In addition, the memory 10010 can also be used to temporarily store various data that have been output or are to be output.

[0089] In some embodiments, the processor 10020 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.

[0090] The network interface 10030 may include a wireless network interface or a wired network interface, and the network interface 10030 is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and to establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be a wireless or wired network such as an intranet, the Internet, the Global System of Mobile communication (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, etc.

[0091] It should be pointed out that Fig.11 Only a computer device having components 10010 - 10030 is shown, but it should be understood that implementation of all of the components shown is not a requirement, and more or fewer components may alternatively be implemented.

[0092] In this embodiment, the image processing method stored in the memory 10010 can also be divided into one or more program modules and executed by one or more processors (such as processor 10020) to complete the embodiment of the present application.

[0093] Embodiment 4 An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the image processing method in the embodiment are implemented.

[0094] In this embodiment, the computer-readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the computer-readable storage medium can be an internal storage unit of a computer device, such as a hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium can also be an external storage device of a computer device, such as a plug-in hard disk equipped on the computer device, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the computer-readable storage medium can also include both the internal storage unit of the computer device and its external storage device. In this embodiment, the computer-readable storage medium is generally used to store an operating system and various application software installed on the computer device, such as the program code of the image processing method in the embodiment, etc. In addition, the computer-readable storage medium can also be used to temporarily store various types of data that have been output or are to be output.

[0095] Embodiment 5 An embodiment of the present application also provides a computer program product, including a computer program, which implements the method in the above embodiment when executed by a processor.

[0096] Obviously, those skilled in the art should understand that the modules or steps of the above-mentioned embodiments of the present application can be implemented by general-purpose computer devices, they can be concentrated on a single computer device, or distributed on a network composed of multiple computer devices, optionally, they can be implemented by executable program codes of computer devices, so that they can be stored in a storage device and executed by the computer device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0097] It should be noted that the above are only preferred embodiments of the present application, and the patent protection scope of the present application is not limited thereto. Any equivalent structure or equivalent process transformation made using the contents of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. An image processing method, characterized in that: The method comprises: Get the first video frame of the low-quality video; Repairing the first video frame based on a diffusion model to obtain a second video frame; constructing a training data pair based on the first video frame and the second video frame; Training a target repair model based on the training data to obtain a trained target repair model; The video to be repaired is input into the trained target repair model for repair, to obtain a repaired video.

2. The method according to claim 1, characterized in that The repairing the first video frame based on the diffusion model to obtain the second video frame includes: The first video frame is divided into N frames according to the length of the first video frame to obtain a plurality of groups of first video frame sequences, where N is a natural number greater than 1; Inputting each group of the first video frame sequences into the diffusion model for restoration, respectively, to obtain a second video frame sequence corresponding to each group of the first video frame sequences; Correspondingly, constructing a training data pair based on the first video frame and the second video frame includes: Pairing each group of the first video frame sequence with the corresponding second video frame sequence to form preliminary data pairs; A training data pair is constructed based on the preliminary data pair.

3. The method according to claim 2, characterized in that The constructing a training data pair based on the preliminary data pair comprises: Evaluating the preliminary data pairs, and screening the preliminary data pairs to obtain remaining data pairs based on the evaluation results; The training data pairs are obtained based on the remaining data pairs.

4. The method according to claim 3, characterized in that The step of evaluating the preliminary data pairs and selecting the remaining data pairs from the preliminary data pairs based on the evaluation results includes: evaluating a large multimodal model employing image quality perception based on said preliminary data; Remaining data pairs are obtained based on the evaluation results and the preliminary data pairs.

5. The method according to claim 4, characterized in that The multimodal large model is evaluated based on information of image modality and information of text modality, the information of image modality includes information of the first video frame sequence and the second video frame sequence, and the information of text modality is guiding information for evaluating, screening or eliminating input images.

6. The method according to claim 4, characterized in that The obtaining of the remaining data pairs based on the evaluation results and the preliminary data pairs comprises: Based on the evaluation results, preliminary screening data pairs are obtained from the preliminary data pairs; receiving subjective evaluation information of the preliminary screening data; The preliminary screened data pairs are screened a second time based on the subjective evaluation information to obtain the remaining data pairs.

7. An image processing device, characterized in that: The device comprises: An acquisition module, used for acquiring a first video frame of a low-quality video; A first restoration module, configured to restore the first video frame based on a diffusion model to obtain a second video frame; A construction module, configured to construct a training data pair based on the first video frame and the second video frame; A training module, used for training a target repair model based on the training data to obtain a trained target repair model; The second restoration module is used to input the video to be restored into the trained target restoration model for restoration, so as to obtain a restored video.

8. A computer device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.