An Image and Video Reconstruction Method, System, Terminal, and Storage Medium Based on a Diffusion Model

Through the lightweight time-step prediction model and reinforcement learning optimized diffusion model, the problem of insufficient reconstruction capability of diffusion model in image and video editing is solved, and the generation effect of higher accuracy and consistency is achieved, which is suitable for diversified editing tasks.

CN119888014BActive Publication Date: 2025-07-22GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510365205.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-22
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

When editing images and videos, the generated content does not exactly match the style, structure or context of the original image or video, resulting in insufficient reconstruction capabilities and inability to meet user editing needs.

Method used

By designing a lightweight time-step prediction model and reinforcement learning optimization, combining the diffusion model for reverse generation processing, reducing reconstruction errors, and adaptively adjusting parameters according to the editing scene to ensure that the generated content is consistent with the original data.

Benefits of technology

It significantly improves the reconstruction ability of the diffusion model, the generated images or videos are more consistent with the original content, adapt to diverse editing scenarios, reduces multi-step denoising errors, and improves generation quality and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888014B_ABST
    Figure CN119888014B_ABST
Patent Text Reader

Abstract

The present invention discloses an image and video reconstruction method, system, terminal and storage medium based on a diffusion model. The method includes: obtaining original image data and an original text description, and inputting them into the diffusion model to obtain noise data and intermediate features; determining a time step prediction model, and performing an inverse generation process on the noise data, intermediate features and the original text description through the time step prediction model and the diffusion model to obtain initial image data; updating the time step prediction model according to the initial image data to obtain an updated time step prediction model; obtaining an updated text description, and performing an inverse generation process on the noise data, intermediate features and the updated text description through the updated time step prediction model and the diffusion model to obtain target image data. The present invention can effectively improve the reconstruction ability of the diffusion model and achieve accurate reconstruction of the original image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to an image and video reconstruction method, system, terminal and computer-readable storage medium based on a diffusion model. Background Art

[0002] With the rapid development of digital content creation, the demand for image and video editing is increasing day by day. Traditional image and video editing methods usually rely on professional software. Although these software are powerful, they require users to have certain professional skills and experience, and the learning cost is relatively high. In addition to requiring a large number of manual operations when dealing with complex editing tasks, which is time-consuming and cumbersome, these tools also have certain limitations when implementing some complex creative effects (such as style transfer, content replacement, etc.). In recent years, with the rapid development of image generation models and video generation models, image and video editing methods based on diffusion models have gradually become a research hotspot.

[0003] However, the existing technology faces the problem of content consistency when using the diffusion model for image and video editing. During the editing process of an image or video, especially when modifying a local area of the image or video, the generated content may not fully match the style, structure or context of the original image or video, thus failing to meet the requirements of users for image and video editing.

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] The main purpose of the present invention is to provide an image and video reconstruction method, system, terminal and computer-readable storage medium based on a diffusion model, aiming to solve the problem that when the existing technology uses the diffusion model to edit an image or video, the generated content does not fully match the style, structure or context of the original image or video, thus failing to meet the requirements of users for image and video editing.

[0006] To achieve the above purpose, the present invention provides an image and video reconstruction method based on a diffusion model, and the image and video reconstruction method based on a diffusion model includes the following steps:

[0007] Obtain the original image data and the original text description, and input the original image data and the original text description into the diffusion model to obtain noise data and intermediate features;

[0008] Determine the time step prediction model, and perform reverse generation processing on the noise data, the intermediate features and the original text description through the time step prediction model and the diffusion model to obtain the initial image data;

[0009] Calculate the reconstruction error between the original image data and the initial image data, and update the time step prediction model according to the reconstruction error to obtain an updated time step prediction model;

[0010] Obtain updated text descriptions, and perform inverse generation processing on the noise data, the intermediate features, and the updated text descriptions through the updated time step prediction model and the diffusion model to obtain target image data.

[0011] Optionally, in the image and video reconstruction method based on a diffusion model, the obtaining of the original image data and the original text description, and inputting the original image data and the original text description into the diffusion model to obtain noise data and intermediate features specifically includes:

[0012] Obtain the original image data and the original text description, and input the original image data and the original text description into the diffusion model;

[0013] Perform forward diffusion processing on the original image data and the original text description through the diffusion model to obtain noise data;

[0014] Determine a preset time step, extract the key-value cache of the attention mechanism in the diffusion model according to the preset time step, and obtain the intermediate features corresponding to the preset time step according to the key-value cache.

[0015] Optionally, in the image and video reconstruction method based on a diffusion model, the determining of the time step prediction model, and performing inverse generation processing on the noise data, the intermediate features, and the original text description through the time step prediction model and the diffusion model to obtain the initial image data specifically includes:

[0016] Input the noise data, the intermediate features, and the original text description into the diffusion model, and perform inverse generation processing on the noise data, the intermediate features, and the original text description through the diffusion model;

[0017] Determine the time step prediction model, and perform prediction processing on each time step in the inverse generation processing through the time step prediction model;

[0018] When any one of the time steps in the inverse generation processing is 0, the inverse generation processing is completed, and the diffusion model outputs the initial image data.

[0019] Optionally, in the above-mentioned image and video reconstruction method based on a diffusion model, when calculating the reconstruction error between the original image data and the initial image data and updating the time step prediction model according to the reconstruction error to obtain an updated time step prediction model, it specifically includes:

[0020] Calculate the reconstruction error between the original image data and the initial image data, and determine whether the reconstruction error is greater than or equal to a preset threshold;

[0021] If so, update the time step prediction model to obtain an updated time step prediction model.

[0022] Optionally, in the above-mentioned image and video reconstruction method based on a diffusion model, when updating the time step prediction model to obtain an updated time step prediction model, it specifically includes:

[0023] Obtain the opposite number of the reconstruction error, and use the opposite number as a reward signal;

[0024] Obtain historical generated data, and update the time step prediction model according to the historical generated data and the reward signal to obtain an updated time step prediction model.

[0025] Optionally, in the above-mentioned image and video reconstruction method based on a diffusion model, after calculating the reconstruction error between the original image data and the initial image data and updating the time step prediction model according to the reconstruction error to obtain an updated time step prediction model, it further includes:

[0026] Perform iterative reverse generation processing on the noise data, the intermediate features, and the original text description through the updated time step prediction model and the diffusion model until the reconstruction error is less than the preset threshold.

[0027] Optionally, in the above-mentioned image and video reconstruction method based on a diffusion model, when obtaining an updated text description and performing reverse generation processing on the noise data, the intermediate features, and the updated text description through the updated time step prediction model and the diffusion model to obtain target image data, it specifically includes:

[0028] Determine the user's needs, and perform text update processing on the original text description according to the user's needs to obtain an updated text description, where the text update processing includes keyword replacement and content detail modification;

[0029] Input the updated text description, the noise data, and the intermediate features into the diffusion model again, and perform reverse generation processing on the updated text description, the noise data, and the intermediate features through the updated time step prediction model and the diffusion model to obtain target image data.

[0030] In addition, to achieve the above object, the present invention further provides an image and video reconstruction system based on a diffusion model, wherein the image and video reconstruction system based on a diffusion model includes:

[0031] A forward diffusion processing module, configured to obtain original image data and an original text description, and input the original image data and the original text description into a diffusion model to obtain noise data and intermediate features;

[0032] A time step prediction model processing module, configured to determine a time step prediction model, and perform reverse generation processing on the noise data, the intermediate features, and the original text description through the time step prediction model and the diffusion model to obtain initial image data;

[0033] A time step prediction model update module, configured to calculate a reconstruction error between the original image data and the initial image data, and update the time step prediction model according to the reconstruction error to obtain an updated time step prediction model;

[0034] A reverse generation processing module, configured to obtain an updated text description, and perform reverse generation processing on the noise data, the intermediate features, and the updated text description through the updated time step prediction model and the diffusion model to obtain target image data.

[0035] In the present invention, original image data and an original text description are obtained, and the original image data and the original text description are input into a diffusion model to obtain noise data and intermediate features; a time step prediction model is determined, and reverse generation processing is performed on the noise data, the intermediate features, and the original text description through the time step prediction model and the diffusion model to obtain initial image data; a reconstruction error between the original image data and the initial image data is calculated, and the time step prediction model is updated according to the reconstruction error to obtain an updated time step prediction model; an updated text description is obtained, and reverse generation processing is performed on the noise data, the intermediate features, and the updated text description through the updated time step prediction model and the diffusion model to obtain target image data. By using a time step prediction model to assist the diffusion model in reverse generation, the present invention can effectively improve the reconstruction ability of the diffusion model and reduce the reconstruction error in the reverse generation process of the diffusion model, so as to ensure that the generated target image data is more consistent with the original image data and effectively improve the reconstruction effect of the original image data. Brief Description of the Drawings

[0036] Figure 1 is a flowchart of a preferred embodiment of the method for image and video reconstruction based on the diffusion model of the present invention;

[0037] Figure 2 is a schematic diagram of the overall structural implementation process of a preferred embodiment of the method for image and video reconstruction based on the diffusion model of the present invention;

[0038] Figure 3 is a structural diagram of a preferred embodiment of the image and video reconstruction system based on the diffusion model of the present invention;

[0039] Figure 4 is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Description of the Preferred Embodiment

[0040] To make the objectives, technical solutions and advantages of the present invention more clear and definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only for explaining the present invention and are not used to limit the present invention.

[0041] With the rapid development of digital content creation, the demand for image and video editing is increasing day by day. Traditional editing methods usually rely on professional software, such as Adobe Photoshop and Premiere (both Adobe Photoshop and Premiere are image processing software). Although these software are powerful, they require users to have certain professional skills and experience, and the learning cost is relatively high. In addition to the need for a large number of manual operations when dealing with complex editing tasks, which are time-consuming and cumbersome, these tools also have certain limitations in achieving some complex creative effects (such as style transfer, content replacement, etc.).

[0042] In recent years, with the rapid development of image generation models and video generation models, the method for image and video editing based on the diffusion model has gradually become a research hotspot. The diffusion model is a type of generation model, and its working principle is to gradually add noise to the data and learn how to denoise in reverse to generate new data. This process mainly includes two stages: the forward diffusion stage and the reverse generation stage. In the forward diffusion stage, the diffusion model starts from the original data and gradually adds noise to generate a series of gradually degraded data; in the reverse generation stage, the diffusion model starts from the noisy data and gradually denoises and restores the original data.

[0043] However, when using diffusion models for image and video editing, there is a problem of content consistency. During the editing process, especially when modifying local regions of an image or video, the generated content may not fully match the style, structure, or context of the original image, thus failing to meet the requirements of image and video editing. This is because each part of the content generated by the diffusion model is gradually denoised based on noise data and control conditions, making it difficult to ensure consistency with the details of the original image or video, especially in complex backgrounds or diverse scenes. Due to the multiple denoising operations and the introduction of text descriptions, the generated images often fail to be consistent with the original images, that is, the diffusion model has the weakness of poor reconstruction ability, which limits its editing ability for images and videos.

[0044] The image and video editing method based on diffusion models has the following two main drawbacks: 1. Lack of adaptability: Most existing methods rely on unified and deterministic training-free strategies and cannot be adaptively adjusted according to different actual editing scenarios. This means that when dealing with different types of images or videos, there may still be unavoidable reconstruction errors. Although these methods provide a certain degree of improvement, they fail to be personalized and optimized for specific editing tasks, thus limiting their effectiveness in practical applications. 2. Failure to utilize the training ability of neural networks: Existing technologies do not fully exploit the potential of neural networks to adaptively address the editing requirements of different scenarios through parameter training. Due to the lack of a training-based dynamic adjustment mechanism, these methods cannot effectively reduce or eliminate reconstruction errors, thus affecting the performance and generation effect of the diffusion model in complex editing tasks. They fail to fully utilize the powerful capabilities of neural networks to improve the generation quality by learning to optimize and reduce reconstruction errors.

[0045] To solve the above problems, the present invention provides an image and video editing method for enhancing the reconstruction ability of diffusion models. Through a training-based method, the diffusion model can adaptively adjust the neural network parameters according to the actual image and video editing scenarios, thereby more accurately reconstructing the original image or video content and significantly improving the editing effect of the diffusion model. The present invention optimizes the reconstruction ability of the diffusion model through the training process, enabling it to handle the editing requirements in different scenarios, solving the problem of reconstruction errors existing in existing methods in diverse editing tasks, and providing a more efficient and flexible solution for the field of image and video editing.

[0046] The image and video reconstruction method based on diffusion models according to a preferred embodiment of the present invention, as Figure 1 shown, the image and video reconstruction method based on diffusion models comprises the following steps:

[0047] Step S10: Obtain the original image data and the original text description, and input the original image data and the original text description into the diffusion model to obtain the noise data and the intermediate features. Among them, the original image data can refer to either image data or video data.

[0048] The technical fields involved in the present invention include: 1. AIGC technology (Artificial Intelligence Generated Content): Using artificial intelligence algorithms, especially deep learning and natural language processing technologies, to automatically generate various forms of content, including text, images, audio, and video, etc. AIGC trains models through a large amount of data, enabling computers to simulate the process of human creation, and thus automatically generating creative and personalized content. 2. Diffusion Model Optimization technology: Through a series of methods and strategies, improve the effectiveness and efficiency of diffusion models in generation tasks. Diffusion models generate high-quality samples from noise by simulating the process of gradually adding noise and denoising, and are widely used in fields such as image generation, image restoration, super-resolution, and text generation. The core goal of the optimization technology is to improve the generation quality, accelerate the generation process, reduce the computational cost, and enhance the stability of the model. Common optimization technologies include improvements in network architecture, optimization of noise scheduling strategies, learning rate adjustment, pre-training and fine-tuning, sampling strategy optimization, and regularization methods, etc. These technologies help diffusion models achieve better results in multiple generation tasks and make them more efficient and practical in practical applications. 3. Intelligent Image and Video Editing Systems: Using artificial intelligence and deep learning technologies to automate or enhance traditional image and video editing processes. Through intelligent algorithms, these systems can automatically repair image defects, enhance image quality, perform style transformation, and object recognition and editing, etc. They can also intelligently edit videos, automatically generate background music, and add special effects in real time, significantly improving the creation efficiency and reducing the editing difficulty. These technologies are widely used in fields such as social media, advertising, entertainment, and art creation, promoting the automation and personalization of content production. 4. Real-time Editing & Acceleration technology: Through hardware and software optimization, ensure the real-time processing and editing of multimedia content such as images and videos. These technologies rely on means such as GPU acceleration, deep learning inference optimization, cloud computing, and edge computing to improve the editing efficiency and reduce latency. For example, through GPU and deep learning acceleration, the processing process of images and videos can be quickly completed, supporting real-time special effect applications, dynamic adjustment, and efficient rendering. Real-time Editing & Acceleration technology is widely used in fields such as video production, live broadcast, and virtual reality, greatly enhancing the creation and interaction experience.

[0049] Such as Figure 2As shown in the figure, the use of diffusion models for image and video editing in the present invention mainly consists of two steps: First, through the forward diffusion process, real images or videos are converted into noise data, and the intermediate features of the diffusion model at each time step are retained during this process; when the time step prediction model is updated, the text description is adjusted according to the user's editing preferences, and the edited text description and the converted noise data are input into the diffusion model, and the previously retained intermediate features are injected at each time step of the reverse generation process to generate the edited image or video.

[0050] It can be understood that the prerequisite for a diffusion model to achieve good image and video editing effects is to have strong reconstruction ability. For example, in the reverse generation process of the diffusion model, when the input text is the original unedited text description, the generated image or video should be consistent with the original content. Only on the basis of perfectly reconstructing the content consistent with the original image or video can the text description be flexibly edited and the design of the feature injection method be carried out.

[0051] Specifically, obtain the original image data and the original text description, and input the original image data and the original text description into the diffusion model; perform forward diffusion processing on the original image data and the original text description through the diffusion model to obtain noise data; determine a preset time step, and extract the key-value cache of the attention mechanism in the diffusion model according to the preset time step, and obtain the intermediate feature corresponding to the preset time step according to the key-value cache.

[0052] The methods for improving the reconstruction ability of the diffusion model can be mainly optimized from two perspectives. From the perspective of the forward diffusion process, the noise added in the forward diffusion stage is predicted by jointly inputting the current noise and the text description into the model, which ensures the consistency of the noise added in the forward diffusion stage and the noise removed in the reverse generation stage because the noise is predicted by the model. At the same time, the intermediate features of the model in the forward process can be injected into the model in the reverse generation process to promote the consistency of the generated image and the original image through the consistency of the model features. From the perspective of the reverse generation process, optimize the reverse solution process of the diffusion model through mathematical derivation to gradually reduce the error of multi-step denoising.

[0053] First, input the original image or original video (i.e., the original image data in the present invention) into the diffusion model and generate noise data and intermediate features through forward diffusion. The specific process is as follows: Input the original image or video and the original text description , and iteratively generate noise data through the forward diffusion process . At a specific time step (where is the time step, , (where is the number of time steps), extract intermediate features from the diffusion model (i.e., the KV cache of the attention mechanism in the diffusion model), which reflects the structural information of the latent variable

[0054] It can be understood that the process of forward expansion is as follows: (assign the original image or video to the latent variable of time step , where , is the th time step, is the latent variable corresponding to time step ); : ; where represents the value order of time steps, is the latent variable corresponding to the updated time step, is the latent variable corresponding to the time step, is the time step, is the updated time step, is the diffusion model, is the original text description (the latent variable corresponding to the updated time step is ); ; (the noise generated by the forward diffusion process is the latent variable of time step ).

[0055] Step S20, determine a time step prediction model, and perform reverse generation processing on the noise data, the intermediate features, and the original text description through the time step prediction model and the diffusion model to obtain initial image data.

[0056] Different from the method of fine-tuning the parameters of the diffusion model, the present invention focuses on designing an additional lightweight model to predict the time steps in the reverse iterative generation process. The core idea of this design is that fine-tuning the diffusion model may lead to the loss of the ability of the original model to generate images or videos. Especially after fine-tuning, the model may overfit to a specific editing task and lose its generality in general generation tasks. In addition, fine-tuning a model with a large number of parameters usually requires a large amount of training data sets. In actual image and video editing tasks, usually only a single image or a single video is edited, and the data volume is far from sufficient to support large-scale fine-tuning.

[0057] Therefore, the present invention proposes to design a lightweight additional model (i.e., the time step prediction model) to specifically predict the time steps during the iterative process. The time step prediction model is optimized only for specific editing tasks, without the need to retrain the entire diffusion model, nor will it affect the general generation ability of the model. In this way, without disturbing the original model's generation ability, the evolution of each step in the editing process can be precisely controlled, thereby enhancing the reconstruction ability of the diffusion model.

[0058] Specifically, input the noise data, the intermediate features, and the original text description into the diffusion model, and perform reverse generation processing on the noise data, the intermediate features, and the original text description through the diffusion model; determine the time step prediction model, and perform prediction processing on each time step in the reverse generation process through the time step prediction model (i.e., injecting the previously reserved intermediate features at each time step in the reverse generation process to generate the edited image or video); when any one of the time steps in the reverse generation process is 0, the reverse generation process is completed, and the diffusion model outputs the initial image data.

[0059] However, due to the approximation of the update iteration formula, that is, the update iteration formula should be converted to , the diffusion model cannot reconstruct the original image or video well due to error accumulation.

[0060] Among them, the reverse generation process ( ) is as follows: (Assign the noise data to the latent variable of time step ); : ; (Update the latent variable corresponding to time step ); (The generated image or video in the reverse generation process is the latent variable of time step ).

[0061] To reduce the reconstruction error, the present invention performs Taylor expansion on , and the iterative formula By changing to an expression taking the second-order term, a better reconstruction effect is obtained. However, the iterative formula is still an approximate formula, and there are still cumulative errors in the reverse generation process of the diffusion model. Therefore, although the training-free method for image and video editing using the diffusion model can reduce the reconstruction error and improve the generation effect to a certain extent, it still essentially relies on the approximate formula and cannot completely eliminate the cumulative error, which limits the accuracy of the final reconstruction effect. In addition, since these methods are not designed to consider the diversity and complexity of actual image and video editing scenarios, the trained models usually cannot generalize well to new and unseen editing tasks.

[0062] To overcome these problems, the present invention can make the diffusion model automatically adjust parameters according to different editing scenarios by introducing a training mechanism, thereby reducing the reconstruction error and improving the generalization ability of the model in various practical applications.

[0063] The present invention performs reverse generation based on noise data and text descriptions, and the specific implementation process is as follows: The noise data , the original text description and the intermediate features are input into the diffusion model to perform the reverse generation process. The reverse generation gradually denoises through an iterative mathematical expression to restore the image or video , where the KV cache of the attention mechanism in the diffusion model is replaced by the intermediate features . In the reverse generation at each time step, the next time step is dynamically predicted by the time step prediction model , being the prediction model parameters. When the time step decreases to 0, the image or video is generated. In particular, the noise data and the intermediate features are used as the input of the diffusion model to maintain consistency. is completely noised data generated by forward diffusion and serves as the starting point for reverse generation, ensuring that the generation process starts from a noise distribution related to the original data, thereby maintaining content consistency; captures the structural information in the diffusion process and serves as an additional input to constrain the generation process, making aligned with in details.

[0064] Step S30: Calculate the reconstruction error between the original image data and the initial image data, and update the time step prediction model according to the reconstruction error to obtain an updated time step prediction model.

[0065] Specifically, calculate the reconstruction error between the original image data and the initial image data, and determine whether the reconstruction error is greater than or equal to a preset threshold; if so, obtain the opposite number of the reconstruction error, and use the opposite number as the reward signal; obtain the historical generated data, and update the time step prediction model according to the historical generated data and the reward signal to obtain an updated time step prediction model. Further, perform iterative reverse generation processing on the noise data, the intermediate features, and the original text description through the updated time step prediction model and the diffusion model until the reconstruction error is less than the preset threshold.

[0066] Further, the present invention introduces the RLOO reinforcement learning training method, regards the opposite number of the reconstruction error as the reward signal, and regards the entire time step sequence as an action sequence, so as to optimize the parameters of the time step prediction model. Through the training process of reinforcement learning, the model can adaptively adjust the time step prediction, so as to minimize the reconstruction error in each iteration step and improve the consistency of the finally generated content.

[0067] As Figure 2 shown, the present invention evaluates the reconstruction error and decides the subsequent steps. The specific implementation process is as follows: 1. Calculate the reconstruction error between the generated image or video and the original image or video , which is the square of the second norm, representing the calculation of the difference between the two. 2. If the reconstruction error is lower than the preset threshold (generally set to 0.01), then directly jump to step S40; if the reconstruction error is higher than the preset threshold , then optimize the time step prediction model through reinforcement learning.

[0068] Among them, the specific implementation steps of optimizing the time step prediction model through reinforcement learning are as follows: 1. Use the opposite number of the reconstruction error as the reward signal , defined as the reward signal . 2. Adopt the RLOO method to optimize the parameters of the time step prediction model by using the historical generated data and the reward signal , and the optimization goal is , where is the optimization goal, is the optimization process. 3. Return to the step of inputting the noise data, the intermediate features, and the original text description into the diffusion model, and perform reverse generation processing according to the diffusion model and the time step prediction model (at this time, the updated time step prediction model is adopted), and perform a new round of reverse generation iteration.

[0069] Step S40: Obtain the updated text description, and perform reverse generation processing on the noise data, the intermediate features, and the updated text description through the updated time step prediction model and the diffusion model to obtain the target image data.

[0070] In the reverse generation stage, the user can provide additional control conditions (such as text descriptions) for the diffusion model, making the generation process more flexible and controllable. This enables the diffusion model to generate diverse high-quality image or video content, and the user can guide the generation effect with a simple text description. This flexibility and efficiency make the diffusion model show great potential in the field of digital content creation and editing, especially when achieving complex creative effects, it is more convenient and efficient than traditional editing tools.

[0071] Specifically, determine the user's needs, and perform text update processing on the original text description according to the user's needs to obtain the updated text description, where the text update processing includes keyword replacement and content detail modification; input the updated text description, the noise data, and the intermediate features into the diffusion model again, and perform reverse generation processing on the updated text description, the noise data, and the intermediate features through the updated time step prediction model and the diffusion model to obtain the target image data.

[0072] The present invention can adjust the original text description according to the user's needs to generate an edited text description (i.e., the updated text description in the present invention), and the editing includes keyword replacement and / or content detail modification.

[0073] As Figure 2 shown, the present invention generates new results based on the edited text description, and the specific implementation process is as follows: Input the noise data , the edited text description , and the intermediate features into the diffusion model again, and perform reverse generation to obtain the edited image or video. At this time, the next time step is predicted by the optimized time step prediction model , and finally the edited image or video is output.

[0074] The beneficial effects of the present invention:

[0075] 1. The present invention designs a lightweight time-step prediction model and optimizes it in combination with reinforcement learning. This model significantly improves the reconstruction ability of the diffusion model. By reducing the cumulative error in the reverse generation process, it achieves a higher-precision reconstruction of the original content. Compared with traditional methods, it can more accurately control each generation step, avoid the error accumulation in the multi-step denoising process, and ensure that the generated image or video content is more consistent with the original data.

[0076] 2. Adjustment of the adaptive generation process: The present invention has an adaptive ability and can adjust the generation process according to different editing tasks, so as to better adapt to diverse editing scenarios. Through this adaptive mechanism, the present invention effectively avoids the reconstruction errors that occur in the prior art in complex scenarios, can handle different types of image or video editing tasks, and ensures the consistency and high quality of the generated content.

[0077] 3. No need to modify the original diffusion model: The present invention does not require any modification to the original diffusion model and retains the general performance of the diffusion model in tasks such as text-to-image or text-to-video. This feature enables the present invention to improve the performance of the model in actual editing tasks without changing the original model architecture, taking into account both flexibility and effect.

[0078] 4. Low training data requirements and strong applicability: Compared with traditional methods, the present invention has lower requirements for training data, can be effectively trained under limited data conditions, quickly adapt to specific scenarios or tasks, and can flexibly adapt to personalized editing tasks for a single image or a single video. It can still maintain good performance in a small-sample environment and has a wide range of applications.

[0079] In summary, the innovations of the present invention include:

[0080] 1. The present invention designs a lightweight time-step prediction model for optimizing the reverse generation process of the diffusion model. By predicting the time-step sequence in the iterative process, it reduces the cumulative error in multi-step denoising, thereby improving the accuracy and consistency of the generated content. The present invention uses reinforcement learning to optimize the time-step prediction model, takes the time-step sequence as the action sequence, and trains with the negative value of the reconstruction error as the reward signal, realizing the function of dynamically adjusting the time-step sequence according to different editing tasks. Through the above dynamic adjustment mechanism, the present invention minimizes the reconstruction error in each generation step, thereby significantly improving the quality of the final generated result.

[0081] 2. The present invention sets up an innovative process for adaptive adjustment of the generation process: The present invention dynamically adjusts the generation process according to the actual editing scenario, integrates time step prediction, intermediate feature injection, and text description guidance to ensure that the generation results are highly consistent with user requirements in complex editing tasks (such as style transfer and local modification); the present invention automates the improvement of generation quality through an iterative optimization mechanism, combining threshold judgment and parameter update; the present invention can enhance the model performance in diverse scenarios, making the generated images or videos highly consistent with real data.

[0082] 3. Uniqueness of the technology combination: The present invention combines a diffusion model with a lightweight time step prediction model to solve the reconstruction error problem in traditional methods through a training mechanism; the present invention does not require large-scale fine-tuning of the original diffusion model and maintains its general performance in text-to-image and text-to-video tasks; the present invention avoids the overfitting problem in fine-tuning through an additional model design and achieves dedicated optimization for specific editing tasks; the present invention takes into account the efficiency and wide applicability of the original diffusion model and provides highly customized editing functions in practical applications.

[0083] In addition, the application scenarios of the present invention may include:

[0084] 1. Film and television post-production: The editing method of the present invention can achieve efficient shot modification, scene reconstruction, and special effect generation in the film and television industry, especially suitable for projects that need to complete a large number of editing tasks in a short time. In traditional film and television post-production, the modification of shots and scenes usually requires a lot of manual intervention, consuming a lot of time and effort. However, the present invention can significantly improve work efficiency, reduce manual operations, and quickly complete complex editing work through advanced automated editing technology. The application of this technology not only shortens the production cycle of the project but also ensures that, while maintaining the creative quality, it reduces labor costs and improves productivity.

[0085] 2. Advertising and marketing content generation: Advertising companies can use the editing method of the present invention to quickly generate image and video content that meets specific advertising needs. By analyzing users' editing preferences, market trends, and the needs of the target audience, personalized advertising materials can be customized. This process greatly shortens the advertising creation cycle, enabling advertising production to respond more quickly and flexibly to market changes and consumer demands. In addition, the present invention can help advertising companies quickly generate advertising materials of various specifications and styles on different media platforms, improving the efficiency and effectiveness of advertising placement and enhancing the market adaptability of advertising creativity.

[0086] 3. Game Development and Design: During the game development process, especially when frequent updates of image or video materials are required, developers can use the editing method of the present invention to quickly generate and adjust image, scene, and character materials in the game. This technology can ensure that the landscapes, characters, and scene styles in the game are in line with the overall game design and player needs, thereby enhancing the immersion and player experience of the game. At the same time, using the automated editing method of the present invention can significantly reduce the workload of manual modeling and rendering, lower the development cost, and remarkably shorten the game development cycle. The present invention is particularly suitable for game development models with rapid iteration and testing, improving development efficiency and flexibility.

[0087] 4. Personalized Entertainment Applications: In the field of personalized entertainment, such as customized video production, virtual idol creation, etc., the editing method of the present invention can quickly generate content and perform personalized customization to meet the diverse needs of users for entertainment products. Users can customize video content according to their personal preferences, select the appearance and behavioral characteristics of virtual idols, etc., thereby obtaining a more personalized entertainment experience. Using this technology, entertainment products can not only be launched more quickly but also achieve higher flexibility and diversity in terms of creativity, meeting the growing demand of modern users for personalized content. At the same time, this technology can also promote innovation in the entertainment industry, opening up new business models and application scenarios.

[0088] Furthermore, as Figure 3 shown, based on the above image and video reconstruction method based on the diffusion model, the present invention also correspondingly provides an image and video reconstruction system based on the diffusion model, wherein the image and video reconstruction system based on the diffusion model includes:

[0089] A forward diffusion processing module 51, configured to obtain original image data and an original text description, and input the original image data and the original text description into the diffusion model to obtain noise data and intermediate features;

[0090] A time step prediction model processing module 52, configured to determine a time step prediction model, and perform reverse generation processing on the noise data, the intermediate features, and the original text description through the time step prediction model and the diffusion model to obtain initial image data;

[0091] A time step prediction model update module 53, configured to calculate the reconstruction error between the original image data and the initial image data, and update the time step prediction model according to the reconstruction error to obtain an updated time step prediction model;

[0092] The reverse generation processing module 54 is used to obtain the updated text description, and perform reverse generation processing on the noise data, the intermediate features, and the updated text description through the updated time step prediction model and the diffusion model to obtain the target image data.

[0093] Further, as Figure 4 shown, based on the above image and video reconstruction method and system based on the diffusion model, the present invention also correspondingly provides a terminal, which includes a processor 10, a memory 20, and a display 30. Figure 4 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0094] The memory 20 may be an internal storage unit of the terminal in some embodiments, such as the hard disk or memory of the terminal. The memory 20 may also be an external storage device of the terminal in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the terminal. Further, the memory 20 may also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software installed on the terminal and various types of data, such as the program code for installing the terminal. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a program 40 for image and video reconstruction based on the diffusion model is stored on the memory 20, and the program 40 for image and video reconstruction based on the diffusion model can be executed by the processor 10 to implement the image and video reconstruction method based on the diffusion model in the present application.

[0095] The processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chips in some embodiments, and is used to run the program code stored in the memory 20 or process data, such as executing the image and video reconstruction method based on the diffusion model.

[0096] The display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. in some embodiments. The display 30 is used to display information on the terminal and to display a visual user interface.

[0097] In one embodiment, when the processor 10 executes the image and video reconstruction program 40 based on the diffusion model in the memory 20, the steps of the image and video reconstruction method based on the diffusion model are implemented.

[0098] In summary, the present invention provides an image and video reconstruction method, system and terminal based on a diffusion model. The method includes: obtaining original image data and an original text description, and inputting the original image data and the original text description into the diffusion model to obtain noise data and intermediate features; determining a time step prediction model, and performing reverse generation processing on the noise data, the intermediate features and the original text description through the time step prediction model and the diffusion model to obtain initial image data; calculating a reconstruction error between the original image data and the initial image data, and updating the time step prediction model according to the reconstruction error to obtain an updated time step prediction model; obtaining an updated text description, and performing reverse generation processing on the noise data, the intermediate features and the updated text description through the updated time step prediction model and the diffusion model to obtain target image data. By using the time step prediction model to assist the diffusion model in reverse generation, the present invention can effectively improve the reconstruction ability of the diffusion model and reduce the reconstruction error in the reverse generation process of the diffusion model, thereby ensuring that the generated target image data is more consistent with the original image data and effectively improving the reconstruction effect of the original image data.

[0099] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or terminal including that element.

[0100] Of course, those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium readable by a computer. When the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0101] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. An image and video reconstruction method based on a diffusion model, characterized in that, The image and video reconstruction method based on the diffusion model includes: Obtain the original image data and the original text description, and input the original image data and the original text description into the diffusion model to obtain noise data and intermediate features; Determine the time step prediction model, and perform reverse generation processing on the noise data, the intermediate features, and the original text description through the time step prediction model and the diffusion model to obtain the initial image data; Calculate the reconstruction error between the original image data and the initial image data, and update the time step prediction model according to the reconstruction error to obtain the updated time step prediction model; The calculating the reconstruction error between the original image data and the initial image data, and updating the time step prediction model according to the reconstruction error to obtain the updated time step prediction model specifically includes: Calculate the reconstruction error between the original image data and the initial image data, and determine whether the reconstruction error is greater than or equal to a preset threshold; If so, update the time step prediction model to obtain the updated time step prediction model; Wherein, the expression of the reconstruction error is: ; wherein, is the reconstruction error, is the initial image data, is the original image data, is the square of the second norm; The updating the time step prediction model to obtain the updated time step prediction model specifically includes: Obtain the opposite number of the reconstruction error, and use the opposite number as the reward signal, where the expression of the reward signal is: , where is the reward signal; Obtain historical generated data, and update the time step prediction model according to the historical generated data and the reward signal to obtain the updated time step prediction model; Obtaining the updated time step prediction model specifically includes: Use the RLOO method to optimize the time-step prediction model parameters using historical generated data and reward signals , and the optimization objective is as follows: ; Among them, is the optimization objective, is the optimization process, is the time-step prediction model, is the updated time step, is the conditional relationship, is the latent variable corresponding to the time step, is the time step; Obtain the updated text description, and perform reverse generation processing on the noise data, the intermediate features, and the updated text description through the updated time step prediction model and the diffusion model to obtain the target image data.

2. The method for image and video reconstruction based on a diffusion model according to claim 1, wherein The obtaining the original image data and the original text description, and inputting the original image data and the original text description into the diffusion model to obtain noise data and intermediate features specifically includes: Obtain the original image data and the original text description, and input the original image data and the original text description into the diffusion model; Perform forward diffusion processing on the original image data and the original text description through the diffusion model to obtain noise data; Determine a preset time step, extract the key-value cache of the attention mechanism in the diffusion model according to the preset time step, and obtain the intermediate features corresponding to the preset time step according to the key-value cache.

3. The method for image and video reconstruction based on a diffusion model according to claim 1, wherein The determining the time step prediction model, and performing reverse generation processing on the noise data, the intermediate features, and the original text description through the time step prediction model and the diffusion model to obtain the initial image data specifically includes: Input the noise data, the intermediate features, and the original text description into the diffusion model, and perform reverse generation processing on the noise data, the intermediate features, and the original text description through the diffusion model; Determine the time step prediction model, and perform prediction processing on each time step in the reverse generation processing through the time step prediction model; When any time step in the reverse generation process is 0, the reverse generation process is completed, and the diffusion model outputs the initial image data.

4. The method for image and video reconstruction based on a diffusion model according to claim 1, characterized in that, Calculating the reconstruction error between the original image data and the initial image data, and updating the time step prediction model according to the reconstruction error to obtain an updated time step prediction model, and then further including: Performing iterative reverse generation processing on the noise data, the intermediate features, and the original text description through the updated time step prediction model and the diffusion model until the reconstruction error is less than the preset threshold.

5. The method for image and video reconstruction based on a diffusion model according to claim 1, wherein, Obtaining the updated text description, and performing reverse generation processing on the noise data, the intermediate features, and the updated text description through the updated time step prediction model and the diffusion model to obtain the target image data, specifically including: Determining the user requirements, and performing text update processing on the original text description according to the user requirements to obtain the updated text description, where the text update processing includes keyword replacement and content detail modification; Inputting the updated text description, the noise data, and the intermediate features into the diffusion model again, and performing reverse generation processing on the updated text description, the noise data, and the intermediate features through the updated time step prediction model and the diffusion model to obtain the target image data.

6. An image and video reconstruction system based on a diffusion model, characterized in that, The image and video reconstruction system based on the diffusion model is used to implement the image and video reconstruction method based on the diffusion model according to any one of claims 1-5. The image and video reconstruction system based on the diffusion model includes: A forward diffusion processing module, configured to obtain the original image data and the original text description, and input the original image data and the original text description into the diffusion model to obtain the noise data and the intermediate features; A time step prediction model processing module, configured to determine the time step prediction model, and perform reverse generation processing on the noise data, the intermediate features, and the original text description through the time step prediction model and the diffusion model to obtain the initial image data; A time step prediction model update module, configured to calculate the reconstruction error between the original image data and the initial image data, and update the time step prediction model according to the reconstruction error to obtain an updated time step prediction model; A reverse generation processing module, configured to obtain the updated text description, and perform reverse generation processing on the noise data, the intermediate features, and the updated text description through the updated time step prediction model and the diffusion model to obtain the target image data.

7. A terminal, characterized in that, The terminal includes: a memory, a processor, and an image and video reconstruction program based on the diffusion model stored on the memory and executable on the processor. When the image and video reconstruction program based on the diffusion model is executed by the processor, the steps of the image and video reconstruction method based on the diffusion model according to any one of claims 1-5 are implemented.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an image and video reconstruction program based on a diffusion model. When the image and video reconstruction program based on the diffusion model is executed by a processor, the steps of the image and video reconstruction method based on the diffusion model according to any one of claims 1-5 are implemented.

Citation Information

Patent Citations

  • Brain decoding system based on functional magnetic resonance image and potential diffusion model

    CN119273783A

  • Medical image super-resolution method and device based on edge enhancement diffusion model

    CN119624779A