Consistency model creation method and system based on artificial intelligence
By combining consistency loss, segmentation loss, training influence analysis and time-step adjustment strategies, the training and inference process of the consistency model is optimized, and the problems of degraded generation quality, low training efficiency and large computing overhead are solved, and efficient and high-quality image generation is achieved.
Patent Information
- Application Number
- CN202510286621.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
During the training and inference process, the consistency model generates problems such as quality reduction, low training efficiency, large calculation overhead, and insufficient inference time step optimization.
By combining consistency loss with segmentation loss, training impact analysis and time-step adjustment strategies, the training stability, inference accuracy and computational efficiency of the model are optimized.
It improves the training stability and inference accuracy of the model, reduces the consumption of computing resources, and improves the quality and efficiency of image generation.
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method and system for creating a consistency model based on artificial intelligence. Background Art
[0002] The consistency model is a new type of generative model that can introduce consistency constraints during the training process, enabling the model to generate high-quality samples with fewer inference steps. This model is widely used in tasks such as image generation and speech synthesis, and demonstrates efficient and high-quality performance in the field of generative artificial intelligence, which is of great significance for improving the practicality of generative models.
[0003] Currently, diffusion models (Diffusion Models) have made significant progress in generation quality. A typical example is Stable Diffusion, which can generate high-quality images through large-scale supervised training and is widely used in fields such as AI drawing, object detection, and image segmentation. However, diffusion models require multiple time steps for sampling during the inference stage, resulting in high computational resource consumption and slow inference speed, making it difficult to meet scenarios with high real-time requirements.
[0004] To improve the inference speed, researchers have proposed latent consistency models (Latent Consistency Models). Through techniques such as consistency distillation, the model can generate high-quality samples in fewer time steps, thus significantly improving the inference efficiency. However, when reducing the inference steps, this model is difficult to accurately capture the complex distribution of data, resulting in problems such as blurred, missing, or misaligned details in the generated images. At the same time, the training of existing consistency models relies on consistency loss, the training process is relatively slow, and the convergence effect is unstable. In addition, the selection of different time steps during the inference process has a great impact on the final generated result. If the time steps are adjusted improperly, the generation quality may decline. Summary of the Invention
[0005] In view of the above existing problems, the present invention is proposed.
[0006] The present invention aims to solve problems such as the decline in generation quality, low training efficiency, high computational overhead, and insufficient optimization of inference time steps in the training and inference processes of consistency models. By combining consistency loss and segment loss, training influence analysis, and time step adjustment strategies, the training stability, inference accuracy, and computational efficiency of the model are improved, realizing efficient and high-quality image generation.
[0007] To solve the above technical problems, the present invention provides the following technical solutions: In the first aspect, an embodiment of the present invention provides a method for creating a consistency model based on artificial intelligence, which includes, Step S1: Obtain the original image and encode the original image using an encoder to obtain the encoded image data; Step S2: Perform noise addition processing on the encoded image data to generate images with multiple different noise levels; Step S3: Input the noisy image into the neural network in the consistency model for processing to obtain the corresponding image output; Step S4: Calculate the consistency loss to keep the images under different noise levels consistent during the generation process; Step S5: Calculate the segment loss and optimize the images for different time steps respectively to improve the generation quality; Step S6: Combine the consistency loss and the segment loss to construct an optimization objective function to improve the training stability of the model; Step S7: Use the training influence analysis method to calculate the contribution of training samples to the model and dynamically adjust the weight of the segment loss based on this; Step S8: During the inference process, use the time step adjustment strategy to optimize the inference process so that the model can generate more accurately; Step S9: Generate high-quality images through the optimized consistency model to achieve the consistency generation task.
[0008] As a preferred solution of the method for creating a consistency model based on artificial intelligence according to the present invention, wherein: the consistency model adopts a neural network structure and is used to maintain the consistency of images under different noise levels.
[0009] As a preferred solution of the method for creating a consistency model based on artificial intelligence according to the present invention, wherein: during the training process of the consistency model, an adversarial loss is further introduced to enhance the details and realism of the model generation results; The adversarial loss is obtained by adding a discriminator network and performing adversarial training with the generator of the consistency model, and is used to optimize the image generation quality and reduce artifacts.
[0010] As a preferred solution of the method for creating a consistency model based on artificial intelligence according to the present invention, wherein: the calculation of the segment loss includes: Obtain the image data for different time steps and calculate the differences between each time step, Map the features between time steps to ensure the consistency of the images, Combine the training influence analysis method to optimize the calculation method of the segment loss and improve the training stability.
[0011] As a preferred solution of the method for creating a consistency model based on artificial intelligence according to the present invention, wherein: the training influence analysis method is used to calculate the influence of training samples on the model and dynamically adjust the loss function during the training process.
[0012] As a preferred solution of the method for creating a consistency model based on artificial intelligence according to the present invention, wherein: the time step adjustment strategy includes: Calculating the deviation between the current time step and the ideal time step during the inference process, Determining the optimal time step length through a sliding window method, Dynamically adjusting the next inference time step according to the distribution of time steps.
[0013] As a preferred solution of the method for creating a consistency model based on artificial intelligence according to the present invention, wherein: the optimization objective function is jointly composed of a consistency loss and a piecewise loss, so as to improve the training efficiency while maintaining the generation quality.
[0014] In a second aspect, the present invention provides a system for creating a consistency model based on artificial intelligence, including, An encoding module, configured to encode the original image to generate encoded data; A noise addition module, configured to add different levels of noise to the encoded data; A consistency calculation module, including a neural network structure, configured to process images with different noise levels and generate corresponding output images; A loss calculation module, including a consistency loss calculation unit and a piecewise loss calculation unit, configured to optimize the training process; A dynamic adjustment module, configured to calculate the influence of training samples on the model and dynamically adjust the weight of the loss function based on this to improve the stability of training; An inference optimization module, configured to adjust the time step during the inference process to optimize the inference accuracy of the consistency model.
[0015] As a preferred solution of the system for creating a consistency model based on artificial intelligence according to the present invention, wherein: the consistency calculation module is based on a neural network structure and combines a consistency loss and a piecewise loss.
[0016] As a preferred solution of the system for creating a consistency model based on artificial intelligence according to the present invention, wherein: the inference optimization module adopts a time step adjustment strategy to optimize the accuracy during the multi-step inference process.
[0017] The beneficial effects of the present invention are as follows: After reducing the inference steps, the quality of the generated results by traditional consistency models may decline, resulting in problems such as blurriness or missing details. However, by combining consistency loss and segment loss, the present invention enables the model to maintain high consistency between different time steps, while optimizing the image features of intermediate states, improving the clarity and logic of the generated results, and reducing image distortion and artifacts.
[0018] Existing consistency models converge slowly during the training process, resulting in high training costs. However, the present invention dynamically adjusts the contribution of training samples to the model through a training influence analysis method, making the training process more adaptive, effectively reducing the time required for training, and improving the optimization efficiency of the model.
[0019] Traditional diffusion models and consistency models require a large amount of computing resources during the inference stage, especially when generating high-resolution images, the computational cost increases significantly. However, the present invention adopts a time step adjustment strategy during the inference process, dynamically optimizing the inference time steps, reducing redundant calculations, improving the utilization rate of computing resources, and reducing hardware requirements.
[0020] Since the adaptability of consistency models may vary in different tasks, the present invention enables the model to be dynamically optimized for different training datasets by introducing a training influence analysis method, improving the adaptability to diverse data, and thus enhancing the generalization ability of the consistency model in different application scenarios.
[0021] During the inference process, the time step selection of traditional consistency models is mostly fixed, which easily leads to unstable generated results. However, the present invention adopts a time step adjustment strategy to optimize the distribution of time steps based on the actual generation situation, making the inference process more accurate, and improving the detail expressiveness and overall consistency of the generated results.
[0022] Although existing consistency models can reduce the inference steps, they still need to optimize the output quality during the multi-step inference process. However, the present invention introduces a time offset sampler to accurately determine the optimal choice of the current time step, reducing computational redundancy, while ensuring that the generated results are more in line with the data distribution and improving the inference efficiency.
[0023] Traditional consistency loss usually only focuses on the final generated results, ignoring the learning effects of the model at different stages. However, the present invention combines segment loss, enabling the model to better learn the feature distributions at different stages, and improving the training stability and generation consistency.
[0024] By introducing adversarial loss during the training process, the present invention enhances the competition between the generator and the discriminator, improves the model's ability to handle details, makes the generated images more realistic, and enhances the practical application value of the model.
[0025] In summary, the present invention realizes an optimized method for creating a consistency model by combining consistency loss, segment loss, training influence analysis, and time step adjustment strategy. This method improves the training efficiency, generation quality, and computational resource utilization rate of the model while ensuring efficient inference. It is applicable to artificial intelligence application scenarios such as image generation, object detection, and image restoration, and has broad practical value. Detailed implementation manners
[0026] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following provides a detailed description of the specific implementation manners of the present invention.
[0027] In the following description, many specific details are set forth to facilitate a thorough understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0028] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it an individual or alternative embodiment that is mutually exclusive with other embodiments.
[0029] Based on the background problem, the present invention proposes an improved method and system for creating a consistency model. By introducing segment loss, training influence analysis method, and time step adjustment strategy, it improves the model training efficiency and inference accuracy, thereby optimizing the image generation quality, reducing the computational overhead, and enhancing the practical application value of the consistency model.
[0030] Embodiment 1 provides a training method for a consistency model based on artificial intelligence, including: First, prepare and load the data. The training loop processes the data in batches. For each batch, it loads the images and their corresponding text prompts. The images are transferred to an appropriate computing device (such as a GPU) to improve processing efficiency. The text prompts are encoded by a text encoder to generate embeddings for conditioning the model output.
[0031] Next, the images are encoded into the latent space. A variational autoencoder (VAE) is used to encode the images into latent representations, and at the same time, random time steps are sampled. During the segmentation process, the termination time step is also sampled. The code calculates the intermediate steps and the corresponding time steps according to the schedule of the solver (such as the DDIM schedule).
[0032] After that, the boundary scaling is calculated. And random noise is sampled from the standard normal distribution. According to the noise magnitude at each sampling time step, the sampled noise is added to the latent representation. This simulates the forward diffusion process, creating a noisy input for the model to denoise.
[0033] Subsequently, a random guidance scale is sampled. And when preparing the prompt embedding and conditions, embeddings are extracted from the encoded text prompt for conditioning the output of the model. Also, any additional conditions or context required for the model to run are prepared.
[0034] Using the prediction of the online student model and the previously calculated boundary scaling, the noisy input and the refined prediction are combined to generate the final output for this iteration. For the guided teacher model prediction, the teacher model (pre-trained and frozen) uses the noisy input and the conditional prompt embedding to predict the noise component. Also, the teacher model uses the unconditional embedding (e.g., the embedding representing an empty prompt) for prediction. In the ODE solver step, using the CFG estimate of the teacher model, the ODE solver performs one step to estimate the next point on the reverse diffusion trajectory. This simulates the process of gradually denoising over time.
[0035] When calculating the loss, first, the difference between the final prediction of the online student model and the final prediction of the target student model is calculated as the main loss. This loss encourages the online model to generate outputs consistent with the target model. Then, a piecewise consistency loss based on the difference between the original latent estimates of the online and target models is calculated. This piecewise consistency loss focuses on specific parts of the output to ensure fine-grained consistency.
[0036] In the influence calculation, the TracIn method is used to evaluate the impact of individual training samples on the model parameters by calculating the gradient of the piecewise consistency loss with respect to certain model parameters. By calculating the gradient and its magnitude, this method determines the degree of influence of a specific training sample on the learning process.
[0037] In the backpropagation and optimization phase, the total loss is backpropagated through the online student model to calculate the gradients of all learnable parameters. Then, the optimizer uses the calculated gradients and the learning rate schedule to update the parameters of the online student model.
[0038] The parameters of the target student model are updated using the exponential moving average of the online student model parameters. Also, in logging and checkpoint saving, the training progress is recorded, including metrics such as loss values and learning rates. The model is periodically evaluated on the validation set to assess its performance.
[0039] Example 2 provides a method for creating a consistency model based on artificial intelligence, aiming to optimize the training and inference processes of existing consistency models, and improve the quality of generated images and computational efficiency. This method includes three stages: data processing, model training, and inference optimization. (1) Data processing Obtain the original images, and use an encoder to encode the original images and convert them into feature representations, thereby reducing the data dimension and improving the model's computational efficiency.
[0040] Add noise to the encoded image data to generate multiple images with different noise levels, simulating the inputs at different time steps.
[0041] (2) Model training Input the noisy images into the neural network in the consistency model for processing to obtain the predicted output. This neural network adopts the Unet structure to enhance the image consistency at different noise levels.
[0042] Calculate the consistency loss to keep the input images at different time steps consistent during the generation process and improve the training stability.
[0043] Calculate the segment loss to locally optimize the images at different time steps, thereby enhancing the model's fitting ability to complex data distributions.
[0044] Combine the consistency loss and the segment loss to construct an optimized objective function to improve the training efficiency and generation quality.
[0045] Adopt a training influence analysis method to calculate the influence of training samples on the model, and dynamically adjust the weight of the segment loss based on this to make the optimization of the model more accurate at different training stages.
[0046] (3) Inference optimization During the inference process, adopt a time step adjustment strategy to dynamically optimize the time steps of model inference, reduce the computational overhead, and improve the inference accuracy.
[0047] Generate high-quality images through the optimized consistency model to achieve the consistency generation task.
[0048] Example 3 provides a method for implementing a consistency model system, aiming to optimize the training and inference processes of existing consistency models, and improve the quality of generated images and computational efficiency. This method includes three stages: data processing, model training, and inference optimization. (1) Composition of the consistency model system This system includes: Encoding module: used to receive the original images and convert them into low-dimensional feature representations to reduce the computational complexity.
[0049] Noise addition module: used to add different levels of noise to the feature representation to simulate the inputs at different time steps.
[0050] Consistency calculation module: includes a neural network structure that is used to generate consistent outputs at different time steps, improving the stability and generalization ability of the model.
[0051] Loss calculation module: includes a consistency loss calculation unit and a segmented loss calculation unit, used to optimize the training objective and improve the convergence speed of the model.
[0052] Dynamic adjustment module: used to calculate the impact of training samples on the model and dynamically adjust the weights of the loss function to make the training process more adaptive.
[0053] Inference optimization module: used to optimize the adjustment strategy of inference time steps to enable the model to achieve better computational efficiency and generation quality during the inference process.
[0054] (2) System working process 1. Data preprocessing: The system receives the input image, encodes it, and adds different levels of noise to generate a training dataset.
[0055] 2. Model training: Use a neural network structure for data processing and optimize the model parameters based on the consistency loss and segmented loss.
[0056] Through the training influence analysis method, optimize the calculation method of the segmented loss to improve the adaptability of the model at different training stages.
[0057] 3. Model inference: Adopt a time step adjustment strategy to dynamically optimize the time steps during the inference process to improve the generation accuracy of the model.
[0058] Generate high-quality images through the optimized neural network to complete the consistency generation task.
[0059] Example 4, this example provides a training influence analysis method. During the training process of the consistency model, different training samples have different impacts on the final model. Therefore, the present invention introduces a training influence analysis method to quantify the contribution of training samples to the model and dynamically adjust the weights of the loss function.
[0060] (1) Calculation process During the training process, track the contribution of each training sample to the change of model parameters to measure its impact on the training results of the model.
[0061] Use the gradient analysis method to calculate the importance of training samples and dynamically adjust the segmented loss weights in the loss function based on this importance.
[0062] (2)Optimization method When the training samples have a greater impact on the model, increase their weights in the loss function so that the model focuses on learning the features of these samples.
[0063] When the training samples have a smaller impact on the model, reduce their weights to minimize interference with the overall training process.
[0064] By dynamically adjusting the loss weights, improve the training efficiency of the model and enable it to converge to the optimal solution faster.
[0065] Example 5 provides an inference time step adjustment strategy. During the inference process of the consistency model, the selection of different time steps has a significant impact on the generation effect. Therefore, the present invention introduces a time step adjustment strategy to optimize the inference accuracy.
[0066] (1)Time step adjustment method Adopt a sliding window method to calculate the deviation between the current time step and the optimal time step.
[0067] Determine the optimal time step length through probability distribution calculation.
[0068] Dynamically adjust the next inference time step according to the distribution of the model generation results.
[0069] (2)Optimization effect By adjusting the time step, reduce the computational redundancy during the inference process and improve the computational efficiency.
[0070] On the premise of ensuring the generation quality, accelerate the inference speed and improve the adaptability in real-time application scenarios.
[0071] In summary, the present invention: Adopts a combination of consistency loss and segment loss to improve the stability of training and enable the model to learn the data distribution more efficiently.
[0072] Adopts a training influence analysis method to optimize the weights of training samples and improve the self-adaptability of the training process.
[0073] Adopts a time step adjustment strategy to optimize the selection of inference time steps and improve the inference accuracy and computational efficiency.
[0074] Based on the Stable Diffusion model and the latent consistency model, the consistency loss function is improved. Structural optimization analysis is carried out on the diffusion model and its derived consistency model. In the existing consistency models, usually a single consistency loss function is used to directly fit the generated images, or the model is trained by means of multiple piecewise fittings. This innovation point is analyzed through experiments, using the method of consistency loss + n times of piecewise fitting loss, where the consistency loss is the latent space tensor of the final image after two-stage solution, n is the solution obtained by the method in part 2 of the innovation point, and the piecewise fitting loss is the front control tensor of the intermediate state image after one-stage solution. By adopting the above method, the image of the improved consistency model is clearer and more logical.
[0075] Based on the method called TracIn (Training Example Influence), the influence of piecewise loss on the model is explained. This method explains the prediction results of the model on the test set through the loss and gradient of the model. In the present invention, the influence degree of each sample in the current training subset on the trained model, that is, the model after gradient update and backpropagation through a larger training set, is calculated, so as to explain the decision-making process of the model. First, the idealized influence is defined, that is, the change amount of the loss function of the model on the test set when a certain training example is used. Then, the first-order approximation formula based on the gradient descent algorithm is used, and TracIn is introduced. The method based on TracIn parameterizes the influence of piecewise loss and introduces it into the loss function.
[0076] Based on the method of Time-Shift Sampler, it is improved and applied to the multi-step inference of the consistency model. This method introduces a sampling window, and by calculating the probability distribution of each time step in the window and comparing it with the distribution of the actual obtained image latent space variables, the time step of the model is updated to make it more accurate.
[0077] In summary, the present invention can effectively improve the training efficiency and inference quality of the consistency model, and is applicable to tasks such as image generation, object detection, and image restoration.
[0078] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for creating a consistency model based on artificial intelligence, characterized in that: include, Step S1, obtaining an original image, and encoding the original image using an encoder to obtain encoded image data; Step S2, performing noise addition processing on the encoded image data to generate a plurality of images with different noise levels; Step S3, inputting the noisy image into the neural network in the consistency model for processing to obtain the corresponding image output; Step S4, calculating the consistency loss to make the images with different noise levels consistent during the generation process; Step S5, calculating the segmentation loss and optimizing the images at different time steps respectively; Step S6, combining consistency loss and segmentation loss to construct an optimization objective function; Step S7, using a training influence analysis method to calculate the contribution of the training sample to the model, and dynamically adjusting the weight of the segmentation loss based on this; Step S8, during the reasoning process, optimizing the reasoning process using a time step adjustment strategy; Step S9, generating a high-quality image through the optimized consistency model.
2. The method for creating a consistency model based on artificial intelligence as claimed in claim 1, characterized in that: The consistency model adopts a neural network structure to maintain the consistency of images under different noise levels.
3. The method for creating a consistency model based on artificial intelligence as claimed in claim 2, characterized in that: Adversarial loss is further introduced during the training process of the consistency model to enhance the details and realism of the model generation results; The adversarial loss is used to optimize image generation quality and reduce artifacts by adding a discriminator network to perform adversarial training with the generator of the consistency model.
4. The method for creating a consistency model based on artificial intelligence as claimed in claim 3, characterized in that: The calculation of the segment loss includes: Get image data at different time steps and calculate the difference between each time step, Mapping features between time steps, Combined with the training influence analysis method, the calculation method of segmentation loss is optimized.
5. The method for creating a consistency model based on artificial intelligence as claimed in claim 4, characterized in that: The training influence analysis method is used to calculate the impact of training samples on the model and dynamically adjust the loss function during the training process.
6. The method for creating a consistency model based on artificial intelligence as claimed in claim 5, characterized in that: The time step adjustment strategy includes: During inference, the deviation between the current time step and the ideal time step is calculated. The optimal time step is determined by a sliding window method, According to the distribution of time steps, the next inference time step is dynamically adjusted.
7. The method for creating a consistency model based on artificial intelligence according to claim 6, characterized in that: The optimization objective function is composed of consistency loss and segmentation loss.
8. A system for creating a consistency model based on artificial intelligence, based on the method for creating a consistency model based on artificial intelligence according to any one of claims 1 to 7, characterized in that: include: An encoding module, used for encoding the original image to generate encoded data; A noise adding module, used to add different levels of noise to the encoded data; A consistency calculation module, including a neural network structure, is used to process images with different noise levels and generate corresponding output images; The loss calculation module includes a consistency loss calculation unit and a segmentation loss calculation unit, which are used to optimize the training process; The dynamic adjustment module is used to calculate the impact of training samples on the model and dynamically adjust the weight of the loss function based on this to improve the stability of training; The reasoning optimization module is used to adjust the time step in the reasoning process and optimize the reasoning accuracy of the consistency model.
9. The system for creating a consistency model based on artificial intelligence as claimed in claim 8, characterized in that: The consistency calculation module is based on a neural network structure, combining consistency loss and segmentation loss.
10. The artificial intelligence-based consistency model creation system according to claim 9, characterized in that: The reasoning optimization module adopts a time step adjustment strategy to optimize the accuracy in the multi-step reasoning process.
Citation Information
Cited By
Image consistency model adaptive discretization method and device for optimizing visual angle inspiration, computer readable storage medium and computer program product
CN121072614A