Visual content generation method and device, equipment and medium

By using formula calculations in the visual content generation model instead of partial inference calculations, the problems of slow and high cost of visual content generation in the prior art are solved, and the inference acceleration and generation speed are achieved.

CN120070242APending Publication Date: 2025-05-30BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510220566.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, the generation speed of visual content based on the diffusion model is very slow, the generation cost is high, and it is difficult to meet the real-time generation needs.

Method used

By performing denoising processing of multiple denoising steps in the visual content generation model, formula calculations are used instead of partial inference calculations, saving inference calculation time and improving generation speed.

Benefits of technology

It has achieved a significant inference acceleration, improved the speed of visual content generation, reduced the generation cost, and improved the overall processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070242A_ABST
    Figure CN120070242A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a visual content generation method and device, equipment and a medium. The method comprises the steps that target input information is acquired; the target input information is input into the visual content generation model to execute denoising processing of a plurality of denoising steps, target visual content is output and obtained, and the denoising processing comprises reasoning calculation and formula calculation. According to the technical scheme, when the visual content generation model is used for executing the denoising processing of the multiple denoising steps based on the input information, the denoising processing comprises formula calculation, and in the partial denoising step, reasoning calculation is replaced by formula calculation, and the visual content is generated, so that the denoising step of partial reasoning calculation is saved, and the denoising efficiency is improved. According to the method, the inference acceleration is greatly realized, so that the visual content generation speed is effectively improved, and the generation cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method, apparatus, device, and medium for generating visual content. Background Art

[0002] With the development of technologies, visual content of images and videos can be generated through models, mainly based on the inference method of diffusion models. In related technologies, when generating visual content using a visual content generation model, the speed is very slow and the generation cost is relatively high, which needs to be improved. Summary of the Invention

[0003] To solve the above technical problems, the present disclosure provides a method, apparatus, device, and medium for generating visual content.

[0004] An embodiment of the present disclosure provides a method for generating visual content, the method including:

[0005] Obtaining target input information;

[0006] Inputting the target input information into a visual content generation model to perform denoising processing of multiple denoising steps, and outputting target visual content, where the denoising processing includes inference calculation and formula calculation.

[0007] An embodiment of the present disclosure further provides a visual content generation apparatus, the apparatus including:

[0008] An obtaining module, configured to obtain target input information;

[0009] A generating module, configured to input the target input information into a visual content generation model to perform denoising processing of multiple denoising steps, and output target visual content, where the denoising processing includes inference calculation and formula calculation.

[0010] An embodiment of the present disclosure further provides an electronic device, the electronic device including: a processor; a memory for storing executable instructions of the processor; the processor, configured to read the executable instructions from the memory and execute the executable instructions to implement the method for generating visual content provided in the embodiment of the present disclosure.

[0011] An embodiment of the present disclosure further provides a computer-readable storage medium, the storage medium storing a computer program, the computer program being used to execute the method for generating visual content provided in the embodiment of the present disclosure.

[0012] The technical solutions provided by the embodiments of the present disclosure have the following advantages compared with the prior art: The visual content generation solution provided by the embodiments of the present disclosure obtains target input information; inputs the target input information into a visual content generation model to perform denoising processing with multiple denoising steps, and outputs the target visual content, where the denoising processing includes inference calculation and formula calculation. By adopting the above technical solutions, when the visual content generation model performs denoising processing with multiple denoising steps based on the input information, since the denoising processing includes formula calculation, some denoising steps use formula calculation to replace inference calculation to generate visual content, thereby saving some denoising steps of inference calculation, achieving a significant acceleration of inference, and effectively improving the visual content generation speed and reducing the generation cost. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the original components and elements are not necessarily drawn to scale.

[0014] Figure 1 It is a flowchart of a method for generating visual content provided by an embodiment of the present disclosure;

[0015] Figure 2 It is a flowchart of another method for generating visual content provided by an embodiment of the present disclosure;

[0016] Figure 3 It is a flowchart of still another method for generating visual content provided by an embodiment of the present disclosure;

[0017] Figure 4 It is a schematic structural diagram of a device for generating visual content provided by an embodiment of the present disclosure;

[0018] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0020] It should be understood that the various steps described in the method embodiments of the present disclosure can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0021] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0022] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependent relationships.

[0023] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0024] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0025] In the related art, when generating visual content of an image or video based on the inference method of a diffusion model, by setting a relatively large number of denoising steps, a step-by-step denoising process is performed on the initialized noise to obtain the final image or video. However, the speed is very slow and the generation cost is relatively high. For example, to generate a 4s video of 720p, even on an 8-card A100 machine, without optimization, it takes 5 minutes to 20 minutes, the speed is very slow, and it is almost impossible to directly experience.

[0026] To solve the above problems, the embodiments of the present disclosure provide a visual content generation method, which will be introduced below in combination with specific embodiments.

[0027] Figure 1 It is a schematic flowchart of a visual content generation method provided by an embodiment of the present disclosure. This method can be executed by a visual content generation device, where the device can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 1 shown, this method includes:

[0028] Step 101, obtain target input information.

[0029] The visual content generation method according to the embodiments of the present disclosure can be executed by an application that supports a large model. The application can interact with the user to generate visual content. The large model according to the embodiments of the present disclosure is a visual content generation model. Visual content can be content mainly presented in a visual form, and specifically, visual content can include images and / or videos. In the embodiments of the present disclosure, the visual content is described by taking a video as an example. The target input information can be reference information input by the user currently for generating a visual content, and the target input information can include text, images, etc. For example, the target input information can include text information of "a person running" and a reference image of the first frame.

[0030] Specifically, the visual content generation device can obtain the target input information in response to the user's input operation.

[0031] Step 102: Input the target input information into the visual content generation model to perform denoising processing of multiple denoising steps, and output the target visual content. Among them, the denoising processing includes inference calculation and formula calculation.

[0032] Among them, the visual content generation model can be a model for generating visual content based on the user's input information. The visual content generation model can be a large model constructed based on a diffusion model. The diffusion model is a type of generative model that generates high-quality images or videos by gradually denoising a random noise vector. In the embodiments of the present disclosure, the visual content generation model can include a text-to-image model, a text-to-video model, an image-to-video model, an image-to-image model, etc. The specific type is determined according to the actual situation, and the number of visual content generation models is not limited. The denoising step can be a step for the visual content generation model to remove noise through multiple forward passes during the visual content generation process. The number of denoising steps is called the number of denoising steps. The target visual content can be an image or video corresponding to the target input information generated by using the visual content generation model and some denoising steps including formula calculation. In the generation process of the target visual content, not all are inference processes, but formula calculation is used to replace part of the inference process to achieve inference acceleration.

[0033] Specifically, after the visual content generation device obtains the target input information, it can generate a prompt word for the target input information, input the prompt word into the visual content generation model to perform formula calculation on the steps to be optimized in multiple denoising steps, and perform inference calculation on other steps except the steps to be optimized in multiple denoising steps, and output the target visual content.

[0034] Among them, the step to be optimized can be a step in multiple denoising steps that can skip the inference process and perform denoising processing through formula calculation. The step to be optimized can be determined according to the fitting error between the current step and the previous step, and the number of steps to be optimized can be one or more. Other steps except the step to be optimized in the multiple denoising steps can perform denoising processing according to the original inference calculation.

[0035] Exemplarily, Figure 2 is a schematic flowchart of another visual content generation method provided by an embodiment of the present disclosure. As Figure 2 shown, in a feasible implementation manner, inputting the target input information into the visual content generation model to perform formula calculation on the step to be optimized in multiple denoising steps, and performing inference calculation on other steps except the step to be optimized in the multiple denoising steps, and outputting the target visual content may include the following steps:

[0036] Step 201: Sequentially determine multiple denoising steps as steps to be processed for the target input information through the visual content generation model.

[0037] Among them, the step to be processed can be the denoising step that the visual content generation model is currently performing.

[0038] Specifically, after the visual content generation device inputs the target input information into the visual content generation model, the visual content generation model can perform denoising processing on multiple denoising steps. Specifically, when processing, multiple denoising steps can be sequentially determined as steps to be processed for processing according to the execution order.

[0039] Step 202: For the step to be processed, determine whether the step to be processed is a step to be optimized. If so, execute step 203; otherwise, execute step 204.

[0040] Specifically, the visual content generation device can use the input error between the step to be processed and the previous denoising step to fit and estimate the output error for the step to be processed to obtain an estimated output error, and determine whether the step to be processed is a step to be optimized according to the estimated output error. If so, execute step 203; otherwise, execute step 204. The formula for the above fitting and estimation can be set according to the actual situation.

[0041] In some embodiments, determining whether a step to be processed is a step to be optimized may include: determining the input error between the step to be processed and the previous denoising step; determining the estimated output error between the step to be processed and the previous denoising step based on the input error, the number of denoising steps of multiple denoising steps, the denoising order of the step to be processed, and a pre-determined output error estimation formula; and determining whether the step to be processed is a step to be optimized based on the estimated output error and an inference skip threshold. Optionally, determining whether the step to be processed is a step to be optimized based on the estimated output error and the inference skip threshold may include: determining the current cumulative output error based on the estimated output error and the previous cumulative output error; determining whether the current cumulative output error is less than the inference skip threshold, and if so, determining that the step to be processed is a step to be optimized; otherwise, determining that the step to be processed is not a step to be optimized.

[0042] The input error may be the relative input error between the input fusion information of the step to be processed and the input fusion information of the previous denoising step. The determination of the relative input error may be determined, for example, by the following formula: Rel in =(mod i -mod i-1 ).abs.mean / (mod i-1 ).abs.mean, where Rel in represents the input error between the i-th denoising step (i.e., the step to be processed) and the (i - 1)-th denoising step (the previous denoising step), mod i represents the input fusion information of the i-th denoising step, mod i-1 represents the input fusion information of the (i - 1)-th denoising step, abs represents taking the absolute value, mean represents taking the average value, and calculating the relative error can avoid the defect of poor generalization performance caused by the weight change of the model during the calculation of the absolute error.

[0043] Here, the input fusion information of a denoising step may be the information obtained by fusing the input information of the denoising step and the time step information, and the fusion method is the same as that of the visual content generation model. Exemplarily, the input fusion information of the i-th denoising step may be obtained by the following fusion formula: mod i =func(In i ,ts i ), where mod i represents the input fusion information of the i-th denoising step, In i represents the input information of the i-th denoising step, ts i represents the time step information of the i-th denoising step, which may be the time step embedding feature (timestep embedding) obtained by processing the time step through the embedding layer.

[0044] The denoising order of the step to be processed can be the specific execution position of the step to be processed when multiple denoising steps are executed, that is, the order of execution. The step to be processed can be the position corresponding to the step to be processed in the sorting result after sorting multiple denoising steps according to the iterative order or execution order. The output error prediction formula can be a formula used to fit and calculate the output error according to the input error for two adjacent denoising steps. The specific formula adopted by the output error prediction formula can be set according to the actual situation. In the embodiments of the present disclosure, the output error prediction formula adopts a polynomial fitting formula as an example, and the input information, denoising order, and denoising steps of the denoising step are considered simultaneously during the fitting of the output error prediction formula, avoiding the problem that different output values need to be mapped for the same input during the fitting process between input and output.

[0045] The predicted output error is the output error calculated based on the input error, denoising steps, and denoising order of the step to be processed and the previous denoising step. It is a predicted error, not the real error. The inference skip threshold can be the threshold of the output error set for a denoising step to determine whether its inference process can be skipped, that is, the threshold to determine whether the denoising step is a step to be optimized. The specific value is set according to the actual situation. The cumulative output error can be the error obtained by cumulatively adding the predicted output errors of each denoising step in the execution order when multiple consecutive denoising steps are steps to be optimized. By comparing the current cumulative output error of a denoising step with the inference skip threshold, it is determined whether the denoising step is a step to be optimized, avoiding the inaccuracy of the output result caused by the error calculated by the formula being very small for multiple consecutive denoising steps. The previous cumulative output error can be the cumulative output error determined when the previous denoising step of the step to be processed is a step to be optimized. When the previous denoising step is other steps that are not steps to be optimized, the previous cumulative output error is 0; the current cumulative output error can be the cumulative output error obtained by adding the previous cumulative output error and the predicted output error for the step to be processed.

[0046] Specifically, when the visual content generation device determines whether the step to be processed is a step to be optimized, it can first subtract the input fusion information of the previous denoising step from the input fusion information of the step to be processed to obtain an input error. Then, the input error, the number of denoising steps, and the denoising order of the step to be processed are input into the input-output error prediction formula for calculation to obtain the predicted output error between the step to be processed and the previous denoising step. After that, the sum of the predicted output error and the previous cumulative output error can be determined to obtain the current cumulative output error. It is judged whether the current cumulative output error is less than the inference skip threshold. If so, it means that the error between the predicted step to be processed and the output information of the previous denoising step is small, and the inference process can be skipped, and the step to be processed is a step to be optimized. If the current cumulative output error is greater than or equal to the inference skip threshold, it is determined that the step to be processed is not a step to be optimized.

[0047] Step 203, perform formula calculation in the step to be processed.

[0048] Formula calculation can be a calculation method for determining output information through a formula for a denoising step. The specific formula used can be set according to the actual situation. For example, a linear formula can be used.

[0049] Specifically, when the visual content generation device performs formula calculation in the step to be processed, in response to the step to be processed being a step to be optimized, it can calculate the output information corresponding to the step to be processed based on the output information of the previous denoising step and the previous output error using a preset linear formula in the step to be processed. After that, the current cumulative output error can be updated to the previous cumulative output error.

[0050] The above preset linear formula can be set according to the actual situation. For example, it can be a regression linear formula, such as it can be expressed as hs i =hs i-1 +diff hs , where hs i represents the output information of the step to be processed, hs i-1 represents the output information of the previous denoising step. The output information of the previous denoising step can be obtained through formula calculation or inference calculation. diff hs represents the previous output error. The previous output error can be the output error between the output information of the previous denoising step and the output information of the step before the previous denoising step. Specifically, it can be expressed as diff hs =hs i-1 -hs i-2 , hs i-2 represents the output information of the step before the previous denoising step.

[0051] In the above solution, when a denoising step determines that the inference calculation can be skipped based on the estimated cumulative output error and the inference skip threshold, the output can be determined by using a fast formula calculation method. Compared with the inference calculation, this method greatly improves the calculation efficiency, reduces the cost, helps to improve the processing efficiency of the entire model, and significantly reduces the inference time of the visual content generation model.

[0052] After step 203, step 205 can be executed.

[0053] Step 204, perform inference calculation on the step to be processed.

[0054] The inference calculation can be the calculation of the inference process of the model through the denoising module set by the visual content generation model for a denoising step. Similar to the inference process of the denoising step in the related technology, it will not be elaborated here.

[0055] Specifically, when the visual content generation device performs formula calculation on the step to be processed, in response to the step to be processed being other steps that are not steps to be optimized, inference calculation can be performed to obtain the output information of the step to be processed, and then the previous cumulative output error can be updated to 0.

[0056] After step 204, step 205 can be executed.

[0057] Step 205, until all denoising steps are completed, and the target visual content is output.

[0058] Since in the process of generating visual content by the visual content generation model, the feature changes of adjacent denoising steps in the early and late stages of denoising are very large, while in the middle iteration process, the output feature maps of different denoising steps have small differences. However, when directly reducing the number of denoising steps, the generation effect will also be very poor.

[0059] In this solution, for multiple denoising steps, the output error is fitted and estimated based on the input error between the current denoising step and the previous denoising step, and its cumulative error is used as the judgment criterion to determine whether to use the formula calculation method to quickly determine the output of the current denoising step based on the output of the previous denoising step. Some denoising steps use formula calculation instead of inference calculation, effectively improving the calculation speed, so as to achieve the goal of inference acceleration.

[0060] The visual content generation solution provided by the embodiments of the present disclosure obtains target input information; inputs the target input information into a visual content generation model for denoising processing that performs multiple denoising steps, and outputs the target visual content, where the denoising processing includes inference calculation and formula calculation. By adopting the above technical solution, when the visual content generation model performs denoising processing with multiple denoising steps based on the input information, since the denoising processing includes formula calculation, some denoising steps use formula calculation to replace inference calculation to generate visual content, thereby saving some denoising steps of inference calculation, achieving a significant acceleration of inference, effectively improving the visual content generation speed, and reducing the generation cost.

[0061] In some embodiments, the visual content generation method may further include: determining the coefficients of the output error estimation formula of the visual content generation model. The number of visual content generation models may be one or more, which is specifically determined according to the actual situation. The coefficient is a key component in the output error estimation formula, which determines the shape and characteristics of the formula, and is a constant or variable used for calculation, representing a numerical value of the relationship.

[0062] Exemplarily, Figure 3 is a flowchart of another visual content generation method provided by the embodiments of the present disclosure. As Figure 3 shown, in a feasible implementation manner, determining the coefficients of the output error estimation formula of the visual content generation model may include the following steps:

[0063] Step 301, obtain multiple calibration input information.

[0064] Among them, the calibration input information may be the input information for generating visual content used in the calibration process. Here, the calibration process refers to the process of fitting and determining the coefficients of the output error estimation formula. The specific number of calibration input information is not limited and is determined according to the actual situation. For example, the number of calibration input information may be 100.

[0065] Step 302, input each calibration input information into the visual content generation model for inference calculation of multiple denoising steps, and obtain multiple step information of multiple denoising steps, where each step information includes input fusion information and output information.

[0066] The step information may be the detailed information recorded for a denoising step, which may specifically include the input fusion information and output information of this denoising step. The input fusion information of a denoising step may be the information obtained by fusing the input information and the time step information of this denoising step, and the fusion method is the same as the fusion method of the visual content generation model.

[0067] Specifically, after the visual content generation device obtains multiple calibration input information, for each calibration input information, it can input the calibration input information into the visual content generation model, use the visual content generation model to perform inference calculations for multiple denoising steps. The entire process is an inference process of the large model, obtaining the corresponding calibrated visual content, and extracting the step information of each denoising step during the inference process to obtain multiple step information of multiple denoising steps.

[0068] Step 303: Determine the input error and output error of two adjacent steps of multiple denoising steps of each calibration input information to obtain multiple error pairs. Among them, each error pair includes the input error and output error of the subsequent denoising step and the previous denoising step in the adjacent steps.

[0069] The error pair can include the input error and output error of two adjacent denoising steps. The input error can be the relative input error between the input fusion information of the subsequent denoising step and the input fusion information of the previous denoising step in the adjacent denoising steps. The determination of the relative input error can be determined, for example, by the following formula: Rel in =(mod i -mod i-1 ).abs.mean / (mod i-1 ).abs.mean, where Rel in represents the input error of the i-th denoising step (i.e., the subsequent denoising step) and the (i - 1)-th denoising step (i.e., the previous denoising step), mod i represents the input fusion information of the i-th denoising step, mod i-1 represents the input fusion information of the (i - 1)-th denoising step, abs represents taking the absolute value, and mean represents taking the average value; the output error can be the relative output error between the output information of the subsequent denoising step and the output information of the previous denoising step in two adjacent denoising steps, and can be determined by the following formula: Rel out =(hs i -hs i-1 ).abs.mean / (hs i-1 ).abs.mean, where Rel out represents the output error of the i-th denoising step (i.e., the subsequent denoising step) and the (i - 1)-th denoising step (i.e., the previous denoising step), hs i represents the output information of the i-th denoising step, hs i-1 represents the output information of the (i - 1)-th denoising step, abs represents taking the absolute value, and mean represents taking the average value. Multiple denoising steps can obtain the number of error pairs minus one. For example, 100 denoising steps can obtain 99 error pairs.

[0070] Specifically, for multiple denoising steps corresponding to each calibration input information, the visual content generation device extracts two adjacent steps, determines the input error and output error based on the step information of the two adjacent steps to obtain an error pair, and further obtains multiple error pairs, and determines corresponding multiple error pairs for each calibration input information.

[0071] Step 304: Based on multiple error pairs of each calibration input information among multiple calibration input information, perform polynomial fitting calculation on the output error prediction formula of the visual content generation model to obtain the coefficients of the output error prediction formula.

[0072] The output error prediction formula can be a formula for fitting and calculating the output error according to the input error for two adjacent denoising steps. The specific formula adopted by the output error prediction formula can be set according to the actual situation. In this embodiment of the present disclosure, the output error prediction formula adopts a polynomial fitting formula as an example.

[0073] Specifically, after the visual content generation device determines the error pairs of multiple calibration input information, for each error pair, based on the input error and the output error prediction formula in the error pair, calculate the fitted output error. Based on all error pairs and the fitted output error of each error pair, perform polynomial fitting calculation using the least squares method. Finally, the coefficients of the output error prediction formula can be obtained. The output error prediction formula can be expressed as Rel′ out = poly(Rel in , n a ) + poly(i / S, n b ), where Rel out represents the predicted output error between the i-th denoising step and the i - 1-th denoising step, poly represents polynomial fitting calculation, Rel in represents the input error between the i-th denoising step and the i - 1-th denoising step, i represents the denoising order of the i-th denoising step, S represents the number of denoising steps, n a and n b are respectively the powers of polynomial fitting regarding the input fusion information, denoising order, and number of denoising steps. According to all error pairs of multiple calibration input information for polynomial fitting, coefficients a and b can be obtained. The coefficients a and b are only examples with a power of 1, and the specific number may vary according to the power of the output error prediction formula. For example, when the output error prediction formula is a cubic polynomial, y = ax 3 + bx 2 + cx + d, the coefficients may include a, b, c, and d.

[0074] In the above solution, the coefficients of the output error prediction formula are determined by fitting. When fitting, the input information, denoising order, and denoising steps are taken into account simultaneously, thus avoiding the problem that different output values need to be mapped for the same input during the fitting process between input and output. Moreover, the above fitting can fit multiple visual content generation models in a series based on a batch of data at one time, thus avoiding the complex process of adjusting the weights of different models and simplifying the inference steps.

[0075] Figure 4 FIG. 4 is a schematic structural diagram of a visual content generation device provided by an embodiment of the present disclosure. The device can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 4 shown, the device includes:

[0076] An acquisition module 401, configured to acquire target input information;

[0077] A generation module 402, configured to input the target input information into a visual content generation model to perform denoising processing for multiple denoising steps, and output target visual content, where the denoising processing includes inference calculation and formula calculation.

[0078] Optionally, the generation module 402 is configured to:

[0079] Input the target input information into the visual content generation model to perform the formula calculation for the steps to be optimized in multiple denoising steps, and perform the inference calculation for the other steps except the steps to be optimized in multiple denoising steps, and output the target visual content.

[0080] Optionally, the generation module 402 specifically includes:

[0081] A first unit, configured to sequentially determine the multiple denoising steps as steps to be processed for the target input information through the visual content generation model;

[0082] A second unit, configured to, for the step to be processed, determine whether the step to be processed is a step to be optimized. If so, perform the formula calculation for the step to be processed; otherwise, perform the inference calculation for the step to be processed;

[0083] A third unit, configured to output the target visual content until all the multiple denoising steps are completed.

[0084] Optionally, the second unit includes:

[0085] A first sub-unit, configured to determine the input error between the step to be processed and the previous denoising step;

[0086] A second sub-unit, configured to determine a predicted output error between the to-be-processed step and the previous denoising step based on the input error, the number of denoising steps of the multiple denoising steps, the denoising order of the to-be-processed step, and a pre-determined output error prediction formula.

[0087] A third sub-unit, configured to determine whether the to-be-processed step is a step to be optimized based on the predicted output error and an inference skip threshold.

[0088] Optionally, the third sub-unit is configured to:

[0089] Determine a current cumulative output error based on the predicted output error and the previous cumulative output error;

[0090] Determine whether the current cumulative output error is less than the inference skip threshold. If so, determine that the to-be-processed step is a step to be optimized; otherwise, determine that the to-be-processed step is not a step to be optimized.

[0091] Optionally, the second unit further includes a fourth sub-unit, configured to:

[0092] Calculate the output information corresponding to the to-be-processed step based on the output information of the previous denoising step and the previous output error by using a preset linear formula at the to-be-processed step.

[0093] Optionally, the apparatus further includes a calibration module, configured to:

[0094] Obtain a plurality of calibration input information;

[0095] Input each of the calibration input information into the visual content generation model for inference calculation of multiple denoising steps, to obtain multiple step information of the multiple denoising steps, where each piece of step information includes input fusion information and output information;

[0096] Determine the input error and output error between two adjacent steps of the multiple denoising steps of each of the calibration input information, to obtain a plurality of error pairs, where each error pair includes the input error and output error of the subsequent denoising step and the previous denoising step in the adjacent steps;

[0097] Perform polynomial fitting calculation on the output error prediction formula of the visual content generation model based on the multiple error pairs of each calibration input information in the plurality of calibration input information, to obtain the coefficients of the output error prediction formula.

[0098] The visual content generation apparatus provided in the embodiments of the present disclosure can execute the visual content generation method provided in any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.

[0099] Embodiments of the present disclosure also provide a computer program product, including computer programs / instructions which, when executed by a processor, implement the visual content generation method provided by any embodiment of the present disclosure.

[0100] Specifically, referring to Figure 5 , which shows a schematic structural diagram of an electronic device 500 suitable for implementing embodiments of the present disclosure. The electronic device 500 in embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable media players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of embodiments of the present disclosure.

[0101] As Figure 5 shown, the electronic device 500 may include a processing device 501 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0102] Generally, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 5 the electronic device 500 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be implemented or included alternatively.

[0103] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by a processing device 501, the above functions defined in the visual content generation method of the embodiment of the present disclosure are performed.

[0104] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination of the above.

[0105] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as the HyperText Transfer Protocol (HTTP), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0106] The above computer-readable medium can be included in the above electronic device; it can also exist separately and not be assembled into the electronic device.

[0107] The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: obtain target input information; input the target input information into a visual content generation model to perform denoising processing of multiple denoising steps, and output target visual content, where the denoising processing includes inference calculation and formula calculation.

[0108] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++; and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network or a wide area network, or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0109] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0110] The units described in the embodiments of the present disclosure can be implemented in software or in hardware. Among them, the name of the unit does not, in some cases, constitute a limitation on the unit itself.

[0111] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field-Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.

[0112] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories, flash memories, optical fibers, portable compact disk read-only memories, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0113] It is to be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the information involved in the present disclosure should be informed to users and the authorization of the users should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0114] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0115] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0116] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for generating visual content, characterized in that: include: Get target input information; The target input information is input into a visual content generation model to perform denoising processing of multiple denoising steps, and the target visual content is obtained as output, wherein the denoising processing includes inference calculation and formula calculation.

2. The method according to claim 1, characterized in that: The target input information is input into the visual content generation model to perform denoising processing of multiple denoising steps, and the target visual content is output, including: The target input information is input into the visual content generation model, the formula calculation is performed on the step to be optimized in the multiple denoising steps, and the reasoning calculation is performed on the other steps in the multiple denoising steps except the step to be optimized, and the target visual content is output.

3. The method according to claim 2, characterized in that Inputting the target input information into the visual content generation model, performing the formula calculation on the step to be optimized in the multiple denoising steps, and performing the inference calculation on the other steps in the multiple denoising steps except the step to be optimized, and outputting the target visual content, including: Determining the plurality of denoising steps as steps to be processed in sequence for the target input information by using the visual content generation model; For the step to be processed, determine whether the step to be processed is a step to be optimized, if so, perform the formula calculation in the step to be processed; otherwise, perform the inference calculation in the step to be processed; Until all the multiple denoising steps are completed, the target visual content is output.

4. The method according to claim 3, characterized in that Determining whether the step to be processed is a step to be optimized includes: Determine the input error between the step to be processed and the previous denoising step; Determine an estimated output error between the step to be processed and the previous denoising step based on the input error, the number of denoising steps of the multiple denoising steps, the denoising order of the step to be processed, and a predetermined output error estimation formula; Based on the estimated output error and the inference skip threshold, it is determined whether the step to be processed is a step to be optimized.

5. The method according to claim 4, characterized in that Judging whether the step to be processed is a step to be optimized based on the estimated output error and the inference skip threshold includes: Determining a current accumulated output error based on the estimated output error and a previous accumulated output error; It is determined whether the current accumulated output error is less than the inference skip threshold; if so, it is determined that the step to be processed is a step to be optimized; otherwise, it is determined that the step to be processed is not a step to be optimized.

6. The method according to claim 3, characterized in that Executing the formula calculation in the pending step includes: In the step to be processed, the output information corresponding to the step to be processed is calculated using a preset linear formula based on the output information of the previous denoising step and the previous output error.

7. The method according to claim 1, characterized in that The method further comprises: Obtain multiple calibration input information; Inputting each of the calibration input information into the visual content generation model to perform inference calculations of multiple denoising steps to obtain multiple step information of the multiple denoising steps, wherein each of the step information includes input fusion information and output information; Determine the input errors and output errors of two adjacent steps of the plurality of denoising steps of each calibration input information to obtain a plurality of error pairs, wherein each error pair includes the input error and output error of a subsequent denoising step and a previous denoising step in the adjacent steps; Based on multiple error pairs of each calibration input information in the multiple calibration input information, a polynomial fitting calculation is performed on the output error prediction formula of the visual content generation model to obtain coefficients of the output error prediction formula.

8. A visual content generation device, characterized in that: include: An acquisition module is used to obtain target input information; A generation module is used to input the target input information into a visual content generation model to perform a denoising process of multiple denoising steps, and output the target visual content, wherein the denoising process includes inference calculation and formula calculation.

9. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the executable instructions to implement the visual content generation method described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and the computer program is used to execute the visual content generation method described in any one of claims 1 to 7.

Citation Information

Cited By

  • Diffusion model generation acceleration method and device, electronic equipment and storage medium

    CN121233891A