Information processing method and apparatus, and device and storage medium

By performing progressive knowledge distillation training on the diffusion model, segmented training and combining human feedback, the problem of high resource and time consumption in the diffusion model generation process is solved, and an efficient single-step denoising process is achieved.

WO2025214019A1PCT designated stage Publication Date: 2025-10-16BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/081149
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-11
Filing Date
2025-03-06
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Existing diffusion models require multiple steps of denoising during the generation process, resulting in high resource consumption and long time.

Method used

The second diffusion model is trained through progressive knowledge distillation based on the reference state of the pre-trained diffusion model. The training process is segmented and combined with human feedback and evaluation information to gradually approach the target diffusion model.

Benefits of technology

The model training efficiency is improved, and a target diffusion model that can generate results in a single-step denoising process is obtained, which reduces resource consumption and time cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025081149_16102025_PF_FP_ABST
    Figure CN2025081149_16102025_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure relate to an information processing method and apparatus, and a device and a storage medium. The method provided herein comprises: on the basis of a first set of reference states associated with a pre-trained first diffusion model, executing a first training process on a second diffusion model, wherein the first set of reference states corresponds to denoising results of a first set of denoising stages among a plurality of denoising stages of the first diffusion model (210); on the basis of a second set of reference states associated with the first diffusion model, executing a second training process on the second diffusion model, wherein the second set of reference states corresponds to denoising results of a second set of denoising stages among the plurality of denoising stages of the first diffusion model, and the quantity of the second set of denoising stages is less than that of the first set of denoising stages (220); and at least on the basis of the first training process and the second training process of the second diffusion model, acquiring a target diffusion model (230). Thus, the embodiments of the present disclosure can improve the model training efficiency by means of a progressive distillation process.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, equipment and storage medium for information processing

[0001] The present application claims priority to the Chinese patent application No. 202410437727.X, filed on April 11, 2024, entitled “Method, device, equipment and storage medium for information processing”, the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The example embodiments of the present disclosure generally relate to the field of computer, and in particular, to a method, device, equipment and computer readable storage medium for information processing. BACKGROUND

[0003] With the development of computer technology, various models are gradually applied to various aspects of people's daily life. For example, diffusion models can be used to generate images based on text. Generally speaking, diffusion models need to use a multi-step denoising process to perform the generation process, which consumes a lot of resources and a long time. SUMMARY

[0004] In a first aspect of the present disclosure, a method for information processing is provided. The method comprises: performing a first training process on a second diffusion model based on a first set of reference states associated with a pre-trained first diffusion model, the first set of reference states corresponding to denoising results of a first set of denoising stages in a plurality of denoising stages of the first diffusion model; performing a second training process on the second diffusion model based on a second set of reference states associated with the first diffusion model, the second set of reference states corresponding to denoising results of a second set of denoising stages in the plurality of denoising stages of the first diffusion model, wherein the number of the second set of denoising stages is less than the first set of denoising stages; and obtaining a target diffusion model based on at least the first training process and the second training process of the second diffusion model.

[0005] In a second aspect of the present disclosure, a device for information processing is provided. The device comprises: a first training module configured to perform a first training process on a second diffusion model based on a first set of reference states associated with a pre-trained first diffusion model, the first set of reference states corresponding to denoising results of a first set of denoising stages in a plurality of denoising stages of the first diffusion model; a second training module configured to perform a second training process on the second diffusion model based on a second set of reference states associated with the first diffusion model, the second set of reference states corresponding to denoising results of a second set of denoising stages in the plurality of denoising stages of the first diffusion model, wherein the number of the second set of denoising stages is less than the first set of denoising stages; and a model obtaining module configured to obtain a target diffusion model based on at least the first training process and the second training process of the second diffusion model.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium has stored thereon a computer program, the computer program being executable by a processor to implement the method of the first aspect.

[0008] In a fifth aspect of the present disclosure, a computer program product is provided. The computer program product includes computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method of the first aspect.

[0009] It should be understood that nothing in the Summary is intended to limit the scope of the embodiments of the present disclosure or the claims herein. Other aspects, features, and advantages of the present disclosure will become apparent to those of ordinary skill in the art from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other features, aspects, and advantages of embodiments of the present disclosure will become more apparent from the following description in conjunction with the accompanying drawings. In the drawings:

[0011] FIG. 1 illustrates an example environment capable of implementing in real time according to some embodiments of the present disclosure;

[0012] FIG. 2 illustrates a flowchart of an information processing process according to some embodiments of the present disclosure;

[0013] FIGS. 3A and 3B illustrate schematic diagrams of a training process according to some embodiments of the present disclosure;

[0014] FIG. 4 illustrates a schematic structural block diagram of an example apparatus for information processing according to some embodiments of the present disclosure; and

[0015] FIG. 5 illustrates a block diagram of an electronic device capable of implementing a plurality of embodiments of the present disclosure. DETAILED DESCRIPTION

[0016] Embodiments of the present disclosure will be described herein below with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein, but rather, the embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are merely for illustrative purposes and are not intended to limit the scope of protection of the present disclosure.

[0017] It should be noted that the headings provided herein are for convenience only and are not to be construed as limiting. Various embodiments are described herein, and any type of embodiment can be included under any heading. Further, embodiments described in any heading can be combined with any other embodiment described in the same heading and / or in a different heading in any manner.

[0018] In the description of embodiments of the disclosure, the term "includes" and its derivatives, such as "including" should be understood as open-ended, i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions can also be included below. The terms "first", "second", etc. can refer to different or the same objects. Other explicit and implicit definitions can also be included below.

[0019] Data related to users, acquisition and / or use of data, etc. can be involved in embodiments of the disclosure. These aspects all comply with corresponding laws and regulations and relevant provisions. In embodiments of the disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user is aware of and confirms. Accordingly, when implementing various embodiments of the disclosure, the type of data or information that can be involved, the scope of use, the use scenario, etc. should be notified to the user and the authorization of the user should be obtained according to relevant laws and regulations through appropriate means. The specific notification and / or authorization method can vary according to the actual situation and application scenario, and the scope of the disclosure is not limited in this respect.

[0020] In the specification and embodiments of the present disclosure, if personal information processing is involved, it will be processed on the premise of legality (for example, obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and only within the scope of the provisions or agreement. The user refuses to process personal information other than the necessary information required for basic functions, which does not affect the user's use of basic functions.

[0021] As briefly mentioned above, the diffusion model usually needs to perform a multi-step noise adding and denoising process during the training process. Accordingly, it also needs to perform a multi-step denoising process in the generation phase, which consumes a lot of resources and takes a long time.

[0022] Embodiments of the present disclosure propose a scheme for information processing. According to the scheme, a first training process is performed on a second diffusion model based on a first set of reference states associated with a pre-trained first diffusion model, the first set of reference states corresponding to denoising results of a first set of denoising stages among a plurality of denoising stages of the first diffusion model; a second training process is performed on the second diffusion model based on a second set of reference states associated with the first diffusion model, the second set of reference states corresponding to denoising results of a second set of denoising stages among the plurality of denoising stages of the first diffusion model, wherein a number of the second set of denoising stages is less than the first set of denoising stages; and a target diffusion model is obtained based at least on the first training process and the second training process of the second diffusion model.

[0023] In this way, embodiments of the present disclosure can improve the efficiency of model training through a progressive distillation process to obtain a target diffusion model with higher execution efficiency.

[0024] Various example implementations of the scheme are described in further detail below in conjunction with the accompanying drawings.

[0025] Example Environment

[0026] FIG. 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. As shown in FIG. 1, the example environment 100 can include an electronic device 110.

[0027] In this example environment 100, the electronic device 110 can train a student diffusion model 130 based on a teacher diffusion model 120. In some embodiments, the student diffusion model 130 can have fewer inference steps than the teacher diffusion model 120.

[0028] As an example, the student diffusion model 130 may, for example, have the ability to perform single-step inference by knowledge distillation of the teacher diffusion model 120.

[0029] The specific process of training the student diffusion model 130 will be described in detail below.

[0030] Example Process

[0031] An information processing process according to some embodiments of the present disclosure will be described below with reference to FIG. 2. FIG. 2 illustrates a flowchart of an example process 200 for information processing according to some embodiments of the present disclosure. The process 200 may, for example, be implemented at the electronic device 110 as shown in FIG. 1.

[0032] As shown in FIG. 2, at block 210, the electronic device 110 performs a first training process on a second diffusion model based on a first set of reference states associated with a pre-trained first diffusion model, the first set of reference states corresponding to denoising results of a first set of denoising stages among a plurality of denoising stages of the first diffusion model.

[0033] In some embodiments, the first diffusion model can also be referred to as a teacher diffusion model. The teacher diffusion model may, for example, include 1000 denoising steps. To implement knowledge distillation on the teacher diffusion model, the electronic device 110 can divide the 1000 denoising steps into multiple segments.

[0034] As an example, the electronic device 110 can perform a four-stage progressive knowledge distillation, where the first stage may, for example, correspond to eight segments, each of which can correspond to 125 denoising steps out of the 1000 denoising steps; the second stage may, for example, correspond to four segments, each of which can correspond to 250 denoising steps out of the 1000 denoising steps; the third stage may, for example, correspond to two segments, each of which can correspond to 500 denoising steps out of the 1000 denoising steps; and the fourth stage may, for example, correspond to a single segment, which can correspond to all 1000 denoising steps.

[0035] For ease of description, the third stage is taken as an example of the first training process in block 210 below.

[0036] In some embodiments, at each training stage, the electronic device 110 can maintain three models, namely the student model (i.e., the second diffusion model) to be trained, the pre-trained teacher model, and the moving average model. The moving average model is initialized the same as the student model, and its updating manner is the moving average of the student model. As an example, the sliding coefficient can be set to 0.999, and each training iteration of the moving average model is updated to 0.999*original parameters of the student model + 0.001*updated parameters of the student model.

[0037] In the stage alternation process, the teacher model remains unchanged at all times, and the student model and the moving average model are initialized and inherit the student model parameters of the previous stage.

[0038] FIG. 3A illustrates a training process corresponding to two segments. As shown in FIG. 3A, the electronic device 110 can divide the trajectory into two segments. The trajectory refers to the process of the pre-trained teacher model from a Gaussian noise (step 1000) to a high-quality, high-resolution image (step 0) that matches the user instruction after step-by-step denoising, i.e., including all intermediate states of steps 1000, 999, …, 2, 1.

[0039] In the first training process, the electronic device 110 can train the second diffusion model such that the second diffusion model outputs a result that approximates the first set of reference states of the first diffusion model at the corresponding denoising stage.

[0040] As an example, FIG. 3A shows that for the t1 time step in the first segment, the electronic device 110 can cause the student diffusion model to approximate the position of the teacher diffusion model at T / 2 at the t1 time step. For the t2 time step in the second segment, the electronic device 110 can cause the student diffusion model to approximate the real image (i.e., the position at time step 0) at the t1 time step.

[0041] In this way, embodiments of the present disclosure can alleviate the difficulty of model fitting and improve the robustness of training.

[0042] In some embodiments, the electronic device 110 can achieve the above approximation through a distillation process. Specifically, when the electronic device 110 wants the t1 time step (also referred to as the first time step) to approximate the n1 time step (also referred to as the third time step) of the teacher diffusion model, the electronic device 110 can use the teacher diffusion model to perform inference to obtain the point of the next interval time step (also referred to as the second time step), for example, t1-s.

[0043] Further, the electronic device 110 can use a moving average model to obtain the intermediate state at the n1 time step, so that the prediction of the student diffusion model to be trained at the t1 time step learns the prediction of the moving average model at the t1-s time step for the n1 time step. For example, as shown in FIG. 3A, the arrow at x t1 -s and the arrow at x t1 may point to the same position, i.e., x n1 .

[0044] In this way, embodiments of the present disclosure reduce the difficulty of model training.

[0045] As shown in FIG. 3A, the electronic device 110 can perform the same type of processing on the t2 time step, so that the prediction of the student diffusion model to be trained at the t2 time step learns the prediction of the moving average model at the t2-s time step for the n2 time step. For example, as shown in FIG. 3A, the arrow at x t2 -s and the arrow at x t2 may point to the same position, i.e., x n2 .

[0046] After each stage of trajectory distillation, the performance of the model is inevitably subject to some loss. To compensate for the loss of performance, the electronic device 110 can also introduce a human feedback learning. Specifically, after each stage of distillation, we additionally perform a human feedback learning.

[0047] In some embodiments, the electronic device 110 can determine the first evaluation information based on the first generation result of the second diffusion model. In some examples, the first generation result may, for example, include a first image.

[0048] Accordingly, the electronic device 110 may, for example, generate first evaluation information about the first image based on the structure information, the style information, and / or the quality information of the first image.

[0049] As an example, the electronic device 110 may, after respectively performing inference on the real image and the generated image using the instance segmentation model, let the prediction of the generated image approximate the prediction of the real image, so that the student diffusion model can generate a reasonable object contour as much as possible and optimize the structure of the image.

[0050] As another example, the electronic device 110 may, using a network that extracts style features, extract features of the real image and the generated image as much as possible to make them close to each other, so that the style of the generated image can be closer to that of the real image.

[0051] As yet another example, the electronic device 110 may, further train a reward model that evaluates the overall quality of a picture using some data, to generate a score of the image. Further, the electronic device 110 may, make the score of the generated image close to that of the real image.

[0052] In this way, the electronic device 110 may, after each stage of trajectory distillation, further adjust the parameters of the second diffusion model based on the first evaluation information.

[0053] With reference back to FIG. 2, at block 220, the electronic device 110 performs a second training process on the second diffusion model based on a second set of reference states associated with the first diffusion model, the second set of reference states corresponding to denoising results of a second set of denoising stages of the plurality of denoising stages of the first diffusion model, wherein the number of the second set of denoising stages is less than the first set of denoising stages.

[0054] As an example, the first training process may, for example, correspond to a two-segmented distillation process, and the second training process may, for example, correspond to a single-segmented distillation process.

[0055] As shown in FIG. 3B, unlike the two-segmented distillation process, in the second training process, the electronic device 110 may, make the student diffusion model approximate the real image at time step t (i.e., the position of time step 0).

[0056] Similar to the process described above, the electronic device 110 may, perform approximation processing on the time step t, so that the prediction of the student diffusion model to be trained at the time step t learns the prediction of the sliding average model at time step t-2s for the time step n. For example, as shown in FIG. 3B, the arrow at x t and the arrow at x t-2s may point to the same position, i.e., X n .

[0057] With reference back to FIG. 2, at block 230, the electronic device 110 obtains a target diffusion model based on at least the first training process and the second training process of the second diffusion model.

[0058] It should be appreciated that the electronic device 110 can perform a two-stage or more-stage trajectory distillation process. In some embodiments, the multi-stage trajectory distillation process may, for example, include at least the single-segment trajectory distillation process corresponding to FIG. 3B.

[0059] In some embodiments, after completing the multi-stage progressive trajectory distillation process, the electronic device 110 can further perform a score distillation process. Specifically, after performing the multiple rounds of training process on the second diffusion model based on the first diffusion model, the electronic device 110 can determine second evaluation information based on the second generated results (e.g., second images) of the second diffusion model after the multiple rounds of training and the reference generated results (e.g., real images).

[0060] Further, the electronic device 110 can adjust the parameters of the second diffusion model based on the second evaluation information to obtain a target diffusion model.

[0061] Specifically, in the score distillation process, the electronic device 110 can utilize the score model to determine score information of the generated images and the real images. The score may, for example, represent the direction that the model predicts the real image at a certain time step.

[0062] Based on the trajectory distillation process described above, the estimation of the direction (i.e., score) that the model predicts the real image (step 0) at any time step will be more accurate. In addition, because in the last training stage, the electronic device 110 will sample to any time step and approximate the position of time step 0 (the target will be more consistent in general), it is possible to narrow down the scores of these time step pairs (t1, t1-s) or (t2, t2-s).

[0063] Thus, the electronic device 110 can further adjust the parameters of the student diffusion model through the score to obtain a final target diffusion model.

[0064] In some embodiments, after the score distillation, the electronic device 110 can implement one-step inference of the student diffusion model by training a low-rank matrix (Low-Rank Adaptation). Since the parameter amount of the low-rank matrix is relatively small, this can greatly save the training overhead.

[0065] In some embodiments, the target diffusion model obtained above can perform a single-step inference process, for example, based on a single-step denoising process to obtain a generated result. In some examples, such a target diffusion model may, for example, also perform a multi-step inference process.

[0066] Based on the above-described process, embodiments of the present disclosure can improve the efficiency of model training through a progressive distillation process to obtain a target diffusion model with higher execution efficiency.

[0067] Example apparatus and device

[0068] Embodiments of the present disclosure also provide a corresponding apparatus for implementing the above-described method or process. FIG. 4 shows a schematic structural block diagram of an example apparatus 400 for information processing according to certain embodiments of the present disclosure. The apparatus 400 can be implemented as or included in an electronic device. Various modules / components in the apparatus 400 can be implemented by hardware, software, firmware, or any combination thereof.

[0069] As shown in FIG. 4, the apparatus 400 includes a first training module 410 configured to perform a first training process on a second diffusion model based on a first set of reference states associated with a pre-trained first diffusion model, the first set of reference states corresponding to denoising results of a first set of denoising stages in a plurality of denoising stages of the first diffusion model; a second training module 420 configured to perform a second training process on the second diffusion model based on a second set of reference states associated with the first diffusion model, the second set of reference states corresponding to denoising results of a second set of denoising stages in the plurality of denoising stages of the first diffusion model, wherein a number of the second set of denoising stages is less than the first set of denoising stages; and a model obtaining module 430 configured to obtain a target diffusion model based at least on the first training process and the second training process of the second diffusion model.

[0070] In some embodiments, in the first training process, the second diffusion model is trained such that the second diffusion model outputs results approximating the first set of reference states at corresponding denoising stages; and / or in the second training process, the second diffusion model is trained such that the second diffusion model outputs results approximating the second set of reference states at corresponding denoising stages.

[0071] In some embodiments, the first training module 410 is configured to, for a first time step of the second diffusion model, determine a second time step associated with the first time step using the first diffusion model; and perform the first training process on the second diffusion model based on a first prediction of a moving average model from the second time step to a third time step to be approximated, such that the second training model learns the first prediction of the moving average model at a second prediction of the first time step, wherein the moving average model is determined based on model parameters of the second diffusion model.

[0072] In some embodiments, the apparatus 400 further includes an adjusting module configured to: after the first training process, determine first evaluation information of the first generation result based on the first generation result of the second diffusion model; and adjust parameters of the second diffusion model based on the first evaluation information.

[0073] In some embodiments, the first generation result includes a first image, and the first evaluation information is determined based on at least one of: structure information of the first image; style information of the first image; and quality information of the first image.

[0074] In some embodiments, the model obtaining module 430 is further configured to: after performing the multiple rounds of training processes on the second diffusion model based on the first diffusion model, determine second evaluation information based on a second generation result of the second diffusion model after the multiple rounds of training and a reference generation result; and adjust parameters of the second diffusion model based on the second evaluation information to obtain a target diffusion model.

[0075] In some embodiments, the second evaluation information includes score information determined by processing the second generation result and the reference generation result using a gradient model.

[0076] In some embodiments, the target diffusion model is used to generate images based on text.

[0077] In some embodiments, the target diffusion model is configured to perform a single-step inference process.

[0078] FIG. 5 illustrates a block diagram of an electronic device 500 in which one or more embodiments of the disclosure can be implemented. It should be understood that the electronic device 500 illustrated in FIG. 5 is merely exemplary and should not be construed as limiting on the functionality and scope of the embodiments described herein. The electronic device 500 illustrated in FIG. 5 can be used for the electronic device 110 as illustrated in FIG. 1.

[0079] As shown in FIG. 5, the electronic device 500 is in the form of a general electronic device. The components of the electronic device 500 can include, but are not limited to, one or more processors or processing units 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 can be an actual or virtual processor and is capable of performing various processing according to programs stored in the memory 520. In a multi-processor system, multiple processing units perform computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 500.

[0080] The electronic device 500 typically includes a plurality of computer storage media. Such media can be any available media that is accessible by the electronic device 500 and includes both volatile and nonvolatile media, removable and non-removable media. The memory 520 can be volatile (such as register, cache, RAM), non-volatile (such as ROM, EEPROM, flash memory), or some combination of the two. The storage device 530 can be a removable or non-removable media, and can include machine-readable media, such as flash drives, magnetic disks, or any other media that can be used to store information and / or data and that can be accessed by the electronic device 500.

[0081] The electronic device 500 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 5, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a removable, non-volatile optical disk (e.g., a CD-ROM) can be provided. In such instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. The memory 520 can include a computer program product 525 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.

[0082] The communication unit 540 enables communication with other electronic devices over communication media. Additionally, the functionality of the components of the electronic device 500 can be implemented in a single computing cluster or a plurality of computer machines capable of communication over a communication connection. Thus, the electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes.

[0083] The input device 550 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 560 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through the communication unit 540, as needed, one or more devices that enable a user to interact with the electronic device 500, or any devices (e.g., a network card, a modem, etc.) that enable the electronic device 500 to communicate with one or more other electronic devices. Such communication can be carried out via an input / output (I / O) interface (not shown).

[0084] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.

[0085] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0086] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium. The instructions stored on the computer readable storage medium can be used to program a computer, a programmable data processing apparatus, and / or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0087] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0088] The computer program product of the present disclosure can have a signal including said computer program. This signal can be electronic, electromagnetic, optical, or any other suitable type of signal. Such a signal can be provided through a communication connection, such as electrical wiring, optical fiber, wireless interface, etc. Examples of computer program products include computer program implemented on a personal computer, server, or other networked device. A non-transitory computer readable medium, such as a floppy disk, CD-ROM, DVD-ROM, Blu-ray Disc, hard disk drive, or any other suitable non-transitory computer readable medium can store the computer program product.

[0089] Various implementations of the disclosure have been described in detail above. The foregoing description is exemplary and explanatory only, and is not intended to be exhaustive or to limit various implementations of the disclosure to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings without departing from the scope and spirit of the disclosure. It is intended that the scope of the disclosure be limited only by the claims and the equivalents thereof. The use of the terms "including," "containing," "comprising," "having," "in involving," "portions," "elements," "components," "steps," "phases," "processes," "operations," "steps," "stages," "procedures," "methods," "mechanisms," "devices," "systems," "apparatuses," "units," "means," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "apparatuses," "units," "devices," "systems," "

Claims

1. A method for information processing, comprising: performing a first training process on a second diffusion model based on a first set of reference states associated with a pre-trained first diffusion model, the first set of reference states corresponding to denoising results of a first set of denoising stages among a plurality of denoising stages of the first diffusion model; performing a second training process on the second diffusion model based on a second set of reference states associated with the first diffusion model, the second set of reference states corresponding to denoising results of a second set of denoising stages among the plurality of denoising stages of the first diffusion model, wherein the second set of denoising stages has a smaller number than the first set of denoising stages; as well as A target diffusion model is acquired based on at least the first training process and the second training process of the second diffusion model.

2. The method according to claim 1, wherein: In the first training process, the second diffusion model is trained so that a result output by the second diffusion model in a corresponding denoising phase approaches the first set of reference states; and / or In the second training process, the second diffusion model is trained so that a result output by the second diffusion model in a corresponding denoising phase approaches the second set of reference states.

3. The method according to claim 2, wherein performing a first training process on the second diffusion model comprises: For a first time step of the second diffusion model, determining a second time step associated with the first time step using the first diffusion model; as well as The first training process is performed on the second diffusion model based on the first prediction of the sliding average model from the second time step to the third time step to be approximated, so that the second training model learns the first prediction of the sliding average model at the second prediction of the first time step, wherein the sliding average model is determined based on model parameters of the second diffusion model.

4. The method according to claim 1, further comprising: After the first training process, determining first evaluation information of the first generation result based on the first generation result of the second diffusion model; as well as Based on the first evaluation information, parameters of the second diffusion model are adjusted.

5. The method according to claim 4, wherein the first generated result comprises a first image, and the first evaluation information is determined based on at least one of the following: structural information of the first image; style information of the first image; quality information of the first image.

6. The method according to claim 1, wherein acquiring a target diffusion model based on at least the first training process and the second training process of the second diffusion model comprises: After performing multiple rounds of training on the second diffusion model based on the first diffusion model, determining second evaluation information based on a second generation result of the second diffusion model after the multiple rounds of training and a reference generation result; as well as Parameters of the second diffusion model are adjusted based on the second evaluation information to obtain the target diffusion model.

7. The method according to claim 6, wherein the second evaluation information comprises: Score information determined by processing the second generation result and the reference generation result using a gradient model. The method of claim 1 , wherein the target diffusion model is used to generate an image based on text.

9. The method of claim 1, wherein the target diffusion model is configured to perform a single-step reasoning process.

10. An apparatus for information processing, comprising: a first training module configured to perform a first training process on a second diffusion model based on a first set of reference states associated with a pre-trained first diffusion model, the first set of reference states corresponding to denoising results of a first set of denoising stages among a plurality of denoising stages of the first diffusion model; a second training module configured to perform a second training process on the second diffusion model based on a second set of reference states associated with the first diffusion model, the second set of reference states corresponding to denoising results of a second set of denoising stages among the plurality of denoising stages of the first diffusion model, wherein the number of the second set of denoising stages is smaller than the number of the first set of denoising stages; as well as The model acquisition module is configured to acquire a target diffusion model based on at least the first training process and the second training process of the second diffusion model.

11. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 9 when executed by the at least one processing unit.

12. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method according to any one of claims 1 to 9.

13. A computer program product comprising computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Information processing method, device and equipment and computer readable storage medium

    CN117196039A

  • Image generation method and device, equipment and storage medium

    CN117707400A

  • Diffusion model training method and device, image generation method and device and medium

    CN117809135A

  • Video generation with latent diffusion probabilistic models

    US20240087179A1