Information processing method and device, equipment and storage medium

By progressively training the diffusion model and utilizing trajectory distillation and fractional distillation techniques, the problems of high resource consumption and long time in the diffusion model generation process are solved, and an efficient single-step inference process is achieved.

CN120823402APending Publication Date: 2025-10-21BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410437727.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-11
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing diffusion models require multiple steps of denoising during the generation process, resulting in high resource consumption and long time.

Method used

By performing a progressive training process on the second diffusion model based on the reference state of the pre-trained diffusion model, including trajectory distillation and score distillation, the denoising result of the target diffusion model is gradually approached to obtain an efficient target diffusion model.

Benefits of technology

The efficiency of model training is improved, a target diffusion model capable of performing single-step reasoning is obtained, and resource consumption and generation time are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823402A_ABST
    Figure CN120823402A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to an information processing method and device, equipment and a storage medium. The method proposed herein includes: performing a first training process on a second diffusion model based on a first set of reference states associated with a pre-trained first diffusion model, the first set of reference states corresponding to denoising results of a first set of denoising phases of a plurality of denoising phases of the first diffusion model; based on a second group of reference states associated with the first diffusion model, executing a second training process on the second diffusion model, the second group of reference states corresponding to denoising results of a second group of denoising stages in the plurality of denoising stages of the first diffusion model, and the number of the second group of denoising stages being smaller than that of the first group of denoising stages; and obtaining a target diffusion model at least based on the first training process and the second training process of the second diffusion model. Therefore, according to the embodiment of the invention, the efficiency of model training can be improved through a progressive distillation process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to information processing methods, apparatuses, devices, and computer-readable storage media. Background Art

[0002] With the development of computer technology, various models have been gradually applied to various aspects of our daily lives. For example, diffusion models can be used to generate images based on text. Generally speaking, diffusion models require a multi-step denoising process to perform the generation process, which consumes a lot of resources and is time-consuming. Summary of the Invention

[0003] In a first aspect of the present disclosure, a method for information processing is provided. The method includes: performing a first training process on a second diffusion model based on a first set of reference states associated with a pre-trained first diffusion model, the first set of reference states corresponding to denoising results of a first set of denoising stages among a plurality of denoising stages of the first diffusion model; performing a second training process on the second diffusion model based on a second set of reference states associated with the first diffusion model, the second set of reference states corresponding to denoising results of a second set of denoising stages among a plurality of denoising stages of the first diffusion model, wherein the second set of denoising stages has a smaller number of denoising stages than the first set; and obtaining a target diffusion model based on at least the first training process and the second training process for the second diffusion model.

[0004] In a second aspect of the present disclosure, a device for information processing is provided. The device includes: a first training module configured to perform a first training process on a second diffusion model based on a first set of reference states associated with a pre-trained first diffusion model, the first set of reference states corresponding to denoising results of a first set of denoising stages among multiple denoising stages of the first diffusion model; a second training module configured to perform a second training process on the second diffusion model based on a second set of reference states associated with the first diffusion model, the second set of reference states corresponding to denoising results of a second set of denoising stages among multiple denoising stages of the first diffusion model, wherein the number of denoising stages in the second set is smaller than that in the first set; and a model acquisition module configured to acquire a target diffusion model based on at least the first training process and the second training process of the second diffusion model.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.

[0007] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0009] Figure 1 An example environment is shown that can implement some embodiments of the present disclosure in real time;

[0010] Figure 2 A flowchart illustrating an information processing process according to some embodiments of the present disclosure is shown;

[0011] Figure 3A and Figure 3B A schematic diagram illustrating a training process according to some embodiments of the present disclosure is shown;

[0012] Figure 4 A schematic structural block diagram showing an example apparatus for information processing according to some embodiments of the present disclosure; and

[0013] Figure 5 A block diagram of an electronic device capable of implementing various embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0014] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0015] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment may be included under any section / subsection. Furthermore, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or in different sections / subsections.

[0016] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.

[0017] The embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. These aspects shall comply with the corresponding laws, regulations and relevant provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user is aware of and confirms them. Accordingly, when implementing the various embodiments of the present disclosure, the types, scope of use, and usage scenarios of the data or information that may be involved should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with the relevant laws and regulations. The specific notification and / or authorization method may vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this respect.

[0018] If this specification and the solutions in the examples involve the processing of personal information, such processing will be done only with a legitimate basis (such as with the consent of the subject of personal information or as necessary for the performance of a contract) and only within the prescribed or agreed scope. A user's refusal to process personal information other than that required for basic functions will not affect the user's use of basic functions.

[0019] As briefly mentioned above, diffusion models typically require multiple steps of denoising and adding noise during training. Consequently, they also require multiple steps of denoising during the generation phase, which consumes significant resources and takes a long time.

[0020] Embodiments of the present disclosure provide a scheme for information processing. According to the scheme, a first training process is performed on a second diffusion model based on a first set of reference states associated with a pre-trained first diffusion model, the first set of reference states corresponding to denoising results of a first set of denoising stages among multiple denoising stages of the first diffusion model; a second training process is performed on the second diffusion model based on a second set of reference states associated with the first diffusion model, the second set of reference states corresponding to denoising results of a second set of denoising stages among multiple denoising stages of the first diffusion model, wherein the second set of denoising stages has fewer stages than the first set; and a target diffusion model is acquired based on at least the first and second training processes of the second diffusion model.

[0021] Therefore, the embodiments of the present disclosure can improve the efficiency of model training through a progressive distillation process to obtain a target diffusion model with higher execution efficiency.

[0022] Various example implementations of this solution are described in detail below in conjunction with the accompanying drawings.

[0023] Sample Environment

[0024] Figure 1 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Figure 1 As shown, example environment 100 may include electronic device 110 .

[0025] In this example environment 100, the electronic device 110 may train a student diffusion model 130 based on the teacher diffusion model 120. In some embodiments, the student diffusion model 130 may have fewer inference steps than the teacher diffusion model 120.

[0026] As an example, the student diffusion model 130 may be capable of performing single-step reasoning by, for example, distilling knowledge from the teacher diffusion model 120 .

[0027] The specific process of training the student diffusion model 130 will be described in detail below.

[0028] Example Process

[0029] The following will refer to Figure 2 To describe the information processing process according to some embodiments of the present disclosure. Figure 2 FIG. 2 is a flow chart illustrating an example process 200 for information processing according to some embodiments of the present disclosure. The process 200 may be implemented, for example, in Figure 1 At the electronic device 110 shown.

[0030] like Figure 2 As shown, in box 210, the electronic device 110 performs a first training process on the second diffusion model based on a first set of reference states associated with a pre-trained first diffusion model, and the first set of reference states corresponds to the denoising results of the first set of denoising stages in multiple denoising stages of the first diffusion model.

[0031] In some embodiments, the first diffusion model may also be referred to as a teacher diffusion model. The teacher diffusion model may, for example, include 1000 denoising steps. To implement knowledge distillation of the teacher diffusion model, the electronic device 110 may divide the 1000 denoising steps into multiple segments.

[0032] As an example, the electronic device 110 may perform four stages of progressive knowledge distillation, wherein the first stage may, for example, correspond to eight segments, each of which may correspond to 125 denoising steps out of 1000 denoising steps; the second stage may, for example, correspond to four segments, each of which may correspond to 250 denoising steps out of 1000 denoising steps; the third stage may, for example, correspond to two segments, each of which may correspond to 500 denoising steps out of 1000 denoising steps; and the fourth stage may, for example, correspond to a single segment, which may correspond to all 1000 denoising steps.

[0033] For the convenience of description, the third stage is taken as an example of the first training process in block 210 .

[0034] In some embodiments, in each training phase, the electronic device 110 can maintain three models, namely, a student model to be trained (i.e., a second diffusion model), a pre-trained teacher model, and a sliding average model. The sliding average model is initialized the same as the student model, and its update method is the sliding average of the student model. As an example, the sliding coefficient can be set to 0.999, and each training iteration of the sliding average model is updated to 0.999*the original parameters of the student model+0.001*the updated parameters of the student model.

[0035] During the stage alternation process, the teacher model remains unchanged, and the student model and sliding average model are initialized and inherit the student model parameters of the previous stage.

[0036] Figure 3A The training process corresponding to two segments is shown. Figure 3A As shown, the electronic device 110 can divide the trajectory into two segments. The trajectory refers to the process in which the pre-trained teacher model gradually denoises a Gaussian noise (step 1000) and obtains a high-quality, high-resolution image that matches the user instruction (step 0), that is, all intermediate states including steps 1000, 999, ..., step 2, and step 1.

[0037] In the first training process, the electronic device 110 may train the second diffusion model so that the result output by the second diffusion model in the corresponding denoising phase approaches the first set of reference states of the first diffusion model.

[0038] by Figure 3A For example, for the first segment at time step t1, the electronic device 110 may cause the student diffusion model to approach the position of the teacher diffusion model at time step t1 / 2. For the second segment at time step t2, the electronic device 110 may cause the student diffusion model to approach the true image at time step t1 (i.e., the position of time step 0).

[0039] In this way, the embodiments of the present disclosure can alleviate the difficulty of model fitting and improve the robustness of training.

[0040] In some embodiments, the electronic device 110 can achieve the above approximation through a distillation process. Specifically, when the electronic device 110 wishes to approximate the t1 time step (also referred to as the first time step) to the n1 time step (also referred to as the third time step) of the teacher diffusion model, the electronic device 110 can use the teacher diffusion model to perform inference to obtain the point at the next interval time step (also referred to as the second time step), for example, t1-s.

[0041] Furthermore, the electronic device 110 can use the sliding average model to obtain the intermediate state at the same time step n1, so that the prediction of the student diffusion model to be trained at the time step t1 can learn the prediction of the sliding average model at the time step t1-s for the time step n1. Figure 3A As shown, x t1 Arrow and x t1-s The arrows at can point to the same position, that is, x n1 .

[0042] In this way, the embodiment of the air-open cover reduces the fitting difficulty of model training.

[0043] like Figure 3A As shown, the electronic device 110 can perform a type of processing on the t2 time step so that the prediction of the student diffusion model to be trained at the t2 time step can learn the prediction of the sliding average model t2-s time step for the n2 time step. Figure 3A As shown, x t2 Arrow and x t2-s The arrows at can point to the same position, that is, x n2 .

[0044] After each stage of trajectory distillation, the model's performance will inevitably suffer some degradation. To compensate for this performance loss, the electronic device 110 can also introduce a round of human feedback learning. Specifically, after each stage of distillation, we will perform an additional round of human feedback learning.

[0045] In some embodiments, the electronic device 110 may determine the first evaluation information based on the first generated result of the second diffusion model. In some examples, the first generated result may include, for example, a first image.

[0046] Accordingly, the electronic device 110 may generate first evaluation information about the first image based on the structural information, style information and / or quality information of the first image, for example.

[0047] As an example, the electronic device 110 can use the instance segmentation model to infer the real image and the generated image respectively, and let the prediction of the generated image approach the prediction of the real image, so that the student diffusion model can generate reasonable object contours as much as possible and optimize the image structure.

[0048] As another example, the electronic device 110 may use a style feature extraction network to extract features of the real image and the generated image to make them as close as possible, so that the style of the generated image can be closer to the real image.

[0049] As another example, the electronic device 110 may also use some data to train a reward model for evaluating the overall quality of the image to generate a score for the image. Furthermore, the electronic device 110 may make the score of the generated image close to the score of the real image.

[0050] Therefore, the electronic device 110 may further adjust the parameters of the second diffusion model based on the first evaluation information after the trajectory distillation at each stage.

[0051] Continue to refer Figure 2 In block 220, the electronic device 110 performs a second training process on the second diffusion model based on a second set of reference states associated with the first diffusion model, where the second set of reference states corresponds to denoising results of a second set of denoising stages among the plurality of denoising stages of the first diffusion model, wherein the number of the second set of denoising stages is less than that of the first set of denoising stages.

[0052] As an example, the first training process may correspond to a two-segment distillation process, and the second training process may correspond to a single-segment distillation process.

[0053] like Figure 3B As shown, unlike the distillation process based on two segments, in the second training process, the electronic device 110 can make the student diffusion model approach the real image at time step t (ie, the position of time step 0).

[0054] Similar to the process described above, the electronic device 110 can perform approximation processing on the t time step so that the prediction of the student diffusion model to be trained at the t time step can learn the prediction of the sliding average model at the t-2s time step for the n time step. Figure 3B As shown, x t Arrow and x t-2s The arrows at can point to the same position, that is, x n .

[0055] Continue to refer Figure 2 In block 230 , the electronic device 110 obtains a target diffusion model based on at least the first training process and the second training process of the second diffusion model.

[0056] It should be understood that the electronic device 110 may perform a trajectory distillation process in two or more stages. In some embodiments, the trajectory distillation process in multiple stages may include at least Figure 3B The corresponding single-segment trajectory distillation process.

[0057] In some embodiments, after completing multiple stages of the progressive trajectory distillation process, the electronic device 110 may further perform a score distillation process. Specifically, after performing multiple rounds of training on the second diffusion model based on the first diffusion model, the electronic device 110 may determine second evaluation information based on a second generation result (e.g., a second image) of the second diffusion model after the multiple rounds of training and a reference generation result (e.g., a real image).

[0058] Furthermore, the electronic device 110 may adjust parameters of the second diffusion model based on the second evaluation information to obtain a target diffusion model.

[0059] Specifically, during the score distillation process, the electronic device 110 may use the score model to determine score information of the generated image and the real image. The score may, for example, represent the direction of the model's prediction of the real image at a certain time step.

[0060] Based on the trajectory distillation process described above, the model's estimate of the predicted direction (i.e., score) for the true image (step 0) at any time step is more accurate. Furthermore, because the electronic device 110 samples at any time step in the final training phase and approaches the position at time step 0 (the target is relatively consistent), it is possible to bring the scores of time step pairs (t1, t1-s) or (t2, t2-s) closer together.

[0061] Therefore, the electronic device 110 can further adjust the parameters of the student diffusion model according to the scores to obtain the final target diffusion model.

[0062] In some embodiments, after score distillation, the electronic device 110 can implement one-step reasoning of the student diffusion model by training a low-rank matrix (Low-Rank Adaptation). Since the number of parameters of the low-rank matrix is ​​relatively small, this can greatly save training overhead.

[0063] In some embodiments, the target diffusion model obtained above can perform a single-step reasoning process, for example, a generation result can be obtained based on a single-step denoising process. In some examples, such a target diffusion model can also perform a multi-step reasoning process.

[0064] Based on the process described above, the embodiments of the present disclosure can improve the efficiency of model training through a progressive distillation process to obtain a target diffusion model with higher execution efficiency.

[0065] Example devices and equipment

[0066] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 4 1 shows a schematic block diagram of an example apparatus 400 for information processing according to certain embodiments of the present disclosure. Apparatus 400 may be implemented as or included in an electronic device. Each module / component in apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.

[0067] like Figure 4 As shown, the device 400 includes a first training module 410, which is configured to perform a first training process on the second diffusion model based on a first set of reference states associated with a pre-trained first diffusion model, the first set of reference states corresponding to the denoising results of a first set of denoising stages in multiple denoising stages of the first diffusion model; a second training module 420, which is configured to perform a second training process on the second diffusion model based on a second set of reference states associated with the first diffusion model, the second set of reference states corresponding to the denoising results of a second set of denoising stages in multiple denoising stages of the first diffusion model, wherein the number of the second set of denoising stages is less than that of the first set of denoising stages; and a model acquisition module 430, which is configured to acquire a target diffusion model based on at least the first training process and the second training process of the second diffusion model.

[0068] In some embodiments, during the first training process, the second diffusion model is trained so that the results output by the second diffusion model in the corresponding denoising stage approach the first set of reference states; and / or during the second training process, the second diffusion model is trained so that the results output by the second diffusion model in the corresponding denoising stage approach the second set of reference states.

[0069] In some embodiments, the first training module 410 is configured to: for a first time step of the second diffusion model, determine a second time step associated with the first time step using the first diffusion model; and perform a first training process on the second diffusion model based on the first prediction of the sliding average model from the second time step to the third time step to be approximated, so that the second training model learns the first prediction of the sliding average model based on the second prediction of the first time step, wherein the sliding average model is determined based on the model parameters of the second diffusion model.

[0070] In some embodiments, the apparatus 400 further includes an adjustment module configured to: determine first evaluation information of the first generation result based on the first generation result of the second diffusion model after the first training process; and adjust parameters of the second diffusion model based on the first evaluation information.

[0071] In some embodiments, the first generated result includes a first image, and the first evaluation information is determined based on at least one of the following: structural information of the first image; style information of the first image; and quality information of the first image.

[0072] In some embodiments, the model acquisition module 430 is further configured to: after performing multiple rounds of training on the second diffusion model based on the first diffusion model, determine second evaluation information based on the second generation results and reference generation results of the second diffusion model after multiple rounds of training; and adjust the parameters of the second diffusion model based on the second evaluation information to obtain the target diffusion model.

[0073] In some embodiments, the second evaluation information includes score information determined by processing the second generation result and the reference generation result using a gradient model.

[0074] In some embodiments, a target diffusion model is used to generate images based on text.

[0075] In some embodiments, the target diffusion model is configured to perform a single-step inference process.

[0076] Figure 5 1 shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. Figure 5 The illustrated electronic device 500 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 5 The electronic device 500 shown can be used for Figure 1 An electronic device 110 is shown.

[0077] like Figure 5 As shown, electronic device 500 is in the form of a general electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of electronic device 500.

[0078] The electronic device 500 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 500.

[0079] The electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 5 As shown in FIG, a magnetic disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. Memory 520 may include a computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0080] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 500 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 500 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0081] Input device 550 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 560 may be one or more output devices, such as a display, a speaker, or a printer. Electronic device 500 may also communicate with one or more external devices (not shown) via communication unit 540 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with electronic device 500, or with any device that allows electronic device 500 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0082] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0083] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0084] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0085] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0086] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0087] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for information processing, comprising: performing a first training process on a second diffusion model based on a first set of reference states associated with a pre-trained first diffusion model, the first set of reference states corresponding to denoising results of a first set of denoising stages among a plurality of denoising stages of the first diffusion model; performing a second training process on the second diffusion model based on a second set of reference states associated with the first diffusion model, the second set of reference states corresponding to denoising results of a second set of denoising stages among the plurality of denoising stages of the first diffusion model, wherein the second set of denoising stages has a smaller number than the first set of denoising stages; as well as A target diffusion model is acquired based on at least the first training process and the second training process of the second diffusion model.

2. The method according to claim 1, wherein: In the first training process, the second diffusion model is trained so that a result output by the second diffusion model in a corresponding denoising phase approaches the first set of reference states; and / or In the second training process, the second diffusion model is trained so that a result output by the second diffusion model in a corresponding denoising phase approaches the second set of reference states.

3. The method according to claim 2, wherein performing a first training process on the second diffusion model comprises: For a first time step of the second diffusion model, determining a second time step associated with the first time step using the first diffusion model; as well as The first training process is performed on the second diffusion model based on the first prediction of the sliding average model from the second time step to the third time step to be approximated, so that the second training model learns the first prediction of the sliding average model at the second prediction of the first time step, wherein the sliding average model is determined based on model parameters of the second diffusion model.

4. The method according to claim 1, further comprising: After the first training process, determining first evaluation information of the first generation result based on the first generation result of the second diffusion model; as well as Based on the first evaluation information, parameters of the second diffusion model are adjusted.

5. The method according to claim 4, wherein the first generated result comprises a first image, and the first evaluation information is determined based on at least one of the following: structural information of the first image; style information of the first image; quality information of the first image.

6. The method according to claim 1, wherein acquiring a target diffusion model based on at least the first training process and the second training process of the second diffusion model comprises: After performing multiple rounds of training on the second diffusion model based on the first diffusion model, determining second evaluation information based on a second generation result of the second diffusion model after the multiple rounds of training and a reference generation result; as well as Parameters of the second diffusion model are adjusted based on the second evaluation information to obtain the target diffusion model.

7. The method according to claim 6, wherein the second evaluation information comprises: Score information determined by processing the second generation result and the reference generation result using a gradient model. The method of claim 1 , wherein the target diffusion model is used to generate an image based on text.

9. The method of claim 1, wherein the target diffusion model is configured to perform a single-step reasoning process.

10. An apparatus for information processing, comprising: a first training module configured to perform a first training process on a second diffusion model based on a first set of reference states associated with a pre-trained first diffusion model, the first set of reference states corresponding to denoising results of a first set of denoising stages among a plurality of denoising stages of the first diffusion model; a second training module configured to perform a second training process on the second diffusion model based on a second set of reference states associated with the first diffusion model, the second set of reference states corresponding to denoising results of a second set of denoising stages among the plurality of denoising stages of the first diffusion model, wherein the number of the second set of denoising stages is smaller than the number of the first set of denoising stages; as well as The model acquisition module is configured to acquire a target diffusion model based on at least the first training process and the second training process of the second diffusion model.

11. An electronic device comprising: at least one processing unit; as well as At least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 9 when executed by the at least one processing unit.

12. A computer-readable storage medium having a computer program stored thereon, wherein the computer program can be executed by a processor to implement the method according to any one of claims 1 to 9.