Method, apparatus, device and storage medium for generating media content
Patent Information
- Application Number
- CN202510353440.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2026-09-25
AI Technical Summary
但是由于扩散模型在理论上能够逼近复杂的数据分布,且生成的内容往往具有更高的质量和一致性,因此扩散模型也需要较高的计算成本,导致内容生成效率较低
[0008]应当理解,本内容部分中所描述的内容并非旨在限定本公开的实施例的关键特征或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的描述而变得容易理解。
Smart Images

Figure CN122824918A_ABST
Abstract
Description
Technical Field
[0001] The exemplary embodiments disclosed herein relate generally to the field of computers, and more particularly to methods, apparatus, devices, and computer-readable storage media for generating media content. Background Technology
[0002] The diffusion model is a content generation model that generates new samples by simulating the process of data gradually transforming from noise into a target distribution. Its core idea is to gradually add noise to transform the data distribution into a simple distribution, and then recover the data from the noise through a reverse process. However, because the diffusion model can theoretically approximate complex data distributions and the generated content often has higher quality and consistency, it also requires high computational costs, resulting in low content generation efficiency. Summary of the Invention
[0003] In a first aspect of this disclosure, a method for generating media content is provided. The method includes: receiving a media generation request associated with a diffusion model; and generating media content for the media generation request based on a denoising process of the diffusion model, the denoising process including at least a first stage, a second stage, and a third stage; the first stage is configured to perform at least one round of denoising on first noise data to determine a first feature representation; the second stage is configured to perform at least one round of denoising on second noise data to determine a second feature representation, the second noise data being constructed based on downsampling the first feature representation; and the third stage is configured to perform at least one round of denoising on third noise data, the third noise data being constructed based on upsampling the second feature representation.
[0004] In a second aspect of this disclosure, an apparatus for generating media content is provided. The apparatus includes: a receiving module configured to receive a media generation request associated with a diffusion model; and a generation module configured to generate media content for the media generation request based on a denoising process of the diffusion model, the denoising process including at least a first stage, a second stage, and a third stage; the first stage being configured to perform at least one round of denoising on first noise data to determine a first feature representation; the second stage being configured to perform at least one round of denoising on second noise data to determine a second feature representation, the second noise data being constructed based on downsampling the first feature representation; and the third stage being configured to perform at least one round of denoising on third noise data, the third noise data being constructed based on upsampling the second feature representation.
[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes computer-executable instructions that, when executed by a processor, implement the method according to a first aspect of this disclosure.
[0008] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0010] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;
[0011] Figure 2 A flowchart illustrating a process for generating media content according to some embodiments of the present disclosure is shown;
[0012] Figure 3 A flowchart of a noise reduction process according to some embodiments of the present disclosure is shown;
[0013] Figure 4 A schematic structural block diagram of an apparatus for generating media content according to certain embodiments of the present disclosure is shown;
[0014] Figure 5 A block diagram of an electronic device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation
[0015] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0016] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.
[0017] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0018] The embodiments of this disclosure may involve user data, data acquisition, and / or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
[0019] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
[0020] Traditionally, the diffusion model is a content generation model that generates new samples by simulating the process of data gradually transforming from noise into a target distribution. Its core idea is to gradually add noise to transform the data distribution into a simple distribution, and then recover the data from the noise through a reverse process. However, the efficiency of current media content generation is relatively low.
[0021] Embodiments of this disclosure propose a scheme for generating media content. According to the scheme, a media generation request associated with a diffusion model is received; and a denoising process based on the diffusion model is used to generate media content for the media generation request. The denoising process includes at least a first stage, a second stage, and a third stage. The first stage is configured to perform at least one round of denoising on first noise data to determine a first feature representation; the second stage is configured to perform at least one round of denoising on second noise data to determine a second feature representation, the second noise data being constructed based on downsampling the first feature representation; and the third stage is configured to perform at least one round of denoising on third noise data, the third noise data being constructed based on upsampling the second feature representation. Since downsampling can reduce data dimensionality and computational load, denoising the downsampled second noise data can improve the efficiency of media content generation.
[0022] Example Environment
[0023] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, example environment 100 may include electronic device 110.
[0024] In some embodiments, electronic device 110 may receive a media generation request associated with diffusion model 120; and generate media content for the media generation request based on a denoising process of diffusion model 120, the denoising process including at least a first stage, a second stage, and a third stage; the first stage is configured to perform at least one round of denoising on first noise data to determine a first feature representation; the second stage is configured to perform at least one round of denoising on second noise data to determine a second feature representation, the second noise data being constructed based on downsampling the first feature representation; the third stage is configured to perform at least one round of denoising on third noise data, the third noise data being constructed based on upsampling the second feature representation. Diffusion model 120 may be deployed on electronic device 110 or on other devices, which will not be elaborated here.
[0025] In some embodiments, the electronic device 110 may be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device 110 may also support any type of interface for the target user (such as "wearable" circuitry).
[0026] Electronic device 110 can also be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Electronic device 110 may include, for example, computing systems / servers, such as mainframes, edge computing nodes, computing devices in cloud environments, and so on.
[0027] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0028] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.
[0029] Example process
[0030] Figure 2 A flowchart of a process 200 for generating media content according to some embodiments of the present disclosure is shown. Process 200 can be implemented at electronic device 110. Reference is made below. Figure 1 Describe the process 200.
[0031] In box 210, electronic device 110 receives a media generation request associated with a diffusion model.
[0032] In some embodiments, a media generation request is a request provided to the network model to guide its content generation process. A media generation request may include a request to generate image content or a request to generate video content. For example, a media generation request could be "generate a cartoon-style image." Alternatively, a media generation request may indicate reference information, such as a reference image.
[0033] In box 220, electronic device 110 generates media content for a media generation request based on a diffusion model-based noise reduction process.
[0034] In some embodiments, the diffusion model is a content generation model that generates new samples by simulating the process of data gradually transforming from noise into a target distribution. Its core idea is to gradually add noise to transform the data distribution into a simple distribution, and then recover the data from the noise through a reverse process. In embodiments of this disclosure, reference is made to... Figure 3 The diffusion model can perform progressive noise reduction operations based on media generation requests, ultimately generating media content 308.
[0035] In some embodiments, reference Figure 3 The noise reduction process may include multiple stages, each corresponding to a specific resolution. The electronic device 110 may configure the noise reduction process to include at least a first stage, a second stage, and a third stage. The first stage is configured to perform at least one round of noise reduction on the first noise data 301 to determine a first feature representation 302; the second stage is configured to perform at least one round of noise reduction on the second noise data 304 to determine a second feature representation 305, the second noise data 304 being constructed based on downsampling the first feature representation 302; the third stage is configured to perform at least one round of noise reduction on the third noise data 307, the third noise data 307 being constructed based on upsampling the second feature representation 305.
[0036] Because a downsampling operation was performed on the first feature representation 302 in the first stage, the resolution processed in the second stage is lower than that in the first stage. Conversely, because an upsampling operation was performed on the second feature representation 305 in the second stage, the resolution processed in the third stage is higher than that in the second stage. For example, the first stage corresponds to a resolution of 1024, the second stage to a resolution of 512, and the third stage to a resolution of 1024. In this way, the diffusion model progressively reduces noise across multiple stages, achieving efficient inference while maintaining quality.
[0037] In some embodiments, the electronic device 110 may determine a preset velocity field corresponding to the diffusion model, where the preset velocity field indicates the feature of the next denoising time step generated after denoising processing is performed on the feature of the current denoising time step. Then, denoising processing is performed on the first noise data 301 based on the preset velocity field to generate media content in response to the media generation request.
[0038] For example, the size of the media content to be generated is h×w, where h represents height and w represents width. Then the first noise data 301 can be expressed as x0∈R b×c×h×w , where b represents the batch size and c represents the number of channels. The denoising process includes K stages, and in these stages, h1>h2>...>h K-1 <h K and w1>w2>...>w K-1 <w K are satisfied. Then if the denoising process includes 3 stages, h1>h2<h3 and w1>w3<w2 are satisfied.
[0039] In addition, the number of inference steps included in each stage can be defined as N1,N2,...,N K , then the total number of inferences is In each stage i, discrete denoising time steps can be constructed The denoising process can be as follows:
[0040]
[0041] u θ represents the preset velocity field, y represents the feature representation corresponding to the media generation request, and j∈{0,1,…,N i}.
[0042] In some embodiments, with reference to Figure 3 , the electronic device 110 may configure the first stage to perform at least one round of denoising on the first noise data 301 to determine the first feature representation 302. Since semantic information and specific details are difficult to obtain at low resolution, and the first stage performs inference at high resolution, inference at high resolution can facilitate the network model to capture semantic information of the input media content.
[0043] In some embodiments, with reference to Figure 3 the electronic device 110 configures the second stage to perform at least one round of denoising on the second noise data 304 to determine the second feature representation 305. Since the second stage targets feature representations with lower resolution, and performing inference at low resolution during the denoising process can reduce computational overhead, this method can improve the denoising efficiency of the diffusion model.
[0044] In some embodiments, by introducing preset noise at the resolution transition, a completely new noise reduction process from high noise to low noise can be performed at the new resolution, ensuring a more coherent and effective diffusion process. Therefore, the electronic device 110 can determine a third feature representation 303 by downsampling a first feature representation 302, and then add first random noise to the third feature representation 303 to construct second noise data 304. Downsampling is a media content generation technique used to reduce the sampling rate or resolution of data. Through downsampling operations, the main features of the data can be preserved while reducing the amount of data and computational complexity.
[0045] For example, the formula for obtaining the third feature representation 303 by downsampling the first feature representation 302 can be as follows:
[0046]
[0047] i represents a stage. The final potential representation of the (i-1)th stage is downsampled to match the resolution used in the ith stage. After adjusting the resolution, first random noise can be added to the third feature representation 303 to construct the second noisy data 304.
[0048] In some embodiments, the second stage corresponds to a first number of inference steps, and the first number is determined based on a first weight of a first random noise. Furthermore, the first number is positively correlated with the first weight. For example, it can be applied... The weight of the noise is then determined by the target time step τ for noise injection. i The formula corresponding to the noise addition process can be as follows:
[0049]
[0050] That is, noise is added to the latent representation in the previous stage, and only the first number N is needed in the later stages. i w i One reasoning step, not the complete w i One reasoning step.
[0051] This paper proposes a novel method for denoising from high noise to low noise at a new resolution. This method ensures that the training distribution learned by the network model during training aligns with the inference distribution encountered during inference, avoiding mismatches caused by directly transmitting information or features between different resolution stages. Furthermore, it allows the application of multi-resolution priors from the network model.
[0052] In some embodiments, adjusting the resolution not only alters the spatial characteristics of the data but also directly affects the signal-to-noise ratio (SNR) of the latent variable region, thereby weakening the effective signal preserved in the previous step. Furthermore, noise is reintroduced during resolution adjustment, essentially reverting each latent representation to a low SNR state. To mitigate this impact, an additional scheduler offset can be introduced during resolution adjustment, resulting in a more stable noise reduction effect.
[0053] Therefore, the electronic device 110 can construct a first sampling scheduler corresponding to the second stage based on a first offset coefficient. Then, based on the first sampling scheduler, it determines multiple denoising time steps corresponding to a first number of inference steps. Finally, based on these multiple denoising time steps, it performs at least one round of denoising on the second noisy data 304. For example, the offset coefficient can be expressed as... The adjustment process for the noise reduction time step can be represented by the following formula:
[0054]
[0055] t i,n t represents the initial noise reduction time step. i,m This indicates the adjusted denoising time step. This method ensures that the denoising process remains consistent with the varying signal-to-noise ratio at different resolutions, thereby improving the performance of the diffusion model in both image and video generation.
[0056] In some embodiments, the electronic device 110 can determine a fourth feature representation 306 by upsampling a second feature representation 305, and then add second random noise to the fourth feature representation 306 to construct third noise data 307. Upsampling is a media content generation technique used to increase the sampling rate or resolution of data. Upsampling expands data by inserting new data points between the original data, which can maintain a good visual effect when the image is enlarged and avoid obvious pixelation or distortion.
[0057] For example, the formula for obtaining the fourth feature representation 306 by upsampling the second feature representation 305 can be as follows:
[0058]
[0059] i represents a stage. The final potential representation of the (i-1)th stage is upsampled to match the resolution used in the ith stage. After adjusting the resolution, a second random noise can be added to the fourth feature representation 306 to construct the third noisy data 307.
[0060] In some embodiments, the third stage corresponds to a second number of inference steps, and the second number is determined based on a second weight of a second random noise. Furthermore, the second number is positively correlated with the second weight. For example, it can be applied... The weight of the noise is then determined by the target time step τ for noise injection. i The formula corresponding to the noise addition process can be as follows:
[0061]
[0062] That is, adding noise to the latent representation in the previous stage, so that only the second number N is needed in the later stages. i w i One reasoning step, not the complete w i One reasoning step.
[0063] In some embodiments, the electronic device 110 may construct a second sampling scheduler corresponding to the third stage based on a second offset coefficient corresponding to the third stage. Then, based on the second sampling scheduler, multiple denoising time steps corresponding to a second number of inference steps are determined. Finally, based on the multiple denoising time steps, at least one round of denoising is performed on the third noisy data 307. For example, the offset coefficient can be expressed as... The adjustment process for the noise reduction time step can be represented by the following formula:
[0064]
[0065] t i,n t represents the initial noise reduction time step. i,m This indicates the adjusted denoising time step. This method ensures that the denoising process remains consistent with the varying signal-to-noise ratio at different resolutions, thereby improving the performance of the diffusion model in both image and video generation.
[0066] In some embodiments, the noise reduction process may further include a fourth stage and a fifth stage. The fourth stage is configured to perform at least one round of noise reduction on the fourth noise data to determine a fifth feature representation, whereby the first noise data 301 is constructed based on the downsampled fifth feature representation. The fifth stage is configured to perform at least one round of noise reduction on the fifth noise data, which is constructed based on the upsampled sixth feature representation generated by the third stage. For example, the fourth stage corresponds to a resolution of 1024, the first stage to a resolution of 512, the second stage to a resolution of 256, the third stage to a resolution of 512, and the fifth stage to a resolution of 1024. In practical applications, the number of stages can be added or reduced according to actual needs, which will not be elaborated here.
[0067] By dividing the diffusion model into multiple stages and following a high-resolution-low-resolution-high-resolution denoising process, high-resolution denoising is performed at the beginning and final stages, while low-resolution denoising is performed in the intermediate stages. This allows for reduced computational overhead through low-resolution priors, improving content generation efficiency while maintaining the fidelity of the network model output. Furthermore, to mitigate issues such as aliasing and blur artifacts in the network model output, the resolution transition points are further optimized, and the denoising steps are adaptively adjusted at each stage.
[0068] Example devices and equipment
[0069] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 4 A schematic structural block diagram of an apparatus 400 for generating media content according to certain embodiments of the present disclosure is shown. The apparatus 400 may be implemented as or included in the electronic device 110 discussed above. The various modules / components in the apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.
[0070] like Figure 4 As shown, the apparatus 400 includes a receiving module 410 configured to receive a media generation request associated with a diffusion model; and a generation module 420 configured to generate media content for the media generation request based on a denoising process of the diffusion model, the denoising process including at least a first stage, a second stage, and a third stage; the first stage is configured to perform at least one round of denoising on first noise data to determine a first feature representation; the second stage is configured to perform at least one round of denoising on second noise data to determine a second feature representation, the second noise data being constructed based on downsampling the first feature representation; and the third stage is configured to perform at least one round of denoising on third noise data, the third noise data being constructed based on upsampling the second feature representation.
[0071] In some embodiments, the generation module 420 includes a downsampling submodule configured to determine a third feature representation by downsampling a first feature representation; and a first noise addition submodule configured to add first random noise to the third feature representation to construct second noise data.
[0072] In some embodiments, the second stage corresponds to a first number of inference steps, and the first number is determined based on a first weight of a first random noise.
[0073] In some embodiments, the first number is positively correlated with the first weight.
[0074] In some embodiments, the generation module 420 is further configured to construct a first sampling scheduler corresponding to the second stage based on a first offset coefficient corresponding to the second stage; determine a plurality of noise reduction time steps corresponding to a first number of inference steps based on the first sampling scheduler; and perform at least one round of noise reduction on the second noise data based on the plurality of noise reduction time steps.
[0075] In some embodiments, the generation module 420 further includes an upsampling submodule configured to determine a fourth feature representation by upsampling the second feature representation; and a second noise addition submodule configured to add second random noise to the fourth feature representation to construct third noise data.
[0076] In some embodiments, the third stage corresponds to a second number of inference steps, and the second number is determined based on a second weight of the second random noise.
[0077] In some embodiments, the second number is positively correlated with the second weight.
[0078] In some embodiments, the generation module 420 is further configured to construct a second sampling scheduler corresponding to the third stage based on a second offset coefficient corresponding to the third stage; determine multiple denoising time steps corresponding to a second number of inference steps based on the second sampling scheduler; and perform at least one round of denoising on the third noise data based on the multiple denoising time steps.
[0079] In some embodiments, the generation module 420 is further configured to perform at least one round of noise reduction on the fourth noise data to determine a fifth feature representation, the fourth noise data being constructed based on a downsampled first feature representation; and the fifth stage is configured to perform at least one round of noise reduction on the fifth noise data to determine a sixth feature representation, the fifth noise data being constructed based on an upsampled second feature representation.
[0080] The units included in device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units may be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 400 may be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that may be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.
[0081] Figure 5A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 5 The electronic device 500 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 5 The electronic device 500 shown can be used to achieve Figure 1 The electronic device 110 shown.
[0082] like Figure 5 As shown, electronic device 500 is in the form of a general-purpose electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 500.
[0083] Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 530 can be a removable or non-removable medium and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within electronic device 500.
[0084] Electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 5 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.
[0085] Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0086] Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0087] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0088] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0089] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0090] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0091] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0092] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for generating media content, comprising: Receive media generation requests associated with the diffusion model; as well as Based on the denoising process of the diffusion model, media content for the media generation request is generated, and the denoising process includes at least a first stage, a second stage, and a third stage. The first stage is configured to perform at least one round of noise reduction on the first noisy data to determine a first feature representation; The second stage is configured to perform at least one round of noise reduction on the second noise data to determine a second feature representation, the second noise data being constructed based on downsampling the first feature representation; The third stage is configured to perform at least one round of noise reduction on the third noise data, which is constructed based on upsampling of the second feature representation.
2. The method according to claim 1, further comprising: The third feature representation is determined by downsampling the first feature representation; as well as A first random noise is added to the third feature representation to construct the second noise data.
3. The method of claim 2, wherein the second stage corresponds to a first number of inference steps, and the first number is determined based on a first weight of the first random noise.
4. The method of claim 3, wherein the first number is positively correlated with the first weight.
5. The method of claim 3, wherein performing at least one round of noise reduction on the second noise data comprises: Based on the first offset coefficient corresponding to the second stage, a first sampling scheduler corresponding to the second stage is constructed; Based on the first sampling scheduler, multiple noise reduction time steps corresponding to the first number of inference steps are determined; as well as Based on the multiple noise reduction time steps, at least one round of noise reduction is performed on the second noise data.
6. The method according to claim 1, further comprising: The fourth feature representation is determined by upsampling the second feature representation; as well as A second random noise is added to the fourth feature representation to construct the third noise data.
7. The method of claim 6, wherein the third stage corresponds to a second number of inference steps, and the second number is determined based on a second weight of the second random noise.
8. The method of claim 7, wherein the second number is positively correlated with the second weight.
9. The method of claim 7, wherein performing at least one round of noise reduction on the third noise data comprises: Based on the second offset coefficient corresponding to the third stage, a second sampling scheduler corresponding to the third stage is constructed; Based on the second sampling scheduler, multiple noise reduction time steps corresponding to the second number of inference steps are determined; as well as Based on the multiple noise reduction time steps, at least one round of noise reduction is performed on the third noise data.
10. The method according to claim 1, wherein the noise reduction process further comprises a fourth stage and a fifth stage. The fourth stage is configured to perform at least one round of noise reduction on the fourth noise data to determine a fifth feature representation, wherein the first noise data is constructed based on the fifth feature representation by downsampling; The fifth stage is configured to perform at least one round of noise reduction on the fifth noise data, which is based on a sixth feature representation constructed by upsampling, the sixth feature representation being generated by the third stage.
11. An apparatus for generating media content, comprising: The receiving module is configured to receive media generation requests associated with the diffusion model; as well as The generation module is configured to generate media content for the media generation request based on the noise reduction process of the diffusion model, wherein the noise reduction process includes at least a first stage, a second stage, and a third stage. The first stage is configured to perform at least one round of noise reduction on the first noisy data to determine a first feature representation; The second stage is configured to perform at least one round of noise reduction on the second noise data to determine a second feature representation, the second noise data being constructed based on downsampling the first feature representation; The third stage is configured to perform at least one round of noise reduction on the third noise data, which is constructed based on upsampling of the second feature representation.
12. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.
13. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method according to any one of claims 1 to 10.
14. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 10.