Content generation method and device, computer program product and electronic equipment

By obtaining user intention encoding information in the content generation method and performing conditional modulation processing, the problem of not being able to fully utilize conditional input in the prior art is solved, and fine control and quality improvement of generated content is achieved.

CN120144918APending Publication Date: 2025-06-13NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510209968.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing content generation methods based on diffusion models cannot make timely and fully utilize conditional input, affecting the accuracy and quality of generated content.

Method used

By obtaining user intention coding information, using the preset diffusion model to process the noise latent variable, extract the latent variable feature map, and perform conditional modulation processing to generate a gating weight map, fusing the feature map as the input of the subsequent network layer, denoising processing, and finally generating the target content.

Benefits of technology

It realizes fine control of generated content, improves the quality and interactivity of content generation, and improves user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144918A_ABST
    Figure CN120144918A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, in particular to a content generation method and device, a computer program product and electronic equipment. The content generation method comprises the following steps: acquiring user intention coding information, and processing a noise latent variable at the current moment through a preset diffusion model to extract a latent variable feature map output by a target intermediate layer of the preset diffusion model; performing conditional modulation processing on the latent variable feature map according to the user intention coding information to obtain a modulated latent variable feature map, and performing prediction based on the user intention coding information to obtain a gating weight map; fusing the modulated latent variable feature map and the gated weight map to obtain a fused feature map, taking the fused feature map as an input of a subsequent network layer of the target intermediate layer, and performing denoising processing on the noise latent variable to obtain a denoised noise latent variable; and generating target content according to the denoised noise latent variable. According to the invention, fine control on the generated content can be realized according to the action information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and more particularly, to a content generation method, a content generation device, a computer program product, and an electronic device. Background Art

[0002] With the development of the field of computer technologies, using pre-trained models for content generation can promote the improvement of content quality and generation efficiency. Currently, in the method of content generation based on diffusion models, the generation process can be guided by conditional inputs. However, there is still a problem that conditional inputs cannot be utilized timely and fully, which to a certain extent affects the accuracy of the generated content and results in insufficient quality of the generated content.

[0003] It should be noted that the information disclosed in the above background art is only used to enhance the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0004] An object of the present disclosure is to provide a content generation method, a content generation device, a computer program product, and an electronic device, so as to realize fine control of the generated content according to action information and improve the quality and interactivity of content generation.

[0005] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be partially learned through the practice of the present disclosure.

[0006] According to one aspect of the present disclosure, there is provided a content generation method, including: obtaining user intent encoding information, and processing the noise latent variable at the current moment through a preset diffusion model to extract the latent variable feature map output by the target intermediate layer of the preset diffusion model; performing conditional modulation processing on the latent variable feature map according to the user intent encoding information to obtain a modulated latent variable feature map, and predicting a gating weight map based on the user intent encoding information, where the gating weight map is used to indicate the degree of attention to different feature regions; fusing the modulated latent variable feature map and the gating weight map to obtain a fused feature map, and using the fused feature map as the input of the subsequent network layer of the target intermediate layer to perform denoising processing on the noise latent variable to obtain a denoised noise latent variable; generating target content according to the denoised noise latent variable.

[0007] In an exemplary embodiment of the present disclosure, obtaining user intent encoding information includes: receiving action instruction information, and extracting object category information and object attribute information from the action instruction information; respectively performing encoding processing on the object category information and the object attribute information to obtain a plurality of encoding results, and fusing the plurality of encoding results to obtain user intent encoding information.

[0008] In an exemplary embodiment of the present disclosure, an action-aware gating network is embedded in at least one target intermediate layer of a preset diffusion model. The action-aware gating network includes a conditional modulation network and a gating network. Conditionally modulating the latent variable feature map according to the user intention encoding information to obtain a modulated latent variable feature map, and predicting a gating weight map based on the user intention encoding information includes: based on the user intention encoding information, using the conditional modulation network to perform an affine transformation on the latent variable feature map to obtain a modulated latent variable feature map; performing feature selection processing on the user intention encoding information through the gating network to generate a gating weight map.

[0009] In an exemplary embodiment of the present disclosure, based on the user intention encoding information, using the conditional modulation network to perform an affine transformation on the latent variable feature map to obtain a modulated latent variable feature map includes: based on the user intention encoding information, using the first multi-layer perceptron network in the conditional modulation network to predict a scaling factor, where the scaling factor is used to adjust the amplitude of the latent variable feature map in different dimensions; through the second multi-layer perceptron network in the conditional modulation network, predicting a bias term according to the user intention encoding information, where the bias term is used to adjust the value distribution of the latent variable feature map; performing an affine transformation on the latent variable feature map according to the scaling factor and the bias term to obtain a modulated latent variable feature map.

[0010] In an exemplary embodiment of the present disclosure, the action-aware gating network further includes a feature fusion network. Fusing the modulated latent variable feature map and the gating weight map includes: multiplying the modulated latent variable feature and the gating weight map element-wise through the feature fusion network to obtain a fused feature map.

[0011] In an exemplary embodiment of the present disclosure, using the fused feature map as the input of the subsequent network layer of the target intermediate layer to perform denoising processing on the noisy latent variable to obtain a denoised noisy latent variable includes: fusing the fused feature map with the original features of the subsequent network layer to obtain a target feature; predicting noise based on the noisy latent variable, the user intention encoding information, the time step information, and the target feature to obtain the noise at the current moment, and performing denoising processing on the noisy latent variable according to the noise at the current moment to obtain a denoised noisy latent variable.

[0012] In an exemplary embodiment of the present disclosure, generating a target content according to the denoised noisy latent variable includes: if the current denoising iteration number is reached, performing latent variable decoding on the denoised noisy latent variable to obtain the target content; if the current denoising iteration number is not reached, determining the denoised noisy latent variable as the noisy latent variable at the next moment and performing the denoising process.

[0013] According to one aspect of the present disclosure, there is provided a content generation device, including: an information acquisition module, configured to acquire user intention encoding information and process the noise latent variable at the current moment through a preset diffusion model to extract the latent variable feature map output by the target intermediate layer of the preset diffusion model; a feature processing module, configured to perform conditional modulation processing on the latent variable feature map according to the user intention encoding information to obtain a modulated latent variable feature map, and predict a gating weight map based on the user intention encoding information, where the gating weight map is used to indicate the degree of attention to different feature regions; a generation control module, configured to fuse the modulated latent variable feature map and the gating weight map to obtain a fused feature map, and use the fused feature map as the input of the subsequent network layer of the target intermediate layer to perform denoising processing on the noise latent variable to obtain a denoised noise latent variable; and a content generation module, configured to generate target content according to the denoised noise latent variable.

[0014] According to one aspect of the present disclosure, there is provided a computer program product, including a computer program, which when executed by a processor, implements the method according to any one of the above.

[0015] According to one aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the method according to any one of the above by executing the executable instructions.

[0016] In the content generation method in the exemplary embodiments of the present disclosure, the noise latent variable at the current moment is processed through a preset diffusion model to extract the latent variable feature map output by the target intermediate layer of the preset diffusion model, conditional modulation processing is performed on the latent variable feature map according to the user intention encoding information to obtain a modulated latent variable feature map, a gating weight map is predicted based on the user intention encoding information, the modulated latent variable feature map and the gating weight map are fused, and the obtained fused feature map is used as the input of the subsequent network layer of the target intermediate layer to perform denoising processing on the noise latent variable to obtain a denoised noise latent variable. On the one hand, the user intention encoding information can be incorporated into the content generation process, and the features of the content can be adaptively adjusted according to the user intention during the denoising process of the preset diffusion model, so as to achieve the guidance of the generated content and avoid simple random generation, realizing the refined control of the generated content. On the other hand, by fully integrating the user intention into the processing process of the preset diffusion model, the generated content can be made highly consistent with the user intention, improving the interactivity of content generation and being beneficial to enhancing the user's interactive experience during content generation. On the further hand, during the process of controlling content generation based on the user intention, the processing process of the preset diffusion model is carried out in the latent space, which can reduce the computational complexity and improve the content generation efficiency.

[0017] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] By referring to the accompanying drawings and reading the following detailed description, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown in an exemplary rather than restrictive manner, where:

[0019] Figure 1 FIG. shows an application environment diagram related to a content generation method according to an exemplary embodiment of the present disclosure.

[0020] Figure 2 FIG. shows a flowchart of a content generation method according to an exemplary embodiment of the present disclosure.

[0021] Figure 3 FIG. shows a schematic structural diagram of a preset diffusion model according to an exemplary embodiment of the present disclosure.

[0022] Figure 4 FIG. shows a flowchart of a method for obtaining user intent encoding information according to an exemplary embodiment of the present disclosure.

[0023] Figure 5 FIG. shows a schematic structural diagram of an action perception gating network according to an exemplary embodiment of the present disclosure.

[0024] Figure 6 FIG. shows a flowchart of a conditional modulation process according to an exemplary embodiment of the present disclosure.

[0025] Figure 7 FIG. shows a complete flowchart of a content generation method according to an exemplary embodiment of the present disclosure.

[0026] Figure 8 FIG. shows a schematic composition diagram of a content generation device according to an exemplary embodiment of the present disclosure.

[0027] Figure 9 FIG. shows a block diagram of an electronic device according to an exemplary embodiment of the present disclosure.

[0028] In the drawings, the same or corresponding reference numerals indicate the same or corresponding parts. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. Like reference numerals in the figures denote like or similar structures, and thus their detailed description will be omitted.

[0030] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, well-known structures, methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.

[0031] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more software-hardened modules, or in different networks and / or processor devices and / or microcontroller devices.

[0032] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.

[0033] Artificial intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0034] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and the theory of algorithm complexity. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning generally include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.

[0035] The technical solution provided by the exemplary embodiments of the present disclosure relates to the machine learning technology of artificial intelligence, and uses the machine learning technology to effectively integrate the user intention (such as an action instruction) into the content generation process to achieve fine-grained control of the generated content.

[0036] It should be noted that the content generation method of the exemplary embodiments of the present disclosure can be applied to application scenarios such as virtual reality, augmented reality, game development, and robot training. For example, 3D animation assets can be generated according to the user intention, and there is no limitation in this regard.

[0037] The content generation method provided by the exemplary embodiments of the present disclosure can be applied to an application environment as Figure 1 shown. Among them, the terminal 101 communicates with the server 102 through the network. The data storage system can store the data that the server 102 needs to process. The data storage system can be integrated on the server 102, or can be placed in the cloud or on other network servers.

[0038] In one exemplary embodiment, the content generation method provided by the exemplary embodiments of the present disclosure can be executed by the server 102. Correspondingly, the content generation device is set in the server 102. Correspondingly, in this manner executed by the server 102, the server 102 can start executing the steps in the technical solution of the exemplary embodiments of the present disclosure in response to a trigger execution order, where the trigger execution order can be sent by the terminal used by the user, or can be locally triggered by the server in response to some automated events.

[0039] The server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server 102 can execute background tasks.

[0040] Furthermore, in another exemplary embodiment, the terminal 101 may also have a similar function to the server 102, so as to execute the content generation method provided by the exemplary embodiment of the present disclosure.

[0041] Among them, the terminal 101 may be a smart phone, a tablet computer, a laptop computer, or a desktop computer. The terminal 101 may also be referred to as a mobile terminal, a terminal device, a mobile device, etc. The exemplary embodiment of the present disclosure does not limit the type of the terminal 101.

[0042] In addition, the technical solution of the exemplary embodiment of the present disclosure may also be executed jointly by the terminal 101 and the server 102. In this way of joint execution by the terminal 101 and the server 102, some steps in the technical solution provided by the exemplary embodiment of the present disclosure are executed by the terminal 101, while some other steps are executed by the server 102. It should be noted that, in this way of joint execution by the terminal 101 and the server 102, the steps respectively executed by the terminal 101 and the server 102 can be dynamically adjusted according to the actual situation, and no special limitation is imposed thereon.

[0043] Among them, the terminal 101 and the server 102 may be directly or indirectly connected through a wireless communication method, and the exemplary embodiment of the present disclosure does not impose any special limitation thereon.

[0044] Currently, when generating content based on a diffusion model, the denoising process of the model can be controlled by a conditional input (such as a text description or a global feature vector). However, only the conditional input is used as a guiding condition of the model, such as a positive guiding condition or a negative guiding condition, and the conditional input cannot be fully incorporated into the content generation process. That is, the features of the content cannot be adaptively adjusted according to the user's intention during the generation process, and thus the more refined generation of the generated content cannot be achieved, which affects the content generation quality and the interactive experience.

[0045] Based on this, the exemplary embodiment of the present disclosure provides a content generation method, which intervenes in the features of the content generation process of the diffusion model according to the user's intention, so that the generated content fully conforms to the user's intention and realizes the refined control of the content generation result.

[0046] As Figure 2 shown is a flowchart of the content generation method of the exemplary embodiment of the present disclosure. Referring to Figure 2 shown, the content generation method includes steps S210 to S240, which are specifically as follows:

[0047] In step S210, obtain the user intention encoding information, and process the noise latent variable at the current moment through a preset diffusion model to extract the latent variable feature map output by the target intermediate layer of the preset diffusion model.

[0048] In an exemplary embodiment of the present disclosure, the user intention encoding information is the encoding information corresponding to the user's action instruction information. That is to say, the user intention encoding information is generated based on the user's input, so as to integrate the user intention into the content generation process subsequently. The action instruction information reflects the actual needs of the user for content generation. For example, the action instruction information is "jump", "move forward and turn around", "place a red sofa". In addition, to further enable the model to efficiently process data in a low-dimensional space, the user intention encoding information can also be mapped to a low-dimensional dense vector space to obtain an embedding vector. During the dimensionality reduction process, the embedding vector retains the main information of the original data but can greatly reduce the computational complexity. The subsequent description of the user intention encoding information in the exemplary embodiment of the present disclosure is based on the example that it has been mapped to a low-dimensional dense vector space.

[0049] Among them, the diffusion model: is a deep learning technology used to generate high-quality images, audio, text, or other data types. It belongs to a class of generative models, and its purpose is to learn the distribution of a given data set so as to be able to generate new, similar instances. The diffusion model works by simulating a step-by-step diffusion process. First, the data is gradually converted into unstructured random noise, and then this process is reversed to recover meaningful data from the noise. In the exemplary embodiment of the present disclosure, an action-aware gating unit (AAGU) is embedded in a preset diffusion model to refine the latent variable features during the denoising process by using the action-aware gating network with the user intention.

[0050] Specifically, as Figure 3 schematically shows a structural diagram of a preset diffusion model. As Figure 3 , the preset diffusion model includes an encoder, a decoder, and a bottleneck layer between the encoder and the decoder, and an action-aware gating network is embedded in at least one target intermediate layer of the preset diffusion model. The action-aware gating network is similar to a switch and can control the flow of information. In the exemplary embodiment of the present disclosure, through the action-aware gating network, which features information affects the generation result is controlled according to the user intention.

[0051] The preset diffusion model includes an encoder, a decoder, and a bottleneck layer between the encoder and the decoder. This part is similar to the structure of the U-Net-based diffusion model, and no excessive restrictions are imposed on this. The target intermediate layer of the preset diffusion model can be the bottleneck layer between the encoder and the decoder, where an action-aware gating network is embedded in the bottleneck layer. Alternatively, an action-aware gating network can also be embedded between certain corresponding layers of the encoder and the decoder. Or, an action-aware gating network can also be inserted into multiple intermediate layers of the diffusion model (such as U-Net), enabling the user's intention to play a role in features at different scales. Among them, it should be understood that shallower layers extract low-level features (such as edges, textures), and deeper layers extract high-level semantic features. Selecting appropriate layers to insert the action-aware gating network allows the user's intention to play a role in features at different levels of abstraction. Moreover, inserting the action-aware gating network in deeper layers, the size of the feature maps processed is smaller, and the computational cost is relatively lower. Therefore, specifically which intermediate layer(s) to insert the action-aware gating network can be determined according to the actual task scenario and requirements, and the exemplary embodiments of the present disclosure do not impose restrictions on this.

[0052] It is worth noting that the denoising process of the diffusion model is iterative. The output of each step (the denoised noise latent variable corresponding to each moment t) will be used as the input for the next step (the next moment), gradually removing the noise until a clear latent variable representation is finally obtained. Therefore, in each denoising step, the noise latent variable at the current moment is obtained with the denoised noise latent variable corresponding to the previous moment as the input. The process of processing the noise latent variable at the current moment through the preset diffusion model is similar to the process of processing the noise latent variable by a conventional diffusion model, and this will not be elaborated here. However, since the exemplary embodiments of the present disclosure need to incorporate the user intention encoding information into the denoising process, it is necessary to extract the latent variable features output by the target intermediate layer of the preset diffusion model to adjust the latent variable features using the user intention encoding.

[0053] In addition, the preset diffusion model first adds noise to the data step by step (diffusion process), and then learns how to reverse this process to generate new data samples (reverse generation process, i.e., denoising process). Specifically, in the forward diffusion process, noise is gradually added to the original data until the data is completely perturbed into pure noise. A Markov chain can be used, and a small amount of Gaussian noise is added to the data at each step. For example, Formula 1 is a Markov chain that describes how to gradually add noise to the clear latent variable:

[0054]

[0055] where, N represents the probability density function of the Gaussian distribution, β t is the noise scheduling parameter, I is the identity matrix, z t is the noise latent variable at the current moment, q(zt |z t-1 ) is the conditional probability distribution, representing the distribution of z t-1 given that z t is known. β t is the noise schedule parameter that controls the amount of noise added to the latent variable in each step of the forward diffusion process. The larger the value of β t , the more noise is added. β t can be preset according to the task requirements.

[0056] Exemplarily, β t can be set by the linear schedule method, t such that the value of β t increases linearly from a smaller value to a larger value. Alternatively, β t can be set by the cosine schedule method, i.e., the value of β t changes according to the shape of the cosine function. Of course, other non-linear functions can also be used to control the change of the value of β

[0057] Correspondingly, the reverse generation process starts from pure noise and gradually denoises to restore new data samples. This process requires predicting and removing the noise in each step. The exemplary embodiments of the present disclosure mainly improve the denoising process by using an action-aware gating network as part of the denoising network structure to guide the denoising process according to the user's intention, and finally predict and generate a "clean" latent variable that better conforms to the user's intention.

[0058] As Figure 3 shown, the latent variable feature map output by the target intermediate layer of the preset diffusion model refers to the latent variable feature map obtained by processing the current moment's noise latent variable through the preset diffusion model to the target intermediate layer.

[0059] In step S220, the latent variable feature map is conditionally modulated according to the user intention encoding information to obtain a modulated latent variable feature map, and a gating weight map is predicted based on the user intention encoding information. The gating weight map is used to indicate the degree of attention to different feature regions.

[0060] In an exemplary embodiment of the present disclosure, conditional modulation processing of the latent variable feature map refers to adjusting the numerical characteristics of the latent variable feature map conditional on the user intention encoded information. For example, features related to red color and sofa shape are emphasized more. Predicting the gating weight map based on the user intention encoded information means that, based on the user intention encoded information, the model can selectively focus on different regions of the image. It can be understood that the core role of the gating mechanism is to enable the model to focus more on the feature information related to the user intention (such as action instructions) according to the user intention, thereby avoiding excessive processing of irrelevant information. The gating weight map is like a mask that can be used to selectively retain or suppress information in the modulated latent variable feature map and can also improve the potential computational efficiency.

[0061] Among them, the conditional adjustment can affect attributes such as the color and texture of the generated content, and the gating mechanism can affect the object composition and spatial layout in the generated scene, etc. The action-aware gating network is used to implement the conditional modulation processing of the latent variable feature map and predict the gating weight map.

[0062] In step S230, the modulated latent variable feature map and the gating weight map are fused to obtain a fused feature map, and the fused feature map is used as the input of the subsequent network layer of the target intermediate layer to perform denoising processing on the noise latent variable to obtain the denoised noise latent variable.

[0063] In an exemplary embodiment of the present disclosure, after fusing the modulated latent variable feature map and the gating weight map, a fused feature map is obtained, and then the fused feature map is passed to the subsequent network layer of the target intermediate layer to participate in the denoising process to obtain the denoised noise latent variable, that is, the denoised noise latent variable corresponding to the current moment.

[0064] In step S240, the target content is generated according to the denoised noise latent variable.

[0065] In an exemplary embodiment of the present disclosure, since the denoising process is iterative, if the current denoising iteration count is reached, the denoised noise latent variable is decoded to obtain the target content. Specifically, a decoder can be used to decode the denoised noise latent variable to map the features in the latent space to the target content, such as 3D data content or 2D image content.

[0066] If the current denoising iteration count is not reached, the denoised noise latent variable is determined as the noise latent variable for the next moment and the denoising process is performed. That is to say, with the noise latent variable at the current moment as the model input, the next denoising step is continued.

[0067] In the content generation method of the exemplary embodiments of the present disclosure, on the one hand, the user intention encoding information can be incorporated into the content generation process, and during the denoising process of the preset diffusion model, the features of the content can be adaptively adjusted according to the user intention, realizing the guidance of the generated content, avoiding simple random generation, and achieving fine control of the generated content. On the other hand, by fully integrating the user intention into the processing process of the preset diffusion model, it can also make the generated content highly consistent with the user intention, improve the interactivity of content generation, and is beneficial to enhancing the user's interactive experience during content generation. On the further hand, during the process of controlling content generation based on the user intention, the processing process of the preset diffusion model is carried out in the latent space, which can reduce the computational complexity and improve the content generation efficiency.

[0068] In one exemplary embodiment, as Figure 4 shown, obtaining the user intention encoding information may include:

[0069] Step S410: Receive action instruction information, and extract object category information and object attribute information from the action instruction information.

[0070] The action instruction information refers to the relevant information provided by the user or the intelligent agent for guiding content generation. The action instruction information may include different types of object information or attribute information. The object category information and the object attribute information can be extracted from the action instruction information respectively. Among them, for example, in the action instruction information of "place a red sofa", "place" and "sofa" belong to the object category information (action category, object type), and "red" belongs to the attribute category information.

[0071] Step S420: Perform encoding processing on the object category information and the object attribute information respectively to obtain multiple encoding results, and fuse the multiple encoding results to obtain the user intention encoding information.

[0072] The object category information and the object attribute information can be encoded and fused respectively to obtain the user intention encoding information, and the fusion method can be splicing.

[0073] Exemplarily, for a set A = {a 1 , a 2 ,..., a N} containing N discrete actions, the action a iIt can be encoded as a one-hot vector vi ∈ {0, 1}^N, where vi[i] = 1 and the remaining elements are 0. Here, the set A refers to the set of all possible user or agent actions. Each time it is input as "action instruction information", it can be one of the following two cases: "a single discrete action" and "a combination of multiple discrete actions". That is to say, the action instruction information of the exemplary embodiments of the present disclosure can be a single discrete action or a combination of multiple discrete actions according to the actual interaction complexity and task requirements, with relatively high flexibility.

[0074] "A single discrete action". For example, when the user performs a "jump" action, the input action instruction information corresponds to the encoding of the "jump" action.

[0075] "A combination of multiple discrete actions". For example, in a more complex scenario, the user may perform multiple actions simultaneously or within a very short time, such as "move forward and turn around". At this time, these actions can be combined as an action instruction information for encoding. Among them, the combination method of these actions can be multi-label one-hot encoding. If multiple actions can occur simultaneously, multi-label one-hot encoding can be used, that is, the corresponding actions are marked as 1 in the encoding vector. Sequence encoding. If the actions have a sequence, a sequence model (such as RNN (Recurrent Neural Network) or Transformer) can be used to encode the action sequence. In addition, for one-hot encoding, if the action has semantic information, a word embedding model (such as Word2Vec) can also be used to encode it into a dense vector, and for continuous actions (such as movement speed, rotation angle), their numerical values can be directly used for encoding and normalization processing. The exemplary embodiments of the present disclosure can select the encoding method according to the actual task situation.

[0076] By encoding different aspects of information separately, it is possible to more clearly represent the different aspects of information contained in the action instruction information, so that instructions containing different combinations of objects and attributes can be conveniently processed, and the encoding results are also easier to understand and analyze.

[0077] In an exemplary embodiment, as Figure 5 shown, the action perception gating network includes a conditional modulation network and a gating network. Among them, performing conditional modulation processing on the latent variable feature map according to the user intention encoding information to obtain a modulated latent variable feature map, and predicting the gating weight map based on the user intention encoding information may include:

[0078] Based on the user intention encoding information, using the conditional modulation network to perform an affine transformation on the latent variable feature map to obtain a modulated latent variable feature map;

[0079] The gating network performs feature selection processing by encoding information according to the user's intention to generate a gating weight map.

[0080] Among them, the conditional modulation network is used to adjust the latent variable feature map according to the encoded information of the user's intention. A multilayer perceptron network (MLP, Multilayer Perceptron) can be adopted. The gating network can adopt a CNN network. By learning the mapping relationship between the action and the gating weight map, a gating weight map that can selectively activate or inhibit features can be generated according to the encoded information of the user's intention. Based on the gating network, the action-aware gating network can pay more attention to the features related to the current user's intention, reduce unnecessary calculations, improve the generation efficiency, and accelerate the convergence of model training.

[0081] Specifically, as Figure 6 shown, based on the encoded information of the user's intention, an affine transformation is performed on the latent variable feature map by using the conditional modulation network. The obtained modulated latent variable feature map may include:

[0082] Step S610: Based on the encoded information of the user's intention, use the first multilayer perceptron network in the conditional modulation network to predict the scaling factor.

[0083] The scaling factor is used to adjust the amplitude of the latent variable feature map in different dimensions. The prediction of the scaling factor by the first multilayer perceptron network can be represented by Formula 2:

[0084]

[0085] where MLP γ represents using the first multilayer perceptron network to predict the scaling factor according to the encoded information of the user's intention, and γ ea is the predicted scaling factor, and e a represents the encoded information of the user's intention (after being embedded and represented, not repeated hereinafter).

[0086] Step S620: Through the second multilayer perceptron network in the conditional modulation network, predict the bias term according to the encoded information of the user's intention.

[0087] The bias term is used to adjust the value distribution of the latent variable feature map. The prediction of the bias term by the second multilayer perceptron network can be represented by Formula 3:

[0088]

[0089] where MLP β represents using the second multilayer perceptron network to predict the bias term according to the encoded information of the user's intention, and β ea is the predicted bias term.

[0090] Step S630: Perform an affine transformation on the latent variable feature map according to the scaling factor and the bias term to obtain a modulated latent variable feature map.

[0091] Performing an affine transformation on the latent variable feature map according to the scaling factor and the bias term is to adjust the feature values of the latent variable feature map so that the modulated latent variable feature map can better adapt to the user intention encoded information.

[0092] Among them, the affine transformation of the latent variable feature map can be represented by Formula 4:

[0093]

[0094] Among them, F is the latent variable feature map, ⊙ represents element-wise multiplication, and F’ is the modulated latent variable feature map.

[0095] Taking the user intention encoded information as the input to implement conditional adjustment of the latent variable feature map, so that the modulated latent variable feature map can better adapt to the user intention and achieve fine-grained control of the generated content.

[0096] In an exemplary embodiment, continue to refer to Figure 5 , the action-aware gating network further includes a feature fusion network. The fusion of the modulated latent variable feature map and the gating weight map may include:

[0097] Multiply the modulated latent variable feature and the gating weight map element-wise through the feature fusion network to obtain a fused feature map.

[0098] Among them, the feature fusion network is used to fuse the outputs of the conditional modulation network and the gating network. The fusion process can be represented by Formula 5:

[0099] F out = G⊙F′ Formula 5

[0100] Among them, F out is the fused feature map, G is the gating weight map, and ⊙ represents element-wise multiplication.

[0101] The gating weight map has the same spatial dimension as the feature map, enabling the action-aware gating network to achieve spatial action awareness. The model can selectively focus on different regions of the image according to the user intention. Moreover, through decoupled conditional modulation (adjusting feature values) and gating mechanism (selecting features), the model can flexibly process different types of actions and features. Since the scaling factor, bias term, and gating weight are all dynamically generated based on the user intention, the model can adaptively adjust the internal feature processing method according to different user intentions (the influence of user intention on features), achieve fine-grained control of the generated content, and improve the content generation quality.

[0102] In an exemplary embodiment, taking the obtained fused feature map as the input of the subsequent network layer of the target intermediate layer and performing denoising processing on the noise latent variable to obtain the denoised noise latent variable may include:

[0103] Fusing the fused feature map with the original features of the subsequent network layer to obtain target features;

[0104] Performing noise prediction based on the noise latent variable, user intention encoding information, time step information, and target features to obtain the noise at the current moment, and performing denoising processing on the noise latent variable according to the noise at the current moment to obtain the denoised noise latent variable.

[0105] See also Figure 5 , after obtaining the fused feature map through the action-aware gating network, the fused feature map can be inserted back into the target intermediate layer as the input of the subsequent network layer of the target intermediate layer (i.e., re-embedded into the backbone network of the preset diffusion model) and participate in the subsequent denoising calculation process.

[0106] The noise prediction in the denoising network part of the preset diffusion model can be represented by Equation 6:

[0107] ε θ (z t , t, e a ) Equation 6

[0108] where εθ is the denoising network, z t is the noise latent variable, t is the time step, e a is the user intention encoding information. After obtaining the fused feature map, the fused feature map is fused with other features of the subsequent network layer to control the noise prediction process, that is, jointly used to predict the noise εθ(z t , t, e a ) at the current moment. Then, using the predicted noise, the noise is removed from the noise latent variable z t at the current moment to obtain the denoised noise latent variable.

[0109] Specifically, the goal of the denoising process is to gradually remove the noise and predict the latent variable at the previous moment:

[0110] P θ (z t-1 |z t ) = N(z t-1 ; μ θ (z t , t, e a ), ∑ θ (z t , t)) Equation 7

[0111] where F θ (zt-1 |z t ) is the conditional probability distribution of the denoising process, indicating that given z t Under the condition of t-1 The probability of N(z t-1 ;μ θ (z t ,t,e a ),∑ θ (z t , t)) describes the model prediction of the previous moment latent variable z t-1 , subject to the mean μ θ , covariance∑ θ Gaussian distribution of , θ is the parameter related to the denoising network.

[0112] The user intention encoding information is integrated into the content generation process, and the content features are adaptively adjusted according to the user intention during the denoising process of the preset diffusion model, so as to guide the generated content and achieve refined control over the generated content.

[0113] In an exemplary embodiment, in the denoising process of the preset diffusion model, the time step t starts from the preset maximum value T and gradually decreases to 1, which can be understood as the progress bar of the denoising process. The discrete time step t can be encoded into a continuous vector representation, and each time step corresponds to a learnable vector. Then, the time step information can be provided to the denoising network so that the denoising network knows which stage of the denoising process it is currently in. For example, in the early stage of denoising (t is large), there is more noise, and the network needs to learn to remove a large amount of noise; in the late stage of denoising (t is small), there is less noise, and the network needs to learn to finely restore image details. Among them, the preset maximum value T can be set according to the actual task requirements. For example, if a longer denoising time is required, the maximum value T can be increased. On the contrary, if fewer denoising times are required, the maximum value T can be reduced. The exemplary embodiments of the present disclosure do not limit the specific value of the maximum value T.

[0114] In addition, it should be noted that the dimension and structure of the latent space of the exemplary embodiment of the present disclosure are determined based on the design of the autoencoder, which can be a 2-dimensional feature map, a 3-dimensional feature map, or even an abstract representation of a higher dimension. The key is that the encoder part can map the data to be processed to the latent space, and the decoder part can restore the representation of the latent space back to data of the corresponding dimension.

[0115] For example, an encoder can effectively map 3D data to a latent space, while a decoder can restore the representation in the latent space back to 3D data. Therefore, the content generation method of the exemplary embodiments of the present disclosure can be applicable to generating 2D and 3D content. If used to generate 2D content, the data for training the model will be 2D image data, and the latent space feature maps processed by the autoencoder and the diffusion model will also be 2D. Correspondingly, the operations involved in each network layer in the action-aware gating network are adapted to the dimensions of the 2D feature maps, and the decoder can decode the latent space representation into 2D image content. Similarly, if used to generate 3D content, the relevant content involved will be 3D.

[0116] Such as Figure 7 Fig. shows a complete flowchart of a content generation method. Taking the scenario where a user interacts with a virtual indoor environment as an example, the content generation method of the exemplary embodiments of the present disclosure will be described. Among them, the user needs to place a red sofa in the virtual living room. This is an illustration taking the generation of 3D content as an example.

[0117] First, obtain the user's action instruction information "place a red sofa".

[0118] Among them, object category information and object attribute information are extracted from the action instruction information "place a red sofa", and are respectively encoded and fused to obtain user intention encoding information.

[0119] Specifically, "place" and "sofa" are object category information, and "red" is object attribute information. If there is a preset set containing action types such as "place", "move", "delete", etc., the "place" action can use One-Hot encoding, for example, a_type = [1, 0, 0], "sofa" can be represented by a pre-trained word vector, for example, v_category = [0.2, -0.5, 0.8,...], and "red" can be represented by the embedding of RGB values or color names, for example, v_color = [0.9, 0.1, 0.1]. The encoding results of the object category information and the object attribute information can be concatenated to obtain the user intention encoding information a = [1, 0, 0, 0.2, -0.5, 0.8,..., 0.9, 0.1, 0.1].

[0120] Secondly, assume a denoising step t (i.e., the current moment is t), and the input of the model is the noise latent variable z at the current moment t .

[0121] Among them, the noise latent variable z at the current moment t reaches the action-aware gating network through the target intermediate layer of the preset diffusion model, and the output of the action-aware gating network will be used as the input of the subsequent network layers of the target intermediate layer.

[0122] Specifically, the action perception gating network may include an action embedding network, a conditional modulation network, a gating network, and a feature fusion network. The specific process includes:

[0123] a. Use the action embedding network to perform action embedding on the user intention encoding information, where the action embedding network can be a multi-layer perceptron network, and the embedding process can be represented by Formula 8:

[0124] e a' = MLP q emb(a) Formula 8

[0125] where e a' is the user intention encoding information after the embedding representation, MLP q represents performing an embedding representation on the user intention encoding information, and a is the user intention encoding information before the embedding representation.

[0126] b. Based on the user intention encoding information after the embedding representation, use the first multi-layer perceptron network in the conditional modulation network to predict the scaling factor, use the second multi-layer perceptron network in the conditional modulation network to predict the bias term according to the user intention encoding information after the embedding representation, and perform an affine transformation on the latent variable feature map according to the scaling factor and the bias term to obtain the modulated latent variable feature map.

[0127] For example, for a certain feature channel representing the shape of an object, if the corresponding value of the scaling factor is 1.2 and the corresponding value of the bias term is 0.1, then the modulated result of this channel is: F’ = 1.2 * F + 0.1. Based on this, the features related to the shape of the sofa can be enhanced. For example, for the feature channel representing red, the scaling factor and the bias term may be adjusted to values that encourage red activation.

[0128] c. Send the user intention encoding information after the embedding representation to the gating network. The gating network can adopt a CNN network and can generate a gating weight map that can selectively activate or inhibit features according to the user intention encoding information, which can be represented by Formula 9:

[0129] G = CNN gate (e a ) Formula 9

[0130] where G is the gating weight map, and CNN() represents that the gating network performs feature selection processing on the user intention encoding information after the embedding representation.

[0131] For example, if the gating network learns that the action of "placing the sofa" needs to focus on the central area of the ground, then the value of the gating weight map G at the pixel position corresponding to the central area of the living room floor is close to 1, and the values at other positions are close to 0.

[0132] d. Fuse the modulated latent variable feature map and the gating weight map to obtain a fused feature map, and use the obtained fused feature map as the input for the subsequent network layers of the target intermediate layer.

[0133] The modulated latent variable feature map and the gating weight map can be multiplied element-wise. Based on this, only the feature information related to the shape and color of the sofa in the central area of the ground can be retained, while the feature information in other areas or unrelated feature information is suppressed. In addition, through the conditional modulation network and the gating network, users can control the attributes (color, shape) and spatial layout (placement position) of the generated object. For example, the user can further instruct "move the sofa 0.5 meters to the left", and the action-aware gating network will adjust the features again according to the new user intention information to achieve fine control of the scene. The implementation method is similar and will not be elaborated here.

[0134] Next, the obtained fused feature map is passed to the subsequent layer of the target intermediate layer to participate in predicting the noise at this step, obtaining the predicted noise, and performing denoising processing based on the predicted noise to obtain the denoised noise latent variable. The model will tend to predict the noise that can generate a red sofa at this position.

[0135] Finally, generate the target content according to the denoised noise latent variable.

[0136] Among them, if the current denoising iteration number is reached, the denoised noise latent variable is decoded for the latent variable to obtain the target content. For example, if the denoised noise latent variable is decoded by the decoder into a visual 3D world state, the user can see "a red sofa is placed in the center of the living room floor". Of course, if the current denoising iteration number is not reached, the denoised noise latent variable is determined as the noise latent variable at the next moment and the denoising process is performed, that is, the above steps are re-executed until the stop condition is met (such as reaching the denoising iteration number).

[0137] It should be noted that the specific formulas and other details involved in each step have been described in the above exemplary embodiments and will not be elaborated here.

[0138] The content generation method in the exemplary embodiments of the present disclosure, on the one hand, can incorporate user intention encoding information into the content generation process, adaptively adjust the features of the content according to the user intention during the denoising process of the preset diffusion model, realize the guidance of the generated content, avoid simple random generation, and achieve fine-grained control of the generated content. On the other hand, by fully integrating the user intention into the processing process of the preset diffusion model, it can also make the generated content highly consistent with the user intention. The user intention issued by the user can directly affect the generation result, improve the interactivity of content generation, and is beneficial to enhancing the user's interactive experience during content generation. On the other hand, during the process of controlling content generation based on the user intention, the processing process of the preset diffusion model is carried out in the latent space, which can reduce the computational complexity and improve the content generation efficiency.

[0139] In the exemplary embodiments of the present disclosure, a content generation device is also provided. Refer to Figure 8 As shown, the content generation device 800 may include an information acquisition module 810, a feature processing module 820, a generation control module 830, and a content generation module 840. Specifically:

[0140] The information acquisition module 810 is configured to acquire user intention encoding information and process the noise latent variable at the current moment through a preset diffusion model to extract the latent variable feature map output by the target intermediate layer of the preset diffusion model; the feature processing module 820 is configured to perform conditional modulation processing on the latent variable feature map according to the user intention encoding information to obtain a modulated latent variable feature map, and predict a gating weight map based on the user intention encoding information, where the gating weight map is used to indicate the degree of attention to different feature regions; the generation control module 830 is configured to fuse the modulated latent variable feature map and the gating weight map to obtain a fused feature map, and use the fused feature map as the input of the subsequent network layer of the target intermediate layer to perform denoising processing on the noise latent variable to obtain a denoised noise latent variable; the content generation module 840 is configured to generate target content according to the denoised noise latent variable.

[0141] In an exemplary embodiment of the present disclosure, the information acquisition module 810 is configured to execute: receive action instruction information, and extract object category information and object attribute information from the action instruction information; respectively perform encoding processing on the object category information and the object attribute information to obtain a plurality of encoding results, and fuse the obtained plurality of encoding results to obtain user intention encoding information.

[0142] In an exemplary embodiment of the present disclosure, an action perception gating network is embedded in at least one target intermediate layer of a preset diffusion model. The action perception gating network includes a conditional modulation network and a gating network; the feature processing module 820 is configured to perform: based on the user intention encoding information, use the conditional modulation network to perform an affine transformation on the latent variable feature map to obtain a modulated latent variable feature map; perform feature selection processing according to the user intention encoding information through the gating network to generate a gating weight map.

[0143] In an exemplary embodiment of the present disclosure, the feature processing module 820 is configured to perform: based on the user intention encoding information, use the first multi-layer perceptron network in the conditional modulation network to predict a scaling factor, and the scaling factor is used to adjust the amplitude of the latent variable feature map in different dimensions; through the second multi-layer perceptron network in the conditional modulation network, predict a bias term according to the user intention encoding information, and the bias term is used to adjust the value distribution of the latent variable feature map; perform an affine transformation on the latent variable feature map according to the scaling factor and the bias term to obtain a modulated latent variable feature map.

[0144] In an exemplary embodiment of the present disclosure, the action perception gating network further includes a feature fusion network; the generation control module 830 is configured to perform: multiply the modulated latent variable feature and the gating weight map element-wise through the feature fusion network to obtain a fused feature map.

[0145] In an exemplary embodiment of the present disclosure, the generation control module 830 is configured to perform: fuse the fused feature map with the original features of the subsequent network layer to obtain target features; perform noise prediction based on the noise latent variable, the user intention encoding information, the time step information, and the target features to obtain the noise at the current moment, and perform denoising processing on the noise latent variable according to the noise at the current moment to obtain a denoised noise latent variable.

[0146] In an exemplary embodiment of the present disclosure, the content generation module 840 is configured to perform: if the current denoising iteration number is reached, perform latent variable decoding on the denoised noise latent variable to obtain the target content; if the current denoising iteration number is not reached, determine the denoised noise latent variable as the noise latent variable at the next moment and perform the denoising process.

[0147] Since the detailed content of each functional module of the content generation device in the exemplary embodiment of the present disclosure has been described in the exemplary embodiment of the above content generation method, it will not be repeated here.

[0148] It should be noted that although several modules or units of the content generation device are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0149] The exemplary embodiments of the present disclosure also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the above-described content generation method.

[0150] In one embodiment, the computer program product may be a tangible product containing a computer program, such as a computer-readable storage medium storing the computer program. The readable storage medium may be a storage medium based on signals such as electricity, magnetism, light, electromagnetic, infrared, etc., including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory (Flash), hard disk drive (HDD), solid state drive (SSD), and so on. Exemplarily, the computer program product may be implemented as a non-volatile storage medium storing the computer program, such as read-only memory, NAND flash memory, etc.

[0151] In one embodiment, the computer program product may be an intangible product containing a computer program. Exemplarily, the computer program product may be implemented as a virtual digital product, such as an executable file storing the computer program, an installation package, and other digital files.

[0152] The code of the computer program can be written in one or more programming languages. The program code can be executed entirely on the user's computing device, or partially on the user's computing device, or executed as an independent software package, or partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, such as a local area network (LAN), a wide area network (WAN), etc., or can be connected to an external computing device (for example, through an Internet connection provided by an operator).

[0153] The computer program can be carried or transmitted by signals such as electricity, magnetism, light, electromagnetic, infrared, etc. The electronic device can convert the signal carrying the computer program into a digital signal and then run the computer program. When the computer program runs on the electronic device, its code is used to cause the electronic device to execute (more specifically, can cause the processor of the electronic device to execute) the method steps of various exemplary embodiments of the present disclosure, such as can execute the steps of the above-described content generation method, for example:

[0154] Obtain the user intention encoding information, and process the noise latent variable at the current moment through a preset diffusion model to extract the latent variable feature map output by the target intermediate layer of the preset diffusion model; perform conditional modulation processing on the latent variable feature map according to the user intention encoding information to obtain the modulated latent variable feature map, and predict a gating weight map based on the user intention encoding information, where the gating weight map is used to indicate the degree of attention to different feature regions; fuse the modulated latent variable feature map and the gating weight map to obtain a fused feature map, and use the fused feature map as the input of the subsequent network layer of the target intermediate layer to perform denoising processing on the noise latent variable to obtain the denoised noise latent variable; generate the target content according to the denoised noise latent variable.

[0155] In addition, in the exemplary embodiments of the present disclosure, an electronic device capable of implementing the above method is also provided. Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here.

[0156] The following refers to Figure 9 to describe the electronic device 900 according to this embodiment of the present disclosure. Figure 9 The electronic device 900 shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0157] As Figure 9 shown, the electronic device 900 is presented in the form of a general-purpose computing device. The components of the electronic device 900 may include, but are not limited to: at least one of the above-mentioned processing units 910, at least one of the above-mentioned storage units 920, a bus 930 connecting different system components (including the storage unit 920 and the processing unit 910), and a display unit 940.

[0158] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 910, so that the processing unit 910 executes the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification, for example:

[0159] Obtain the user intention encoding information, and process the noise latent variable at the current moment through a preset diffusion model to extract the latent variable feature map output by the target intermediate layer of the preset diffusion model; perform conditional modulation processing on the latent variable feature map according to the user intention encoding information to obtain the modulated latent variable feature map, and predict a gating weight map based on the user intention encoding information, where the gating weight map is used to indicate the degree of attention to different feature regions; fuse the modulated latent variable feature map and the gating weight map to obtain a fused feature map, and use the fused feature map as the input of the subsequent network layer of the target intermediate layer to denoise the noise latent variable to obtain the denoised noise latent variable; generate the target content according to the denoised noise latent variable.

[0160] The storage unit 920 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 921 and / or a cache storage unit 922, and may further include a read-only storage unit (ROM) 923.

[0161] The storage unit 920 may also include a program / utilities 924 having a set (at least one) of program modules 925. Such program modules 925 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0162] The bus 930 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of a variety of bus structures.

[0163] The electronic device 900 may also communicate with one or more external devices 1000 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 900, and / or may communicate with any device that enables the electronic device 900 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be through the input / output (I / O) interface 950. And, the electronic device 900 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 960. As shown in the figure, the network adapter 960 communicates with other modules of the electronic device 900 through the bus 930. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0164] From the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0165] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed, for example, synchronously or asynchronously in multiple modules.

[0166] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

Claims

1. A content generation method, characterized in that: include: Obtaining user intention encoding information, and processing the noise latent variable at the current moment through a preset diffusion model to extract a latent variable feature map output by a target intermediate layer of the preset diffusion model; Conditionally modulating the latent variable feature map according to the user intention encoding information to obtain a modulated latent variable feature map, and predicting a gating weight map based on the user intention encoding information, wherein the gating weight map is used to indicate the degree of attention paid to different feature areas; Fusing the modulated latent variable feature map and the gated weight map to obtain a fused feature map, and using the fused feature map as an input of a subsequent network layer of the target intermediate layer, denoising the noise latent variable, and obtaining a denoised noise latent variable; Generate target content according to the denoised noise latent variable.

2. The method according to claim 1, characterized in that The obtaining of user intention encoding information includes: receiving action instruction information, and extracting object category information and object attribute information from the action instruction information; The object category information and the object attribute information are respectively encoded to obtain a plurality of encoding results, and the plurality of encoding results are merged to obtain the user intended encoding information.

3. The method according to claim 1, characterized in that Embedding a motion-aware gating network in at least one target intermediate layer of the preset diffusion model, wherein the motion-aware gating network includes a conditional modulation network and a gating network; The conditionally modulating the latent variable feature map according to the user intention encoding information to obtain the modulated latent variable feature map, and predicting the gating weight map based on the user intention encoding information includes: Based on the user intention encoding information, using the conditional modulation network to perform an affine transformation on the latent variable feature map to obtain the modulated latent variable feature map; The gating network performs feature selection processing according to the user intention encoding information to generate the gating weight map.

4. The method according to claim 3, characterized in that The method of performing an affine transformation on the latent variable feature map based on the user intention encoding information by using the conditional modulation network to obtain the modulated latent variable feature map includes: Based on the user intention encoding information, using a first multi-layer perceptron network in the conditional modulation network to predict a scaling factor, wherein the scaling factor is used to adjust the amplitude of the latent variable feature map in different dimensions; encoding information according to the user intention to predict a bias term through a second multi-layer perceptron network in the conditional modulation network, wherein the bias term is used to adjust the value distribution of the latent variable feature map; According to the scaling factor and the bias term, an affine transformation is performed on the latent variable feature map to obtain the modulated latent variable feature map.

5. The method according to claim 3, characterized in that: The motion-aware gating network further includes a feature fusion network; the step of fusing the modulated latent variable feature map and the gating weight map includes: The modulated latent variable features and the gated weight map are element-wise multiplied by the feature fusion network to obtain the fused feature map.

6. The method according to claim 1, characterized in that The step of using the fused feature map as an input of a subsequent network layer of the target intermediate layer, performing denoising on the noise latent variable, and obtaining the denoised noise latent variable includes: Fusing the fused feature map with the original features of the subsequent network layer to obtain target features; Noise prediction is performed according to the noise latent variable, the user intention encoding information, the time step information and the target feature to obtain the noise at the current moment, and denoising is performed on the noise latent variable according to the noise at the current moment to obtain the denoised noise latent variable.

7. The method according to any one of claims 1 to 6, characterized in that: The generating target content according to the noise latent variable after denoising includes: If the current denoising iteration number is reached, performing latent variable decoding on the denoised noise latent variable to obtain the target content; If the number of denoising iterations has not been reached currently, the denoised noise latent variable is determined as the noise latent variable at the next moment and the denoising process is performed.

8. A content generating device, characterized in that: include: An information acquisition module, used to acquire the user intention encoding information, and process the noise latent variable at the current moment through a preset diffusion model to extract the latent variable feature map output by the target intermediate layer of the preset diffusion model; A feature processing module, configured to perform conditional modulation processing on the latent variable feature map according to the user intention encoding information to obtain a modulated latent variable feature map, and to predict a gating weight map based on the user intention encoding information, wherein the gating weight map is used to indicate the degree of attention paid to different feature areas; A generation control module is used to fuse the modulated latent variable feature map and the gated weight map to obtain a fused feature map, and use the fused feature map as an input of a subsequent network layer of the target intermediate layer, denoise the noise latent variable, and obtain a denoised noise latent variable; A content generation module is used to generate target content according to the noise latent variable after denoising.

9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 7 by executing the executable instructions.

Citation Information

Cited By

  • Graph-graph conversion method and device for fine tuning by using noise control module and diffusion model principle

    CN121073754A

  • Method, device, medium and product for forming intelligent data set with body

    CN121212193A

  • Body-aware data aggregation method, device, medium and product

    CN121212193B

  • Image restoration method and related equipment

    CN121563838A