Content generation method, content generation model training method and corresponding device

Through global attention processing and masked attention maps, the problem of insufficient temporal and spatial modeling capabilities of existing models is solved, the quality of media content generation and computational efficiency are improved, and the efficient generation of multimodal input data is supported.

CN120434485BActive Publication Date: 2025-09-23BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510941607.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-09-23
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

Existing content generation models lack temporal and spatial modeling capabilities, resulting in poor quality of generated media content. In particular, blurring and artifacts are prone to occur during complex motions and scene changes, making it difficult to maintain temporal continuity and consistency.

Method used

Global attention processing and masked attention maps are used to generate target media content by encoding multiple input data and unifying the spatiotemporal dimension indexes, avoiding the influence between different types of time series data frames and interference between data from the same media channel.

Benefits of technology

It improves the quality of content generation, enhances spatiotemporal modeling capabilities, reduces computing and storage costs, supports the generation of multiple media content, and has stronger controllability and generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120434485B_ABST
    Figure CN120434485B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose a content generation method, a training method for a content generation model, and a corresponding device. The content generation method includes: obtaining multiple input data, wherein the multiple input data include at least two types of time series data; encoding the multiple input data to obtain feature representations of the multiple input data; using a masked attention map, performing global attention processing on the feature representations of the multiple input data to obtain an intermediate representation of the multiple input data; the value of each element in the attention map represents the degree of association between corresponding positions in the multiple input data, wherein elements corresponding to different frames between different types of time series data are masked, and elements corresponding to input data that do not have the same media channel are masked; based on the intermediate representation of the input data, target media content is generated. The present application can effectively improve the generation quality of media content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a content generation method, a content generation model training method, and corresponding devices. Background Art

[0002] Media content generation, such as image and video generation, is a key research area in computer vision and artificial intelligence. Driven by deep learning, diffusion models, and the Transformer architecture, it has experienced rapid progress in recent years. Current technological developments are focused on controllable content generation, for example, by introducing control conditions such as images, audio, and motion to guide video generation. However, these control conditions often involve temporal features, such as audio, video, and motion vectors. This requires content generation models to possess not only spatial modeling capabilities but also temporal modeling capabilities, thereby improving the quality of media content generation. Summary of the Invention

[0003] In view of this, the present application provides a content generation method, a content generation model training method and corresponding devices to improve the quality of media content generation.

[0004] This application provides the following solutions:

[0005] In a first aspect, a content generation method is provided, the method comprising:

[0006] Acquire a plurality of input data, wherein the plurality of input data includes at least two types of time series data;

[0007] Performing encoding processing on the plurality of input data to obtain feature representations of the plurality of input data;

[0008] Using a masked attention map, global attention processing is performed on the feature representations of the multiple input data to obtain an intermediate representation of the multiple input data; the value of each element in the attention map represents the degree of association between corresponding positions in the multiple input data, wherein elements corresponding to different frames between different types of time series data are masked, and elements corresponding to input data that do not have the same media channel are masked;

[0009] Target media content is generated based on the intermediate representation of the input data.

[0010] According to an achievable method in an embodiment of the present application, encoding the plurality of input data to obtain feature representations of the plurality of input data includes:

[0011] Performing lemma processing on the plurality of input data respectively to obtain lemma sequences of the plurality of input data;

[0012] Indexing the word-gram sequences of the plurality of input data in a unified spatiotemporal dimension to obtain a positional encoding representation;

[0013] Based on the positional encoding representation of the word-gram sequence of the plurality of input data, feature representations of the plurality of input data are obtained.

[0014] According to an achievable method in an embodiment of the present application, indexing the word-gram sequences of the plurality of input data in a unified spatiotemporal dimension includes at least one of the following:

[0015] If the multiple input data include non-time series data, copy the word sequence of the non-time series data in the time dimension so that the length of the copied non-time series data is consistent with that of the time series data in the time dimension, and index the time series data and non-time series data in the time dimension;

[0016] If the spatial dimensions of the at least two types of time series data are inconsistent, copying the position index of each word in the word sequence of the time series data with a lower spatial dimension, so as to make the spatial dimensions of the at least two types of time series data consistent;

[0017] If the at least two types of time series data correspond to different modalities, the position indexes of the word sequences of the time series data of different modalities in the spatial dimension are based on different index ranges.

[0018] According to an implementable manner in an embodiment of the present application, the multiple input data include control conditions and a noise image, and the control conditions include at least two types of time series data; generating target media content based on the intermediate representation of the input data includes: denoising the noise image based on the intermediate representation of the input data to obtain a target image; or

[0019] The multiple input data include control conditions and noisy videos, and the control conditions include at least one type of time series data; generating target media content based on the intermediate representation of the input data includes: denoising the noisy image based on the intermediate representation of the input data to obtain the target video; or

[0020] The multiple input data include control conditions and noise audio, and the control conditions include at least one type of time series data; generating target media content based on the intermediate representation of the input data includes: denoising the noise audio based on the intermediate representation of the input data to obtain target audio; or

[0021] The multiple input data include control conditions, and the control conditions include at least two types of time series data; generating target media content based on the intermediate representation of the input data includes: performing regression prediction based on the intermediate representation of the input data to generate target text.

[0022] In a second aspect, a method for training a content generation model is provided, the method comprising:

[0023] Acquire first training data including a plurality of first training samples, wherein the first training samples include a first input data sample and a target media sample, the first input data sample includes a plurality of first input data, and the plurality of first input data includes at least two types of time series data;

[0024] A first content generation model is obtained by training using the first training data; wherein the first content generation model encodes multiple first input data included in the first input data sample to obtain feature representations of the multiple first input data; a masked attention map is used to perform global attention processing on the feature representations of the multiple first input data to obtain intermediate representations of the multiple first input data; the value of each element in the masked attention map represents the degree of association between corresponding positions in the multiple first input data, wherein elements corresponding to different frames between different types of time series data are masked, and elements corresponding to first input data that do not have the same media channel are masked; target media content is generated based on the intermediate representation of the first input data;

[0025] The training objective includes minimizing the difference between the target media content and the target media sample.

[0026] According to an achievable method in an embodiment of the present application, encoding the plurality of first input data included in the first input data sample to obtain feature representations of the plurality of first input data includes:

[0027] Performing lemma processing on the plurality of first input data respectively to obtain lemma sequences of the plurality of first input data;

[0028] Indexing the word-gram sequences of the plurality of first input data in a unified spatiotemporal dimension to obtain positional encoding representations;

[0029] Based on the positional encoding representations of the word-gram sequences of the plurality of first input data, feature representations of the plurality of first input data are obtained.

[0030] According to an achievable method in an embodiment of the present application, indexing the word-gram sequences of the plurality of first input data in a unified spatiotemporal dimension includes at least one of the following:

[0031] If the plurality of first input data include non-time series data, copying the word sequence of the non-time series data in the time dimension so that the length of the copied non-time series data is consistent with that of the time series data in the time dimension, and indexing the time series data and the non-time series data in the time dimension;

[0032] If the spatial dimensions of the at least two types of time series data are inconsistent, copying the position index of each word in the word sequence of the time series data with a lower spatial dimension, so as to make the spatial dimensions of the at least two types of time series data consistent;

[0033] If the at least two types of time series data correspond to different modalities, the position indexes of the word sequences of the time series data of different modalities in the spatial dimension are based on different index ranges.

[0034] According to an achievable manner in an embodiment of the present application, the plurality of first input data include control conditions and noise media, where the noise media is obtained by performing noise processing on the target media sample;

[0035] Generating target media content based on the intermediate representation of the first input data includes: performing denoising on the noisy media based on the intermediate representation of the input data to obtain target media content;

[0036] determining a loss function using the noise predicted in the denoising process and the noise added in the denoising process, and updating model parameters of the first content generation model using the loss function;

[0037] The target media sample and the target media content are: image, video or video.

[0038] According to an achievable method in an embodiment of the present application, the target media sample is a target text sample; the plurality of first input data include control conditions, and the control conditions include at least two types of time series data;

[0039] Generating target media content based on the intermediate representation of the first input data includes: performing regression prediction based on the intermediate representation of the input data to generate target text;

[0040] A loss function is determined using the target text and the target text sample, and the model parameters of the first content generation model are updated using the loss function.

[0041] According to an achievable manner in an embodiment of the present application, before using the first training data to train a first content generation model, the method further includes:

[0042] A second content generation model is trained using second training data comprising a plurality of second training samples, wherein the second training samples comprise second input data samples and target media samples, and the second input data samples comprise only non-time series data as a control condition;

[0043] fine-tuning the second content generation model using third training data comprising a plurality of third training samples to obtain a third content generation model, wherein the third training samples include third input data samples and target media samples, the third input data samples only including time series data as a control condition, and the third input data samples include fewer data types than the first input data samples;

[0044] Using the first training data to train a first content generation model includes: using the first training data to fine-tune the third content generation model to obtain the first content generation model.

[0045] According to a third aspect, a content generation device is provided, the device comprising:

[0046] a data acquisition unit, configured to acquire a plurality of input data, wherein the plurality of input data includes at least two types of time series data;

[0047] an encoding processing unit, configured to perform encoding processing on the plurality of input data to obtain feature representations of the plurality of input data;

[0048] an attention processing unit configured to perform global attention processing on the feature representations of the plurality of input data using a masked attention map to obtain an intermediate representation of the plurality of input data; the value of each element in the attention map represents the degree of association between corresponding positions in the plurality of input data, wherein elements corresponding to different frames between different types of time series data are masked, and elements corresponding to input data not having the same media channel are masked;

[0049] The content generation unit generates target media content based on the intermediate representation of the input data.

[0050] In a fourth aspect, a training device for a content generation model is provided, the device comprising:

[0051] a sample acquisition unit configured to acquire first training data including a plurality of first training samples, wherein the first training samples include a first input data sample and a target media sample, wherein the first input data sample includes a plurality of first input data, and the plurality of first input data includes at least two types of time series data;

[0052] A model training unit is configured to use the first training data to train a first content generation model; wherein the first content generation model encodes multiple first input data included in the first input data sample to obtain feature representations of the multiple first input data; uses a masked attention map to perform global attention processing on the feature representations of the multiple first input data to obtain intermediate representations of the multiple first input data; the value of each element in the masked attention map represents the degree of association between corresponding positions in the multiple first input data, wherein elements corresponding to different frames between different types of time series data are masked, and elements corresponding to first input data that do not have the same media channel are masked; based on the intermediate representation of the first input data, target media content is generated; the goal of the training includes minimizing the difference between the target media content and the target media sample.

[0053] In a fifth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the steps of the method described in the first aspect or the second aspect are implemented.

[0054] According to a sixth aspect, an electronic device is provided, including:

[0055] one or more processors; and

[0056] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in the first aspect or the second aspect above.

[0057] In a seventh aspect, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the method described in the first or second aspect above.

[0058] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0059] 1) This application uses global attention processing to obtain an intermediate representation of multiple input data. Since global attention processing calculates the similarity between any two spatiotemporal position features in multiple input data and directly models spatiotemporal correlations, it has stronger spatiotemporal modeling capabilities than cross-attention processing that models correlations between features from different sources, and can effectively improve the quality of content generation. In addition, to avoid signal conflicts between multiple time series data in global attention processing, this application uses masked attention maps to avoid the influence between different frames of different types of time series data, as well as the influence between input data that do not share the same media channel, thereby ensuring strict synchronization and correspondence between time series signals, as well as accurate interaction between data from the same media channel, thereby improving the quality of media content generation.

[0060] 2) This application positionally encodes the word-unit sequences of multiple input data, integrates the multiple input data in the sequence dimension, and performs global attention processing without adding a large number of additional parameters, thereby reducing computing and storage costs and improving computing efficiency.

[0061] 3) This application can replicate the word sequence of non-time series data in the time dimension, thereby making the length of the non-time series data and the time series data consistent in the time dimension; replicate the position index of each word in the time series data with lower spatial dimensions in the time series data with inconsistent spatial dimensions in the spatial dimension, thereby making the time series data consistent in spatial dimensions; and base the position index of the word sequence of time series data of different modalities in the spatial dimension on different index ranges, thereby distinguishing the position encoding representation of different modalities. These methods can index the word sequence of multiple input data in a unified spatiotemporal dimension, thereby eliminating the representation gap between modalities and improving the end-to-end model's ability to capture complex spatiotemporal features.

[0062] 4) The content generation method and model training method provided in this application can be applied to the generation of various types of media content, such as video, audio, image, text, etc., and supports full-modal input data, making the model more controllable.

[0063] 5) This application introduces a three-stage training approach to training the content generation model. First, the initial content generation model (the second content generation model) is trained using only non-time series data as a control condition. Then, the second content generation model is further optimized using only time series data as a control condition to obtain the third content generation model. Finally, the third content generation model is further optimized using input data containing more types to obtain the final content generation model (the first content generation model). This approach gradually improves the model's spatiotemporal understanding and content generation capabilities, avoiding poor training results due to complex tasks from the outset and improving the model's generalization ability.

[0064] Of course, any product implementing the present application does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0066] Figure 1 2 is a system architecture diagram applicable to the embodiments of the present application.

[0067] Figure 2 A flowchart of a content generation method provided in an embodiment of the present application.

[0068] Figure 3a A schematic diagram of generating a target video provided in an embodiment of the present application.

[0069] Figure 3b A schematic diagram of generating a target image provided in an embodiment of the present application.

[0070] Figure 3c A schematic diagram of generating target text provided in an embodiment of the present application.

[0071] Figure 4 A schematic diagram of indexing video and audio in the spatial dimension provided in an embodiment of the present application.

[0072] Figure 5 Another schematic diagram of indexing video and audio in the spatial dimension provided in an embodiment of the present application.

[0073] Figure 6 A schematic diagram of a masked attention map provided in an embodiment of the present application.

[0074] Figure 7 A structural diagram of the video generation model provided in an embodiment of the present application.

[0075] Figure 8 A structural diagram of a global attention-based generative network provided in an embodiment of the present application.

[0076] Figure 9 A flowchart of a training method for a content generation model provided in an embodiment of the present application.

[0077] Figure 10 A flowchart of another training method for a content generation model provided in an embodiment of the present application.

[0078] Figure 11 A schematic block diagram of a content generation device provided in an embodiment of the present application.

[0079] Figure 12 A schematic block diagram of a training device for a content generation model provided in an embodiment of the present application.

[0080] Figure 13 A schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0081] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0082] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "an", "the" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.

[0083] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.

[0084] The word "if," as used herein, may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.

[0085] Figure 1 This is a system architecture diagram applicable to the embodiments of the present application, such as Figure 1 As shown in , the system architecture may include: a user device, a content generation device located on the server side, and a device for training a content generation model.

[0086] The user equipment and the server can communicate with each other. The user equipment and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in this application.

[0087] User devices include, but are not limited to, smart mobile terminals, smart home devices, wearable devices, and personal computers (PCs). Smart mobile devices include mobile phones, tablets, laptops, PDAs (Personal Digital Assistants), and internet-connected cars. Smart home devices include smart TVs and smart refrigerators. Wearable devices include smart watches, smart glasses, virtual reality devices, augmented reality devices, and mixed reality devices (i.e., devices that support both virtual reality and augmented reality).

[0088] A server can be a standalone server, a server cluster, or even a cloud server. A cloud server, also known as a cloud computing server or cloud host, is a hosting product within the cloud computing service ecosystem. It addresses the management difficulties and limited scalability of traditional physical hosting and virtual private server (VPS) services.

[0089] Before performing a content generation task, the device for training a content generation model can adopt the method provided in the embodiment of the present application to train the content generation model.

[0090] Users can input data through their user devices, which then transmit the input data over the network to a content generation device on the server side. The content generation device uses a trained content generation model to generate target media content and returns this target media content to the user device over the network. The user device then displays the received target media content to the user.

[0091] Apart from Figure 1 In addition to the illustrated architecture, a computer terminal device with strong computing capabilities may also use the method provided in the embodiments of the present application to train a content generation model and / or generate target media content.

[0092] It should be understood that Figure 1 The number of user devices, apparatuses for training content generation models, content generation apparatuses, and content generation models in the embodiment is merely illustrative. Any number of user devices, apparatuses for training content generation models, content generation apparatuses, and content generation models may be provided as needed.

[0093] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0094] Currently, the vast majority of content generation models are based on diffusion models. These models fuse temporal features extracted from audio and video data using a cross-attention mechanism, based on a traditional convolutional U-network. However, this approach is weak in both temporal and spatial modeling. As a result, the generated images and videos often suffer from quality issues such as blurring and artifacts. Complex motion and scene changes can lead to unnatural deformation and distortion, making it difficult to maintain temporal continuity and consistency.

[0095] With the emergence of the Transformer architecture, diffusion models are using Transformer networks to replace U-shaped networks based on convolutional neural networks. However, traditional diffusion models based on Transformer networks often introduce adapters to process temporal features. These adapters typically use a cross-attention mechanism. However, the cross-attention mechanism's temporal and spatial modeling capabilities still need improvement. Furthermore, the introduction of these adapters adds a large number of parameters to the base model, which affects computational performance.

[0096] In view of this, this application provides a new idea. Figure 2 This is a flowchart of a content generation method provided in an embodiment of the present application. The method can be Figure 1 The content generation device in the system shown in FIG. Figure 2 As shown in , the method may include the following steps:

[0097] Step 201: Acquire multiple input data, where the multiple input data include at least two types of time series data.

[0098] Step 203: Encode the multiple input data to obtain feature representations of the multiple input data.

[0099] Step 205: Using a masked attention map, perform global attention processing on the feature representations of multiple input data to obtain intermediate representations of the multiple input data; the value of each element in the attention map represents the degree of correlation between corresponding positions in the multiple input data, wherein elements corresponding to different frames between different types of time series data are masked, and elements corresponding to input data that do not have the same media channel are masked.

[0100] Step 207: Generate target media content based on the intermediate representation of the input data.

[0101] It can be seen from the above process that the present application uses global attention processing to obtain an intermediate representation of multiple input data. Since global attention processing calculates the similarity of any two spatiotemporal position features and directly models spatiotemporal correlations, it has stronger spatiotemporal modeling capabilities than cross-attention processing that models the correlations of different source features, and can effectively improve the quality of content generation. In addition, in order to avoid signal conflicts between multiple time series data in global attention processing, the present application uses a masked attention map to avoid the influence between different frames of different types of time series data, and to avoid the influence between input data that do not have the same media channel, thereby ensuring strict synchronization and correspondence between time series signals, and accurate interaction between data from the same media channel, thereby improving the quality of content generation.

[0102] The following describes in detail the steps in the above process and the effects that can be further produced in conjunction with the embodiments.

[0103] First, the above step 201, namely "obtaining multiple input data, where the multiple input data include at least two types of time series data", is described in detail with reference to an embodiment.

[0104] The method provided in the embodiment of the present application can be applied to a variety of content generation tasks, such as video generation tasks, image generation tasks, audio generation tasks, text generation tasks, etc. For different generation tasks, the corresponding input data may be different.

[0105] Taking the video generation task as an example, the above-mentioned multiple input data may include control conditions and noise videos, wherein the control conditions are usually content input or selected by the user. In an embodiment of the present application, the control conditions include at least one type of time series data, and may further include non-time series data. The so-called time series data refers to data with a time dimension, that is, a series of data recorded in chronological order, such as video, motion vector, camera trajectory or audio. Non-time series data refers to data that does not have a time dimension, such as images, text, etc.

[0106] like Figure 3a As shown in the figure, the control conditions include a reference video, motion vectors, camera trajectory, audio, images, and text. In addition to the control conditions, the input data also includes a noisy video. The subsequent content generation model denoises the noisy video to generate the target video. Motion vectors typically indicate the direction of motion of each token between adjacent frames. The camera trajectory records the camera's motion patterns over time, helping the model generate the corresponding target video based on the camera trajectory.

[0107] Taking the audio generation task as an example, the multiple input data mentioned above can include control conditions and noise audio. In the embodiment of the present application, the control conditions include at least one type of time series data and can further include non-time series data. For example, the control conditions can include reference video, images, and text, and also include noise audio. This situation is similar to video generation.

[0108] Taking the image generation task as an example, the multiple input data may include control conditions and noise images. In the embodiment of the present application, the control conditions include at least two types of time series data, and may further include non-time series data.

[0109] by Figure 3b As shown in Figure 2, the control conditions are reference video, audio, and text. The input data includes a noisy image in addition to the control conditions. The subsequent content generation model denoises the noisy image to obtain the target image.

[0110] Taking the text generation task as an example, the plurality of input data mentioned above only include control conditions. In the embodiment of the present application, the control conditions include at least two types of time series data, and may further include non-time series data.

[0111] by Figure 3c As shown in Figure 2, the control conditions are reference video, audio, and text. The subsequent content generation model generates the target text through regression prediction.

[0112] The above step 203, i.e., "encoding the multiple input data to obtain feature representations of the multiple input data", is described in detail below with reference to an embodiment.

[0113] In an embodiment of the present application, multiple input data can be first lemmatized to obtain lemma sequences of the multiple input data; the lemma sequences of the multiple input data can be indexed on a unified time and space dimension to obtain a position coding representation; based on the position coding representation of the lemma sequences of the multiple input data, a feature representation of the multiple input data can be obtained.

[0114] When tokenizing multiple input data points, you can use a corresponding type of Tokenizer. The essence of a Tokenizer is to map raw data to discrete semantic units, converting input data into a sequence of discrete tokens (lemmas) that the model can process. Tokens are discrete semantic units and the foundation of model building.

[0115] For example, a text tokenizer can be used to lemmatize text, generating tokens corresponding to the text. Each token typically corresponds to a basic semantic unit in natural language, typically represented by a word and some special markers (such as classification markers and separators). The text tokenizer can be implemented using pre-trained language models, such as BERT (Bidirectional Encoder Representation from Transformers), XLNet (an autoregressive model that implements bidirectional contextual information by permuting a language model), the GPT (Generative Pre-Training) model, and CLIP (an encoding model that enables multimodal encoding).

[0116] For images, you can use an image tokenizer to tokenize them and generate tokens corresponding to the image. Each token typically corresponds to a pixel block in the image. Image tokenizers can use models such as VAE (Variational Autoencoder), VIT (Vision Transformer), and CLIP.

[0117] For videos, a video tokenizer can be used to perform tokenization and generate corresponding tokens. Each token is essentially a pixel block in each video frame, with the frame as the time dimension. Video tokenizers can use methods such as VAE and VideoViT.

[0118] After obtaining the tokens corresponding to multiple input data, the input data can be converted into sequences and then concatenated. For time series data, the tokens of each frame can be arranged in sequence to convert it into a sequence. This method allows the tokens corresponding to multiple input data to be integrated in the sequence dimension, providing a foundation for subsequent global attention processing.

[0119] In addition, in order to help the model accurately understand the spatiotemporal relationship of the input data, the word units of multiple input sequences can be further encoded based on position. In the embodiment of the present application, since time series data and non-time series data are involved, it is necessary to index on a unified spatiotemporal dimension to obtain a positional encoding representation.

[0120] The unified spatiotemporal dimension used in the embodiments of the present application can be the maximum spatiotemporal dimension of the input data. Assuming that the input data includes video, image, text, etc., where the reference video has the largest spatiotemporal dimension, that is, three dimensions (number of time frames, spatial length, and spatial width), then the spatiotemporal dimension is used as the spatiotemporal dimension used for position coding. Taking video as an example, a three-segment index can be used as the position code, where the first segment index represents the number of time frames, the second segment index represents the spatial length, and the third segment index represents the spatial width. For example, "0,1,1" represents the index of the 2nd row and 2nd column Token in the 1st frame.

[0121] Since there is no time dimension for non-time series data, such as images and text, in an embodiment of the present application, non-time series data can be copied in the time dimension so that the length of the copied non-time series data in the time dimension is consistent with that of the time series data, and the time series data and non-time series data can be indexed in the time dimension. Taking images as an example, assuming that the input data has n frames of video, the image is copied n-1 frames so that the image also has n frames in the time dimension, and then indexed according to the above three-segment method. There may also be a situation where there are multiple non-time series data in the input data, and the lengths of these multiple non-time series data are inconsistent. For example, if the video has n frames and the audio has m frames, then for the image in the input data, the image can be copied n-1 frames so that the image has the same length in the time dimension as the video, or the image can be copied m-1 frames so that the image has the same length in the time dimension as the audio. In this way, the synchronous correspondence problem of time series signals and non-time series signals can be solved, and the temporal and spatial correlation relationship between time series data and non-time series data can be fully considered.

[0122] Since time series data all have time dimensions, their spatial dimensions may be inconsistent. For example, video has two spatial dimensions, while audio has only one. Therefore, the position index of the time series data with lower spatial dimensions can be copied to make the spatial dimensions of all time series data consistent. Taking video and audio as an example, assuming that the index of video in the spatial dimension can be as follows Figure 4 As shown in (a), the audio can copy the index of only one dimension in the spatial dimension so that the indexes of the last two segments are equal, as shown in Figure 4 As shown in (b), it is equivalent to the diagonal index in the two-dimensional state.

[0123] Furthermore, for time series data, there may be time series data corresponding to different modalities in the input data. For example, audio and video are two completely different modalities. In order to allow the model to better understand their differences in spatial dimensions, as one of the more preferred methods, the index ranges of time series data of different modalities in space can be staggered, so that time series data of different modalities can be indexed based on different index ranges in spatial dimensions. Figure 5 As shown in , videos are indexed in the index range of (0~99) and audios are indexed in the index range of (100~199).

[0124] It should be noted here that in addition to the positional encoding representation, the word units of multiple input data can also be encoded based on content, and then the content encoding representation and the positional encoding representation obtained by the content-based encoding can be fused (for example, spliced) to obtain the feature representation of multiple input data.

[0125] The above step 205, i.e., "using a masked attention map to perform global attention processing on feature representations of multiple input data to obtain intermediate representations of the multiple input data," is described in detail below in conjunction with an embodiment.

[0126] In the embodiment of the present application, the Adapter method, which is often used in scenarios with multiple input data, is no longer used. This method requires the additional setting of an Adapter, which performs cross-attention processing on the feature representations of each input data. Instead, the embodiment of the present application directly uses a global attention mechanism in the content generation model. The so-called global attention mechanism is a core mechanism in deep learning. Its core idea is to enable the model to simultaneously focus on all positions in the entire input sequence when processing input data, rather than being limited to local areas. This mechanism calculates the correlation between each position in the input sequence and all other positions and dynamically assigns weights to capture the global context.

[0127] In the embodiment of the present application, the tokens of multiple input data are spliced ​​into an input sequence. Then the feature representation of the multiple input data (i.e., the input sequence) is essentially a feature matrix, assuming it is represented by X, where each row corresponds to the feature vector of a token.

[0128] Among them, self-attention is a typical method of global attention. Taking self-attention as an example, we can use X to obtain query (query vector), key (key vector) and value (value vector), which are expressed as 、 and , for example, the following formula can be used:

[0129] = X (1)

[0130] = X (2)

[0131] = X (3)

[0132] in, 、 and is the weight matrix, which is the parameter to be learned by the model.

[0133] The attention mechanism can be described as the process of mapping a query and a series of key-value pairs to an output. This output is the sum of the weights calculated based on the query and key and applied to the value, expressed as:

[0134] (4)

[0135] in, represents the attention processing process, for spatial dimension. Refers to the normalization process performed by the softmax function.

[0136] As can be seen, global attention associates the query vector representing a particular spatiotemporal location feature with the key and value vectors representing all spatiotemporal location features in a spatiotemporal manner. Its core purpose is to model spatiotemporal associations. Cross-attention, on the other hand, associates the query vector representing a particular feature source with the key and value vectors representing another feature source. Its core purpose is to model associations between features from different sources. Therefore, global attention has better spatiotemporal modeling capabilities.

[0137] Because the input data includes at least two types of time-series data, the presence of this time-series data makes it difficult for the content generation model to converge. This is because time-series data such as audio and video need to be synchronized, and directly using global attention processing will cause conflicts. In view of this, an embodiment of the present application provides a solution that uses a masked Attention Map to perform global attention processing and obtain an intermediate representation of multiple input data.

[0138] The Attention Map is essentially a weight distribution matrix dynamically calculated based on the query and key. The value of each element represents the degree of association between corresponding positions in the input sequence (i.e., multiple input data). The Attention Map can be expressed as follows:

[0139] (5)

[0140] In an embodiment of the present application, the elements of different frames corresponding to different types of time series data are masked to avoid the influence between different frames of different types of time series data. For example, it is assumed that the input data includes images, audio and noisy videos, where the images and audio are used as control conditions to denoise the noisy videos to obtain the target video. Among them, the audio and noisy videos are time series data. The noisy video is assumed to be expressed as F×M1, the image is expressed as K×M1, and the audio is expressed as F×M2, where F represents the number of frames of the video, M1 represents the number of video tokens in space, and M2 represents the number of audio tokens in space. K is the number of images. The Attention Map can be as follows Figure 6 As shown in , the left side of the Attention Map represents the query, and the top side represents the key. The white square in the figure indicates that the corresponding element is masked, and the gray part indicates that the corresponding element is not masked. Figure 6 As shown in , a frame of noisy image includes 4 tokens, and a frame of audio includes 3 tokens. This masking method preserves the association between the noisy video and itself, as well as the association between the noisy video and the entire image. That is, the model can learn and understand the association between the noisy video and the entire image, but only retains the noisy video and the audio of the corresponding frame. This ensures strict synchronization and correspondence between the time series data, thereby improving the quality of content generation.

[0141] In the Attention Map of the embodiment of the present application, the corresponding elements between the input data that do not have the same media channel are masked. For the input data, there may be a situation where the media channels are completely different, where the media channel refers to an independent signal component in the multimedia data. For example, video has an audio channel and an image channel, audio only has an audio channel, and image only has an image channel. Therefore, if Figure 6 As shown in , for audio, the association between the audio and itself, as well as the audio and the corresponding noisy video frame, is preserved. However, since the audio and image do not share any common media channels, the corresponding elements between the audio and image are masked. This masking method can avoid the influence of input data that do not share the same media channels, thereby ensuring accurate interaction between data with the same media channels, thereby improving the quality of content generation.

[0142] The above step 207 , ie “generating target media content based on the intermediate representation of the input data”, is described in detail below with reference to an embodiment.

[0143] In the embodiments of the present application, different generation mechanisms can be used depending on the type of content generation task. For example, a diffusion model can be used for video generation tasks, audio generation tasks, and image generation tasks. Based on the intermediate representation of the input data, a noisy video is denoised to obtain a target video. Based on the intermediate representation of the input data, a noisy audio is denoised to obtain a target audio. Based on the intermediate representation of the input data, a noisy image is denoised to obtain a target image.

[0144] For another example, for text generation tasks, a regression model can be used to perform regression prediction based on the intermediate representation of the input data to obtain the target video.

[0145] Take the video generation task as an example, Figure 7 As shown in , the content generation model is actually a video generation model, which can include P global attention-based generation networks and decoding networks, where P is a positive integer. Each global attention-based generation network is as follows Figure 8 As shown in , the above-mentioned global attention processing is performed on the feature representations of multiple input data using a masked attention map to obtain intermediate representations of multiple input data, including an intermediate representation of the control condition and an intermediate representation of the noisy video. The intermediate representation of the control condition is output to the next global attention-based generative network. After predicting the noise features based on the intermediate representations of the multiple input data, the predicted noise features are removed from the intermediate representation of the noisy video to obtain an intermediate representation of a new noise image, which is output to the next global attention-based generative network. For the last global attention-based generative network, the intermediate representation of the noise image is output to the decoding network. The decoding network uses the intermediate representation of the input noisy image to decode and obtain the target video. Learning accurate spatiotemporal correlations between different inputs through masked global attention can improve the clarity, temporal consistency, and detail expression of video generation.

[0146] As can be seen, each of the aforementioned global attention-based generative networks is essentially a diffusion model based on a masked global attention mechanism. Each global attention-based generative network can perform denoising for T time steps, where T is a preset positive integer. Setting a larger time step yields higher quality for the generated target video, but this consumes more resources and increases inference time. Therefore, a balance is typically struck between the two.

[0147] Figure 9 This is a flowchart of a method for training a content generation model provided in an embodiment of the present application. The method can be performed by Figure 1 The system shown in FIG. Figure 9 As shown in , the method may include the following steps:

[0148] Step 901: Acquire first training data including multiple first training samples, where the first training samples include first input data samples and target media samples, and the first input data samples include multiple first input data, where the multiple first input data include at least two types of time series data.

[0149] In an embodiment of the present application, the target media sample may be a video sample, an image sample, an audio sample, or a text sample. The first input data sample includes at least a plurality of first input data, each of which includes at least a control condition, wherein the control condition may be manually annotated for the target media sample. If the target media sample is a video sample, an image sample, or an audio sample, the first input data may also include a noise sample, such as a noisy video, a noisy image, or a noisy audio. The noise sample may be obtained by adding noise to the target media sample (i.e., adding noise for T time steps).

[0150] Step 903: A first content generation model is obtained by training using the first training data; wherein the first content generation model encodes the multiple first input data included in the first input data sample to obtain feature representations of the multiple first input data; a masked attention map is used to perform global attention processing on the feature representations of the multiple first input data to obtain intermediate representations of the multiple first input data; the value of each element in the masked attention map represents the degree of association between corresponding positions in the multiple first input data, wherein elements corresponding to different frames between different types of time series data are masked, and elements corresponding to first input data that do not exist in the same media channel are masked; target media content is generated based on the intermediate representation of the first input data; the training objectives include minimizing the difference between the target media content and the target media sample.

[0151] In this step, when the first content generation model encodes the multiple first input data included in the first input data sample, it can perform lemma processing on the multiple first input data respectively to obtain multiple tokens of the first input data; index the tokens of the multiple first input data on a unified time and space dimension to obtain position coding representation; based on the position coding representation of the tokens of the multiple first input data, obtain feature representations of the multiple first input data.

[0152] Similar to the previous content generation method embodiment, if multiple first input data include non-time series data, the non-time series data is copied in the time dimension so that the length of the copied non-time series data is consistent with the time series data in the time dimension, and the time series data and non-time series data are indexed in the time dimension.

[0153] If the spatial dimensions of at least two types of time series data are inconsistent, the position index of the time series data with a lower spatial dimension in the spatial dimension is copied to make the spatial dimensions of at least two types of time series data consistent.

[0154] If at least two types of time series data correspond to different modalities, the position indexes of the time series data of different modalities in the spatial dimension are based on different index ranges.

[0155] The specific details of the above processing of the first content generation model can be found in the relevant records in the content generation method embodiment, and will not be repeated here.

[0156] The first content generation model can be an image generation model, a video generation model, an audio generation model, or a text generation model. If the first content generation model is an image generation model, a video generation model, or an audio generation model, the first content generation model can be implemented based on a diffusion model, for example Figure 6 As shown in , a global attention-based diffusion network can be used for implementation. In this case, the multiple first input data may include control conditions and noise media, where the noise media is obtained by adding noise to the target media sample. The first content generation model denoises the noise media based on the intermediate representation of the input data to obtain the target media content.

[0157] Taking the video generation model as an example, the first input data sample may include multiple first input data and target video samples. The multiple first input data may include control conditions and noise videos. The control conditions include at least one type of time series data, such as reference video or audio. The noise video is obtained by adding noise to the target video sample. This is done by gradually adding Gaussian noise through a forward diffusion process. Assume that the noise video is obtained by adding noise for T time steps, and the noise added at each time step is recorded.

[0158] The training process of the video generation model is a reverse denoising process. The video generation model (corresponding to the first content generation model mentioned above) denoises the noisy video based on the intermediate feature representation of the input data. The denoising process includes noise prediction for T time steps. If the difference between the target video obtained by the final denoising process and the target video sample is to be minimized, the loss function can be determined using the noise predicted during the denoising process and the noise added during the denoising process. For example, the loss function can be used to ensure that the predicted noise and the added noise are consistent through the L2 distance constraint, and the model parameters of the video generation model (i.e., the first content generation model) are updated using the loss function. In other words, if the noise predicted by the first content generation model is as consistent as possible with the noise added during the forward diffusion process, then the target video obtained by the denoising process can be as consistent as possible with the target video sample.

[0159] If the first content generation model is a text generation model, it can be implemented based on a regression model. In this case, the first input data samples may include multiple first input data and target text samples. The multiple first input data only include control conditions, which include at least two types of time series data, such as reference video and reference audio. The first content generation model performs regression prediction based on the intermediate representation of the input data to generate the target text.

[0160] In this case, the generated target text and target text samples can be used to determine a loss function, such as a mean square error loss, a mean absolute error loss, or other types of loss functions; and the loss function can be used to update the model parameters of the first content generation model.

[0161] When using the loss function to update the model parameters of the first content generation model, the model parameters may be updated using methods such as, but not limited to, gradient descent, until a preset training termination condition is met. The training termination condition may include, for example, the loss function value being less than or equal to a preset loss function threshold, the number of iterations reaching a preset threshold, etc.

[0162] In the embodiment of this application, it is possible to directly use Figure 9 To improve the efficiency and performance of model training, you can use Figure 10 The three-stage training method shown in Figure 10 As shown in , the following steps may be included:

[0163] Step 1001: A second content generation model is trained using second training data including a plurality of second training samples, wherein the second training samples include second input data samples and target media samples, and the second input data samples only include non-time series data as a control condition.

[0164] In the first stage, only non-time series data is used for model training. For example, non-time series data such as text and images are used as control conditions. Combined with target video samples, the initial video generation model is trained as the above-mentioned second content generation model.

[0165] When the control conditions do not include non-time series data, it is relatively easy to obtain samples and perform model training. Therefore, a basic content generation model can be pre-trained as the second content generation model.

[0166] Step 1003: Fine-tune the second content generation model using third training data comprising a plurality of third training samples to obtain a third content generation model, wherein the third training samples comprise third input data samples and target media samples, the third input data samples comprise only time series data, and the third input data samples comprise fewer data types than the first input data samples.

[0167] In the second stage, the second content generation model is fine-tuned using a third input data sample consisting only of time series data. For example, time series data such as reference video, audio, and motion vectors are used as control conditions. Combined with the target video sample, the second content generation model trained in the first stage is fine-tuned, and the resulting video generation model is used as the third content generation model.

[0168] Step 1005: Use the first training data including multiple first training samples to fine-tune the third content generation model to obtain the first content generation model, the first training sample includes a first input data sample and a target media sample, the first input data sample includes multiple first input data, and the multiple first input data include at least two types of time series data.

[0169] In the third stage, non-time series signals are gradually added on the basis of the second stage, that is, the third input data sample used in the second stage contains less data types than the first input data sample, that is, the third stage gradually changes to data of more modal types, including at least two types of time series data, and may also include non-time series data. In this process, the type ratio and data volume of non-time series data and time series data can be adjusted. For example, in the third stage, reference video, audio, motion vector, image, text and other time series data and non-time series data are used to fine-tune the video generation model obtained in the second stage, so as to obtain the final video generation model, namely the third content generation model.

[0170] It should be noted that the training methods and loss functions of the above three stages can all be used Figure 9 The training method and loss function described in the illustrated embodiment are not described in detail here.

[0171] The above-mentioned method provided in the embodiment of the present application can realize controllable media content generation and can be applied to content creation (news production, film and television production, animation production, etc.), education and training (for example, generating teaching courseware, case analysis, etc.), artistic creation (for example, generating conversation, music, dance and other works of art), and data visualization (for example, generating visual charts, images, etc. from text, video, etc.).

[0172] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0173] According to another embodiment, a content generation device is provided. Figure 11 As shown, the device 1100 includes: a data acquisition unit 1101, an encoding processing unit 1102, an attention processing unit 1103 and a content generation unit 1104. The main functions of each component unit are as follows:

[0174] The data acquisition unit 1101 is configured to acquire a plurality of input data, where the plurality of input data includes at least two types of time series data.

[0175] The encoding processing unit 1102 is configured to perform encoding processing on the multiple input data to obtain feature representations of the multiple input data.

[0176] The attention processing unit 1103 is configured to use a masked attention map to perform global attention processing on the feature representations of the multiple input data to obtain an intermediate representation of the multiple input data; the value of each element in the attention map represents the degree of association between corresponding positions in the multiple input data, wherein elements corresponding to different frames between different types of time series data are masked, and corresponding elements between input data that do not have the same media channel are masked.

[0177] The content generation unit 1104 generates target media content based on the intermediate representation of the input data.

[0178] As one of the feasible ways, the encoding processing unit 1102 can be specifically configured as follows: performing tokenization processing on the multiple input data respectively to obtain tokens of the multiple input data; indexing the tokens of the multiple input data on a unified spatiotemporal dimension to obtain positional encoding representations; and obtaining feature representations of the multiple input data based on the positional encoding representations of the tokens of the multiple input data.

[0179] Specifically, when the encoding processing unit 1102 indexes the word-grams of the plurality of input data in a unified spatiotemporal dimension, it may be specifically configured as follows:

[0180] If the multiple input data include non-time series data, copy the non-time series data in the time dimension so that the length of the copied non-time series data is consistent with that of the time series data in the time dimension, and index the time series data and the non-time series data in the time dimension;

[0181] If the spatial dimensions of the at least two types of time series data are inconsistent, copying the position index of the time series data with a lower spatial dimension in the spatial dimension to make the spatial dimensions of the at least two types of time series data consistent;

[0182] If the at least two types of time series data correspond to different modalities, the position indexes of the time series data of different modalities in the spatial dimension are based on different index ranges.

[0183] As a first feasible method, the above-mentioned multiple input data include control conditions and noise images, and the control conditions include at least two types of time series data; the content generation unit 1104 can be configured to: denoise the noise image based on the intermediate representation of the input data to obtain a target image.

[0184] As a second possible implementation, the plurality of input data includes a control condition and a noisy video, wherein the control condition includes at least one type of time series data. The content generation unit 1104 may be configured to: perform denoising on the noisy image based on an intermediate representation of the input data to obtain a target video.

[0185] As a third possible implementation, the plurality of input data includes control conditions and noisy audio, wherein the control conditions include at least one type of time series data. The content generation unit 1104 may be configured to: perform denoising on the noisy audio based on an intermediate representation of the input data to obtain the target audio.

[0186] As a fourth possible implementation, the plurality of input data include control conditions, wherein the control conditions include at least two types of time series data. The content generation unit 1104 may be configured to perform regression prediction based on the intermediate representation of the input data to generate the target text.

[0187] Figure 12 A schematic block diagram of a training device for a content generation model provided in an embodiment of the present application, such as Figure 12 As shown in FIG, the apparatus 1200 may include: a sample acquisition unit 1201 and a model training unit 1202. The main functions of each component unit are as follows:

[0188] The sample acquisition unit 1201 is configured to acquire first training data including multiple first training samples, where the first training samples include first input data samples and target media samples, where the first input data samples include multiple first input data, and where the multiple first input data include at least two types of time series data.

[0189] The model training unit 1202 is configured to use the first training data to train and obtain a first content generation model; wherein, the first content generation model encodes the multiple first input data included in the first input data sample to obtain feature representations of the multiple first input data; uses a masked attention map to perform global attention processing on the feature representations of the multiple first input data to obtain an intermediate representation of the multiple first input data; the value of each element in the masked attention map represents the degree of association between corresponding positions in the multiple first input data, wherein elements corresponding to different frames between different types of time series data are masked, and elements corresponding to first input data that do not have the same media channel are masked; based on the intermediate representation of the first input data, target media content is generated; the goal of the training includes minimizing the difference between the target media content and the target media sample.

[0190] As one possible implementation, the multiple first input data include control conditions and noisy media, where the noisy media is obtained by adding noise to the target media sample. The first content generation model denoises the noisy media based on the intermediate representation of the first input data to obtain target media content. In this approach, the model training unit 1202 uses the noise predicted during the denoising process and the noise added during the noisy process to determine a loss function, and uses the loss function to update the model parameters of the first content generation model; wherein the target media sample and the target media content are: images, videos, or videos.

[0191] In another possible implementation, the target media sample is a target text sample; the plurality of first input data includes control conditions, which include at least two types of time series data. The first content generation model performs regression prediction based on the intermediate representation of the first input data to generate the target text. In this implementation, the model training unit 1202 uses the target text and the target text sample to determine a loss function, and uses the loss function to update the model parameters of the first content generation model.

[0192] Furthermore, before using the first training data to train the first content generation model, the model training unit 1202 is also configured to: use second training data including multiple second training samples to train the second content generation model, the second training samples include second input data samples and target media samples, and the second input data samples only include non-time series data as a control condition; use third training data including multiple third training samples to fine-tune the second content generation model to obtain a third content generation model, the third training samples include third input data samples and target media samples, the third input data samples only include time series data as a control condition, and the third input data samples contain less data types than the first input data samples.

[0193] In this case, when the model training unit 1202 uses the first training data to train the first content generation model, it is specifically configured to: use the first training data to fine-tune the third content generation model to obtain the first content generation model.

[0194] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiment. The device embodiment described above is merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0195] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.

[0196] And an electronic device comprising:

[0197] one or more processors; and

[0198] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the aforementioned method embodiments.

[0199] The present application also provides a computer program product, comprising a computer program, which implements the steps of any one of the methods described in the aforementioned method embodiments when executed by a processor.

[0200] in, Figure 13 The electronic device architecture is shown as an example, and may include a processor 1310, a video display adapter 1311, a disk drive 1312, an input / output interface 1313, a network interface 1314, and a memory 1320. The processor 1310, the video display adapter 1311, the disk drive 1312, the input / output interface 1313, the network interface 1314, and the memory 1320 may be communicatively connected via a communication bus 1330.

[0201] The processor 1310 may be implemented as a general-purpose CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and may be used to execute relevant programs to implement the technical solutions provided in this application.

[0202] The memory 1320 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1320 can store an operating system 1321 for controlling the operation of the electronic device 1300, and a basic input and output system (BIOS) 1322 for controlling the low-level operations of the electronic device 1300. In addition, a web browser 1323, a data storage management system 1324, and a content generation device 1100 / a device for training a content generation model 1200, etc. can also be stored. The above-mentioned content generation device 1100 / a device for training a content generation model 1200 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided by the present application is implemented by software or firmware, the relevant program code is stored in the memory 1320 and is called and executed by the processor 1310.

[0203] The input / output interface 1313 is used to connect to input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown) or externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, and various sensors. Output devices may include a display, speaker, vibrator, indicator light, etc.

[0204] The network interface 1314 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, Wi-Fi, Bluetooth, etc.).

[0205] The bus 1330 comprises a pathway for transmitting information between the various components of the device (eg, the processor 1310 , the video display adapter 1311 , the disk drive 1312 , the input / output interface 1313 , the network interface 1314 , and the memory 1320 ).

[0206] It should be noted that although the above device only shows a processor 1310, a video display adapter 1311, a disk drive 1312, an input / output interface 1313, a network interface 1314, a memory 1320, a bus 1330, etc., in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the solution of the present application, and does not necessarily include all the components shown in the figure.

[0207] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus the necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer program product. The computer program product can be stored in a storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.

[0208] The above is a detailed introduction to the technical solutions provided by this application. Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the contents of this specification should not be understood as limiting this application.

Claims

1. A content generation method, characterized in that: The method comprises: Acquire a plurality of input data, wherein the plurality of input data includes at least two types of time series data; Performing encoding processing on the plurality of input data to obtain feature representations of the plurality of input data; Using a masked attention map, global attention processing is performed on the feature representations of the multiple input data to obtain an intermediate representation of the multiple input data; the value of each element in the attention map represents the degree of correlation between corresponding positions in the multiple input data, wherein elements corresponding to different frames between different types of time series data are masked, and if two input data do not have the same media channel, then the corresponding elements between the two input data are masked; Target media content is generated based on the intermediate representation of the input data.

2. The method according to claim 1, characterized in that Encoding the plurality of input data to obtain feature representations of the plurality of input data includes: Performing lemma processing on the plurality of input data respectively to obtain lemma sequences of the plurality of input data; Indexing the word-gram sequences of the plurality of input data in a unified spatiotemporal dimension to obtain a positional encoding representation; Based on the positional encoding representation of the word-gram sequence of the plurality of input data, feature representations of the plurality of input data are obtained.

3. The method according to claim 2, characterized in that Indexing the word-gram sequences of the plurality of input data in a unified spatiotemporal dimension includes at least one of the following: If the multiple input data include non-time series data, copy the word sequence of the non-time series data in the time dimension so that the length of the copied non-time series data is consistent with that of the time series data in the time dimension, and index the time series data and non-time series data in the time dimension; If the spatial dimensions of the at least two types of time series data are inconsistent, copying the position index of each word in the word sequence of the time series data with a lower spatial dimension, so as to make the spatial dimensions of the at least two types of time series data consistent; If the at least two types of time series data correspond to different modalities, the position indexes of the word sequences of the time series data of different modalities in the spatial dimension are based on different index ranges.

4. The method according to any one of claims 1 to 3, characterized in that The plurality of input data includes a control condition and a noise image, wherein the control condition includes at least two types of time series data; Generating target media content based on the intermediate representation of the input data includes: performing denoising on the noisy image based on the intermediate representation of the input data to obtain a target image; or The multiple input data include control conditions and noisy videos, and the control conditions include at least one type of time series data; generating target media content based on the intermediate representation of the input data includes: denoising the noisy image based on the intermediate representation of the input data to obtain the target video; or The multiple input data include control conditions and noise audio, and the control conditions include at least one type of time series data; generating target media content based on the intermediate representation of the input data includes: denoising the noise audio based on the intermediate representation of the input data to obtain target audio; or The multiple input data include control conditions, and the control conditions include at least two types of time series data; generating target media content based on the intermediate representation of the input data includes: performing regression prediction based on the intermediate representation of the input data to generate target text.

5. A training method for a content generation model, characterized in that: The method comprises: Acquire first training data including a plurality of first training samples, wherein the first training samples include a first input data sample and a target media sample, the first input data sample includes a plurality of first input data, and the plurality of first input data includes at least two types of time series data; A first content generation model is obtained by training using the first training data; wherein the first content generation model encodes multiple first input data included in the first input data sample to obtain feature representations of the multiple first input data; a masked attention map is used to perform global attention processing on the feature representations of the multiple first input data to obtain intermediate representations of the multiple first input data; the value of each element in the masked attention map represents the degree of association between corresponding positions in the multiple first input data, wherein elements corresponding to different frames between different types of time series data are masked, and if there is no common media channel between two first input data, then the corresponding elements between the two first input data are masked; based on the intermediate representation of the first input data, target media content is generated; The training objective includes minimizing the difference between the target media content and the target media sample.

6. The method according to claim 5, characterized in that The encoding process is performed on the plurality of first input data included in the first input data sample to obtain feature representations of the plurality of first input data, including: Performing lemma processing on the plurality of first input data respectively to obtain lemma sequences of the plurality of first input data; Indexing the word-gram sequences of the plurality of first input data in a unified spatiotemporal dimension to obtain positional encoding representations; Based on the positional encoding representations of the word-gram sequences of the plurality of first input data, feature representations of the plurality of first input data are obtained.

7. The method according to claim 6, characterized in that Indexing the word-gram sequences of the plurality of first input data in a unified spatiotemporal dimension includes at least one of the following: If the plurality of first input data include non-time series data, copying the word sequence of the non-time series data in the time dimension so that the length of the copied non-time series data is consistent with that of the time series data in the time dimension, and indexing the time series data and the non-time series data in the time dimension; If the spatial dimensions of the at least two types of time series data are inconsistent, copying the position index of each word in the word sequence of the time series data with a lower spatial dimension, so as to make the spatial dimensions of the at least two types of time series data consistent; If the at least two types of time series data correspond to different modalities, the position indexes of the word sequences of the time series data of different modalities in the spatial dimension are based on different index ranges.

8. The method according to any one of claims 5 to 7, characterized in that The plurality of first input data include control conditions and noise media, wherein the noise media is obtained by performing noise processing on the target media sample; Generating target media content based on the intermediate representation of the first input data includes: performing denoising on the noisy media based on the intermediate representation of the first input data to obtain target media content; determining a loss function using the noise predicted in the denoising process and the noise added in the denoising process, and updating model parameters of the first content generation model using the loss function; The target media sample and the target media content are: image, video or video.

9. The method according to any one of claims 5 to 7, characterized in that The target media sample is a target text sample; the plurality of first input data include control conditions, and the control conditions include at least two types of time series data; Generating target media content based on the intermediate representation of the first input data includes: performing regression prediction based on the intermediate representation of the first input data to generate target text; A loss function is determined using the target text and the target text sample, and the model parameters of the first content generation model are updated using the loss function.

10. The method according to any one of claims 5 to 7, before using the first training data to train a first content generation model, the method further comprises: A second content generation model is trained using second training data comprising a plurality of second training samples, wherein the second training samples comprise second input data samples and target media samples, and the second input data samples comprise only non-time series data as a control condition; fine-tuning the second content generation model using third training data comprising a plurality of third training samples to obtain a third content generation model, wherein the third training samples include third input data samples and target media samples, the third input data samples only including time series data as a control condition, and the third input data samples include fewer data types than the first input data samples; Using the first training data to train a first content generation model includes: using the first training data to fine-tune the third content generation model to obtain the first content generation model.

11. A content generating device, characterized in that: The device comprises: a data acquisition unit, configured to acquire a plurality of input data, wherein the plurality of input data includes at least two types of time series data; an encoding processing unit, configured to perform encoding processing on the plurality of input data to obtain feature representations of the plurality of input data; an attention processing unit configured to perform global attention processing on the feature representations of the plurality of input data using a masked attention map to obtain an intermediate representation of the plurality of input data; the value of each element in the attention map represents the degree of association between corresponding positions in the plurality of input data, wherein elements corresponding to different frames between different types of time series data are masked, and if two input data do not have the same media channel, then the corresponding elements between the two input data are masked; The content generation unit generates target media content based on the intermediate representation of the input data.

12. A training device for a content generation model, characterized in that: The device comprises: a sample acquisition unit configured to acquire first training data including a plurality of first training samples, wherein the first training samples include a first input data sample and a target media sample, wherein the first input data sample includes a plurality of first input data, and the plurality of first input data includes at least two types of time series data; A model training unit is configured to use the first training data to train a first content generation model; wherein the first content generation model encodes multiple first input data included in the first input data sample to obtain feature representations of the multiple first input data; uses a masked attention map to perform global attention processing on the feature representations of the multiple first input data to obtain intermediate representations of the multiple first input data; the value of each element in the masked attention map represents the degree of association between corresponding positions in the multiple first input data, wherein elements corresponding to different frames between different types of time series data are masked, and if there is no same media channel between two first input data, the corresponding elements between the two first input data are masked; based on the intermediate representation of the first input data, target media content is generated; the goal of the training includes minimizing the difference between the target media content and the target media sample.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

14. An electronic device, characterized in that: include: one or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, perform the steps of the method according to any one of claims 1 to 10.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Video generation method and device and storage medium

    CN110162667A

  • Data processing method and device, equipment and medium

    CN116246213A