Content generation model training method and device, equipment, medium and program product
By training a T2V model to enhance the recognition of the characteristics of different game characters, the problem of lack of differentiation in animated GIFs or videos of skin effects of different game characters is solved, and high-quality dynamic content generation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2024-11-21
- Publication Date
- 2026-05-22
AI Technical Summary
Existing T2V models lack differentiation and have low quality when generating animated GIFs or videos of different game character skin effects, and cannot effectively distinguish the visual presentation of different game characters.
By acquiring sample dynamic content and descriptive text of at least two different objects, the sample content generation model is trained to enhance the model's recognition of the characteristics of different objects, learn the unique visual presentation of objects, and generate dynamic content with distinctiveness.
It improves the relevance and accuracy of the generated dynamic content, enhances the quality of the dynamic content, and enables the model to generate distinctive dynamic content based on differences in object characteristics.
Smart Images

Figure CN122072846A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a training method, apparatus, device, medium, and program product for a content generation model. Background Technology
[0002] The T2V (Text-to-Video) model is a model that can transform text descriptions into video content. It generates video content that matches the text description by understanding elements such as scenes, actions, and characters in the text.
[0003] In related technologies, T2V models can be applied to the gaming industry. Designers input a descriptive text about the skin effects of game characters into the T2V model, and the T2V model will generate animated GIFs or videos about the skin effects of game characters.
[0004] However, if the same descriptive terms are provided for different game characters (e.g., hero A wields a spear and hero B wields a spear), the skin effect animations or videos generated by the T2V model are similar, lacking differentiation across different game characters and of low quality. Summary of the Invention
[0005] This application provides a training method, apparatus, device, medium, and program product for a content generation model. The technical solution is as follows:
[0006] On the one hand, a method for training a content generation model is provided, the method comprising:
[0007] Obtain a sample content generation model to be trained. The sample content generation model is used to generate dynamic content based on text description. The dynamic content refers to content containing at least two consecutive image frames.
[0008] Obtain at least two first data sets corresponding to each object; wherein, the first data set corresponding to the i-th object includes sample dynamic content and descriptive text containing the i-th object, the sample dynamic content is used to characterize the visual presentation of the i-th object, and the descriptive text is used to describe the visual presentation of the i-th object in text form, where i is a positive integer;
[0009] The sample content generation model is used to analyze the descriptive text corresponding to the at least two objects to obtain the predicted dynamic content corresponding to the at least two objects.
[0010] The sample content generation model is trained based on the predicted dynamic content and sample dynamic content corresponding to the at least two objects respectively to obtain a first content generation model. The first content generation model is used to generate dynamic content with different visual presentation methods based on the object features of different objects according to the text description.
[0011] On the other hand, a training apparatus for a content generation model is provided, the apparatus comprising:
[0012] The acquisition module is used to acquire a sample content generation model to be trained. The sample content generation model is used to generate dynamic content based on text description. The dynamic content refers to the content containing at least two consecutive image frames. The module acquires first data sets corresponding to at least two objects respectively. The first data set corresponding to the i-th object includes sample dynamic content containing the i-th object and descriptive text. The sample dynamic content is used to characterize the visual presentation of the i-th object, and the descriptive text is used to describe the visual presentation of the i-th object in text form. i is a positive integer.
[0013] The analysis module is used to analyze the descriptive text corresponding to the at least two objects respectively through the sample content generation model to obtain the predicted dynamic content corresponding to the at least two objects respectively;
[0014] The training module is used to train the sample content generation model based on the predicted dynamic content and sample dynamic content corresponding to the at least two objects respectively, to obtain a first content generation model. The first content generation model is used to generate dynamic content with different visual presentation methods based on the object features of different objects according to the text description.
[0015] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the training method of any of the above-described content generation models.
[0016] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored therein, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement a training method for any of the content generation models described above.
[0017] On the other hand, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the training method for any of the content generation models described above.
[0018] The beneficial effects of the technical solutions provided in this application include at least the following:
[0019] By training the sample content generation model with dynamic content and descriptive text from at least two different objects, the model can continuously strengthen its understanding of the characteristics of different objects during training. It learns the unique features of each object and the relationship between them and their corresponding visual presentation methods. This allows the final content generation model to generate distinctive dynamic content based on the differences in the characteristics of different objects when faced with the same descriptive words, thus improving the targeting and accuracy of the generated dynamic content and ultimately enhancing its quality. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a schematic diagram of a computer system provided in an exemplary embodiment of this application;
[0022] Figure 2 This is a flowchart of a training method for a content generation model provided in an exemplary embodiment of this application;
[0023] Figure 3 This is a flowchart of a training method for a content generation model provided in another exemplary embodiment of this application;
[0024] Figure 4 This is a flowchart of a training method for a content generation model provided in yet another exemplary embodiment of this application;
[0025] Figure 5 This is a flowchart of a training method for a content generation model provided in an exemplary embodiment of this application;
[0026] Figure 6 This is a schematic diagram of a computer system for an application content generation model provided in an exemplary embodiment of this application;
[0027] Figure 7 This is a structural block diagram of a training apparatus for a content generation model provided in an exemplary embodiment of this application;
[0028] Figure 8 This is a structural block diagram of a computer device provided in an exemplary embodiment of this application. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0030] In this application, the terms "first" and "second" are used to distinguish between identical or similar items that have essentially the same function. It should be understood that there is no logical or temporal dependency between "first" and "second", nor is there any limitation on the quantity or execution order.
[0031] It should be noted that this application may display prompt interfaces, pop-ups, or output voice prompts before and during the collection of user data. These prompt interfaces, pop-ups, or voice prompts are used to inform the user that their data is being collected. This ensures that the application only begins the steps for collecting user data after receiving confirmation from the user regarding the prompt interface or pop-up; otherwise (i.e., without user confirmation), the steps for collecting user data end, meaning no user data is collected. In other words, all user data collected in this application is collected with the user's consent and authorization, and the collection, use, and processing of related user data must comply with relevant laws, regulations, and standards.
[0032] First, the terms used in the embodiments of this application will be introduced.
[0033] Game skins: In video games, these are decorative elements used to change the appearance of interactive elements, such as game characters and items.
[0034] In games, different game skins have different effects. These effects are usually showcased through videos or GIFs. These videos or GIFs demonstrate how a specific game skin changes the appearance, color scheme, and decorative details of corresponding interactive elements (such as game characters or items) when applied to them. In some embodiments, the game skin effect videos or GIFs can also place the interactive element wearing the skin in typical in-game scenes for display, such as showing the running, attacking, and skill-casting actions of a game character wearing a specific skin in a combat scenario.
[0035] The T2V (Text-to-Video) model is a model that can transform text descriptions into video content. It generates video content that matches the text description by understanding elements such as scenes, actions, and characters within the text. In related technologies, the T2V model can be applied to the gaming industry. Designers input descriptive text about game character skin effects into the T2V model, which then generates animated GIFs or videos of those skin effects. However, if the same descriptive terms are provided for different game characters (e.g., Hero A wielding a spear and Hero B wielding a spear), the T2V model generates similar animated GIFs or videos of skin effects, lacking differentiation across different game characters and resulting in low quality.
[0036] Based on this, this application provides a training method for a content generation model. The sample content generation model is trained using dynamic content and descriptive text of at least two different objects. During the training process, the model can continuously strengthen its understanding of the characteristics of different objects, learn the unique features of each object and the relationship between them and their corresponding visual presentation methods. This enables the final first content generation model to generate distinctive dynamic content based on the differences in the characteristics of the objects themselves when faced with the same descriptive words but different objects. This improves the targeting and accuracy of the generated dynamic content, thereby enhancing the quality of the generated dynamic content.
[0037] The computer system that implements the training method of the content generation model of this application is described below.
[0038] Figure 1 This is a structural block diagram of a computer system provided in an exemplary embodiment of this application. The computer system can implement a system architecture for a training method of a content generation model. The computer system includes: a terminal 110 and a server 120.
[0039] The terminal 110 and the server 120 communicate through a communication network 130, which can be a wired network or a wireless network, and is not limited here.
[0040] Terminal 110 can be an electronic device such as a mobile phone, tablet computer, vehicle terminal (vehicle system), wearable device, or PC (Personal Computer). A client application for the first application can be installed and run on terminal 110. This first application can be an application for training the content generation model, or an application that provides training functions for the content generation model; this application does not limit the specific application. Furthermore, this application does not limit the form of the first application, including but not limited to apps, mini-programs, etc., installed on terminal 110, and can also be in web page form.
[0041] Server 120 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, a cloud server providing basic cloud computing services, or a node in a blockchain system. Server 120 can be the backend server of the aforementioned first application, used to provide backend services to the clients of the first application.
[0042] Terminal 110 and server 120 can communicate via a network, such as a wired or wireless network.
[0043] The training method for the content generation model provided in this application embodiment can be executed by a computer device, which refers to an electronic device with data computing, processing, and storage capabilities. Figure 1 Taking the implementation environment of the scheme shown as an example, the training method of the content generation model can be executed by the terminal 110 (such as the client of the first application installed and running in the terminal 110 executing the training method of the content generation model), or the training method of the content generation model can be executed by the server 120, or the terminal 110 and the server 120 can interact and cooperate to execute it. This application does not limit this.
[0044] Those skilled in the art will understand that the number of terminals 110 described above can be more or less. For example, there may be only one terminal 110, or there may be dozens or hundreds of terminals 110, or even more. This application does not limit the number or type of terminals 110 in its embodiments.
[0045] The following explanation uses the training method of the content generation model executed by server 120 as an example.
[0046] As an illustration, server 120 will have at least two objects ( Figure 1Given n objects (where n is an integer greater than 1), the first data set corresponding to each object is used as the training dataset. The first content generation model is obtained by fine-tuning the sample content generation model using this training dataset. The sample content generation model can be implemented as a pre-trained T2V model, such as a video generation model based on a diffusion model or a Transformer structure. A pre-trained T2V model is a basic model obtained after initial training with a large amount of data, possessing basic video generation capabilities. Fine-tuning the sample content generation model to obtain the first content generation model means further optimizing the pre-trained T2V model using the training dataset, enabling it to generate high-quality videos with object discrimination.
[0047] The first data set corresponding to the i-th object includes sample dynamic content (as a label) containing the i-th object and descriptive text, where i is a positive integer and i≤n. For example, the first data set corresponding to object 1 includes a skin effect video containing object 1 and descriptive text for the skin effect video. During the fine-tuning of the sample content generation model, for the first data set corresponding to the i-th object, the descriptive text corresponding to the i-th object is input into the sample content generation model. The model analyzes the descriptive text corresponding to the i-th object and outputs predicted dynamic content that matches the description in the descriptive text. Then, based on the difference between the preset dynamic content and the sample dynamic content, the target loss is calculated. The model parameters to be adjusted in the sample content generation model are then adjusted according to the target loss, thus completing one fine-tuning of the sample content generation model.
[0048] The sample content generation model is fine-tuned using the first data group corresponding to at least two objects in the training dataset. This completes one iteration of training for the sample content generation model. After multiple iterations, if the model meets the training requirements (e.g., the loss is less than or equal to the preset loss, the number of iterations reaches the preset number, etc.), the fine-tuning stops. The model obtained at this point is the first content generation model. The first content generation model can generate dynamic content with different visual presentation methods based on the object features of different objects according to the text description.
[0049] In some embodiments, a content generation model (such as a first content generation model) is trained using the method provided in this application. This content generation model can be applied to a second application installed on the terminal 110. The second application is used to generate dynamic content with different visual presentation methods based on the object characteristics of different objects according to the text description using the content generation model. Furthermore, this application does not limit the form of the second application, including but not limited to apps, mini-programs, etc., installed on the terminal 110, and can also be in webpage form.
[0050] Next, the training process of the content generation model provided in this application will be introduced.
[0051] Based on the above introduction, Figure 2 This is a flowchart illustrating a training method for a content generation model provided in an embodiment of this application, which is applied to, for example... Figure 1 The method is illustrated using the server shown as an example, and steps 210 to 240 are as follows.
[0052] Step 210: Obtain the sample content generation model to be trained.
[0053] The sample content generation model is used to generate dynamic content based on text descriptions.
[0054] In some embodiments, the sample content generation model includes a diffusion-based video generation model. The diffusion-based video generation model extracts a text feature representation of the input text, which characterizes the semantic information of the input text. Then, using this semantic representation as conditional information, the model progressively removes noise from noisy sample dynamic content (e.g., noisy video), transforming the noisy sample dynamic content into clear dynamic content that conforms to the text description.
[0055] Optionally, the sample content generation model includes a text encoding layer, a denoising network layer, and a video decoding layer. The text encoding layer is used to extract the text feature representation of the input text; the denoising network layer is used to use the text feature representation as conditional information to denoise the dynamic content with added noise and obtain the denoising result; and the video decoding layer is used to decode the denoising result to obtain the predicted dynamic content.
[0056] In some embodiments, the sample content generation model is a pre-trained model. Optionally, the pre-training method includes at least one of unsupervised pre-training and supervised pre-training. The pre-trained sample content generation model has basic dynamic content generation capabilities, and when faced with new input text, it has a certain ability to generate dynamic content that conforms to the text description.
[0057] Dynamic content refers to content that contains at least two consecutive image frames.
[0058] Optionally, dynamic content includes at least one of video, GIF, animation, etc., and is not limited here. In this embodiment, video is used as an example to illustrate the dynamic content.
[0059] Optionally, the sample content generation model is stored on other devices, and the server retrieves the sample content generation model from those other devices; or, the sample content generation model is stored on the server, and the server retrieves it from local storage.
[0060] Step 220: Obtain the first data group corresponding to at least two objects respectively.
[0061] The first data group corresponding to the i-th object includes the sample dynamic content and descriptive text of the i-th object.
[0062] The object is used to indicate the main subject in the sample dynamic content. Illustratively, the video showcasing the skin effects of game character A demonstrates the character's appearance, visual presentation of skill activations, action effects, and related special effects when equipped with a specific skin. The main subject of the video is game character A.
[0063] Optionally, the number of subjects in the i-th object can be one or more. For example, a skin effect video is used to show the appearance, skill release visual effects, cooperative action effects, and other video content of multiple game characters with related relationships when they have a specific skin. The video subjects include multiple game characters.
[0064] The sample dynamic content is used to characterize the visual presentation of the i-th object, where i is a positive integer.
[0065] Optionally, the sample dynamic content includes at least one of video, GIF, animation, etc.
[0066] The visual presentation method refers to the way in which the i-th object is displayed using visual elements. In other embodiments, the visual presentation method may also be the way in which the i-th object is displayed using sensory elements such as tactile elements and auditory elements that are associated with visual elements. This application does not limit this aspect.
[0067] Indicatively, visual elements include the appearance, actions, position, scene, and style of the i-th object; auditory elements associated with visual elements include sound effects when the i-th object performs a specific action and background sounds of the scene in which the i-th object is located; and tactile elements associated with visual elements include vibration effects when the i-th object performs a specific action. No further limitations are imposed here.
[0068] The descriptive text is used to describe the visual presentation of the i-th object in text form.
[0069] In other words, the descriptive text is used to describe the details of the sample's dynamic content in a textual way. Schematic, the descriptive text is used to describe the appearance features, action effects, position, scene content, style features, sound effects when the i-th object performs a specific action, background sound effects of the scene, vibration effects when the i-th object performs a specific action, etc., without being limited here.
[0070] Optionally, the above at least two objects are at least two different objects.
[0071] Methods for obtaining dynamic content of samples and corresponding descriptive text:
[0072] Taking a sample video as an example of dynamic sample content, the system can obtain a sample video containing the i-th object and its corresponding descriptive text. The descriptive text can be manually written by relevant personnel for the sample video; alternatively, the sample video can be input into a large language model, which automatically generates the corresponding video descriptive text; or, relevant personnel can refine the video content descriptive text output by the large language model to obtain the descriptive text; or, the manually written video descriptive text and the sample video can be input into the large language model, which refines the manually written video descriptive text based on the sample video's content to obtain the descriptive text. This application does not limit this approach.
[0073] Step 230: Analyze the descriptive text corresponding to at least two objects using the sample content generation model to obtain the predicted dynamic content corresponding to at least two objects.
[0074] After obtaining the first data set corresponding to at least two objects, each descriptive text in the first data set is sequentially input into the sample content generation model. The sample content generation model analyzes each descriptive text sequentially to obtain the predicted dynamic content corresponding to each descriptive text. The following explanation uses the analysis of the descriptive text of the i-th object by the sample content generation model as an example:
[0075] Optionally, the sample content generation model is implemented as a video generation model based on a diffusion model. After inputting the descriptive text corresponding to the i-th object into the sample content generation model, the text feature representation of the descriptive text is extracted through a text encoding layer. Then, the text feature representation is used as conditional information. Guided by the conditional information, the sample dynamic content with added random noise is denoised through a denoising network layer to obtain the denoised result. Finally, the denoised result is decoded through a video decoding layer to obtain the predicted dynamic content corresponding to the i-th object.
[0076] In some embodiments, the aforementioned conditional information may include other information besides textual feature representations, such as visual feature information. Optionally, the sample content generation model also includes a visual encoder, which compresses the sample dynamic content into a low-dimensional latent space representation. The low-dimensional latent space representation is an abstract and compressed form of the sample dynamic content, containing its main features and structural information. Using the low-dimensional latent space representation and textual feature representation as conditional information, the denoising network layer, guided by this conditional information, performs denoising processing on the sample dynamic content with added random noise to obtain the denoised result. The combined use of the low-dimensional latent space representation and textual feature representation as conditional information allows the model to better combine semantic information with actual visual representation during the generation process, enhancing the accuracy of the generated predicted dynamic content.
[0077] In the above embodiments, the random noise added to the sample dynamic content refers to random values that conform to a specific probability distribution (such as Gaussian distribution, uniform distribution, etc.). Adding random noise to the sample dynamic content means changing the pixel values of the image frames based on random noise in at least two image frames corresponding to the sample dynamic content.
[0078] Step 240: Train the sample content generation model based on the predicted dynamic content and sample dynamic content corresponding to at least two objects respectively, to obtain the first content generation model.
[0079] Optionally, based on the difference between the predicted dynamic content and the sample dynamic content corresponding to at least two objects, the target loss corresponding to at least two objects is determined; the sample content generation model is trained based on the target loss corresponding to at least two objects to obtain the first content generation model.
[0080] To illustrate, the sample dynamic content corresponding to each object is used as a label, and the target loss for each object is determined based on the difference between the predicted dynamic content corresponding to each object and the sample dynamic content corresponding to each object.
[0081] For the i-th object, after generating the predicted dynamic content corresponding to the i-th object, the sample dynamic content containing the i-th object is used as the label. The difference between the predicted dynamic content and the sample dynamic content is calculated to obtain the target loss. The function used to calculate the target loss includes at least one of the following: mean squared error loss function, cross-entropy loss function, mean absolute error loss function, etc., which is not limited here. After obtaining the target loss, the update gradient is calculated through the backpropagation algorithm. The model parameters in the sample content generation model are updated according to the calculated update gradient, thereby completing one training process of the sample content generation model.
[0082] The sample content generation model is trained sequentially using the target losses corresponding to at least two objects, thus completing one iteration of training. Optionally, training stops when the sample content generation model meets the training conditions; the resulting model is the first content generation model. Meeting the training conditions means that the target loss obtained by the model is less than a preset loss, or the model has undergone a preset number of iterations.
[0083] In some embodiments, the sample content generation model includes a first model parameter and a second model parameter, wherein the first model parameter is a model parameter that does not need to be updated (frozen) when training the sample content generation model, and the second model parameter is a model parameter that needs to be updated (adjusted) when training the sample content generation model.
[0084] Optionally, the first model parameters in the sample content generation model are frozen, and the second model parameters in the sample content generation model are adjusted based on the predicted dynamic content and sample dynamic content corresponding to at least two objects, respectively, to obtain the first content generation model.
[0085] To illustrate, after obtaining the target loss, the update gradient is calculated using the backpropagation algorithm. The values of the second model parameters are then updated based on the calculated update gradient, thus completing one training process for the sample content generation model.
[0086] Optionally, the second model parameters include at least one of the following parameters:
[0087] (1) Hidden layer parameters.
[0088] Hidden layer parameters refer to the parameters of the hidden layer in the sample content generation model.
[0089] Indicatively, the sample content generation model includes an input layer, at least one hidden layer, and an output layer. The input layer is used to receive raw data (e.g., descriptive text), the hidden layer is used to transform and extract features from the raw data, and the output layer is used to generate the final prediction result (e.g., predicting dynamic content).
[0090] Optionally, the parameters of the last hidden layer in the sample content generation model can be used as the parameters of the second model. During the pre-training phase, the sample content generation model learns a large number of general and relatively abstract features through its preceding hidden layers. These general features capture the basic structure and patterns in the data. By focusing on adjusting the parameters of the last hidden layer, the model can be appropriately modified based on the original general features to quickly and accurately adapt to new tasks, such as generating a video of skin effects that match the description of a game character's skin effects from the text.
[0091] (2) Input parameters.
[0092] In a schematic representation, an input parameter is added to the description text, and the input parameter is randomly initialized. Optionally, the input parameter includes at least one of an input prefix parameter, an input suffix parameter, etc., wherein the input prefix parameter refers to a trainable random parameter added at the beginning of the description text, and the input suffix parameter refers to a trainable random parameter added at the end of the description text. It should be noted that the embodiments of this application do not limit the position of the input parameter in the description text; for example, the input parameter can also be added in the middle of the description text.
[0093] Optionally, the input prefix parameters added to the descriptive text can be used as second model parameters. Illustratively, this training scheme of adjusting the input prefix parameters can be called prefix fine-tuning. Prefix fine-tuning adds a set of trainable parameters to the front of the descriptive text and updates these prefix parameters during training, thereby enabling the model to better capture key information and contextual relationships in the text.
[0094] In the above embodiments, when updating parameters for dynamic sample content, it is not necessary to train all parameters in the model; only certain specified model parameters (i.e., the second model parameters) need to be updated, which reduces the consumption of computing resources and improves training efficiency.
[0095] The first content generation model, which has been trained, is used to generate dynamic content with different visual presentation methods based on the object features of different objects according to the text description.
[0096] For example: Obtain the skin effect description text 1 for game character a, input the skin effect description text 1 into the content generation model, and the content generation model will output a skin effect video a that contains the character image of game character a and matches the description of the skin effect description text 1; Obtain the skin effect description text 2 for game character b, input the skin effect description text 2 into the content generation model, and the content generation model will output a skin effect video b that contains the character image of game character b and matches the description of the skin effect description text 2, where different game characters in skin effect video a and skin effect video b have different behavior patterns and action characteristics.
[0097] In summary, the training method for the content generation model provided in this application trains the sample content generation model using dynamic content and descriptive text from at least two different objects. This allows the model to continuously strengthen its understanding of the characteristics of different objects during training, learn the unique features of each object and the relationship between them and their corresponding visual presentation methods. As a result, when the final first content generation model encounters the same descriptive words but different objects, it can generate dynamic content with distinctive features based on the differences in the characteristics of the objects themselves, thereby improving the targeting and accuracy of the generated dynamic content and thus enhancing the quality of the generated dynamic content.
[0098] In some embodiments, to further improve the dynamic content generation quality of the trained first content generation model on a specified object, a two-stage training method can be used to train the sample content generation model. The training dataset used in the first stage of training consists of a first data set corresponding to at least two objects, and the training dataset used in the second stage of training consists of a second data set corresponding to the i-th object. For illustrative purposes, please refer to [reference needed]. Figure 3 The above Figure 2 The illustrated embodiment can also be implemented as follows: steps 301 to 307.
[0099] Step 301: Obtain the sample content generation model to be trained.
[0100] The sample content generation model is used to generate dynamic content based on text descriptions. Dynamic content refers to content that contains at least two consecutive image frames.
[0101] In some embodiments, the sample content generation model is a pre-trained model. Optionally, the pre-training includes at least one of unsupervised pre-training and supervised pre-training. The following example illustrates the implementation of the sample content generation model as a video generation model based on a diffusion model:
[0102] (1) Unsupervised pre-training.
[0103] A large amount of dynamic content is acquired as a training dataset. For a specific piece of dynamic content, noise is added to it to obtain noisy dynamic content. This noisy dynamic content is then input into a sample diffusion model, which denoises the noisy content to obtain predicted dynamic content. The model parameters of the sample diffusion model are updated based on the difference between the predicted dynamic content and the original dynamic content. The above steps are repeated on a large amount of dynamic content until the sample diffusion model meets the training requirements for unsupervised pre-training. Training stops when the accuracy of the predicted dynamic content generated by the sample diffusion model is greater than or equal to a first accuracy. Optionally, the model obtained at the point of stopping training is used as the sample content generation model.
[0104] (2) Supervised pre-training.
[0105] A large number of dynamic content-text description data pairs are acquired as the training dataset, where each data pair includes a dynamic content and a descriptive text, with the descriptive text referring to the text describing the dynamic content. For a given data pair, the sample diffusion model adds noise to the dynamic content, resulting in noisy dynamic content. The sample diffusion model extracts the text feature representation of the descriptive text, and uses this text feature representation as conditional information. Guided by this conditional information, the sample diffusion model denoises the noisy dynamic content to obtain predicted dynamic content. The model parameters of the sample diffusion model are updated based on the difference between the predicted dynamic content and the original dynamic content. The above steps are repeated on a large number of dynamic content-text description data pairs until the sample diffusion model meets the training requirements of supervised pre-training. Training stops when the accuracy of the predicted dynamic content generated by the sample diffusion model is greater than or equal to a second accuracy (the second accuracy is greater than the first accuracy). Optionally, the model obtained at the point of stopping training is used as the sample content generation model.
[0106] Optionally, the sample content generation model includes a text encoding layer, a denoising network layer, and a video decoding layer.
[0107] Step 302: Obtain the first data group corresponding to at least two objects respectively.
[0108] The first data group corresponding to the i-th object includes sample dynamic content and descriptive text containing the i-th object. The sample dynamic content is used to characterize the visual presentation of the i-th object, and the descriptive text is used to describe the visual presentation of the i-th object in text form, where i is a positive integer.
[0109] In some embodiments, for the first data group corresponding to the i-th object, the first data group may include one or more data pairs, each data pair including one sample dynamic content containing the i-th object and one descriptive text corresponding to the sample dynamic content.
[0110] When the first data set contains multiple data pairs, the sample dynamic content corresponding to each data pair is used to represent different visual presentation methods corresponding to the i-th object. For example, for virtual character a, the effect video and corresponding description text of skin 1 of virtual character a, as well as the effect video and corresponding description text of skin 2 of virtual character a, are obtained as data in the first data set.
[0111] The first dataset is defined as the first data group corresponding to at least two objects.
[0112] Step 303: For the first data group corresponding to at least two objects respectively, analyze the descriptive text corresponding to at least two objects respectively through the sample content generation model to obtain the predicted dynamic content corresponding to at least two objects respectively.
[0113] This is illustrated by analyzing each descriptive text in the first dataset using a sample content generation model to obtain the predicted dynamic content corresponding to each descriptive text. The process of analyzing each descriptive text in the first dataset using a sample content generation model to obtain the predicted dynamic content corresponding to each descriptive text can be referred to in step 230, and will not be elaborated here.
[0114] Step 304: Train the sample content generation model based on the predicted dynamic content and sample dynamic content corresponding to at least two objects respectively, and obtain the candidate content generation model.
[0115] Optionally, based on the difference between the predicted dynamic content and the sample dynamic content corresponding to at least two objects, a first loss corresponding to at least two objects is determined; a sample content generation model is trained based on the first loss corresponding to at least two objects to obtain a candidate content generation model.
[0116] To illustrate, the sample dynamic content corresponding to each object is used as a label, and the first loss for each object is determined based on the difference between the predicted dynamic content corresponding to each object and the sample dynamic content corresponding to each object.
[0117] For the i-th object, after generating the predicted dynamic content corresponding to the i-th object, the sample dynamic content containing the i-th object is used as the label. The difference between the predicted dynamic content and the sample dynamic content is calculated to obtain the first loss. The function used to calculate the first loss includes at least one of the following: mean squared error loss function, cross-entropy loss function, mean absolute error loss function, etc., which is not limited here. After obtaining the first loss, the update gradient is calculated through the backpropagation algorithm. The model parameters in the sample content generation model are updated according to the calculated update gradient, thereby completing one training process of the sample content generation model.
[0118] Traverse the first training set and train the sample content generation model based on at least two obtained first losses, which completes one iteration of training on the first training set.
[0119] In the above embodiments, in scenarios involving at least two objects, a first loss corresponding to each object is determined and trained, enabling the model to pay attention to the differences between different objects, thereby more accurately presenting the uniqueness of each role in the generated dynamic content and avoiding confusion or homogenization of the features of different roles.
[0120] In some embodiments, the loss used to train the sample content generation model further includes a second loss, which represents the difference between the action feature representation corresponding to the i-th object and the action feature representations corresponding to other objects. Optionally, after obtaining the predicted dynamic content corresponding to at least two objects, the method for training the sample content generation model further includes the following steps:
[0121] Step 1: Extract action feature representations of the predicted dynamic content for at least two objects.
[0122] Here, the action feature corresponding to the i-th object represents the action performance of the i-th object in the predicted dynamic content corresponding to the i-th object.
[0123] Among these, the expression of movement includes, but is not limited to: the type of movement (e.g., walking, running, jumping, etc.), the range of motion (e.g., the height of a jump, the range of motion of a wave of the hand, etc.), the speed of the movement, and the trajectory of the movement.
[0124] To illustrate, for the i-th object, the predicted dynamic content corresponding to the i-th object is input into the action feature extraction model to extract the action feature representation corresponding to the i-th object. The action feature extraction model is a pre-trained model, which can be implemented as a convolutional neural network, a recurrent neural network, etc.
[0125] In some embodiments, the above-mentioned object refers to a virtual character, and the predicted dynamic content corresponding to at least two objects includes the predicted dynamic content corresponding to the i-th virtual character.
[0126] Optionally, for the predicted dynamic content corresponding to the i-th virtual character, determine the skeletal key points of the character model in each image frame corresponding to the predicted dynamic content, where the character model refers to the three-dimensional model corresponding to the i-th virtual character; and determine the action feature representation based on the skeletal key points of the character model in at least two image frames corresponding to the predicted dynamic content.
[0127] Among them, the skeletal key points include the positions of key points such as the head, shoulders, elbows, wrists, hips, knees, and ankles of the i-th virtual character, which are not limited here.
[0128] This diagram illustrates how analyzing the positional information of skeletal keypoints in each image frame yields motion feature representations. By comparing the positional changes of the same skeletal keypoint in different image frames, the character's motion trajectory and motion type can be preliminarily inferred. Statistical analysis of the displacements in multiple adjacent image frames reveals the speed of the character's motion. The relative positional relationships between different skeletal keypoints and their changes during the motion reflect the amplitude of the character's movements. For example, in a waving motion, the change in distance between shoulder and hand skeletal keypoints reflects the magnitude of the wave.
[0129] In the above embodiments, obtaining action feature representations based on the location information of skeletal key points can improve the accuracy of the obtained action feature representations.
[0130] Step 2: Based on the action feature representations corresponding to at least two objects, determine the corresponding actions of at least two objects. The second loss.
[0131] The second loss corresponding to the i-th object is used to characterize the similarity between the action feature representation of the i-th object and the action feature representation of other objects.
[0132] Optionally, the similarity between the above features can be implemented as at least one of Euclidean distance, Manhattan distance, cosine similarity, etc., without limitation here.
[0133] In some embodiments, the above action feature representation includes multiple sub-feature representations, and different sub-feature representations are associated with different action descriptive words in the descriptive text.
[0134] To illustrate, for each object, its descriptive text may contain a series of action descriptive words, such as "running," "jumping," "waving," etc. The action feature representation contains multiple sub-feature representations, each of which is associated with a specific action descriptive word in the descriptive text. For example, a sub-feature representation may be specifically used to describe features such as speed and amplitude of the "running" action.
[0135] The example will be illustrated using at least two objects, including a first object and a second object. The first object and the second object are distinct objects.
[0136] In some embodiments, during the process of determining the second loss corresponding to the first object, it is necessary to determine the similarity between the action feature representation corresponding to the first object and the action feature representations corresponding to other objects, wherein the second object (one of the other objects) is:
[0137] The action feature representation corresponding to the first object includes multiple sub-feature representations, and the action feature representation corresponding to the second object also includes multiple sub-feature representations. The process iterates through the multiple sub-feature representations corresponding to the first object. For the j-th sub-feature representation, the similarity between the j-th sub-feature representation and each of the multiple sub-feature representations corresponding to the second object is calculated to obtain the sub-similarity corresponding to the j-th sub-feature representation. For example, a weighted summation is performed on the similarities between the j-th sub-feature representation and each of the multiple sub-feature representations corresponding to the second object to obtain the sub-feature similarity corresponding to the j-th sub-feature representation. Based on the multiple sub-feature similarities corresponding to the first object, the similarity between the action feature representation corresponding to the first object and the action feature representation corresponding to the second object is determined. For example, the average of the multiple sub-feature similarities corresponding to the first object is used as the similarity between the action feature representation corresponding to the first object and the action feature representation corresponding to the second object. Here, j is a positive integer.
[0138] When performing a weighted summation of the similarities between the j-th sub-feature representation and the multiple sub-feature representations corresponding to the second object, it is necessary to determine the weights of the multiple sub-feature representations corresponding to the second object. The method for determining the weights of the multiple sub-feature representations corresponding to the second object is described below:
[0139] If the same first action descriptor exists in the description text corresponding to the first object and the description text corresponding to the second object, determine the first similarity between the first sub-feature representation of the first object and the second sub-feature representation of the second object; the first sub-feature representation refers to the feature representation in the action feature representation corresponding to the first object that is associated with the first action descriptor, and the second sub-action feature representation refers to the feature representation in the action feature representation corresponding to the second object that is associated with the first action descriptor.
[0140] Specifically, when the first similarity is greater than or equal to the preset similarity, the sub-feature similarity corresponding to the first sub-feature representation is determined based on the first weight and the first similarity.
[0141] Indicatively, the preset similarity is a pre-defined similarity threshold. If the first similarity between the first sub-feature representation and the second sub-feature representation of the second object is greater than or equal to the preset similarity, it means that the actions generated by the model for different objects are highly similar for the same first action descriptor. A larger first weight is used as the weight between the first and second sub-feature representations. For example, for the same action descriptor "running", when the similarity of the "running" sub-feature representations of object 1 and object 2 is high, a larger weight makes this similar part more prominent in the overall calculation. The model will be more penalized in subsequent training, thus prompting it to generate more differentiated "running" actions.
[0142] Optionally, the third weight is used as the weight between the first sub-feature representation and other sub-feature representations (excluding the second sub-feature representation) corresponding to the second object; based on the weights corresponding to the multiple sub-feature representations corresponding to the second object, a weighted summation is performed on the similarities between the first sub-feature representation and the multiple sub-feature representations corresponding to the second object to obtain the sub-feature similarity corresponding to the first sub-feature representation. The third weight is less than the first weight; for example, the third weight is 1 and the first weight is 1.2.
[0143] Specifically, if the first similarity is less than the preset similarity, the similarity of the sub-features corresponding to the first sub-feature representation is determined based on the second weight and the first similarity.
[0144] To illustrate, if the first similarity between the first sub-feature representation and the second sub-feature representation of the second object is less than a preset similarity, it indicates that the actions generated by the model for different objects are dissimilar for the same first action descriptor. The second weight is then used as the weight between the first and second sub-feature representations, and the third weight is used as the weight between the first sub-feature representation and other sub-feature representations corresponding to the second object (excluding the second sub-feature representation). Based on the weights corresponding to the multiple sub-feature representations corresponding to the second object, a weighted summation is performed on the similarities between the first sub-feature representation and the multiple sub-feature representations corresponding to the second object to obtain the sub-feature similarity corresponding to the first sub-feature representation. Here, the first weight is greater than the second weight, and the second weight is less than or equal to the third weight; for example, the first weight is 1.2, the second weight is 0.9, and the third weight is 1.
[0145] After determining the sub-feature similarity corresponding to the first sub-feature representation, the second loss corresponding to the first object is determined based on the sub-feature similarity corresponding to the first sub-feature representation.
[0146] Indicatively, a process similar to that used to determine the sub-feature similarity corresponding to the first sub-feature representation is adopted to determine the sub-feature similarity corresponding to other sub-feature representations (excluding the first sub-feature representation) corresponding to the first object. The average of the sub-feature similarities corresponding to the multiple sub-feature representations corresponding to the first object is calculated as the similarity between the action feature representation corresponding to the first object and the action feature representation corresponding to the second object. This similarity is used as the loss between the first object and the second object.
[0147] In the above scheme, by finely measuring and weighting the similarity of actions with the same action descriptor on different objects, the problem of action homogenization when the model generates actions for multiple objects is effectively prevented.
[0148] Based on the losses between the first object and all other objects (second objects), a second loss corresponding to the first object is determined. For example, assuming objects 1, 2, and 3, for object 1, the similarity between the action feature representation of object 1 and the action feature representation of object 2 is calculated as loss 1; the similarity between the action feature representation of object 1 and the action feature representation of object 3 is calculated as loss 2; the average of loss 1 and loss 2 is used as the second loss, or the maximum value of loss 1 and loss 2 is used as the second loss, or the weighted average of loss 1 and loss 2 is used as the second loss, etc.
[0149] Step 3: Generate a model based on the training sample content corresponding to at least two objects using the first loss and the second loss respectively. The candidate content generation model is obtained by using the model.
[0150] Indicatively, for the i-th object, after obtaining the first loss and the second loss, the sample content generation model is trained with the goal of minimizing the first loss and maximizing the second loss, thus obtaining the candidate content generation model.
[0151] Optionally, a first target loss is determined based on a first loss and a second loss, and a sample content generation model is trained based on the first target loss to obtain a candidate content generation model.
[0152] To illustrate, we construct a first objective function, which can be set as: J1 = aL1 - bL2 (where L1 represents the first loss, L2 represents the second loss, and a and b are weights used to balance the importance of the first and second losses in the first objective function). The goal of training is to minimize the first objective function.
[0153] After obtaining the first target loss, the update gradient is calculated using the backpropagation algorithm. The model parameters in the sample content generation model are then updated based on the calculated update gradient, thus completing one training cycle of the sample content generation model. The first training set is then traversed, and the sample content generation model is trained using at least two of the obtained first target losses, thus completing one iterative training cycle on the first training set.
[0154] In the above embodiments, not only is the difference between the predicted dynamic content and the sample dynamic content considered (reflected by the first loss), but also the similarity of the action feature representations between different objects is considered (reflected by the second loss). The second loss can avoid the problem that the actions of all objects are too similar, i.e., homogenization, thereby improving the distinguishability of the actions of different objects in the dynamic content generated by the model.
[0155] Optionally, training can be stopped when the sample content generation model meets the training conditions, and the resulting model is the candidate content generation model. Meeting the training conditions means that the model's first loss or first target loss is less than a first preset loss, or the model has undergone a first preset number of training iterations.
[0156] In some embodiments, the sample content generation model includes a first model parameter and a second model parameter. The above relates to updating the model parameters in the sample content generation model, that is, adjusting the first model parameter in the sample content generation model. Illustratively, the first model parameter in the sample content generation model is frozen, and the second model parameter in the sample content generation model is adjusted according to a first loss or a first target loss to obtain a candidate content generation model.
[0157] Optionally, the second model parameters include at least one of the parameters of the hidden layer of the sample content generation model and the input parameters.
[0158] Step 305: Obtain the second data group corresponding to the i-th object.
[0159] The second data group corresponding to the i-th object includes sample dynamic content and descriptive text of the i-th object. It should be noted that the specific data content contained in the second data group corresponding to the i-th object may be the same as or different from the specific data content contained in the first data group corresponding to the i-th object. If the data content contained in the second data group corresponding to the i-th object is different from the specific data content contained in the first data group, then a new sample dynamic content and descriptive text of the i-th object needs to be obtained as the second data group. In some embodiments, the data quality contained in the second data group is higher than the data quality contained in the first data group. For example, the resolution of the sample skin effect video in the first data group is lower than the resolution of the sample skin effect video in the second data group; or, for example, the sample skin effect video in the first data group refers to a video clip of the skin 1 effect video corresponding to virtual character a, or the sample skin effect video in the first data group refers to the complete video of the skin 1 effect video corresponding to virtual character a, etc. The sample dynamic content is used to characterize the visual presentation method of the i-th object, where i is a positive integer. Optionally, the sample dynamic content includes at least one of video, GIF, animation, etc.
[0160] The descriptive text is used to describe the visual presentation of the i-th object in text form. In other words, the descriptive text is used to describe the content details of the sample's dynamic content in text form.
[0161] The second data group corresponding to the i-th object is used as the second dataset.
[0162] Step 306: For the second data group corresponding to the i-th object, analyze the descriptive text corresponding to the i-th object through the candidate content generation model to obtain the predicted dynamic content corresponding to the i-th object.
[0163] This is illustrated by analyzing each descriptive text in the second dataset using a candidate content generation model to obtain the predicted dynamic content corresponding to each descriptive text. The process of analyzing each descriptive text in the second dataset using a candidate content generation model to obtain the predicted dynamic content corresponding to each descriptive text can be referenced from the process of analyzing descriptive text using a sample content generation model in step 230, and will not be elaborated here.
[0164] In some embodiments, the i-th object includes the i-th virtual character. Optionally, the method for obtaining the predicted dynamic content corresponding to the i-th object further includes the following steps:
[0165] Step 1: Obtain the character setting parameters corresponding to the i-th virtual character.
[0166] Optionally, character setting parameters refer to the basic setting parameters of the i-th virtual character in the game application. Character setting parameters include at least one of the following: body size, character type, skills, and equipment of the i-th virtual character.
[0167] Step 2: For the second data group corresponding to the i-th virtual character, perform dynamic sampling of the i-th virtual character. The noise-adding process is executed to obtain the dynamic content of the sample with added noise.
[0168] Optionally, random noise is added to the sample dynamic content of the i-th virtual character to obtain sample dynamic content with added noise. Illustratively, random noise refers to random values that conform to a specific probability distribution (such as Gaussian distribution, uniform distribution, etc.). Adding random noise to the sample dynamic content means changing the pixel values of at least two image frames corresponding to the sample dynamic content based on random noise.
[0169] Step 3: Extract the text feature table corresponding to the descriptive text of the i-th object using the candidate content generation model. The character feature representation corresponding to the character setting parameters of the i-th virtual character is shown and extracted.
[0170] The character setting parameters are input into the candidate content generation model in text form. The candidate content generation model includes a text encoder, which extracts the text feature representation of the description text corresponding to the i-th object and the character feature representation of the character setting parameters corresponding to the i-th virtual character.
[0171] Step 4: Using text feature representations and role feature representations as conditional information, the candidate content generation model... Guided by conditional information, a denoising process is performed on the dynamic content of the sample with added noise to obtain the pre-denoising result for the i-th object. Test dynamic content.
[0172] Optionally, the candidate content generation model includes a denoising network layer and a video decoder. Conditional information and noisy sample dynamic content are input into the denoising network layer, and the denoising result is input. The denoising network layer uses text feature representation and role feature representation as conditional information to gradually remove noise from the noisy sample dynamic content, thereby obtaining the denoising result. Finally, the denoising result is decoded by the video decoder to obtain the predicted dynamic content corresponding to the i-th object.
[0173] In the above scheme, text feature representation and character feature representation are used as conditional information to guide the denoising process, providing a clear and precise direction for the model to generate and predict dynamic content. Text feature representation encompasses descriptive information about the object, while character feature representation includes detailed settings for the virtual character. The combination of the two allows the model to accurately determine which parts need repair and adjustment, and how to perform these operations, during the denoising process. For example, when recreating a character's skill release animation, the model can precisely adjust the speed, trajectory, and special effects of the animation based on the expected skill effect in the text description and the specific specifications of the skill in the character's settings, thereby improving the quality of the generated dynamic content and making the generated content more in line with the expected requirements.
[0174] Step 307: Train the candidate content generation model based on the predicted dynamic content and sample dynamic content corresponding to the i-th object to obtain the first content generation model corresponding to the i-th object.
[0175] Optionally, a third loss is determined based on the difference between the predicted dynamic content and the sample dynamic content corresponding to the i-th object; a sample content generation model is trained based on the third loss to obtain a candidate content generation model.
[0176] Schematic, the difference between the predicted dynamic content corresponding to the i-th object and the sample dynamic content corresponding to the i-th object is calculated to obtain the third loss. The function used to calculate the third loss includes at least one of the following: mean squared error loss function, cross-entropy loss function, and mean absolute error loss function, etc., which are not limited here. After obtaining the third loss, the update gradient is calculated through the backpropagation algorithm. The model parameters in the candidate content generation model are updated according to the calculated update gradient, thereby completing one training process of the candidate content generation model.
[0177] In some embodiments, the second data group may include multiple data pairs, each data pair including one sample dynamic content containing the i-th object and one descriptive text corresponding to the sample dynamic content.
[0178] Traverse the second training set and train the candidate content generation model based on at least one third loss obtained, which completes one iteration of training on the second training set.
[0179] In the above embodiments, for each object, the candidate content generation model is used to analyze its descriptive text to obtain predicted dynamic content. A third loss is determined based on the difference between the predicted dynamic content and the sample dynamic content corresponding to the object. Then, the candidate content generation model is trained again based on the third loss to obtain the final first content generation model. This can more accurately correct the deviation that the model may have when generating content for a specific object, continuously optimize the performance of the model, and make the generated dynamic content of higher quality.
[0180] In some embodiments, the loss used to train the sample content generation model further includes a fourth loss, which represents the semantic similarity between the predicted dynamic content corresponding to the i-th object and the descriptive text. Optionally, after obtaining the predicted dynamic content corresponding to the i-th object, the method for training the candidate content generation model further includes the following steps:
[0181] Optionally, a fourth loss is determined based on the semantic similarity between the predicted dynamic content and the descriptive text corresponding to the i-th object; a candidate content generation model is trained based on the third and fourth losses to obtain a first content generation model corresponding to the i-th object.
[0182] To obtain the predicted dynamic content corresponding to the i-th object, the first semantic feature representation corresponding to the predicted dynamic content is extracted through the first semantic feature extraction model. The first semantic feature representation is used to characterize the semantics of the predicted dynamic content. The second semantic feature representation corresponding to the descriptive text (corresponding to the predicted dynamic content) is extracted through the second semantic feature extraction model. The second semantic feature representation is used to characterize the semantics of the descriptive text.
[0183] A fourth loss is determined based on the similarity between the first semantic feature representation and the second semantic feature representation. Optionally, the similarity between the first semantic feature representation and the second semantic feature representation is determined as the fourth loss. The similarity between the features can be at least one of Euclidean distance, Manhattan distance, cosine similarity, etc., and is not limited here.
[0184] Optionally, a second target loss is determined based on the third and fourth losses, and a candidate content generation model is trained based on the second target loss to obtain a first content generation model.
[0185] To illustrate, a second objective function is constructed, which can be set as: J2 = cL3 + dL4 (where L3 represents the third loss, L4 represents the fourth loss, and c and d are weights used to balance the importance of the third and fourth losses in the second objective function). The goal of training is to minimize the second objective function.
[0186] After obtaining the second objective loss, the update gradient is calculated using the backpropagation algorithm. The model parameters in the candidate content generation model are then updated based on the calculated update gradient, thus completing one training cycle of the candidate content generation model. The second training set is then traversed, and the candidate content generation model is trained using at least one obtained second objective loss, thus completing one iterative training cycle on the second training set.
[0187] In the above embodiments, the model is trained by combining the third and fourth losses. This approach considers both the difference between the predicted dynamic content and the sample dynamic content, as well as the semantic similarity between the predicted dynamic content and the descriptive text. This ensures that the generated content accurately conveys the intent of the descriptive text at the semantic level, thereby further improving the quality of the dynamic content generated by the model.
[0188] Optionally, training is stopped when the candidate content generation model meets the training conditions, and the resulting model is the first content generation model. Meeting the training conditions means that the model's third loss or second objective loss is less than the second preset loss, or the model has undergone a second preset number of iterations.
[0189] In some embodiments, the candidate content generation model includes a third model parameter (corresponding to the first model parameter of the sample generation model) and a fourth model parameter (corresponding to the second model parameter of the sample generation model). The third model parameter is a model parameter that does not need to be updated (frozen) when training the candidate content generation model, and the fourth model parameter is a model parameter that needs to be updated (adjusted) when training the candidate content generation model.
[0190] Optionally, the third model parameters in the candidate content generation model are frozen, and the fourth model parameters in the candidate content generation model are adjusted based on the third loss or the second target loss to obtain the first content generation model.
[0191] Optionally, the fourth model parameters include at least one of the parameters of the hidden layer of the candidate content generation model and the input parameters. For a description of the fourth model parameters, please refer to the description of the second model parameters; no further limitations are specified here.
[0192] The first content generation model corresponding to the i-th object can generate dynamic content with different visual presentation methods based on the object features of different objects according to the text description. When generating dynamic content for the i-th object, the first content generation model corresponding to the i-th object can generate dynamic content containing the detailed features of the i-th object according to the text description.
[0193] In summary, the training method for the content generation model provided in this application, by using a first data set corresponding to at least two objects for training, helps the sample content generation model understand the basic rules of dynamic content generation and the differences in behavioral patterns between different objects; by using a second data set corresponding to a specified object for training, the model can learn more deeply the behavioral characteristics and appearance change details of the specified object, thereby improving the accuracy of the trained first content generation model in generating dynamic content containing the specified object.
[0194] In some embodiments, the training method of the above-described content generation model can train a content generation model for generating skin effect videos of game characters. (Illustratively, the above...) Figure 2 or Figure 3 The illustrated embodiment further includes the following steps:
[0195] Step 401: Obtain the sample content generation model to be trained.
[0196] The sample content generation model is used to generate videos based on text descriptions.
[0197] In illustrative terms, this sample content generation model can be implemented as a pre-trained video generation model. The sample content generation model has basic video generation capabilities and can generate videos that conform to the text description based on the input text.
[0198] In this embodiment, a two-stage fine-tuning method is used to train the sample content generation model. For illustrative purposes, please refer to [the relevant documentation / reference]. Figure 5 The diagram illustrates a content generation model training method. For the sample content generation model 501, the first data set corresponding to at least two game characters is used as the first dataset to fine-tune the sample content generation model to obtain the candidate content generation model 502. Then, the second data set corresponding to the i-th game character is used as the second dataset to fine-tune the candidate content generation model to obtain the first content generation model 503. The first content generation model 503 is used to generate a skin effect video that matches the description of the skin effect of the specified character.
[0199] The training framework used in the first-stage and second-stage fine-tuning is similar. First, data preparation is performed, and either a first or second training set is obtained. Then, data augmentation is performed on the data in the first and second training sets. Next, model training continues. Finally, model prediction (i.e., model inference) can be performed based on the model obtained from one iteration of training. Data is selected from the data obtained from the model prediction as data in the training set, which is to perform data augmentation. After performing data augmentation, the model is trained again based on the data in the augmented data set. This process of data augmentation, model training, and model prediction is repeated until the trained model meets the training conditions, at which point training stops.
[0200] The training steps for the first-stage and second-stage fine-tuning are explained in detail below. Steps 402 to 405 describe the data preparation, data augmentation, model training, and model prediction processes in the first-stage fine-tuning; and steps 406 to 409 describe the data preparation, data augmentation, model training, and model prediction processes in the second-stage fine-tuning.
[0201] Step 402: Obtain the first data set corresponding to at least two game characters.
[0202] The first dataset is the first set of data corresponding to at least two game characters.
[0203] The first data group corresponding to the i-th game character includes a sample skin effect video of the i-th game character and a description text of the skin effect.
[0204] In some embodiments, for the first data group corresponding to the i-th game character, the first data group may include one or more data pairs, each data pair including one sample skin effect video containing the i-th game character and one skin effect description text corresponding to the sample skin effect video.
[0205] In illustrative purposes, the sample skin effect video obtained above may be a promotional video for a certain game skin of the i-th game character; or the sample skin effect video may refer to a game video of the i-th game character wearing a certain game skin in a game match, etc. This application embodiment does not limit this.
[0206] The data augmentation methods include at least one of the following:
[0207] (1) Video splicing processing: For at least one sample skin effect video corresponding to the i-th game character, the at least one sample skin effect video can be divided into multiple video segments, and then the multiple video segments are spliced according to different splicing strategies. Assuming that the resulting spliced video 1 and spliced video 2 are obtained, the skin effect description text corresponding to spliced video 1 and the sample skin effect video is added to the first data group as the training data pair corresponding to the i-th game character; the skin effect description text corresponding to spliced video 2 and the sample skin effect video is added to the first data group as the training data pair corresponding to the i-th game character.
[0208] (2) Noise addition: Random noise is added to the sample skin effect video to obtain the sample skin effect video with added noise. The sample skin effect video with added noise and the skin effect description text corresponding to the sample skin effect video are used as training data pairs and added to the first dataset.
[0209] (3) Scaling process: The sample skin effect video is scaled according to random scaling parameters to obtain the scaled sample skin effect video. The scaled sample skin effect video and the skin effect description text corresponding to the sample skin effect video are added to the first dataset as training data pairs.
[0210] (4) Speed-up processing: The sample skin effect video is speed-up processed according to the random speed-up playback parameters to obtain the speed-up sample skin effect video. The speed-up sample skin effect video and the skin effect description text corresponding to the sample skin effect video are used as training data pairs and added to the first dataset.
[0211] For the skin effect description text corresponding to the sample skin effect video, data augmentation processing can also be performed on the skin effect description text, where the data augmentation processing method includes at least one of the following methods:
[0212] (1) Text transformation processing: Input the skin effect description text into the large language model, transform the skin effect description text through the large language model to obtain the transformed skin effect description text, and add the transformed skin effect description text and the sample skin effect video corresponding to the original skin effect description text as training data pairs into the first dataset.
[0213] (2) Text back translation processing: The text describing the skin effect is translated into various other languages, such as English and French, using translation software. The translated text is then translated back into the original language to obtain the back-translated text. The sample skin effect videos corresponding to the back-translated text and the original skin effect description text are used as training data pairs and added to the first dataset.
[0214] The examples of data augmentation methods described above are merely illustrative and are not intended to be limiting.
[0215] Step 403: For the first data group corresponding to at least two game characters, analyze the skin effect description text corresponding to at least two game characters through the sample content generation model to obtain the predicted skin effect video corresponding to at least two game characters.
[0216] This is illustrated using the example of analyzing the skin effect description text of the i-th game character through a sample content generation model:
[0217] After inputting the skin effect description text of the i-th game character into the sample content generation model, the text encoding layer extracts the text feature representation of the skin effect description text with the first input prefix. Then, the text feature representation is used as conditional information. Guided by this conditional information, a denoising network layer denoises the sample skin effect video with added random noise to obtain the denoised result. Finally, a video decoding layer decodes the denoised result to obtain the predicted skin effect video corresponding to the i-th object. In some embodiments, the aforementioned conditional information may include other information besides the text feature representation. For example, the sample content generation model may also include a visual encoder, which compresses the sample skin effect video into a low-dimensional latent space representation, and uses this low-dimensional latent space representation and the text feature representation as conditional information.
[0218] Step 404: Based on the difference between the predicted skin effect video and the sample predicted skin effect video corresponding to at least two game characters, determine the first loss corresponding to at least two game characters.
[0219] To illustrate, we obtain the predicted skin effect video corresponding to each descriptive text, use the sample skin effect video corresponding to each descriptive text as a label, and calculate the difference between the predicted skin effect video and the sample skin effect video to obtain the first loss.
[0220] Step 405: Based on the first loss training sample content generation model corresponding to at least two game characters, a candidate content generation model is obtained.
[0221] The sample content generation model includes a first model parameter and a second model parameter. Optionally, the second model parameter includes the prefix parameter of the first input prefix and the parameters of the last hidden layer in the sample content generation model.
[0222] To illustrate, the first model parameters in the sample content generation model are frozen, and the second model parameters in the sample content generation model are adjusted according to the first loss.
[0223] The sample content generation model is trained sequentially using the first loss corresponding to at least two game characters, which completes one iteration of training on the first dataset.
[0224] Optionally, after completing one iteration of training on the first dataset, a trained sample content generation model is obtained. The trained sample content generation model is then used to analyze at least two skin effect description texts in the first dataset to obtain predicted skin effect videos corresponding to the at least two skin effect description texts. From the at least two predicted skin effect videos, predicted skin effect videos whose first loss with the label (the sample skin effect video corresponding to the predicted skin effect video) is less than or equal to a first preset value are selected and added to the first dataset. That is, the predicted skin effect video and its corresponding skin effect description text are used as training data to expand the first dataset. The sample content generation model is then trained a second time using the expanded first dataset.
[0225] Optionally, training can be stopped when the sample content generation model meets the training conditions, and the resulting model is the candidate content generation model. Meeting the training conditions means that the model's first loss is less than a first preset loss, or the model has undergone a first preset number of training iterations.
[0226] Step 406: Obtain the second data set corresponding to the i-th game character.
[0227] The second data group corresponding to the i-th game character includes a sample skin effect video of the i-th game character and a text describing the skin effect.
[0228] In some embodiments, for the second data group corresponding to the i-th game character, the second data group may include multiple data pairs, each data pair including one sample skin effect video containing the i-th game character and one skin effect description text corresponding to the sample skin effect video.
[0229] Optionally, after obtaining the sample skin effect video and skin effect description text, data augmentation processing can be performed on the sample skin effect video and skin effect description text. The specific process of data augmentation processing can be found in step 402, which will not be elaborated here.
[0230] Step 407: For the second data group corresponding to the i-th game character, analyze the skin effect description text corresponding to the i-th game character through the candidate content generation model to obtain the predicted skin effect video corresponding to the i-th game character.
[0231] The second data set corresponding to the i-th game character is used as the second dataset.
[0232] In a schematic manner, after inputting the skin effect description text of the i-th game character into the candidate content generation model, the text feature representation of the skin effect description text with the second input prefix is extracted through the text encoding layer. Then, the text feature representation is used as conditional information. Under the guidance of the conditional information, the sample skin effect video with added random noise is denoised through the denoising network layer to obtain the denoising result. Finally, the denoising result is decoded through the video decoding layer to obtain the predicted skin effect video corresponding to the i-th object.
[0233] In some embodiments, the conditional information may include other information besides the text feature representation. For example, the candidate content generation model may also include a visual encoder, which is used to compress the sample skin effect video into a low-dimensional latent space representation, and the low-dimensional latent space representation and the text feature representation are used as conditional information.
[0234] Step 408: Determine the third loss based on the difference between the predicted skin effect video and the sample skin effect video corresponding to the i-th game character.
[0235] To illustrate, after obtaining the predicted skin effect video, the sample skin effect video is used as the label, and the difference between the predicted skin effect video and the sample skin effect video is calculated to obtain the third loss.
[0236] Step 409: Train the candidate content generation model based on the third loss to obtain the first content generation model corresponding to the i-th game character.
[0237] The candidate content generation model includes a first model parameter and a second model parameter. Optionally, the second model parameter includes the prefix parameter of the second input prefix and the parameters of the last hidden layer in the candidate content generation model.
[0238] To illustrate, the first model parameters in the candidate content generation model are frozen, and the second model parameters in the candidate content generation model are adjusted according to the third loss.
[0239] The candidate content generation model is trained using at least one third loss calculated based on the second dataset, which completes one iteration of training on the second dataset.
[0240] Optionally, after completing one iteration of training on the second dataset, a trained candidate content generation model is obtained. The trained candidate content generation model is then used to analyze at least two skin effect description texts in the second dataset to obtain predicted skin effect videos corresponding to at least two skin effect description texts. From the at least two predicted skin effect videos, predicted skin effect videos whose third loss with the label (the sample skin effect video corresponding to the predicted skin effect video) is less than or equal to a second preset value are selected and added to the second dataset. That is, the predicted skin effect video and its corresponding skin effect description text are used as training data to expand the second dataset, and the candidate content generation model is trained a second time using the expanded second dataset.
[0241] Optionally, training is stopped when the candidate content generation model meets the training conditions, and the resulting model is the first content generation model. Meeting the training conditions means that the model's third loss is less than the second preset loss, or the model has undergone a second preset number of training iterations.
[0242] In some embodiments, after obtaining the first content generation model, the first content generation model can be applied to the game industry. For illustrative examples, please refer to [reference needed]. Figure 6 The diagram illustrates a computer system applying a first content generation model, comprising a server 610, a gateway 620, and a mobile client 630. The mobile client 630 can be implemented as a client for a game application, which can be any of the following: first-person shooter (FPS), third-person shooter (TPS), multiplayer online battle arena (MOBA), strategy game (SLG), party game, building game, or game community application.
[0243] Server-side 610 is used to provide background services for the game application, and server-side 610 includes at least one server. Figure 6 Several types of servers are shown, including:
[0244] (1) Lobby server, which is used to handle player login to the game application and to manage various functions within the game lobby.
[0245] (2) Load management server, used to manage the resource usage of each server, including processor utilization, memory usage, network bandwidth, etc.
[0246] (3) Battle server, used to manage data related to game matches.
[0247] The first content generation model can be set on the lobby server. When a player triggers a login operation in the mobile client 630, the mobile client 630 sends a login request to the gateway 620 through the wireless communication network. The gateway 620 forwards the login request to the lobby server on the server side 610. After receiving the login request, the lobby server sends the login page display data to the mobile client 630 after confirming that the player's login information has been verified. The display data includes a game character skin effect video generated by the first content generation model. Optionally, the game character skin effect video is a video automatically generated after the developers input a description of the game character's skin effect into the first content generation model.
[0248] After receiving the displayed data, the mobile client 630 will display the login page based on the data, which includes a video of the game character's skin effects.
[0249] This is illustrative; please refer to it. Figure 7 It shows a structural block diagram of a training device for a content generation model, such as Figure 7 As shown, the device includes:
[0250] The acquisition module 710 is used to acquire a sample content generation model to be trained. The sample content generation model is used to generate dynamic content based on text description. The dynamic content refers to the content containing at least two consecutive image frames. The module acquires first data groups corresponding to at least two objects respectively. The first data group corresponding to the i-th object includes sample dynamic content containing the i-th object and descriptive text. The sample dynamic content is used to characterize the visual presentation of the i-th object, and the descriptive text is used to describe the visual presentation of the i-th object in text form. i is a positive integer.
[0251] Analysis module 720 is used to analyze the descriptive text corresponding to the at least two objects respectively through the sample content generation model to obtain the predicted dynamic content corresponding to the at least two objects respectively;
[0252] The training module 730 is used to train the sample content generation model based on the predicted dynamic content and sample dynamic content corresponding to the at least two objects respectively, to obtain a first content generation model. The first content generation model is used to generate dynamic content with different visual presentation methods based on the object features of different objects according to the text description.
[0253] In some embodiments, the acquisition module 710 is further configured to acquire a second data set corresponding to the i-th object, wherein the second data set corresponding to the i-th object includes sample dynamic content and descriptive text containing the i-th object; the training module 730 is configured to train the sample content generation model based on the predicted dynamic content and sample dynamic content corresponding to the at least two objects respectively, to obtain a candidate content generation model; for the second data set corresponding to the i-th object, analyze the descriptive text corresponding to the i-th object through the candidate content generation model to obtain the predicted dynamic content corresponding to the i-th object; train the candidate content generation model based on the predicted dynamic content and sample dynamic content corresponding to the i-th object to obtain a first content generation model corresponding to the i-th object.
[0254] In some embodiments, the training module 730 is configured to determine a first loss corresponding to each of the at least two objects based on the difference between the predicted dynamic content and the sample dynamic content corresponding to each of the at least two objects; and to train the sample content generation model based on the first loss corresponding to each of the at least two objects to obtain the candidate content generation model.
[0255] In some embodiments, the training module 730 is configured to extract action feature representations of the predicted dynamic content corresponding to the at least two objects, wherein the action feature representation corresponding to the i-th object is used to characterize the action performance of the i-th object in the predicted dynamic content corresponding to the i-th object; based on the action feature representations corresponding to the at least two objects, determine a second loss corresponding to the at least two objects, wherein the second loss corresponding to the i-th object is used to characterize the similarity between the action feature representation corresponding to the i-th object and the action feature representations corresponding to other objects; and train the sample content generation model based on the first loss and the second loss corresponding to the at least two objects to obtain the candidate content generation model.
[0256] In some embodiments, the object includes a virtual character; the predicted dynamic content corresponding to the at least two objects includes the predicted dynamic content corresponding to the i-th virtual character; the training module 730 is used to determine the skeletal key points of the character model in each image frame corresponding to the predicted dynamic content corresponding to the i-th virtual character, wherein the character model refers to the three-dimensional model corresponding to the i-th virtual character; and to determine the motion feature representation based on the skeletal key points of the character model in at least two image frames corresponding to the predicted dynamic content.
[0257] In some embodiments, the action feature representation includes multiple sub-feature representations, and different sub-feature representations are associated with different action descriptive words in the description text; the at least two objects include a first object and a second object; the training module 730 is configured to determine a first similarity between the first sub-feature representation of the first object and the second sub-feature representation of the second object when the same first action descriptive word exists in the description text corresponding to the first object and the description text corresponding to the second object; the first sub-feature representation refers to the feature representation in the action feature representation corresponding to the first object that is associated with the first action descriptive word, and the second sub-action feature representation refers to the feature representation in the action feature representation corresponding to the second object that is associated with the first action descriptive word; when the first similarity is greater than or equal to a preset similarity, the sub-feature similarity corresponding to the first sub-feature representation is determined based on a first weight and the first similarity; when the first similarity is less than the preset similarity, the sub-feature similarity corresponding to the first sub-feature representation is determined based on a second weight and the first similarity; wherein, the first weight is greater than the second weight; and a second loss corresponding to the first object is determined based on the sub-feature similarity corresponding to the first sub-feature representation.
[0258] In some embodiments, the training module 730 is used to determine a third loss based on the difference between the predicted dynamic content and the sample dynamic content corresponding to the i-th object; and to train the candidate content generation model based on the third loss to obtain a first content generation model corresponding to the i-th object.
[0259] In some embodiments, the training module 730 is used to determine a fourth loss based on the semantic similarity between the predicted dynamic content and the descriptive text corresponding to the i-th object; and to train the candidate content generation model based on the third loss and the fourth loss to obtain a first content generation model corresponding to the i-th object.
[0260] In some embodiments, the i-th object includes the i-th virtual character; the acquisition module 710 is further configured to acquire the character setting parameters corresponding to the i-th virtual character; the character setting parameters include at least one of the body shape, character type, skills, and equipment of the i-th virtual character; the analysis module 720 is configured to perform a noise-adding process on the sample dynamic content of the i-th virtual character for the second data group corresponding to the i-th virtual character, to obtain sample dynamic content with added noise; extract the text feature representation corresponding to the descriptive text corresponding to the i-th object and extract the character feature representation corresponding to the character setting parameters corresponding to the i-th virtual character through the candidate content generation model; use the text feature representation and the character feature representation as conditional information, and under the guidance of the conditional information, perform a noise-removing process on the sample dynamic content with added noise through the candidate content generation model to obtain the predicted dynamic content corresponding to the i-th object.
[0261] In some embodiments, the sample content generation model includes a first model parameter and a second model parameter; the training module 730 is used to freeze the first model parameter in the sample content generation model, and adjust the second model parameter in the sample content generation model based on the predicted dynamic content and sample dynamic content corresponding to the at least two objects respectively, to obtain the first content generation model.
[0262] In some embodiments, the second model parameters include at least one of the following parameters:
[0263] Input prefix parameters, which are trainable random parameters added at the beginning of the descriptive text;
[0264] Hidden layer parameters, which refer to the parameters of the hidden layer in the sample content generation model.
[0265] In summary, the training device for the content generation model provided in this application training method uses sample dynamic content and descriptive text of at least two different objects to train the sample content generation model. This allows the model to continuously strengthen its understanding of the characteristics of different objects during training, learn the unique features of each object and the relationship between them and their corresponding visual presentation methods. This enables the final first content generation model to generate distinctive dynamic content based on the differences in the characteristics of the objects themselves when faced with the same descriptive words but different objects, thereby improving the targeting and accuracy of the generated dynamic content and thus enhancing the quality of the generated dynamic content.
[0266] It should be noted that the specific limitations of the embodiments of the training device for one or more content generation models provided above can be found in the limitations of the training method for content generation models above, and will not be repeated here. Each module of the above device can be implemented entirely or partially by software, hardware, or a combination thereof. Each module can be embedded in the processor of the computer device in hardware form or independent of the processor, or it can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0267] This application also provides a computer device, which includes: a processor and a memory, wherein the memory stores a computer program; the processor is used to execute the computer program in the memory to implement the training method of the content generation model provided in the above method embodiments.
[0268] For example, Figure 8 This is a structural block diagram of a computer device 800 provided in an exemplary embodiment of this application. Optionally, the computer device 800 is a server 800.
[0269] Typically, server 800 includes a processor 801 and memory 802.
[0270] Processor 801 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 801 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). Processor 801 may also include a main processor and a coprocessor. The main processor, also known as a Central Processing Unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 801 may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 801 may also include an Artificial Intelligence (AI) processor, which is used to handle computational operations related to machine learning.
[0271] The memory 802 may include one or more computer-readable storage media, which may be non-transitory. The memory 802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 802 are used to store at least one instruction, which is executed by the processor 801 to implement the training method of the content generation model provided in the various method embodiments of this application.
[0272] In some embodiments, the server 800 may also optionally include an input interface 803 and an output interface 804. The processor 801, memory 802, and input interfaces 803 and 804 can be connected via a bus or signal lines. Various peripheral devices can be connected to the input interfaces 803 and 804 via a bus, signal lines, or a circuit board. The input interfaces 803 and 804 can be used to connect at least one input / output (I / O) related peripheral device to the processor 801 and memory 802. In some embodiments, the processor 801, memory 802, and input interfaces 803 and 804 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 801, memory 802, and input interfaces 803 and 804 can be implemented on separate chips or circuit boards, and this application does not limit this aspect.
[0273] Those skilled in the art will understand that Figure 8 The structure shown does not constitute a limitation on the computer device 800, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0274] In an exemplary embodiment, this application provides a chip including programmable logic circuits and / or program instructions, which, when run on a computer device, are used to implement the training method of the content generation model provided in the above-described method embodiments.
[0275] In an exemplary embodiment, this application provides a computer-readable storage medium storing a computer program that is loaded and executed by a processor to implement the training method of the content generation model provided in the above-described method embodiments.
[0276] In an exemplary embodiment, this application provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the processor of the computer device to load and execute the training method for the content generation model provided in the above-described method embodiments.
[0277] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0278] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0279] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0280] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A training method for a content generation model, characterized in that, The method includes: Obtain a sample content generation model to be trained. The sample content generation model is used to generate dynamic content based on text description. The dynamic content refers to content containing at least two consecutive image frames. Obtain at least two first data sets corresponding to each object; wherein, the first data set corresponding to the i-th object includes sample dynamic content and descriptive text containing the i-th object, the sample dynamic content is used to characterize the visual presentation of the i-th object, and the descriptive text is used to describe the visual presentation of the i-th object in text form, where i is a positive integer; The sample content generation model is used to analyze the descriptive text corresponding to the at least two objects to obtain the predicted dynamic content corresponding to the at least two objects. The sample content generation model is trained based on the predicted dynamic content and sample dynamic content corresponding to the at least two objects respectively to obtain a first content generation model. The first content generation model is used to generate dynamic content with different visual presentation methods based on the object features of different objects according to the text description.
2. The method according to claim 1, characterized in that, The method further includes: Obtain the second data group corresponding to the i-th object, wherein the second data group corresponding to the i-th object includes sample dynamic content and descriptive text containing the i-th object; The step of training the sample content generation model based on the predicted dynamic content and sample dynamic content corresponding to the at least two objects respectively to obtain the first content generation model includes: The sample content generation model is trained based on the predicted dynamic content and sample dynamic content corresponding to the at least two objects respectively, to obtain the candidate content generation model; For the second data group corresponding to the i-th object, the candidate content generation model is used to analyze the descriptive text corresponding to the i-th object to obtain the predicted dynamic content corresponding to the i-th object; The candidate content generation model is trained based on the predicted dynamic content and sample dynamic content corresponding to the i-th object to obtain the first content generation model corresponding to the i-th object.
3. The method according to claim 2, characterized in that, The step of training the sample content generation model based on the predicted dynamic content and sample dynamic content corresponding to the at least two objects respectively, to obtain the candidate content generation model, includes: Based on the difference between the predicted dynamic content and the sample dynamic content corresponding to the at least two objects respectively, the first loss corresponding to the at least two objects is determined; The sample content generation model is trained based on the first loss corresponding to each of the at least two objects to obtain the candidate content generation model.
4. The method according to claim 3, characterized in that, The method further includes: Extract the action feature representations of the predicted dynamic content corresponding to the at least two objects respectively, wherein the action feature representation corresponding to the i-th object is used to characterize the action performance of the i-th object in the predicted dynamic content corresponding to the i-th object; Based on the action feature representations corresponding to the at least two objects respectively, a second loss corresponding to the at least two objects is determined, wherein the second loss corresponding to the i-th object is used to characterize the similarity between the action feature representation corresponding to the i-th object and the action feature representations corresponding to other objects; The step of training the sample content generation model based on the first loss corresponding to each of the at least two objects to obtain the candidate content generation model includes: The sample content generation model is trained based on the first loss and the second loss corresponding to the at least two objects respectively, to obtain the candidate content generation model.
5. The method according to claim 4, characterized in that, The object includes a virtual character; the predicted dynamic content corresponding to the at least two objects includes the predicted dynamic content corresponding to the i-th virtual character. The step of extracting the action feature representations of the predicted dynamic content corresponding to the at least two objects includes: For the predicted dynamic content corresponding to the i-th virtual character, determine the skeletal key points of the character model in each image frame corresponding to the predicted dynamic content, where the character model refers to the three-dimensional model of the i-th virtual character. The motion feature representation is determined based on the skeletal key points of the character model in at least two image frames corresponding to the predicted dynamic content.
6. The method according to claim 4, characterized in that, The action feature representation includes multiple sub-feature representations, and different sub-feature representations are associated with different action descriptive words in the descriptive text; the at least two objects include a first object and a second object; The step of determining the second loss corresponding to each of the at least two objects based on the action feature representations corresponding to the at least two objects includes: If the same first action descriptor exists in the description text corresponding to the first object and the description text corresponding to the second object, a first similarity is determined between the first sub-feature representation of the first object and the second sub-feature representation of the second object; the first sub-feature representation refers to the feature representation in the action feature representation corresponding to the first object that is associated with the first action descriptor, and the second sub-action feature representation refers to the feature representation in the action feature representation corresponding to the second object that is associated with the first action descriptor; If the first similarity is greater than or equal to a preset similarity, the sub-feature similarity corresponding to the first sub-feature representation is determined based on the first weight and the first similarity; if the first similarity is less than the preset similarity, the sub-feature similarity corresponding to the first sub-feature representation is determined based on the second weight and the first similarity; wherein, the first weight is greater than the second weight. Based on the sub-feature similarity corresponding to the first sub-feature representation, the second loss corresponding to the first object is determined.
7. The method according to claim 2, characterized in that, The step of training the candidate content generation model based on the predicted dynamic content and sample dynamic content corresponding to the i-th object to obtain the first content generation model corresponding to the i-th object includes: Based on the difference between the predicted dynamic content and the sample dynamic content corresponding to the i-th object, a third loss is determined; The candidate content generation model is trained based on the third loss to obtain the first content generation model corresponding to the i-th object.
8. The method according to claim 7, characterized in that, The method further includes: The fourth loss is determined based on the semantic similarity between the predicted dynamic content and the descriptive text corresponding to the i-th object; The step of training the candidate content generation model based on the third loss to obtain the first content generation model corresponding to the i-th object includes: The candidate content generation model is trained based on the third loss and the fourth loss to obtain the first content generation model corresponding to the i-th object.
9. The method according to claim 2, characterized in that, The i-th object includes the i-th virtual character; Obtain the character setting parameters corresponding to the i-th virtual character; The character setting parameters include at least one of the following: the body size, character type, skills, and equipment of the i-th virtual character; The step of analyzing the descriptive text corresponding to the i-th object through the candidate content generation model for the second data group corresponding to the i-th object to obtain the predicted dynamic content corresponding to the i-th object includes: For the second data group corresponding to the i-th virtual character, a noise-adding process is performed on the sample dynamic content of the i-th virtual character to obtain sample dynamic content with added noise. The candidate content generation model is used to extract the text feature representation of the description text corresponding to the i-th object and the character feature representation of the character setting parameters corresponding to the i-th virtual character. Using the text feature representation and the role feature representation as conditional information, the candidate content generation model performs a denoising process on the noisy sample dynamic content under the guidance of the conditional information to obtain the predicted dynamic content corresponding to the i-th object.
10. The method according to any one of claims 1 to 9, characterized in that, The sample content generation model includes a first model parameter and a second model parameter; The step of training the sample content generation model based on the predicted dynamic content and sample dynamic content corresponding to the at least two objects respectively to obtain the first content generation model includes: Freeze the first model parameters in the sample content generation model, and adjust the second model parameters in the sample content generation model based on the predicted dynamic content and sample dynamic content corresponding to the at least two objects respectively, to obtain the first content generation model.
11. The method according to claim 10, characterized in that, The second model parameters include at least one of the following parameters: Input prefix parameters, which are trainable random parameters added at the beginning of the descriptive text; Hidden layer parameters, which refer to the parameters of the hidden layer in the sample content generation model.
12. A training device for a content generation model, characterized in that, The device includes: The acquisition module is used to acquire a sample content generation model to be trained. The sample content generation model is used to generate dynamic content based on text description. The dynamic content refers to the content containing at least two consecutive image frames. The module acquires first data sets corresponding to at least two objects respectively. The first data set corresponding to the i-th object includes sample dynamic content containing the i-th object and descriptive text. The sample dynamic content is used to characterize the visual presentation of the i-th object, and the descriptive text is used to describe the visual presentation of the i-th object in text form. i is a positive integer. The analysis module is used to analyze the descriptive text corresponding to the at least two objects respectively through the sample content generation model to obtain the predicted dynamic content corresponding to the at least two objects respectively; The training module is used to train the sample content generation model based on the predicted dynamic content and sample dynamic content corresponding to the at least two objects respectively, to obtain a first content generation model. The first content generation model is used to generate dynamic content with different visual presentation methods based on the object features of different objects according to the text description.
13. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one program, which is loaded and executed by the processor to implement the training method of the content generation model as described in any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that, The storage medium stores at least one program segment, which is loaded and executed by a processor to implement the training method of the content generation model as described in any one of claims 1 to 11.
15. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements a training method for the content generation model as described in any one of claims 1 to 11.