Data processing method for task processing model and virtual character animation generation method

The data processing method for a task processing model integrates multimodal sample guide information to train a processing model, enhancing the accuracy and versatility of motion generation for virtual characters in video, game, and digital human applications.

JP2025536026AActive Publication Date: 2025-10-30ALIBABA INNOVATION PRIVATE LIMITED
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025526440
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-15
Filing Date
2023-12-14
Publication Date
2025-10-30
Estimated Expiration
2043-12-14

AI Technical Summary

Technical Problem

Current motion generation methods in video, game, and digital human applications rely on inefficient real-world character motion collection, requiring high environmental and hardware demands and result in low accuracy motions that need refinement.

Method used

A data processing method for a task processing model that integrates multimodal sample guide information to train an initial processing model, obtaining predicted task features, and transmitting model parameters to an end device for generating virtual character animations, utilizing a cloud-device and end-device interaction.

Benefits of technology

Enables accurate and versatile motion generation by integrating multiple tasks, improving the accuracy and versatility of motion generation for virtual characters in various applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536026000001_ABST
    Figure 2025536026000001_ABST
Patent Text Reader

Abstract

[0013] The present disclosure provides a data processing method for a task processing model and a virtual character animation generation method, the data processing method for the task processing model including the steps of: acquiring a first sample set, the first sample set including multimodal sample guide information; inputting the sample guide information and a sample task sequence into an initial processing model to acquire predicted task features corresponding to the sample guide information; training the initial processing model based on the predicted task features and the sample task features corresponding to the sample task sequence; and, when a first predetermined stopping condition is reached, acquiring model parameters of the processing model obtained by the training, the sample task features being obtained by quantizing and encoding the sample task sequence; and transmitting the model parameters of the trained task processing model to an end device. Because the task processing model is trained based on the multimodal sample guide information, multimodal task integration can be achieved, improving the accuracy and versatility of the model.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority to a Chinese patent application filed with the State Intellectual Property Office of the People's Republic of China on December 15, 2022, bearing application number 202211611176.1 and entitled "Data processing method for task processing model and virtual character animation generation method," the entire contents of which are incorporated herein by reference.

[0002] FIELD OF THE INVENTION The present disclosure relates to the field of computer technology, and in particular to a data processing method for a task processing model. One or more embodiments of the present disclosure simultaneously relate to a virtual character animation generation method, a data processing system for a task processing model, a data processing device for a task processing model, a virtual character animation generation device, a computing device, and a computer-readable storage medium. [Background technology]

[0003] With the development of computer technology, motion generation has gradually become a key step in video, game, and digital human applications, such as motion generation of characters in videos, characters in games, customer objects in web pages or application software, and virtual characters in film production. The realism of motion is one of the important factors that reflect the reality and naturalness of the interaction between characters and the environment.

[0004] Currently, motion generation relies on motion collection of many real characters, which is inefficient, places high demands on the environment and collection hardware, and the generated motions are low in accuracy and require refinement at a later stage. Therefore, there is a strong demand for a versatile and accurate motion generation method. Summary of the Invention [Problem to be solved by the invention]

[0005] In view of the above, the embodiments of the present specification provide a data processing method for a task processing model. One or more embodiments of the present specification simultaneously relate to a virtual character animation generation method, a data processing system for a task processing model, a data processing device for a task processing model, a virtual character animation generation device, a computing device, a computer-readable storage medium, and a computer program, which solve the technical shortcomings of the prior art. [Means for solving the problem]

[0006] According to a first aspect of an embodiment of the present specification, there is provided a data processing method of a task processing model executed by a cloud device connected to a plurality of end devices, the method comprising: obtaining a first sample set, the first sample set including multimodal sample guide information; inputting sample guide information and a sample task sequence into an initial processing model to obtain predicted task features corresponding to the sample guide information; training an initial process model based on the predicted task features and sample task features corresponding to the sample task sequence, and obtaining model parameters of the trained process model when a first predetermined stopping condition is reached, wherein the sample task features are obtained by quantizing and encoding the sample task sequence; and transmitting model parameters of the task processing model obtained by training to the end device.

[0007] According to a second aspect of the present disclosure, there is provided a data processing method for a task processing model executed by an end device connected to a cloud device, the method comprising: receiving model parameters of the task processing model sent by the cloud device, and constructing the task processing model based on the model parameters; receiving a task processing request input by a user, the task processing request including task guide information; inputting the task guide information and the entire mask task sequence into a task processing model, and acquiring target task features corresponding to the task guide information through processing by the task processing model, wherein the task processing model is obtained by training based on multimodal sample guide information, the sample task sequence, and the sample task features, and the sample task features are obtained by quantizing and encoding the sample task sequence; quantizing and decoding the target task features to obtain task processing results corresponding to the task guide information.

[0008] According to a third aspect of the present disclosure, there is provided a virtual character animation generation method, the method comprising: receiving a virtual character animation generation request sent by the front end, the virtual character animation generation request including animation guide information; a step of inputting the video guide information and the entire mask video sequence into a virtual character video generation model, and acquiring virtual character video features corresponding to the video guide information by processing the virtual character video generation model, wherein the virtual character video generation model is obtained by training based on a plurality of sample video guide information, sample video sequences and sample video features, and the sample video features are obtained by quantizing and encoding the sample video sequences; quantizing and decoding the virtual character animation features to obtain a virtual character action sequence corresponding to the animation guide information; and generating a virtual character animation based on the virtual character action sequence and transmitting the generated animation to the front end, thereby displaying the virtual character animation on the front end.

[0009] According to a fourth aspect of the present disclosure, there is provided a data processing device of a task processing model executed by a cloud device connected to a plurality of end devices, the data processing device comprising: an acquisition module configured to acquire a first sample set, the first sample set including multimodal sample guide information; a first input module configured to input sample guide information and a sample task sequence into an initial processing model to obtain predicted task features corresponding to the sample guide information; a training module configured to train an initial process model based on the predicted task features and sample task features corresponding to the sample task sequence, and to obtain model parameters of the trained process model when a first predetermined stopping condition is reached, wherein the sample task features are obtained by quantizing and encoding the sample task sequence; and a transmission module configured to transmit model parameters of the task processing model obtained by training to the end device.

[0010] According to a fifth aspect of the embodiments of the present specification, there is provided a data processing device of a task processing model executed by an end device connected to a cloud device, the device comprising: a construction module configured to receive model parameters of the task processing model sent by the cloud device and construct the task processing model based on the model parameters; a first receiving module configured to receive a task processing request input by a user, the task processing request including task guide information; a second input module configured to input the task guide information and the entire mask task sequence into a task processing model and obtain target task features corresponding to the task guide information through processing by the task processing model, wherein the task processing model is obtained by training based on multimodal sample guide information, the sample task sequence, and the sample task features, and the sample task features are obtained by quantizing and encoding the sample task sequence; a first decoding module configured to quantize and decode the target task features to obtain task processing results corresponding to the task guide information.

[0011] According to a sixth aspect of the present disclosure, there is provided a virtual character animation generation device, the device comprising: a second receiving module configured to receive a virtual character animation generation request sent by the front end, the virtual character animation generation request including animation guide information; a third input module configured to input the video guide information and the full mask video sequence into a virtual character video generation model, and obtain virtual character video features corresponding to the video guide information through processing of the virtual character video generation model, wherein the virtual character video generation model is obtained by training based on a plurality of sample video guide information, sample video sequences and sample video features, and the sample video features are obtained by quantizing and encoding the sample video sequences; a first decoding module configured to quantize and decode the virtual character animation features to obtain a virtual character action sequence corresponding to the video guide information; and a generating module configured to generate a virtual character animation based on the virtual character action sequence and send it to the front end, thereby causing the front end to display the virtual character animation.

[0012] According to a seventh aspect of the present disclosure, there is provided a data processing system of a task processing model, the system comprising: an end device for constructing a first sample set and transmitting the first sample set to a cloud device, the first sample set including multimodal sample guide information; The method includes: inputting sample guide information and a sample task sequence into an initial processing model; obtaining predicted task features corresponding to the sample guide information; training the initial processing model based on the predicted task features and sample task features corresponding to the sample task sequence; and, when a first predetermined stopping condition is reached, obtaining model parameters of the processing model obtained by training, the sample task features being obtained by quantizing and encoding the sample task sequence; and a cloud device for transmitting the model parameters of the task processing model obtained by training to the end device.

[0013] According to an eighth aspect of the present disclosure, there is provided a computing device, the computing device comprising: a memory and a processor, The memory is adapted to store computer-executable instructions, and the processor is adapted to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method provided by the first, second or third aspect.

[0014] According to a ninth aspect of embodiments herein, there is provided a computer-readable storage medium having stored thereon computer-executable instructions which, when executed by a processor, implement the steps of the method provided by the first, second or third aspect above.

[0015] According to a tenth aspect of the present disclosure, there is provided a computer program, which, when run on a computer, causes the computer to perform the steps of the method provided in the first, second or third aspect. [Effects of the Invention]

[0016] A data processing method for a task processing model according to one embodiment of the present specification includes: acquiring a first sample set, the first sample set including multimodal sample guide information; inputting the sample guide information and a sample task sequence into an initial processing model; acquiring predicted task features corresponding to the sample guide information; training the initial processing model based on the predicted task features and the sample task features corresponding to the sample task sequence; and, when a first predetermined stopping condition is reached, acquiring model parameters of the processing model obtained by training, the sample task features being obtained by quantizing and encoding the sample task sequence; and transmitting the model parameters of the trained task processing model to an end device. Because the task processing model is trained based on the multimodal sample guide information, multimodal task integration can be achieved, improving the accuracy and versatility of the model. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a configuration diagram of a data processing system of a task processing model according to one embodiment of the present specification. [Figure 2] FIG. 10 is a diagram illustrating the configuration of a data processing system of another task processing model according to one embodiment of the present specification. [Figure 3] 1 is a flowchart of a data processing method of a task processing model according to one embodiment of the present specification; [Figure 4] 10 is a flowchart of a data processing method of another task processing model according to an embodiment of the present specification; [Figure 5] 1 is a flowchart of a virtual character animation generation method according to one embodiment of the present specification. [Figure 6] 1 is a schematic diagram of a process of a virtual character animation generation method according to one embodiment of the present specification; [Figure 7] FIG. 2 is a schematic diagram of a virtual character animation generation interface according to one embodiment of the present disclosure. [Figure 8]1 is a flowchart of a data processing method for a quantized generative model according to one embodiment of the present specification. [Figure 9] 1 is a flowchart of training a task processing model and a quantized generative model according to one embodiment of the present specification. [Figure 10] 1 is a schematic diagram illustrating the structure of a data processing device of a task processing model according to one embodiment of the present specification; [Figure 11] FIG. 2 is a schematic diagram illustrating the structure of a data processing device of another task processing model according to one embodiment of the present specification. [Figure 12] 1 is a schematic diagram illustrating the structure of a virtual character animation generating device according to one embodiment of the present specification. [Figure 13] FIG. 2 is a block diagram illustrating the structure of a computing device according to one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0018] In order to facilitate a thorough understanding of the present specification, numerous specific details are set forth in the following description. However, the present specification can be embodied in many other forms different from those described herein, and those skilled in the art can deduce the same without departing from the spirit of the present specification. Therefore, the present specification is not limited to the specific implementations disclosed below.

[0019] The terminology used in one or more examples herein is for the purpose of describing a particular example only and is not intended to limit one or more examples herein. As used in one or more examples herein and in the appended claims, the singular forms "a," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Additionally, the term "and / or" as used in one or more examples herein should be understood to refer to and include any and all possible combinations of one or more associated listed items.

[0020] In one or more embodiments herein, various pieces of information may be described with terms such as "first," "second," etc., but it should be understood that such information should not be limited by these terms. These terms are used only to distinguish between the same types of information. For example, a first piece of information may be referred to as a second piece of information, and similarly, a second piece of information may be referred to as a first piece of information, without departing from the scope of one or more embodiments herein. Depending on the context, the word "if" used herein may be interpreted as "when," "when," or "responsive to determining."

[0021] First, the terms used in one or more embodiments of this specification will be interpreted.

[0022] A Transformer is a neural network structure.

[0023] Seq2seq is a coding-decoding structure in which the input and output may be sequences of unequal length.

[0024] Human motion is skeletal data that drives a human body model to move its limbs in 3D (three-dimensional) digital humans, 3D games, and 3D videos. It is composed of multiple frames, and the data for each frame describes the orientation and displacement of the entire body, as well as the rotation angle of each joint in the body.

[0025] Diffusion models are a novel deep generative model. During training, samples are given noise of different intensities (intensity = 0, 1, ..., or T), and the model is asked to reconstruct the original sample based on the noise intensity and the resulting sample. During inference, starting with a single random noise sample (noise intensity = T), the model is asked to gradually remove the noise to obtain samples with noise intensities of T-1, T-2, ..., 0, respectively. The sample with noise intensity 0 is the final sample generated by the model.

[0026] With the development of computer technology, motion generation has gradually become a key step in video, game, and digital human applications. Traditional motion generation methods rely on collecting motions from many real-world characters, which is inefficient, places high demands on the environment and collection hardware, and requires refinement at a later stage. Currently, learning-based methods have made progress in motion generation, achieving high-quality continuous motion generation under different conditions by learning from large amounts of motion capture data. However, currently, to meet the needs of diverse motion generation tasks, specific frameworks are often used to train each corresponding task individually, and they are unable to benefit from different tasks or types of datasets.

[0027] To solve the above problems, embodiments of the present specification provide a task processing framework that integrates multiple tasks, eliminates the gap between different tasks, learns higher-level semantic information from different cross-modal data signals, and enables efficient representation and learning of different generation tasks, thereby realizing a variety of different generation requirements and providing abundant action materials for video, games, digital human applications, etc. at low cost, efficiently, and accurately. Specifically, a first sample set is obtained, the first sample set including multimodal sample guide information. The sample guide information and sample task sequences are input into an initial processing model. Predicted task features corresponding to the sample guide information are obtained. The initial processing model is trained based on the predicted task features and the sample task features corresponding to the sample task sequences. When a first predetermined stopping condition is reached, model parameters of the trained processing model are obtained. The sample task features are obtained by quantizing and encoding the sample task sequences. The model parameters of the trained task processing model are sent to the end device. Because the task processing model is trained based on multimodal sample guide information, multimodal task integration is achieved, improving the accuracy and versatility of the model.

[0028] This specification provides a data processing method for a task processing model, and also relates to a virtual character animation generation method, a data processing system for a task processing model, a data processing device for a task processing model, a virtual character animation generation device, a computing device, a computer-readable storage medium, and a computer program, which will be described in detail in the following examples.

[0029] Referring to FIG. 1, FIG. 1 illustrates a configuration diagram of a task processing model data processing system according to one embodiment of the present specification, where the task processing model data processing system includes a cloud device 100 and an end device 200; The end device 200 constructs a first sample set and transmits the first sample set to the cloud device 100, the first sample set including multimodal sample guide information; The cloud device 100 inputs sample guide information and a sample task sequence into an initial processing model, obtains predicted task features corresponding to the sample guide information, trains the initial processing model based on the predicted task features and sample task features corresponding to the sample task sequence, and when a first predetermined stopping condition is reached, obtains model parameters of the processing model obtained by training, where the sample task features are obtained by quantizing and encoding the sample task sequence, and transmits the model parameters of the task processing model obtained by training to the end device 200.

[0030] Using the solution of the embodiments of the present specification, a first sample set is obtained, the first sample set including multimodal sample guide information, the sample guide information and the sample task sequence are input into an initial processing model, predicted task features corresponding to the sample guide information are obtained, the initial processing model is trained based on the predicted task features and the sample task features corresponding to the sample task sequence, and when a first predetermined stopping condition is reached, model parameters of the processing model obtained by training are obtained, the sample task features are obtained by quantizing and encoding the sample task sequence, and the model parameters of the trained task processing model are sent to the end device. Because the task processing model is trained based on the multimodal sample guide information, multimodal task integration can be achieved and the accuracy and versatility of the model can be improved.

[0031] 2, Fig. 2 shows a configuration diagram of a data processing system of another task processing model according to one embodiment of the present specification, which may include a cloud device 100 and a plurality of end devices 200. A communication connection can be established between the plurality of end devices 200 via the cloud device 100, and in a task processing scenario, the cloud device 100 is used to provide task processing services between the plurality of end devices 200, and the plurality of end devices 200 can realize real-time communication via the cloud device 100 as a sender or a receiver, respectively.

[0032] A user can interact with the cloud device 100 through the end device 200 to receive data transmitted by other end devices 200 or transmit data to other end devices 200. In a task processing scenario, a user can transmit a data stream to the cloud device 100 through the end device 200, and the cloud device 100 can generate an action based on the data stream and push the action generation result to other end devices with which communication has been established.

[0033] A connection between the end device 200 and the cloud device 100 is established via a network. The network is a medium that provides a communication link between the end device and the cloud device. The network may include various connection types, such as wired, wireless communication links, or fiber optic cables. Data transmitted by the end device 200 may require processing, such as encoding, transcoding, or compression, before being delivered to the cloud device 100.

[0034] The end device 200 may be a web page application such as a browser, an APP (Application), or an H5 (HyperText Markup Language 5) application, a mini-application (also called an applet, a small application program), or a cloud application. The end device 200 can be developed and acquired based on a software development kit (SDK) of a corresponding service provided by a cloud device, for example, a real-time communication (RTC) SDK. The end device 200 can be deployed in an electronic device and needs to run or depend on an APP in the device. The electronic device may have a display screen and support information browsing, and may be a personal mobile terminal such as a mobile phone, tablet, or personal computer. Various other types of applications, such as a man-machine interaction application, a model training application, a text processing application, a web page browser application, a shopping application, a search application, an instant communication tool, a mailbox end device, or social platform software, may also be deployed in the electronic device.

[0035] The cloud device 100 may include servers providing various services, such as a server providing communication services to multiple end devices, a server for background training to support models used by the end devices, and a server processing data sent by the end devices. The cloud device 100 may be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server may be a server in a distributed system or a server coupled to a blockchain. The server may be a cloud server for basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), big data, and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud hosting equipped with artificial intelligence technology.

[0036] Referring to FIG. 3, FIG. 3 shows a flowchart of a data processing method of a task processing model according to one embodiment of the present specification, where the data processing method of the task processing model is applied to an end device connected to a cloud device, and specifically includes the following steps 302 to 308:

[0037] In step 302, the model parameters of the task processing model sent by the cloud device are received, and the task processing model is constructed based on the model parameters.

[0038] In the embodiments of the present specification, in order to perform task processing accurately and efficiently, model parameters of a task processing model sent by a cloud device can be received, a task processing model can be constructed based on the model parameters, and task processing can be realized by the task processing model.

[0039] In step 304, a task processing request input by a user is received, the task processing request including task guide information.

[0040] In one or more embodiments of the present specification, task guide information is obtained at an early stage of job processing, and the task processing process is guided based on the task guide information, thereby efficiently and accurately generating processing results that match the task guide information.

[0041] Specifically, the task processing request may be a processing request for a different task, and the task may include, but is not limited to, a text generation task, a motion generation task, a voice processing task, etc. The task guide information is multimodal and may include, but is not limited to, text guide information, image guide information, audio guide information, trajectory guide information, video guide information, etc., and may be selected according to actual circumstances, and the embodiments of the present specification are not limited thereto. Taking the motion generation task as an example, the motion guide information may be, for example, motion guide information for different digital objects such as virtual characters, virtual animals, and virtual vehicles. The text guide information may include, but is not limited to, motion categories and natural language description text. The image guide information may be understood as a visual signal, such as a reference image, which may include a reference action, an action at a partial time point, etc., and aims to ensure that the actions at corresponding times in the generated action sequence, such as action insertion (in-betweening) and action filling (in-filling), are consistent. The trajectory guide information may be understood as a trajectory signal, which means positions corresponding to different points in time in a motion sequence, and can be used to control the direction of travel of an object and for motion control tasks.

[0042] In actual use, there are various ways to receive a task processing request input by a user, and specific methods are selected according to actual situations, and the embodiments of this specification are not limited thereto. In one possible implementation form of this specification, a task processing request voluntarily sent by a user can be received, and the task processing request includes task guide information. In another possible implementation form of this specification, the task processing request includes an information identifier of the task guide information, and task guide information corresponding to the information identifier can be obtained from a guide information base, and the guide information base includes multiple task guide information.

[0043] In step 306, the task guide information and the entire mask task sequence are input into a task processing model, and target task features corresponding to the task guide information are obtained through processing by the task processing model, where the task processing model is obtained by training based on the multimodal sample guide information, the sample task sequence, and the sample task features, and the sample task features are obtained by quantizing and encoding the sample task sequence.

[0044] In one or more embodiments of the present specification, after receiving a task processing request input by a user, task guide information and the entire mask task sequence can be further input into a task processing model, and target task features corresponding to the task guide information can be obtained through processing by the task processing model.

[0045] Specifically, the task processing model is a Transformer model capable of extracting features from task guide information, and the task guide information can be processed as discrete target task features through processing by the task processing model.

[0046] The task processing model includes a first encoder and a first decoder, and the step of inputting the task guide information and the entire mask task sequence into the task processing model and obtaining the target task features corresponding to the task guide information through processing by the task processing model includes: inputting task guide information into a first encoder and obtaining task guide features corresponding to the task guide information; inputting the task guide features and the entire mask task sequence into a first decoder to obtain target task features corresponding to the task guide information.

[0047] Specifically, a task sequence is a continuous representation of the task itself; for example, a motion sequence includes 23 main joint points, and the motion is mainly represented by the rotation angles of each skeleton, and the angles are continuous.

[0048] The task guide information can be input to a first encoder to encode the task guide information and generate an encoded representation corresponding to the task guide information, i.e., task guide features. After obtaining the task guide features, unlike training a task processing model, there is no sample task sequence to refer to in actual use, so masking is repeatedly performed based on the full mask task sequence to achieve task processing with a task sequence without reference, and the complete task sequence can be predicted, and the target task features corresponding to the task guide information can be obtained.

[0049] Using the solution of the embodiments of this specification, the task guide information is input to a first encoder to obtain the task guide function corresponding to the task guide information, and the task guide feature and the entire mask task sequence are input to a first decoder to obtain the target task feature corresponding to the task guide information, thereby realizing efficient and accurate acquisition of the target task feature.

[0050] In step 308, the target task features are quantized and decoded to obtain task processing results corresponding to the task guide information.

[0051] In one or more embodiments of the present specification, a task processing request input by a user is received, the task processing request includes task guide information, the task guide information and the entire mask task sequence are input into a task processing model, and target task features corresponding to the task guide information are obtained through processing by the task processing model, and then the target task features are quantized and decoded to obtain task processing results corresponding to the task guide information.

[0052] Using the solution of the embodiments of the present specification, a task processing model is received as a model parameter of a task processing model sent by a cloud device, a task processing model is constructed based on the model parameter, a task processing request input by a user is received, the task processing request includes task guide information, the task guide information and the entire mask task sequence are input into the task processing model, and target task features corresponding to the task guide information are obtained through processing by the task processing model, the task processing model is trained and obtained based on multimodal sample guide information, sample task sequence, and sample task features, the sample task features are obtained by quantizing and encoding the sample task sequence, and the target task features are quantized and decoded to obtain a task processing result corresponding to the task guide information. Because the task processing model is trained and obtained based on multimodal sample guide information, multimodal task integration can be realized, and the task processing model can efficiently and accurately generate target task features, further improving the accuracy of the task processing result.

[0053] In the embodiments of this specification, the task processing model includes a first encoder and a first decoder, and the training methods of the first encoder and the first decoder will be described in detail in the following embodiments.

[0054] Referring to FIG. 4, FIG. 4 shows a flowchart of another task processing model data processing method according to one embodiment of the present specification, where the task processing model data processing method is applied to a cloud device connected to multiple end devices, and specifically includes the following steps 402 to 408:

[0055] In step 402, a first sample set is obtained, the first sample set including multimodal sample guide information.

[0056] Specifically, the sample guide information is a multimodal control signal, and includes at least two of the following information: sample text information, sample screen information, sample trajectory information, sample audio information, and sample video information.

[0057] There are several ways to obtain the first sample set. The first sample set may be constructed by artificially inputting a large amount of multimodal sample guide information, or by reading a large amount of multimodal sample guide information from another data acquisition device or database. The method for obtaining the first sample set is specifically selected according to the actual situation, and the examples in this specification are not limited thereto.

[0058] In step 404, the sample guide information and the sample task sequence are input into an initial processing model to obtain predicted task features corresponding to the sample guide information.

[0059] In step 406, an initial processing model is trained based on the predicted task features and the sample task features corresponding to the sample task sequence. When a first predetermined stopping condition is reached, model parameters of the processing model obtained by the training are obtained, and the sample task features are obtained by quantizing and encoding the sample task sequence.

[0060] In step 408, the model parameters of the task processing model obtained through training are transmitted to the end device.

[0061] Specifically, the sample task sequence is a sequential representation of the sample task itself. The initial processing model includes a first encoder and a first decoder, and the first predetermined stopping condition includes a first stopping sub-condition.

[0062] Using the solution of the embodiments of the present specification, a first sample set is obtained, the first sample set including multimodal sample guide information, the sample guide information and the sample task sequence are input into an initial processing model, predicted task features corresponding to the sample guide information are obtained, the initial processing model is trained based on the predicted task features and the sample task features corresponding to the sample task sequence, and when a first predetermined stopping condition is reached, model parameters of the processing model obtained by training are obtained, the sample task features are obtained by quantizing and encoding the sample task sequence, and the model parameters of the trained task processing model are sent to the end device. Because the task processing model is trained based on the multimodal sample guide information, multimodal task integration can be achieved and the accuracy and versatility of the model can be improved.

[0063] In one alternative embodiment of the present specification, the step of inputting the sample guide information and the sample task sequence into an initial processing model and obtaining predicted task features corresponding to the sample guide information includes: A step of extracting first sample guide information from a first sample set, wherein the first sample guide information is any one of the sample guide information in the first sample set; inputting first sample guide information into a pre-trained first encoder to obtain first sample guide features; masking the sample task sequence to obtain a mask task sequence; inputting the first sample guide feature and the mask task sequence into a first decoder to obtain a first predicted task feature; the step of training an initial process model based on the predicted task features and the sample task features corresponding to the sample task sequence, and obtaining model parameters of the trained process model when a first predetermined stopping condition is reached; calculating a decoding loss value based on the first predicted task features and sample task features corresponding to the sample task sequence; and adjusting parameters of the first decoder based on the decoding loss value, returning to extract first sample guide information from the first sample set, and obtaining a trained first decoder if a first stopping sub-condition is reached.

[0064] The pre-trained first encoder refers to the first encoder of the pre-trained initial processing model. When training the first decoder, a training method similar to the feature extraction model (Vision Transformer) can be used, in which the sample task sequence corresponding to the sample guide information is decomposed, and adjacent time points and key points are considered as one patch, thereby obtaining multiple patches, and each patch is further smoothed and mapped to the input embedding by one linear layer, and a learnable [Motion MASK] token is added to represent the masked position.

[0065] Furthermore, the training process of the first decoder can adopt an iterative mask training method, which includes a shrinkage and restoration process, specifically shrinking at discrete code levels, and each iteration replaces the feature (embedding) in the sample task sequence with the embedding of the mask representation (mask token) with a predetermined probability. As for the sampling method, the network can be used to predict the sampling position, thereby improving the convergence speed and making the prediction result more stable.

[0066] In one possible implementation of the present specification, the first stopping sub-condition includes: the decoding loss value is less than or equal to a first predetermined threshold, and the first predetermined threshold is specifically selected according to actual circumstances, and the embodiments of the present specification are not limited thereto. Input the first sample guide feature and the mask task sequence into a first decoder to obtain a first predicted task feature. After obtaining the first predicted task feature, calculate a decoding loss value based on the first predicted task feature and the sample task feature corresponding to the sample task sequence, and compare the decoding loss value with the first predetermined threshold.

[0067] Specifically, if the decoding loss value is greater than the first predetermined threshold, it indicates that the difference between the first predicted task feature and the sample task feature corresponding to the sample task sequence is large, and the first decoder's prediction ability for the first sample guide feature and the mask task sequence is low. At this time, the parameters of the first decoder can be adjusted, and the step of extracting the first sample guide information from the multimodal sample guide information can be performed again to continue training the first decoder. If the decoding loss value is less than or equal to the first predetermined threshold, it indicates that the difference between the first predicted task feature and the sample task feature corresponding to the sample task sequence is small, and the first stopping sub-condition is reached, and the trained first decoder is obtained.

[0068] Using the method of the embodiments of this specification, the first sample guide feature and the mask task sequence are input into the first decoder, a first predicted task feature is obtained, a decoding loss value is calculated based on the first predicted task feature and the sample task feature corresponding to the sample task sequence, the decoding loss value is compared with a first predetermined threshold, if it is greater than the first predetermined threshold, the first decoder continues to be trained until the decoding loss value is less than or equal to the first predetermined threshold, the training for the first decoder is completed, and the parameters of the first decoder are continuously adjusted to make the finally obtained first decoder more accurate.

[0069] In another possible implementation of this specification, in addition to comparing the magnitude relationship between the decoding loss value and the first predetermined threshold, it can be determined whether the current first decoder has been trained according to the number of iterations.

[0070] Specifically, if the decoding loss value is greater than the first predetermined threshold, adjust the parameters of the first decoder, return to the step of extracting the first sample guide information from the multi-modal sample guide information, continue to train the first decoder, and when the first predetermined number of iterations is reached, stop the iteration and obtain a trained first decoder, and the first predetermined number of iterations is specifically selected according to the actual situation, and the embodiments of this specification are not limited thereto.

[0071] Using the method of the embodiments of this specification, the first sample guide feature and the mask task sequence are input into the first decoder, a first predicted task feature is obtained, a decoding loss value is calculated based on the first predicted task feature and the sample task feature corresponding to the sample task sequence, the decoding loss value is compared with a first predetermined threshold, and if the decoding loss value is greater than the first predetermined threshold, the first decoder continues to be trained until a first predetermined number of iterations is reached, the training for the first decoder is completed, and the parameters of the first decoder are continuously adjusted to make the finally obtained first decoder more accurate.

[0072] In actual use, there are many functions for calculating the decoding loss value, such as a cross-entropy loss function, an L1 norm loss function, a maximum loss function, a mean square error loss function, a logarithmic loss function, etc., and the specific function to be selected depends on the actual situation, and the examples in this specification are not limited thereto.

[0073] In one alternative embodiment of the present specification, the first predetermined stopping condition includes a second stopping sub-condition, and the training manner of the first encoder of the task processing model is: obtaining a second sample set, the second sample set including multimodal sample guide information, the sample guide information including sample guide features; A step of extracting second sample guide information from the second sample set, wherein the second sample guide information is any one of the sample guide information in the second sample set; inputting second sample guide information into a first encoder to obtain first predicted guide features corresponding to the second sample guide information; Calculating a coding loss value based on the first predicted guide feature and the second sample guide feature included in the second sample guide information; and adjusting parameters of the first encoder based on the encoding loss value, returning to extract second sample guide information from the second sample set, and obtaining a trained first encoder if a second stopping sub-condition is reached.

[0074] Specifically, the task processing model structurally adopts an encoder-decoder framework, where the first encoder and the first decoder are connected via a cross-attention layer to realize seq2seq association and no longer rely on the compression of a single implicit variable. Both the first encoder and the first decoder are multi-layer bidirectional transformer structures. The first encoder inputs a multimodal control signal, i.e., multimodal sample guide information, which includes at least two of the following information: sample text information, sample image information, sample trajectory information, sample audio information, and sample video information.

[0075] In addition, when training the first encoder, a large-scale pre-training model, such as a multimodal generative model (OFA, One-For-All) or an image-text correlation matching model (CLIP, Contrastive Language-Image Pre-training), may be adopted. The specific selection is based on the actual situation, and the embodiments of this specification are not limited thereto.

[0076] In actual use, there are multiple ways to obtain the second sample set. The second sample set can be constructed by artificially inputting a large amount of multimodal sample guide information, or by reading a large amount of multimodal sample guide information from other data acquisition devices or databases. The way to obtain the second sample set is specifically selected according to the actual situation, and the examples in this specification are not limited thereto.

[0077] In one possible implementation form of the present specification, the second stopping sub-condition includes: the encoding loss value is less than or equal to a second predetermined threshold, and the second predetermined threshold is specifically selected according to actual circumstances, and the embodiments of the present specification are not limited thereto. Input the second sample guide information into the first decoder, obtain a first predicted guide feature corresponding to the second sample guide information, and after obtaining the first predicted guide feature, calculate the encoding loss value according to the first predicted guide feature and the second sample guide feature included in the second sample guide information, and compare the encoding loss value with the second predetermined threshold.

[0078] Specifically, if the encoding loss value is greater than the second predetermined threshold, it indicates that the difference between the first predicted guide features and the second sample guide features contained in the second sample guide information is large, and the first encoder's predictive ability for the second sample guide information is low. At this time, the parameters of the first encoder can be adjusted, and the step of extracting the second sample guide information from the multimodal sample guide information can be executed again to continue training the first encoder. If the encoding loss value is less than or equal to the second predetermined threshold, it indicates that the difference between the first predicted guide features and the second sample guide features contained in the second sample guide information is small, and the second stopping sub-condition is reached, and the trained first encoder is obtained.

[0079] Using the method of the embodiments of this specification, the second sample guide information is input into the first encoder, and the first predicted guide feature corresponding to the second sample guide information is obtained. Based on the first predicted guide feature and the second sample guide feature contained in the second sample guide information, a coding loss value is calculated. The coding loss value is compared with a second predetermined threshold. If the coding loss value is greater than the second predetermined threshold, the first encoder continues to be trained until the coding loss value is less than or equal to the second predetermined threshold, completing the training for the first encoder. The parameters of the first encoder are continuously adjusted to make the final first encoder more accurate.

[0080] In another possible implementation of the present specification, in addition to comparing the magnitude relationship between the encoding loss value and the second predetermined threshold, it can be determined whether the current first encoder has been trained according to the number of iterations.

[0081] Specifically, if the encoding loss value is greater than the second predetermined threshold, adjust the parameters of the first encoder, return to the step of extracting the second sample guide information from the multimodal sample guide information, continue to train the first encoder, and when the second predetermined number of iterations is reached, stop the iteration and obtain the trained first encoder, and the second predetermined number of iterations is specifically selected according to the actual situation, and the embodiments in this specification are not limited thereto.

[0082] Using the method of the embodiments of this specification, the second sample guide information is input into the first encoder, and the first predicted guide feature corresponding to the second sample guide information is obtained. Based on the first predicted guide feature and the second sample guide feature contained in the second sample guide information, a coding loss value is calculated. The coding loss value is compared with a second predetermined threshold. If the coding loss value is greater than the second predetermined threshold, the first encoder continues to be trained until a second predetermined number of iterations is reached, and the training for the first encoder is completed. The parameters of the first encoder are continuously adjusted to make the final first encoder more accurate.

[0083] In actual use, there are many functions for calculating the encoding loss value, such as a cross-entropy loss function, an L1 norm loss function, a maximum loss function, a root mean square error loss function, a logarithmic loss function, etc., and the specific function is selected according to the actual situation, and the embodiments of this specification are not limited thereto. Preferably, the encoding loss value can be calculated using a cross-entropy loss function, and the cross-entropy loss function is used to calculate the cross-entropy of the first predicted guide feature and the second sample guide feature included in the second sample guide information as the encoding loss value, thereby improving the efficiency of calculating the encoding loss value and thereby improving the training efficiency of the first encoder.

[0084] In one alternative embodiment of the present specification, the sample task features corresponding to the sample task sequence are obtained by quantizing and encoding the sample task sequence, i.e., before the step of calculating the decoding loss value based on the first predicted task features and the sample task features corresponding to the sample task sequence, The method may further include inputting the sample task sequence into a second encoder of the pre-trained quantized generative model, and obtaining sample task features corresponding to the sample task sequence through an encoding process of the second encoder.

[0085] The quantized generative model includes a second encoder and a second decoder, where the second encoder quantizes a continuous task sequence to obtain discrete task features, and the second decoder quantizes and decodes the discrete task features to reconstruct the task sequence. The second encoder of the quantized generative model (VQ, Vector Quantization) is an autoencoder structure for image generation, and when training the quantized generative model, a discrete coding dictionary (codebook) is introduced in the intermediate process of the network to target the reconstruction task, thereby iteratively updating the model parameters of the quantized generative model to obtain the quantized generative model.

[0086] By using the solution of the embodiments of this specification, the sample task sequence is input to the second encoder of the pre-trained quantized generative model, and the sample task features corresponding to the sample task sequence are obtained through the encoding process of the second encoder, thereby improving the efficiency and accuracy of obtaining the sample task features corresponding to the sample task sequence.

[0087] In one alternative embodiment of the present specification, when training a task processing model, the accuracy of the task processing model can be improved by inputting corresponding time step features (time step embedding) as training supplements in each iteration process, that is, the above step of inputting the first sample guide features and the mask task sequence into the first decoder to obtain the first predicted task features can be: acquiring pre-defined time step features, wherein the time step features and the number of training iterations correspond to each other one-to-one; inputting the time step features, the first sample guide features, and the mask task sequence into a first decoder to obtain a first predicted task feature.

[0088] Furthermore, since different time steps correspond to different learnable features, when training a task processing model, a timestep embedding corresponding to the current iteration number can be obtained, and by inputting the timestep embedding, the first sample guide feature, and the mask task sequence together into the first decoder, an accurate first predicted task feature can be obtained.

[0089] For example, suppose the current training iteration number is 2, determine the time step feature corresponding to the current training iteration number 2 as X, and then input the time step feature X, the first sample guide feature, and the mask task sequence together into the first decoder to obtain the first prediction task feature.

[0090] When using the embodiments of the present specification, a predetermined time step feature is obtained, the time step feature and the number of training iterations correspond to each other, and the time step feature, the first sample guide feature, and the mask task sequence are input to a first decoder to obtain a first predicted task feature. The time step feature is used as a supplement to the training task processing model to improve the accuracy of the task processing model.

[0091] In one alternative embodiment of the present specification, the quantization generation model includes a second encoder and a second decoder, and the training method of the quantization generation model is: obtaining a third sample set, the third sample set including a plurality of training sample sequences; extracting a first training sample sequence from a third sample set, the first training sample sequence being any one of the training sample sequences in the third sample set; inputting a first training sample sequence into a second encoder to obtain a first test feature; inputting the first test feature and the predetermined presentation information into a second decoder to obtain a first test sequence; calculating a quantization loss value based on the first training sample sequence and the first test sequence; adjusting parameters of the second encoder and the second decoder according to the quantization loss value, and returning to extracting the first training sample sequence from the third sample set; and obtaining model parameters of the trained quantized generative model when a second predetermined stopping condition is reached; and transmitting the model parameters of the trained quantized generative model to the end device.

[0092] Specifically, predetermined presentation information control can be introduced during the training process of the quantized generative model, i.e., by introducing predetermined presentation information into the second decoder side of the quantized generative model, the second decoder reconstructs the task sequence according to the predetermined presentation information and the first test feature compressed by the second encoder. Taking the motion generation task as an example, the predetermined presentation information is predetermined trajectory information, which can realize trajectory decoupling control. Furthermore, the motion direction can be randomly selected, i.e., different predetermined trajectory information can be set, and the quantized generative model can learn the association between motion and trajectory direction, thereby intuitively realizing a simple motion control function at the feature level.

[0093] In addition, in the embodiments of the present specification, the multiple training sample sequences in the third sample set may be sample sequences corresponding to multimodal sample information, and the multimodal sample information includes at least two of information such as sample text information, sample image information, sample trajectory information, sample audio information, and sample video information.

[0094] In actual use, there are multiple ways to obtain the third sample set. For example, the third sample set can be constructed by artificially inputting a large number of training sample sequences, or by reading a large number of training sample sequences from other data acquisition devices or databases. The way to obtain the third sample set is specifically selected according to the actual situation, and the examples in this specification do not limit this.

[0095] In one possible implementation of the present specification, the third predetermined stopping condition includes that the quantization loss value is less than or equal to a third predetermined threshold, and the third predetermined threshold is specifically selected according to actual circumstances, and the embodiments of the present specification are not limited thereto. The first test feature and the predetermined presentation information are input to the second decoder to obtain a first test sequence. After obtaining the first test sequence, the quantization loss value is calculated based on the first test sequence and the first training sample sequence, and the quantization loss value is compared with the third predetermined threshold.

[0096] Specifically, if the quantization loss value is greater than the third predetermined threshold, it indicates that the difference between the first test sequence and the first training sample sequence is large, and the predictive ability of the quantization generative model for the first test feature and the predetermined presentation information is low. At this time, the model parameters of the quantization generative model are adjusted, and the step of extracting the first training sample sequence from the multiple training sample sequences is performed again to continue training the quantization generative model. When the quantization loss value is less than or equal to the third predetermined threshold, it indicates that the difference between the first test sequence and the first training sample sequence is small, and the third predetermined stopping condition is reached, and a trained quantization generative model is obtained.

[0097] Using the method of the embodiments of this specification, a first training sample sequence is input into a second encoder to obtain a first test feature, the first test feature and predetermined presentation information are input into a second decoder to obtain a first test sequence, a quantization loss value is calculated based on the first training sample sequence and the first test sequence, the quantization loss value is compared with a third predetermined threshold, and if the quantization loss value is greater than the third predetermined threshold, the quantization generative model continues to be trained until the quantization loss value is less than or equal to the third predetermined threshold, completing the training of the quantization generative model, and continuously adjusting the model parameters of the quantization generative model to make the finally obtained quantization generative model more accurate.

[0098] In another possible implementation of the present specification, in addition to comparing the magnitude relationship between the quantization loss value and the third predetermined threshold, it can be determined whether the current quantization generative model has been trained according to the number of iterations.

[0099] Specifically, if the quantization loss value is greater than the third predetermined threshold, adjust the model parameters of the quantization generative model, return to the step of extracting a first training sample sequence from the plurality of training sample sequences, and continue to train the quantization generative model. When the third predetermined number of iterations is reached, stop the iteration and obtain a trained quantization generative model. The third predetermined number of iterations is specifically selected according to the actual situation, and the embodiments in this specification are not limited thereto.

[0100] Using the method of the embodiments of this specification, a first training sample sequence is input into a second encoder to obtain a first test feature, the first test feature and predetermined presentation information are input into a second decoder to obtain a first test sequence, a quantization loss value is calculated based on the first training sample sequence and the first test sequence, the quantization loss value is compared with a third predetermined threshold, and if the quantization loss value is greater than the third predetermined threshold, the quantization generative model continues to be trained until a third predetermined number of iterations is reached, and the training for the quantization generative model is completed, and the model parameters of the quantization generative model are continuously adjusted to make the finally obtained quantization generative model more accurate.

[0101] In actual use, there are many functions for calculating the quantization loss value, such as a cross-entropy loss function, an L1 norm loss function, a maximum loss function, a mean square error loss function, and a logarithmic loss function. The specific function to be selected depends on the actual situation, and the examples in this specification are not limited thereto.

[0102] In one alternative embodiment of the present specification, in order to ensure time continuity in the task sequence, after the step of inputting the first test feature and the predetermined presentation information into a second decoder to obtain the first test sequence, splitting the first test sequence at a random time to obtain a first test subsequence and a second test subsequence; calculating a continuity loss value based on the first test subsequence and the second test subsequence; The step of adjusting parameters of the second encoder and the second decoder according to the quantization loss value, and returning to extract the first training sample sequence from the third sample set, and obtaining model parameters of the trained quantization generative model when a second predetermined stopping condition is reached, comprises: The method may include adjusting parameters of the second encoder and the second decoder based on the quantization loss value and the continuity loss value, returning to extract the first training sample sequence from the third sample set, and obtaining model parameters of the trained quantized generative model when a second predetermined stopping condition is reached.

[0103] Taking the action generation task as an example, in the action generation process, even if two discrete action features are joined together, a continuous action sequence can be generated, which has the effect of a smooth transition. Therefore, in order to avoid the confusion of the time series in the action generation process and meet the requirement of temporal continuity of the action sequence, when training the quantized generative model, a continuous action sequence is divided into a first test subsequence (seq1) and a second test subsequence (seq2) at random time points, and a time series reconstruction consistency loss function of the action sequence is introduced to calculate the continuity loss value L, which is shown in the following equation (1).

[0104] JPEG2025536026000002.jpg15170Using the method of the embodiments of this specification, the first test sequence is divided at a random time point to obtain a first test subsequence and a second test subsequence, a continuity loss value is calculated based on the first test subsequence and the second test subsequence, the parameters of the second encoder and the second decoder are adjusted based on the quantization loss value and the continuity loss value, and the step of returning to extracting the first training sample sequence from the third sample set is performed. When a second predetermined stopping condition is reached, the model parameters of the quantization generation model obtained by training are obtained, so that the generated task sequence meets the temporal continuity requirement, and the accuracy of the task processing model is further improved.

[0105] The motion generation method according to the embodiments of the present specification can be applied to different motion generation scenes, such as motion generation of objects such as virtual humans and virtual animals in game scenes, and motion generation of virtual customers in e-commerce scenes, and the specifics are selected according to the actual situation, and the embodiments of the present specification are not limited thereto.

[0106] Hereinafter, the data processing method of the task processing model according to the present specification will be further described with reference to Fig. 5, taking the application of the data processing method of the task processing model according to the present specification in the field of virtual character animation generation as an example. Referring to Fig. 5, Fig. 5 shows a flowchart of the virtual character animation generation method according to one embodiment of the present specification, specifically including the following steps 502 to 508.

[0107] In step 502, a virtual character animation generation request sent by the front end is received, and the virtual character animation generation request includes animation guide information.

[0108] In step 504, the video guide information and the entire mask video sequence are input into a virtual character video generation model, and the virtual character video characteristics corresponding to the video guide information are obtained through processing by the virtual character video generation model.

[0109] The virtual character animation generation model is obtained by training based on a plurality of sample animation guide information, sample animation sequences, and sample animation features, and the sample animation features are obtained by quantizing and encoding the sample animation sequences.

[0110] In step 506, the virtual character animation features are quantized and decoded to obtain the virtual character action sequence corresponding to the video guide information.

[0111] In step 508, a virtual character animation is generated based on the virtual character action sequence and sent to the front end, thereby causing the front end to display the virtual character animation.

[0112] Specifically, video guide information is information for guiding the generation of virtual character video, and video guide information includes, but is not limited to, video guide text information, video guide image information, video guide audio information, video guide video information, and video guide trajectory information, and is specifically selected according to actual situations, and the examples in this specification are not limited in this regard.

[0113] The implementation of steps 502, 504, and 506 is the same as the implementation of steps 404, 406, and 408 described above, and will not be described in detail in the embodiments of this specification.

[0114] Furthermore, after acquiring the virtual character action sequence corresponding to the video guide information, the virtual character video can be generated by controlling the skeleton motion of the virtual character based on the time order of the virtual character action sequence.

[0115] Using the solution of the embodiments of the present specification, a virtual character animation generation request sent by a front end is received, the virtual character animation generation request includes animation guide information, the animation guide information and the full mask animation sequence are input into a virtual character animation generation model, and the virtual character animation generation model processes to obtain virtual character animation features corresponding to the animation guide information, the virtual character animation generation model is trained and obtained based on a plurality of sample animation guide information, sample animation sequences, and sample animation features, the sample animation features are obtained by quantizing and encoding the sample animation sequences, the virtual character animation features are quantized and decoded to obtain virtual character movement sequences corresponding to the animation guide information, and a virtual character animation is generated based on the virtual character movement sequences and sent to the front end, and the virtual character animation is displayed on the front end. Since the virtual character animation generation model is trained and obtained based on a plurality of sample animation guide information, sample animation sequences, and sample animation features, and the sample animation features are obtained by quantizing and encoding the sample animation sequences, the virtual character animation generation model can efficiently and accurately generate virtual character movement sequences, and accurate virtual character animation can be generated based on the virtual character movement sequences.

[0116] In one alternative embodiment of the present specification, before the step of inputting the video guide information and the entire mask video sequence into a virtual character video generation model, and obtaining the virtual character video features corresponding to the video guide information through processing of the virtual character video generation model, The method may further include receiving designated presentation information input by a user, the designated presentation information including designated presentation text and / or designated presentation audio; The step of inputting the video guide information and the entire mask video sequence into a virtual character video generation model, and obtaining the virtual character video characteristics corresponding to the video guide information through processing of the virtual character video generation model, The method may include a step of inputting the specified presentation information, video guide information, and all mask video sequences into a virtual character video generation model, and obtaining virtual character video features corresponding to the video guide information by processing the virtual character video generation model.

[0117] Specifically, the designated presentation information is presentation information for video guide information, and the designated presentation information may be in text format, audio format, or of course, a combination of text and audio format. Specifically, it is selected according to the actual situation, and the examples in this specification are not limited thereto.

[0118] For example, assuming that the video guide information is a video guide image, the specified presentation information may be text that presents the video guide image, such as "The character in the video guide image is dancing."

[0119] Using the solution of the embodiments of this specification, specific presentation information input by a user is received, and the specific presentation information includes specific presentation text and / or specific presentation audio. The specific presentation information, video guide information, and the entire mask video sequence are input into a virtual character animation generation model, and the virtual character animation features corresponding to the video guide information are obtained through processing by the virtual character animation generation model, so as to improve the accuracy of the virtual character animation features.

[0120] In one alternative embodiment of the present specification, after the above step of quantizing and decoding the virtual character animation features to obtain the virtual character action sequence corresponding to the animation guide information, receiving scene information for a current scene input by a user; The method can further include adjusting the virtual character action sequence using the virtual character animation generation model based on the scene information, and obtaining the adjusted virtual character action sequence.

[0121] Specifically, the scene information is an animated scene in a virtual character animation including background information, weather information, geographic information, etc., and is specifically selected according to the actual situation, and the embodiments of this specification are not limited thereto.

[0122] In addition, when adjusting a virtual character movement sequence using a virtual character animation generation model based on scene information, the scene information and animation guide information can be input into the encoder of the virtual character animation generation model to obtain scene coding features and animation guide features, and then the scene coding features, animation guide features and full mask animation sequence can be input into the decoder of the virtual character animation generation model to obtain the adjusted virtual character movement sequence, and finally the adjusted virtual character animation features can be quantized and decoded to obtain the adjusted virtual character movement sequence.

[0123] The solution of the embodiments of this specification receives current scene information input by a user, and adjusts the virtual character's motion sequence based on the scene information using a virtual character animation generation model to obtain the adjusted virtual character's motion sequence. By adjusting the virtual character's motion sequence based on the current scene information, the virtual character's motion sequence becomes more realistic, and the realism of the virtual character's animation is further improved.

[0124] In one alternative embodiment of this specification, the virtual character animation features can be input to the second decoder of the pre-trained quantization generation model, and the second decoder can decode the virtual character action sequence corresponding to the video guide information.Furthermore, in order to generate a more accurate virtual character action sequence that meets the user's needs, the user can input target trajectory information and use the target trajectory information to control the direction of action generation, that is, the above steps of quantizing and decoding the virtual character animation features and obtaining the virtual character action sequence corresponding to the video guide information can be: receiving user-entered target trajectory information; The method may include a step of inputting the virtual character animation features and target trajectory information into a second decoder of the quantized generation model, and obtaining a virtual character action sequence corresponding to the animation guide information through a decoding process of the second decoder, wherein the quantized generation model is obtained by training based on a plurality of training sample sequences.

[0125] The target trajectory information can be expressed as the relative displacement between two adjacent frames. The quantized generative model learns a codebook, such as a text dictionary, to convert continuous motion sequences into discrete motion features. It can also control the direction of motion generation by trajectory decoupling based on input of different trajectory information. The virtual character motion sequence includes motion-related parameters and can drive the virtual character. For example, a T×N motion sequence is used, where T is the number of motion frames and N is the number of motion skeletons. Each skeleton in each frame is represented by a 6D rotation, i.e., the motion sequence is expressed in T×N×6 dimensions.

[0126] Using the solution of the embodiments of this specification, target trajectory information input by a user is received, and the virtual character video features and target trajectory information are input into a second decoder of the quantization generation model, and the second decoder decodes to obtain a virtual character action sequence corresponding to the video guide information, and the quantization generation model is trained based on multiple training sample sequences. Because the motion trajectory of a movement is often related to the orientation of a virtual object, data augmentation can be performed by simultaneously rotating the orientation and trajectory of the virtual object, allowing the model to learn related information such as the orientation of the movement from the trajectory information, ensuring that a virtual character action sequence that meets the user's needs is generated and improving the accuracy of the virtual character video.

[0127] Referring to Figure 6, Figure 6 shows a schematic diagram of the processing process of a virtual character animation generation method according to one embodiment of the present specification. As shown in Figure 6, a user inputs video guide information, which includes an action trajectory (3D displacement), an action type (text), some actions (6D display and 3D position of the action of key frames), text description, image, and any one of various combinations of images, text description, and trajectory. Each video guide information is input into a virtual character animation generation model, and the virtual character animation generation model processes the video guide information to obtain virtual character animation features corresponding to the video guide information. The virtual character animation generation model is trained and obtained based on multiple sample video guide information, sample animation sequences, and sample animation features. The sample animation features are obtained by quantizing and encoding the sample animation sequences. The virtual character animation features are quantized and decoded to obtain a virtual character action sequence corresponding to the video guide information.

[0128] By using the solutions of the embodiments of this specification to convert different modal video guide information into an integrated format and simultaneously accept the integrated input of different information in the learning process, the integration of multiple multimodal virtual character video generation tasks can be realized and the accuracy of the virtual character movement sequences can be improved.

[0129] Referring to FIG. 7, FIG. 7 shows a schematic diagram of a virtual character animation generation interface according to one embodiment of the present disclosure.

[0130] The virtual character animation generation interface includes a video guide information upload box, a "confirm" control, a "cancel" control, and a virtual character animation display box. When the user uploads video guide information through the video guide information upload box displayed on the front end and clicks the "confirm" control, the server side inputs the video guide information into the virtual character animation generation model, obtains a virtual character action sequence corresponding to the video guide information through processing of the virtual character animation generation model, generates a virtual character animation based on the virtual character action sequence, sends it to the front end, and displays the virtual character animation corresponding to the video guide information in the virtual character animation display box of the front end.

[0131] The manner in which the user operates the control may include any one of clicking, double-clicking, touching, mouse hovering, sliding, long pressing, voice control, shaking, etc., and may be selected according to the actual situation, and the embodiments of this specification are not limited thereto.

[0132] Referring to Figure 8, Figure 8 shows a flowchart of a data processing method for a quantization generation model according to one embodiment of the present specification, where the quantization generation model includes a second encoder and a second decoder, and the data processing method for the quantization generation model is applied to a cloud device, and specifically includes the following steps 802 to 810.

[0133] In step 802, a third sample set is obtained, where the third sample set includes a plurality of training sample sequences.

[0134] In step 804, the training sample sequence is input to a second encoder to obtain test motion features.

[0135] In step 806, the test motion characteristics and predetermined trajectory information are input to a second decoder to obtain a test motion sequence.

[0136] In step 808, a quantized generative model is trained based on the training sample sequence and the test motion sequence, and when a third predetermined stopping condition is reached, model parameters of the trained quantized generative model are obtained.

[0137] In step 810, the model parameters of the trained quantized generative model are transmitted to the end device.

[0138] The specific implementation of steps 802 to 808 is the same as the implementation of data processing in the task processing model according to FIG. 4, and will not be described in detail in the embodiments of this specification.

[0139] In actual use, after the cloud device transmits the model parameters of the quantized generative model to the end device, the end device can construct the quantized generative model based on the model parameters of the quantized generative model, thereby realizing behavior generation on the end side.

[0140] It should be noted that, since the model of the quantized generative model is small, the data processing method of the quantized generative model can be executed by an end device.

[0141] In accordance with the solution of the embodiments of the present specification, the cloud device acquires a third sample set, the third sample set including a plurality of training sample sequences, inputs the training sample sequence into a second encoder to acquire test motion features, inputs the test motion features and predetermined trajectory information into a second decoder to acquire the test motion sequence, trains a quantized generative model based on the training sample sequence and the test motion sequence, and when a third predetermined stopping condition is reached, acquires model parameters of the trained quantized generative model, and transmits the model parameters of the trained quantized generative model to the end device. The model parameters of the quantized generative model are continuously adjusted to make the model parameters of the finally acquired quantized generative model more accurate.

[0142] Referring to Figure 9, Figure 9 shows a flowchart of training a task processing model and a quantization generation model according to one embodiment of the present specification. Taking the action generation task as an example, the task processing model is an action generation model, and specifically includes:

[0143] To train the quantized generative model, the network learns a codebook. The network first converts the multimodal training sample sequence into several test motion features (embeddings), where the number of test motion features is less than the number of original frames multiplied by the number of keypoints. Then, each test motion feature is discretized into the latest embedding in the codebook according to its similarity. Finally, the network decodes the embedding corresponding to the codebook and the predetermined trajectory information to obtain the test motion sequence. The entire process is optimized based on the error between the training sample sequence and the test motion sequence, and the distance between the latest embedding in the codebook and the test motion feature.

[0144] Note that the second encoder of the quantized generative model is used only in the training stage, and the output of the second encoder is a label of the output of the first decoder of the motion generative model.

[0145] For training the action generation model, multimodal sample guide information is input into a first encoder of the action generation model to obtain sample guide features corresponding to the sample guide information, the original action sequence is masked, the sample guide features and the original action sequence after masking are input into a first decoder of the action generation model to obtain a decoded action sequence, the decoded action sequence is further discretized to obtain predicted action features, the first decoder is optimized based on the error between the predicted action features and the sample action features corresponding to the original action sequence to obtain an action generation model, and the first encoder and the first decoder are connected via a cross-attention layer.

[0146] The training process for the movement generation model includes a degeneration and restoration process. The degeneration process uses a masking method to train multimodal sample guide information together. If there is a gap in the sample guide information for a certain modality, the corresponding position is replaced with a mask token to perform the predicted masking. Each time, a different step is randomized to correspond to a different mask probability distribution and masking is performed according to the corresponding mask probability. The model is then trained to predict the original movement sequence. An additional mask head is then added and the model prediction result is reinput into the network to allow the model to predict the masked position. In the testing phase, various movement generation tasks are completed by adjusting different input conditions. Initially, a full mask sequence is input, and each prediction is masked based on the mask head structure for each prediction, with the final prediction result being output. Specifically, the iterative process is as follows: first, the original operation sequence of all mask tokens is input to the first decoder, and prediction is performed to obtain the complete sequence features of the first step; then, the second decoder obtains the test operation sequence from the complete sequence features; the test operation sequence is input to the first decoder to predict the mask position and perform masking; then, the complete sequence features of the first step after masking and the timestamp of the second step are input to the first decoder to re-predict (i.e., recover), and the iteration is repeated until completion.

[0147] When actually used, the motion generation model converts the sequence into an embedding position in the codebook, i.e., predicts a discretized motion sequence, and when discretizing the decoded motion sequence, specifically, the motion distances of the X and Y axes included in the motion of each step are directly discretized to obtain predicted motion features. For example, the range of 0 to 1 meter is divided into 200 sections, and 0 to 0.005 meters is used as the first predicted motion feature, and the motion feature prediction based on the discrete motion corresponds to the learned codebook.

[0148] Using the method described in this specification, action sequences are converted into discrete action feature sequences, allowing different control signals to be represented as seq2seq. Differences in different action guide information are replaced with mask tokens, allowing associations with action sequences to be learned from multimodal signals. Converting different modal action guide information into a unified representation ensures improved data volume, supports model scaling, and avoids model overfitting in multitasking. Using a recursive mask modeling method, multiple mask shrinkage and restoration processes are learned at the discrete code level. The mask position is predicted and sampled by the network, ensuring consistency between training and testing in the generation task compared to the method of adding random noise to a diffusion model. Cross-modal learning allows the model to exhibit a degree of migration and zero-shot learning, enabling joint control of multimodal input information.

[0149] In addition, the information and data such as sample guide information, sample task sequence, training sample sequence, predetermined presentation information, video guide information, full mask video sequence, designated presentation information, scene information, target trajectory information, etc. in the above-mentioned method embodiments are all information and data approved by the user or fully approved by each party, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entries are provided for the user to approve or reject.

[0150] The present specification further provides an embodiment of a data processing device of a task processing model, corresponding to the embodiment of the data processing method of the above task processing model executed by a cloud device. FIG. 10 shows a schematic diagram showing the structure of a data processing device of a task processing model according to one embodiment of the present specification, the device is applied to a cloud device connected to multiple cloud devices, and as shown in FIG. 10, the device: an acquisition module 1002 configured to acquire a first sample set, the first sample set including multimodal sample guide information; a first input module 1004 configured to input sample guide information and a sample task sequence into the initial processing model to obtain predicted task features corresponding to the sample guide information; a training module 1006 configured to train an initial processing model based on the predicted task features and sample task features corresponding to the sample task sequence, and to obtain model parameters of the trained processing model when a first predetermined stopping condition is reached, wherein the sample task features are obtained by quantizing and encoding the sample task sequence; and a transmission module 1008 configured to transmit the model parameters of the task processing model obtained by training to the end device.

[0151] Optionally, the initial processing model includes a first encoder and a first decoder, the first predetermined stopping condition includes a first stopping sub-condition, the first input module 1004 is further configured to obtain first sample guide information from the first sample set, where the first sample guide information is any one of the sample guide information in the first sample set, input the first sample guide information to a pre-trained first encoder, obtain first sample guide features, mask the sample task sequence, obtain the masked task sequence, input the first sample guide features and the masked task sequence to the first decoder, and obtain first predicted task features, the training module 1006 is further configured to calculate a decoding loss value based on the first predicted task features and the sample task features corresponding to the sample task sequence, adjust parameters of the first decoder based on the decoding loss value, and perform the steps of returning to extract the first sample guide information from the first sample set, and obtain the trained first decoder if the first stopping sub-condition is reached.

[0152] Optionally, the first predetermined stopping condition includes a second stopping sub-condition, and the first input module 1004 is further configured to perform the steps of obtaining a second sample set, the second sample set including multimodal sample guide information, the sample guide information having sample guide features, extracting second sample guide information from the second sample set, the second sample guide information being any one of the sample guide information in the second sample set, inputting the second sample guide information into a first encoder, obtaining first predicted guide features corresponding to the second sample guide information, calculating an encoding loss value based on the first predicted guide feature and the second sample guide feature included in the second sample guide information, adjusting parameters of the first encoder based on the encoding loss value, and returning to extract the second sample guide information from the second sample set, and obtaining a trained first encoder if the second stopping sub-condition is reached.

[0153] Optionally, the first input module 1004 is further configured to input the sample task sequence into a second encoder of the pre-trained quantized generative model, and obtain sample task features corresponding to the sample task sequence through the encoding process of the second encoder.

[0154] Optionally, the first input module 1004 is further configured to acquire predetermined time step features, where the time step features and the number of training iterations correspond one-to-one, and input the time step features, the first sample guide features, and the mask task sequence to a first decoder to acquire first prediction task features.

[0155] Optionally, the quantized generative model includes a second encoder and a second decoder, and the apparatus further includes a quantized generative model training module, wherein the quantized generative model training module is configured to obtain a third sample set, the third sample set including a plurality of training sample sequences, extract a first training sample sequence from the third sample set, the first training sample sequence being any one of the training sample sequences in the third sample set, input the first training sample sequence to the second encoder, obtain a first test feature, input the first test feature and predetermined presentation information to the second decoder, obtain a first test sequence, calculate a quantization loss value based on the first training sample sequence and the first test sequence, adjust parameters of the second encoder and the second decoder based on the quantization loss value, and return to extract the first training sample sequence from the third sample set; and when a second predetermined stopping condition is reached, obtain model parameters of the quantized generative model obtained by training, and send the model parameters of the quantized generative model obtained by training to the end device.

[0156] Optionally, the quantized generative model training module is further configured to perform the steps of splitting the first test sequence at a random time point to obtain a first test subsequence and a second test subsequence, calculating a continuity loss value based on the first test subsequence and the second test subsequence, adjusting parameters of the second encoder and the second decoder based on the quantization loss value and the continuity loss value, and returning to extract the first training sample sequence from the third sample set; and obtaining model parameters of the trained quantized generative model when a second predetermined stopping condition is reached.

[0157] Using the solution of the embodiments of the present specification, a first sample set is obtained, the first sample set including multimodal sample guide information, the sample guide information and the sample task sequence are input into an initial processing model, predicted task features corresponding to the sample guide information are obtained, the initial processing model is trained based on the predicted task features and the sample task features corresponding to the sample task sequence, and when a first predetermined stopping condition is reached, model parameters of the processing model obtained by training are obtained, the sample task features are obtained by quantizing and encoding the sample task sequence, and the model parameters of the trained task processing model are sent to the end device. Because the task processing model is trained based on the multimodal sample guide information, multimodal task integration can be achieved and the accuracy and versatility of the model can be improved.

[0158] The above is a schematic solution of the task processing model data processing device in this embodiment, and the technical solution of the task processing model data processing device belongs to the same idea as the technical solution of the task processing model data processing method executed by the cloud device, and for details not explained in detail in the technical solution of the task processing model data processing device, please refer to the description of the technical solution of the task processing model data processing method executed by the cloud device.

[0159] This specification further provides an embodiment of a data processing device of a task processing model, corresponding to the embodiment of the data processing method of the above task processing model executed by an end device. Figure 11 shows a schematic diagram showing the structure of another data processing device of a task processing model according to one embodiment of this specification, which is executed by an end device connected to a cloud device. As shown in Figure 11, the device includes: a construction module 1102 configured to receive model parameters of a task processing model sent by a cloud device and construct a task processing model based on the model parameters; a first receiving module configured to receive a task processing request input by a user, the task processing request including task guide information 1104; a second input module 1106 configured to input task guide information and the entire mask task sequence into a task processing model and obtain target task features corresponding to the task guide information through processing by the task processing model, wherein the task processing model is obtained by training based on multimodal sample guide information, the sample task sequence, and the sample task features, and the sample task features are obtained by quantizing and encoding the sample task sequence; a first decoding module 1108 configured to quantize and decode the target task features to obtain task processing results corresponding to the task guide information.

[0160] Optionally, the task processing model includes a first encoder and a first decoder, and the second input module 1106 is further configured to input task guide information to the first encoder to obtain task guide features corresponding to the task guide information, and input the task guide features and the entire mask task sequence to the first decoder to obtain target task features corresponding to the task guide information.

[0161] Using the solution of the embodiments of the present specification, a task processing model is received as a model parameter of a task processing model sent by a cloud device, a task processing model is constructed based on the model parameter, a task processing request input by a user is received, the task processing request includes task guide information, the task guide information and the entire mask task sequence are input into the task processing model, and target task features corresponding to the task guide information are obtained through processing by the task processing model, the task processing model is trained and obtained based on multimodal sample guide information, sample task sequence, and sample task features, the sample task features are obtained by quantizing and encoding the sample task sequence, and the target task features are quantized and decoded to obtain a task processing result corresponding to the task guide information. Because the task processing model is trained and obtained based on multimodal sample guide information, multimodal task integration can be realized, and the task processing model can efficiently and accurately generate target task features, further improving the accuracy of the task processing result.

[0162] The above is an exemplary solution of the task processing model data processing device of this embodiment, and the technical solution of the task processing model data processing device belongs to the same idea as the technical solution of the task processing model data processing method executed by the end device, and for details not explained in detail in the technical solution of the task processing model data processing device, please refer to the description of the technical solution of the task processing model data processing method executed by the end device.

[0163] This specification further provides an embodiment of a virtual character animation generation device corresponding to the embodiment of the virtual character animation generation method described above, and Figure 12 shows a schematic diagram illustrating the structure of a virtual character animation generation device according to one embodiment of this specification. As shown in Figure 12, the device: a second receiving module 1202 configured to receive a virtual character animation generation request sent by the front end, the virtual character animation generation request including animation guide information; a third input module 1204 configured to input the video guide information and the full mask video sequence into a virtual character video generation model, and obtain the virtual character video features corresponding to the video guide information through processing of the virtual character video generation model, wherein the virtual character video generation model is obtained by training based on a plurality of sample video guide information, sample video sequences and sample video features, and the sample video features are obtained by quantizing and encoding the sample video sequences; a first decoding module 1206 configured to quantize and decode the virtual character animation features to obtain a virtual character action sequence corresponding to the video guide information; and a generating module 1208 configured to generate a virtual character animation based on the virtual character action sequence and send it to the front end, thereby causing the front end to display the virtual character animation.

[0164] Optionally, the device further includes a fourth input module configured to receive specified presentation information input by a user, the specified presentation information including specified presentation text and / or specified presentation audio, and a third input module 1204 further configured to input the specified presentation information, video guide information, and the entire mask video sequence into a virtual character animation generation model, and obtain virtual character animation features corresponding to the video guide information by processing the virtual character animation generation model.

[0165] Optionally, the device further includes an adjustment module configured to receive scene information of a current scene input by a user, the adjustment module configured to adjust a virtual character action sequence using a virtual character animation generation model based on the scene information to obtain an adjusted virtual character action sequence.

[0166] Optionally, the first decoding module 1206 is further configured to receive target trajectory information input by a user, input the virtual character animation features and the target trajectory information into a second decoder of the quantized generation model, and obtain a virtual character action sequence corresponding to the animation guide information through the decoding process of the second decoder, and the quantized generation model is obtained by training based on a plurality of training sample sequences.

[0167] Using the solution of the embodiments of the present specification, a virtual character animation generation request sent by a front end is received, the virtual character animation generation request includes animation guide information, the animation guide information and the full mask animation sequence are input into a virtual character animation generation model, and the virtual character animation generation model processes to obtain virtual character animation features corresponding to the animation guide information, the virtual character animation generation model is trained and obtained based on a plurality of sample animation guide information, sample animation sequences, and sample animation features, the sample animation features are obtained by quantizing and encoding the sample animation sequences, the virtual character animation features are quantized and decoded to obtain virtual character movement sequences corresponding to the animation guide information, and a virtual character animation is generated based on the virtual character movement sequences and sent to the front end, and the virtual character animation is displayed on the front end. Since the virtual character animation generation model is trained and obtained based on a plurality of sample animation guide information, sample animation sequences, and sample animation features, and the sample animation features are obtained by quantizing and encoding the sample animation sequences, the virtual character animation generation model can efficiently and accurately generate virtual character movement sequences, and accurate virtual character animation can be generated based on the virtual character movement sequences.

[0168] The above is an exemplary solution of the virtual character animation generation device of this embodiment. Note that the technical solution of the virtual character animation generation device is based on the same idea as the technical solution of the virtual character animation generation method described above, and for details not described in detail in the technical solution of the virtual character animation generation device, please refer to the description of the technical solution of the virtual character animation generation method described above.

[0169] 13 shows a block diagram of a computing device according to one embodiment of the present disclosure. Components of the computing device 1300 include, but are not limited to, a memory 1310 and a processor 1320. The processor 1320 and the memory 1310 are connected via a bus 1330, and a database 1350 stores data.

[0170] Computing device 1300 further includes access device 1340, which enables computing device 1300 to communicate over one or more networks 1360. Examples of these networks include a combination of communication networks, such as a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or the Internet. Access device 1340 may include one or more of any type of network interface, wired or wireless (e.g., a network interface card (NIC)), such as an IEEE 802.11 wireless local area network (WLAN) radio interface, a Wi-MAX (World Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth (Bluetooth is a registered trademark) interface, a near field communication (NFC) interface, etc.

[0171] In one embodiment of the present specification, the above components of computing device 1300, as well as other components not shown in Figure 13, may be connected to each other, for example, via a bus. It should be understood that the block diagram of the structure of a computing device shown in Figure 13 is for illustrative purposes only and is not intended to limit the scope of the present specification. Those skilled in the art may add or substitute other components as desired.

[0172] Computing device 1300 may be any type of fixed or mobile computing device, including a mobile computer or mobile computing device (e.g., tablet computer, personal digital assistant, laptop computer, notebook computer, netbook, etc.), a mobile phone (e.g., smartphone), a wearable computing device (e.g., smart watch, smart glasses, etc.), or other type of mobile device, or a fixed computing device such as a desktop computer or personal computer (PC). Computing device 1300 may also be a mobile or fixed server.

[0173] The processor 1320 is used to execute computer-executable instructions that, when executed by the processor, implement the steps of the data processing method or the virtual character animation generation method of the task processing model.

[0174] The above is an exemplary solution of the computing device in this embodiment, and the technical solution of the computing device belongs to the same idea as the technical solutions of the data processing method of the task processing model and the virtual character animation generation method described above, and for details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the data processing method of the task processing model or the virtual character animation generation method described above.

[0175] One embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement steps of the data processing method or virtual character animation generation method of the above task processing model.

[0176] The above is an exemplary solution of a computer-readable storage medium in this embodiment. The technical solution of the storage medium is based on the same concept as the technical solutions of the data processing method for task processing model and the virtual character animation generation method described above. For details not described in detail in the technical solution of the storage medium, please refer to the descriptions of the technical solutions of the data processing method for task processing model and the virtual character animation generation method described above.

[0177] Furthermore, one embodiment of the present specification further provides a computer program that, when executed on a computer, causes the computer to execute steps of the data processing method or the virtual character animation generation method of the above task processing model.

[0178] The above is an exemplary solution of the computer program of this embodiment. The technical solution of the computer program is based on the same concept as the technical solutions of the data processing method of the task processing model and the virtual character animation generation method described above. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solutions of the data processing method of the task processing model or the virtual character animation generation method described above.

[0179] The foregoing describes specific embodiments of the present specification. Other embodiments are within the scope of the following claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the examples and still achieve desirable results. Furthermore, processes depicted in the figures do not necessarily require the particular order shown or sequential order, as desired results cannot be achieved. In some embodiments, multitasking and parallel processing may also be possible or advantageous.

[0180] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or any intermediate form, etc. The computer-readable medium may include any entity or device capable of having the computer program code, such as a recording medium, a U-disk, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier wave signal, an electric communication signal, and a software distribution medium.

[0181] Although the above-described embodiments of the methods are described as a combination of a series of operations for ease of explanation, those skilled in the art should understand that the embodiments herein are not limited by the order of the operations described, and that certain steps may be performed in other orders or simultaneously according to the embodiments of the present specification. Next, those skilled in the art should understand that the embodiments described herein are all preferred embodiments, and that the operations and modules described are not necessarily required for the embodiments of the present specification.

[0182] In the above embodiments, the description of each embodiment has its own emphasis, and for the parts not described in detail in one embodiment, reference can be made to the relevant descriptions of other embodiments.

[0183] The preferred embodiments disclosed herein are provided merely to aid in the description of the present specification. The selectable embodiments do not include all details and are not intended to limit the invention to the specific embodiments. Obviously, various modifications and variations can be made based on the content of the embodiments described herein. The present specification specifically describes and selects these embodiments to better understand the principles and practical applications of the embodiments described herein, thereby enabling those skilled in the art to fully understand and utilize the present specification. The present specification is limited only by the claims, the full scope of which is defined below, and equivalents thereto.

Claims

1. A data processing method for a task processing model executed by a cloud device connected to a plurality of end devices, comprising: obtaining a first sample set, the first sample set including multimodal sample guide information; inputting the sample guide information and a sample task sequence into an initial processing model to obtain predicted task features corresponding to the sample guide information; training the initial process model based on the predicted task features and sample task features corresponding to the sample task sequence, and obtaining model parameters of the trained process model when a first predetermined stopping condition is reached, wherein the sample task features are obtained by quantizing and encoding the sample task sequence; transmitting model parameters of the task processing model obtained by the training to an end device; Including, Data processing methods for task processing models.

2. the initial processing model includes a first encoder and a first decoder, and the first predetermined stopping condition includes a first stopping sub-condition; The step of inputting the sample guide information and the sample task sequence into an initial processing model and obtaining predicted task features corresponding to the sample guide information includes: A step of acquiring first sample guide information from the first sample set, wherein the first sample guide information is any one of the sample guide information in the first sample set; inputting the first sample guide information into a pre-trained first encoder to obtain first sample guide features; masking the sample task sequence to obtain a mask task sequence; inputting the first sample guide feature and the mask task sequence into the first decoder to obtain a first predicted task feature; Including, the step of training the initial process model based on the predicted task features and sample task features corresponding to the sample task sequence, and obtaining model parameters of the trained process model when a first predetermined stopping condition is reached, calculating a decoding loss value based on the first predicted task features and sample task features corresponding to the sample task sequence; According to the decoding loss value, adjust the parameters of the first decoder, and return to the step of extracting first sample guide information from the first sample set, and obtain a trained first decoder if a first stopping sub-condition is reached; Including, The method of claim 1.

3. the first predetermined stopping condition includes a second stopping sub-condition; The training method of the first encoder is obtaining a second sample set, the second sample set including multimodal sample guide information, the sample guide information including sample guide features; A step of extracting second sample guide information from the second sample set, wherein the second sample guide information is any one of the sample guide information in the second sample set; inputting the second sample guide information into the first encoder to obtain first predicted guide features corresponding to the second sample guide information; calculating a coding loss value based on the first predicted guide feature and the second sample guide feature included in the second sample guide information; adjusting parameters of the first encoder based on the encoding loss value, and returning to the step of extracting second sample guide information from the second sample set, and obtaining a trained first encoder if a second stopping sub-condition is reached; Including, The method of claim 2.

4. prior to the step of calculating a decoding loss value based on the first predicted task features and sample task features corresponding to the sample task sequence, inputting a sample task sequence into a second encoder of the pre-trained quantized generative model, and obtaining sample task features corresponding to the sample task sequence through an encoding process of the second encoder; The method of claim 2.

5. inputting the first sample guide feature and the mask task sequence into the first decoder to obtain a first predicted task feature, acquiring predetermined time step features, wherein the time step features and the number of training iterations correspond to each other one-to-one; inputting the time step features, the first sample guide features, and the mask task sequence into the first decoder to obtain first predicted task features; Including, The method of claim 2.

6. the quantized generative model includes a second encoder and a second decoder; The training method of the quantized generative model is as follows: obtaining a third sample set, the third sample set including a plurality of training sample sequences; extracting a first training sample sequence from the third sample set, the first training sample sequence being any one of the training sample sequences of the third sample set; inputting the first training sample sequence into the second encoder to obtain first test features; inputting the first test feature and predetermined presentation information into the second decoder to obtain a first test sequence; calculating a quantization loss value based on the first training sample sequence and the first test sequence; adjusting parameters of the second encoder and the second decoder based on the quantization loss value, and returning to the step of extracting a first training sample sequence from the third sample set, and obtaining model parameters of the trained quantized generative model when a second predetermined stopping condition is reached; transmitting model parameters of the trained quantized generative model to an end device; Including, The method of claim 4.

7. After the step of inputting the first test feature and predetermined presentation information into the second decoder to obtain a first test sequence, splitting the first test sequence at a random time to obtain a first test subsequence and a second test subsequence; calculating a continuity loss value based on the first test subsequence and the second test subsequence; further comprising adjusting parameters of the second encoder and the second decoder based on the quantization loss value, returning to extracting a first training sample sequence from the third sample set, and obtaining model parameters of the trained quantized generative model when a second predetermined stopping condition is reached, adjusting parameters of the second encoder and the second decoder based on the quantization loss value and the continuity loss value, and returning to the step of extracting a first training sample sequence from the third sample set, and obtaining model parameters of the trained quantized generative model when a second predetermined stopping condition is reached; The method of claim 6.

8. A data processing method for a task processing model executed by an end device connected to a cloud device, comprising: receiving model parameters of a task processing model sent by the cloud device, and constructing a task processing model based on the model parameters; receiving a task processing request input by a user, the task processing request including task guide information; inputting the task guide information and the entire mask task sequence into the task processing model, and acquiring target task features corresponding to the task guide information through processing by the task processing model, wherein the task processing model is obtained by training based on multimodal sample guide information, a sample task sequence, and sample task features, and the sample task features are obtained by quantizing and encoding the sample task sequence; quantizing and decoding the target task features to obtain a task processing result corresponding to the task guide information; Including, Data processing methods for task processing models.

9. the task processing model includes a first encoder and a first decoder; inputting the task guide information and the entire mask task sequence into a task processing model, and obtaining target task features corresponding to the task guide information through processing by the task processing model; inputting the task guide information into the first encoder to obtain task guide features corresponding to the task guide information; inputting the task guide features and the entire mask task sequence into the first decoder to obtain target task features corresponding to the task guide information; Including, The method of claim 8.

10. receiving a virtual character animation generation request sent by the front end, the virtual character animation generation request including animation guide information; a step of inputting the video guide information and the entire mask video sequence into a virtual character video generation model, and acquiring virtual character video features corresponding to the video guide information by processing the virtual character video generation model, wherein the virtual character video generation model is obtained by training based on a plurality of sample video guide information, sample video sequences and sample video features, and the sample video features are obtained by quantizing and encoding the sample video sequences; quantizing and decoding the virtual character animation features to obtain a virtual character action sequence corresponding to the video guide information; generating a virtual character animation based on the virtual character action sequence, transmitting the generated virtual character animation to the front end, and displaying the virtual character animation on the front end; Including, A method for generating virtual character animation.

11. before the step of inputting the video guide information and the entire mask video sequence into a virtual character video generation model, and obtaining a virtual character video feature corresponding to the video guide information through processing of the virtual character video generation model; receiving designated presentation information input by a user, the designated presentation information including designated presentation text and / or designated presentation audio; The step of inputting the video guide information and the entire mask video sequence into a virtual character video generation model, and obtaining a virtual character video feature corresponding to the video guide information through processing of the virtual character video generation model, inputting the designated presentation information, video guide information, and all mask video sequences into a virtual character video generation model, and acquiring virtual character video features corresponding to the video guide information through processing of the virtual character video generation model; The method of claim 10.

12. After the step of quantizing and decoding the virtual character animation features to obtain a virtual character action sequence corresponding to the video guide information, receiving scene information for a current scene input by a user; adjusting the virtual character motion sequence using the virtual character animation generation model based on the scene information, and obtaining an adjusted virtual character motion sequence; further comprising:

12. The method according to claim 10 or 11.

13. The step of quantizing and decoding the virtual character animation features to obtain a virtual character action sequence corresponding to the animation guide information includes: receiving user-entered target trajectory information; inputting the virtual character animation features and the target trajectory information into a second decoder of a quantized generation model, and obtaining a virtual character action sequence corresponding to the animation guide information through a decoding process of the second decoder, wherein the quantized generation model is obtained by training based on a plurality of training sample sequences; Including, The method of claim 10.

14. an end device for constructing a first sample set and transmitting the first sample set to a cloud device, the first sample set including multimodal sample guide information; the cloud device inputs the sample guide information and the sample task sequence into an initial processing model, obtains predicted task features corresponding to the sample guide information, trains the initial processing model based on the predicted task features and sample task features corresponding to the sample task sequence, and, when a first predetermined stopping condition is reached, obtains model parameters of the processing model obtained by training, wherein the sample task features are obtained by quantizing and encoding the sample task sequence, and transmits the model parameters of the task processing model obtained by training to an end device; Including, Task processing model data processing system.

15. a memory and a processor, the memory is used to store computer-executable instructions; the processor is adapted to execute the computer-executable instructions; The computer-executable instructions, when executed by a processor, perform the steps of the method of any one of claims 1 to 7, or any one of claims 8 to 9, or any one of claims 10 to 13. Computing devices.

16. storing computer executable instructions which, when executed by a processor, implement the steps of the method of any one of claims 1 to 7, or any one of claims 8 to 9, or any one of claims 10 to 13; A computer-readable storage medium.

Citation Information

Patent Citations

  • Data processing system and method and electronic equipment

    CN115170840A