Animation generation method and device, computer equipment and storage medium
The speech text is analyzed through a large language model, and combined with the animation library and animation generation model, it generates semantic matching and rhythmic animations, which solves the problem of lack of content in the existing technology and improves the playback effect of animations.
Patent Information
- Application Number
- CN202311659321.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-06
AI Technical Summary
Although the animations generated by the prior art have a sense of rhythm, they lack rich content, resulting in poor playback effects.
The speech text of virtual characters is analyzed through a large language model, and the first text segment with semantic labels and the second text segment without semantic labels are divided. The matching animations are retrieved from the animation library and the rhythm matching animations are generated through the animation generation model to synthesize the target animation.
A semantic matching and rhythmic animation is generated, which enriches the amount of information in the animation and improves the playback effect of the animation.
Smart Images

Figure CN120107425A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technology, and in particular to an animation generation method, apparatus, computer equipment, and storage medium. Background Art
[0002] With the development of computer technology, animation needs to be generated in many fields such as games, cartoons, and movies. A method for generating animation based on audio data is proposed in the related art. Animation production personnel produce animations with a rhythm consistent with the rhythm in the audio data based on the audio data, and store the audio data and the animation in an animation library. Subsequently, the animation can be played synchronously while the audio data is played on the terminal.
[0003] However, animation producers only focus on making sure that the rhythm of the animation is consistent with the rhythm of the audio data. Although the generated animation has a sense of rhythm, it still lacks richer content, which results in poor animation playback effect. Summary of the invention
[0004] The embodiments of the present application provide an animation generation method, apparatus, computer equipment and storage medium, which generate a semantically matched and rhythmic animation, enrich the amount of information contained in the animation, and improve the animation playback effect. The technical solution is as follows:
[0005] In one aspect, a method for generating an animation is provided, the method comprising:
[0006] Performing semantic analysis on a speech text of a virtual character by using a large language model, determining a first text segment in the speech text and a first semantic tag corresponding to the first text segment, wherein the first semantic tag represents the semantics of the first text segment;
[0007] Retrieving a first animation matching the first semantic tag from an animation library, wherein the semantics expressed by the virtual character in the first animation matches the semantics represented by the first semantic tag;
[0008] Processing an audio segment corresponding to a second text segment through an animation generation model to obtain a second animation, wherein the second text segment is a text segment without a corresponding semantic tag in the spoken text, the second animation includes an animation corresponding to the second text segment, and a rhythm of the virtual character in the animation corresponding to the second text segment matches a speaking rhythm of the second text segment;
[0009] A target animation of the virtual character is generated based on the first animation and the second animation.
[0010] Optionally, before selecting the first text segment from the at least two third text segments based on the weights of the at least two third text segments, the method further includes:
[0011] Displaying the semantic labels and weights of the at least two third text segments in an animation generation interface;
[0012] In response to the editing operation in the animation generation interface, the semantic label and weight of the edited third text segment are obtained.
[0013] Optionally, the performing semantic analysis on the speech text of the virtual character by using the large language model to determine a first text segment in the speech text and a first semantic tag corresponding to the first text segment includes:
[0014] By using the large language model, semantic analysis is performed on the spoken text based on the prompt words to determine the first text segment and the first semantic label corresponding to the first text segment;
[0015] The prompt word represents a semantic analysis task sent to the large language model.
[0016] Optionally, the step of processing the audio segment corresponding to the second text segment by using the animation generation model to obtain the second animation includes:
[0017] Processing the speech audio corresponding to the speech text by the animation generation model to obtain the second animation, wherein the second animation includes an animation corresponding to the first text segment and an animation corresponding to the second text segment, and the rhythm of the virtual character in the second animation matches the speaking rhythm of the speech text;
[0018] The step of generating a target animation of the virtual character based on the first animation and the second animation includes:
[0019] The animation in the second animation corresponding to the first text segment is replaced with the first animation to obtain the target animation.
[0020] Optionally, before performing semantic analysis on the speech text of the virtual character by using the large language model to determine a first text segment in the speech text and a first semantic tag corresponding to the first text segment, the method further includes:
[0021] Obtaining the speech audio of the virtual character;
[0022] Converting the spoken audio into the spoken text;
[0023] After performing semantic analysis on the speech text of the virtual character by using the large language model to determine the first text segment in the speech text and the first semantic label corresponding to the first text segment, the method further includes:
[0024] An audio segment corresponding to the second text segment is extracted from the spoken audio.
[0025] Optionally, before performing semantic analysis on the speech text of the virtual character by using the large language model to determine a first text segment in the speech text and a first semantic tag corresponding to the first text segment, the method further includes:
[0026] Obtaining the speech text of the virtual character;
[0027] After performing semantic analysis on the speech text of the virtual character by using the large language model to determine the first text segment in the speech text and the first semantic label corresponding to the first text segment, the method further includes:
[0028] The second text segment is converted into an audio segment.
[0029] In another aspect, an animation generation device is provided, the device comprising:
[0030] A semantic analysis module, configured to perform semantic analysis on a speech text of a virtual character by using a large language model, and determine a first text segment in the speech text and a first semantic tag corresponding to the first text segment, wherein the first semantic tag represents the semantics of the first text segment;
[0031] A retrieval module, configured to retrieve a first animation matching the first semantic tag from an animation library, wherein the semantics expressed by the virtual character in the first animation matches the semantics represented by the first semantic tag;
[0032] a first animation generation module, configured to process an audio segment corresponding to a second text segment through an animation generation model to obtain a second animation, wherein the second text segment is a text segment without a corresponding semantic tag in the spoken text, the second animation includes an animation corresponding to the second text segment, and a rhythm of the virtual character in the animation corresponding to the second text segment matches a speaking rhythm of the second text segment;
[0033] The second animation generating module is used to generate a target animation of the virtual character based on the first animation and the second animation.
[0034] Optionally, the semantic analysis module includes:
[0035] a semantic analysis unit, configured to perform semantic analysis on the spoken text by using the large language model to obtain semantic labels and weights of at least two third text segments in the spoken text, wherein the weights represent the importance of the third text segments to the spoken text;
[0036] a text selection unit, configured to select the first text segment from the at least two third text segments based on weights of the at least two third text segments, wherein the weight of the first text segment is greater than the weight of the third text segment that is not selected;
[0037] A determination unit is used to determine the unselected third text segment as the second text segment.
[0038] Optionally, the semantic analysis module further includes:
[0039] A display unit, configured to display the semantic labels and weights of the at least two third text segments in an animation generation interface;
[0040] The editing unit is used to obtain the semantic label and weight of the edited third text segment in response to the editing operation in the animation generation interface.
[0041] Optionally, the semantic analysis module is used to perform semantic analysis on the spoken text based on the prompt words through the large language model to determine the first text segment and the first semantic label corresponding to the first text segment;
[0042] The prompt word represents a semantic analysis task sent to the large language model.
[0043] Optionally, the training process of the large language model includes:
[0044] Acquire a sample text, wherein the sample text includes semantic keywords;
[0045] Acquire a semantic label and a weight of the semantic keyword, wherein the weight indicates the importance of the semantic keyword to the sample text;
[0046] The large language model is trained based on the sample text, the semantic labels and weights of the semantic keywords.
[0047] Optionally, the retrieval module includes:
[0048] A retrieval unit, configured to retrieve a plurality of candidate animations matching the first semantic tag from the animation library;
[0049] A first animation selection unit, configured to select the first animation from the multiple candidate animations based on the playing time of the first text segment and the playing time of the multiple candidate animations;
[0050] The difference between the playing time of the first animation and the first text segment is less than the target threshold; or, the difference between the playing time of the first animation and the first text segment is less than the difference between the playing time of the other candidate animations and the first text segment.
[0051] Optionally, the retrieval module includes:
[0052] A retrieval unit, which retrieves a plurality of candidate animations matching the first semantic tag from the animation library;
[0053] An adjacent animation acquisition unit, used for acquiring an adjacent animation of the first text segment, wherein the adjacent animation is an animation of an adjacent text segment of the first text segment;
[0054] a difference determination unit, used to respectively determine the difference information between each candidate animation and the adjacent animation;
[0055] The second animation selection unit is used to select the first animation from the multiple candidate animations based on the difference information of the multiple candidate animations.
[0056] Optionally, the animation generation model includes a first encoder, a second encoder and a conditional autoregressive model; the first animation generation module includes:
[0057] A first feature extraction unit, configured to extract an audio feature of the audio segment through the first encoder, wherein the audio feature represents a speaking rhythm of the second text segment;
[0058] A second feature extraction unit is used to extract a style feature of the template style animation through the second encoder, wherein the style feature indicates the template style to which the template style animation belongs, and the virtual character in the template style animation performs an action belonging to the template style;
[0059] An animation generating unit is used to generate the second animation based on the audio feature and the style feature through the conditional autoregressive model, wherein the virtual character in the second animation performs an action belonging to the style of the template.
[0060] Optionally, the first feature extraction unit is used to:
[0061] Extracting a first audio feature of the audio segment, where the first audio feature includes at least one of the following: a logarithmic amplitude spectrum, a Mel frequency, or energy of each audio frame included in the audio segment;
[0062] Sampling the first audio feature;
[0063] The sampled first audio feature is processed by the first encoder to obtain a second audio feature of the audio segment.
[0064] Optionally, the second feature extraction unit is used to:
[0065] Extracting parameters of each joint of the virtual character from the template style animation;
[0066] splicing the parameters of each joint of the virtual character to obtain a first style feature;
[0067] Normalizing the first style feature;
[0068] The normalized first style feature is processed by the second encoder to obtain a second style feature of the template style animation.
[0069] Optionally, the animation generating unit is used to:
[0070] Acquire an adjacent animation of the audio segment, where the adjacent animation is an animation of an adjacent text segment of the second text segment;
[0071] The second animation is generated based on the audio feature, the style feature and the adjacent animation by using the conditional autoregressive model.
[0072] Optionally, the conditional autoregressive model includes a neural network and a recurrent decoder, the recurrent decoder includes at least two layers of gated recurrent neural units, and the animation generation unit is used to:
[0073] Processing the style features and the adjacent animations through the neural network to obtain hidden layer features;
[0074] Processing the audio features and the hidden layer features through the loop decoder to obtain parameters of each joint of the virtual character in each animation frame;
[0075] The second animation is generated based on the parameters of each joint of the virtual character in each animation frame.
[0076] Optionally, the second animation is an animation corresponding to the second text segment, and the second animation generation module is used to splice the first animation and the second animation to obtain the target animation.
[0077] Optionally, the first animation generation module is used to process the speech audio corresponding to the speech text through the animation generation model to obtain the second animation, wherein the second animation includes an animation corresponding to the first text segment and an animation corresponding to the second text segment, and the rhythm of the virtual character in the second animation matches the speaking rhythm of the speech text;
[0078] The second animation generating module is used to replace the animation in the second animation corresponding to the first text segment with the first animation to obtain the target animation.
[0079] Optionally, the device further comprises:
[0080] An audio acquisition module, used to acquire the speech audio of the virtual character;
[0081] A first conversion module, used for converting the speech audio into the speech text;
[0082] The device also includes:
[0083] The audio extraction module is used to extract an audio segment corresponding to the second text segment from the speech audio.
[0084] Optionally, the device further comprises:
[0085] A text acquisition module, used to acquire the speech text of the virtual character;
[0086] A second conversion module is configured to convert the second text segment into an audio segment.
[0087] On the other hand, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the animation generation method described in the above aspects.
[0088] On the other hand, a computer-readable storage medium is provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to implement the operations performed by the animation generation method described in the above aspects.
[0089] On the other hand, a computer program product is provided, comprising a computer program, wherein the computer program is loaded and executed by a processor to implement the operations performed by the animation generation method as described in the above aspects.
[0090] In an embodiment of the present application, a large language model is used to perform semantic analysis on a speech text of a virtual character, thereby dividing the speech text into a first text segment with a first semantic tag and a second text segment without a semantic tag. For the first text segment, a first animation matching the first semantic tag is retrieved from an animation library, and the semantics expressed by the virtual character in the first animation matches the semantics of the first text segment, thereby matching the semantics of the speech text. For the second text segment, a second animation is generated through an animation generation model, and in the animation corresponding to the second text segment included in the second animation, the rhythm of the virtual character matches the speech rhythm of the second text segment. In the target animation generated based on the first animation and the second animation, the semantics expressed by the virtual character matches the semantics of the speech text, and the rhythm of the virtual character matches the speech rhythm of the second text segment, that is, a semantically matching and rhythmic animation is generated, which enriches the amount of information contained in the animation and improves the animation playback effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0092] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present application;
[0093] Figure 2 is a flow chart of an animation generation method provided in an embodiment of the present application;
[0094] Figure 3 is a flowchart of another animation generation method provided in an embodiment of the present application;
[0095] Figure 4 is a schematic diagram of semantic tags and animations in an animation library provided in an embodiment of the present application;
[0096] Figure 5 It is a flowchart of an animation generation method provided in an embodiment of the present application;
[0097] Figure 6 is a schematic diagram of an animation generation interface provided in an embodiment of the present application;
[0098] Figure 7 It is a structural schematic diagram of an animation generation model provided in an embodiment of the present application;
[0099] Figure 8 It is a processing flow diagram of an animation generation model provided in an embodiment of the present application;
[0100] Fig. 9 is a structural schematic diagram of another animation generation model provided in an embodiment of the present application;
[0101] Fig.10 is a schematic diagram of an animation frame in an animation generated by a related technology;
[0102] Fig.11 is a schematic diagram of an animation frame in an animation generated by an embodiment of the present application;
[0103] Fig.12 is a structural schematic diagram of an animation generating device provided in an embodiment of the present application;
[0104] Fig.13 is a structural schematic diagram of another animation generating device provided in an embodiment of the present application;
[0105] Fig.14 is a schematic diagram of the structure of a terminal provided in an embodiment of the present application;
[0106] Fig.15 It is a structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0107] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the present application will be further described in detail below in conjunction with the accompanying drawings.
[0108] It is understood that the terms "first", "second", etc. used in this application can be used in this article to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another concept. For example, without departing from the scope of this application, a first animation can be referred to as a second animation, and similarly, a second animation can be referred to as a first animation.
[0109] Here, at least two refers to two or more than two, for example, at least two animations can be two animations, three animations, or any integer greater than or equal to two. Each refers to each of the at least two, for example, each animation refers to each animation of the at least two animations, and if the at least two animations are three animations, each animation refers to each animation of the three animations.
[0110] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) involved in this application are all fully authorized by users or relevant parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0111] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0112] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, pre-trained models are also called large models and basic models. After fine-tuning, they can be widely used in downstream tasks in various major directions of artificial intelligence. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, as well as machine learning / deep learning, autonomous driving, smart transportation and other major directions.
[0113] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless cars, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence-generated content, conversational interaction, smart medical care, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0114] Pre-training model (PTM), also known as cornerstone model or big model, refers to a deep neural network (DNN) with large parameters. It is trained on massive unlabeled data, and the function approximation ability of large-parameter DNN is used to enable PTM to extract common features from the data. After fine tuning, parameter-efficient fine-tuning (PEFT), prompt-tuning and other technologies, it is suitable for downstream tasks. Therefore, the pre-training model can achieve ideal results in few-shot or zero-shot scenarios. PTM can be divided into language model, visual model, speech model, multimodal model, etc. according to the data modality processed. Among them, the multimodal model refers to a model that establishes two or more data modality feature representations. The pre-training model is an important tool for outputting AI-generated content, and can also be used as a general interface to connect multiple specific task models.
[0115] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between people and computers using natural language. Natural language processing involves natural language, which is the language people use in daily life, and is closely related to linguistics research; it also involves computer science and mathematics. The pre-training model, an important technology for model training in the field of artificial intelligence, is developed from the large language model (LLM) in the field of NLP. After fine-tuning, the large language model can be widely used in downstream tasks. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0116] The solution provided in the embodiments of the present application involves technologies such as artificial intelligence natural language processing, which is specifically described by the following embodiments:
[0117] First, the terms involved in the embodiments of the present application are explained as follows:
[0118] 1. Big Language Model: An AI model designed to understand and generate human language. Big Language Models are trained on large amounts of text and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, and more.
[0119] 2. Action matching: Search the animation library to find matching animations, and then combine the determined animation with the found animation to form a new animation.
[0120] 3. Speech To Gesture (S2G): A deep learning model that can generate speaker gesture animations based on input speaker audio.
[0121] 4. Prompt: Information used to prompt the large language model in the large language model, which can represent the task assigned to the large language model. Prompts are used to guide and inspire the large language model to generate specific content. Prompts are crucial to obtaining high-quality, accurate, and useful output. Prompts help the large language model better understand user needs and generate more relevant content.
[0122] 5. Animation library: The animation library is a database that stores a large number of animation clips and the semantic tags of each animation clip. The matching animation clips can be retrieved from the animation library through the semantic tags to generate animations.
[0123] 6. Text To Speech (TTS): A technology that converts text into audio, which can convert text into speech, making the speech generated by computer devices sound as natural as human voice.
[0124] 7. Convolutional Neural Networks (CNN): Convolutional neural networks are a type of artificial neural network that includes convolution calculations and has a deep structure. Convolutional neural networks include input layer, convolution layer, activation function layer and pooling layer.
[0125] 8. UE (Unreal Engine): A game engine, a complete game development platform for the next generation of game consoles and personal computers, providing a large number of core technologies, data generation tools and basic support needed by game developers.
[0126] 9. Logarithmic frequency spectrum: A frequency spectrum diagram in which logarithmic calculations are performed on each spectral line so that the lower amplitude components can be raised relative to the high amplitude components in order to observe periodic signals hidden in low-amplitude noise.
[0127] 10. Mel frequency: also known as the Mel scale, is an indicator that describes the frequency perception of sound heard by the human ear. It is a nonlinear frequency scale determined based on the human ear's sensory judgment of equidistant pitch changes. Its principle is that the Mel frequency filter group has a high resolution in the low-frequency part, which is consistent with the auditory characteristics of the human ear.
[0128] 11. Conditional autoregressive model: a stationary time series model that uses statistical analysis of past time series data to infer the development trend of things.
[0129] 12. Gated Recurrent Neural Unit: An improved recurrent neural network architecture that incorporates some gating mechanisms to better capture long-term dependencies in time series data.
[0130] The animation generation method provided in the embodiment of the present application is used in a computer device. Optionally, the computer device is a terminal or a server. Optionally, the terminal is a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart TV, an intelligent voice interaction device, a smart home appliance, etc., but is not limited to this. Optionally, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network, content distribution network) and big data and artificial intelligence platforms and other basic cloud computing services. The embodiment of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc.
[0131] In one possible implementation, the computer program involved in the embodiments of the present application may be deployed and executed on a computer device, or on multiple computer devices located at one location, or on multiple computer devices distributed at multiple locations and interconnected through a communication network. Multiple computer devices distributed at multiple locations and interconnected through a communication network can constitute a blockchain system.
[0132] In one possible implementation, the computer device in the embodiment of the present application is a node in a blockchain system, which can store spoken text and animation of spoken text, or spoken audio and animation of spoken audio in the blockchain, and then the node or the node corresponding to other devices in the blockchain can query the spoken text or animation of spoken audio by accessing the blockchain.
[0133] Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application, see Figure 1 The implementation environment includes: a terminal 101 and a server 102. The terminal 101 and the server 102 are connected via a communication network 103.
[0134] The terminal 101 runs an application, and the server 102 is associated with the application, thereby providing computing services for the application. The application is an animation production program, a game application, etc., which is not limited in the embodiment of the present application.
[0135] The terminal 101 can obtain the spoken text of the virtual character and send the spoken text to the server 102. The server 102 uses the method provided in the embodiment of the present application to perform semantic analysis on the spoken text through a large language model. For the first text segment with the first semantic tag, a matching first animation is retrieved from the animation library. For the second text segment without the semantic tag, a second animation is generated through the animation generation model. In the target animation generated based on the first animation and the second animation, the semantics expressed by the virtual character matches the semantics of the spoken text, and the rhythm of the virtual character matches the speaking rhythm of the second text segment. The server 102 returns the target animation to the terminal 101. The semantics expressed by the virtual character in the animation matches the semantics of the spoken text. The terminal 101 can convert the spoken text into spoken audio, and play the animation while playing the spoken audio, simulating the scene of the virtual character making actions or expressions while speaking. Alternatively, the terminal 101 may obtain the speech audio of the virtual character and send the speech audio to the server 102. The server 102 uses the method provided in the embodiment of the present application to convert the speech audio into speech text, and performs semantic analysis on the speech text through a large language model. For the first text segment with the first semantic tag, a matching first animation is retrieved from the animation library, and for the second text segment without the semantic tag, a second animation is generated through the animation generation model. In the target animation generated based on the first animation and the second animation, the semantics expressed by the virtual character matches the semantics of the speech text, and the rhythm of the virtual character matches the speaking rhythm of the second text segment. The server 102 returns the target animation to the terminal 101. The semantics expressed by the virtual character in the target animation matches the semantics of the speech audio. The terminal 101 can play the target animation while playing the speech audio, simulating a scene where the virtual character speaks while making actions or expressions.
[0136] The animation generation method provided in the embodiment of the present application can be applied to a variety of scenarios.
[0137] For example, in a game development scenario, one or more virtual characters are set in a game application, and it is necessary to play an animation of the virtual character performing actions while speaking. Therefore, the game developer determines the speech text of the virtual character and inputs it into the animation production program. The computer device uses the method provided in the embodiment of the present application through the animation production program to generate an animation based on the speech text, wherein the speech text represents what the virtual character wants to say, and the generated animation is the animation played when the virtual character speaks, and the animation can present the effect of the virtual character performing corresponding actions while speaking, and the actions performed match the words spoken. After the animation is generated, the speech text and the animation are stored correspondingly for use by the game engine (such as UE) in the game application when displaying the virtual character.
[0138] For another example, in an animation production scenario, if a production staff wants to generate an animation, the lines of the cartoon character are input into the animation production program as speech text. The computer device uses the method provided in the embodiment of the present application through the animation production program to generate an animation based on the speech text. The generated animation can present the effect of the cartoon character speaking the lines while performing corresponding actions, and the actions performed match the spoken lines.
[0139] Figure 2 is a flowchart of an animation generation method provided in an embodiment of the present application. The embodiment of the present application is executed by a computer device. Figure 2 , the method comprising:
[0140] 201. A computer device performs semantic analysis on a speech text of a virtual character through a large language model to determine a first text segment in the speech text and a first semantic tag corresponding to the first text segment, wherein the first semantic tag represents the semantics of the first text segment.
[0141] Among them, virtual characters include various types of characters such as virtual people and virtual animals. In the embodiment of the present application, the computer device generates animations for the virtual characters, and the generated animations include virtual characters.
[0142] The speech text is the text that the virtual character needs to speak, so the animation generated based on the speech text includes the scene of the virtual character speaking the speech text. For example, when displaying a virtual character in a game application, it is necessary to play the scene of the virtual character speaking, or when making an animation, it is necessary to produce a scene of the virtual character speaking lines.
[0143] The spoken text may be input into the computer device by the user. Optionally, the computer device displays an animation generation interface, the animation generation interface displays a text input area, the user inputs the spoken text in the text input area, and the computer device obtains the input spoken text.
[0144] Alternatively, the speech text may be obtained by converting the speech audio. Optionally, the computer device obtains the input speech audio of the virtual character and converts the speech audio into speech text. For example, the computer device displays an animation generation interface, and the animation generation interface displays an audio input area. After the user triggers the audio input area, the user selects the storage location where the speech audio is located from the storage location selection window displayed by the computer device. The computer device obtains the speech audio located at the storage location and converts the speech audio into speech text.
[0145] In addition, the large language model is used to perform semantic analysis on the text, and may include GPT (Generative Pre-Trained Transformer), GLM (Generalized Linear Models), CPM (Chinese Pre-trained Models), mixed-element large models, etc. The embodiments of the present application do not limit the large language model. The spoken text is input into the large language model, and the large language model can be used to perform semantic analysis on the spoken text to identify the first text segment in the spoken text and the first semantic label corresponding to the first text segment. Among them, the first text segment is a text segment in the spoken text that has a corresponding semantic label, that is, a text segment containing semantics, and the first semantic label represents the semantics of the first text segment.
[0146] The first text segment is all text segments in the spoken text, or is part of the text segments in the spoken text. The first text segment may be one text segment or multiple text segments, which is not limited in the embodiment of the present application.
[0147] Optionally, the first text segment is a partial text segment in the spoken text. In addition to the first text segment, the spoken text also includes a second text segment. The second text segment is a text segment without a corresponding semantic label, that is, a text segment that does not contain semantics. The computer device can also recognize the second text segment in the spoken text through a large language model. The second text segment can be one text segment or multiple text segments. Since the second text segment cannot express semantics, there is no need to obtain the semantic label corresponding to the second text segment.
[0148] The big language model is an artificial intelligence model that needs to be trained on a large amount of text to ensure that semantic analysis can be accurately performed through the big language model. The training process of the big language model includes: obtaining sample text, in which the sample text is annotated with semantic keywords and semantic labels corresponding to the semantic keywords, wherein the semantic keyword is a text segment representing semantics in the sample text, and the semantics of the semantic keyword will affect the overall semantics of the sample text. Then, based on the sample text, the semantic keyword and the semantic label, the big language model is trained so that the big language model has the ability to identify text segments with semantic labels in any text and determine the semantic labels.
[0149] The semantic keywords and the semantic tags can be annotated by a technician, who counts the words that may be mentioned in daily conversation texts as preset semantic tags, and annotates the words in the spoken text with appropriate preset semantic tags. For example, if the spoken text includes a text segment, and the text segment is a preset semantic tag itself, the text segment is annotated as a semantic keyword, and the preset semantic tag is used as the semantic tag corresponding to the semantic keyword. Alternatively, if the spoken text includes a text segment, and the text segment is a similar word to a preset semantic tag, the text segment is annotated as a semantic keyword, and the preset semantic tag is used as the semantic tag corresponding to the semantic keyword.
[0150] By collecting multiple sample texts and iteratively training the large language model, the accuracy of the large language model can be improved until the large language model meets the stop training condition, indicating that the accuracy of the large language model meets the requirement, and then the large language model can be used to perform semantic analysis on any spoken text. Among them, the stop training condition includes: the number of iterative training reaches the target number, or the accuracy of the large language model is not less than the target accuracy, or the loss value of the large language model is less than the target loss value, etc.
[0151] Optionally, the large language model includes model parameters, for example, the large language model includes multiple network layers, and the model parameters of the large language model include parameters of each network layer, such as the large language model includes network layers such as input layer, hidden layer and output layer, and the parameters of the network layer include weight parameters and bias parameters, etc. However, the initialized model parameters are not accurate enough, resulting in low accuracy of the large language model. In the process of training the large language model based on sample text, the sample text is input into the large language model, and the large language model processes the sample text according to the current model parameters, identifies the predicted semantic keywords and the predicted semantic labels corresponding to the predicted semantic keywords from the sample text, and compares the predicted semantic keywords and the predicted semantic labels with the semantic keywords and semantic labels marked in the sample text to determine the loss value, which indicates the degree of deviation of the semantic analysis result of the large language model. Based on the loss value, the model parameters of the large language model are adjusted to improve the accuracy of the large language model after adjustment. This is an iterative training process. After multiple iterative trainings, the accuracy of the large language model can be guaranteed to meet the requirements. Afterwards, if the spoken text is input into the large language model, the large language model processes the spoken text according to the trained model parameters, and identifies a first text segment and a first semantic label corresponding to the first text segment from the spoken text.
[0152] 202. The computer device retrieves a first animation matching a first semantic tag from an animation library, wherein the semantics expressed by the virtual character in the first animation matches the semantics represented by the first semantic tag.
[0153] The computer device is provided with an animation library, in which at least one animation is stored, and each animation is annotated with a matching semantic tag, wherein the matching of the animation and the semantic tag means that the semantics expressed by the virtual character in the animation matches the semantics represented by the semantic tag. Therefore, in the first animation retrieved from the animation library that matches the first semantic tag, the semantics expressed by the virtual character matches the semantics represented by the first semantic tag, that is, the semantics matches the first text segment.
[0154] There are many ways for virtual characters in animations to express semantics. For example, if a virtual character makes an action that expresses a certain semantics, then matching the animation with the semantic label means that the semantics of the action made by the virtual character in the animation matches the semantics represented by the semantic label. For example, if the first text segment is "Hello", then the virtual character in the first animation makes a waving motion. Or if the first text segment is "No", then the virtual character in the first animation makes a shaking motion. Or, if the virtual character makes an expression that expresses a certain semantics, then matching the animation with the semantic label means that the semantics of the expression made by the virtual character in the animation matches the semantics represented by the semantic label. For example, if the first text segment is "Angry", then the virtual character in the first animation makes an angry expression.
[0155] The animations in the animation library can be made by technicians. For example, the technicians determine the semantic tags, and the assistants make movements that match the semantic tags while wearing motion capture equipment, so as to record the animations through the motion capture equipment, or the assistants make expressions that match the semantic tags and shoot the assistants through a camera to obtain the animations, and then store the animations and semantic tags in the animation library. In order to ensure the universality of the animation library, the determined semantic tags need to cover as many semantic tags that may be mentioned in daily conversation texts as possible, so as to ensure the diversity of the animations.
[0156] 203. The computer device processes the audio segment corresponding to the second text segment through the animation generation model to obtain a second animation.
[0157] Among them, the second text segment is a text segment without a corresponding semantic tag in the spoken text. The semantic tag corresponding to the second text segment is not obtained through the large language model, and the matching animation cannot be retrieved from the animation library. Therefore, in an embodiment of the present application, an animation corresponding to the second text segment is generated by an animation generation model.
[0158] The animation generation model is used to generate animation based on audio, and the rhythm of the virtual character in the generated animation matches the speaking rhythm in the audio. The animation generation model includes an S2G model or other models, which are not limited in the embodiments of the present application.
[0159] The computer device processes the audio segment corresponding to the second text segment through the animation generation model to obtain a second animation. The second animation includes an animation corresponding to the second text segment, that is, the animation corresponding to the second text segment includes a picture of a virtual character speaking the second text segment, and the rhythm of the virtual character in the animation corresponding to the second text segment matches the speaking rhythm of the second text segment, for example, the action rhythm of the virtual character matches the speaking rhythm of the second text segment, and when there is a pause in the middle of the second text segment, the virtual character in the animation corresponding to the second text segment also stops performing actions.
[0160] 204. The computer device generates a target animation of the virtual character based on the first animation and the second animation.
[0161] The first animation is an animation corresponding to the first text segment, and the second animation includes an animation corresponding to the second text segment. Based on the first animation and the second animation, a target animation of a virtual character can be generated, and the target animation includes a picture of the virtual character performing an action, and the semantics of the action performed matches the semantics of the spoken text. The target animation also includes an animation corresponding to the second text segment. Since the rhythm of the virtual character in the animation corresponding to the second text segment matches the speaking rhythm of the second text segment, the rhythm of the virtual character in the target animation matches the speaking rhythm of the second text segment.
[0162] In the embodiment of the present application, the second animation includes an animation corresponding to the second text segment, which means that the second animation itself is an animation corresponding to the second text segment, or the second animation includes an animation corresponding to the second text segment and also includes other animations. In view of the difference of the second animation, the embodiment of the present application includes the following two possible implementation methods:
[0163] In a first possible implementation, the animation generation model only processes the audio segment corresponding to the second text segment, and no longer processes the audio segment corresponding to the first text segment, that is, step 203 includes: the computer device processes the audio segment corresponding to the second text segment through the animation generation model, and the obtained second animation is the animation corresponding to the second text segment. Step 204 includes: the computer device splices the first animation and the second animation to obtain the target animation of the virtual character.
[0164] When the computer device obtains the input speech text and converts the speech text into speech audio, after the computer device determines the first text segment and the second text segment, the second text segment is converted into an audio segment, and the audio segment corresponding to the second text segment can be processed by the animation generation model. Alternatively, when the computer device obtains the input speech audio and converts the speech audio into speech text, after the computer device determines the first text segment and the second text segment, based on the position of the second text segment in the speech text, the audio segment corresponding to the second text segment is extracted from the speech audio, and the audio segment corresponding to the second text segment can be processed by the animation generation model.
[0165] Since the spoken text includes the first text segment and the second text segment, the first animation and the second animation are spliced according to the order of the first text segment and the second text segment in the spoken text to obtain the target animation corresponding to the spoken text.
[0166] In a second possible implementation, step 203 includes: the computer device processes the speech audio corresponding to the speech text through the animation generation model to obtain a second animation. Step 204 includes: replacing the animation corresponding to the first text segment in the second animation with the first animation to obtain a target animation.
[0167] Since the spoken text includes a first text segment and a second text segment, the spoken audio corresponding to the spoken text includes an audio segment corresponding to the first text segment and an audio segment corresponding to the second text segment. Accordingly, the second animation includes an animation corresponding to the first text segment and an animation corresponding to the second text segment, wherein the animation corresponding to the first text segment includes a picture of a virtual character speaking the first text segment, and the rhythm of the virtual character matches the speaking rhythm of the first text segment, and the animation corresponding to the second text segment includes a picture of a virtual character speaking the second text segment, and the rhythm of the virtual character matches the speaking rhythm of the second text segment.
[0168] The spoken text includes a first text segment and a second text segment. Since the second text segment does not contain semantics, the semantics of the first text segment can represent the semantics of the spoken text, and the semantics of the first animation matches the semantics of the first text segment, so the semantics of the first animation also matches the semantics of the spoken text. However, the animation corresponding to the first text segment in the second animation cannot accurately reflect the semantics of the first text segment. In order to obtain a target animation that can represent the semantics of the spoken text, the animation corresponding to the first text segment in the second animation needs to be replaced with the first animation.
[0169] It should be noted that the virtual character in the embodiment of the present application is a template virtual character set by default in the computer device, the animation stored in the animation library includes the template virtual character, and the animation generated by the animation generation model also includes the template virtual character, then the generated target animation also includes the template virtual character. If an animation including another target virtual character is to be generated subsequently, the target animation is redirected according to the character data of the template virtual character and the target virtual character, so that the template virtual character in the target animation is modified to the target virtual character. For example, the redirection process includes: according to the difference information between the character data of the template virtual character and the character data of the target virtual character, the template virtual character in the target animation is adjusted according to the difference information, so that the adjusted virtual character becomes the target virtual character.
[0170] Alternatively, the user inputs the character data of the target virtual character in the computer device, and the animation stored in the animation library includes the template virtual character. After retrieving the first animation from the animation library, the first animation is redirected according to the character data of the template virtual character and the target virtual character, thereby modifying the template virtual character in the first animation to the target virtual character. The character data of the target virtual character is input into the animation generation model, and the character data of the target virtual character is processed by the animation generation model so that the generated second animation includes the target virtual character. The target animation generated based on the first animation and the second animation can include the target virtual character.
[0171] The target virtual character may be a cartoon character in an animation, or an NPC (Non-Player Character) in a game application, etc.
[0172] In an embodiment of the present application, a large language model is used to perform semantic analysis on a speech text of a virtual character, thereby dividing the speech text into a first text segment with a first semantic tag and a second text segment without a semantic tag. For the first text segment, a first animation matching the first semantic tag is retrieved from an animation library, and the semantics expressed by the virtual character in the first animation matches the semantics of the first text segment, thereby matching the semantics of the speech text. For the second text segment, a second animation is generated through an animation generation model, and in the animation corresponding to the second text segment included in the second animation, the rhythm of the virtual character matches the speech rhythm of the second text segment. In the target animation generated based on the first animation and the second animation, the semantics expressed by the virtual character matches the semantics of the speech text, and the rhythm of the virtual character matches the speech rhythm of the second text segment, that is, a semantically matching and rhythmic animation is generated, which enriches the amount of information contained in the animation and improves the animation playback effect.
[0173] Furthermore, assuming that a target animation is to be obtained by processing speech audio through a certain model, if the semantics of the target animation is to match the semantics of the speech text contained in the speech audio, it is necessary to ensure that the training data of the model includes sample speech audio and sample animation, and the actions performed by the virtual character in the sample animation match the semantics of the sample speech text contained in the sample speech audio. Even if the technicians collect a sufficient amount of training data to train the model, after the training is completed, when the model is actually used to generate animations, it depends entirely on the internal processing process of the model. However, semantic labels are ambiguous and contingent, that is, the same semantic label corresponds to a variety of different actions, and different semantic labels correspond to similar actions. The actual situation is complex and changeable, resulting in the difficulty in achieving strict control of the final generation result of the model. It is very likely that there are semantic keywords in the speech text but there are no actions matching the semantic keywords in the generated animation, or there are no actions matching the semantic keywords in the generated animation after the semantic keywords in the speech text are replaced with similar words. Therefore, the accuracy and diversity of the above scheme are poor.
[0174] However, the animation generation method provided in the embodiment of the present application is different. In the embodiment of the present application, the spoken text is first semantically analyzed through a large language model to ensure that the first text segment with semantic tags and the second text segment without semantic tags in the spoken text can be accurately identified, and then the animation library and the animation generation model are used to process the first text segment and the second text segment respectively to generate a rhythmic and semantically matching animation. It does not rely entirely on the internal processing process of the animation generation model. Therefore, the animation generation method provided in the embodiment of the present application has strong interpretability and controllability, improves accuracy and diversity, and does not have errors in the generation results.
[0175] In the above Figure 2 Based on the embodiment shown, the embodiment of the present application also provides another animation generation method. Figure 3 is a flowchart of another animation generation method provided in an embodiment of the present application. The embodiment of the present application is executed by a computer device. Figure 3 , the method comprising:
[0176] 301. The computer device performs semantic analysis on the spoken text through a large language model to obtain semantic labels and weights of at least two third text segments in the spoken text.
[0177] The large language model is used to perform semantic analysis on spoken text and identify the semantic labels and weights of text segments in the spoken text. The semantic labels represent the semantics of the text segments, and the weights represent the importance of the text segments to the spoken text, that is, the importance of the semantics of the text segments to the semantics of the spoken text.
[0178] In the embodiment of the present application, taking the example of the large language model identifying the semantic labels and weights of at least two third text segments, in one possible implementation, the at least two third text segments are all the text segments in the spoken text, that is, the large language model identifies the semantic labels and weights of each text segment in the spoken text. In another possible implementation, at least two third text segments are some text segments in the spoken text, that is, the large language model identifies the semantic labels and weights of some text segments in the spoken text, and the remaining text segments are text segments without semantic labels and weights.
[0179] The big language model is an artificial intelligence model that needs to be trained on a large amount of text to ensure that the big language model can accurately perform semantic analysis. The training process of the big language model includes: obtaining sample text, which includes semantic keywords; obtaining the semantic labels and weights of the semantic keywords, and the weights represent the importance of the semantic keywords to the sample text; based on the sample text, the semantic labels and weights of the semantic keywords, the big language model is trained to enable the big language model to have the ability to identify a text segment with a semantic label in any text and determine the semantic label and weight of the text segment.
[0180] Among them, the semantic keyword is a text segment that represents semantics in the sample text, and the semantics of the semantic keyword will affect the overall semantics of the sample text. Among them, the semantic keyword, the semantic label and weight of the semantic keyword can be annotated by a technician, and the technician counts the words that may be mentioned in the daily conversation text as preset semantic labels, and annotates the words in the spoken text with appropriate preset semantic labels. For example, the spoken text includes a text segment, and the text segment is a preset semantic label itself, then the text segment is annotated as a semantic keyword, and the preset semantic label is used as the semantic label corresponding to the semantic keyword, and the technician measures the degree of influence of the semantic keyword on the spoken text based on empirical knowledge, thereby determining the weight of the semantic keyword. Alternatively, the spoken text includes a text segment, and the text segment is a similar word to a preset semantic label, then the text segment is annotated as a semantic keyword, and the preset semantic label is used as the semantic label corresponding to the semantic keyword, and the technician measures the degree of influence of the semantic keyword on the spoken text based on empirical knowledge, thereby determining the weight of the semantic keyword. In addition, the words without semantics in the spoken text can also be annotated, so as to obtain a large number of positive samples and negative samples for training.
[0181] By collecting multiple sample texts and iteratively training the large language model, the accuracy of the large language model can be improved until the large language model meets the conditions for stopping training, which means that the accuracy of the large language model has reached the requirements. After that, the large language model can be used to perform semantic analysis on any spoken text.
[0182] Optionally, the large language model includes model parameters, for example, the large language model includes multiple network layers, and the model parameters of the large language model include parameters of each network layer, such as the large language model includes network layers such as input layer, hidden layer and output layer, and the parameters of the network layer include weight parameters and bias parameters, etc. However, the initialized model parameters are not accurate enough, resulting in low accuracy of the large language model. In the process of training the large language model based on sample text, the sample text is input into the large language model, and the large language model processes the sample text according to the current model parameters, identifies the predicted semantic keywords and the predicted semantic labels and predicted weights corresponding to the predicted semantic keywords from the sample text, and compares the predicted semantic keywords, the predicted semantic labels and the predicted weights with the semantic keywords, semantic labels and weights marked in the sample text to determine the loss value, which indicates the degree of deviation of the semantic analysis result of the large language model, and adjusts the model parameters of the large language model based on the loss value to improve the accuracy of the adjusted large language model. After multiple iterations of training, the accuracy of the large language model can be guaranteed to meet the requirements. Afterwards, if the spoken text is input into the large language model, the large language model processes the spoken text according to the trained model parameters, and identifies the third text segment, the semantic label and weight of the third text segment from the spoken text.
[0183] 302. The computer device selects a first text segment from at least two third text segments based on weights of the at least two third text segments.
[0184] The weight of the third text segment indicates the importance of the third text segment to the spoken text. Therefore, the higher the weight, the greater the influence of the semantics of the third text segment on the semantics of the spoken text. The first text segment is selected from at least two third text segments, and the semantic label of the first text segment can be called the first semantic label. The weight of the first text segment is greater than the weight of the unselected third text segment, that is, the text segment with a higher weight is selected. Subsequently, the animation with semantic matching will be retrieved for the text segment with a higher weight, so as to ensure that the semantics expressed by the virtual character in the generated animation matches the semantics of the spoken text as much as possible.
[0185] 303. The computer device determines the unselected third text segment as the second text segment.
[0186] The third text segments that are not selected are text segments with lower weights, and the semantics of these text segments have little impact on the semantics of the spoken text. In order to avoid the generated animation being too redundant, these text segments are determined as second text segments without semantic labels, that is, the text segments with lower importance are discarded. Subsequently, there is no need to retrieve the animation matching the second text segment from the animation library, nor is there any need to generate an animation matching the semantics of the second text segment. Instead, the animation corresponding to the second text segment is generated through the animation generation model.
[0187] Optionally, the computer device determines the input target number, which represents the number of text segments to be selected, and then selects the target number of first text segments from at least two third text segments in descending order of weight, and determines the unselected third text segments as the second text segments. Alternatively, the computer device determines the input target ratio, which represents the ratio of the number of text segments to be selected to the total number of text segments in the spoken text, and then determines the target number according to the total number of text segments in the spoken text and the target ratio, and the target number represents the number of text segments to be selected, and then selects the target number of first text segments from at least two third text segments in descending order of weight, and determines the unselected third text segments as the second text segments.
[0188] In addition, when at least two third text segments are partial text segments in the spoken text, in addition to the at least two third text segments recognized by the large language model, the spoken text also includes other text segments, then the other text segments are determined as second text segments without semantic tags.
[0189] Optionally, before step 302, the method further includes: the computer device displays semantic labels and weights of at least two third text segments in the animation generation interface, and obtains the semantic labels and weights of the edited third text segments in response to the editing operation in the animation generation interface.
[0190] The editing operation includes at least one of the following: an operation of changing the weight of the third text segment, an operation of changing the semantic label of the third text segment, an operation of adding the semantic label and weight of a new text segment, an operation of deleting the third text segment, the semantic label and weight of the third text segment, etc. After performing the editing operation, the computer device may execute steps 302 and 303 for the semantic label and weight of the edited third text segment.
[0191] The embodiment of the present application provides a secondary editing function, where the user edits the semantic label and weight of the third text segment determined by the large language model based on his or her own experience and knowledge, so that the subsequently generated animation better meets the user's requirements.
[0192] It should be noted that the above steps 301-303 are optional steps, and the computer device can also use other methods through the large language model to determine the first text segment in the spoken text and the first semantic label corresponding to the first text segment.
[0193] Optionally, the speech text of the virtual character is semantically analyzed by the large language model to determine the first text segment in the speech text and the first semantic tag corresponding to the first text segment, including: the speech text is semantically analyzed by the large language model based on the prompt word to determine the first text segment and the first semantic tag corresponding to the first text segment; wherein the prompt word represents the semantic analysis task sent to the large language model, and represents the task requirements of the semantic analysis task. The large language model can also process the prompt word during the processing process, understand the semantics of the prompt word, and thus generate an animation according to the semantics of the prompt word to ensure that the generated animation meets the requirements of the prompt word. For example, the prompt word includes "the division should be as detailed as possible", "divide the whole sentence into as many words as possible, each word has a unique meaning", etc. The prompt word can be set by default by the computer device and used every time an animation needs to be generated, or it can be input by the user of the computer device when an animation needs to be generated at the moment.
[0194] In order to ensure that the large language model can accurately understand the meaning of the prompt word, a corpus can be created during the training of the large language model. The corpus includes multiple sentences. The large language model is trained based on the corpus. The large language model can learn the relationship between words in the corpus, as well as the relationship between sentences, etc., so as to understand the meaning of words and sentences. After the prompt word is input into the large language model, the large language model can use previous experience and knowledge to understand the meaning of the prompt word, and thus perform semantic analysis according to the requirements of the prompt word.
[0195] 304. The computer device retrieves a first animation matching the first semantic tag from an animation library.
[0196] The animation library can store multiple animations matching the first semantic tag. The computer device can randomly select an animation from the multiple animations as the first animation, or display the multiple animations to the user, and the user selects one of the animations as the first animation, or the computer device can select a suitable animation from the multiple animations as the first animation.
[0197] In one possible implementation, step 304 includes: retrieving multiple candidate animations matching the first semantic tag from an animation library, and selecting a first animation from the multiple candidate animations based on the playback duration of the first text segment and the playback durations of the multiple candidate animations.
[0198] The difference between the playing time of the first animation and the first text segment is smaller than the target threshold; or the difference between the playing time of the first animation and the first text segment is smaller than the difference between the playing time of other candidate animations and the first text segment.
[0199] The computer device determines a playback duration for the first text segment, where the playback duration represents the duration of the animation corresponding to the first text segment in the target animation of the virtual character. The playback duration can be determined based on the total duration of the target animation and the proportion of the number of words in the first text segment to the total number of words in the spoken text, or based on the speaking speed of an average person and the number of words in the first text segment, or based on the playback duration of the audio segment corresponding to the first text segment in the spoken audio, and this embodiment of the present application does not limit this.
[0200] Multiple alternative animations all match the first semantic tag, that is, the semantics expressed by the virtual characters in the multiple alternative animations all match the semantics represented by the first semantic tag. Considering that the playback duration of the animation and the first text segment cannot differ too much, otherwise the playback duration of the target animation finally generated will differ too much, therefore, an animation whose playback duration is less different from that of the first text segment is selected from the multiple alternative animations as the first animation.
[0201] In another possible implementation, step 304 includes: retrieving multiple alternative animations matching the first semantic tag from the animation library, obtaining adjacent animations of the first text segment, where the adjacent animations are animations of adjacent text segments of the first text segment, respectively determining difference information between each alternative animation and the adjacent animation, and selecting a first animation from the multiple alternative animations based on the difference information of the multiple alternative animations.
[0202] The spoken text includes multiple text segments, wherein the text segment adjacent to the first text segment is the adjacent text segment of the first text segment, such as the text segment before the first text segment, or the text segment after the first text segment. Considering that after determining the animation of each text segment in the spoken text, the animations of the multiple text segments need to be spliced in the order of the multiple text segments to obtain the target animation, in order to ensure that the target animation is played smoothly and naturally, the animations of adjacent text segments are required to be smoothly connected.
[0203] Therefore, when multiple candidate animations matching the first semantic tag are retrieved from the animation library, and the animation of the adjacent text segment of the first text segment has been determined, an animation with a smaller difference from the adjacent animation can be selected from the multiple candidate animations as the first animation. Among them, the text segment adjacent to the first text segment is a text segment with a semantic tag, or a text segment without a semantic tag, then the adjacent animation may include an animation retrieved from the animation library, and may also include an animation generated by an animation generation model. The generation method of the adjacent animation is the same as the generation method of the first animation or the second animation, and the embodiments of the present application will not be repeated here.
[0204] For example, the animation includes character skeleton data of at least one animation frame, and the character skeleton data is used to describe the position of the skeleton of the virtual character in the animation frame, and the character skeleton data includes the position coordinates, rotation vector, movement speed, movement phase, etc. of each skeleton of the virtual character. The posture of the virtual character in the animation frame can be determined by the character skeleton data. And determining the difference information between the alternative animation and the adjacent animation includes: calculating the sum, average value or variance of the difference between each data in the character skeleton data of the alternative animation and the adjacent animation as the difference information between the alternative animation and the adjacent animation.
[0205] In a possible implementation, the alternative animation includes multiple animation frames, and the adjacent animation is the animation of the text segment before the first text segment. The first N animation frames are selected from the alternative animation, and the last N animation frames are selected from the adjacent animation. The sum, average or variance of the difference between the first N animation frames in the alternative animation and the last N animation frames in the adjacent animation are calculated as the difference information between the alternative animation and the adjacent animation. Wherein, N is an integer greater than 1. The adjacent animation needs to be used as the preceding animation of the first animation to be generated. The posture of the virtual character in the last N animation frames in the adjacent animation can represent the starting posture of the virtual character who is going to speak the first text segment. Selecting the animation with the smallest difference from the starting posture from multiple alternative animations as the first animation can ensure that the first animation and the adjacent animation are smoothly connected and the posture changes of the virtual character are natural.
[0206] Alternatively, the alternative animation includes multiple animation frames, and the adjacent animation is the animation of the text segment after the first text segment, then the last N animation frames are selected from the alternative animation, and the first N animation frames are selected from the adjacent animation, and the sum, average value or variance of the difference between the last N animation frames in the alternative animation and the first N animation frames in the adjacent animation are calculated as the difference information between the alternative animation and the adjacent animation. Wherein, N is an integer greater than 1. The first animation to be generated needs to be the preceding animation of the adjacent animation, and the posture of the virtual character in the last N animation frames in the first animation represents the starting posture of the virtual character in the adjacent animation. Selecting an animation with a smaller difference from the first N animation frames in the adjacent animation as the first animation from multiple alternative animations can ensure that the first animation is smoothly connected with the adjacent animation, and the posture change of the virtual character is natural.
[0207] 305. The computer device processes the audio segment corresponding to the second text segment through the animation generation model to obtain a second animation.
[0208] 306. The computer device generates a target animation of the virtual character based on the first animation and the second animation.
[0209] The process of steps 305-306 is the same as that of steps 203-204, and will not be described in detail in this embodiment of the present application.
[0210] For example, animation libraries such as Figure 4 As shown in the figure, the animation library stores animations corresponding to three semantic tags "hello", "over there" and "me" ( Figure 4 Only one animation corresponding to each semantic label is shown). The spoken text is "I saw Xiao Hei sneaking around in the garden today", see Figure 5 , the computer device uses a large language model to perform semantic analysis based on the prompt words and the spoken text, and determines that the spoken text includes two first text segments: "I" and "over there", and the first semantic tags corresponding to these two first text segments are: "I" and "over there", respectively. Then, based on the determined semantic tags, animation matching is performed in the animation library, and animation 1 matching "I" and animation 3 matching "over there" are retrieved. Animation 1 includes a picture of a virtual character pointing to himself, and animation 3 includes a picture of a virtual character pointing to "over there". In addition, the computer device uses TTS technology to convert the spoken text into spoken audio, and according to the semantic analysis results of the large language model, it is determined that "I saw Xiaohei in the garden today" and "sneaky" are second text segments, and the audio segments corresponding to the two second text segments are extracted from the spoken audio, and animation 2 and animation 4 are generated respectively through the S2G model. Then, according to the arrangement order of each text segment in the spoken text, animation 1 to animation 4 are spliced to obtain the target animation.
[0211] The embodiment of the present application implements an animation generation scheme based on a large language model and multimodal drive. Animation is created through a motion capture device, thereby constructing a large-scale animation library containing semantic tags. By summarizing the semantic tags that may be mentioned in daily conversations, prompt words are designed and a large language model is trained as a semantic analysis system. For any spoken text, the spoken text is first divided and labeled with semantic tags through a large language model. For text segments with semantic tags in the spoken text, animations are retrieved from the animation library through animation matching. For text segments without semantic tags in the spoken text, the audio segments corresponding to the text segments without semantic tags are processed through the animation generation model to generate rhythmic animations, and the animations of the divided text segments are spliced in chronological order to obtain the target animation.
[0212] In an embodiment of the present application, a large language model is used to perform semantic analysis on a speech text of a virtual character, thereby dividing the speech text into a first text segment with a first semantic tag and a second text segment without a semantic tag. For the first text segment, a first animation matching the first semantic tag is retrieved from an animation library, and the semantics expressed by the virtual character in the first animation matches the semantics of the first text segment, thereby matching the semantics of the speech text. For the second text segment, a second animation is generated through an animation generation model, and in the animation corresponding to the second text segment included in the second animation, the rhythm of the virtual character matches the speech rhythm of the second text segment. In the target animation generated based on the first animation and the second animation, the semantics expressed by the virtual character matches the semantics of the speech text, and the rhythm of the virtual character matches the speech rhythm of the second text segment, that is, a semantically matching and rhythmic animation is generated, which enriches the amount of information contained in the animation and improves the animation playback effect.
[0213] Furthermore, in the embodiment of the present application, a semantic analysis of the spoken text is first performed through a large language model to ensure that a first text segment with a semantic label and a second text segment without a semantic label in the spoken text can be accurately identified, and then the animation library and the animation generation model are used to process the first text segment and the second text segment respectively to generate a rhythmic and semantically matching animation, which does not entirely rely on the internal processing process of the animation generation model. Therefore, the animation generation method provided in the embodiment of the present application has strong interpretability and controllability, improves accuracy and diversity, and does not generate errors in the results.
[0214] Furthermore, in the embodiment of the present application, the large language model can not only identify the semantic tags of the text segments in the spoken text, but also identify the weights of the text segments, and use the weights to represent the importance of the text segments to the spoken text, so that the first text segment and the second text segment can be determined based on the weights of the text segments, and the more important text segment can be used as the first text segment, and the less important text segment can be used as the second text segment, thereby improving the accuracy of the semantic analysis and thus improving the accuracy of the generated animation.
[0215] The following is an example of the application scenario of the embodiment of the present application:
[0216] The user opens the animation production program in the computer device, and the computer device displays the animation production program as shown in FIG. Figure 6 The animation generation interface shown includes an audio input area, a skeleton input area, a parameter selection area, an animation output area and generation controls.
[0217] Among them, the audio input area is used to input spoken audio. The user selects the storage location of the audio in the storage location bar in the audio input area, and determines the range of the spoken audio in the audio by sliding and selecting in the sliding bar below the storage location bar, thereby determining the spoken audio.
[0218] The skeleton input area is used to input the skeleton data of the virtual character, such as the parameters of each joint of the virtual character. The user selects the storage location of the skeleton data in the storage location bar in the skeleton input area, and the name of the selected skeleton data is displayed in the name bar.
[0219] The parameter selection area is used to determine the type of virtual character and the style and type of animation. There are two types of virtual characters: standing alone and talking while walking. If the user selects one of the types, the virtual character in the subsequent animation will conform to the selected type. The animation style can be selected from the template style. Figure 6 In the example of only one template style, multiple template styles can be set for users to choose in other embodiments. The type of animation varies according to different parts of the body. For example, if the animation type is a gesture type, it means that a gesture animation needs to be generated, and in the gesture animation, the virtual character expresses different semantics by making different gestures. Alternatively, the animation type is a head type, which means that a head animation needs to be generated, and in the head animation, the virtual character expresses different semantics by making different gestures with the head. Alternatively, the animation type is a facial type, which means that a facial animation needs to be generated, and in the facial animation, the virtual character expresses different semantics by making different expressions with the face.
[0220] The animation output area is used to input the storage location, name, and frame rate of the generated animation. The user selects the storage location of the generated animation in the storage location column, enters the name of the animation in the name column, and selects one of the three frame rates provided by the computer device: 30 Hz (Hertz), 60 Hz, and 120 Hz. The computer device obtains the input storage location, and stores the animation to the storage location after the animation is subsequently generated. In addition, the computer device obtains the name of the input animation, and names the animation to the name after the animation is subsequently generated. In addition, the selected frame rate is obtained, and the animation is played at the frame rate after the animation is subsequently generated.
[0221] The generation control is used to instruct the generation of an animation. The user triggers the generation control after inputting various parameters in the animation generation interface. The computer device responds to the triggering operation of the generation control, obtains the various parameters input in the animation generation interface, and generates an animation based on the various parameters obtained. In addition, the intermediate data generated in the process of generating the animation can also be displayed in the animation generation interface, such as the identified semantic tags and corresponding weights, the retrieved first animation, the generated second animation, etc., for the user to view or edit. The generated target animation can also be displayed in the animation generation interface for the user to preview or edit.
[0222] In the embodiment of the present application, by inputting various parameters in the animation generation interface, the computer device can be triggered to generate an animation that meets the requirements, which greatly improves the efficiency of animation production and saves labor costs and time costs.
[0223] Figure 7 This is a schematic diagram of the structure of an animation generation model provided in an embodiment of the present application, see Figure 7 ,The animation generation model includes a first encoder, a second encoder and a conditional autoregressive model. Figure 8 This is a schematic diagram of the processing flow of an animation generation model provided in an embodiment of the present application, see Figure 8 The process of processing the audio segment corresponding to the second text segment through the animation generation model to obtain the second animation includes:
[0224] 801. A computer device extracts audio features of an audio segment corresponding to a second text segment through a first encoder, where the audio features represent a speaking rhythm of the second text segment.
[0225] In the embodiment of the present application, the animation generation model only processes the audio segment corresponding to the second text segment, and no longer processes the audio segment corresponding to the first text segment, so the generated second animation is the animation corresponding to the second text segment.
[0226] Optionally, a first audio feature of the audio segment is extracted, where the first audio feature includes at least one of the following: a logarithmic amplitude spectrum, a Mel frequency, or the energy of each audio frame contained in the audio segment. The first audio feature is sampled, and the sampled first audio feature is processed by a first encoder to obtain a second audio feature of the audio segment.
[0227] In which, when the first audio feature includes at least two of the logarithmic amplitude spectrum, the Mel frequency, or the energy of each audio frame contained in the audio segment, the at least two acquired audio features can be concatenated in a fixed order to obtain the first audio feature.
[0228] Among them, the sampling number can be set by the default of the computer device or determined by the structure of the animation generation model. By sampling the first audio feature, the number of features contained in the sampled first audio feature can be made to be the target number, so that the target number of features is processed by the first encoder to obtain the second audio feature.
[0229] In addition, the first encoder is a CNN or other types of neural networks, which is not limited in the embodiments of the present application.
[0230] 802. The computer device extracts style features of the template-style animation through a second encoder, where the style features represent the template style to which the template-style animation belongs.
[0231] In the template style animation, the virtual character makes movements or expressions that belong to the template style, and the template style animation can be set by default by the computer device, or selected by the user from multiple template style animations. The template style animation can be produced by a technician, for example, the technician determines the template style, and the assistant makes movements that conform to the template style while wearing a motion capture device, so as to record the template style animation through the motion capture device, or the assistant makes expressions that conform to the template style and shoots the assistant through a camera to obtain the template style animation.
[0232] For example, the computer device pre-sets one or more template style animations corresponding to styles, such as old man style, child style, neutral style, angry style, tired style, etc. The animation generation interface is displayed, and the animation generation interface displays a style input area. The computer device obtains the style input by the user in the style input area, thereby obtaining the template style animation corresponding to the style.
[0233] Optionally, parameters of each joint of the virtual character are extracted from the template style animation, the parameters of each joint of the virtual character are spliced to obtain a first style feature, the first style feature is normalized, and the normalized first style feature is processed by a second encoder to obtain a second style feature of the template style animation.
[0234] Among them, the parameters of the joints include global displacement, local displacement, global rotation speed, local rotation speed, global displacement speed, local displacement speed, displacement and rotation speed of the root joint, etc. The global displacement of the joint refers to the displacement of the joint in the template style animation, and the local displacement of the joint refers to the relative displacement of the joint relative to the root joint. The other parameters are similar and will not be repeated here. The arrangement order can be determined for the multiple joints of the virtual character, and the parameters of each joint are spliced according to the arrangement order of the multiple joints to obtain the first style feature. And by normalizing the first style feature, the influence of other factors other than style in different template style animations (such as the size of the joint, etc.) can be removed, so that the encoded second style feature can more accurately represent the style of the template style animation.
[0235] In addition, the second encoder is a CNN or other type of neural network, which is not limited in the embodiments of the present application.
[0236] 803. The computer device generates a second animation based on the audio features and the style features through a conditional autoregressive model.
[0237] The virtual character in the second animation performs actions belonging to the template style, and the rhythm of the virtual character in the animation corresponding to the second text segment in the second animation matches the speaking rhythm of the second text segment.
[0238] Optionally, adjacent animations of the audio segment are obtained, and a second animation is generated based on audio features, style features and adjacent animations through a conditional autoregressive model. The adjacent animation is an animation of an adjacent text segment of the second text segment, including at least one of the following: an animation of a text segment before the second text segment and an animation of a text segment after the second text segment. The conditional autoregressive model will not only generate an animation with audio rhythm and style that meets the requirements based on audio features and style features, but will also consider adjacent animations to ensure that the generated animation is smoothly connected with the adjacent animation.
[0239] For example, the adjacent animation includes character skeleton data of at least one animation frame, and the adjacent animation is an animation of a text segment located before a second text segment, then the last N animation frames are selected from the adjacent animation, and the audio features, style features, and character skeleton data of the last N animation frames in the adjacent animation are processed through the animation generation model to obtain a second animation, so that the animation playback effect obtained by splicing the last N animation frames in the adjacent animation with the second animation is smooth and natural.
[0240] Alternatively, if the adjacent animation is an animation of a text segment located after the second text segment, the first N animation frames are selected from the adjacent animation, and the audio features, style features, and character skeleton data of the first N animation frames in the adjacent animation are processed through the animation generation model to obtain a second animation, so that the animation playback effect obtained by splicing the second animation with the first N animation frames in the adjacent animation is smooth and natural.
[0241] In a possible implementation, the structural diagram of the animation generation model is as follows: Fig. 9 As shown, the conditional autoregressive model includes a neural network and a recurrent decoder, and the recurrent decoder includes at least two layers of gated recurrent neural units. The conditional autoregressive model is used to generate a second animation based on audio features, style features, and adjacent animations, including: processing the style features and adjacent animations through a neural network to obtain hidden layer features; processing the audio features and hidden layer features through a recurrent decoder to obtain parameters of each joint of the virtual character in each animation frame; and generating the second animation based on the parameters of each joint of the virtual character in each animation frame.
[0242] Among them, the parameters of the joint include global displacement, local displacement, global rotation speed, local rotation speed, global displacement speed, local displacement speed, displacement and rotation speed of the root joint, etc. The global displacement of the joint refers to the displacement of the joint in the template style animation, and the local displacement of the joint refers to the relative displacement of the joint relative to the root joint. The other parameters are the same and will not be repeated here.
[0243] In addition, the neural network is a CNN or other types of neural networks, which is not limited in the embodiments of the present application.
[0244] In another embodiment, the animation generation model processes the speech audio corresponding to the speech text, and the processing process is the same as the above steps 801-803. The generated second animation includes the animation corresponding to the first text segment and the animation corresponding to the second text segment. Subsequently, the animation corresponding to the first text segment in the second animation is replaced with the first animation to obtain the target animation.
[0245] Optionally, the process of training the animation generation model includes: obtaining sample audio and sample animation, where the sample animation can be produced by a technician, and the rhythm of the virtual character in the sample animation matches the speaking rhythm in the sample audio. Based on the sample audio and sample animation, the animation generation model is trained so that the animation generation model has the ability to generate animations with matching rhythm based on the audio.
[0246] By collecting multiple sample audios and corresponding sample animations, and iteratively training the animation generation model, the accuracy of the animation generation model can be improved until the animation generation model meets the stop training condition, indicating that the accuracy of the animation generation model meets the requirement, and then any audio can be processed by the animation generation model to obtain a rhythm-matched animation. Among them, the stop training condition includes: the number of iterative training reaches the target number, or the accuracy of the animation generation model is not less than the target accuracy, or the loss value of the animation generation model is less than the target loss value, etc.
[0247] Optionally, the animation generation model includes model parameters, for example, the model parameters of the animation generation model include parameters of the first encoder, the second encoder, the conditional autoregressive model, etc., but the initialized model parameters are not accurate enough, resulting in low accuracy of the animation generation model. In the process of training the animation generation model based on the sample audio, the sample audio is input into the animation generation model, and the animation generation model processes the sample audio according to the current model parameters to generate a predicted animation. By comparing the predicted animation with the sample animation, a loss value is determined. The loss value indicates the degree of deviation of the animation generated by the animation generation model. Based on the loss value, the model parameters of the animation generation model are adjusted to improve the accuracy of the adjusted animation generation model. This is an iterative training process. After multiple iterative trainings, it can be ensured that the accuracy of the animation generation model meets the requirements. Afterwards, if the speech audio or the audio segment corresponding to the second text segment is input into the animation generation model, the animation generation model processes the input speech audio or audio segment according to the trained model parameters to obtain the second animation.
[0248] In an embodiment of the present application, a large language model is used to perform semantic analysis on a speech text of a virtual character, thereby dividing the speech text into a first text segment with a first semantic tag and a second text segment without a semantic tag. For the first text segment, a first animation matching the first semantic tag is retrieved from an animation library, and the semantics expressed by the virtual character in the first animation matches the semantics of the first text segment, thereby matching the semantics of the speech text. For the second text segment, a second animation is generated through an animation generation model, and in the animation corresponding to the second text segment included in the second animation, the rhythm of the virtual character matches the speech rhythm of the second text segment. In the target animation generated based on the first animation and the second animation, the semantics expressed by the virtual character matches the semantics of the speech text, and the rhythm of the virtual character matches the speech rhythm of the second text segment, that is, a semantically matching and rhythmic animation is generated, which enriches the amount of information contained in the animation and improves the animation playback effect.
[0249] Furthermore, the audio segment corresponding to the second text segment, the template-style animation and the adjacent animation are processed through the animation generation model, thereby ensuring that in the generated second animation, the rhythm of the virtual character matches the speaking rhythm of the second text segment, the style of the virtual character belongs to the template style, and the connection between the second animation and the adjacent animation is smooth and natural, thereby improving the playback effect of the target animation finally generated.
[0250] Take the spoken text "I don't know" as an example. Fig.10 is a schematic diagram of three animation frames of an animation generated by a related technology, see Fig.10 In the related art, even if the animation is generated through a model, it can only ensure that the rhythm of the animation matches the rhythm of the spoken audio. However, the actions performed by the virtual character in the animation are random, and it cannot be guaranteed that the actions performed match the semantics of the spoken text "unknown". Fig.11 is a schematic diagram of three animation frames of the animation generated by the embodiment of the present application, see Fig.11 The embodiment of the present application generates a high-quality shaking head animation by performing semantic analysis on the spoken text "I don't know" and then searching it through an animation matching method. That is, the virtual character in the animation makes a shaking head movement, which matches the semantics of the spoken text. The generated animation is closer to the movements of people in daily conversations.
[0251] Fig.12 is a structural diagram of an animation generation device provided in an embodiment of the present application. Fig.12 , the device comprises:
[0252] The semantic analysis module 1201 is used to perform semantic analysis on the speech text of the virtual character through a large language model to determine a first text segment in the speech text and a first semantic tag corresponding to the first text segment, where the first semantic tag represents the semantics of the first text segment;
[0253] A retrieval module 1202 is used to retrieve a first animation matching a first semantic tag from an animation library, wherein the semantics expressed by the virtual character in the first animation matches the semantics represented by the first semantic tag;
[0254] The first animation generation module 1203 is used to process the audio segment corresponding to the second text segment through the animation generation model to obtain a second animation, wherein the second text segment is a text segment without a corresponding semantic tag in the spoken text, and the second animation includes an animation corresponding to the second text segment, and the rhythm of the virtual character in the animation corresponding to the second text segment matches the speaking rhythm of the second text segment;
[0255] The second animation generating module 1204 is used to generate a target animation of the virtual character based on the first animation and the second animation.
[0256] In an embodiment of the present application, a large language model is used to perform semantic analysis on a speech text of a virtual character, thereby dividing the speech text into a first text segment with a first semantic tag and a second text segment without a semantic tag. For the first text segment, a first animation matching the first semantic tag is retrieved from an animation library, and the semantics expressed by the virtual character in the first animation matches the semantics of the first text segment, thereby matching the semantics of the speech text. For the second text segment, a second animation is generated through an animation generation model, and in the animation corresponding to the second text segment included in the second animation, the rhythm of the virtual character matches the speech rhythm of the second text segment. In the target animation generated based on the first animation and the second animation, the semantics expressed by the virtual character matches the semantics of the speech text, and the rhythm of the virtual character matches the speech rhythm of the second text segment, that is, a semantically matching and rhythmic animation is generated, which enriches the amount of information contained in the animation and improves the animation playback effect.
[0257] Alternatively, see Fig.13 , the semantic analysis module 1201 includes:
[0258] A semantic analysis unit 1211 is used to perform semantic analysis on the spoken text through a large language model to obtain semantic labels and weights of at least two third text segments in the spoken text, where the weights represent the importance of the third text segments to the spoken text;
[0259] A text selection unit 1212, configured to select a first text segment from at least two third text segments based on weights of the at least two third text segments, wherein the weight of the first text segment is greater than the weight of the unselected third text segment;
[0260] The determining unit 1213 is configured to determine the unselected third text segment as the second text segment.
[0261] Alternatively, see Fig.13 , the semantic analysis module 1201 further includes:
[0262] A display unit 1214, configured to display semantic labels and weights of at least two third text segments in the animation generation interface;
[0263] The editing unit 1215 is used to obtain the semantic label and weight of the edited third text segment in response to the editing operation in the animation generation interface.
[0264] Optionally, the semantic analysis module 1201 is used to perform semantic analysis on the spoken text based on the prompt word through a large language model to determine a first text segment and a first semantic tag corresponding to the first text segment;
[0265] Among them, the prompt word represents the semantic analysis task sent to the large language model.
[0266] Optionally, the training process of the large language model includes:
[0267] Obtaining sample text, where the sample text includes semantic keywords;
[0268] Obtaining semantic labels and weights of semantic keywords, where the weights represent the importance of semantic keywords to sample texts;
[0269] Train a large language model based on sample text, semantic labels and weights of semantic keywords.
[0270] Alternatively, see Fig.13 , the retrieval module 1202 includes:
[0271] A retrieval unit 1221 is used to retrieve a plurality of candidate animations matching the first semantic tag from an animation library;
[0272] A first animation selection unit 1222, configured to select a first animation from a plurality of candidate animations based on a playback duration of the first text segment and a playback duration of the plurality of candidate animations;
[0273] The difference between the playing time of the first animation and the first text segment is smaller than the target threshold; or the difference between the playing time of the first animation and the first text segment is smaller than the difference between the playing time of other candidate animations and the first text segment.
[0274] Optionally, the retrieval module 1202 includes:
[0275] A retrieval unit 1221 retrieves a plurality of candidate animations matching the first semantic tag from an animation library;
[0276] The adjacent animation obtaining unit 1223 is used to obtain the adjacent animation of the first text segment, where the adjacent animation is the animation of the adjacent text segment of the first text segment;
[0277] A difference determination unit 1224, used to respectively determine difference information between each candidate animation and an adjacent animation;
[0278] The second animation selection unit 1225 is configured to select a first animation from a plurality of candidate animations based on difference information of the plurality of candidate animations.
[0279] Alternatively, see Fig.13 The animation generation model includes a first encoder, a second encoder and a conditional autoregressive model; the first animation generation module 1203 includes:
[0280] A first feature extraction unit 1231 is used to extract an audio feature of the audio segment through a first encoder, where the audio feature represents a speaking rhythm of the second text segment;
[0281] A second feature extraction unit 1232 is used to extract a style feature of the template style animation through a second encoder, where the style feature indicates the template style to which the template style animation belongs, and in the template style animation, the virtual character performs an action belonging to the template style;
[0282] The animation generating unit 1233 is used to generate a second animation based on the audio features and the style features through a conditional autoregressive model, in which the virtual character performs actions belonging to the template style.
[0283] Optionally, the first feature extraction unit 1231 is used to:
[0284] Extracting a first audio feature of the audio segment, the first audio feature comprising at least one of the following: a logarithmic amplitude spectrum, a Mel frequency, or energy of each audio frame included in the audio segment;
[0285] sampling a first audio feature;
[0286] The sampled first audio feature is processed by the first encoder to obtain a second audio feature of the audio segment.
[0287] Optionally, the second feature extraction unit 1232 is used to:
[0288] Extract the parameters of each joint of the virtual character from the template style animation;
[0289] The parameters of each joint of the virtual character are spliced to obtain a first style feature;
[0290] Normalizing the first style feature;
[0291] The normalized first style feature is processed by a second encoder to obtain a second style feature of the template style animation.
[0292] Optionally, the animation generating unit 1233 is used to:
[0293] Obtaining an adjacent animation of the audio segment, where the adjacent animation is an animation of an adjacent text segment of the second text segment;
[0294] A second animation is generated based on audio features, style features and adjacent animations through a conditional autoregressive model.
[0295] Optionally, the conditional autoregressive model includes a neural network and a recurrent decoder, the recurrent decoder includes at least two layers of gated recurrent neural units, and the animation generation unit 1233 is used to:
[0296] The style features and adjacent animations are processed through a neural network to obtain hidden layer features;
[0297] The audio features and hidden layer features are processed by a loop decoder to obtain the parameters of each joint of the virtual character in each animation frame;
[0298] A second animation is generated based on the parameters of each joint of the virtual character in each animation frame.
[0299] Optionally, the second animation is an animation corresponding to the second text segment, and the second animation generation module 1204 is used to splice the first animation and the second animation to obtain a target animation.
[0300] Optionally, the first animation generation module 1203 is used to process the speech audio corresponding to the speech text through the animation generation model to obtain a second animation, the second animation includes an animation corresponding to the first text segment and an animation corresponding to the second text segment, and the rhythm of the virtual character in the second animation matches the speaking rhythm of the speech text;
[0301] The second animation generating module 1204 is used to replace the animation corresponding to the first text segment in the second animation with the first animation to obtain a target animation.
[0302] Alternatively, see Fig.13 , the device further comprises:
[0303] The audio acquisition module 1205 is used to acquire the speech audio of the virtual character;
[0304] A first conversion module 1206, configured to convert the speech audio into speech text;
[0305] The device also includes:
[0306] The audio extraction module 1207 is used to extract an audio segment corresponding to the second text segment from the spoken audio.
[0307] Alternatively, see Fig.13 , the device further comprises:
[0308] A text acquisition module 1208 is used to acquire the speech text of the virtual character;
[0309] The second conversion module 1209 is configured to convert the second text segment into an audio segment.
[0310] It should be noted that the animation generation device provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the animation generation device provided in the above embodiment and the animation generation method embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0311] An embodiment of the present application also provides a computer device, which includes a processor and a memory, wherein at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed in the animation generation method of the above embodiment.
[0312] Optionally, the computer device is provided as a terminal. Fig.14 A schematic diagram of the structure of a terminal 1400 provided by an exemplary embodiment of the present application is shown.
[0313] The terminal 1400 includes a processor 1401 and a memory 1402 .
[0314] The processor 1401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 1401 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1401 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0315] The memory 1402 may include one or more computer-readable storage media, which may be non-transitory. The memory 1402 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1402 is used to store at least one computer program, which is used by the processor 1401 to implement the animation generation method provided in the method embodiment of the present application.
[0316] In some embodiments, the terminal 1400 may further optionally include: a peripheral device interface 1403 and at least one peripheral device. The processor 1401, the memory 1402 and the peripheral device interface 1403 may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface 1403 via a bus, a signal line or a circuit board. Optionally, the peripheral device includes: at least one of a radio frequency circuit 1404, a display screen 1405, a camera assembly 1406 and a power supply 1407.
[0317] The peripheral device interface 1403 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 1401 and the memory 1402. In some embodiments, the processor 1401, the memory 1402, and the peripheral device interface 1403 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1401, the memory 1402, and the peripheral device interface 1403 may be implemented on a separate chip or circuit board, which is not limited in this embodiment.
[0318] The radio frequency circuit 1404 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1404 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1404 converts the electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1404 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. The radio frequency circuit 1404 can communicate with other devices through at least one wireless communication protocol. The wireless communication protocol includes, but is not limited to: a metropolitan area network, various generations of mobile communication networks (2G, 3G, 4G and 5G), a wireless local area network and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1404 may also include circuits related to NFC (Near Field Communication), which is not limited in this application.
[0319] The display screen 1405 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1405 is a touch display screen, the display screen 1405 also has the ability to collect touch signals on the surface or above the surface of the display screen 1405. The touch signal can be input to the processor 1401 as a control signal for processing. At this time, the display screen 1405 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, the display screen 1405 can be one, set on the front panel of the terminal 1400; in other embodiments, the display screen 1405 can be at least two, respectively set on different surfaces of the terminal 1400 or in a folding design; in other embodiments, the display screen 1405 can be a flexible display screen, set on a curved surface or a folding surface of the terminal 1400. Even, the display screen 1405 can also be set to a non-rectangular irregular shape, that is, a special-shaped screen. The display screen 1405 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0320] The camera assembly 1406 is used to capture images or videos. Optionally, the camera assembly 1406 includes a front camera and a rear camera. The front camera is arranged on the front panel of the terminal 1400, and the rear camera is arranged on the back of the terminal 1400. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize the panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 1406 may also include a flash. The flash can be a monochrome temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.
[0321] The power supply 1407 is used to power various components in the terminal 1400. The power supply 1407 can be an alternating current, a direct current, a disposable battery, or a rechargeable battery. When the power supply 1407 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0322] Those skilled in the art will understand that Fig.14The structure shown in the figure does not constitute a limitation on the terminal 1400, and the terminal 1400 may include more or less components than those shown in the figure, or combine some components, or adopt a different component arrangement.
[0323] Optionally, the computer device is provided as a server. Fig.15 This is a schematic diagram of the structure of a server provided in an embodiment of the present application. The server 1500 may have relatively large differences due to different configurations or performances, and may include one or more processors (Central Processing Units, CPU) 1501 and one or more memories 1502, wherein the memory 1502 stores at least one computer program, and the at least one computer program is loaded and executed by the processor 1501 to implement the methods provided in the above-mentioned various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output, and the server may also include other components for implementing device functions, which will not be described in detail here.
[0324] An embodiment of the present application further provides a computer-readable storage medium, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to implement the operations performed by the animation generation method of the above embodiment.
[0325] The embodiment of the present application also provides a computer program product, including a computer program, which is loaded and executed by a processor to implement the operations performed by the animation generation method of the above embodiment.
[0326] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0327] The above description is only an optional embodiment of the embodiments of the present application and is not intended to limit the embodiments of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the protection scope of the present application.
Claims
1. An animation generation method, It is characterized in that The method comprises: Performing semantic analysis on a speech text of a virtual character by using a large language model, determining a first text segment in the speech text and a first semantic tag corresponding to the first text segment, wherein the first semantic tag represents the semantics of the first text segment; Retrieving a first animation matching the first semantic tag from an animation library, wherein the semantics expressed by the virtual character in the first animation matches the semantics represented by the first semantic tag; Processing an audio segment corresponding to a second text segment through an animation generation model to obtain a second animation, wherein the second text segment is a text segment without a corresponding semantic tag in the spoken text, the second animation includes an animation corresponding to the second text segment, and a rhythm of the virtual character in the animation corresponding to the second text segment matches a speaking rhythm of the second text segment; A target animation of the virtual character is generated based on the first animation and the second animation.
2. The method according to claim 1, It is characterized in that The performing semantic analysis on the speech text of the virtual character by using the large language model to determine a first text segment in the speech text and a first semantic label corresponding to the first text segment includes: Performing semantic analysis on the spoken text by using the large language model to obtain semantic labels and weights of at least two third text segments in the spoken text, wherein the weights represent the importance of the third text segments to the spoken text; Based on the weights of the at least two third text segments, selecting the first text segment from the at least two third text segments, the weight of the first text segment being greater than the weight of the third text segment that is not selected; The unselected third text segment is determined as the second text segment.
3. The method according to claim 1, It is characterized in that The training process of the large language model includes: Acquire a sample text, wherein the sample text includes semantic keywords; Acquire a semantic label and a weight of the semantic keyword, wherein the weight indicates the importance of the semantic keyword to the sample text; The large language model is trained based on the sample text, the semantic labels and weights of the semantic keywords.
4. The method according to claim 1, It is characterized in that The retrieving a first animation matching the first semantic tag from an animation library includes: Retrieving a plurality of candidate animations matching the first semantic tag from the animation library; Selecting the first animation from the multiple candidate animations based on the playing duration of the first text segment and the playing duration of the multiple candidate animations; The difference between the playing time of the first animation and the first text segment is less than the target threshold; or, the difference between the playing time of the first animation and the first text segment is less than the difference between the playing time of the other candidate animations and the first text segment.
5. The method according to claim 1, It is characterized in that The retrieving a first animation matching the first semantic tag from an animation library includes: Retrieving a plurality of candidate animations matching the first semantic tag from the animation library; Acquire an adjacent animation of the first text segment, where the adjacent animation is an animation of an adjacent text segment of the first text segment; respectively determining difference information between each candidate animation and the adjacent animation; Based on the difference information of the multiple candidate animations, the first animation is selected from the multiple candidate animations.
6. The method according to claim 1, It is characterized in that The animation generation model includes a first encoder, a second encoder and a conditional autoregressive model; the audio segment corresponding to the second text segment is processed by the animation generation model to obtain a second animation, including: extracting, by the first encoder, audio features of the audio segment, wherein the audio features represent a speaking rhythm of the second text segment; extracting, by the second encoder, style features of the template-style animation, wherein the style features represent the template style to which the template-style animation belongs, and in the template-style animation, the virtual character performs an action belonging to the template style; The second animation is generated based on the audio features and the style features through the conditional autoregressive model, and the virtual character in the second animation performs actions belonging to the style of the template.
7. The method according to claim 6, It is characterized in that The extracting the audio feature of the second audio segment by the first encoder includes: Extracting a first audio feature of the audio segment, where the first audio feature includes at least one of the following: a logarithmic amplitude spectrum, a Mel frequency, or energy of each audio frame included in the audio segment; Sampling the first audio feature; The sampled first audio feature is processed by the first encoder to obtain a second audio feature of the audio segment.
8. The method according to claim 6, It is characterized in that The extracting the style features of the template style animation by the second encoder includes: Extracting parameters of each joint of the virtual character from the template style animation; splicing the parameters of each joint of the virtual character to obtain a first style feature; Normalizing the first style feature; The normalized first style feature is processed by the second encoder to obtain a second style feature of the template style animation.
9. The method according to claim 6, It is characterized in that The step of generating the second animation based on the audio feature and the style feature by using the conditional autoregressive model includes: Acquire an adjacent animation of the audio segment, where the adjacent animation is an animation of an adjacent text segment of the second text segment; The second animation is generated based on the audio feature, the style feature and the adjacent animation by using the conditional autoregressive model.
10. The method according to claim 9, It is characterized in that The conditional autoregressive model includes a neural network and a recurrent decoder, the recurrent decoder includes at least two layers of gated recurrent neural units, and the second animation is generated based on the audio feature, the style feature and the adjacent animation by the conditional autoregressive model, including: Processing the style features and the adjacent animations through the neural network to obtain hidden layer features; Processing the audio features and the hidden layer features through the loop decoder to obtain parameters of each joint of the virtual character in each animation frame; The second animation is generated based on the parameters of each joint of the virtual character in each animation frame.
11. The method according to claim 1, It is characterized in that The second animation is an animation corresponding to the second text segment, and generating a target animation of the virtual character based on the first animation and the second animation includes: The first animation and the second animation are spliced together to obtain the target animation.
12. An animation generating device, It is characterized in that The device comprises: A semantic analysis module, configured to perform semantic analysis on a speech text of a virtual character by using a large language model, and determine a first text segment in the speech text and a first semantic tag corresponding to the first text segment, wherein the first semantic tag represents the semantics of the first text segment; A retrieval module, configured to retrieve a first animation matching the first semantic tag from an animation library, wherein the semantics expressed by the virtual character in the first animation matches the semantics represented by the first semantic tag; a first animation generation module, configured to process an audio segment corresponding to a second text segment through an animation generation model to obtain a second animation, wherein the second text segment is a text segment without a corresponding semantic tag in the spoken text, the second animation includes an animation corresponding to the second text segment, and a rhythm of the virtual character in the animation corresponding to the second text segment matches a speaking rhythm of the second text segment; The second animation generating module is used to generate a target animation of the virtual character based on the first animation and the second animation.
13. A computer device, It is characterized in that The computer device includes a processor and a memory, wherein at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the animation generation method according to any one of claims 1 to 11.
14. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by a processor to implement the operations performed by the animation generation method according to any one of claims 1 to 11.
15. A computer program product comprising a computer program, It is characterized in that The computer program is loaded and executed by a processor to implement the operations performed by the animation generation method according to any one of claims 1 to 11.
Citation Information
Cited By
Animation character motion track management system and method based on semantic analysis
CN122244247A
Animation character motion trajectory management system and method based on semantic analysis
CN122244247B