Animation generation method and apparatus, computer device, storage medium, and program product

By acquiring descriptive text, and using animation to guide agents and decision-making agents to select function libraries and functions from multiple function libraries, animation generation code is generated, solving the problem of low animation generation efficiency and realizing intelligent and automated animation generation.

WO2025223132A1PCT designated stage Publication Date: 2025-10-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/084666
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-26
Filing Date
2025-03-25
Publication Date
2025-10-30

Smart Images

  • Figure CN2025084666_30102025_PF_FP_ABST
    Figure CN2025084666_30102025_PF_FP_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of computers. Disclosed are an animation generation method and apparatus, a computer device, a storage medium, and a program product. The method comprises: acquiring a first description text; on the basis of the first description text, determining a first function library from among a plurality of function libraries, wherein a function in the function library is used for generating an animation; on the basis of the first description text, determining a first function from the first function library; on the basis of the first function, generating an animation generation code corresponding to the first description text, wherein the animation generation code is used for calling the first function to generate an animation; and running the animation generation code corresponding to the first description text to obtain a first animation.
Need to check novelty before this filing date? Find Prior Art

Description

Animation generation methods, apparatus, computer equipment, storage media and program products

[0001] Cross-references to related applications

[0002] This application is based on and claims priority to Chinese Patent Application No. 202410540501.2, filed on April 26, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of computer technology, and in particular to an animation generation method, apparatus, computer equipment, storage medium, and program product. Background Technology

[0004] With the rapid development of computer technology, video games and films are attracting more and more attention, and animation production has become one of the important technologies in related fields.

[0005] In related technologies, animators typically need to design scenes, characters, and character movements in animations using animation production software, and then render manually configured animation resources to obtain the animation. This process requires a lot of manpower and time, resulting in low efficiency in animation generation. Summary of the Invention

[0006] This application provides an animation generation method, apparatus, computer device, storage medium, and program product, which can improve the efficiency of animation generation. The technical solution is as follows:

[0007] This application provides an animation generation method applied to a computer device, the method comprising:

[0008] Obtain the first description text, which is used to describe the animation;

[0009] Based on the first description text, a first function library is determined from multiple function libraries, and the functions in the function library are used to generate animations;

[0010] Based on the first description text, determine the first function in the first function library;

[0011] Based on the first function, animation generation code corresponding to the first descriptive text is generated, and the animation generation code is used to call the first function to generate an animation;

[0012] Run the animation generation code corresponding to the first descriptive text to obtain the first animation, the content of which matches the content described in the first descriptive text.

[0013] This application provides an animation generation apparatus, applied to a computer device, the apparatus comprising:

[0014] The text acquisition module is configured to acquire a first descriptive text, which is used to describe the animation.

[0015] The function library determination module is configured to determine a first function library from multiple function libraries based on the first description text, wherein the functions in the function library are used to generate animation;

[0016] The function determination module is configured to determine a first function from the first function library based on the first description text;

[0017] The code generation module is configured to generate animation generation code corresponding to the first descriptive text based on the first function using a code generation agent. The animation generation code is used to call the first function to generate an animation, and the code generation agent is used to generate the code.

[0018] The animation generation module is configured to run the animation generation code corresponding to the first descriptive text to obtain a first animation, the content of which matches the content described in the first descriptive text.

[0019] This application provides a computer device including a processor and a memory. The memory stores at least one computer program, which is loaded and executed by the processor to perform the operations performed by the animation generation method described above.

[0020] This application provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to perform the operations of the animation generation method described above.

[0021] This application provides a computer program product, including a computer program loaded and executed by a processor to perform the operations performed by the animation generation method described above.

[0022] The solution provided in this application, when generating animation, only requires providing descriptive text to describe the animation. Based on the descriptive text, it intelligently predicts which function libraries to use for animation generation, and then intelligently predicts which functions from those libraries to use. Finally, it intelligently generates executable animation generation code based on the selected first function. This code calls the intelligently selected function to generate the animation. Therefore, by running the animation generation code, the corresponding animation can be obtained. Since the function used to generate the animation is intelligently selected based on the descriptive text, the animation generated by calling that function matches the content described in the descriptive text. This achieves automatic generation of matching animations based on the descriptive text, making the animation production process more intelligent and automated, and improving the efficiency of animation generation. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 is a schematic diagram of an implementation environment provided in an embodiment of this application;

[0025] Figure 2 is a flowchart of an animation generation method provided in an embodiment of this application;

[0026] Figure 3 is a flowchart of another animation generation method provided in an embodiment of this application;

[0027] Figure 4 is a schematic diagram of an animation generation method provided in an embodiment of this application;

[0028] Figure 5 is a flowchart of another animation generation method provided in an embodiment of this application;

[0029] Figure 6 is a schematic diagram of another animation generation method provided in an embodiment of this application;

[0030] Figure 7 is a schematic diagram of the result of an animation generation method provided in an embodiment of this application;

[0031] Figure 8 is a flowchart of another animation generation method provided in an embodiment of this application;

[0032] Figure 9 is a schematic diagram of another animation generation method provided in an embodiment of this application;

[0033] Figure 10 is a flowchart of another animation generation method provided in an embodiment of this application;

[0034] Figure 11 is a schematic diagram of an animation generation device provided in an embodiment of this application;

[0035] Figure 12 is a schematic diagram of another animation generation device provided in an embodiment of this application;

[0036] Figure 13 is a schematic diagram of the structure of a terminal provided in an embodiment of this application;

[0037] Figure 14 is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0039] It is understood that the terms "first," "second," etc., used in this application may be used to describe various concepts herein, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of this application, a first animation may be referred to as a second animation, and similarly, a second animation may be referred to as a first animation.

[0040] "At least one" refers to one or more animation segments. For example, at least one animation segment can be one animation segment, two animation segments, three animation segments, or any integer number of animation segments greater than or equal to one. "Multiple" refers to two or more animation segments. For example, multiple animation segments can be two animation segments, three animation segments, or any integer number of animation segments greater than or equal to two. "Each" refers to each of the at least one animation segment. For example, each animation segment refers to each of the multiple animation segments. If the multiple animation segments are three animation segments, then each animation segment refers to each of the three animation segments.

[0041] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals (including but not limited to signals transmitted between user terminals and other devices) involved in this application have all been fully authorized by the user or relevant parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0042] 1) Artificial Intelligence (AI): This refers to the theories, methods, technologies, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.

[0043] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0044] 2) Natural Language Processing (NLP): This is an important area within computer science and artificial intelligence. It studies various theories and methods that enable effective communication between humans and computers using natural language. NLP involves natural language, that is, the language people use in daily life, and is closely related to linguistics research; it also involves important techniques for model training in computer science and mathematics, such as pre-trained models (PTMs), which evolved from large language models (LLMs) in NLP.

[0045] 3) Pre-trained models: Also known as foundational models or large models, these refer to deep neural networks (DNNs) with large parameters. They are trained on massive amounts of unlabeled data. Leveraging the function approximation capabilities of large-parameter DNNs, PTMs extract common features from the data. Through techniques such as fine-tuning, efficient parameter fine-tuning (PEFT), and prompt-tuning, they are suitable for downstream tasks. Therefore, pre-trained models can achieve ideal results in small-shot or zero-shot scenarios. PTMs are categorized according to the data modality they process, including language models, visual models, speech models, and multimodal models. Examples of language models include ELMO (Embeddings from Language Model), BERT (Bidirectional Encoder Representations from Transformers), and GPT (Generative Pre-trained Transformer).

[0046] 4) Large Language Models: These are large-scale deep learning models that typically use autoregressive loss as the training objective. They enable large language models to predict the next word in a given context, thereby learning to generate grammatically correct and semantically coherent text to understand and generate human language. By learning from massive amounts of text data, large language models can understand the complex patterns and contextual relationships of language, thus generating coherent and relevant text. A key characteristic of large language models is their scale, often with billions or even trillions of parameters, allowing them to capture and simulate the rich diversity and complexity of human language. Large language models excel in various language tasks, including but not limited to text generation, text understanding, machine translation, and sentiment analysis. Multimodal models refer to models that establish feature representations of two or more data modalities. Pre-trained models are important tools for outputting Artificial Intelligence Generated Content (AIGC) and can also serve as a general interface connecting multiple task-specific models. After fine-tuning, large language models can be widely applied to downstream tasks. Natural language processing technologies typically include text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.

[0047] 5) Animation: Also known as the first and second animations described in the embodiments of this application, animation is a visual art form that uses a series of continuous images or frames to create the illusion of motion or change by playing the video at a certain rate per second. These images or frames can be hand-drawn, computer-generated, or captured by photography or other technical means. In the fields of computer science and digital media, animation typically involves using software tools and programming languages ​​to create and control these image sequences. Animation can be used for various purposes, including entertainment, education, advertising, simulation, and visualization. The first and second animations described in the embodiments of this application are the final animation products obtained by running animation generation code. The content of the first animation is completely consistent with the description in the first descriptive text, accurately reproducing the desired visual effects and storyline.

[0048] Pre-trained models (PTMs), also known as foundational models or large models, refer to deep neural networks (DNNs) with large parameters. These are trained on massive amounts of unlabeled data. Leveraging the function approximation capabilities of large-parameter DNNs, PTMs extract common features from the data and are then fine-tuned using techniques such as fine-tuning, parameter-efficient fine-tuning (PEFT), and prompt-tuning, making them suitable for downstream tasks. Therefore, pre-trained models can achieve ideal results in few-shot or zero-shot scenarios. PTMs can be categorized according to the data modality they process, including language models, visual models, speech models, and multimodal models. Multimodal models refer to models that represent features from two or more data modalities. Pre-trained models are important tools for outputting Artificial Intelligence Generated Content (AIGC) and can also serve as a general interface connecting multiple specific task models.

[0049] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning. Pre-trained models are the latest development in deep learning, integrating all of these techniques.

[0050] The emergence of large language models has propelled content creation, including film, animation, and games, into a new era. This application proposes an animation generation method based on an LLM-Agent (Large Language Model-Agent) system. This method aims to convert given descriptive text into 3D rendered animation. It not only precisely controls the virtual objects and scene elements required in the animation but also ensures that the content in the animation matches the content described in the descriptive text, while maintaining narrative coherence and scene consistency. The following will provide a detailed description of the animation generation method provided in this application based on artificial intelligence technology.

[0051] The animation generation method provided in this application can be executed by a computer device. In some embodiments, the computer device is a terminal or a server.

[0052] In some embodiments, the server is a standalone physical server, or a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In some embodiments, the terminal is a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, smart voice interaction device, smart home appliance, or in-vehicle terminal, but is not limited to these.

[0053] In one possible implementation, the computer program involved in the embodiments of this application may be deployed and executed on a computer device, or executed on multiple computer devices located in one location, or executed on multiple computer devices distributed in multiple locations and interconnected through a communication network. Multiple computer devices distributed in multiple locations and interconnected through a communication network can form a blockchain system.

[0054] Figure 1 is a schematic diagram of an implementation environment provided in an embodiment of this application. Referring to Figure 1, the implementation environment includes a terminal 101 and a server 102, which are connected via a wireless or wired network.

[0055] To facilitate understanding, a brief introduction will first be given to the multiple function libraries, animation guidance agents, decision agents corresponding to each function library, and code generation agents involved in the embodiments of this application. The functions in the function libraries are used to generate animations, and each function has a different purpose, responsible for different tasks in the animation generation process. The animation guidance agent is used to select the function library to use when generating the animation corresponding to the descriptive text; the decision agent is used to select the function to use from the function library when generating the animation corresponding to the descriptive text; and the code generation agent is used to generate executable code, which is then used to generate the animation.

[0056] In one possible implementation, as shown in Figure 1, server 102 stores the aforementioned multiple function libraries, animation guidance agents, decision agents corresponding to each function library, and code generation agents. In this implementation, the user inputs a first descriptive text into terminal 101, and terminal 101 sends an animation generation request carrying the first descriptive text to server 102. Upon receiving the animation generation request, server 102, through the animation guidance agent, decision agent, and code generation agent, generates corresponding animation generation code based on the first descriptive text. This animation generation code is used to call functions from the multiple function libraries to generate the animation. Server 102 runs the animation generation code, obtains the first animation, and returns the first animation to terminal 101. Upon receiving the first animation, terminal 101 can then display the first animation to the user.

[0057] In one possible implementation, as shown in Figure 1, an animation creation application runs on terminal 101, and multiple function libraries are stored in terminal 101. Server 102 stores an animation guidance agent, decision agents corresponding to each function library, and a code generation agent. In this implementation, the user inputs a first descriptive text in terminal 101, and terminal 101 sends an animation generation request carrying the first descriptive text to server 102. Upon receiving the animation generation request, server 102, through the animation guidance agent, decision agent, and code generation agent, generates corresponding animation generation code based on the first descriptive text. This animation generation code is used to call functions in the multiple function libraries to generate the animation. Server 102 returns the animation generation code to terminal 101. After receiving the animation generation code, terminal 101 runs the animation generation code in the animation creation application to obtain the first animation, which is then displayed to the user.

[0058] Figure 2 is a flowchart of an animation generation method provided in an embodiment of this application. This embodiment of the application is executed by a computer device. Referring to Figure 2, the method includes:

[0059] 201. The computer device obtains the first description text, which is used to describe the animation.

[0060] The first descriptive text is used to describe the content that needs to be presented in the animation to be generated.

[0061] In some embodiments, the first descriptive text can be input by the user into the computer device. For example, when the user has conceived what content they want to present in the animation, they input descriptive text that describes that content into the computer device. Alternatively, the first descriptive text can also be identified or extracted by the computer device from multimedia data. This application embodiment does not limit the method of obtaining the first descriptive text.

[0062] In some embodiments, the first descriptive text refers to text information provided by the user or computer device to describe the content to be presented in the animation to be generated. This text can be manually entered by the user or automatically identified or extracted by the computer device from other multimedia data. Its main purpose is to provide a clear guide and description for animation generation, helping the computer or animation software understand the user's needs, thereby generating animation content that meets the user's expectations.

[0063] 202. The computer device determines a first function library from multiple function libraries based on the first description text, and the functions in the function library are used to generate animation.

[0064] In some embodiments, each function library corresponds to its own decision agent, with the animation guidance agent used to determine the function library and the decision agent used to determine the function.

[0065] In some embodiments, after obtaining the first description text, the computer device provides the first description text to the animation guidance agent, and the animation guidance agent can then determine the first function library from multiple function libraries based on the first description text.

[0066] In some embodiments, the plurality of function libraries are predefined function libraries. For example, if an animation production application is running on a computer device, multiple function libraries can be predefined in the animation production application so that they can be directly called when generating animations. The first function library is the function library required to generate an animation adapted to the first descriptive text. The number of the first function libraries can be one or more. For example, the plurality of function libraries includes function library 1 and function library 2. If the animation guidance agent determines that function library 1 is the first function library, it means that generating an animation adapted to the first descriptive text only requires using function library 1, and function library 2 is not needed.

[0067] In some embodiments, the animation guidance agent is used to determine the function library corresponding to any descriptive text. Each function library also corresponds to its own decision agent, which is used to determine the function corresponding to any descriptive text in that function library. In some embodiments, both the animation guidance agent and the decision agent can belong to a large language model.

[0068] In some embodiments, a computer device contains multiple function libraries, each containing a set of functions for generating animations. These functions may have different functionalities, such as creating animation objects, setting animation properties, and controlling animation playback. Each function library corresponds to a decision agent. The decision agent is an intelligent component that can select an appropriate function to generate an animation based on the input descriptive text. It may be a machine learning model or a rule-based system. The animation guidance agent is a higher-level intelligent component responsible for selecting the most suitable function library from multiple function libraries for the current task. The animation guidance agent may also be a machine learning model or a rule-based system. The computer device obtains a first descriptive text, which describes the content the user wants to present in the animation. This text can be manually entered by the user or automatically identified or extracted by the computer device from other multimedia data. The computer device provides the first descriptive text to the animation guidance agent. Based on this text, the animation guidance agent analyzes the user's needs and selects the most suitable function library from multiple function libraries for the current task. The animation guidance agent determines the first function library from multiple function libraries based on the first descriptive text. This function library contains a set of functions that can meet the user's needs. Once the initial function library is determined, the computer device can use the decision-making agent within that library to select specific functions to generate the animation. The decision-making agent selects an appropriate function to generate the animation based on the initial descriptive text.

[0069] In some embodiments, the animation guidance agent is a high-level intelligent component responsible for understanding the user's needs and intentions, and selecting the most suitable function library from multiple function libraries for the current task. It typically receives descriptive text provided by the user, analyzes the text content, and then guides the animation generation process based on the analysis results. The main function of the animation guidance agent is to select an appropriate function library to generate animation based on the user's needs. It may use natural language processing techniques to understand the text, and then use machine learning models or rule-based systems to make decisions.

[0070] In some embodiments, the decision agent is an intelligent component associated with a specific function library, responsible for selecting a specific function from the library to generate the animation. Each function library has its own decision agent, which understands the functionality and applicable scenarios of the functions in that library. The main function of the decision agent is to select an appropriate function from the selected function library to perform the animation generation task, based on the instructions of the animation guidance agent and the user's needs. It may use a machine learning model to predict the best function or a rule-based system to make the decision.

[0071] In some embodiments, the Animation Guidance Agent and the Decision Agent are two intelligent components used in computer graphics and artificial intelligence to generate animations. They each have different responsibilities and functions, and work together to enable the computer device to automatically select appropriate function libraries and functions to generate animations based on user needs. The Animation Guidance Agent is responsible for macro-level decisions, selecting function libraries, while the Decision Agent is responsible for micro-level decisions, selecting specific functions.

[0072] As an example, suppose we have a computer device with multiple function libraries, each containing a set of functions for generating different types of animations. There's a character animation library: containing functions for creating and controlling character animations, such as walking, running, and jumping. An environment animation library: containing functions for creating and controlling environmental element animations, such as weather changes, water flow, and fire. An effects animation library: containing functions for creating various visual effects, such as explosions, lighting effects, and particle effects. Each function library has a corresponding decision-making agent that understands the functions' capabilities and applicable scenarios. Now, a user inputs a descriptive text: "I need an animation showing a character walking in a forest and suddenly encountering a heavy rain." The computer device provides this text to the animation guidance agent. The animation guidance agent uses natural language processing to understand the text content and analyzes the user's desired animation, determining that it includes a character walking and a rain scene. Based on this information, the animation guidance agent selects from multiple function libraries. It might prioritize the character animation library because the text mentions walking, while also considering the environment animation library because the text mentions rain. Once the animation guidance agent has identified the character animation library and the environment animation library as the first function libraries, it passes information from these libraries to the corresponding decision-making agents. The decision-making agent for the character animation library will select functions related to character walking, such as the "character walking animation" function. The decision-making agent for the environment animation library will select functions related to rain, such as the "rain animation" function. These two decision-making agents may further adjust the parameters of the functions to ensure the animation is natural and smooth. For example, the character walking animation function might be adjusted to accommodate slippery ground during rain, while the rain animation function might be adjusted to match the visual effects of a forest environment. Ultimately, the computer device uses these selected functions to generate an animation that matches the user's description, showing a character walking in a forest and encountering heavy rain.

[0073] Thus, the method of determining the first function library from multiple function libraries based on the first descriptive text using computer equipment can significantly improve the efficiency and accuracy of animation generation. Through the collaborative work of the animation guidance agent and the decision-making agent, intelligent management of the animation generation process is achieved. The animation guidance agent can understand the user's needs and select the most suitable function library by analyzing the first descriptive text, reducing the tedious process of manually selecting function libraries and improving work efficiency. The decision-making agent, within the selected function library, chooses the most suitable function to generate the animation based on the user's specific needs, ensuring the animation quality and meeting user expectations. It can also adapt to different animation requirements, improving the flexibility and scalability of animation generation. The method of determining the first function library based on the first descriptive text not only improves the efficiency and accuracy of animation generation but also enhances the system's flexibility and scalability, providing a more intelligent and efficient solution for animation production.

[0074] 203. The computer device determines a first function from a first function library based on the first description text.

[0075] In some embodiments, after determining a first function library, the computer device provides a first description text to the decision agent corresponding to the first function library, and the decision agent determines a first function in the first function library based on the first description text.

[0076] In some embodiments, the first function library includes multiple functions, each used to generate animations, but each function is responsible for different functions during the animation generation process. The first function is the function required to generate an animation adapted to the first descriptive text. The number of first functions can be one or more. For example, if the first function library includes functions 1-5, and the decision agent determines that functions 1 and 5 are the first functions, then generating an animation adapted to the first descriptive text only requires functions 1 and 5, and does not require functions 2, 3, and 4.

[0077] In some embodiments, the computer device, through a decision agent corresponding to a first function library, determines a first function from the first function library based on a first descriptive text. The computer device needs to determine which function library is most suitable for the current animation generation task. This is typically done by an animation guidance agent, which analyzes the first descriptive text, understands the user's needs, and selects the most suitable function library. Once the first function library is determined, the computer device passes the first descriptive text to the decision agent corresponding to that function library. This decision agent is specifically responsible for selecting a suitable function from the library. The decision agent receives the first descriptive text and analyzes it in depth. Based on the analysis of the first descriptive text, the decision agent searches and selects the most suitable function from the first function library. This selection process involves matching the function's functional description with the user's needs, or using a machine learning model to predict the best function. After selecting the first function, the decision agent may further adjust the function's parameters to ensure that the generated animation more accurately reflects the user's needs. This may include adjusting the animation's speed, intensity, duration, etc. The computer device uses the selected first function to generate the animation. This process may involve calling other relevant functions in the function library to complete the animation production. The intelligent decision-making process utilizes natural language processing and machine learning technologies, enabling computer devices to automatically select and adjust functions based on user needs, thereby generating high-quality animations.

[0078] As an example, suppose we have a computer device with multiple function libraries, each containing a set of functions for generating different types of animations. The character animation library contains functions for creating and controlling character animations, such as walking, running, and jumping. The environment animation library contains functions for creating and controlling environmental element animations, such as weather changes, water flow, and fire. The special effects animation library contains functions for creating various visual effects, such as explosions, lighting effects, and particle effects. Each function library has a corresponding decision-making agent that understands the functionality and applicable scenarios of the functions in its respective library. A user inputs a descriptive text: "I need an animation showing a character walking in a forest and suddenly encountering a heavy rain." The computer device provides this text to the animation guidance agent. The animation guidance agent uses natural language processing technology to understand the text content and analyzes the user's desired animation, determining that it includes a scene of a character walking and rain. The animation guidance agent then selects from multiple function libraries. It might prioritize the character animation library because the text mentions a character walking, while also considering the environment animation library because the text mentions rain. Once the animation guidance agent has identified the character animation library and the environment animation library as the first function libraries, it passes information from these libraries to the corresponding decision-making agents. The decision-making agent for the character animation library will select functions related to character walking, such as the "character walking animation" function. The decision-making agent for the environment animation library will select functions related to rain, such as the "rain animation" function. These two decision-making agents may further adjust the parameters of the functions to ensure the animation is natural and smooth. For example, the character walking animation function might be adjusted to accommodate slippery ground during rain, while the rain animation function might be adjusted to match the visual effects of a forest environment. The computer device uses these selected functions to generate an animation that matches the user's description, showing a character walking in a forest and encountering heavy rain.

[0079] Thus, by utilizing a decision-making agent corresponding to the first function library and determining the first function from the first function library based on the first descriptive text, the intelligence and personalization of animation generation can be significantly improved. This process, through in-depth analysis and understanding by the decision-making agent, can accurately match user needs with specific functions in the function library, ensuring that the generated animation not only meets user expectations but also achieves professional standards in quality and effect. This not only reduces the tedious process of manually selecting and adjusting functions, improving work efficiency, but also enhances the system's flexibility and scalability, enabling the computer equipment to adapt to various complex animation needs. The method of determining the first function based on the first descriptive text not only improves the efficiency and accuracy of animation generation but also enhances the system's flexibility and scalability, providing a more intelligent and efficient solution for animation production.

[0080] 204. The computer device generates animation generation code corresponding to the first descriptive text based on the first function. The animation generation code is used to call the first function to generate animation.

[0081] In some embodiments, after determining a first function, the computer device provides the first function to a code generation agent, which generates animation generation code based on the first function. This animation generation code is executable code. The code generation agent is used to generate code.

[0082] In some embodiments, the code-generating agent is used to generate corresponding animation generation code based on any function. This code-generating agent belongs to a large language model.

[0083] In some embodiments, the computer device determines a first function most suitable for the current animation generation task through a decision-making agent. This function is selected from a first function library based on the analysis and understanding of a first descriptive text. Once the first function is determined, the computer device provides it to the code generation agent. The code generation agent is a component specifically responsible for converting functions into executable code. After receiving the first function, the code generation agent generates corresponding animation generation code based on the function's definition and functionality. This code is executable and can be directly used to call the first function to generate animation. The generated animation generation code is executed by the computer device, thereby calling the first function and initiating animation generation. The code generation agent can be a large language model. Large language models have powerful natural language processing capabilities, enabling them to understand function definitions and contexts and generate code that meets the requirements. The computer device can automatically translate user requirements into executable code, thereby generating animations that meet the user's expectations.

[0084] As an example, suppose we have a computer device with a function library containing functions for generating various animations. Now, a user inputs a descriptive text: "I need an animation showing a character walking in a forest and suddenly encountering a heavy rain." The computer device analyzes this descriptive text through a decision-making agent and selects the most suitable function library and function. Then, a code-generating agent generates the corresponding animation generation code based on the selected function. The computer device runs this animation generation code, calling the corresponding function and generating the first animation. This animation shows a character walking in a forest and suddenly encountering a heavy rain, matching the content described in the first descriptive text. This demonstrates how the computer device runs the animation generation code corresponding to the first descriptive text to obtain a first animation that matches the content of the descriptive text.

[0085] 205. The computer device runs the animation generation code corresponding to the first description text to obtain the first animation. The content of the first animation matches the content described in the first description text.

[0086] In some embodiments, after obtaining the animation generation code, the computer device runs the animation generation code to call a first function and generate a first animation. Since the first function is the function required to generate an animation that matches the first animation description, the content of the first animation generated by calling the first function matches the content described by the first description text.

[0087] In some embodiments, the computer device runs the animation generation code corresponding to the first descriptive text to obtain a first animation, ensuring that the content of the first animation matches the content of the first descriptive text. Through previous steps, the computer device has already generated the animation generation code corresponding to the first descriptive text. This code is automatically generated based on the user's needs and the selected first function, containing all the instructions and parameters required to generate the animation. The computer device needs a suitable execution environment to run this animation generation code. This may include necessary software libraries, API interfaces, rendering engines, etc., to ensure that the code can correctly call the first function and generate the animation. The computer device runs this animation generation code. During code execution, the first function and other related functions are called in a predetermined order to execute various parts of the animation. For example, the code may first call a function to create the animation background, then create the character, set the animation's actions, and finally add effects and save the animation. A crucial part of the code execution process is calling the first function. This function is analyzed based on the first descriptive text and can generate animation content that matches the user's needs. By calling this function, the computer device can ensure that the generated animation content matches the content described in the first descriptive text.

[0088] As an example, suppose there is a computer device with multiple function libraries, each containing a set of functions for generating different types of animations. The character animation library contains functions for creating and controlling character animations, such as walking, running, and jumping. The environment animation library contains functions for creating and controlling environmental element animations, such as weather changes, water flow, and fire. The effects animation library contains functions for creating various visual effects, such as explosions, lighting effects, and particle effects. A user inputs a first description: "I need an animation showing a character walking in a forest and suddenly encountering a heavy rain." The computer device first obtains the first description text, which describes the desired animation content. The computer device analyzes the first description text to understand the user's needs. Based on the character and forest background mentioned in the text, the computer determines that the character animation library and the environment animation library are needed. After determining the first function libraries, the computer device further analyzes the first description text to determine the specific functions to be used within these libraries. For example, it selects the "character walking" function from the character animation library and the "rain" function from the environment animation library. The computer device generates the corresponding animation generation code based on the selected functions. This code contains instructions to call these functions to generate animations that meet the user's needs. The computer device runs the generated animation code, calling the corresponding functions to begin generating the animation. This process may involve interaction with other functions or components to ensure the integrity and quality of the animation. The computer device generates the first animation, which shows a character walking in a forest and suddenly encountering a heavy rain, consistent with the content described in the first descriptive text.

[0089] The animation generation method provided in this application only requires a descriptive text to describe the animation when needed. An animation guidance agent intelligently predicts which function libraries to use to generate the animation based on the descriptive text. Then, a decision agent corresponding to each function library intelligently predicts which functions from that library to use to generate the animation based on the descriptive text. Finally, a code generation agent intelligently generates executable animation generation code based on the selected functions. This animation generation code calls the intelligently selected functions to generate the animation. Therefore, by running the animation generation code, the corresponding animation can be obtained. Since the functions used to generate the animation are intelligently selected based on the descriptive text, the animation generated by calling these functions matches the content described in the descriptive text. This achieves automatic generation of matching animations based on the descriptive text, making the animation production process more intelligent and automated, and improving the efficiency of animation generation.

[0090] In some embodiments, a detailed description of the desired animation is collected or written. The description text should include information such as the animation's theme, scene, characters, actions, effects, duration, and style. The description text may come from animators, directors, clients, or be automatically generated using natural language processing technology.

[0091] In some embodiments, a first function library is determined from multiple function libraries based on a first descriptive text. This can be achieved through in-depth analysis of the first descriptive text to understand the specific requirements of the animation. This includes the animation's theme, style, required special effects, character movements, scene complexity, etc. Key functionalities that may be needed in animation production are identified, such as 3D modeling, skeletal animation, particle effects, physics simulation, rendering techniques, etc. All available function libraries are listed, which may include graphics libraries, animation libraries, game engines, physics engines, etc. Each function library is evaluated, considering factors such as whether its functionality covers the animation requirements, performance, ease of use, community support, documentation completeness, and license fees. The animation requirements are matched with the functionality of each function library. It is determined which function libraries can provide the required functionality, or whether multiple function libraries need to be combined to meet all requirements. For example, if the animation requires advanced 3D rendering and physics simulation, a game engine such as Unity or Unreal Engine may be chosen, as they provide rich 3D functionality and physics simulation tools. The technical feasibility of the selected function libraries is analyzed, including hardware requirements, programming language compatibility, development environment settings, etc. Ensure the selected library integrates with existing development tools and processes, or assess migration costs and time. Perform performance testing on candidate libraries, especially when handling complex animation scenes. Testing may include key performance indicators such as frame rate, memory usage, and loading time. Consider factors such as library license fees, development costs, and maintenance costs. Evaluate the total cost and potential benefits of using the library long-term. Consider the library's community activity and support. An active community and good technical support can significantly reduce problems during development. Based on the above analysis, select the library most suitable for the requirements of the initial description as the primary library. Determine the rationale for selection, including technical advantages, cost-effectiveness, and community support. Integrate the primary library into the development environment. Conduct preliminary testing to ensure the library functions correctly and meets the basic requirements of animation production. This allows for the determination of the most suitable library from multiple options based on the initial description, laying a solid foundation for subsequent animation production.

[0092] In some embodiments, based on the first descriptive text, a first function is determined in the first function library. Through in-depth analysis of the first descriptive text, the specific requirements of the animation are clarified, including details such as scenes, characters, actions, and special effects. The key functions and effects required to achieve these requirements are determined. The documentation of the first function library is studied to understand all the functions and functional modules it provides. The functions in the function library are categorized and organized to quickly find functions related to the animation requirements. Functions related to the animation requirements are selected. This may include graphics rendering functions, animation transition functions, physics simulation functions, particle system functions, etc. The applicability of the functions, the flexibility of parameter configuration, and compatibility with other functions are considered. The selected functions are evaluated, considering factors such as performance, ease of use, and customizability. It is assessed whether the functions can achieve the required effects or whether multiple functions need to be combined to achieve the goal. Several key functions are selected, and simple prototype code is written for testing. The output of the functions is observed to see if it meets expectations, and parameters are adjusted to optimize the effect. Based on the results of the prototype testing, it is determined which functions can be combined to achieve more complex animation effects. The calling order and interaction logic between functions are designed. The selected functions are optimized to improve the performance and quality of the animation. Adjust the function parameters to ensure smoothness and realism of the animation. Refer to the function library's documentation and sample code to learn best practices and advanced usage. Ensure a deep understanding of the function's usage to avoid common errors and pitfalls. Consider all factors to determine the most suitable function to meet the requirements of the initial description. Determine the selection criteria, including the function's functionality, performance, and ease of use. Integrate the selected function into the animation generation code. Run the code to verify that the function can correctly generate the animation effects that meet the description requirements. This allows you to determine the most suitable function from the initial function library based on the initial description, providing crucial technical support for animation generation.

[0093] In some embodiments, the computer device first needs to understand the interface of the first function, including input parameters, output results, and function behavior. Based on the first description text, the logical flow of the animation is designed. This includes determining the animation's start state, transition states, and end state. The motion trajectories, timelines, and interaction logic of each element in the animation are planned. Based on the animation design, the parameters of the first function are configured. This may include setting the animation's duration, motion path, speed curve, special effects parameters, etc. The parameter configuration is ensured to achieve the expected animation effect. Based on the configured parameters, the computer device generates code that calls the first function. This includes the function name, parameter list, and any necessary context settings. The code should be able to correctly initialize the animation environment, call the first function, and process the function's output. The generated function call code is integrated with the code of other animation elements. This may include scene setting, character animation, special effects rendering, etc. All elements are ensured to work together to form a complete animation. The generated code is analyzed to identify potential performance bottlenecks. The code structure and algorithms are optimized to improve the animation's running efficiency and smoothness. The generated animation code is run on the computer device to test whether the animation effect meets the requirements of the first description text. Parameters and code are adjusted until the animation effect achieves the expected result. Once the animation effect meets the requirements, the computer device outputs the generated code as an animation file, such as a video file or a sequence of animated images. It ensures that the output animation file format meets the requirements and that the quality reaches the standard. The computer device can generate animation generation code corresponding to the first description text based on the first function, and ultimately output animation content that meets the requirements.

[0094] The above embodiments are merely brief descriptions of the animation generation method. Based on these embodiments, multiple function libraries, including action function libraries and scene element function libraries, are used. For a detailed description of the animation generation method, please refer to the embodiment shown in Figure 3 below. Figure 3 is a flowchart of another animation generation method provided by an embodiment of this application. This embodiment is executed by a computer device. Referring to Figure 3, the method includes:

[0095] 301. The computer device acquires the first description text, which is used to describe the animation.

[0096] Step 301 is the same as step 201 above, and will not be repeated here.

[0097] 302. The computer device guides the intelligent agent through animation, determines the first function library from multiple function libraries based on the first descriptive text, and the functions in the function library are used to generate animation. Each function library corresponds to its own decision intelligent agent. The animation guiding intelligent agent is used to determine the function library, and the decision intelligent agent is used to determine the function.

[0098] The animation guidance agent can perform a preliminary analysis of the descriptive text from a comprehensive perspective to determine which function libraries(s) should be used to generate the animation. The animation guidance agent first determines which function library to use, and then the decision agent corresponding to that function library determines which functions(s) within that library should be used. Compared to directly using the decision agent corresponding to each function library to determine whether to use functions from that library and which functions to use, the method in this embodiment reduces the number of calls to the decision agents, thereby improving decision-making efficiency. For example, if there are three function libraries, and the animation guidance agent selects only one, then only the decision agent corresponding to that one function library needs to be called subsequently, without needing to call the decision agents corresponding to the other two function libraries, effectively reducing the number of calls.

[0099] In one possible implementation, the goal of animation generation is to convert descriptive text involving multiple virtual objects into 3D animation. Therefore, during animation generation, it is necessary to design the actions to be performed by each virtual object, as well as decorative scene elements to enhance the animation's appeal and vividness. Based on this, embodiments of this application predefine an action function library and a scene element function library, the details of which are described below.

[0100] (1) Multiple function libraries include an action function library, which includes multiple action functions. Action functions are used to control the actions of virtual objects in the animation.

[0101] In some embodiments, each action function in the action function library corresponds to its own action category. An action function belongs to an action category, and an action category includes at least one action function.

[0102] In some embodiments, the above-described determination of the first function in the first function library based on the first description text can be achieved as follows: when the first function library includes the action function library, the first action function is determined from the plurality of action functions based on the first description text.

[0103] As shown in Table 1 below, the motion categories include special motion, linear motion, curvilinear motion, jumping motion, impact motion, and state recovery operations. Taking linear motion as an example, this linear motion includes motion functions for achieving constant speed motion and motion functions for achieving variable speed motion. In addition, other motion categories also include at least one motion function, and these motion functions all belong to the motion function library.

[0104] In some embodiments, each action function corresponds to a specific action category, and an action function belongs to an action category. An action category includes at least one action function. The step of determining the first action function among the plurality of action functions based on the first description text can be achieved as follows: determining the target action category among the plurality of action categories based on the first description text; and determining the first action function among the action functions belonging to the target action category based on the first description text.

[0105] In some embodiments, action functions in the action function library are categorized into different action categories. For example, action categories might include "movement," "interaction," "attack," "defense," "environmental change," etc. Each action category contains multiple related action functions; for example, the "movement" category might contain action functions such as "walking," "running," and "jumping." The computer device uses natural language processing techniques to analyze the first descriptive text, extracting keywords and phrases related to the actions. These keywords are then matched against predefined action categories to determine which action categories the action described in the text belongs to. For example, if the text mentions "a person walking in a forest," it might match the "movement" category. Based on the results of the text analysis, the computer device determines the target action category most relevant to the first descriptive text. This may involve prioritization; for example, if the text mentions multiple actions simultaneously, the computer device needs to determine which action is primary or performed first. After determining the target action category, the computer device further determines the first action function from among the action functions belonging to that target action category. This can be achieved in various ways, such as selecting the most suitable action function based on the complexity of the action, the applicable scenario, user preferences, etc. Once the first action function is determined, the computer device may configure the function's parameters based on the specific details in the first description text. For example, it might adjust the movement speed, direction, and amplitude. Then, the computer device executes the action function to generate the corresponding animation effect. If the first description text contains multiple actions, the computer device may execute multiple action functions sequentially, generating a series of animation effects and combining them into a complete animation. Finally, the generated animation is output to the user.

[0106] As an example, when the first function library includes the action function library, based on the first descriptive text, the first action function is determined from among the multiple action functions. The computer device uses natural language processing technology to analyze the first descriptive text and extract keywords and phrases related to the action. For example, if the first descriptive text is "A person is walking in a forest and suddenly encounters a heavy rain," then the extracted keywords might include "walking," "encounter," and "heavy rain." The computer device performs semantic understanding on the extracted keywords to determine the action type represented by these keywords. For example, "walking" can correspond to the "movement" action, "encounter" can correspond to "interaction," and "heavy rain" can correspond to the "weather change" action. The computer device matches the understood action type with multiple action functions in the action function library. The action function library may contain various action functions, such as "move the person," "person interacts," and "change the weather." The computer device finds the most suitable action function by comparing the action type with the functional description of the action function. When determining the action function, the computer device also needs to consider the context of the action. For example, if the person is walking in a forest, then it may be necessary to choose a movement function suitable for a forest environment, rather than a movement function suitable for a city or indoor environment. Once the first action function is determined, the computer device may adjust the function's parameters based on specific details in the initial descriptive text. For example, it might adjust the movement speed or the intensity of weather changes to ensure the generated animation is more realistic and meets the user's needs. The computer device confirms the selected first action function and uses it to generate animation generation code. This process may involve combining multiple action functions to achieve complex animation effects. The computer device can determine the first action function from an action function library based on the initial descriptive text, thus preparing to generate an animation that meets the user's requirements. This process demonstrates how computer devices utilize natural language processing and semantic understanding technologies to transform text descriptions into specific action functions, enabling intelligent animation generation.

[0107] (2) Multiple function libraries include a scene element function library, which includes multiple element functions used to control scene elements in the animation.

[0108] In some embodiments, scene elements are used to decorate or embellish the scene in the animation. For example, scene elements include lighting elements, special effects elements, camera focus switching, etc. As shown in Table 1 below, the scene element function library includes element functions for switching cameras, element functions for implementing lighting, element functions for implementing particle effects, element functions for implementing beam effects, element functions for drawing rainbow rain, element functions for adjusting sunlight, etc.

[0109] In some embodiments, the aforementioned plurality of function libraries include a scene element function library, which includes a plurality of element functions used to control scene elements in an animation. The determination of the first function in the first function library based on the first description text can be achieved as follows: when the first function library includes the scene element function library, the first element function is determined from the plurality of element functions based on the first description text.

[0110] In some embodiments, the determination of the first element function among the plurality of element functions based on the first descriptive text can be achieved as follows: the decision agent corresponding to the scene element function library determines the first element function corresponding to the at least one virtual object identifier based on the first descriptive text and at least one virtual object identifier; the virtual object identifier indicates a virtual object appearing in the animation, and the first element function corresponding to the virtual object identifier is used to add scene elements to the virtual object.

[0111] Table 1

[0112] In one possible implementation, the following processing can also be performed: guide the agent through animation to determine duration information based on the first descriptive text.

[0113] In some embodiments, duration information indicates the duration of the animation to be generated, and the duration of the animation to be generated can be determined subsequently based on this duration information.

[0114] In some embodiments, the animation guidance agent outputs guidance information and duration information. This guidance information indicates which function library among multiple function libraries belongs to the first function library and which does not. For example, the guidance information includes an identifier for each function library and a corresponding indicator identifier. If the indicator identifier corresponding to the function library identifier is a first indicator identifier, it indicates that the function library indicated by that identifier belongs to the first function library. If the indicator identifier corresponding to the function library identifier is a second indicator identifier, it indicates that the function library indicated by that identifier does not belong to the first function library. For example, the first indicator identifier is "true" and the second indicator identifier is "false".

[0115] In some embodiments, a plurality of duration markers are predefined in the computer device, each duration marker corresponding to a different frame duration, and the duration information includes one of the plurality of duration markers. For example, the plurality of duration markers include “fast,” “moderate,” “slow,” and “emphasis.”

[0116] In this embodiment, in addition to determining the function library to be called based on the description text, the animation guidance agent can also determine the duration information of the animation to be generated based on the description text, thereby realizing intelligent design of the required animation duration and improving the flexibility and diversity of animation generation.

[0117] In one possible implementation, step 302 includes: guiding the agent through animation to determine candidate function libraries from multiple function libraries based on a first descriptive text; guiding the agent through animation to generate a detection result based on the first descriptive text and candidate function libraries, the detection result indicating whether the candidate function library is accurate, and if the detection result indicates that the candidate function library is incorrect, the detection result also includes the reason for the error; if the detection result indicates that the candidate function library is accurate, the candidate function library is determined as the first function library; if the detection result indicates that the candidate function library is incorrect, guiding the agent through animation to determine the next candidate function library based on the first descriptive text and the reason for the error, until the detection result indicates that the currently obtained candidate function library is accurate, then the currently obtained candidate function library is determined as the first function library.

[0118] The animation guidance agent's processing includes a decision-making process and a self-checking process. During the decision-making process, the animation guidance agent identifies candidate function libraries for intelligent selection. During the self-checking process, the animation guidance agent checks its selected candidate function libraries to determine their accuracy. If accurate, the currently selected candidate function library is adopted as the final function library. If inaccurate, the animation guidance agent needs to re-enter the decision-making process, combining the initial description text and the reason for the error to reselect candidate function libraries, and then re-enter the self-checking process to check the newly selected candidate function libraries until a selected candidate function library successfully passes the self-check. The candidate function library that successfully passes the self-check is then adopted as the final function library.

[0119] In other words, the animation-guided intelligent agent can reflect on and correct its output results. The decision-making process and the self-checking process will alternate until it is confirmed that its output results are error-free.

[0120] In this embodiment, after the animation guidance agent determines the function library based on the description text, it also performs a self-check on the selected function library to verify the accuracy of its decision. If the decision is incorrect, the animation guidance agent re-determines the function library based on the description text and the reason for the error, until the currently determined function library passes the self-check. Therefore, by adding a self-check process, it is possible to effectively ensure that the animation guidance agent selects a more accurate function library, thereby ensuring the adaptability of the subsequently generated animation to the description text.

[0121] To facilitate understanding, the following examples illustrate the inputs and outputs of the animation-guided agent's decision-making and self-checking processes.

[0122] During the first decision-making process, the animation guides the agent's input: the descriptive text: "The sky darkens, and a car equipped with a spotlight slowly drives towards the vase, the car under the focus of the camera."

[0123] During the first decision-making process, the animation guides the agent's output: action function library - "true"; scene element function library - "false"; duration information - "slow".

[0124] During the first self-check, the animation guides the agent's input: Description text: "The sky darkens, a car equipped with a spotlight slowly drives towards the vase, the car is in the camera's focus." Response results: Action function library - "true"; Scene element function library - "false"; Duration information - "slow".

[0125] During the first self-check, the animation-guided agent output: Error; Error reason: "According to the description text, the ambient light should be darkened, but the scene element function library was not selected."

[0126] During the second decision-making process, the animation guides the agent's input: Description text: "The sky darkens, a car equipped with a spotlight slowly drives towards the vase, the car is in focus by the camera"; Error reason: "According to the description text, the ambient light should darken, but the scene element function library was not selected."

[0127] During the second decision-making process, the animation guides the agent's output: action function library - "true"; scene element function library - "true"; duration information - "slow".

[0128] During the second self-check, the animation guides the agent's input: Description text: "The sky darkens, a car equipped with a spotlight slowly drives towards the vase, the car is in the camera's focus." Response results: Action function library - "true"; Scene element function library - "true"; Duration information - "slow".

[0129] During the second self-check, the animation guides the agent's output: accurate.

[0130] 303. When the first function library includes an action function library, the computer device determines the first action function from among multiple action functions based on the first description text by using the decision agent corresponding to the action function library.

[0131] The system includes multiple function libraries, including an action function library. Each action function library contains multiple action functions used to control the actions of virtual objects in the animation. If the first function library includes an action function library—meaning that generating an animation adapted to the first descriptive text requires the use of this library—then the decision-making agent corresponding to that action function library selects the first action function from that library. This decision-making agent can be called the action decision-making agent, and the first function selected from the action function library can be called the first action function.

[0132] In this embodiment, a predefined action function library is provided. The action functions in the action function library can control the actions of virtual objects in the animation. Through the decision-making agent, a suitable action function can be selected from the action function library according to the description text to ensure that the content presented after controlling the action of the virtual object by calling the action function conforms to the content described in the description text. There is no need for manual consideration of which action function is appropriate, thus realizing intelligent design of the actions of virtual objects in the animation.

[0133] In one possible implementation, each action function corresponds to a specific action category, with one action function belonging to one action category, and one action category including at least one action function. Then step 303 includes: using the decision agent corresponding to the action function library, determining the target action category from among multiple action categories based on the first description text; and using the decision agent corresponding to the action function library, determining the first action function from among the action functions belonging to the target action category based on the first description text.

[0134] In other words, the action decision-making agent uses a hierarchical approach to select action functions. First, it selects the action category of the virtual object (e.g., the curvilinear motion of a car), and then selects the specific action function within that category (e.g., the action function for S-curve movement). This hierarchical approach is anthropomorphic, reflecting human thought processes. This anthropomorphic thinking pattern makes the action decision-making agent's decision-making process more accurate and persuasive.

[0135] In this embodiment, the action functions in the action function library correspond to their respective action categories. When selecting an action function, the decision-making agent first selects a suitable target action category from multiple action categories, and then selects a suitable action function from multiple action functions under the target action category. Therefore, the action function selection process is divided into two levels. On the one hand, it can reduce the number of action functions that the decision-making agent needs to screen and improve processing efficiency. On the other hand, it can enable the decision-making agent to imitate human thinking, first determine what type of action to take, and then determine which action to take, thereby improving the accuracy of the decision-making agent.

[0136] In some embodiments, the above-mentioned determination of the target action category among multiple action categories based on the first descriptive text can be achieved in the following manner: obtaining the mapping relationship between the preset descriptive text and the action category, and searching for the preset descriptive text whose similarity to the first descriptive text is greater than a similarity threshold from the mapping relationship; determining the action category corresponding to the found preset descriptive text as the target action category.

[0137] In some embodiments, step 303 includes: determining a first action function corresponding to at least one virtual object identifier based on a first descriptive text and at least one virtual object identifier by a decision agent corresponding to the action function library, wherein the virtual object identifier indicates a virtual object appearing in the animation, and the first action function corresponding to the virtual object identifier is used to control the action of the virtual object.

[0138] In some embodiments, the at least one virtual object identifier is a pre-selected virtual object identifier, which indicates a virtual object that the user expects to appear in the animation to be generated.

[0139] In some embodiments, a decision agent corresponding to an action function library determines a first action function corresponding to at least one virtual object identifier based on a first descriptive text and at least one virtual object identifier. The user selects at least one virtual object identifier, which indicates the virtual object the user expects to appear in the animation to be generated. For example, the user might select virtual object identifiers such as "person," "animal," or "vehicle." The computer device analyzes the first descriptive text using natural language processing technology to extract keywords and phrases related to the action. For example, if the first descriptive text is "A person is walking in a forest and suddenly encounters a heavy rain," then the extracted keywords might include "walking," "encountering," and "heavy rain." Based on the results of the text analysis, the computer device determines the target action category related to the first descriptive text. For example, the text mentioning "walking" might correspond to the "movement" category, and "encountering heavy rain" might correspond to the "environmental change" category. The computer device searches for action functions in the action function library that match the target action category. For example, under the "movement" category, there might be action functions such as "walking," "running," and "jumping," and under the "environmental change" category, there might be action functions such as "rain," "thunder," and "lightning." The computer device associates at least one virtual object identifier with a corresponding action function. For example, it associates the "person" identifier with the "walking" action function and the "forest" identifier with the "raining" action function. The computer device runs the generated animation generation code, calls the first action function, and generates an animation. This animation demonstrates the actions performed by the virtual object, which correspond to the content described in the first descriptive text.

[0140] In some embodiments, the virtual object includes dynamic virtual objects and static virtual objects. Dynamic virtual objects refer to movable virtual objects, such as people, cats, dogs, or cars appearing in animations. Static virtual objects refer to immovable virtual objects, such as trees, houses, or pools appearing in animations.

[0141] As an example, suppose we have animation software with a rich library of motion functions, containing various preset motion functions such as "walk," "run," "jump," "attack," and "defend." A user wants to use this software to generate an animation showing a knight fighting on horseback in battle. The user first selects virtual object identifiers in the software, such as "knight" and "warhorse." These identifiers indicate the virtual objects the user expects to appear in the animation. The user enters the initial descriptive text: "The knight charges at the enemy on horseback, brandishing his sword in attack." The software analyzes the text using natural language processing technology, extracting the keywords "horseback," "charge," "brandish," and "attack." Based on these keywords, the software determines the target motion categories as "movement" and "attack." The software searches the motion function library for motion functions that match the target motion categories. For example, it finds the "horseback" motion function under the "movement" category and the "sword-wielding" motion function under the "attack" category. The software associates the "knight" identifier with the "sword-wielding" motion function and the "warhorse" identifier with the "horseback" motion function. Based on the association between virtual object identifiers and action functions, the software determines that the first action function corresponding to "knight" is "swing sword" and the first action function corresponding to "warhorse" is "ride horse". The software runs the generated animation code, calling the "ride horse" and "swing sword" action functions to generate the animation. This animation shows a scene of a knight riding a horse towards an enemy and attacking with his sword, which matches the content described in the user's first description text. This demonstrates how the computer device, through the decision-making agent corresponding to the action function library, determines the first action function corresponding to at least one virtual object identifier based on the first description text and at least one virtual object identifier, thereby generating an animation that meets the user's requirements.

[0142] Thus, by using the decision agent corresponding to the action function library, based on the first descriptive text and at least one virtual object identifier, the decision agent determines the first action function corresponding to at least one virtual object identifier. The decision agent then uses natural language processing technology to deeply analyze the first descriptive text, extracting keywords and phrases related to the action, thereby accurately identifying the animation scene and action type desired by the user. Simultaneously, the pre-selection of virtual object identifiers ensures that the decision agent can quickly locate the action function matching the user's needs. By associating virtual object identifiers with corresponding action functions, the decision agent can automatically generate code to control the virtual object's actions, reducing the complexity and error rate of manually writing code. Furthermore, this method improves the flexibility and customizability of animation generation; users can select different virtual object identifiers and action functions according to their needs, thereby creating a rich variety of animation effects. In summary, by using the decision agent corresponding to the action function library, based on the first descriptive text and at least one virtual object identifier, to determine the first action function corresponding to at least one virtual object identifier, an efficient, accurate, flexible, and customizable technical means is provided for animation generation.

[0143] In this embodiment, while providing descriptive text to the decision-making agent, the decision-making agent is also provided with virtual objects that need to appear in the animation. The decision-making agent can decide which action function to select for each virtual object based on the descriptive text, thereby realizing the intelligent design of each virtual object's own action in the animation, which helps to ensure the richness of the generated animation.

[0144] In one possible implementation, step 303 includes: using the decision agent corresponding to the action function library, determining candidate action functions from multiple action functions based on the first description text; using the decision agent corresponding to the action function library, generating a detection result based on the first description text and the candidate action functions, the detection result indicating whether the candidate action function is accurate, and if the detection result indicates that the candidate action function is incorrect, the detection result also includes the reason for the error; if the detection result indicates that the candidate action function is accurate, determining the candidate action function as the first action function; if the detection result indicates that the candidate action function is incorrect, determining the next candidate action function using the decision agent based on the first description text and the reason for the error, until the detection result indicates that the currently obtained candidate action function is accurate, then determining the currently obtained candidate action function as the first action function.

[0145] In some embodiments, the action decision agent's processing includes a decision-making process and a self-checking process. During the decision-making process, the action decision agent determines candidate action functions for intelligent selection. During the self-checking process, the action decision agent checks its selected candidate action functions to determine if the selection is accurate. If accurate, the currently selected candidate action function is adopted as the final action function. If inaccurate, the action decision agent needs to re-enter the decision-making process, combine the first description text and the error reason to re-select candidate action functions, and re-enter the self-checking process to check the newly selected candidate action functions until a selected candidate action function successfully passes the self-check. The candidate action function that successfully passes the self-check is then adopted as the final action function.

[0146] In some embodiments, the action decision-making agent first performs natural language processing on the first descriptive text to extract keywords and phrases related to the action. Based on the extracted keywords, the agent searches for matching candidate action functions in an action function library. These candidate functions may be directly related to the keywords or indirectly related through semantic analysis. The agent tests the candidate action functions to determine whether they accurately reflect the action described in the first descriptive text. This may involve checking multiple aspects such as the function's functionality, parameter settings, and expected output. If the test result indicates that the candidate action function is inaccurate, the agent analyzes the reason for the error. This may include problems such as inappropriate function selection, incorrect parameter settings, or the function's functionality not matching the text description. If the test result indicates that the candidate action function is incorrect, the agent re-enters the decision-making process based on the first descriptive text and the reason for the error, selecting the next candidate action function. The agent re-tests the re-selected candidate action function to determine its accuracy. Once the test result indicates that the currently obtained candidate action function is accurate, the agent determines that function as the first action function. If the candidate action function is still inaccurate, the agent continues the iterative selection and testing process until an accurate candidate action function is found. Through the iterative process of decision-making and self-checking, the action decision agent can ensure that the final selected action function accurately reflects the action described in the first descriptive text, thereby improving the accuracy and quality of animation generation.

[0147] As an example, suppose we have an animation software that a user wants to use to generate an animation showing a character running through a forest and dodging obstacles. The user's initial description is: "The character is running fast through the forest, dodging trees and rocks." The decision agent analyzes the text, extracting the keywords "running," "dodging," "trees," and "rocks." The agent checks whether candidate action functions accurately reflect the text description. For example, it checks whether the "running fast" function can simulate the character's running motion in the forest, and whether the "avoiding obstacles" function can allow the character to avoid trees and rocks. If the detection results show that a candidate action function is inaccurate, the agent analyzes the reason for the error. For example, it might be because the "avoiding obstacles" function defaults to avoiding a specific type of obstacle, while trees and rocks in the forest may require specific avoidance methods. Based on the reason for the error, the agent reselects candidate action functions. For example, it selects an "advanced avoidance" function that allows for customized avoidance methods. The agent re-detects the reselected candidate action function to ensure that it accurately simulates the character's action of dodging trees and rocks in the forest. Once the detection results show that the candidate action function is accurate, the agent determines that function as the first action function. If the candidate action function is still inaccurate, the agent continues the iterative selection and detection process until an accurate candidate action function is found.

[0148] Thus, the decision-making agent corresponding to the action function library determines candidate action functions from multiple action functions based on the first descriptive text. It then uses a self-checking process to assess the accuracy of these candidate action functions. The agent employs natural language processing technology to deeply understand the first descriptive text, extracting key action information and intelligently selecting candidate action functions from the action function library. The self-checking process ensures that the selected action functions accurately reflect the text description. Through a feedback mechanism, the agent can identify and correct errors in the selection process, and then reselect a more suitable action function. This iterative decision-making and self-checking process not only improves the accuracy of action function selection but also enhances the system's adaptability and flexibility, making animation generation more efficient and precise.

[0149] In this embodiment, after the decision-making agent determines the function based on the description text, it also performs a self-check on the selected function to verify the accuracy of its decision. If the decision is incorrect, the agent re-determines the function based on the description text and the reason for the error, until the currently determined function passes the self-check. Therefore, by adding a self-check process, it is possible to effectively ensure that the decision-making agent selects a more accurate function, thereby ensuring the adaptability of the subsequently generated animation to the description text.

[0150] To facilitate understanding, the following examples illustrate the inputs and outputs of the action decision-making agent.

[0151] Input to the action decision-making agent: Number of virtual objects: 3; Dynamic virtual objects: "cat", "puppy", "car"; Static virtual objects: "vase", "Christmas tree", "stone sculpture", "sun"; Description text: "The sky darkens, and a car equipped with a spotlight slowly drives towards the vase, under the focus of the camera."

[0152] The output of the action decision agent is: {virtual object "cat"; action category "special motion"; action function "do nothing"; function variable "cat"}; {virtual object "puppy"; action category "special motion"; action function "do nothing"; function variable "puppy"}; {virtual object "car"; action category "linear motion"; action function "constant speed motion"; function variable "car"}.

[0153] It should be noted that the embodiments of this application are only illustrated by taking the first function library including the action function library as an example, so step 303 needs to be performed.

[0154] In another embodiment, if the first function library does not include the action function library, then step 303 does not need to be performed.

[0155] 304. When the first function library includes a scene element function library, the computer device determines the first element function from among multiple element functions based on the first description text by using the decision agent corresponding to the scene element function library.

[0156] The system includes multiple function libraries, including a scene element function library. This scene element function library contains multiple element functions used to control scene elements in the animation. If the first function library includes a scene element function library—meaning that generating an animation adapted to the first description text requires the use of this scene element function library—then the decision agent corresponding to that scene element function library selects the first element function from that library. This decision agent can be called the element decision agent, and the first function selected from the scene element function library can be called the first element function.

[0157] In this embodiment, a predefined scene element function library is provided. The element functions in this library can control the scenes in the animation. Through a decision-making agent, appropriate element functions can be selected from the library based on the description text. This ensures that the content presented after controlling the scene elements in the animation by calling the appropriate element function conforms to the description text, thus reasonably decorating the scenes in the animation. Since there is no need for manual selection of which element function is suitable, intelligent design of the scenes in the animation is achieved.

[0158] In some embodiments, the computer device first performs natural language processing on the first descriptive text to extract keywords and phrases related to scene elements. For example, if the text describes "a sunny beach," the extracted keywords might include "sunshine," "beach," and "blue sky." The computer device accesses a scene element function library, which contains multiple functions for controlling scene elements in the animation, such as "sunshine," "waves," and "blue sky and white clouds." Based on the results of the text analysis, the computer device searches the scene element function library for element functions that match the extracted keywords. For example, the function "sunshine" matches "sunshine," the function "waves" matches "beach," and the function "blue sky and white clouds" matches "blue sky." The decision agent corresponding to the scene element function library, i.e., the element decision agent, evaluates the matched element functions to determine which functions most accurately reflect the scene elements in the first descriptive text. Based on the evaluation results, the element decision agent selects the most suitable element function as the first element function. These functions will be used to control the scene elements in the animation to generate an animation environment adapted to the first descriptive text. The computer device runs the generated animation code, calls the first element function, and generates the scene elements in the animation. These elements together constitute the animation environment that matches the first descriptive text description.

[0159] As an example, suppose we have animation software, and a user wants to use it to generate an animation showcasing a tranquil forest morning scene. The user's initial descriptive text is: "Morning sunlight filters through the leaves onto a forest path, and birds sing in the branches." The software's element decision agent analyzes the text, extracting the keywords "morning," "sunlight," "leaves," "forest path," "birds," and "singing." It accesses a scene element function library, which contains multiple functions for controlling scene elements in the animation, such as "sunlight filtering through," "leaf swaying," "path paving," "birds flying," and "birdsong sound effects." Based on the text analysis results, the agent searches the scene element function library for element functions that match the extracted keywords. For example, the function "sunlight filtering through" matches "sunlight," the function "leaf swaying" matches "leaves," the function "path paving" matches "forest path," the function "birds flying" matches "birds," and the function "birdsong sound effects" matches "singing." The element decision agent evaluates the matched element functions to determine which functions most accurately reflect the scene elements in the initial description text. For example, the agent might choose the "sunlight filtering through" function to simulate the effect of sunlight filtering through leaves, the "leaf swaying" function to simulate the movement of leaves in a breeze, the "path paving" function to create the appearance of a forest path, the "birds flying" function to show the activity of birds on branches, and the "bird sound effect" function to add the sound of birds singing. Based on the evaluation results, the element decision agent selects the most suitable element function as the first element function. These functions will be used to control the scene elements in the animation to generate a forest morning scene that matches the initial description text. Running the generated animation generation code calls the first element function to generate the scene elements in the animation. These elements together constitute a tranquil and vivid forest morning scene that matches the user's initial description text. This demonstrates how the element decision agent, through text analysis, element function matching, and selection, determines the first element function, thus preparing for the generation of animated scene elements that meet the user's needs.

[0160] Thus, when the first function library includes a scene element function library, the computer device, through the decision-making agent corresponding to the scene element function library, determines the first element function from multiple element functions based on the first descriptive text, significantly improving the automation and accuracy of animation scene generation. Specifically, the decision-making agent uses natural language processing technology to deeply understand the first descriptive text, extracting keywords and phrases related to scene elements, and intelligently selecting element functions that match the text description from the scene element function library. This selection process not only improves the accuracy and consistency of scene elements but also enhances the immersiveness and realism of the animation scene. Through this method, the computer device can quickly generate animation scenes that meet user needs, reducing the workload of manual adjustments and optimizations, and improving the efficiency and quality of animation production.

[0161] In one possible implementation, step 304 includes: using a decision agent corresponding to the scene element function library, based on the first descriptive text and at least one virtual object identifier, determining a first element function corresponding to at least one virtual object identifier, wherein the virtual object identifier indicates a virtual object appearing in the animation, and the first element function corresponding to the virtual object identifier is used to add scene elements to the virtual object.

[0162] The at least one virtual object identifier is a pre-selected virtual object identifier, and the virtual object indicated by the virtual object identifier is the virtual object that the user expects to appear in the animation to be generated.

[0163] In some embodiments, a decision agent corresponding to the scene element function library determines a first element function corresponding to at least one virtual object identifier based on a first descriptive text and at least one virtual object identifier. The decision agent first performs natural language processing on the first descriptive text to extract keywords and phrases related to scene elements and virtual objects. For example, if the text describes "a character walking in the rain," the extracted keywords might include "character," "rain," and "walking." The decision agent identifies virtual object identifiers mentioned in the text. These identifiers are pre-selected and correspond to virtual objects that the user expects to appear in the animation. For example, virtual object identifiers might be "protagonist," "supporting character," etc. Based on the results of text analysis and virtual object identifiers, the decision agent searches the scene element function library for element functions that match the extracted keywords. For example, the "rain effect" function matches "rain," and the "walking animation" function matches "walking." The decision agent evaluates the matched element functions to determine which functions can most accurately add the corresponding scene elements to the virtual objects. For example, the agent might choose the "rain effect" function to simulate a rainy environment and the "walking animation" function to display the character's walking action. Based on the evaluation results, the decision-making agent selects the most suitable element function as the first element function. These functions are used to add scene elements to the virtual object to generate an animated scene that matches the initial description text. The computer device runs the generated animation generation code, calling the first element function to add scene elements to the virtual object. These elements together constitute an animated scene that matches the user's description, in which the behavior of the virtual object and environmental effects are accurately represented. The computer device can determine the first element function from the scene element function library based on the initial description text and the virtual object identifier, thereby adding appropriate scene elements to the virtual object. This paper explores how the computer device utilizes the scene element function library, the decision-making agent, and natural language processing technology to transform user needs into specific animated scene elements, creating a vivid and descriptive animated environment for the virtual object.

[0164] As an example, suppose there is animation software, and a user wants to use it to generate an animation showing a character running and dodging obstacles in a forest. The user's initial descriptive text is: "The character runs quickly in the forest, dodging trees and rocks." The user's pre-selected virtual object identifier is "protagonist." The decision-making agent analyzes the text, extracting the keywords "running," "dodging," "trees," and "rocks." The agent identifies the virtual object identifier "protagonist" mentioned in the text. Based on the text analysis results and the virtual object identifier, the agent searches for element functions in the scene element function library that match the extracted keywords. For example, the "run quickly" function matches "running," the "dodge obstacles" function matches "dodging," and the "trees and rocks" function matches "trees" and "rocks." The agent evaluates the matched element functions to determine which functions can most accurately add the corresponding scene elements to the "protagonist." For example, the agent might choose the "run quickly" function to simulate the protagonist's running action, the "dodge obstacles" function to show the protagonist dodging trees and rocks, and the "trees and rocks" function to create the obstacle environment in the forest. Based on the evaluation results, the agent selects the most suitable element function as the first element function. These functions will be used to add scene elements to the "protagonist" to generate an animated scene that matches the initial description text. The first element function is called to add scene elements to the "protagonist." These elements together constitute an animated scene of a character running and dodging obstacles in a forest, consistent with the user's initial description text.

[0165] In this embodiment, while providing descriptive text to the decision-making agent, the agent is also provided with virtual objects that need to appear in the animation. The decision-making agent can decide which action function to select for which virtual object based on the descriptive text, thereby realizing the intelligent addition of corresponding scene elements for specific virtual objects in the animation, which helps to ensure the richness of the generated animation.

[0166] In one possible implementation, step 304 includes: using a decision agent corresponding to the scene element function library, determining candidate element functions from multiple action functions based on the first description text; using the decision agent corresponding to the scene element function library, generating a detection result based on the first description text and the candidate element functions, the detection result indicating whether the candidate element functions are accurate, and if the detection result indicates that the candidate element functions are incorrect, the detection result also includes the reason for the error; if the detection result indicates that the candidate element functions are accurate, determining the candidate element functions as the first element functions; if the detection result indicates that the candidate element functions are incorrect, determining the next candidate element functions using the decision agent based on the first description text and the reason for the error, until the detection result indicates that the currently obtained candidate element functions are accurate, then determining the currently obtained candidate element functions as the first element functions.

[0167] The processing of the scene element agent includes a decision-making process and a self-checking process. During the decision-making process, the scene element agent determines candidate action functions for intelligent selection. During the self-checking process, the scene element agent checks its selected candidate function to determine its accuracy. If accurate, the currently selected candidate function is adopted as the final element function. If inaccurate, the scene element agent needs to re-enter the decision-making process, combining the initial description text and the reason for the error to reselect candidate functions, and then re-enter the self-checking process to check the newly selected candidate function until a successfully selected candidate function passes the self-check. The successfully self-checked candidate function is then adopted as the final element function.

[0168] In some embodiments, a decision agent corresponding to the scene element function library determines candidate element functions from multiple action functions based on the first descriptive text. This process involves the decision agent's intelligent selection and self-checking mechanisms. First, the decision agent performs in-depth analysis of the first descriptive text, extracting keywords and phrases related to scene elements. Then, in the scene element function library, the agent intelligently selects matching candidate element functions based on these keywords and phrases. These candidate element functions are functions used to control scene elements in the animation, adding various visual and auditory effects to the animation scene.

[0169] In some embodiments, the decision-making agent performs a self-check on the selected candidate element functions. The self-check process includes checking the accuracy of the candidate element functions to ensure they accurately reflect the scene elements in the first descriptive text. If the check shows the candidate element function is accurate, the agent identifies that function as the first element function for generating the animation scene. If the check shows an error in the candidate element function, the agent records the reason for the error and re-enters the decision-making process, combining the first descriptive text and the reason for the error to reselect a candidate element function.

[0170] In some embodiments, the processing of the scene element agent includes a decision-making process and a self-checking process. During the decision-making process, the agent determines candidate action functions for intelligent selection. During the self-checking process, the agent checks its selected candidate element functions to determine their accuracy. If accurate, the currently selected candidate element function is adopted as the final element function. If inaccurate, the scene element agent needs to re-enter the decision-making process, combining the first description text and the reason for the error to reselect candidate element functions, and then re-enter the self-checking process to check the selected candidate element functions again, until the selected candidate element functions successfully pass the self-check. The candidate element functions that successfully pass the self-check are adopted as the final element functions. Through this iterative process of decision-making and self-checking, the scene element agent can ensure that the final selected element function is accurate, thereby providing a high-quality guarantee for the generation of animation scenes. This process demonstrates how the agent ensures the generation quality and accuracy of animation scenes through intelligent selection and self-checking mechanisms.

[0171] As an example, suppose there is an animation software program. A user wants to use the software to generate an animation showing a character running and dodging obstacles in a forest. The user's initial descriptive text is: "The character runs quickly in the forest, dodging trees and rocks." The virtual object identifier pre-selected by the user is "protagonist." The decision-making agent first analyzes the text, extracting the keywords "running," "dodging," "trees," and "rocks." Then, it searches for element functions matching these keywords in a scene element function library. For example, the function "running quickly" matches "running," the function "dodging obstacles" matches "dodging," and the function "trees and rocks" matches "trees" and "rocks." The agent selects the "running quickly" function as a candidate element function and tests its accuracy. The test results show that this function accurately simulates the character's running motion, so it is determined as the first element function. The agent then selects the "dodging obstacles" function as a candidate element function and tests its accuracy. The test results show that this function accurately simulates the character's behavior of dodging obstacles, so it is determined as the first element function. The agent selects the "trees and rocks" function as a candidate element function and tests its accuracy. The test results show that this function can accurately create an obstacle environment in the forest, so it is determined as the first element function. The first element function is called to add scene elements to the "protagonist". These elements together constitute an animated scene of the character running and dodging obstacles in the forest, which matches the user's initial description text.

[0172] In this embodiment, after the decision-making agent determines the function based on the description text, it also performs a self-check on the selected function to verify the accuracy of its decision. If the decision is incorrect, the agent re-determines the function based on the description text and the reason for the error, until the currently determined function passes the self-check. Therefore, by adding a self-check process, it is possible to effectively ensure that the decision-making agent selects a more accurate function, thereby ensuring the adaptability of the subsequently generated animation to the description text.

[0173] Furthermore, since the action function library and the scene element function library have significant semantic differences in functionality, corresponding decision agents were set up for each of them. These two decision agents can analyze their respective function libraries and execute their respective decision tasks independently, which helps to improve the accuracy of decision-making.

[0174] To facilitate understanding, the following examples illustrate the inputs and outputs of the element decision agent's decision-making process.

[0175] Input to the element decision agent: Number of virtual objects: 3; Dynamic virtual objects: "cat", "puppy", "car"; Static virtual objects: "vase", "Christmas tree", "stone sculpture", "sun"; Description text: "The sky darkens, and a car equipped with a spotlight slowly drives towards the vase, under the focus of the camera."

[0176] Output of the element decision agent: {element function "lighting function"; function variables "car", "strong light intensity"}; {element function "switch camera"; function variable "car"}.

[0177] It should be noted that the embodiments of this application are only illustrated by taking the first function library including the scene element function library as an example, so step 304 needs to be performed.

[0178] In another embodiment, if the first function library does not include the scene element function library, then step 304 does not need to be performed.

[0179] It should be noted that steps 303 and 304 above enable the decision-making agent corresponding to the first function library to determine the first function in the first function library based on the first description text.

[0180] In one possible implementation, a decision agent determines candidate functions from a first function library based on a first descriptive text; the decision agent generates a detection result based on the first descriptive text and the candidate functions, the detection result indicating whether the candidate functions are accurate, and if the detection result indicates that the candidate functions are incorrect, the detection result also includes the reason for the error; if the detection result indicates that the candidate functions are accurate, the candidate functions are determined as the first function; if the detection result indicates that the candidate functions are incorrect, the decision agent determines the next candidate function based on the first descriptive text and the reason for the error, until the detection result indicates that the currently obtained candidate functions are accurate, then the currently obtained candidate functions are determined as the first function.

[0181] In some embodiments, the decision agent first performs natural language processing on the first descriptive text to extract keywords and phrases related to the animation scene. For example, if the text describes "a character walking in the rain," the extracted keywords might include "character," "rain," and "walking." Based on the text analysis results, the decision agent searches a first function library for functions that match the extracted keywords. These functions might be "character walking," "rain effect," etc. The decision agent generates a detection result based on the first descriptive text and candidate functions. The detection result indicates whether the candidate function is accurate. If the candidate function is inaccurate, the detection result also includes the reason for the error. For example, if the candidate function is "character running" instead of "character walking," the reason for the error might be "action type mismatch." If the detection result indicates that the candidate function is accurate, the decision agent identifies the candidate function as the first function. If the detection result indicates that the candidate function is incorrect, the decision agent determines the next candidate function based on the first descriptive text and the reason for the error. This continues until the detection result indicates that the currently obtained candidate function is accurate. At this point, the agent identifies the currently obtained candidate function as the first function.

[0182] 305. The computer device generates an intelligent agent through code generation. Based on a first function, it generates animation generation code corresponding to a first descriptive text. The animation generation code is used to call the first function to generate an animation, and the code generation intelligent agent is used to generate the code.

[0183] After determining at least one first function required for generating the animation, the computer device provides this first function to a code generation agent. The code generation agent then generates executable animation generation code based on this first function. This code generation process can adaptively adapt to changes in the number of virtual objects during animation production, and is beneficial for long-term dynamic modeling, significantly reducing manual workload and promoting the automation of the animation production process.

[0184] In one possible implementation, if duration information is also determined in step 302 above, then step 305 includes: generating an intelligent agent through code, generating animation generation code based on the first function and the duration information, the animation generation code being used to call the first function to generate an animation that conforms to the duration information.

[0185] In some embodiments, the code generation agent receives a first function determined by a decision agent and duration information provided by the user. For example, the first function might be "character walks," and the duration information might be "5 seconds." The code generation agent generates a corresponding function call statement based on the first function. For example, if the first function is "character walks," the generated function call statement might be `character.walk()`. The code generation agent adds code to control the animation duration based on the duration information. This might involve setting timers, animation loops, or frame rate control. For example, if the duration information is "5 seconds," the code generation agent might generate a timer so that the `character.walk()` function stops executing after 5 seconds. The code generation agent combines the function call statement and duration control code into complete animation generation code. This might include initialization code, animation logic code, and termination code. The code generation agent outputs the generated animation generation code for execution by a computer device. Executing this code will call the first function to generate an animation that conforms to the duration information. The code generation agent is capable of generating animation generation code based on the first function and duration information. This process demonstrates how a code-generating agent uses function calls, duration control, and code synthesis techniques to transform user requirements into specific animated scene elements.

[0186] As an example, suppose there is an animation software that a user wants to generate an animation showing a character running and dodging obstacles in a forest. The user's initial description is: "The character runs quickly in the forest, dodging trees and rocks." The user has pre-selected the virtual object identifier as "protagonist" and specified the animation duration as 10 seconds. The decision agent analyzes the text, extracts the keywords "running," "dodging," "trees," and "rocks," and searches for matching functions in the scene element function library. The first determined function might be "character running" and "dodging obstacles." The code generation agent receives the first function "character running" and "dodging obstacles" determined by the decision agent, along with the user-specified duration of 10 seconds. The code generation agent generates corresponding function call statements based on the first function. For example, the generated function call statements might be `character.run()` and `character.avoidObstacles()`. The code generation agent adds code based on the duration information to control the animation duration. For example, the code generation agent might generate a timer so that the `character.run()` and `character.avoidObstacles()` functions stop executing after 10 seconds. The code-generating agent synthesizes function call statements and duration control code into complete animation generation code. This may include initialization code, animation logic code, and termination code. The code-generating agent outputs the generated animation generation code for execution by a computer device. Executing this code will call the first functions "Character Running" and "Obstacle Avoidance," generating a 10-second animation showing a character running and avoiding obstacles in a forest.

[0187] In this way, the code-generating agent uses natural language processing technology to deeply understand the first function and duration information, automatically generating the corresponding animation generation code. This code generation process not only improves the automation level of animation production but also reduces the time and effort spent on manually writing code. Furthermore, the code-generating agent can automatically adjust the animation duration based on the duration information, ensuring that the generated animation meets the user's needs and expectations. Through this method, computer equipment can quickly generate animation scenes that meet user requirements, reducing the workload of manual adjustments and optimizations, and improving the efficiency and quality of animation production.

[0188] In one possible implementation, step 305 includes: generating an intelligent agent by code, generating animation generation code based on a first function and at least one virtual object identifier, the virtual object identifier indicating a virtual object appearing in the animation, and the animation generation code being used to call the first function to generate an animation including the virtual object.

[0189] The at least one virtual object identifier is a pre-selected virtual object identifier, and the virtual object indicated by the virtual object identifier is the virtual object that the user expects to appear in the animation to be generated.

[0190] In some embodiments, a code-generating agent generates animation generation code based on a first function and at least one virtual object identifier. The code-generating agent first identifies virtual object identifiers pre-selected by the user. These identifiers indicate the virtual objects the user expects to appear in the animation to be generated. For example, virtual object identifiers might be "protagonist," "supporting character," etc. The code-generating agent determines the animation content to be generated based on the first function and the virtual object identifiers. For example, if the first function is "character running" and the virtual object identifier is "protagonist," then the code-generating agent will generate an animation of the protagonist running. The code-generating agent generates corresponding animation generation code based on the first function and the virtual object identifiers. A computer device runs the generated animation generation code, calls the first function, and generates an animation including the virtual objects. This animation will demonstrate the behavior and actions of the virtual objects within the animation. The code-generating agent is capable of generating animation generation code based on a first function and at least one virtual object identifier. This process demonstrates how a code-generating agent utilizes virtual object identifiers, a first function, and code generation techniques to translate user requirements into specific animation scene elements.

[0191] In this embodiment, the code generation agent generates executable animation generation code based on the intelligently selected function and the determined virtual object, so as to ensure that the generated animation contains the specified virtual object and improve the operability of animation generation.

[0192] In one possible implementation, step 305 includes: generating an intelligent agent by code, generating animation generation code based on a first function, at least one virtual object identifier, and duration information, wherein the virtual object identifier indicates a virtual object appearing in the animation, and the animation generation code is used to call the first function to generate an animation that includes the virtual object and whose duration conforms to the duration information.

[0193] In some embodiments, the code generation agent first parses the user-provided input, including a first function, a virtual object identifier, and duration information. For example, the first function might be "character dancing," the virtual object identifier is "protagonist," and the duration information is "30 seconds." The code generation agent generates the corresponding function call statement based on the first function and the virtual object identifier. For example, the generated function call statement might be `main_character.dance()`. The code generation agent adds code to control the animation duration based on the duration information. This might involve setting timers, animation loops, or frame rate control. For example, if the duration information is "30 seconds," the code generation agent might generate a timer so that the `main_character.dance()` function stops executing after 30 seconds. The code generation agent combines the function call statement and duration control code into complete animation generation code. This might include initialization code, animation logic code, and termination code. The code generation agent outputs the generated animation generation code for execution by a computer device. Executing this code will call the first function "character dancing," generating a 30-second animation showing the protagonist dancing. The code-generating agent can generate animation generation code based on a first function, at least one virtual object identifier, and duration information. This process demonstrates how the code-generating agent uses function calls, duration control, and code synthesis techniques to transform user requirements into specific animation scene elements.

[0194] 306. The computer device runs the animation generation code corresponding to the first description text to obtain the first animation, the content of which matches the content described in the first description text.

[0195] In one possible implementation, an animation creation application runs on the computer device. This application predefines the aforementioned function libraries, and the animation generation code is executable code within the animation creation application. The computer device runs this animation generation code within the animation creation application to perform 3D animation rendering, thereby obtaining the first animation. For example, the animation creation application could be Blender.

[0196] Figure 4 is a schematic diagram of an animation generation method provided in an embodiment of this application. As shown in Figure 4, the animation guidance agent, action decision agent, and element decision agent respond based on the descriptive text, obtaining the response results of each agent. Furthermore, the animation guidance agent, action decision agent, and element decision agent can perform self-checks on their respective response results until a response result successfully passes the self-check is obtained. As shown in Figure 4, the response result of the animation guidance agent indicates that the action function library and scene element function library need to be called, and the duration information is "slow". The response result of the action decision agent indicates the action functions to be called to control the cat, dog, and car, and the response result of the element decision agent indicates the element functions to be called to add scene elements to the animation. Then, based on the response results of the above three agents, the duration information and the required action and element functions are converted into executable animation generation code, and a 3D rendered animation can be obtained by running this animation generation code.

[0197] The method provided in this application, when generating animation, only requires providing descriptive text to describe the animation. Animation guidance agent intelligently predicts which function libraries to use to generate the animation based on the descriptive text. Then, the decision agent corresponding to each function library intelligently predicts which functions from that library to use to generate the animation based on the descriptive text. Finally, a code generation agent intelligently generates executable animation generation code based on the selected functions. This animation generation code calls the intelligently selected functions to generate the animation. Therefore, by running the animation generation code, the corresponding animation can be obtained. Since the functions used to generate the animation are intelligently selected based on the descriptive text, the animation generated by calling these functions matches the content described in the descriptive text. This achieves automatic generation of matching animations based on the descriptive text, making the animation production process more intelligent and automated, and improving the efficiency of animation generation.

[0198] Based on the above embodiments, the descriptive text can also be split into multiple fragment descriptive texts. Animation is generated using these fragment descriptive texts as units, following a shot-by-shot approach. For details, please refer to the embodiment shown in Figure 5 below. Figure 5 is a flowchart of another animation generation method provided by this application embodiment. This application embodiment is executed by a computer device. Referring to Figure 5, the method includes:

[0199] 501. The computer device acquires the first description text, which is used to describe the animation.

[0200] Step 501 is the same as step 201 above, and will not be repeated here.

[0201] 502. The computer device splits the first description text into multiple first segment description texts, which are used to describe different animation segments in the same animation.

[0202] The first descriptive text includes the plurality of first segment descriptive texts, meaning that the plurality of first segment descriptive texts constitute the first descriptive text. It can be understood that the first descriptive text describes the entire animation, while the first segment descriptive texts describe different animation segments within the animation.

[0203] In one possible implementation, the computer device splits the descriptive text into multiple first-segment descriptive texts based on punctuation marks in the descriptive text. For example, the descriptive text is split according to periods, with the descriptive text between two periods forming a single first-segment descriptive text.

[0204] 503. The computer device guides the intelligent agent through animation to determine the first function library corresponding to the first fragment of descriptive text from multiple function libraries based on the first fragment of descriptive text.

[0205] 504. The computer device, through the decision-making agent corresponding to the first function library, determines the first function corresponding to the first fragment of descriptive text in the first function library based on the first fragment of descriptive text.

[0206] The process of determining the first function in steps 503-504 is the same as the process of determining the first function in steps 302-304, and will not be repeated here.

[0207] It should be noted that for each first fragment of descriptive text, the computer device executes the above steps 503-504 respectively, thereby obtaining the first function corresponding to each first fragment of descriptive text.

[0208] 505. The computer device generates an intelligent agent through code, and generates animation generation code corresponding to the first descriptive text based on the first function corresponding to multiple first fragment descriptive texts.

[0209] In this embodiment, the computer device provides multiple first functions corresponding to first fragment description texts to a code generation agent, which directly generates animation generation code for the entire first description text. The animation generation code for each first description text contains animation generation code for each first fragment description text, and this animation generation code is used to generate the animation segment described by that first fragment description text.

[0210] Apart from that, the process of step 505 is the same as that of step 305 above, and will not be repeated here.

[0211] 506. The computer device runs the animation generation code corresponding to the first description text to obtain the first animation, which includes the animation segments described by each first segment description text.

[0212] The process of step 506 is the same as that of step 306 above, and will not be repeated here.

[0213] In some embodiments, the computer device first obtains a first descriptive text provided by the user, which describes the content of the entire animation. For example, the first descriptive text might be: "A character runs in a forest, then encounters a friendly animal, and finally goes on an adventure together." The computer device breaks down the first descriptive text into multiple first segment descriptive texts, each segment corresponding to a different animation segment in the animation. For example, the split segment descriptive texts might be: "A character runs in a forest," "Encounters a friendly animal," and "Goes on an adventure together." The computer device, through an animation guidance agent, determines the corresponding first function library from multiple function libraries based on each first segment descriptive text. For example, for the segment descriptive text "A character runs in a forest," the animation guidance agent might determine the "character action function library" and the "scene function library." The computer device, through a decision agent corresponding to the first function library, determines the corresponding first function from the first function library based on each first segment descriptive text. For example, for the segment descriptive text "A character runs in a forest," the decision agent might determine the "character running function" and the "forest scene function." The computer device, through a code generation agent, generates animation generation code corresponding to the first descriptive text based on the first functions corresponding to the multiple first segment descriptive texts. The computer device can generate corresponding animation generation code based on the initial descriptive text. This process demonstrates how the computer device utilizes text segmentation, animation guidance agents, decision-making agents, and code generation agents to translate user needs into specific animated scene elements.

[0214] As an example, suppose we have animation software, and a user wants to use this software to generate an animation showing a character running in a forest, encountering a friendly animal, and then going on an adventure together. The user's initial descriptive text is: "A character runs in a forest, then encounters a friendly animal, and finally goes on an adventure together." The computer device first retrieves the initial descriptive text provided by the user: "A character runs in a forest, then encounters a friendly animal, and finally goes on an adventure together." The computer device breaks down the initial description text into multiple initial fragment description texts, each corresponding to a different animation segment. These fragment description texts might be: "The character is running in the forest," "Encountering a friendly animal," or "Exploring together." The computer device, through an animation-guided agent, determines the corresponding first function library from multiple function libraries based on each initial fragment description text. For example, for the fragment description text "The character is running in the forest," the animation-guided agent might determine the "character action function library" and the "scene function library." The computer device, through a decision agent corresponding to the first function library, determines the corresponding first function from the first function library based on each initial fragment description text. For example, for the fragment description text "The character is running in the forest," the decision agent might determine the "character running function" and the "forest scene function." The computer device, through a code-generating agent, generates animation generation code corresponding to the initial description text based on the first functions corresponding to the multiple initial fragment description texts. The computer device runs the generated animation generation code, calls the corresponding functions, generates and plays the animation. This animation will show a scene of the character running in the forest, encountering a friendly animal, and exploring together, matching the user's initial description text.

[0215] It should be noted that this embodiment only illustrates the example of directly generating animation generation code corresponding to a first descriptive text based on a first function corresponding to multiple first fragment descriptive texts. In another embodiment, in step 505 above, the computer device can also generate animation generation code corresponding to each first fragment descriptive text based on a first function corresponding to each first fragment descriptive text. Then, in step 506 above, the computer device runs the animation generation code corresponding to multiple first fragment descriptive texts to obtain a first animation. Alternatively, in step 506 above, the computer device runs the animation generation code corresponding to each first fragment descriptive text separately to obtain an animation segment corresponding to each first fragment descriptive text, and splices the obtained multiple animation segments to obtain the first animation.

[0216] Figure 6 is a schematic diagram of another animation generation method provided in an embodiment of this application. As shown in Figure 6, the description text is split into fragment description text 1 to fragment description text 6, each fragment description text describing an animation segment. Taking fragment description text 1 as an example, the animation guidance agent determines the duration information and the action decision agent and element decision agent to be called through decision-making and self-checking. Then, the action decision agent and element decision agent determine the action function and element function to be called through decision-making and self-checking. Taking fragment description text 2 as an example, the animation guidance agent determines the duration information and the action decision agent to be called through decision-making and self-checking. Then, the action decision agent determines the action function to be called through decision-making and self-checking. Finally, the duration information, action function, element function, and virtual object identifier determined for fragment description texts 1 to 6 are provided to the code generation agent, which converts them into executable animation generation code. Then, running the animation generation code in the animation production software yields an animation including multiple animation segments. Figure 7 is a schematic diagram of the result of an animation generation method provided in an embodiment of this application. As shown in Figure 7, according to the processing flow in Figure 6, the animation including animation segments 1 to 6 can be obtained.

[0217] The method provided in this application, when generating animation, only requires providing descriptive text to describe the animation. Animation guidance agent intelligently predicts which function libraries to use to generate the animation based on the descriptive text. Then, the decision agent corresponding to each function library intelligently predicts which functions from that library to use to generate the animation based on the descriptive text. Finally, a code generation agent intelligently generates executable animation generation code based on the selected functions. This animation generation code calls the intelligently selected functions to generate the animation. Therefore, by running the animation generation code, the corresponding animation can be obtained. Since the functions used to generate the animation are intelligently selected based on the descriptive text, the animation generated by calling these functions matches the content described in the descriptive text. This achieves automatic generation of matching animations based on the descriptive text, making the animation production process more intelligent and automated, and improving the efficiency of animation generation.

[0218] Furthermore, the descriptive text is broken down into multiple fragment descriptive texts. For each fragment, an animation generation code is determined, allowing each fragment to generate a corresponding animation clip. This fragment-based approach, similar to storyboarding, helps each agent fully understand the descriptive text, improving the accuracy of animation generation.

[0219] Based on the above embodiments, the first descriptive text can be extended by a text generation agent to obtain a second descriptive text associated with it, and a second animation described by the second descriptive text can be generated, which is equivalent to expanding the content of the generated animation. For details, please refer to the embodiment in Figure 8 below.

[0220] Figure 8 is a flowchart of another animation generation method provided in an embodiment of this application. This embodiment of the application is executed by a computer device. Referring to Figure 8, the method includes:

[0221] 801. The computer device acquires the first description text, which is used to describe the animation.

[0222] 802. The computer device guides the intelligent agent through animation, and determines the first function library from multiple function libraries based on the first descriptive text.

[0223] 803. The computer device determines the first function in the first function library based on the first description text by using the decision agent corresponding to the first function library.

[0224] 804. The computer device generates an intelligent agent through code, and generates animation generation code corresponding to the first descriptive text based on the first function.

[0225] The processes of steps 801-804 above are the same as those of steps 301-305 above, and will not be repeated here.

[0226] 805. A computer device generates a text-based intelligent agent that generates a second descriptive text based on a first descriptive text. The content described in the second descriptive text is associated with the content described in the first descriptive text. The text-based intelligent agent is used to generate the descriptive text.

[0227] In some embodiments, the computer device first obtains a first descriptive text provided by the user, which describes the content of the entire animation. For example, the first descriptive text might be: "A character runs in a forest, then encounters a friendly animal, and finally they go on an adventure together." The computer device, through an animation guidance agent, determines the corresponding first function library from multiple function libraries based on the first descriptive text. For example, the animation guidance agent might determine a "character action function library," a "scene function library," and an "animal interaction function library." The computer device, through a decision agent corresponding to the first function library, determines the corresponding first function from the first function library based on the first descriptive text. For example, the decision agent might determine a "character running function," a "forest scene function," and an "animal interaction function." The computer device, through a code generation agent, generates animation generation code corresponding to the first descriptive text based on multiple first functions. The computer device, through a text generation agent, generates a second descriptive text based on the first descriptive text. The content described in the second descriptive text is related to the content described in the first descriptive text and may include further descriptions of the animation scene, the character's inner thoughts, or other relevant plot points. For example, the second descriptive text might be: "In a dense forest, sunlight filters through the leaves, illuminating a running character. He feels a sense of ease and freedom, as if the whole world is making way for him. Suddenly, he hears a soft call, and a friendly deer appears before him. He stops, meets the deer's gaze, and a warm feeling washes over him. They begin their adventure together, traversing the forest and exploring the unknown world." The computer device runs the generated animation code, calls the corresponding functions, and generates and plays the animation. This animation will show the scene of the character running in the forest, encountering a friendly animal, and exploring together, matching the user's first descriptive text. At the same time, the second descriptive text provides a richer background and emotional depth to the animation, enhancing the user's viewing experience.

[0228] A computer device provides the first descriptive text to a text-generating agent, which then generates a second descriptive text associated with the first descriptive text. The text-generating agent is used to generate another descriptive text associated with any given descriptive text. In some embodiments, the text-generating agent belongs to a large language model.

[0229] In this embodiment of the application, a narrative extension function is provided. A text-generating agent extends a second descriptive text associated with a first descriptive text. The content described by the second descriptive text is consistent with and related to the content described by the first descriptive text, thereby enabling the creation of a continuous animation based on the extended descriptive text.

[0230] In one possible implementation, the computer device acquires conditional information indicating the conditions that the generated descriptive text must satisfy. The computer device then generates a text-generating agent that, based on the first descriptive text and the conditional information, generates a second descriptive text whose content is associated with the content described in the first descriptive text, and which satisfies the conditional information.

[0231] Figure 9 is a schematic diagram of another animation generation method provided in an embodiment of this application. The input to the text generation agent is a first descriptive text and conditional information, such as the conditional information "I hope this story has a happy ending." This conditional information indicates the condition satisfied by the second descriptive text. The output of the text generation agent is the second descriptive text that satisfies the conditional information. As shown in Figure 9, the second descriptive text includes fragment description text 1 to fragment description text 4, each fragment description text being used to describe an animation segment.

[0232] 806. The computer device guides the intelligent agent through animation, and determines the second function library from multiple function libraries based on the second description text.

[0233] 807. The computer device determines the second function in the second function library based on the second description text by using the decision agent corresponding to the second function library.

[0234] 808. The computer device generates an intelligent agent through code, and generates animation generation code corresponding to the second descriptive text based on the second function.

[0235] The process of steps 806-806 above is the same as the process of steps 302-305 above, and will not be repeated here.

[0236] 809. The computer device runs the animation generation code corresponding to the first description text to obtain the first animation, the content of which matches the content described in the first description text.

[0237] 810. The computer device runs the animation generation code corresponding to the second description text to obtain the second animation. The content of the second animation matches the content described in the second description text.

[0238] In some embodiments, the computer device guides an agent through animation to determine a second function library from multiple function libraries based on a second descriptive text. Then, the computer device, through a decision agent corresponding to the second function library, determines a second function from the second function library based on the second descriptive text. Next, the computer device generates animation generation code corresponding to the second descriptive text based on the second function using a code-generating agent. Finally, the computer device runs the animation generation code corresponding to the first descriptive text to obtain a first animation, the content of which matches the content described in the first descriptive text; the computer device also runs the animation generation code corresponding to the second descriptive text to obtain a second animation, the content of which matches the content described in the second descriptive text.

[0239] As an example, in the field of education, computer devices can be used to create educational animations to help students better understand complex concepts. For instance, suppose there is educational software that needs to generate corresponding animations based on different teaching content. The computer device, through an animation-guided agent, determines a second function library from multiple function libraries based on a second descriptive text. The descriptive text might be a description of the cell division process in biology. The animation-guided agent analyzes this descriptive text, understands the key concepts and processes, and then searches for animation function libraries related to cell division in multiple function libraries. The computer device, through a decision agent corresponding to the second function library, determines a second function from the second function library based on the second descriptive text. The decision agent further analyzes the descriptive text to determine which specific animation functions are needed to represent the different stages of cell division, such as interphase, prophase, metaphase, anaphase, and telophase. Based on these requirements, the decision agent determines the specific animation functions in the selected function library. The computer device, through a code-generating agent, generates animation generation code corresponding to the second descriptive text based on the second function. The code-generating agent generates corresponding code based on the selected animation functions, and this code can call these functions to create the animation of cell division. The computer device runs the animation generation code corresponding to the first descriptive text, producing a first animation whose content matches the description in the first descriptive text. The first descriptive text might be a description of force in physics. The computer device runs this code, generating an animation demonstrating the concept of force, whose content matches the descriptive text. The computer device runs the animation generation code corresponding to a second descriptive text, producing a second animation whose content matches the description in the second descriptive text. The computer device runs the previously generated cell division animation code, generating an animation demonstrating the cell division process, whose content matches the descriptive text. Educational software can automatically generate corresponding animations based on different teaching content, helping students understand complex scientific concepts more intuitively.

[0240] Since the content described in the second descriptive text is related to the content described in the first descriptive text, and the first animation includes the content described in the first descriptive text, while the second animation includes the content described in the second descriptive text, the content of the second animation is related to the content of the first animation. It can be understood that the second animation is a sequel to the first animation; therefore, the first and second animations can be merged into a single complete animation.

[0241] The solution provided in this application, in addition to automatically generating a first animation using a first descriptive text provided by the user, can also generate a second animation automatically using a text-based intelligent agent that extends the first descriptive text to obtain a related second descriptive text. Therefore, the user only needs to provide a descriptive text, which can then intelligently expand upon the content and generate corresponding animations, further improving the intelligence of animation generation and making the generated animations richer and more interesting, bringing additional surprises to the user.

[0242] The methods in the above embodiments can be implemented using a large language model. Since a large language model has difficulty directly generating executable animation generation code from descriptive text, it can be guided to perform functions at different stages through procedural modeling, enabling the large language model to make decisions at the cognitive level. Therefore, based on the large language model, the aforementioned animation guidance agent, action decision agent, element decision agent, code generation agent, and text generation agent are introduced. Furthermore, multiple function libraries for generating animations are predefined, and the large language model's own learning capabilities are utilized to enable each agent to learn its assigned function.

[0243] For ease of explanation, the following definitions will be explained first.

[0244] L (Function Library): Describes the goals and functions of the predefined function library.

[0245] F(function): Describes the functionality of functions in the predefined function library.

[0246] V(function variable): Interprets the function variable.

[0247] I (Learning Example): Provides input and output to demonstrate how to use the provided input to reason about the output. The learning example is presented in the form of contextual learning.

[0248] The following section introduces the learning process of the animation guidance agent, decision-making agent (action decision-making agent and element decision-making agent), code generation agent, and text generation agent.

[0249] Animation guidance agent: The above-mentioned animation guidance agent belongs to the large language model. Before determining the first function library from multiple function libraries based on the first descriptive text, the method further includes: inputting first learning information into the animation guidance agent. The first learning information is used to instruct the animation guidance agent to learn and predict the function library corresponding to any descriptive text.

[0250] The first learning information includes introductory information for multiple function libraries and a first learning example. The first learning example includes sample description text and the corresponding sample function library, which belongs to multiple function libraries. The animation-guided agent learns the functionality of each function library through its introductory information. Furthermore, by understanding the functionality of the function libraries, it learns how the first learning example determines the corresponding sample function library based on the sample description text. In addition, the first learning information also includes first prompts, which guide the animation-guided agent on how to learn.

[0251] To facilitate understanding, the following examples illustrate the first learning information.

[0252] (1) First prompt: Assuming you are an animation director expert in 3D animation, you can determine which function libraries to use to create 3D animations based on the description text. I will provide you with two function libraries: the action function library <Laction> and the scene element function library <Ldecoration>, and will introduce them in detail. You should understand their functions to help you with your analysis. In addition, you need to provide the animation duration information. You can choose from four different types of duration information, including "fast", "moderate", "slow", and "emphasis". Additionally, learning examples are provided. <idirector>Let me teach you how to reply.

[0253] (2) Introduction to the action function library <Laction> and the scene element function library <Ldecoration>.

[0254] (3) Input in the first learning example: sample description text.

[0255] (4) Output of the first learning example: sample function library and sample duration information.

[0256] In this embodiment, since the animation guidance agent is a large language model, and the large language model itself has learning capabilities, it is only necessary to provide the animation guidance agent with the first learning information. The animation guidance agent can then use the first learning information to learn how to predict the function library without adjusting the parameters of the animation guidance agent. This is convenient and simple to operate and helps to improve learning efficiency.

[0257] Decision agent: The above-mentioned decision agent belongs to the large language model. Before determining the first function in the first function library based on the first descriptive text, the method further includes: inputting second learning information into the decision agent corresponding to the first function library. The second learning information is used to instruct the decision agent corresponding to the first function library to learn and predict the function corresponding to any descriptive text.

[0258] The second learning information includes introductory information for each function in the first function library and a second learning example. The second learning example includes sample description text and the corresponding sample function, which belongs to the first function library. The decision agent learns the function's functionality through the introductory information in the function library. Furthermore, by understanding the function's functionality, it learns how the second learning example determines the corresponding sample function based on the sample description text. In addition, the second learning information also includes second prompting information to guide the decision agent in its learning process.

[0259] In one possible implementation, the decision-making agent includes an action-making agent, and the action functions in the action function library correspond to their respective action categories. Then, C(action category) is defined: introducing the function of each action category in the action function library. For ease of understanding, the second learning information of the action-making agent is illustrated below with an example.

[0260] (1) Second prompt: Assuming you are a motion decision expert for 3D animation, you can determine which motion functions from the motion function library to use to create 3D animations based on the description text. I will provide a detailed introduction to the motion function library <Laction>, as well as the motion categories <Caction> contained in the motion function library <Laction>, and the motion functions <Faction> contained in each motion category <Caction>. You should learn how to use motion functions with variable interpretations. <vaction>Additionally, we have provided learning examples. <iaction>Let me teach you how to reply.

[0261] (2) Information about the action function library <Laction> and its included action categories <Caction> and action functions <Faction>.

[0262] (3) Input in the second learning example: sample description text.

[0263] (4) Output of the second learning example: sample action category and sample action function.

[0264] In one possible implementation, the decision agent includes an element decision agent. For ease of understanding, the second learning information of the element decision agent is illustrated below.

[0265] (1) Second prompt: Assuming you are very skilled at decorating 3D animations, you can determine which scene elements from the scene element function library to use to decorate your 3D animations based on the description text. I will introduce the scene element function library <Ldecoration> in detail, as well as the element function <Fdecoration> contained within <Ldecoration>. You should learn how to use the element function <Vdecoration> with variable interpretation. Additionally, a learning example <Idecoration> is provided to teach you how to respond.

[0266] (2) Introduction to the scene element function library <Ldecoration> and its contained element function <Fdecoration>.

[0267] (3) Input in the second learning example: sample description text.

[0268] (4) Output of the second learning example: Sample element function.

[0269] In this embodiment, since the decision agent is a large language model, and the large language model itself has learning capabilities, it is only necessary to provide the decision agent with second learning information. The decision agent can then use the second learning information to learn how to predict functions without adjusting the parameters of the decision agent. This is convenient and simple to operate and helps to improve learning efficiency.

[0270] Code generation agent: The above code generation agent belongs to the large language model. Before generating the animation generation code corresponding to the first descriptive text based on the first function, the method also includes: inputting third learning information into the code generation agent. The third learning information is used to instruct the code generation agent to learn and predict the animation generation code corresponding to any descriptive text.

[0271] The third learning information includes a third learning example, which comprises a sample function corresponding to the sample description text and sample animation generation code. The code generation agent learns how to generate sample animation generation code based on the sample function in the third learning example. In addition, this third learning information also includes third prompts, which guide the code generation agent on how to learn.

[0272] In this embodiment, since the code generation agent belongs to a large language model, and the large language model itself has learning capabilities, it is only necessary to provide the code generation agent with third learning information. The code generation agent can then use the third learning information to learn how to generate executable animation generation code. There is no need to adjust the parameters of the code generation agent. The operation is convenient and simple, which is conducive to improving learning efficiency.

[0273] Text generation agent: The above text generation agent belongs to the large language model. Before generating the second descriptive text based on the first descriptive text, the method further includes: inputting fourth learning information into the text generation agent. The fourth learning information is used to instruct the text generation agent to learn and predict other descriptive texts associated with any descriptive text.

[0274] The fourth learning information includes introductory information for multiple function libraries and a fourth learning example. The fourth learning example includes a first sample description text and a second sample description text, with the content described in the first sample description text being related to the content described in the second sample description text. The text generation agent learns the functions of each function library through the introductory information. Furthermore, based on its understanding of the function library functions, it learns how the second sample description text is generated from the first sample description text in the fourth learning example. In addition, this fourth learning information also includes a fourth prompt, which guides the text generation agent on how to learn.

[0275] To facilitate understanding, the fourth learning information will be illustrated with examples below.

[0276] (1) Fourth prompt: Imagine you are very skilled at creating 3D animations, and you can build upon existing stories to create new ones. Your new story must be implemented using the action library <Laction> and the scene element library <Ldecoration>, and we will provide a detailed introduction to these two libraries. Additionally, we provide learning examples. <icontinuation>Let me teach you how to reply.

[0277] (2) Introduction to the action function library <Laction> and the scene element function library <Ldecoration>.

[0278] (3) Input in the fourth learning example: first sample description text and conditional information.

[0279] (4) Output of the fourth learning example: the description text of the second sample.

[0280] In this embodiment, since the text generation agent is a large language model, and the large language model itself has learning capabilities, it is only necessary to provide the text generation agent with fourth learning information. The text generation agent can then learn how to generate descriptive text using the fourth learning information. There is no need to adjust the parameters of the text generation agent. The operation is convenient and simple, which is conducive to improving learning efficiency.

[0281] Figure 10 is a flowchart of another animation generation method provided in an embodiment of this application. As shown in Figure 10, the method includes the following steps.

[0282] 1001. Obtain the first description text, which is used to describe the animation.

[0283] 1002. Split the first description text into multiple first fragment description texts.

[0284] 1003. Using the first descriptive text as a unit, execute the code generation process to obtain the animation generation code corresponding to the first descriptive text.

[0285] The code generation process includes the following steps:

[0286] (1) Guide the agent through animation, and determine the first function library and duration information from multiple function libraries based on the first fragment of the descriptive text.

[0287] (2) If the first function library includes an action function library, the decision agent corresponding to the action function library determines the first action function corresponding to each virtual object identifier among multiple action functions based on the first fragment description text and at least one virtual object identifier.

[0288] (3) If the first function library includes the scene element function library, the decision agent corresponding to the scene element function library determines the first element function from among multiple element functions based on the first fragment of the descriptive text.

[0289] (4) Generate an intelligent agent through code, and generate animation generation code corresponding to the first description text based on at least one virtual object identifier and the first action function, first element function and duration information corresponding to multiple first fragment description texts.

[0290] 1004. Run the animation generation code corresponding to the first description text to obtain the first animation. The first animation includes the animation segments described by each first segment description text.

[0291] 1005. Generate an intelligent agent through text, and generate a second descriptive text based on the first descriptive text. The content described by the second descriptive text is related to the content described by the first descriptive text. The second descriptive text includes multiple second fragment descriptive texts.

[0292] 1006. Using the second descriptive text as a unit, execute the code generation process to obtain the animation generation code corresponding to the second descriptive text.

[0293] The code generation process includes the following steps:

[0294] (1) Guide the agent through animation, and determine the second function library and duration information from multiple function libraries based on the second segment description text.

[0295] (2) When the second function library includes an action function library, the decision agent corresponding to the action function library determines the second action function corresponding to each virtual object identifier among multiple action functions based on the second fragment description text and at least one virtual object identifier.

[0296] (3) When the second function library includes the scene element function library, the decision agent corresponding to the scene element function library determines the second element function from among multiple element functions based on the second fragment description text.

[0297] (4) Generate an intelligent agent through code, and generate animation generation code corresponding to the second description text based on at least one virtual object identifier and the second action function, second element function and duration information corresponding to multiple second fragment description texts.

[0298] 1007. Run the animation generation code corresponding to the second description text to obtain the second animation, which includes the animation segments described by each second segment description text.

[0299] 1008. Combine the first and second animations to obtain the target animation.

[0300] The animation guidance agent, decision-making agent (action decision-making agent and element decision-making agent), code generation agent, and text generation agent involved in the embodiments of this application can be regarded as a whole and collectively referred to as the animation generation agent (Animate3D-Agent). This is a novel LLM-Agent (Large Language Model-Agent) framework used to utilize LLM in the exploration of 3D animation production.

[0301] The method provided in this application offers a novel perspective on animation creation, effectively addressing challenges in 3D animation production, including the choreography of multiple virtual objects and the configuration of various scene elements within a given narrative context. Furthermore, by processing the descriptive text as multiple fragments, it ensures that the generated animation segments adhere to a predetermined storyline timeline, thus maintaining narrative coherence and scene consistency throughout the animation. Moreover, Animate3D-Agent also enables the expansion of animation content, maintaining contextual consistency with the original content, thereby potentially enhancing the expressive capabilities of 3D animation.

[0302] Figure 11 is a schematic diagram of an animation generation device provided in an embodiment of this application. Referring to Figure 11, the device includes:

[0303] The text acquisition module 1101 is configured to acquire the first descriptive text, which is used to describe the animation.

[0304] The function library determination module 1102 is configured to determine a first function library from multiple function libraries based on a first description text, guided by an animation intelligent agent. The functions in the function library are used to generate animations. Each function library has its own decision intelligent agent. The animation guiding intelligent agent is used to determine the function library, and the decision intelligent agent is used to determine the functions.

[0305] The function determination module 1103 is configured to determine the first function in the first function library based on the first description text by the decision agent corresponding to the first function library.

[0306] The code generation module 1104 is configured to generate animation generation code corresponding to the first descriptive text based on the first function by a code generation agent. The animation generation code is used to call the first function to generate animation, and the code generation agent is used to generate code.

[0307] The animation generation module 1105 is configured to run the animation generation code corresponding to the first description text to obtain the first animation. The content of the first animation matches the content described in the first description text.

[0308] The animation generation apparatus provided in this application only requires a descriptive text to describe the animation when animation generation is needed. An animation guidance agent intelligently predicts which function libraries to use to generate the animation based on the descriptive text. Then, a decision agent corresponding to the function library intelligently predicts which functions from that library to use to generate the animation based on the descriptive text. Finally, a code generation agent intelligently generates executable animation generation code based on the selected functions. This animation generation code calls the intelligently selected functions to generate the animation. Therefore, by running the animation generation code, the corresponding animation can be obtained. Since the functions used to generate the animation are intelligently selected based on the descriptive text, the animation generated by calling these functions matches the content described in the descriptive text. This achieves automatic generation of matching animations based on the descriptive text, making the animation production process more intelligent and automated, and improving the efficiency of animation generation.

[0309] In some embodiments, referring to Figure 12, the multiple function libraries include an action function library, which includes multiple action functions used to control the actions of virtual objects in the animation; the function determination module 1103 is configured as follows:

[0310] When the first function library includes an action function library, the first action function is determined from multiple action functions based on the first description text by the decision agent corresponding to the action function library.

[0311] In some embodiments, referring to Figure 12, each action function corresponds to a specific action category. An action function belongs to an action category, and an action category includes at least one action function. The function determination module 1103 is configured as follows:

[0312] Based on the first description text, the target action category is determined from multiple action categories by using the decision agent corresponding to the action function library;

[0313] Based on the first descriptive text, the decision agent corresponding to the action function library determines the first action function from among the action functions belonging to the target action category.

[0314] In some embodiments, referring to Figure 12, the function determination module 1103 is configured as follows:

[0315] Based on the first descriptive text and at least one virtual object identifier, the decision agent corresponding to the action function library determines the first action function corresponding to at least one virtual object identifier. The virtual object identifier indicates the virtual object that appears in the animation, and the first action function corresponding to the virtual object identifier is used to control the action of the virtual object.

[0316] In some embodiments, referring to Figure 12, the multiple function libraries include a scene element function library, which includes multiple element functions used to control scene elements in the animation; the function determination module 1103 is configured as follows:

[0317] When the first function library includes a scene element function library, the decision agent corresponding to the scene element function library determines the first element function from among multiple element functions based on the first description text.

[0318] In some embodiments, referring to Figure 12, the function determination module 1103 is configured as follows:

[0319] Based on the first descriptive text and at least one virtual object identifier, the decision agent corresponding to the scene element function library determines the first element function corresponding to at least one virtual object identifier. The virtual object identifier indicates the virtual object that appears in the animation, and the first element function corresponding to the virtual object identifier is used to add scene elements to the virtual object.

[0320] In some embodiments, referring to FIG12, the function library determination module 1102 is configured to determine the first function library and duration information based on the first descriptive text by guiding the intelligent agent through animation;

[0321] The code generation module 1104 is configured to generate an intelligent agent through code, based on the first function and duration information, to generate animation generation code. The animation generation code is used to call the first function to generate an animation that conforms to the duration information.

[0322] In some embodiments, referring to Figure 12, the code generation module 1104 is configured as follows:

[0323] An intelligent agent is generated by code. Based on a first function and at least one virtual object identifier, animation generation code is generated. The virtual object identifier indicates the virtual object that appears in the animation. The animation generation code is used to call the first function to generate an animation that includes the virtual object.

[0324] In some embodiments, referring to FIG12, the apparatus further includes:

[0325] The text splitting module 1106 is configured to split the first description text into multiple first segment description texts, which are used to describe different animation segments in the same animation; wherein, the animation guidance agent is used to determine the first function library corresponding to each first segment description text, and the decision agent corresponding to the first function library is used to determine the first function corresponding to each first segment description text.

[0326] The code generation module 1104 is configured to generate animation generation code corresponding to the first descriptive text based on the first function corresponding to multiple first fragment descriptive texts.

[0327] In some embodiments, referring to Figure 12, the function library determination module 1102 is configured as follows:

[0328] The agent is guided by animation, and candidate function libraries are determined from multiple function libraries based on the first descriptive text.

[0329] The intelligent agent is guided by animation to generate detection results based on the first descriptive text and the candidate function library. The detection results indicate whether the candidate function library is accurate. If the detection results indicate that the candidate function library is wrong, the detection results also include the reason for the error.

[0330] If the test results indicate that the candidate function library is accurate, the candidate function library will be determined as the first function library;

[0331] If the detection result indicates that the candidate function library is incorrect, the agent is guided by animation to determine the next candidate function library based on the first descriptive text and the reason for the error, until the detection result indicates that the currently obtained candidate function library is accurate, then the currently obtained candidate function library is determined as the first function library.

[0332] In some embodiments, referring to Figure 12, the function determination module 1103 is configured as follows:

[0333] The decision-making agent determines candidate functions from the first function library based on the first descriptive text.

[0334] The decision-making agent generates detection results based on the first descriptive text and candidate functions. The detection results indicate whether the candidate functions are accurate. If the detection results indicate that the candidate functions are incorrect, the detection results also include the reasons for the errors.

[0335] If the detection results indicate that the candidate function is accurate, the candidate function will be determined as the first function.

[0336] If the detection result indicates that the candidate function is incorrect, the decision agent determines the next candidate function based on the first descriptive text and the reason for the error, until the detection result indicates that the currently obtained candidate function is accurate, and then the currently obtained candidate function is determined as the first function.

[0337] In some embodiments, referring to Figure 12, the animation guidance agent belongs to a large language model, and the device further includes:

[0338] The first learning module 1107 is configured to input first learning information into the animation guidance agent. The first learning information is used to instruct the animation guidance agent to learn and predict the function library corresponding to any descriptive text.

[0339] The first learning information includes introductory information for multiple function libraries and a first learning example. The first learning example includes sample description text and the sample function library corresponding to the sample description text. The sample function library belongs to multiple function libraries.

[0340] In some embodiments, referring to Figure 12, the decision agent belongs to a large language model, and the device further includes:

[0341] The second learning module 1108 is configured to input second learning information into the decision agent corresponding to the first function library. The second learning information is used to instruct the decision agent corresponding to the first function library to learn and predict the function corresponding to any descriptive text.

[0342] The second learning information includes the introduction information of each function in the first function library and the second learning example. The second learning example includes sample description text and the sample function corresponding to the sample description text. The sample function belongs to the first function library.

[0343] In some embodiments, referring to Figure 12, the code generation agent belongs to a large language model, and the device further includes:

[0344] The third learning module 1109 is configured to input third learning information into the code generation agent. The third learning information is used to instruct the code generation agent to learn and predict the animation generation code corresponding to any descriptive text.

[0345] The third learning information includes a third learning example, which includes the sample function corresponding to the sample description text and the sample animation generation code.

[0346] In some embodiments, referring to FIG12, the apparatus further includes:

[0347] The text generation module 1110 is configured to generate a second descriptive text based on the first descriptive text using a text generation agent. The content described by the second descriptive text is associated with the content described by the first descriptive text. The text generation agent is used to generate the descriptive text.

[0348] The function library determination module 1102 is also configured to guide the agent through animation to determine a second function library from multiple function libraries based on the second description text;

[0349] The function determination module 1103 is further configured to determine the second function in the second function library based on the second description text by the decision agent corresponding to the second function library;

[0350] The code generation module 1104 is also configured to generate animation generation code corresponding to the second descriptive text based on the second function through a code generation agent;

[0351] The animation generation module 1105 is also configured to run the animation generation code corresponding to the second description text to obtain the second animation.

[0352] In some embodiments, referring to Figure 12, the text-generating agent belongs to a large language model, and the device further includes:

[0353] The fourth learning module 1111 is configured to input fourth learning information into the text generation agent. The fourth learning information is used to instruct the text generation agent to learn and predict other descriptive texts associated with any descriptive text.

[0354] The fourth learning information includes introductory information for multiple function libraries and fourth learning examples. The fourth learning examples include first sample description text and second sample description text, and the content described by the first sample description text is related to the content described by the second sample description text.

[0355] It should be noted that the animation generation device provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the animation generation device and the animation generation method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0356] This application also provides a computer device, which includes a processor and a memory. The memory stores at least one computer program, which is loaded and executed by the processor to perform the operations performed in the animation generation method of the above embodiments.

[0357] In some embodiments, the computer device is provided as a terminal. Figure 13 shows a schematic diagram of the structure of a terminal 1300 provided in an exemplary embodiment of this application.

[0358] Terminal 1300 includes a processor 1301 and a memory 1302.

[0359] Processor 1301 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1301 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1301 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is configured to process data in the wake-up state; the coprocessor is a low-power processor configured to process data in the standby state. In some embodiments, processor 1301 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1301 may also include an AI (Artificial Intelligence) processor configured to handle computational operations related to machine learning.

[0360] Memory 1302 may include one or more computer-readable storage media, which may be non-transitory. Memory 1302 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in memory 1302 are used to store at least one computer program configured to be implemented by processor 1301 to implement the animation generation method provided in the method embodiments of this application.

[0361] In some embodiments, the terminal 1300 may also optionally include a peripheral device interface 1303 and at least one peripheral device. The processor 1301, memory 1302, and peripheral device interface 1303 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1303 via a bus, signal line, or circuit board. In some embodiments, the peripheral device includes at least one of a radio frequency circuit 1304, a display screen 1305, a camera assembly 1306, an audio circuit 1307, and a power supply 1308.

[0362] Peripheral device interface 1303 can be configured to connect at least one I / O (Input / Output) related peripheral device to processor 1301 and memory 1302. In some embodiments, processor 1301, memory 1302 and peripheral device interface 1303 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1301, memory 1302 and peripheral device interface 1303 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0363] The radio frequency (RF) circuit 1304 is configured to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1304 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1304 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. In some embodiments, the RF circuit 1304 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1304 can communicate with other devices via at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1304 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.

[0364] Display screen 1305 is configured to display a user interface (UI). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 1305 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1301 for processing. In this case, display screen 1305 can also be configured to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, display screen 1305 may be a single screen, disposed on the front panel of terminal 1300; in other embodiments, display screen 1305 may be at least two screens, disposed on different surfaces of terminal 1300 or in a folded design; in still other embodiments, display screen 1305 may be a flexible display screen, disposed on a curved or folded surface of terminal 1300. Furthermore, display screen 1305 may also be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1305 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0365] The camera assembly 1306 is configured to capture images or videos. In some embodiments, the camera assembly 1306 includes a front-facing camera and a rear-facing camera. The front-facing camera is disposed on the front panel of the terminal 1300, and the rear-facing camera is disposed on the back of the terminal 1300. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1306 may also include a flash. The flash may be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cool light flash, which can be configured for light compensation at different color temperatures.

[0366] The audio circuit 1307 may include a microphone and a speaker. The microphone is configured to acquire sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1301 for processing, or input to the radio frequency circuit 1304 for voice communication. For stereo acquisition or noise reduction purposes, multiple microphones may be used, each located at a different part of the terminal 1300. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is configured to convert electrical signals from the processor 1301 or the radio frequency circuit 1304 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1307 may also include a headphone jack.

[0367] Power supply 1308 is configured to supply power to the various components in terminal 1300. Power supply 1308 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 1308 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be configured to support fast charging technology.

[0368] Those skilled in the art will understand that the structure shown in FIG13 does not constitute a limitation on the terminal 1300, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0369] In some embodiments, the computer device is provided as a server. Figure 14 is a schematic diagram of the structure of a server provided in an embodiment of this application. The server 1400 can vary considerably due to different configurations or performance, and may include one or more Central Processing Units (CPUs) 1401 and one or more memories 1402. The memory 1402 stores at least one computer program, which is loaded and executed by the processor 1401 to implement the methods provided in the above-described method embodiments. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components configured to implement device functions, which will not be elaborated here.

[0370] This application also provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to implement the operations performed by the animation generation method of the above embodiments.

[0371] This application also provides a computer program product, including a computer program loaded and executed by a processor to perform the operations performed by the animation generation method of the above embodiments.

[0372] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0373] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present application should be included within the protection scope of the present application.< / icontinuation> < / iaction> < / vaction> < / idirector>

Claims

1. An animation generation method, applied to a computer device, the method comprising: Obtain the first description text, which is used to describe the animation; Based on the first description text, a first function library is determined from multiple function libraries, and the functions in the function library are used to generate animations; Based on the first description text, determine the first function in the first function library; Based on the first function, animation generation code corresponding to the first descriptive text is generated, and the animation generation code is used to call the first function to generate an animation; Run the animation generation code corresponding to the first descriptive text to obtain the first animation, the content of which matches the content described in the first descriptive text.

2. The method according to claim 1, wherein, The step of determining the first function library from multiple function libraries based on the first description text includes: The agent is guided by animation to determine the first function library from the plurality of function libraries based on the first descriptive text. The step of generating animation generation code corresponding to the first descriptive text based on the first function includes: An intelligent agent is generated through code. Based on the first function, animation generation code corresponding to the first descriptive text is generated. Each function library has its own decision-making intelligent agent. The step of determining the first function in the first function library based on the first descriptive text includes: Based on the first description text, the decision agent corresponding to the first function library determines the first function in the first function library.

3. The method according to claim 1, wherein, The multiple function libraries include an action function library, which includes multiple action functions, and the action functions are used to control the actions of virtual objects in the animation; The step of determining the first function in the first function library based on the first description text includes: If the first function library includes the action function library, the first action function is determined from the plurality of action functions based on the first description text.

4. The method according to claim 3, wherein, The action functions each correspond to their respective action categories, one action function belongs to one action category, and one action category includes at least one action function; determining the first action function from the plurality of action functions based on the first description text includes: Based on the first description text, the target action category is determined from multiple action categories; Based on the first description text, the first action function is determined from among the action functions belonging to the target action category.

5. The method according to claim 4, wherein, The step of determining the target action category among multiple action categories based on the first description text includes: Obtain the mapping relationship between the preset description text and the action category, and find the preset description text whose similarity to the first description text is greater than the similarity threshold from the mapping relationship; The action category corresponding to the preset description text obtained from the search is determined as the target action category.

6. The method according to claim 4, wherein, The step of determining the first action function from among the plurality of action functions based on the first description text includes: Based on the first descriptive text and at least one virtual object identifier, the decision agent corresponding to the action function library determines the first action function corresponding to the at least one virtual object identifier. The virtual object identifier indicates a virtual object that appears in the animation, and the first action function corresponding to the virtual object identifier is used to control the action of the virtual object.

7. The method according to claim 1, wherein, The plurality of function libraries includes a scene element function library, which includes a plurality of element functions used to control scene elements in the animation; the step of determining a first function in the first function library based on the first description text includes: If the first function library includes the scene element function library, a first element function is determined from the plurality of element functions based on the first description text.

8. The method according to claim 7, wherein, The step of determining the first element function among the plurality of element functions based on the first description text includes: Based on the first description text and at least one virtual object identifier, the decision agent corresponding to the scene element function library determines the first element function corresponding to the at least one virtual object identifier. The virtual object identifier indicates a virtual object that appears in the animation, and the first element function corresponding to the virtual object identifier is used to add scene elements to the virtual object.

9. The method according to claim 1, wherein, Before generating the animation generation code corresponding to the first descriptive text based on the first function, the method further includes: The intelligent agent is guided by animation to determine duration information based on the first descriptive text, wherein the duration information indicates the duration of the animation to be generated; The step of generating animation generation code corresponding to the first descriptive text based on the first function includes: An intelligent agent is generated by generating code. Based on the first function and the duration information, the animation generation code is generated. The animation generation code is used to call the first function to generate an animation that conforms to the duration information.

10. The method according to claim 1, wherein, The step of generating animation generation code corresponding to the first descriptive text based on the first function includes: The animation generation code is generated by generating an intelligent agent through code, based on the first function and at least one virtual object identifier; The virtual object identifier indicates a virtual object that appears in the animation, and the animation generation code is used to call the first function to generate an animation that includes the virtual object.

11. The method according to claim 1, wherein, The method further includes: The first descriptive text is split into multiple first fragment descriptive texts, which are used to describe different animation segments in the same animation. The animation guidance agent is used to determine the first function library corresponding to each first segment description text, and the decision agent corresponding to the first function library is used to determine the first function corresponding to each first segment description text. The step of generating animation generation code corresponding to the first descriptive text based on the first function includes: An intelligent agent is generated by generating code. Based on the first function corresponding to the plurality of first fragment description texts, animation generation code corresponding to the first description text is generated. The animation generation code includes animation generation code corresponding to each of the first fragment description texts.

12. The method according to claim 1, wherein, The step of determining the first function library from multiple function libraries based on the first description text includes: The intelligent agent is guided by animation to determine candidate function libraries from the plurality of function libraries based on the first descriptive text. The agent is guided by animation to generate a detection result based on the first descriptive text and the candidate function library. The detection result indicates whether the candidate function library is accurate. If the detection result indicates that the candidate function library is incorrect, the detection result also includes the reason for the error. Based on the detection result, the first function library is determined.

13. The method according to claim 12, wherein, The step of determining the first function library based on the detection results includes: If the detection result indicates that the candidate function library is accurate, the candidate function library is determined as the first function library; If the detection result indicates that the candidate function library is incorrect, the agent is guided by the animation to determine the next candidate function library based on the first descriptive text and the reason for the error, until the detection result indicates that the currently obtained candidate function library is accurate, then the currently obtained candidate function library is determined as the first function library.

14. The method according to claim 1, wherein, The step of determining the first function in the first function library based on the first description text includes: Based on the first descriptive text, a decision-making agent determines candidate functions from the first function library. The decision agent generates a detection result based on the first descriptive text and the candidate function. The detection result indicates whether the candidate function is accurate. If the detection result indicates that the candidate function is incorrect, the detection result also includes the reason for the error. Based on the detection result, the first function is determined.

15. The method according to claim 14, wherein, Determining the first function based on the detection result includes: If the detection result indicates that the candidate function is accurate, the candidate function is determined as the first function; If the detection result indicates that the candidate function is incorrect, the decision agent determines the next candidate function based on the first description text and the reason for the error, until the detection result indicates that the currently obtained candidate function is accurate, then the currently obtained candidate function is determined as the first function.

16. The method according to any one of claims 1-15, wherein, The first function library is determined by an animation guidance agent, which belongs to a large language model. Before determining the first function library from multiple function libraries based on the first descriptive text, the method further includes: The first learning information is input into the animation guidance agent, and the first learning information is used to instruct the animation guidance agent to learn and predict the function library corresponding to any descriptive text. The first learning information includes introductory information of the plurality of function libraries and a first learning example. The first learning example includes sample description text and a sample function library corresponding to the sample description text, and the sample function library belongs to the plurality of function libraries.

17. The method according to any one of claims 1-15, wherein, The first function is determined by a decision agent, which belongs to a large language model. Before determining the first function from the first function library based on the first descriptive text, the method further includes: The second learning information is input into the decision agent corresponding to the first function library, and the second learning information is used to instruct the decision agent corresponding to the first function library to learn and predict the function corresponding to any descriptive text. The second learning information includes introductory information of each function in the first function library and a second learning example. The second learning example includes sample description text and the sample function corresponding to the sample description text, and the sample function belongs to the first function library.

18. The method according to any one of claims 1-15, wherein, The animation generation code is determined by a code generation agent, which belongs to a large language model. Before generating the animation generation code corresponding to the first descriptive text based on the first function, the method further includes: The third learning information is input into the code generation agent, and the third learning information is used to instruct the code generation agent to learn and predict the animation generation code corresponding to any descriptive text; The third learning information includes a third learning example, which includes a sample function corresponding to the sample description text and sample animation generation code.

19. The method according to any one of claims 1-15, wherein, The method further includes: A text generation agent generates a second descriptive text based on the first descriptive text. The content described in the second descriptive text is associated with the content described in the first descriptive text. The text generation agent is used to generate the descriptive text. The agent is guided by animation to determine a second function library from the plurality of function libraries based on the second descriptive text. Based on the second description text, the decision agent corresponding to the second function library determines the second function in the second function library; The intelligent agent is generated by code, and the animation generation code corresponding to the second descriptive text is generated based on the second function; Run the animation generation code corresponding to the second description text to obtain the second animation.

20. The method according to claim 19, wherein, The text generation agent belongs to a large language model. Before generating the second descriptive text based on the first descriptive text using the text generation agent, the method further includes: The fourth learning information is input into the text generation agent, and the fourth learning information is used to instruct the text generation agent to learn and predict other descriptive texts associated with any descriptive text. The fourth learning information includes introductory information of the plurality of function libraries and a fourth learning example. The fourth learning example includes a first sample description text and a second sample description text, wherein the content described by the first sample description text is associated with the content described by the second sample description text.

21. An animation generation apparatus, applied to a computer device, the apparatus comprising: The text acquisition module is configured to acquire a first descriptive text, which is used to describe the animation. The function library determination module is configured to determine a first function library from multiple function libraries based on the first description text, wherein the functions in the function library are used to generate animation; The function determination module is configured to determine a first function from the first function library based on the first description text; The code generation module is configured to generate animation generation code corresponding to the first descriptive text based on the first function, and the animation generation code is used to call the first function to generate an animation. The animation generation module is configured to run the animation generation code corresponding to the first descriptive text to obtain a first animation, the content of which matches the content described in the first descriptive text.

22. A computer device comprising a processor and a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to perform the operations performed by the animation generation method as claimed in any one of claims 1 to 20.

23. A computer-readable storage medium storing at least one computer program, the at least one computer program being loaded and executed by a processor to perform the operations performed by the animation generation method as claimed in any one of claims 1 to 20.

24. A computer program product comprising a computer program loaded and executed by a processor to perform the operations performed by the animation generation method as described in any one of claims 1 to 20.

Citation Information

Patent Citations

  • Method and device applied to classroom activity animation generation

    CN111210494A

  • Animation generation method and device, electronic equipment, medium and computer program product

    CN114299198A

  • Expression animation generation method and device, equipment, storage medium and program product

    CN117557696A

  • Video generation method and system based on large language model

    CN117676195A

  • Terminal emulator generation method for image forming device, terminal application generation method for image forming device, program for making computer execute these methods, and graphic function library for image forming device

    JP2003280844A