Animation generation method, computing device, electronic device and storage medium
Through the multi-agent animation generation system, multiple proxy models are used to decompose the animation generation process, achieving efficient and personalized animation generation, and solving the problems of low animation generation efficiency and poor effect in the existing technology.
Patent Information
- Application Number
- CN202510638484.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-09-12
AI Technical Summary
The animation generation efficiency in the prior art is low and the generation effect is poor. It is difficult for users to participate in the animation generation process, resulting in the animation content being unable to meet the personalized needs of users.
A multi-agent animation generation system is used. The script settings are input into the first agent model to generate a script outline. The script outline is then input into multiple second agent models to generate character interpretation information. Finally, the script outline and character interpretation information are input into the third agent model to generate the target animation.
The agent model with specialized division of labor significantly shortens the animation generation cycle and production costs, supports real-time interaction between users during the generation process, and meets users' personalized needs for animation content.
Smart Images

Figure CN120635261A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an animation generation method, a computing device, an electronic device, and a storage medium. Background Art
[0002] The current digital entertainment industry is developing rapidly. The animation production and generation in related technologies usually requires the collaboration of multiple professional teams such as screenwriters, character designers, animators and post-production. Each link requires a high degree of professional skills and resource consumption, which makes the animation creation and generation limited by complex and time-consuming workflows, that is, the animation generation cycle is long and the cost is high; at the same time, the animation generation method based on artificial intelligence models in related technologies is essentially a "black box" animation generation method, that is, the user can only simply input the generation animation requirements before the animation generation process begins. The generated animation has problems such as uncontrollable content and large quality fluctuations. Users are usually passive content consumers, and it is difficult for users to actually participate in the various processes of animation generation. As a result, the generated animation content is difficult to meet the user's sense of participation, interactivity and personalized animation generation needs, resulting in low efficiency and generation effect of animation generation in related technologies.
[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0004] The embodiments of the present application provide an animation generation method, a computing device, an electronic device, and a storage medium to at least solve the technical problems of low animation generation efficiency and generation effect in related technologies.
[0005] According to one aspect of an embodiment of the present application, an animation generation method is provided, the method comprising: inputting a script setting into a first proxy model, generating a script outline using the first proxy model, wherein the script outline includes information of multiple characters, the script setting is used to represent the story setting of the script, and the script outline is used to represent the story framework of the script; inputting the script outline into multiple second proxy models respectively, generating role interpretation information of multiple characters using the multiple second proxy models, wherein the role interpretation information is used to represent the interpretation parameters of the characters in different scenarios; inputting the script outline and the role interpretation information into a third proxy model, and generating a target animation using the third proxy model.
[0006] According to another aspect of an embodiment of the present application, an animation generation method is also provided, which includes: responding to an input instruction acting on an operation interface, displaying a script setting on the operation interface; responding to a processing instruction acting on the operation interface, displaying a target animation on the operation interface, wherein the target animation is generated by inputting the script outline and role interpretation information into a third agent model, and using the third agent model; the role interpretation information is generated by inputting the script outline into multiple second agent models respectively, and using multiple second agent models to generate role interpretation information for multiple characters; the script outline is generated by inputting the script setting into a first agent model, and using the first agent model; the script outline contains information about multiple characters; the script setting is used to represent the story plot setting of the script; the script outline is used to represent the story plot framework of the script; and the role interpretation information is used to represent the interpretation parameters of the character in different scenarios.
[0007] According to another aspect of an embodiment of the present application, an animation generation method is also provided, which includes: obtaining script settings by calling a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter includes the script settings; inputting the script settings into a first agent model, and generating a script outline using the first agent model, wherein the script outline includes information of multiple characters, the script settings are used to represent the story settings of the script, and the script outline is used to represent the story framework of the script; inputting the script outline into multiple second agent models respectively, and generating role interpretation information of multiple characters using multiple second agent models, wherein the role interpretation information is used to represent the interpretation parameters of the characters in different scenarios; inputting the script outline and the role interpretation information into a third agent model, and generating a target animation using the third agent model; outputting the target animation by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the target animation.
[0008] According to another aspect of the embodiments of the present application, a computing device is further provided, including: a memory storing an executable program; and a processor for running the program, wherein the method of each embodiment of the present application is executed when the program is running.
[0009] According to another aspect of the embodiments of the present application, an electronic device is also provided, including: a memory storing an executable program; a processor connected to the memory via a bus, and configured to run the program, wherein the method of each embodiment of the present application is executed when the program is running.
[0010] According to another aspect of an embodiment of the present application, a computer-readable storage medium is also provided, which includes a stored executable program, wherein when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the methods in various embodiments of the present application.
[0011] According to another aspect of the embodiments of the present application, a computer program product is further provided, including a computer program, which implements the methods in various embodiments of the present application when executed by a processor.
[0012] According to another aspect of an embodiment of the present application, a computer program product is also provided, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method in each embodiment of the present application is implemented.
[0013] According to another aspect of the embodiments of the present application, a computer program is further provided, which implements the methods in various embodiments of the present application when executed by a processor.
[0014] In this embodiment of the present application, the script settings are first input into a first proxy model, which then generates a script outline. The script settings describe the story's plot settings, provide the story's plot framework, and include information about multiple characters. Next, the script outlines are input into multiple second proxy models, which then generate character interpretation information for each character. This character interpretation information displays the character's interpretation parameters in different scenarios. Finally, the script outlines and character interpretation information are input into a third proxy model, which then generates the target animation. It is easy to notice that the above process decomposes the animation generation process into multiple specialized models, namely the first agent model, multiple second agent models and the third agent model. Each model focuses on handling a specific type of creative task and realizes professional division of labor. Specifically, the first agent model can quickly generate a script outline according to the script settings, quickly capture the core elements of the plot, and construct a story framework that conforms to the narrative logic, providing a clear direction and structure for subsequent animation production. Multiple second agent models can generate detailed interpretation parameters for their respective characters, and realize automation and personalization through intelligent agents. The third agent model can automatically plan the scene layout and lens application based on the script outline and character interpretation information, thereby significantly shortening the animation generation cycle and production cost; at the same time, the use of multiple models with specialized division of labor can support users to provide real-time interactive guidance to each model in each link of animation generation. Each process of animation generation is user-controllable, which can meet the user's personalized needs for animation content, thereby solving the technical problems of low animation generation efficiency and generation effect in related technologies.
[0015] It is easy to notice that the above general description and the following detailed description are merely for the purpose of exemplifying and explaining the present application, and do not constitute a limitation of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0017] Figure 1 is a schematic diagram of an application scenario of an animation generation method according to an embodiment of the present application;
[0018] Figure 2 is a flowchart of an animation generation method according to an embodiment of the present application;
[0019] Figure 3 is a schematic diagram of an optional animation generation process according to an embodiment of the present application;
[0020] Figure 4 is a schematic diagram of playing games and chatting with a user during an optional animation generation process according to an embodiment of the present application;
[0021] Figure 5 is a flowchart of an optional animation generation method according to an embodiment of the present application;
[0022] Figure 6 is a flowchart of another optional animation generation method according to an embodiment of the present application;
[0023] Figure 7 is a structural block diagram of a computing device according to an embodiment of the present application;
[0024] Figure 8 This is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the present invention, the following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0027] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following interpretations:
[0028] The Multi-Agent Animation Generation (MAAG) system can be a technical framework for automatic animation generation based on multi-agent collaborative creation.
[0029] Retrieval-augmented generation (RAG) can be a hybrid technology architecture that combines external knowledge base retrieval and generation models.
[0030] Vector Similarity can refer to a mathematical indicator that calculates the degree of semantic similarity of content in a high-dimensional vector space.
[0031] Real-Time Rendering can be a technology for generating real-time three-dimensional graphics.
[0032] A virtual character is an interactive digital person generated by a computer and driven by an algorithm.
[0033] Intention Alignment can be a technical mechanism to ensure that the behavior of artificial intelligence systems meets the real needs of users.
[0034] The Multi-Agent Orchestrator (MAO) framework can provide intelligent request allocation and conversation state management functions.
[0035] Information transfer (Carryover) can realize the data flow mechanism of context coherence of multi-agent system.
[0036] An agent can refer to an artificial intelligence entity with autonomous decision-making capabilities, which can perform specialized tasks according to specific roles and communicate and collaborate with other agents.
[0037] Text to Speech (TTS) is a technology that converts text information into speech and can be applied to voice assistants, navigation systems, audiobooks and other fields.
[0038] Level of detail (LOD) can be used to adjust the model complexity based on the distance between the object and the camera. When the object is far away from the camera, a low-precision model is used to reduce the amount of calculation; when the object is close to the camera, a high-precision model is switched to improve visual quality.
[0039] Lightmap Bake can pre-calculate the lighting effects of static light sources and store them as textures, reducing the lighting calculation overhead at runtime.
[0040] According to an embodiment of the present application, an animation generation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0041] The above-mentioned animation generation method provided in the embodiment of the present application can be applied to Figure 1 The application scenarios shown are not limited to this. Figure 1 In the illustrated application scenario, the server 10 can be connected to one or more client devices 20 via a local area network, a wide area network, the Internet, or other types of data networks. The client devices 20 herein may include, but are not limited to, smartphones, tablet computers, laptop computers, PDAs, personal computers, smart home devices, and in-vehicle devices. The client devices 20 can interact with users via a graphical user interface to implement the methods provided in the embodiments of the present application.
[0042] In an embodiment of the present application, a system consisting of a client device and a server can perform the following steps: the client device can interact with the server. The server can input script settings into a first proxy model and use the first proxy model to generate a script outline; input the script outlines into multiple second proxy models and use the multiple second proxy models to generate character interpretation information for multiple characters; and input the script outlines and character interpretation information into a third proxy model and use the third proxy model to generate a target animation.
[0043] It should be noted that, with the rapid development of high-performance computing units, in other application scenarios, the above method provided in the embodiment of the present application can also be applied to the model all-in-one machine. In an optional embodiment, a plurality of models are built into the model all-in-one machine, and the user can choose to adjust with a model as needed to obtain the user's own model, so that the high-performance computing unit built into the model all-in-one machine can directly call the adjusted model to execute the above method provided in the embodiment of the present application. In another optional embodiment, a trained model is built into the model all-in-one machine, so that the high-performance computing unit built into the model all-in-one machine can directly call the model to execute the above method provided in the embodiment of the present application.
[0044] Furthermore, when users need to train their own models, they can upload their own datasets through the client. This dataset is then sent to the server, which then adjusts the pre-trained model using the dataset to create the user's own model, which can then be deployed in production. To facilitate user model adjustment needs, the server provides a complete set of adjustment tools, development frameworks, and processes, supporting a variety of adjustment strategies, making the adjusted model more adaptable to different application fields and highly customized.
[0045] Under the above operating environment, this application provides Figure 2 The animation generation method shown. Figure 2 This is a flow chart of the animation generation method according to an embodiment of the present application. Figure 2 As shown, the method may include the following steps:
[0046] In step S202 , the script settings are input into the first agent model, and a script outline is generated using the first agent model.
[0047] Among them, the script outline contains information about multiple characters, the script setting is used to represent the story setting of the script, and the script outline is used to represent the story framework of the script.
[0048] The aforementioned script settings refer to the basic concepts and requirements for the animated story, entered by the user. These can include core story elements such as theme, background, character design, and desired style or atmosphere. The script settings serve as the starting point for the entire creative process, providing direction and constraints for subsequent script generation. For example, if a user wants to create an exploration animated short film set in a mysterious forest, the script settings might include the main character being a brave female explorer and the desire for the story to be filled with a spirit of adventure, wonder, and exploration. Once the user enters this basic information into the system, the script settings are formed. The specific content of the script settings can be determined by the user based on actual needs and is not limited here.
[0049] The aforementioned first agent model can refer to an AI model capable of generating the backbone of a story based on the script settings, i.e., a script outline. This model can also fulfill the screenwriter's task in actual animation production, i.e., a screenwriter agent model. The first agent model can be pre-trained using literature, scripts, and story data. This training process enables the first agent model to understand various types of story structures, character development, and narrative techniques, thereby generating a coherent story framework that meets user requirements. During this process, the first agent model can automatically add conflict, twists, and endings to ensure the story is engaging and complete.
[0050] The script outline described above can be an overview of the basic framework and key events of the entire storyline, generated by the first agent model. The script outline can include the various stages of the animation story, such as the introduction, development, transition, and conclusion, as well as the main characters and their basic activities during these stages, providing information for subsequent character interpretation. For example, the script outline generated by the first agent model can be divided into four parts: the opening section, "The female explorer prepares to set off at the edge of the jungle"; the development section, "The explorer enters a dense forest filled with unknown creatures and uncovers the first clue to the Golden City"; the turning point, "The explorer encounters a ferocious beast, escapes with wisdom and courage, and discovers a secret passage to the ruins"; and the ending, "The explorer solves the ancient mystery in the ruins, reveals the truth about the Golden City, and returns safely with the treasure." The script outline can depict the story's trajectory and provide guidance for the subsequent tasks of each second agent model, helping each second agent model generate specific dialogue and actions for its character based on the instructions in the script outline.
[0051] In an optional embodiment, users can describe the basic elements of an animated story, such as theme, background, character attributes, and style requirements, through text input or voice input. This can define the basic outline of the animated story, set boundary conditions for subsequent creative generation, and ensure that the generated content is consistent with user needs. The first agent model can act as an intelligent content creator, performing in-depth analysis based on the script settings. Specifically, it can analyze key information in the script settings, such as the core conflict of the story, character personalities, and target scenes. Then, using machine learning algorithms, it constructs a story skeleton based on narrative rules such as "introduction, development, turn, and conclusion." This process can include plot weaving, character behavior prediction, and narrative style adaptation. The output of the first agent model can be a structured story outline, which can list the main events of the animated story, the initial settings of the characters, and the expected emotional direction. The generated script outline can include basic plot descriptions and can also annotate specific scene requirements, such as the environmental atmosphere, the interactive relationship between characters, and the overall narrative rhythm.
[0052] Through the above process, the script setting establishes clear boundaries for creation, and the first-agent model creates within this range, achieving a balance between creative flexibility and standardization. The intelligent creation capabilities of the first-agent model significantly improve the efficiency of script generation, significantly shorten the content production cycle, and reduce costs. Script settings are directly input by users, allowing them to deeply participate in the story conception from the early stages of creation, enhancing their creative participation and satisfaction. Through interactive design, users can easily become content creators, which can promote the development of the creative content field in a more open, interactive, and diverse direction.
[0053] In step S204 , the script outline is input into a plurality of second agent models respectively, and role interpretation information of a plurality of characters is generated using the plurality of second agent models.
[0054] Among them, the role interpretation information is used to represent the role's interpretation parameters in different scenarios.
[0055] The multiple second-agent models mentioned above can refer to independent intelligent models designed for each character within the multi-agent interactive animation generation framework, also known as character agent models. These second-agent models can be responsible for understanding the content related to a specific character in the script outline and, based on this information, generate interpretation information for that character, including specific parameters such as the character's movements and voice style in different scenarios. Each second-agent model can be customized based on pre-set attributes such as the character's personality, occupation, and emotional tendencies, enabling it to intelligently interpret the instructions in the script outline and respond appropriately to the character's characteristics. For example, in the aforementioned story about exploring a future world, assume there are two main characters: a brave female explorer and a smart robot assistant. The first agent model has generated a script outline describing their discovery and exploration of a mysterious forest. The female explorer's second agent model can parse the description of the female explorer's character in the outline, such as bravery and adventurousness, and generate a series of exploration-related actions, such as careful exploration gestures. The robot assistant's second agent model can generate corresponding interpretation information based on the robot assistant's intelligence and assistance, such as the voice tone when issuing warnings.
[0056] The aforementioned character interpretation information refers to the specific parameters and guidance set for each character's performance in different scenarios during the multi-agent animation generation process. This character interpretation information may include, but is not limited to, the character's dialogue text, voice intonation, body movements, and emotional responses in specific situations. This character interpretation information, generated by the second-agent model based on the script outline, can be used to guide the real-time rendering engine's rendering module in how to represent the character's behavior and emotional state. For example, when a female explorer discovers a mysterious relic, the character interpretation information could guide her to express curiosity and vigilance, including gestures such as a slight bend to observe carefully, and vocal tones that convey excitement and caution. When analyzing the relic data, the robot assistant could use calm and rational voice modulation to convey the results of its analysis. Within the real-time rendering engine, this interpretation information can be converted into detailed 3D animations, including the character's precise movements and matching voice output, presenting a vivid and coherent animated story to the user.
[0057] In an optional embodiment, the script outline generated by the first agent model can be parsed to extract key behaviors, emotional changes, and interactions with other characters in different scenarios for each character. This information can then be assigned to the corresponding second agent model to guide the second agent model in generating the character's interpretation parameters. Each second agent model can be designed to understand and generate interpretation information for a specific character, including by utilizing deep learning and natural language processing techniques. It can be pre-trained with relevant literature, film scripts, and animation data. The resulting second agent model can capture and reflect the character's personality traits, professional attributes, and emotional tendencies. Based on the script outline input, the second agent model can generate the corresponding character's reactions, including specific actions, dialogue text, and voice style. Fine-tuning can be performed based on the character's preset attributes to ensure consistency with the character's setting. Based on the scene descriptions in the script outline and the character's preset attributes, the second agent model can generate a series of action and voice style instructions. For example, for a courageous detective character in a dangerous investigation scenario, the second agent model can generate action instructions that require swift movement and vigilant observation of the surrounding environment, as well as firm and rhythmic voice delivery.
[0058] Through the above process, the customized design of each second-agent model makes the character's interpretation information more closely aligned with the character's personality setting, enhancing the character's realism and expressiveness, making the animation content more vivid and emotional. Users can fine-tune the character's interpretation information by adjusting the script outline, achieving flexibility and controllability in the creative process and reducing trial and error costs. By providing an interactive interface for character customization and interpretation information adjustment, users can deeply participate in the character's creation, enhancing the personalization and participation of the creation, and improving the overall user experience. The process of generating character interpretation information through the second-agent model can achieve improved animation content quality and improved creation efficiency.
[0059] Step S206: input the script outline and role interpretation information into the third proxy model, and use the third proxy model to generate the target animation.
[0060] The aforementioned third-agent model can be an intelligent model that performs the directorial role in the multi-agent interactive animation generation process. This model, also known as the director agent model, can be responsible for translating the script outline and character interpretation information into specific animation production instructions, guiding the animation synthesis. The third-agent model can understand the story's plot and emotions and master the technical details of animation production, such as character positioning, camera design, scene transitions, and overall narrative rhythm. The third-agent model can utilize a combination of intelligent algorithms, including but not limited to natural language understanding, scene analysis, visual design, and sound matching, to ensure high-quality and coherent animation content. Specifically, the third-agent model can deeply analyze the script outline, understand the story's context and key events, and then plan the character's movement trajectories and performance sequence across different scenes to ensure narrative logic coherence. Based on the character interpretation information, the third-agent model can schedule appropriate sound, movement, and background music resources, while also improving visual effects such as lighting, color, and visual effects to enhance the animation's immersion. The third-agent model can also work with a real-time rendering engine to transform text descriptions into dynamic 3D scenes. It also supports real-time user feedback and adjustments during the creation process, ensuring that the generated animation meets the user's creative vision.
[0061] The target animation described above refers to the final output of the multi-agent interactive animation generation process. It can be a visual representation of the story, orchestrated and refined by a third-agent model based on the script outline and character interpretation information. The target animation can include smooth visual effects, accurate character movements, and even incorporate background music and voice narration to form a complete and engaging animated story. For example, suppose the user inputs a script about a mysterious adventure. The first-agent model generates a detailed script outline describing the explorer's first encounter with a mysterious creature. Agent models for each character, such as the explorer and the mysterious creature, then generate character interpretation information, including the explorer's alert movements. Finally, the third-agent model combines this information with the script outline to plan camera movements for the explorer's cautious approach to the creature, set dynamic lighting effects for the forest background, and create appropriate background music. Ultimately, the animation generation model's real-time rendering technology presents the user with a tense and fantastical animated scene.
[0062] In an optional embodiment, a third proxy model can receive the script outline from the first proxy model and the character interpretation information generated by each second proxy model. The script outline provides the story's structure and context, while the character interpretation information includes parameters such as the specific movements and voice styles of each character in different scenarios. By understanding this information, the third proxy model begins to plan the technical details for converting the script outline and character interpretation information into animation. Specifically, based on the plot description in the script outline, the third proxy model can intelligently design camera movement schemes, including camera angles, dynamic effects such as push-pull pans and tilts, as well as character positioning and movement trajectories. It can also select appropriate camera movement techniques based on the story's rhythm and atmosphere to enhance visual impact and expressiveness. Simultaneously, it can automatically select background environments from a scene library that match the script outline description, and adjust details such as lighting, color, and special effects to create a plot-based atmosphere. Based on the character interpretation information, the third proxy model can intelligently retrieve and schedule appropriate resources from voice, action, and music libraries. The third proxy model can use vector similarity technology to quantitatively evaluate the match between resources and the script content, ensuring that the selected resources accurately convey the characters' emotions and personalities.
[0063] When integrating the script outline and character interpretation information, the third-agent model can also utilize an intent alignment mechanism to ensure that the final animation effect aligns with the user's creative vision and the intent of the script outline. The third-agent model can address the coherence of the overall narrative while also focusing on the details of each scene, such as character positioning adjustments and camera adjustments, to achieve precise and detailed animation. The animation production instructions generated by the third-agent model can be converted into dynamic images through the animation generation model's real-time rendering engine. During this process, users can preview the animation effects in real time and intervene and provide feedback through the agent set mode. The third-agent model can quickly adjust camera movements, positioning, or resource selection based on user feedback, making the creative process interactive and controllable.
[0064] Through the above process, the third-agent model's intelligent design and resource improvement capabilities enable professional-grade visual and auditory quality in animation content, enhancing the story's immersion and enjoyment, and providing users with high-quality creative expression. During the multi-agent animation generation process, the third-agent model can improve the animation's technical quality and enable efficient interaction and resource improvement during the creative process.
[0065] Alternatively, suppose a user wants to create an animation about a brave female explorer searching for ancient ruins in a mysterious forest. The user might enter a scenario setting like this: A brave female explorer journeys deep into a mysterious jungle, attempting to find the legendary Golden City. The first agent model can then generate a script outline based on this scenario setting. For example, the first agent model might generate a script outline with four main parts: At the beginning, the female explorer prepares to depart at the edge of the jungle; in the middle, the explorer enters a dense forest teeming with unknown creatures and uncovers the first clue to the Golden City; in the middle, the explorer encounters a ferocious beast, escapes with wit and courage, and discovers a secret passage leading to the ruins; and finally, the explorer solves the ancient mystery within the ruins, reveals the truth about the Golden City, and returns safely with the treasure. The generated script outline outlines the story's development and includes basic information about multiple characters and key scenes, such as the jungle, the ruins, and the Golden City.
[0066] Next, character interpretation information can be generated. The script outline can be broken down and fed into multiple second-agent models. For example, for a female explorer, the second-agent model would generate the following character interpretation information: personality traits (decisive, resourceful, and courageous in facing challenges); occupational description (experienced explorer, skilled in puzzle solving and wilderness survival); behavioral instructions (expressing fear but remaining brave when facing wild beasts, and demonstrating wisdom when solving puzzles). Similarly, other second-agent models would generate corresponding interpretation information, such as the attacking behavior of wild beasts and the gentle guidance of a wise elder providing clues. Finally, a third-agent model can generate the target animation. This third-agent model integrates the script outline and character interpretation information to plan the visual presentation and dynamic flow of the entire animation. It selects appropriate background and ambient music for each scene and designs camera movements to ensure the audience follows the female explorer's perspective as she experiences the adventure. Furthermore, the third-agent model meticulously plans the character's positioning and movement to ensure natural and smooth movements that align with the character's personality and plot development. For example, in scenes where wild beasts appear, tense camera movements can be used to enhance the sense of urgency. The third-party model can convert the script outline and role interpretation information into a format that can be understood by the real-time rendering engine, and call the real-time rendering function to generate the final animation video.
[0067] Through the above steps, not only can the animation generation cycle be significantly shortened and the animation generation cost be reduced, but it also provides a highly participatory, interactive and personalized creation process. This multi-agent collaborative working method imitates the real-world production team and greatly improves the creation efficiency through automated processes, while also maintaining the consistency and depth of the content.
[0068] In this embodiment of the present application, the script settings are first input into a first proxy model, which then generates a script outline. The script settings describe the story's plot settings, provide the story's plot framework, and include information about multiple characters. Next, the script outlines are input into multiple second proxy models, which then generate character interpretation information for each character. This character interpretation information displays the character's interpretation parameters in different scenarios. Finally, the script outlines and character interpretation information are input into a third proxy model, which then generates the target animation. It is easy to notice that the above process decomposes the animation generation process into multiple specialized models, namely the first agent model, multiple second agent models and the third agent model. Each model focuses on handling a specific type of creative task and realizes professional division of labor. Specifically, the first agent model can quickly generate a script outline according to the script settings, quickly capture the core elements of the plot, and construct a story framework that conforms to the narrative logic, providing a clear direction and structure for subsequent animation production. Multiple second agent models can generate detailed interpretation parameters for their respective characters, and realize automation and personalization through intelligent agents. The third agent model can automatically plan the scene layout and lens application based on the script outline and character interpretation information, thereby significantly shortening the animation generation cycle and production cost; at the same time, the use of multiple models with specialized division of labor can support users to provide real-time interactive guidance to each model in each link of animation generation. Each process of animation generation is user-controllable, which can meet the user's personalized needs for animation content, thereby solving the technical problems of low animation generation efficiency and generation effect in related technologies.
[0069] In the above embodiment of the present application, the script outline and character interpretation information are input into the third agent model, and the target animation is generated using the third agent model, including: inputting the script outline and character interpretation information into the third agent model, and generating script audio-visual information using the third agent model, wherein the script audio-visual information is used to plan the movement trajectories of multiple characters; inputting the script outline, script audio-visual information, and character interpretation information into the fourth agent model, and generating camera movement information using the fourth agent model; inputting the script setting, script outline, script audio-visual information, character interpretation information, and camera movement information into the third agent model, and generating a target script using the third agent model, wherein the target script is used to represent the story plot description; and generating a target animation based on the target script.
[0070] The script audiovisual information mentioned above can refer to the detailed information used to guide character performances, environment construction, and visual and auditory experience during the animation or video creation process. This information can include character movements, audio format of dialogue, scene layout, background music selection, and other content. The specific script audiovisual information can be determined based on actual needs and is not limited here.
[0071] The aforementioned fourth agent model can refer to an intelligent agent or model within a multi-agent animation generation system specifically responsible for planning and designing camera usage and shooting techniques for animation or video content, also known as a camera operator agent model. Based on the script outline and audio-visual information, the fourth agent model can automatically or semi-automatically generate camera strategies, including camera angles, focal lengths, motion trajectories, and specific camera techniques such as when to use wide-angle, close-up, and tracking shots.
[0072] In an optional embodiment, a third-party model can receive a script outline and character interpretation information. The script outline provides the basic structure and plot direction of the story, while the character interpretation information includes details such as the movements and voice styles of each character in different scenes. The third-party model deeply integrates these two pieces of information to ensure that the character's actions and emotional expressions are closely aligned with the plot development. Based on this integrated information, the third-party model can generate audio-visual information about the script, which can include planning the character's movement trajectories in the scene. By analyzing the emotional cues and narrative rhythms in the script outline and character interpretation information, the third-party model intelligently designs the movement trajectories of multiple characters, enhancing the story's expressiveness and the audience's sense of immersion.
[0073] Next, the script outline, audio-visual information, and character interpretation information can be passed to the fourth-proxy model, which is responsible for refining the camera movement plan to ensure the accuracy and artistry of the shots. The fourth-proxy model can adjust the dynamic details of the shots, such as camera speed and focus shifts, based on the story's narrative style and emotional changes, to make the camera movement more consistent with the plot and enhance the viewing experience of the animation. The third-proxy model then aggregates the script outline, audio-visual information, character interpretation information, and camera movement information to generate a target script. The target script can detail the storyline, character behavior, emotional expression, camera movement techniques, and background music selection. Based on the target script, a real-time rendering engine can be used to convert the script's descriptions into dynamic 3D animation, completing the transformation from text to audio-visual content.
[0074] Through this process, the collaborative work of the third- and fourth-agent models ensures the animation's narrative structure and emotional coherence, as well as the artistry and technical precision of its camera work, significantly enhancing its narrative effectiveness and visual appeal. By separating the responsibilities of script understanding, character interpretation information generation, and camera movement design, the multi-agent framework achieves specialized division of labor in the animation production process, improving the efficiency and quality of content generation.
[0075] In the above embodiment of the present application, a target animation is generated based on a target script, including: determining the actor identification information and scene identification information corresponding to multiple roles based on the target script; calling the actor model corresponding to the actor identification information, and calling the scene model corresponding to the scene identification information; performing timing control on the actor model and the scene model based on the target script to obtain an animation script; and rendering the animation script to obtain a target animation.
[0076] The actor identification information mentioned above can refer to the data tags used to uniquely identify and locate specific virtual characters in animation or virtual production environments. In a multi-agent collaborative animation generation system, when a character is created or selected, the system can assign it an identifier that contains all relevant information about the character, such as the character identifier, personality traits, appearance model, action library index, and voice style.
[0077] The above-mentioned scene identification information may refer to data tags used to describe and distinguish specific scenes or environments in animation or video content creation. It may include key information such as scene identifier, environment type, background description, lighting conditions, sound effect settings, etc., so that the corresponding scene models and environmental elements can be quickly identified and called during the animation generation process.
[0078] In an optional embodiment, after obtaining the target script, the third proxy model can analyze the identities and appearance scenarios of each character in the target script to determine actor and scene identification information. This step can be achieved by performing deep semantic analysis of the target script using natural language processing technology, which can identify and label key information such as each character's appearance time, behavior, dialogue, and scene changes. Next, based on the parsed actor and scene identification information, the third proxy model can retrieve corresponding 3D models from a pre-set actor model library and scene model library. The actor model library can contain the appearance and motion capture sequences of various characters, while the scene model library can contain a rich variety of environment models, including indoor and outdoor environments, and buildings in specific cultural contexts. Model retrieval can rely on a search mechanism, using a vector similarity algorithm to match the script description with model characteristics, ensuring that the retrieved model is closely aligned with the script settings, thereby enhancing the realism and immersiveness of the animation. After retrieving the actor and scene models, the third proxy model can perform precise timing control on the actor and scene models based on the timing information in the target script to generate an animation script. Finally, the third proxy model can feed the generated animation script into a real-time rendering engine to achieve the target animation. The real-time rendering engine combines 3D models, actions, and special effects based on the instructions in the script to render a coherent picture in real time.
[0079] Through the above process, scripted and modeled timing control, combined with the powerful rendering capabilities of the real-time rendering engine, can quickly generate high-quality animation content, significantly shortening the conversion time from script to visualization while ensuring professional-level animation visual effects. The combination of animation scripts and real-time rendering supports real-time user interaction through the agent set mode, such as modifying scripts and adjusting character behaviors. Through detailed scene construction and camera movement design, the audience's sense of immersion is enhanced.
[0080] In the above embodiment of the present application, the script setting is input into the first agent model, and the script outline is generated using the first agent model, including: screening multiple preset materials based on the script setting to obtain at least one first material, wherein the at least one first material is a material among the multiple preset materials whose matching degree with the script setting is greater than the first preset matching degree; based on the script setting and the at least one first material, screening the at least one first material to obtain at least one second material, wherein the at least one second material is a material among the at least one first material whose matching degree with the script setting is greater than the second preset matching degree; inputting the at least one second material and the script setting into the adjustment model, adjusting the at least one second material using the adjustment model to obtain at least one material; and generating a script outline based on the at least one material.
[0081] In an optional embodiment, the first agent model can use deep natural language processing technology to extract key themes, emotional tone, and stylistic requirements from the script setting text input by the user, thereby forming a preliminary creative outline for the script. Next, a vector similarity algorithm can be used to compare the script setting with multiple preset materials in the material library to identify materials that match the setting. These materials can include plot outlines, character descriptions, scene layouts, or background music. In the first round of screening, a high matching threshold (i.e., a first preset matching threshold) can be set to select materials that closely match the script setting requirements, forming a first material set. This ensures a high correlation between the basic materials and the script setting, providing a high-quality starting point for subsequent screening. The screening and materials can then be refined for a second round of screening. Based on the first material set, a matching threshold (i.e., a second preset matching threshold) can be set to further refine the materials, further eliminating options that do not match the script setting style, resulting in a second material set. During the screening process, you can consider the similarity of the text content and comprehensively analyze the emotional color, style characteristics and theme fit of the material to ensure that the material is consistent with the script setting in multiple dimensions.
[0082] Next, the second set of assets and the script settings can be input into the adjustment model, which is responsible for calling and integrating the assets from the second set of assets, while also adjusting the assets based on the contextual information of the script settings. The adjustment model can enhance the generation mechanism through dynamic retrieval, combining it with the specific needs of the script settings, such as emotional tendencies and stylistic preferences, to fine-tune the assets to ensure they meet the script's vision. The adjusted assets can then be used by the first agent model to construct a script outline, which details the story structure, character interactions, emotional ups and downs, and scene transitions, laying a solid foundation for subsequent character interpretation information generation and animation production.
[0083] Through the above process, multiple rounds of screening and adjustment mechanisms can ensure a high degree of consistency between the materials and the user's creativity, significantly improving the creative quality and user satisfaction of the animation content. Through the adaptive adjustment model, the materials can be fine-tuned according to the user's specific requirements in the script setting, realizing personalized creation from creativity to script while maintaining a high degree of controllability. The combination of the retrieval enhancement generation mechanism and multi-dimensional matching analysis can improve the accuracy and efficiency of material screening, reduce ineffective creative attempts, and shorten the time cycle from creativity to finished product. By calling on a variety of preset materials and combining them with the user's unique script settings, it is possible to generate creative and diverse script outlines, promoting innovation and diversity in content creation.
[0084] In the above embodiment of the present application, a script outline is generated based on at least one material, including: generating an initial script outline based on at least one material; expanding the initial script outline to obtain script text information, wherein the script text information is used to represent the plot information of the script; determining character actions and character voices that match the script text information; generating script role information based on the character actions and character voices, wherein the script role information is used to represent the interpretation information of the character; and obtaining a script outline based on the script text information and the script role information.
[0085] In an optional embodiment, matching materials can be extracted from a pre-set resource library based on the user's selected genre, plot setting, and character information. These materials can include basic plot frameworks, character settings, or specific scene descriptions. The first agent model uses these materials as input and, using deep learning technology, automatically generates an initial script outline. This outline can include the main plot, character interactions, and key plot turning points, providing a foundational structure for subsequent script expansion. A language model can be used to further expand and enrich the initial script outline, generating detailed script text information, including character dialogue, voiceover descriptions, emotional expressions, and scene details. To ensure the script's coherence and richness, an iterative generation strategy can be employed to gradually refine each plot node. While generating dialogue information, the first agent model can automatically identify points where voiceover is needed, namely, those sections where voiceover is needed to supplement contextual or emotional details. By inserting voiceover information at these points, the script's narrative structure is more complete, significantly enhancing the audience's sense of immersion. Based on the script text information, the action and voice of each character can be matched. By analyzing the character's personality, emotional state, and plot requirements, the second agent model selects appropriate action sequences and voice styles from a library of actions and voices. To ensure the authenticity and coherence of the character's performance, it considers the superficial matching of action and voice. Using sentiment analysis and characterization techniques, it refines character interpretation parameters such as speech rate, intonation, and movement fluency, ensuring that the character's behavior more closely matches the intended characteristics. Finally, the third agent model integrates script text information with script character information to generate a complete script outline containing character performance details, shot variations, and scene layout, providing detailed guidance for subsequent animation generation.
[0086] Through the multi-stage script generation process described above, the expansion of the script outline and the refinement of text information can significantly improve the creative realization and quality of the script, making the story richer and more engaging. The precise matching of character movements and voices, combined with the setting of personalized interpretation parameters, makes the characters in the animation more realistic and natural, enhancing the immersiveness of the story and the emotional resonance of the audience. By breaking down the script generation process into multiple stages, not only can the initial outline be quickly generated, but it can also be targeted and adjusted in subsequent stages, significantly improving the efficiency and controllability of creative content creation. The process of matching movements and voices for each character not only takes into account the maximum utilization of resources, but also supports flexible editing and reuse of content through textual storage strategies, reducing overall creation costs and improving resource utilization efficiency.
[0087] In the above embodiment of the present application, the initial script outline is expanded to obtain script text information, including: expanding the initial script outline to obtain dialogue information and narration information; identifying the dialogue information to obtain narration demand points of the dialogue information; adding narration information to the narration demand points of the dialogue information to obtain script text information.
[0088] The above-mentioned dialogue information may refer to the verbal content of direct conversations between characters in the script, which may include the words spoken by the characters, the order of the dialogues, the tone, the emotional color, and the voice characteristics.
[0089] The above-mentioned narration information can refer to the non-dialogue narrative content in the script used to provide background information, describe scene details, or explain the inner thoughts and motivations of the characters. For example, narration can be used to describe the complex inner feelings of the characters or the historical background of the story.
[0090] In an optional embodiment, the first agent model can deeply expand upon the initial script outline to generate specific character dialogue content. This step takes into account the character's personality, emotional state, and plot context, making the dialogue more closely aligned with the character setting and enriching the story's plot and inter-character interactions. Using natural language processing technology, the first agent model can intelligently identify contextual points in the script dialogue that require narration. These contextual points can include descriptions of the character's inner thoughts, the surrounding atmosphere, or undirected plot details, contributing to the overall narrative's coherence and depth. The first agent model can then generate appropriate narration text based on the context of the narration point to supplement the implicit or unexpressed plot information in the dialogue, ensuring the story's integrity and audience immersion. The first agent model accurately embeds the generated narration information into the corresponding position in the script outline, naturally integrating it with the dialogue information to form the script text information, ensuring a seamless connection between the narration and dialogue, and avoiding abrupt or logical inconsistencies.
[0091] Through the above process, by intelligently identifying narration needs and generating narration information, we can supplement the plot context and character inner thoughts that are not expressed in the dialogue, thereby enriching the story's layers, enhancing the overall narrative integrity and the audience experience. The addition of narration information not only provides story context and explanations of character motivations, but also enhances the audience's immersion by creating atmosphere and context, allowing them to more deeply understand the plot and feel the characters' emotions. The first agent model can reduce the trial-and-error and iteration costs of manual creation by automatically expanding dialogue and narration information. At the same time, through deep learning and consistency assurance technology, it improves the logic and artistry of the generated content, making the expansion of the script outline both fast and high-quality.
[0092] In the above-mentioned embodiment of the present application, determining the character action and character voice that match the script text information includes: matching the script text information with multiple preset actions to obtain an action matching result; matching the script text information with multiple preset voices to obtain a voice matching result; quantitatively evaluating the action matching result based on at least one dimension to obtain a first evaluation result, and quantitatively evaluating the voice matching result based on at least one dimension to obtain a second evaluation result; determining the character action from multiple preset character actions based on the first evaluation result, and determining the character voice from multiple preset character voices based on the second evaluation result.
[0093] In an optional embodiment, a first agent model can perform in-depth analysis of the script text, extracting key emotional states, action requirements, and character traits to form a preliminary description of the character's required actions and voice style. A second agent model then matches this preliminary description with a pre-set action library and, by calculating the vector similarity between the text description and the action data, screens a series of candidate actions that match the plot. Similarly, the second agent model can compare the text description with a pre-set voice library to identify candidate voice samples that embody the character's characteristics and emotional states. Next, a multi-dimensional quantitative evaluation can be performed. For action evaluation, a quantitative evaluation system based on multiple dimensions, such as emotional expression, action coherence, and visual effects, can be introduced to score the action matching results, resulting in a first evaluation result. For voice evaluation, metrics such as voice clarity, tone consistency, and emotional communication can be used to evaluate the voice matching results, resulting in a second evaluation result. Finally, based on the first evaluation result, action sequences that best meet the plot requirements and visual effects can be selected from the pre-set character actions. Based on the second evaluation result, voice samples that best match the character's personality and emotional expression can be selected from the pre-set character voices.
[0094] Through the above process, meticulous matching and quantitative evaluation of movement and voice can generate natural performances that match the character's settings, enhancing the character's realism and ensuring the coherence and logic of the character's behavior and dialogue throughout the story. Based on the evaluation of emotional expression and visual effects, more appropriate movements and more emotionally appropriate voices can be selected, enhancing the visual impact of the plot while ensuring effective emotional communication and enhancing the audience's sense of immersion and emotional resonance. The multi-dimensional quantitative evaluation mechanism accelerates the movement and voice selection process while also ensuring user control over the generated results, allowing users to quickly find suitable matches even in a large preset library.
[0095] In the above embodiment of the present application, the method also includes: obtaining story plot information; in response to receiving a style setting instruction for the story plot information, rewriting the story plot information based on the style setting instruction to obtain a script setting; in response to receiving a plot direction instruction for the story plot information, continuing to write the story plot information based on the plot direction instruction to obtain a script setting; in response to not receiving a style setting instruction and / or a plot direction instruction, determining that the story plot information is a script setting.
[0096] The aforementioned style-setting instructions can refer to instructions issued by users when creating animation or video content to guide the specific artistic style and atmosphere of the story's expression. They can express the user's aesthetic preferences and creative requirements for the target animation in terms of visuals, narrative, and emotion. They can be subdivided into various types, such as visual style, narrative style, and musical style. Style-setting instructions can be input at the beginning of the user's creative process to customize the tone and genre of the story, ensuring that the generated content conforms to the specific style desired by the user. For example, if the user wants to create a suspense-style animation, the user can indicate this in the style-setting instructions. The first agent model will adjust the plot direction based on the style-setting instructions, select plots and settings with suspense elements, and design intriguing character dialogues to ensure that the entire script outline presents the characteristics of a suspenseful style.
[0097] In an optional embodiment, story plot information can be received through a user interface. This story plot information can be a user-created text description. The first agent model can deeply understand and analyze the input information, extracting key plot points, character traits, and scene settings, laying the foundation for subsequent stylistic rewriting and plot development. When the user issues a style setting instruction through the interface, such as requesting a plot change to a suspense, romance, or comedy style, the first agent model can rewrite the story plot information based on this instruction. Specifically, the first agent model can use natural language processing technology to identify stylistic characteristics in the input text, establishing a starting point for subsequent rewriting. Based on different stylistic instructions, the first agent model can reorganize the plot structure, adjust the plot direction, and modify character dialogue and scene descriptions to meet the requirements of the new style. During the rewriting process, a preview function can be provided to allow the user to review the rewriting results and make iterative adjustments based on feedback. When the user specifies a specific plot development, such as a protagonist's counterattack, a falling out, or a reconciliation, the first agent model will continue the plot based on these instructions.
[0098] Then, through the pre-set script generation algorithm and multi-stage reasoning mechanism, the first agent model can predict the plot trend and predict the possible development path of the story based on the plot direction instructions input by the user; automatically generate plots, which can automatically generate detailed plots that match the plot direction instructions, including conflicts, reconciliations, adversities and turning points between characters, making the story richer and more fascinating; dynamically adjust the plot details according to the user's feedback on the generated content to ensure that the story direction meets user expectations; automatically determine the script setting, when the user does not provide a specific style or plot direction instruction, it can automatically analyze the initial story plot information, and generate the script setting based on the internal logic and emotional orientation of the content to ensure the natural development and creative expression of the story.
[0099] Through this process, users can directly participate in the script's style and plot direction, making the generated content more tailored to personal preferences and enhancing the interactivity and personalized experience of the creative process. By rewriting and continuing the storyline information, logically coherent and plot-rich story content can be generated, significantly improving the quality and artistic expression of the animation script. At the same time, the iterative improvement mechanism based on real-time user feedback can quickly respond to changes in user needs, increasing the flexibility and efficiency of content generation. In addition, even if the user does not provide style or plot direction instructions, the system can automatically complete the generation of the script settings, lowering the threshold for user participation and stimulating the creative potential of more people.
[0100] In the above-mentioned embodiment of the present application, obtaining story plot information includes: in response to receiving a role selection instruction from a user, determining at least one first target role corresponding to the role selection instruction from a plurality of preset roles; in response to receiving the user's style setting information and plot development information, generating story plot information based on the at least one first target role, the style setting information and the plot direction information.
[0101] In an optional embodiment, upon receiving a character selection instruction from a user, natural language processing can be used to parse the instruction and identify the type or specific character of interest. This process leverages the deep learning model's ability to understand text semantics, accurately capturing user preferences such as the character's gender, age, personality traits, and descriptions relevant to the specific story context. After capturing the user's style preferences, such as a desired romance or suspense story, and plot development information, such as a desired protagonist's growth, challenges, or ultimate victory, a first agent model can integrate the style preferences and plot development information based on at least one first target character selected by the user to generate story plot information by: utilizing a pre-trained style classification model to extract keywords and descriptions from the style preferences to form style labels; and analyzing the plot development information to predict the story's trajectory and key turning points, providing guidance for plot design. The first agent model matches the style characteristics and plot development with the target character's attributes, such as personality, background story, and skills, to generate the character's behavioral patterns and interactive plots within the specified style and plot context. Finally, text generation techniques, such as generative models, can be used to automatically generate the story framework and plot details based on the characters, style, and plot direction. Subsequently, through an iterative improvement process, the coherence and appeal of the plot can be adjusted to meet higher generation quality standards.
[0102] Through the above process, users can directly choose their favorite characters and set the story style and direction, which can increase the autonomy and personalization of creation, and can significantly improve the satisfaction and retention rate of user participation in creation. By quickly matching preset characters and plot directions, the story framework can be quickly generated, reducing the trial and error cost of creation from scratch, and improving the efficiency of the creation process. Combined with user instructions and intelligent generation technology, it can generate high-quality story plot information with rich plots, rich emotions and meeting specific style requirements.
[0103] In the above embodiment of the present application, the method also includes: in the process of generating a target animation based on a target script, in response to receiving an interaction instruction, performing intent recognition on the interaction instruction to obtain an intent recognition result, and determining a target interaction mode corresponding to the interaction instruction from multiple interaction modes based on the intent recognition result; and generating interaction information between the user and at least one of the multiple characters based on the target interaction mode.
[0104] In an optional embodiment, users can issue specific interactive commands through the system interface, such as changing character behavior, adjusting dialogue, or modifying the plot's direction. The system can use natural language understanding technology to deeply analyze these interactive commands and identify the user's true intent, such as requesting a specific action, adjusting the emotional expression of a dialogue, or changing the plot's direction. Based on the identified user intent, the system can determine a suitable target interaction mode from multiple pre-set interaction modes, which can include modifying the script, adjusting character performance, participating in a puzzle game, or discussing the plot. The system can then apply the target interaction mode to the current script generation process, providing the user with the opportunity to interact with the target character. In chat mode, the system can use a classifier agent within the multi-agent orchestration framework to analyze user speech content and select an appropriate character agent to respond. The generated interactive information includes the character's dialogue and reactions, enabling users to engage in natural and fluent conversations with the virtual character. In game mode, the system can use a host agent to control the order and method of interaction between the user and the virtual character based on game tasks. The generated interactive information includes the user's and the character's in-game behavior and feedback, increasing interactivity and entertainment through puzzle solving or task completion.
[0105] Through the above process, users can directly participate in the creation of animation content, and enhance the immersive and personalized experience of creation by adjusting the script, character behavior or participating in interactive games. By receiving user intentions in real time and adjusting the animation generation process, the system can dynamically improve according to user feedback, making the generated content more flexible and diverse, while stimulating the diversity of creation. Users can adjust unsatisfactory parts in a timely manner, avoiding the time-consuming process of repeated revisions and full regeneration in the creation process, significantly improving the efficiency of creation. The system seamlessly integrates user instructions with intelligently generated content, realizing the transformation of the creation mode from one-way output to two-way interaction, promoting deep collaboration between users and intelligence, and improving the intelligence and interactivity of creation.
[0106] In the above-described embodiment of the present application, generating interaction information between a user and at least one of multiple characters based on a target interaction mode includes: in response to the target interaction mode being a chat interaction mode, obtaining a first speech from the user, inputting the first speech into a classifier proxy model, analyzing the first speech using the classifier proxy model to obtain a first character associated with the first speech; inputting the first speech into a first target proxy model of the first character, and generating a second speech using the first target proxy model; and generating interaction information based on the first speech and the second speech. In an optional embodiment, the user can input the first speech through the system's interactive interface, which can include questions, comments, or suggestions for the plot. The system can pre-process the user's speech, including text cleaning, sentiment analysis, and intent recognition, to ensure the clarity and relevance of the input content. The pre-processed first speech can then be input into the classifier proxy model, which can analyze the relevance of the speech to various characters in the script, including emotional resonance, information needs, or relevance to plot progression. Based on the relevance analysis, a first character can be determined, which can be the character most relevant to the current speech, to respond. The first speech content can then be passed to the first target proxy model of the first character. The first target proxy model can generate the second speech content, that is, the character's response, based on the character's settings, such as personality, occupation, background story, etc. When generating a response, the first target proxy model can maintain emotional expression and dialogue style consistent with the character settings to enhance the character's realism and coherence. Finally, the system can integrate the first speech content and the second speech content to generate interactive information containing user questions or comments and character responses, providing users with a dialogue experience with virtual characters. The interactive information can be recorded in the dialogue history, and the classifier proxy model updates the global dialogue state in real time to ensure the coherence and consistency of subsequent dialogues with historical content.
[0107] Through the above process, through chat-style interaction, users can not only watch the generated animation, but also have conversations with virtual characters, which enhances the user's immersive experience and sense of creative participation, making the user a part of the animation story. The responses generated by the first-target agent model take into account the character settings and dialogue context, significantly improving the realism and personalization of the character reactions, making the interaction of virtual characters more vivid and natural. Through the role selection of the classifier agent model and the personalized responses of the first-target agent model, the system can provide accurate and coherent interactive information, avoiding incoherence in the dialogue process and improving user satisfaction.
[0108] In the above embodiment of the present application, interaction information between the user and at least one of the multiple characters is generated based on the target interaction mode, including: in response to the target interaction mode being the game interaction mode, determining the second target agent model of the second character from multiple second agent models based on the game task; determining the game participation order of the user and the second character; controlling the user and the second character to perform the game task in the first interaction round based on the game participation order, and obtaining an execution result, wherein the execution result is used to indicate whether the game task is successfully performed; in response to the execution result being that the game task is not successfully performed, controlling the user and the second character to perform the game task in multiple interaction rounds based on the game participation order until the execution result is that the game task is successfully performed, and generating interaction information based on the interaction information of the first interaction round and the multiple interaction rounds; in response to the execution result being that the game task is successfully performed, generating interaction information based on the interaction information of the first interaction round.
[0109] In an optional embodiment, based on the current game task, a second target agent model for one or more second characters can be intelligently screened from multiple second agent models and determined. This process can be implemented using retrieval-enhanced generation technology, which retrieves task-related character skills and background information from an external game knowledge base to ensure that the selected character is suitable for the current game context. Next, the user's participation order with the second character can be automatically planned based on the complexity of the game task and the character's characteristics, ensuring a smooth game flow and reasonable character interaction. The participation order can be dynamically adjusted based on game progress and user performance to more efficiently advance the game task. In the first round of the game, the user and the second character can be controlled to perform game tasks, such as solving puzzles, exploring, or selecting dialogue options. Based on the completion of the tasks, an execution result is generated. If the first round of execution results in failure, the system can guide the user to interact with the second character for multiple rounds based on the established participation order until the task is successfully completed. During this process, the system can continuously record and analyze the interaction information between the user and the character, providing data support for subsequent strategy adjustments. When a task is successfully completed, interactive information can be generated based on the interaction information from the first interaction round, providing a concise and efficient interactive experience. For tasks that are not successfully completed on the first try, more detailed interaction information can be generated by combining the interaction information from the first interaction round and subsequent rounds, including task attempts, failure analysis, and success strategies, to enhance the educational and entertainment value of the game.
[0110] Through the above process, the game interaction mode provides users with an opportunity to cooperate with virtual characters to complete tasks, enhances the interactivity and entertainment of the creative process, and enables users to enjoy the fun of the game while participating in the creation, significantly improving the sense of immersion. Through multiple rounds of interactive attempts, the system can more deeply display the characteristics of the characters, such as wisdom, courage and team spirit. At the same time, through the feedback and dialogue of the characters, it enriches the inner world of the characters, making the characters more vivid and three-dimensional. As part of the creative process, game interaction can not only inspire users' creative inspiration, but also allow users to continuously improve their script settings during the interaction, realizing a deep combination of creative generation and user participation.
[0111] In the above embodiment of the present application, the method also includes: obtaining feedback information of the user and the second character performing the game task; inputting the feedback information into the third target agent model, and using the third target agent model to detect the feedback information to obtain a detection result, wherein the detection result is used to indicate whether the feedback information is preset information; in response to the feedback information being preset information, determining that the execution result is successful execution of the game task; in response to the feedback information being non-preset information, determining that the execution result is unsuccessful execution of the game task.
[0112] The above-mentioned third-target agent model can be a host agent model. The third-target agent model plays the role of activity organizer and process controller in the game mode. It can be responsible for guiding and maintaining the progress of the game in the interactive game between the user and the animated character agent, ensuring the continuity and fun of the game experience.
[0113] In an optional embodiment, real-time feedback from the user and the second character during the execution of a game task can be collected. This information may include, but is not limited to, user choices, character reactions, environmental changes, and task progress. The collected feedback information can be input into a third target proxy model, which can be responsible for in-depth analysis and detection of the feedback information. The third target proxy model can check whether the feedback information matches the preset correct execution path or solution, that is, whether it matches the preset information. If the feedback information matches the preset information, the execution result can be determined as successful execution of the game task, and the next step of the creation process or the release of rewards can be entered. If the feedback information does not match the preset information, the execution result can be determined as unsuccessful execution of the game task, and the user and character can be guided to try again or provided with prompts. In the event of unsuccessful execution, the system not only provides error messages but also generates personalized guidance based on the user's performance and game goals, encouraging users to solve the problem from different angles. Based on user feedback, the system can dynamically adjust the game difficulty and strategy to ensure the adaptability and challenge of the gaming experience.
[0114] Through the above process and intelligent processing of feedback, the system provides users with timely and accurate feedback on execution results, enhancing the game's interactivity and challenge, thereby improving user satisfaction and long-term retention. The virtual characters' responses and adjustments based on this feedback make their behavior more logical and contextual, enhancing their realism and the immersive gaming experience. By collecting and analyzing feedback, the system can intelligently adjust individual creative strategies, avoiding ineffective creative attempts and improving overall creative efficiency and content quality.
[0115] The technical solution proposed in this application is described below in conjunction with an optional embodiment. This application proposes a multi-agent interactive animation generation method. The multi-agent interactive animation generation system proposed in this application constructs a full-process animation generation from plot generation, resource call to real-time rendering, solving the industry pain points of long animation production cycles and high costs in related technologies. During the intelligent collaborative creation process, users can adjust the plot in real time, effectively solving the uncontrollable problems commonly found in video models. This application can intelligently call various material libraries to ensure the diversity and richness of creative content. During the animation generation waiting stage, the proposed agent on-set interaction mode provides dual-path interaction of puzzle games and plot discussions. Users can collaborate with actor agents to decipher puzzles, or they can deeply explore character motivations and story development, significantly improving creative immersion and user retention. This application lowers the threshold for animation production to a level that the public can participate in, realizing the paradigm upgrade of intelligent-driven content creation from one-way output to two-way interaction, and opening up a new editable and interactive content production path for the digital entertainment industry. This application builds an interactive animation generation system based on a multi-agent collaborative framework and retrieval enhancement generation technology, allowing ordinary users to independently design plots through simple text input. The system automatically calls the retrieval enhancement generation enhancement technology and the real-time rendering capability of the real-time rendering engine to convert user creativity into visual animation content, turning content consumers into content creators, who can independently create and instantly watch short dramas that suit their personal tastes, bringing an agile and efficient creative experience and a new interactive model to the digital content industry, and reconstructing the production and consumption paradigm of creative content.
[0116] This application utilizes a multi-agent collaborative framework and retrieval-enhanced generation technology to construct a full-process interactive animation creation system, from plot text understanding, character action generation, to 3D scene rendering, enabling ordinary users to independently create animated content. The system first uses the multi-agent framework to expand the plot and characterize based on user-entered plot concepts, providing a structured script for subsequent performances. The system then utilizes a large-scale language model fine-tuned with film and television animation data to generate character dialogue, emotions, and behavioral instructions. Simultaneously, retrieval-enhanced generation technology retrieves appropriate resources from script, music, and action libraries. Finally, a real-time rendering engine renders high-quality 3D character animations in real time. During the waiting phase, the system innovatively introduces an intelligent studio feature, allowing users to choose between game mode to interact with characters through puzzles or chat mode to discuss plot developments with character agents, significantly improving user retention. Users can also intervene in the subsequent creation process to adjust the plot direction or character behavior, and the system will regenerate subsequent content based on the new instructions. This framework enables rapid, customized production of animation content, providing ordinary users with a channel for creative expression without requiring specialized skills, significantly shortening the animation creation cycle. This application currently serves animation producers, supports the output of secondary character creation content, and brings a new production model to the field of digital content creation.
[0117] Figure 3 is a schematic diagram of an optional animation generation process according to an embodiment of the present application, such as Figure 3 As shown, the previous story can be rewritten or continued, and the style can be selected when rewriting, and the plot direction can be set when continuing; then, based on the basic settings: character profiles and scene descriptions, the screenwriter can select background music from the music library, determine the script reference from the script library, and generate a story outline; the director can select scenes / soundtracks based on the generated story outline and generate initial positions; the screenwriter can generate lines / narration based on the generated story outline; the screenwriter can determine the appropriate voice from the voice library, determine the appropriate action from the action library, and select voice actions; the actor can perform role-playing based on the role settings: personality traits, career description; the director performs role-playing, moves the character based on the initial positions generated above, and generates camera movements; the director can perform final proofreading, combine with the animation script, and use the real-time rendering engine to drive real-time generation to obtain animation.
[0118] The above-mentioned basic settings can be the script settings in this application. The agent: screenwriter can be the first agent model; the story outline can be the script outline; the agent: director can be the third agent model; the agent: actor can be multiple second agent models, and the agent: camera operator can be the fourth agent model.
[0119] The specific process steps for the interactive animation generation of multi-agent collaboration in this application are as follows: the overall system architecture is divided into three key modules: a character creation module, an interaction choreography module, and a real-time animation rendering module. This modular design enables the system to achieve professional processing and efficient collaboration in each link while ensuring content consistency. The user interaction process begins in the character and creative selection stage. The user first selects multiple favorite characters from the system character library, and then decides to rewrite or continue the plot. In the rewrite function, the user can specify different styles, such as comedy, suspense, and science fiction, to reconstruct the existing plot; in the continuation function, the screenwriter agent will generate three different plot development directions, such as the protagonist's counterattack, enmity, and reconciliation. The user can choose one as the basis for creation based on his or her interests. This interactive design not only ensures the system's autonomous creative ability, but also provides key decision points for user participation, realizing the organic combination of intelligence and human creativity.
[0120] The character-based creation module, primarily composed of a screenwriter agent and a character agent, is responsible for the entire creative process from user input to script completion. The module's workflow can be divided into four core steps: First, the screenwriter agent creates a script outline based on the user-specified style / plot direction and character information. In this phase, the system introduces an innovative multi-stage novel recommendation framework, which includes three stages: coarse sorting, fine sorting, and language model-based re-ranking. This progressively refined recommendation mechanism ensures the quality and diversity of recommendation results while maintaining computational efficiency, accurately capturing user interests and providing personalized creative references. Subsequently, the director agent combines the recommended novel references, the preceding story, and the user-specified style or plot direction to generate a complete four-act short play outline. Each act can include specific subthemes and brief plot descriptions, ensuring the integrity and logical coherence of the story structure. The preceding story in this application refers to a user-entered summary of the preceding story, which can be in text form. The user can choose to create a sequel to the preceding story (i.e., continue the preceding story) or rewrite the preceding story (i.e., rewrite the preceding story in a different style). Both input and output can be in text form.
[0121] In the second step, the screenwriter agent expands the generated story outline into a complete dialogue. To enhance narrative fluency, the system automatically adds voice-over elements. This narration begins by explaining the background and character relationships, creating a natural connection for subsequent character dialogue. As the story progresses, the voice-over can explain complex situations that are difficult to convey through pure dialogue, facilitating plot transitions and developments. This narration design significantly enhances the narrative integrity and viewing experience of the short play.
[0122] The third step is intelligent matching of actions and voices. The screenwriter agent selects appropriate actions and emotional voices for each line of dialogue, such as sadness, anger, and joy, from a pre-established library of actions and voices. The system introduces a language model-based recommendation mechanism at this stage, using contextual analysis to deeply understand each line and quantitatively assess the matching of actions or voices with the dialogue. Taking into account the length of the entire script, the system can adopt a parallel strategy of splitting the content into acts and blocks: not only is each act processed in parallel, but the dialogue within each act is also divided into multiple blocks, with the same agent completing the conversion process from dialogue to action or voice for each block, significantly improving processing efficiency.
[0123] The fourth step is translating the character into a script. The system configures differentiated character agents based on the character's personality, occupation, and other attributes. Each agent focuses on analyzing and improving its own lines, emotional expression, and voice style, simulating the process of interpreting lines by a real actor. This refined character perspective ensures consistency and realism in the character's performance, making the generated dialogue more consistent with the character's characteristics.
[0124] The global choreography module and the interactive choreography module are composed of the director agent and the camera operator agent, and are responsible for the overall coordination and visual presentation design of the script interpretation. The workflow of this module can be divided into the following three main steps: First, the director agent assigns fixed scenes and character configurations to each scene in the script. At the same time, the system selects the most suitable background music for the plot from the music library. It is worth noting that the music library can adopt a dynamic update mechanism to continuously create music materials suitable for different themes such as suspense, sadness, passion, and love through the music generation model. Each time a script is generated, the system can use the model to create a new music clip and match it to the plot, and then store it in the music library to enrich the resource library. This link makes full use of the retrieval enhancement generation technology to achieve efficient and accurate multimedia resource retrieval and matching.
[0125] Secondly, the Director Agent plans each character's initial positioning and dynamic movement during the performance, based on its comprehensive interpretation of the script and scene layout. The Director Agent coordinates character interactions and plot progression from a global perspective, ensuring the overall performance's fluidity and dynamic flow. Simultaneously, performance suggestions from the Character Agents during the character workflow are simultaneously transmitted to the Director Agent, who integrates these suggestions and adjusts the script to ensure the final performance is more consistent with the character's personality traits. This division of labor between the Character Agent and the Director Agent has a clear theoretical basis: Character Agents tend to focus on their own parts from a local perspective, ignoring the overall coherence and logic of the script; the Director Agent, on the other hand, focuses on regulating the overall script performance, balancing each character's performance with the overall narrative needs. This functional separation allows each agent in the system to focus more closely on their respective responsibilities, improving overall collaborative efficiency.
[0126] Finally, a camera operator can delegate responsibility for automatically configuring various camera combinations. Based on the needs of the plot, the system intelligently switches between close-ups, mid-shots, and long shots, as well as dynamic and static camera movements, enriching the final animation and enhancing visual expression and audience immersion. The collaborative work of the character-focused and interactive modules ultimately generates a complete first draft of the script, laying the foundation for subsequent animation rendering.
[0127] During the animation presentation phase, the real-time animation rendering module preprocesses information based on the script draft within the real-time rendering engine. This includes creating 3D actor models based on actor identifiers, converting dialogue text into actor voices, and retrieving corresponding body and expression animation files based on action tags in the script. These files are then converted into the character's timbre. Based on a voice library, the text can be transformed into the voice of a specific character. Each character can be mapped one-to-one to an identifier. The actors referred to in this application are also characters, specifically three-dimensional (3D) characters. This preprocessing mechanism ensures that the script runs smoothly in a 3D presentation. The real-time rendering engine and code logic enable natural control of movement, expression, and dialogue timing. An animation state machine manages character behavior transitions, ensuring smooth and natural motion integration. The system utilizes GPU-accelerated computing, dynamic level of detail adjustment, resource preloading, and lighting baking to ensure that the script presentation can be rendered at a frame rate of over 60 frames per second on mobile devices. The combined application of these technologies enables high-quality real-time animation rendering on standard hardware.
[0128] Multi-agent collaborative dual-mode on-set interaction. This application is user-oriented, so it takes into account the user's immersive experience and interaction, emphasizing the user's interactivity. It can be used for users to better understand the generated script and for subsequent script adjustments. It can record the conversation between users and characters for the creation of subsequent plots. During the animation production waiting stage, the agent on-set interaction mode can be set to provide dual-path interaction of puzzle games and plot discussions. Users can collaborate with actor agents to decipher puzzles, and can also deeply explore character motivations and story development, significantly improving creative immersion and user retention.
[0129] Figure 4 This is a schematic diagram of an optional animation generation process for interacting with a user in gaming and chatting according to an embodiment of the present application. Figure 4As shown, intent recognition is performed based on user input. In game mode, agents (Actor 1 and the agent host) can speak and provide feedback based on historical conversations and information, and can also transfer status to subsequent processes. Similarly, agents (Actor 2 and the agent host) can speak and provide feedback based on historical conversations and information, and can also transfer status to subsequent processes. Agents (Actor 3 and the agent host) can speak and provide feedback based on historical conversations and information, and can also transfer status to subsequent processes. Users and agent hosts can also speak and provide feedback based on historical conversations and information. In chat mode, the agent classifier can combine user speech, retrieve actor features from agents (Actor 1, Agent 2, and Agent 3), and retrieve agent chat logs to select a speaker. The speaker can interact with actors 1, 2, and 3, speak on behalf of the agent, and save the agent chat logs.
[0130] The above-mentioned agent: actor one can be the first target agent model in this application, agent: actor two can be the second target agent model, agent: actor three can be the third role agent model, agent: host can be the third target agent model, and agent: classifier can be the classifier agent model.
[0131] The specific framework is as follows: In chat mode, a natural and smooth multi-role group chat interaction system is implemented based on the multi-agent orchestration framework. The classifier agent is used as the central orchestration component to dynamically coordinate the multi-role dialogue process. The multi-agent group chat interaction process is as follows: Classifier agent initialization: The system automatically designs a suitable chat topic and expected dialogue development path based on the script outline. The classifier agent analyzes the topic information and role characteristics and selects the initial speaking role suitable for starting the dialogue; the core function of the classifier agent, as the central orchestration component of the system, coordinates and manages the entire multi-role dialogue process; the speaker selection function selects the next most suitable speaking role after each round of dialogue; user participation control ensures that users have the opportunity to participate every 3-5 rounds of dialogue to maintain user experience; the message routing function accurately routes the current speech to the most appropriate role agent; dialogue state maintenance tracks and stores interaction history to ensure contextual coherence; the topic management function monitors the development of the dialogue to ensure that the dialogue revolves around the preset topic. The interaction process is implemented as follows: Request initiation: The speaker, i.e., the user or role agent, sends the conversation content to the classifier agent; Comprehensive analysis: The classifier agent uses a large language model to analyze the speech content, role information, and the complete conversation history; Agent selection: The classifier agent determines the next role suitable for response based on the analysis results; Request forwarding: The classifier agent forwards the previous speech content to the selected role agent; Response generation: The selected role agent retrieves the conversation history and generates a response that matches the role setting; History recording: The classifier agent saves the speech request and response to a storage with a specific session identifier; State update: After a round of conversation is completed, the global conversation state is updated so that agents can access the latest information. Global state management: The system assigns a unique session identifier to each user and script combination. The classifier agent maintains the complete conversation history under this session identifier. Role agents can retrieve the current conversation state through the classifier agent to ensure conversation continuity. In this architecture, the classifier agent also serves as a traditional orchestrator, acting as a central hub connecting users and multiple role agents, coordinating the operation of the entire multi-role dialogue system, ensuring natural and smooth interactions between roles while maintaining full user engagement.
[0132] In game mode, a structured, sequential multi-agent collaboration framework can be employed to achieve an efficient collaborative puzzle-solving mechanism. When a task arrives at the system, it follows a round-robin interaction process led by the host agent to create a coherent puzzle-solving experience. The multi-generational puzzle-solving interaction process is as follows: The host agent is initialized and automatically selects a puzzle at random from the puzzle library; the host agent, acting as the game progress controller, establishes the order in which players participate. In a sequential interaction mechanism, the host agent interacts with different role agents in a preset order. During each round of interaction, the current role agent automatically retrieves the conversation history under the session identifier. Based on historical conversations and accumulated information, the role agent proposes inference-worthy guesses about the puzzle. The role agent also combines preset characteristics such as its own personality and occupation to generate personalized expressions. Host evaluation and feedback: After receiving a character's guess, the host agent performs a multi-level analysis. Correctness testing determines whether the guess meets the answer criteria. If so, the answer is correct and the game ends. Relevance evaluation determines the relevance of the guess to the puzzle. If not, the irrelevant guess count is updated and the game ends. Directional guidance provides guiding hints to characters who repeatedly make irrelevant guesses, pointing them in the right direction. Accuracy feedback: For relevant guesses, a "yes" or "no" response is given based on accuracy. Global state updates: The system stores each round of character guesses and host feedback in a memory pool. After a round of interaction, a message is sent to the global environment to ensure the agent has access to the new dialogue state. Based on the updated global state, the next round of character interaction begins. This game mode adopts the Reasoning-Action (ReAct) paradigm. Upon receiving a task, each character agent first observes historical information, conducts internal reasoning, and then executes the guessed action. This entire process forms a complete Reasoning-Action cycle, ensuring a coherent and cumulative puzzle-solving process. The agent-on-set interactive system transforms the traditional single waiting process into an interesting interactive experience through a carefully designed multi-agent architecture, effectively improving user retention and creative participation, and providing an innovative user experience design solution for intelligent assisted creation tools.
[0133] Interactive script creation iteration, narrative extension mechanism, this application realizes a structured animation short play generation and extension mechanism based on the narrative principle of "introduction, development, turn and conclusion". The system can automatically generate an animation short play unit consisting of four episodes for the user. Each group of four short plays constitutes a complete narrative cycle, ensuring the integrity of the story structure and artistic expression. The algorithm ensures the internal consistency of each short play unit, and at the same time reserves a narrative extension interface in the structural design to provide a technical basis for subsequent content expansion. The system integrates an interactive plot extension function. After watching the initially generated animation short play, the user can modify or extend the existing content through the built-in rewrite or continue writing function module. This mechanism realizes theoretically unlimited plot extension, and users can continue to participate in the content creation process in an iterative manner.
[0134] The creative interaction mechanism breaks through the traditional end-to-end "black box" generation model and realizes transparent processing and modular interactive architecture for the entire creative process. The intermediate conversion results from user intention to narrative outline will be presented to the user in real time and support conversational modification. The screenwriter agent can dynamically adjust the outline content based on user feedback, forming a true human-computer collaborative creation environment. The system encapsulates core functions such as outline rewriting, outline continuation and outline modification into independent asynchronous interfaces, and builds a highly modular interactive architecture. This design enables users to intervene in real time at the early stages of the creative process, significantly reducing the cost of trial and error in creation, and avoiding the waste of resources caused by dissatisfaction after the entire animation rendering is completed.
[0135] This application innovatively adopts a text-based script storage strategy, persistently storing generated scripts in the form of structured text. Users can access historical content at any time and achieve dynamic visual presentation through an integrated real-time rendering engine. This technical architecture, which separates storage and rendering, exponentially reduces data storage requirements while maintaining the full editability of the content, providing a technical foundation for subsequent creative iterations and content reuse.
[0136] Character creation, this application system enhances the character creation capability through professional character, style, emotion, relationship and personality (CSERP) evaluation framework testing. Specifically, the system achieves good performance in the following dimensions: character characteristics (Character), accurately capturing and displaying the core traits and behavior patterns of the character; style consistency (Style), maintaining the consistent expression of the character in words, deeds and thinking; emotion expression (Emotion), achieving rich and reasonable emotion display through precise emotion modeling; relationship dynamics (Relationship), accurately simulating the interaction patterns and relationship development between characters; personality integrity (Personality), building a complete and coherent personality system to ensure the internal consistency of the character's behavior. This enhanced character creation capability stems from the following technical designs of the system: each role agent in the multi-agent architecture independently maintains the role status and characteristics; structured text representation allows for precise description and control of role characteristics.
[0137] This application builds an interactive animation generation system based on a multi-agent collaborative architecture. By combining retrieval enhancement with retrieval enhancement technology, it forms a complete closed loop from plot conception and resource mobilization to real-time rendering. Each agent in the system is dedicated to specific tasks, achieving a division of labor and collaboration model similar to that of a professional animation production team. This significantly improves the diversity and controllability of generated content, ensuring the stability and quality consistency of the animation content.
[0138] The knowledge-enhanced retrieval-enhanced generation mechanism, this application achieves a deep integration of multi-level material libraries and retrieval-enhanced generation technology, breaking through the knowledge limitations of content generation in related technologies. Based on creative needs, the system intelligently retrieves and integrates relevant professional knowledge, including narrative modes, character creation, scene design and other elements, significantly enhancing the professionalism and richness of the animation plot. This technological innovation enables the system to generate diverse content with professional quality, meeting creative needs at different levels and significantly improving user satisfaction.
[0139] Independent character agent and personality consistency maintenance mechanism. This application implements a character creation system based on independent character agents. Each character agent maintains its own personality traits, behavior patterns and emotional states, ensuring the consistency and coherence of the character performance throughout the animation creation process, so that the animated characters can show real, rich and internally consistent personality traits, significantly improving the artistic expression of the content and the audience's empathy.
[0140] The innovative on-set interaction model proposed in this application represents a breakthrough in resolving the waiting experience issues inherent in related intelligent creation systems. During resource-intensive computational processes like animation rendering, the system provides a dual-path interactive experience encompassing puzzle-solving and plot discussion. Users can deeply interact with animated character agents, enhancing both engagement and character development. This innovation transforms the entire creative process into a continuous, immersive experience, significantly improving user retention and boosting creative enthusiasm.
[0141] Rendering separation and text storage architecture. This application adopts a technical architecture that separates rendering and storage, storing creative content in the form of structured text, which can be rendered in real time only when needed through an integrated real-time rendering engine. This design not only significantly reduces storage requirements, but also achieves high editability and reusability of content. The system supports precise modification and incremental rendering based on text representation, avoiding the waste of resources caused by full regeneration, and providing users with an efficient and flexible creative experience.
[0142] To sum up, through the above-mentioned key technological innovations, this application has built an efficient, reliable and creative intelligent animation creation system, significantly expanded the application boundaries of intelligent animation generation, and provided a new technical paradigm for the field of digital content creation.
[0143] According to an embodiment of the present application, an animation generation method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order not used here.
[0144] Figure 5 is a flow chart of an animation generation method according to an embodiment of the present application, such as Figure 5 As shown, the following steps may be specifically included:
[0145] Step S502, responding to the input instruction on the operation interface, and displaying the script settings on the operation interface.
[0146] Step S504, in response to the processing instruction acting on the operation interface, the target animation is displayed on the operation interface, wherein the target animation is generated by inputting the script outline and role interpretation information into the third agent model, and the role interpretation information is generated by the third agent model, and the script outline is inputted into multiple second agent models respectively, and the role interpretation information of multiple roles is generated by multiple second agent models, and the script outline is inputted into the first agent model and generated by the first agent model, and the script outline contains information of multiple roles, the script setting is used to represent the story setting of the script, the script outline is used to represent the story framework of the script, and the role interpretation information is used to represent the interpretation parameters of the role in different scenarios.
[0147] According to an embodiment of the present application, an animation generation method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order not used here.
[0148] Figure 6 is a flow chart of an animation generation method according to an embodiment of the present application, such as Figure 6 As shown, the following steps may be specifically included:
[0149] Step S602: Acquire script settings by calling a first interface, wherein the first interface includes a first parameter, and a parameter value of the first parameter includes the script settings.
[0150] In step S604, the script setting is input into the first proxy model, and a script outline is generated using the first proxy model, wherein the script outline includes information of multiple characters, the script setting is used to represent the story setting of the script, and the script outline is used to represent the story framework of the script.
[0151] In step S606, the script outline is input into a plurality of second agent models respectively, and role interpretation information of a plurality of characters is generated using the plurality of second agent models, wherein the role interpretation information is used to represent interpretation parameters of the characters in different scenarios.
[0152] Step S608: Input the script outline and role interpretation information into the third proxy model, and use the third proxy model to generate the target animation.
[0153] Step S610: Output the target animation by calling the second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the target animation.
[0154] Optionally, Figure 7 This is a structural block diagram of a computing device according to an embodiment of the present application. Figure 7 As shown, the computing device A may include: one or more (only one is shown in the figure) processors 102, a memory 104 and a peripheral interface 106, wherein the processor 102, the memory 104 and the peripheral interface 106 are interconnected via a bus 108.
[0155] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0156] The processor may call the information and application programs stored in the memory through the transmission device to execute the steps in each embodiment.
[0157] An embodiment of the present application may provide an electronic device. Figure 8 This is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 8 As shown, the electronic device may include: an input / output device 802 ; a memory 804 ; and a processor 806 , wherein the processor 806 is connected to the input / output device 802 and the memory 804 via a bus 808 .
[0158] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0159] The processor can call the executable program stored in the memory through the transmission device to execute the following method: input the script setting into the first agent model, and use the first agent model to generate a script outline, wherein the script outline contains information about multiple characters, the script setting is used to represent the story setting of the script, and the script outline is used to represent the story framework of the script; input the script outline into multiple second agent models respectively, and use the multiple second agent models to generate role interpretation information of multiple characters, wherein the role interpretation information is used to represent the interpretation parameters of the characters in different scenarios; input the script outline and the role interpretation information into the third agent model, and use the third agent model to generate the target animation; and execute the methods in each embodiment of the present application.
[0160] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0161] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0162] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, or of course by hardware. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a number of instructions for causing a processing unit to execute the method in each embodiment of the present application.
[0163] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0164] The embodiment of the present application further provides a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the method provided in the above embodiment.
[0165] Optionally, in this embodiment, the above storage medium may be located in a computing device.
[0166] Optionally, in this embodiment, the computer-readable storage medium is configured to store an executable program, and when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the method described in any one of the above embodiments.
[0167] The embodiment of the present application further provides a computer program product. Optionally, in this embodiment, the computer program product may include a computer program, and when the computer program is executed by a processor, the method provided in the embodiment is implemented.
[0168] The above-mentioned computer program product may refer to a software program that has been written, tested and released, which can be run on a computer or other device. The computer program product may include an application, an operating system, tool software, etc., which is used to implement specific functions or solve specific problems.
[0169] The embodiments of the present application further provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium, which may be used to store a computer program that, when executed by a processor, implements the method provided in the embodiments above.
[0170] The above-mentioned non-volatile computer-readable storage medium may refer to a medium for storing data. The non-volatile computer-readable storage medium can keep the data from being lost when the power is off, and can be used to store long-term data, such as operating systems, applications and user files. The non-volatile storage medium may include hard disk drives, solid-state drives, optical disks and flash memory storage devices, etc.
[0171] The embodiment of the present application further provides a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, the method provided in the above embodiment is implemented.
[0172] The above-mentioned computer program may refer to a collection of instructions used to tell a computer to perform a specific task or operation. A computer program may be written by a programmer using a specific programming language and may include algorithms, data structures, logic, and control flows. Computer programs may be used for a variety of purposes, including application software, operating systems, and the like.
[0173] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0174] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0175] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0176] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0177] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0178] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. An animation generation method, characterized in that: include: Inputting a script setting into a first proxy model and generating a script outline using the first proxy model, wherein the script outline includes information about a plurality of characters, the script setting is used to represent the story setting of the script, and the script outline is used to represent the story framework of the script; Inputting the script outline into a plurality of second proxy models respectively, and using the plurality of second proxy models to generate role interpretation information of the plurality of characters, wherein the role interpretation information is used to represent interpretation parameters of the characters in different scenarios; The script outline and the role interpretation information are input into a third proxy model, and the target animation is generated using the third proxy model.
2. The animation generation method according to claim 1, characterized in that: Inputting the script outline and the character interpretation information into a third proxy model, and generating a target animation using the third proxy model, comprising: Inputting the script outline and the character interpretation information into a third proxy model, and using the third proxy model to generate script audio-visual information, wherein the script audio-visual information is used to plan the movement trajectories of the multiple characters; Inputting the script outline, the script audiovisual information, and the character interpretation information into a fourth proxy model, and generating camera movement information using the fourth proxy model; Inputting the script settings, the script outline, the script audio-visual information, the character interpretation information, and the camera movement information into the third proxy model, and generating a target script using the third proxy model, wherein the target script is used to represent a story plot description; The target animation is generated based on the target script.
3. The animation generation method according to claim 2, characterized in that: Generating the target animation based on the target script includes: Determining actor identification information and scene identification information corresponding to the multiple roles based on the target script; Retrieving the actor model corresponding to the actor identification information, and retrieving the scene model corresponding to the scene identification information; Performing timing control on the actor model and the scene model based on the target script to obtain an animation script; The animation script is rendered to obtain the target animation.
4. The animation generation method according to claim 1, wherein: Inputting the script settings into the first proxy model and generating a script outline using the first proxy model includes: Screening a plurality of preset materials based on the script setting to obtain at least one first material, wherein the at least one first material is a material among the plurality of preset materials whose matching degree with the script setting is greater than a first preset matching degree; Based on the script setting and the at least one first material, the at least one first material is screened to obtain at least one second material, wherein the at least one second material is a material in the at least one first material that has a matching degree with the script setting greater than a second preset matching degree; Inputting the at least one second material and the script setting into an adjustment model, and adjusting the at least one second material using the adjustment model to obtain at least one material; The script outline is generated based on the at least one material.
5. The animation generation method according to claim 4, characterized in that: Generating the script outline based on the at least one material includes: generating an initial script outline based on the at least one source material; Expanding the initial script outline to obtain script text information, wherein the script text information is used to represent plot information of the script; Determining character actions and character voices that match the script text information; Generating script role information based on the role action and the role voice, wherein the script role information is used to represent the role's interpretation information; Based on the script text information and the script role information, the script outline is obtained.
6. The animation generation method according to claim 5, characterized in that: Expanding the initial script outline to obtain the script text information includes: Expanding the initial script outline to obtain dialogue information and narration information; Identify the dialogue information to obtain narration requirements of the dialogue information; The narration information is added to the narration requirement point of the dialogue information to obtain the script text information.
7. The animation generation method according to claim 5, characterized in that: Determining character actions and character voices that match the script text information, including: Matching the script text information with a plurality of preset actions to obtain an action matching result; Matching the script text information with a plurality of preset voices to obtain a voice matching result; Performing a quantitative evaluation on the action matching result based on at least one dimension to obtain a first evaluation result, and performing a quantitative evaluation on the voice matching result based on the at least one dimension to obtain a second evaluation result; The character action is determined from a plurality of preset character actions based on the first evaluation result, and the character voice is determined from the plurality of preset character voices based on the second evaluation result.
8. The animation generation method according to claim 1, characterized in that: The method further comprises: Get story plot information; In response to receiving a style setting instruction for the story plot information, rewriting the story plot information based on the style setting instruction to obtain the script setting; In response to receiving a plot direction instruction for the story plot information, continuing the story plot information based on the plot direction instruction to obtain the script setting; In response to not receiving the style setting instruction and / or the plot direction instruction, determining that the story plot information is the script setting.
9. The animation generation method according to claim 8, characterized in that: Get story information, including: In response to receiving a role selection instruction from a user, determining at least one first target role corresponding to the role selection instruction from a plurality of preset roles; In response to receiving the user's style setting information and plot development information, the story plot information is generated based on the at least one first target character, the style setting information, and the plot direction information.
10. The animation generation method according to claim 2, characterized in that: The method further comprises: In the process of generating the target animation based on the target script, in response to receiving an interaction instruction, performing intent recognition on the interaction instruction to obtain an intent recognition result, and determining a target interaction mode corresponding to the interaction instruction from multiple interaction modes based on the intent recognition result; Interaction information between the user and at least one of the multiple characters is generated based on the target interaction mode.
11. The animation generation method according to claim 10, characterized in that: Generating interaction information between the user and at least one of the multiple characters based on the target interaction mode includes: In response to the target interaction mode being a chat interaction mode, obtaining a first speech content of the user, inputting the first speech content into a classifier proxy model, and analyzing the first speech content using the classifier proxy model to obtain a first role associated with the first speech content; inputting the first speech content into a first target agent model of the first character, and generating a second speech content using the first target agent model; The interactive information is generated based on the first speech content and the second speech content.
12. The animation generation method according to claim 10, characterized in that: Generating interaction information between the user and at least one of the multiple characters based on the target interaction mode includes: In response to the target interaction mode being the game interaction mode, determining a second target agent model of the second character from the plurality of second agent models based on a game task; determining a game participation order between the user and the second character; Controlling the user and the second character to perform the game task in a first interaction round based on the game participation order, and obtaining an execution result, wherein the execution result is used to indicate whether the game task is successfully performed; In response to the execution result being that the game task is not successfully executed, controlling the user and the second character to execute the game task in multiple interaction rounds based on the game participation order until the execution result is that the game task is successfully executed, and generating the interaction information based on the interaction information of the first interaction round and the multiple interaction rounds; In response to the execution result being successful execution of the game task, the interaction information is generated based on the interaction information of the first interaction round.
13. The animation generation method according to claim 12, characterized in that: The method further comprises: Obtaining feedback information on the user and the second character performing the game task; Inputting the feedback information into a third target proxy model, and detecting the feedback information using the third target proxy model to obtain a detection result, wherein the detection result is used to indicate whether the feedback information is preset information; In response to the feedback information being the preset information, determining that the execution result is successful execution of the game task; In response to the feedback information not being the preset information, determining the execution result as unsuccessful execution of the game task.
14. An animation generation method, characterized in that: include: In response to an input command applied to an operation interface, displaying a script setting on the operation interface; In response to a processing instruction acting on the operation interface, a target animation is displayed on the operation interface, wherein the target animation is generated by inputting the script outline and role interpretation information into a third agent model, and using the third agent model; the role interpretation information is generated by inputting the script outline into multiple second agent models respectively, and using the multiple second agent models to generate role interpretation information for multiple roles; the script outline is generated by inputting the script setting into a first agent model, and using the first agent model; the script outline contains information on multiple roles; the script setting is used to represent the story plot setting of the script; the script outline is used to represent the story plot framework of the script; and the role interpretation information is used to represent the interpretation parameters of the role in different scenarios.
15. An animation generation method, characterized in that: include: Acquire script settings by calling a first interface, wherein the first interface includes a first parameter, and a parameter value of the first parameter includes the script settings; Inputting a script setting into a first proxy model and generating a script outline using the first proxy model, wherein the script outline includes information about a plurality of characters, the script setting is used to represent the story setting of the script, and the script outline is used to represent the story framework of the script; Inputting the script outline into a plurality of second proxy models respectively, and using the plurality of second proxy models to generate role interpretation information of the plurality of characters, wherein the role interpretation information is used to represent interpretation parameters of the characters in different scenarios; Inputting the script outline and the character interpretation information into a third proxy model, and generating a target animation using the third proxy model; The target animation is output by calling a second interface, wherein the second interface includes a second parameter, and a parameter value of the second parameter includes the target animation.
16. A computing device, characterized in that include: a memory storing an executable program; A processor, configured to run the program, wherein the program, when running, executes the method according to any one of claims 1 to 15.
17. An electronic device, characterized in that: include: a memory storing an executable program; A processor, connected to the memory via a bus, and configured to run the program, wherein the program executes the method according to any one of claims 1 to 15 when running.
18. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 15.
19. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 15.