System, method and equipment for generating animation and video and storage medium
By designing an end-to-end animation generation system, using multimodal interaction and semantic mapping technology, the cumbersome and cost-effective problems of traditional animation production are solved, efficient, high-quality and customizable animation creation is achieved, and the realism and user interactivity of the animation are enhanced.
Patent Information
- Application Number
- CN202510328774.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-24
AI Technical Summary
Traditional animation production process is cumbersome, costly and complex technical requirements. It is difficult for existing automation tools to accurately analyze user needs, dynamic mapping is lacking, and the sense of reality and flexibility are limited.
Design an end-to-end animation generation system, and integrate instruction understanding, dynamic database mapping and interactive editing through multimodal interaction, semantic mapping and intelligent content generation technologies to achieve efficient, high-quality and customizable animation creation.
Through intelligent mapping, the matching degree between user needs and generated content is improved, and the high-realistic animation effect is achieved, and the user can fully controllable editing from global parameters to frame-by-frame details is supported, and it is adapted to a variety of output formats and application scenarios.
Smart Images

Figure FT_1 
Figure FT_2
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and computer graphics, and specifically to a system, method, device, and storage medium for generating animations and videos. Through multi-modal interaction, semantic mapping, and intelligent content generation technologies, it realizes the automated and semi-automated creation of high-quality animations from user instructions. Background Art
[0002] Traditional animation production relies on manual modeling, keyframe drawing, and post-production special effects synthesis, which has problems such as a cumbersome process, high cost, and complex technical requirements. In the prior art, although some automated tools can generate simple animations through template combinations, they have the following defects: Insufficient instruction understanding: It is difficult to accurately parse user requirements; Lack of dynamic mapping: The pre-defined templates have poor adaptability to user requirements and lack intelligent association rules; Limited realism and flexibility: Physical simulation is rough, and users cannot edit the generated content in real time.
[0003] Therefore, there is an urgent need for an intelligent animation generation system that integrates instruction understanding, dynamic database mapping, and interactive editing. Summary of the Invention
[0004] The present invention aims to provide an end-to-end animation generation system that solves the coordination problems of instruction parsing, dynamic content generation, and user interactive editing through artificial intelligence technology, and realizes efficient, high-quality, and customizable animation creation. Technical Solution
[0005] This system includes a user input module, a multimedia database, a data processing engine, a content generation module, an interactive editing module, and an output module. Specifically, it can be divided into the following parts:
[0006] User input module: Supports multi-modal input of text, voice, sketches, and images, and extracts semantic features through natural language processing and computer vision algorithms.
[0007] Multimedia database: Stores a three-dimensional model library, an animation template library, and a physical property data set.
[0008] Information processing unit: Parses the scene, character, and action keywords in the user instructions.
[0009] Dynamic mapping unit: Based on feature matching algorithms and deep learning models, associates user intentions with database resources.
[0010] Content generation module: Synthesizes a basic animation frame sequence.
[0011] Temporal prediction model: Ensures the coherence of inter-frame motion.
[0012] Physics Engine: Simulate lighting, rigid / soft body dynamics, and particle effects.
[0013] Interactive Editing Module: Provide functions for parameter adjustment, keyframe modification, and effect overlay, and support users to make real-time corrections.
[0014] Output Module: Export video files and 3D engineering files.
[0015] The beneficial effects of the present invention are as follows:
[0016] Intelligent Mapping: Through semantic parsing and dynamic rule combination, improve the matching degree between user requirements and generated content.
[0017] High-quality Generation: Combine the generation module with the physics engine to achieve high-fidelity animation effects.
[0018] Flexible Interaction: Support users to perform fully controllable editing from global parameters to frame-by-frame details.
[0019] Multi-scenario Adaptation: The output format covers film and television, game development, and AR / VR applications, and is compatible with mainstream software and hardware platforms. Brief Description of the Drawings
[0020] Figure 1 : The overall system architecture diagram, showing the data flow and interaction logic between modules.
[0021] Figure 2 : The flowchart for establishing dynamic mapping relationships, including steps of semantic parsing, feature matching, and rule optimization. Detailed Description of the Embodiments
[0022] Embodiment 1: The user inputs the instruction "A cat jumps off the roof and rolls after landing" through text.
[0023] Step 1: Multimodal input and semantic parsing.
[0024] Information Processing Unit: Use the NLP model for word segmentation and entity recognition, and extract keywords: subject object, scene, action.
[0025] Analyze the action logic relationship in combination with the context, and construct a scene time sequence chain.
[0026] Step 2: Dynamic resource matching.
[0027] Multimedia Database Call: According to "cat", match the high-precision feline skeletal model in the 3D model library.
[0028] According to "jump", retrieve the basic action sequence in the animation template library, and associate the gravity parameter and ground collision response coefficient in the physical property dataset.
[0029] Call the soft body dynamics template according to "rolling", and set the threshold of the torso bending angle and the joint torque limit.
[0030] Step 3: Content generation and physical simulation.
[0031] Content generation module: Synthesize the basic animation frame sequence, and initially generate the animation of the cat jumping from the edge of the roof.
[0032] The timing prediction model optimizes the frame - to - frame transition to ensure the action coherence from jumping to landing.
[0033] Physics engine: Simulate the soft body deformation at the moment of landing, and generate the special effect of flying dust based on the particle system.
[0034] Calculate the swing amplitude of the cat's tail in real - time during the rolling process to avoid penetration with the ground.
[0035] Step 4: Interactive editing and output.
[0036] User editing operation: Adjust the jump height through the slider, and the system automatically corrects the hang - time and landing impact force parameters.
[0037] Insert key frames in the timeline, reduce the rolling speed by 50%, and add slow - motion special effects.
[0038] Output module: Export a 4K resolution video file, including multi - track layered data.
[0039] Generate a 3D engineering file in FBX format, retaining the bone binding and material node information, which can be directly imported into Maya or Blender for post - rendering.
[0040] Through semantic parsing and dynamic resource matching, the time for converting user instructions into high - fidelity animations is shortened. The soft body deformation and particle special effects simulated by the physics engine significantly enhance the realism. Interactive editing supports users to quickly correct the action logic and lower the threshold of professional animation production.
[0041] Example 2: The user selects a building and inputs the voice command "Generate a collapse animation".
[0042] Step 1: Multimodal input fusion.
[0043] User input: Synchronize the voice command "Generate a collapse animation with explosion effect".
[0044] Input module processing: The computer vision algorithm identifies the key structures in the building.
[0045] The voice recognition model parses the instruction into semantic tags: subject, action, special effect.
[0046] Step 2: Dynamic mapping and resource adaptation.
[0047] Multimedia database call: Match the "reinforced concrete structure" template in the building rigid body model library.
[0048] Associate with the "central point initiation" mode in the blasting animation template.
[0049] Load the brittle fracture parameters in the physical property dataset.
[0050] Step 3: Content generation and physical simulation.
[0051] Content generation module: Generate a 3D building model based on the sketch structure and divide the rigid body mesh. Synthesize the blasting initiation point and generate the initial collapse direction.
[0052] Physical engine: Simulate the process of layer-by-layer collapse.
[0053] Smoke special effect generation: Calculate the smoke particle density in real time according to the collapse volume and simulate the airflow diffusion path.
[0054] Step 4: Interactive correction and output.
[0055] User editing operation: Drag the key frames in the timeline to adjust the collapse direction to a 30° left tilt, and the system automatically recalculates the mechanical balance.
[0056] Output module: Export the Unity programmable script, including rigid body collision parameters, particle system call interfaces, and camera control logic.
[0057] Generate a lightweight GLB format 3D scene file, supporting real-time rendering on the Web side.
[0058] Technical effects: Achieve rapid construction of complex scenes by non-professional users through cross-modal fusion. The high-precision simulation of the physical engine ensures that the collapse process conforms to the laws of engineering mechanics. The programmable script output is directly compatible with game engines, reducing the workload of secondary development.
[0059] The embodiments of the present invention are committed to achieving the deep integration of art and artificial intelligence technologies, and strive to resolve problems such as resource scarcity and high generation costs in the digital industry. The present invention promotes commercial exchanges and resource sharing among multiple industries by breaking down the barriers between industries, and promotes the complementary development of cultural innovation and technological development.
[0060] Even though the preferred embodiments of the present invention have been described, once those skilled in the art have grasped the basic inventive concept, they will be able to make other changes and modifications to these embodiments. Therefore, the appended claims are intended to cover the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Therefore, if these modifications and variations of the present invention are within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
[0061] Finally, it should be noted that the above content is only the preferred embodiment of the present invention and is not used to limit the present invention. Although the present invention has been described in detail according to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements on some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle scope of the present invention shall be included in the protection scope of the present invention.
Claims
1. A system for generating animation and video, characterized in that: Includes the following modules: A user input module, used to receive text, images, UI interactions or voice commands input by users; Multimedia database, which stores 3D models, dynamic animation templates, physical property data and pre-trained mapping rules; The data processing engine includes an instruction processing unit and a feature matching unit, which is used to parse the semantic features of user input and establish a dynamic mapping relationship with the database; The content generation module combines the timing prediction model and the physics engine to generate the initial animation sequence based on the mapping relationship; Interactive editing module, which provides parameter adjustment interface, keyframe editing tools and special effects overlay function, allowing users to modify the generated animation in real time; The output module supports exporting the final animation as a standard video file, programmable script or 3D project file.
2. The system according to claim 1, characterized in that The user input module includes a multimodal interface, supporting at least one input method of text description, UI interaction, sketch drawing, reference image uploading and voice command.
3. The system according to claim 1, characterized in that The multimedia database further comprises: 3D model library based on semantic label classification; Animation template library including motion trajectory and deformation logic; A dataset of physical properties of object materials, gravity, and collision response.
4. The system according to claim 1, characterized in that The establishment of the dynamic mapping relationship is achieved by the following steps: Analyze user input to extract keywords, scene descriptions, and action instructions; Based on the feature matching algorithm, the semantic elements are associated with the model attributes and animation templates in the database; The mapping weights are optimized through deep learning models to generate rule combinations that adapt to user intentions.
5. The system according to claim 1, characterized in that The content generation module comprises: A generator for synthesizing a sequence of basic animation frames based on input features; Temporal convolutional neural network, predicting the coherent motion logic between animation frames; Physics engine that simulates lighting, material reflection, and rigid / soft body dynamics effects.
6. The system according to claim 1, characterized in that The interactive editing module provides: Global parameter adjustment functions based on sliders and buttons, including animation speed, camera angle and light intensity; Draggable keyframe timeline supports frame-by-frame correction of skeletal movements and object deformations; Special effects overlay interface, supports dynamic insertion of particle effects, depth of field blur and environment maps.
7. The system according to claim 1, characterized in that The output module further supports: Render the animation into 4K resolution video and encapsulate it into MP4 and MOV formats; Generate 3D engineering files containing bone binding information, compatible with Blender and Maya software; Output programmable animation scripts and support secondary development calls through API interfaces.
8. An animation generation method, characterized in that: The animation generation method implements the steps of the animation generation system according to any one of claims 1 to 7.
9. An electronic device, characterized in that: The invention comprises a processor, a memory and a computer program stored in the memory and capable of running on the processor, wherein the computer program implements the steps of the animation generation system according to any one of claims 1 to 7 when executed by the processor.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the animation generation system according to any one of claims 1 to 7 are implemented.