Interactive three-dimensional animation generation method and system based on AIGC technology

Through cross-modal understanding and structured scene decoding, users can conduct detailed interactions between animation sketches and natural language descriptions, solving the problem of difficult adjustment of animation generation details in existing technologies, achieving controllability and flexibility in 3D animation generation, and improving the creative experience and efficiency.

CN120672918AInactive Publication Date: 2025-09-19GUANGZHOU HUAXIA VOCATIONAL COLLEGE

Patent Information

Application Number
CN202510765894.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing interactive 3D animation generation solutions based on AIGC technology make it difficult for users to flexibly customize detailed elements such as animation scenes, character actions, and camera movements. They also lack effective understanding and real-time feedback of users' deep creative intentions, resulting in a large deviation between the generated animation effects and the users' inner ideas, affecting the creative experience and efficiency.

Method used

By performing cross-modal joint understanding of the animation sketch and natural language description input by the user, a low-fidelity preview of the 3D animation is generated. A structured scene decoding mechanism is introduced to decompose the animation intent into detailed interaction modules. Users can issue specific instructions to adjust these modules, and the system updates the structured scene description to generate an updated 3D animation preview.

Benefits of technology

The controllability and flexibility of animation generation are enhanced, allowing users to make targeted and refined adjustments, solving the problem of difficult detail adjustments in traditional AIGC solutions and improving the creative experience and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672918A_ABST
    Figure CN120672918A_ABST
Patent Text Reader

Abstract

The invention discloses an interactive three-dimensional animation generation method and system based on an AI GC technology, and the method comprises the steps: carrying out the cross-modal joint understanding of an animation sketch and natural language description inputted by a user, and generating a preliminary three-dimensional animation low-fidelity preview based on the cross-modal joint understanding; then, by introducing a structured scene decoding mechanism, analyzing the complex animation intention multi-mode joint feature into a structured scene description containing a series of refined interaction modules, so as to decompose the macroscopic animation intention into local units capable of being accurately intervened by a user; a user can send a specific refining instruction for the specific refining interaction modules, the system updates the structured scene description after receiving the instruction, and generates an updated three-dimensional animation preview based on the updated structured scene description. Thus, the user can carry out targeted and refined adjustment on the animation, the controllability and flexibility of animation generation are greatly enhanced, and the problem that details are difficult to adjust in a traditional AI GC scheme is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of three-dimensional animation generation, and more specifically, to an interactive three-dimensional animation generation method and system based on AIGC technology. Background Art

[0002] With the vigorous development of the digital content industry, 3D animation has been widely used in film and television, games, advertising, education and other fields due to its vivid and realistic visual effects and rich expressiveness. However, the traditional 3D animation production process usually relies on professional modeling, binding, animation, rendering and other software, which requires the operator to have high professional skills, and the production cycle is long and the cost is high, which greatly limits the popularization and personalized creation of 3D animation. In recent years, the rapid progress of artificial intelligence generated content (AIGC) technology, especially its breakthroughs in image and video generation, has brought new opportunities for the automated and intelligent generation of 3D animation. Therefore, constructing an interactive 3D animation generation solution that can lower the creation threshold, improve generation efficiency, and allow users to easily express and adjust their creative intentions is of great significance to promoting the development and application of 3D animation technology, and is currently a hot topic of concern in the industry and academia.

[0003] In existing technical practices, some solutions have emerged that attempt to utilize AIGC technology for animation generation. Some solutions focus on directly generating video clips from text or images, but these methods often face problems such as insufficient controllability, difficulty in precisely adjusting details, and lack of effective interaction when generating 3D animations with complex scenes, coherent movements, and specific styles. Specifically in the field of interactive 3D animation generation, although existing solutions provide certain user input interfaces, such as allowing users to influence the generation results through simple parameter adjustments or selecting preset templates, these interaction methods are often coarse-grained and cannot meet users' needs for flexible customization of detailed elements such as animation scenes, character movements, and camera movements. More importantly, these solutions often lack effective understanding of the user's deep creative intent and real-time feedback mechanisms. As a result, after multiple attempts, the generated animation effect may still deviate significantly from the user's inner conception. The existence of such deviations not only seriously affects the user's creative experience but also significantly reduces creative efficiency.

[0004] Therefore, an optimized interactive 3D animation generation method based on AIGC technology is expected. Summary of the Invention

[0005] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides an interactive three-dimensional animation generation method and system based on AIGC technology, which performs cross-modal joint understanding of the animation sketch and natural language description input by the user, and generates a preliminary three-dimensional animation low-fidelity preview image based on this; then, by introducing a structured scene decoding mechanism, the complex multi-modal joint features of the animation intention are parsed into a structured scene description containing a series of refined interaction modules, so as to decompose the macro animation intention into local units that can be precisely intervened by the user. The user can issue specific refinement instructions for these specific refined interaction modules. After receiving the instructions, the system updates the structured scene description and generates an updated three-dimensional animation preview image based on the updated structured scene description. In this way, the user can make targeted and refined adjustments to the animation, which greatly enhances the controllability and flexibility of animation generation and effectively solves the problem of difficult detail adjustment in the traditional AIGC solution.

[0006] According to one aspect of the present application, a method for generating an interactive 3D animation based on AIGC technology is provided, which includes:

[0007] Obtaining an animation sketch and a natural language description of the animation effect input by a user;

[0008] Input the animation sketch and natural language description of the animation effect into the AIGC module to obtain a low-fidelity preview of the 3D animation;

[0009] Performing structured scene decoding on the low-fidelity preview image of the three-dimensional animation to obtain a structured scene description, wherein the structured scene description includes a list of detailed interaction modules;

[0010] Obtaining a refinement instruction input by a user for a specific refinement interaction module in the list of refinement interaction modules to obtain an updated structured scene description;

[0011] The updated structured scene description is input into the AIGC module to obtain the updated 3D animation preview.

[0012] According to another aspect of the present application, an interactive 3D animation generation system based on AIGC technology is provided, which includes:

[0013] A user intention acquisition module is used to obtain the animation sketch and natural language description of the animation effect input by the user;

[0014] The 3D animation preliminary generation module is used to input the animation sketch and natural language description of the animation effect into the AIGC module to obtain a low-fidelity preview image of the 3D animation;

[0015] A structured scene decoding module is used to perform structured scene decoding on the 3D animation low-fidelity preview image to obtain a structured scene description, wherein the structured scene description includes a list of detailed interaction modules;

[0016] A scene description updating module, configured to obtain a refinement instruction input by a user for a specific refinement interaction module in the list of refinement interaction modules to obtain an updated structured scene description;

[0017] The 3D animation update module is used to input the updated structured scene description into the AIGC module to obtain a 3D animation update preview image.

[0018] Compared with the existing technology, the present application provides an interactive 3D animation generation method and system based on AIGC technology. It performs cross-modal joint understanding of the animation sketch and natural language description input by the user, and generates a preliminary low-fidelity preview image of the 3D animation based on this. Then, by introducing a structured scene decoding mechanism, the complex multimodal joint features of the animation intention are parsed into a structured scene description containing a series of refined interaction modules, so as to decompose the macro animation intention into local units that can be precisely intervened by the user. The user can issue specific detailed instructions for these specific refined interaction modules. After receiving the instructions, the system updates the structured scene description and generates an updated 3D animation preview image based on the updated structured scene description. In this way, the user can make targeted and refined adjustments to the animation, greatly enhancing the controllability and flexibility of animation generation, and effectively solving the problem of difficult detail adjustment in traditional AIGC solutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0020] Figure 1 Flowchart of an interactive 3D animation generation method based on AIGC technology according to an embodiment of the present application;

[0021] Figure 2 A data flow diagram of an interactive 3D animation generation method based on AIGC technology according to an embodiment of the present application;

[0022] Figure 3 Flowchart of sub-step S2 of the interactive 3D animation generation method based on AIGC technology according to an embodiment of the present application;

[0023] Figure 43D animation generation system based on AIGC technology according to an embodiment of the present application. DETAILED DESCRIPTION

[0024] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.

[0025] As used in this application and the claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.

[0026] Although the present application makes various references to certain modules in the system according to embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are illustrative only, and different aspects of the system and method can use different modules.

[0027] Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the various steps may be processed in reverse order or simultaneously, as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.

[0028] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.

[0029] In the technical solution of the present application, an interactive 3D animation generation method based on AIGC technology is proposed. Figure 1 Flowchart of an interactive 3D animation generation method based on AIGC technology according to an embodiment of the present application. Figure 2 Schematic diagram of data flow of the interactive 3D animation generation method based on AIGC technology according to an embodiment of the present application. Figure 1 and Figure 2As shown, according to an embodiment of the present application, an interactive three-dimensional animation generation method based on AIGC technology includes the following steps: S1, obtaining an animation sketch and a natural language description of an animation effect input by a user; S2, inputting the animation sketch and the natural language description of the animation effect into the AIGC module to obtain a low-fidelity preview image of the three-dimensional animation; S3, performing structured scene decoding on the low-fidelity preview image of the three-dimensional animation to obtain a structured scene description, wherein the structured scene description includes a list of refined interaction modules; S4, obtaining a refinement instruction for a specific refined interaction module in the list of refined interaction modules input by the user to obtain an updated structured scene description; S5, inputting the updated structured scene description into the AIGC module to obtain an updated preview image of the three-dimensional animation.

[0030] In particular, the S1 obtains the animation sketch and natural language description of the animation effect input by the user. Among them, the animation sketch, as an intuitive visual input, can help the user quickly outline the general structure of the scene, the main outline and spatial relationship of the objects, and the preliminary idea of ​​the animation keyframes; the natural language description can supplement the dynamic information, emotional color, style preferences and more abstract animation effect requirements that are difficult to describe in the sketch. In the technical solution of the present application, by obtaining the animation sketch input by the user and the natural language description of the animation effect input by the user, the user can obtain the most original and direct creative inspiration and core needs, and provide sufficient and multi-dimensional guidance information for the subsequent AIGC module to generate the initial three-dimensional animation low-fidelity preview image.

[0031] In practice, the system provides an interactive interface consisting of a drawing area and a text input box. First, users can use a mouse, touchpad, or drawing tablet to draw an animation sketch in the drawing area. This sketch doesn't need to be very detailed, just expressing key visual elements and spatial layout. Next, users enter a natural language description of the animation effect in the text input box. Once the user completes the sketch and text description, they click "Confirm," and the system completes the acquisition of both user inputs.

[0032] In particular, the S2 inputs the animation sketch and the natural language description of the animation effect into the AIGC module to obtain a low-fidelity preview of the three-dimensional animation. Figure 3 As shown, the S2 includes: S21, respectively parsing the natural language description of the animation sketch and the animation effect to obtain the animation sketch visual features and the animation effect semantic features; S22, performing cross-modal joint understanding of the animation intent on the animation effect semantic features and the animation sketch visual features to obtain the animation effect intention multimodal joint features; S23, inputting the animation effect intention multimodal joint features into the AIGC module to obtain a low-fidelity preview image of the three-dimensional animation.

[0033] Specifically, the S21 parses the natural language description of the animation sketch and the animation effect respectively to obtain the visual features of the animation sketch and the semantic features of the animation effect. In an embodiment of the present application, first, the natural language description of the animation effect is parsed using the NLU module to obtain the semantic feature vector of the animation effect as the semantic feature of the animation effect. Since the complex requirements conveyed by the user through text descriptions are often difficult to be effectively understood by the system, the NLU module, as a technical component specifically for processing semantic understanding, has a working principle based on the fusion of deep language models and domain knowledge. Through multi-level processing such as lexical analysis, syntactic parsing, semantic role labeling and contextual reasoning, discrete text input is converted into a computable semantic representation. Therefore, in order to establish a semantic bridge between the user's language description and the animation generation logic, in the technical solution of the present application, the natural language description of the animation effect is parsed using the NLU module to obtain the semantic feature vector of the animation effect as the semantic feature of the animation effect.

[0034] Specifically, the NLU module uses pre-trained language models (such as BERT, GPT, etc.) to perform context-aware encoding of the input text, and combines entity recognition and action logic analysis in the animation field to extract feature vectors containing key semantic information such as scene elements, character behaviors, and time series relationships, forming animation effect semantic features with a clear dimensional structure. In this process, the NLU module not only extracts key words in explicit instructions through deep semantic modeling, but also captures users' deep demands for animation style (such as cartooning, realism), movement smoothness (such as slow-in and slow-out effects), and physical properties (such as gravity coefficients) through intent recognition and implicit demand reasoning. The generation of this structured semantic feature vector enables the subsequent cross-modal joint understanding stage to accurately align the language description with the visual elements of the animation sketch, avoiding generation deviations caused by semantic gaps.

[0035] Furthermore, the CV module is used to parse the animation sketch to obtain an animation sketch visual feature map as the animation sketch visual feature. It should be understood that the animation sketch is an important carrier for users to express their creative intentions, and contains key visual information such as character posture, scene layout, and motion trajectory, but this information exists in the form of unstructured pixels and cannot be directly understood and processed by the generation system. As an algorithm component specifically for processing visual data, the CV module works based on the hierarchical extraction capability of the deep learning model for image features. It uses a convolutional neural network architecture to encode spatial features of the input sketch and converts pixel-level information into a visual feature map with spatial perception. Therefore, in order to achieve semantic alignment between user sketches and three-dimensional animation generation, in the technical solution of the present application, the CV module is used to parse the animation sketch to obtain an animation sketch visual feature map as the animation sketch visual feature.

[0036] Specifically, the CV module first performs standardized preprocessing on the sketch (such as size normalization and noise suppression), and then gradually extracts low-order to high-order visual features such as line contours, shape structures, and spatial relationships through multi-layer convolution operations, and finally forms a visual feature map that can characterize the overall layout and local details of the sketch. In this process, the CV module uses a feature extraction network to perform structured analysis of these implicit visual clues, and converts the hand-drawn two-dimensional plane information into a visual feature map containing parameters such as spatial position, shape topology, and motion trends. The generation of this feature map not only retains the original composition information of the sketch, but also captures implicit associations that the user may not have explicitly drawn (such as the interactive relationship between the character and the scene) through the adaptive learning ability of the deep neural network, providing a computable visual representation basis for subsequent cross-modal joint understanding.

[0037] Specifically, the S22 performs a cross-modal joint understanding of animation intent on the animation effect semantic features and the animation sketch visual features to obtain a multi-modal joint feature of the animation effect intent. It should be understood that the language description may emphasize the temporal relationship of the action logic (such as "the character rolls after jumping from a high place"), while the sketch implies the mechanical balance of the character posture or the scene perspective relationship through the spatial layout. If the system relies solely on semantic or visual modality for generation, the system may ignore the key creative intent due to the fragmentation of modal information, resulting in a contradiction between the action picture and the character description in the generated animation. Therefore, in order to construct a collaborative mapping network of semantic and visual information, and thus explore the deep creative logic implicit in cross-modal data, in the technical solution of the present application, the animation effect semantic features and the animation sketch visual features are subjected to a cross-modal joint understanding of animation intent to obtain a multi-modal joint feature of the animation effect intent.

[0038] In this process, first, the abstract concepts described in the language (such as "tense atmosphere") are semantically aligned with the specific visual elements of the sketch (such as the character's taut muscle lines and the low-angle scene composition), and the correlation between the two in the spatiotemporal dimensions is analyzed; then, by introducing the feature-like magnetic attraction effect, the system can automatically screen out local features of the sketch that are highly relevant to the language description (such as identifying the curved lines of the legs in the sketch that suggest a "leap" action) and suppress irrelevant or interfering visual information (such as decorative elements in the background that are not related to the action); then, the spatiotemporal modeling capabilities of the LSTM-Transformer hybrid architecture are combined to capture short-term local correlations between language and visual features (such as posture matching in specific action frames), establish cross-modal global contextual relationships (such as the coordination between the overall animation rhythm and the scene layout), and finally form a multimodal joint encoding vector of animation effect intention that integrates multi-scale information. The multimodal joint features of animation effect intentions generated through cross-modal joint understanding not only accurately encode the key elements of the user's explicit expression (such as the basic action sequence of the character), but also capture the implicit creative needs through cross-modal reasoning (such as inferring the expected light and shadow effect intensity through the shadow density of the sketch), significantly improving the accuracy of user intention restoration.

[0039] Specifically, first, the visual feature map of the animation sketch is feature decoupled to obtain a set of local visual feature encoding vectors of the animation sketch. It should be understood that visual features such as the density of lines and the distortion of shapes in hand-drawn sketches often imply the user's potential settings for animation parameters such as the character's motion amplitude and the lens motion trajectory, but this information is mixed in an unstructured manner in the original feature map. The set of local feature encoding vectors formed by feature decoupling enables each visual semantic unit of the sketch (such as the swing angle of the character's arm and the perspective relationship of the background building) to be independently encoded as an operable mathematical representation, providing a granularity-controllable adjustment unit for subsequent cross-modal interaction. This discretization process not only retains the topological relationship of the overall layout of the sketch, but also realizes the construction of a semantic bridge from pixel-level visual features to animation parameters. The feature imitation magnetic effect value between the animation effect semantic feature vector and each animation sketch local visual feature encoding vector in the set of animation sketch local visual feature encoding vectors is calculated to obtain a set of animation effect visual feature imitation magnetic effect values.

[0040] In a specific example of the present application, the following formula is used to perform feature decoupling on the animation sketch visual feature map to obtain a set of local visual feature encoding vectors of the animation sketch; wherein the formula is:

[0041] Partition(F 2 )=V={v1,v2,...,v i ,...,v n}

[0042]

[0043] V={v1,v2,...,v i ,...,v n}

[0044] Among them, Partition(·) represents feature decoupling operation, Flatten(·) represents flattening processing, φ(·) represents a one-dimensional convolutional layer or a fully connected layer, [h i ,w i ] represents the block position index, V is the set of local visual feature encoding vectors of the animation sketch, v1, v2, v i ,v n They are respectively the 1st, 2nd, i-th and n-th animation sketch local visual feature coding vectors in the set of the animation sketch local visual feature coding vectors.

[0045] Next, the feature imitation magnetic effect value between the animation effect semantic feature vector and each animation sketch local visual feature coding vector in the set of animation sketch local visual feature coding vectors is calculated to obtain a set of animation effect visual feature imitation magnetic effect values. It should be understood that there is a complex semantic mapping relationship between the animation effect described by the user in natural language (such as "the character glides with an elegant posture") and the scattered visual elements in the hand-drawn sketch (such as limb curves, ground perspective auxiliary lines). It is difficult for traditional methods to automatically identify which sketch local features truly carry the core intention of the language description. Therefore, in order to construct a dynamic screening mechanism for cross-modal semantic alignment, in the technical solution of the present application, the correlation strength between the semantic feature and each visual local feature is quantified by calculating the feature imitation magnetic effect value, and a priority sorting mechanism for cross-modal information interaction is established, thereby breaking through the computational redundancy and noise interference dilemma of the traditional fully connected interaction mode.

[0046] By calculating the simulated magnetic attraction effect, the system automatically assesses the strength of attraction between each local visual feature and the semantic description, forming a set of quantitative indicators reflecting the significance of cross-modal associations. This dynamic assessment not only captures explicit associations (such as the strong correlation between the "gliding" action and foot trajectory characteristics), but also discovers implicit associations (such as the potential connection between "elegance" style and the smoothness of body lines) through distance metrics in vector space, providing an interpretable mathematical basis for subsequent interactive screening.

[0047] In a specific example of the present application, the feature imitation magnetic effect value between the animation effect semantic feature vector and each animation sketch local visual feature coding vector in the set of animation sketch local visual feature coding vectors is calculated using the following formula to obtain a set of animation effect visual feature imitation magnetic effect values; wherein, the formula is:

[0048]

[0049] Wherein, ||·||2 represents the bi-norm of the vector, u is the semantic feature vector of the animation effect, and W align is the trainable permutation weight matrix, ε is the trainable modulation weight vector, s i For u and v i The characteristic magnetic attraction factor between them, max represents the maximum value in the extracted vector, min represents the minimum value in the extracted vector, and mean represents the average value of the calculated vector.

[0050] Then, based on the comparison between the set of simulated magnetic effect values ​​of the animation effect visual features and the preset threshold, a fast interaction set of the animation sketch local visual feature encoding vectors is extracted from the set of animation sketch local visual feature encoding vectors. It should be understood that the decoupled sketch local features may contain hundreds of visual units, and the traditional full interaction mode needs to process the association between all features and semantics, which not only leads to a waste of computing resources, but is also likely to cause the key action parameters to be overwhelmed by secondary features due to interference from noise features. Therefore, in order to construct a sparse attention mechanism for cross-modal interaction, in the technical solution of the present application, the system can intelligently identify the core elements that are strongly associated with the semantic description from the massive visual features, rather than evenly distributing computing resources to all sketch elements. This selective interaction mechanism actively discards redundant feature interaction paths while retaining the core cross-modal information, so that the LSTM-Transformer encoder can focus on modeling key semantic-visual associations, avoiding the dilution of model capacity caused by processing irrelevant features, and thus solving the dual problems of response delay and detail distortion of traditional solutions in complex scenarios.

[0051] In a specific example of the present application, based on the comparison between the set of simulated magnetic effect values ​​of the animation effect visual feature and a preset threshold, a fast interactive set of animation sketch local visual feature encoding vectors is extracted from the set of animation sketch local visual feature encoding vectors using the following formula; wherein, the formula is:

[0052]

[0053] X={u;v1',v2',...,v k '}

[0054] Where τ is a predetermined threshold, V′ is a fast interactive set of local visual feature encoding vectors of the animation sketch, v1', v2', v k They are respectively the 1st, 2nd and kth animation sketch local visual feature coding vectors in the fast interactive set of the animation sketch local visual feature coding vectors.

[0055] Furthermore, a general correlation correction is performed on each animation sketch local visual feature encoding vector in the fast interaction set of animation sketch local visual feature encoding vectors to obtain a fast interaction set of corrected animation sketch local visual feature encoding vectors. It should be understood that when the system screens out a subset of local features that are strongly associated with the language description from a large number of sketch features through the simulated magnetic effect (i.e., the fast interaction set of animation sketch local visual feature encoding vectors), although these features are unidirectionally bound to the language intent, they are still isolated from each other. This feature discreteness will cause a break in the physical logic when generating a coherent animation. For example, if the leg force feature and the torso tilt feature are not mechanically related, the limb movements of the generated character may be disconnected when jumping. Therefore, in order to construct the dynamic field constraints of the feature space, bridge the semantic gaps within the fast interaction set of the animation sketch local visual feature coding vectors, and elevate the visual fragments absorbed by the language intention from a mechanical stack to an organic whole, in a preferred example of the present application, a general correlation correction is performed on each animation sketch local visual feature coding vector in the fast interaction set of the animation sketch local visual feature coding vectors to obtain a corrected fast interaction set of the animation sketch local visual feature coding vectors.

[0056] Here, universal correlation correction reconstructs the weak correlations between local feature vectors through mathematical mapping. This aims to overcome the limitations of traditional isolated feature processing after feature screening and establish a mechanism for the collaborative expression of local features in cross-modal interaction. Specifically, during the correction process, the global deviation of the feature's simulated magnetic effect value from the threshold is quantified through integral operations. Each local feature vector is then algebraically transformed based on a polynomial root mapping. This ensures that the filtered visual features, while maintaining strong correlations with the semantic vectors, form implicit associations with each other that conform to physical laws or the logic of artistic expression. For example, after correction, the vector space distribution of the foot thrust feature and the aerial posture feature in a character's jump movement automatically converges to a correlation pattern that conforms to kinematic constraints, ensuring that the generated animation meets the "explosive power" requirements of language descriptions while also satisfying the biomechanical plausibility of the limb movement. This correlation correction mechanism fundamentally addresses the motion distortion problem caused by traditional solutions due to isolated local feature processing. This ensures that the generated animation retains the user's creative individuality while possessing the physical plausibility and artistic expression of professional-grade animation.

[0057] Thus, the animation sketch local visual feature encoding vectors v′1, v′2, ..., v′ in the fast interactive set of the animation sketch local visual feature encoding vectors are k In the overall cross-modal interaction space, it is essentially in the neighborhood space centered on the animation effect semantic feature vector u. Then, while ensuring that the animation sketch local visual feature encoding vector v′1, v′2, ..., v′ kWhile the neighborhood correlation property of the animation effect semantic feature vector u is also expected, the animation sketch local visual feature encoding vector v′1, v′2, ..., v′ k They can have the commonality of proximity association, thus facilitating the extraction of subsequent contextual joint semantics.

[0058] First, for each v′ i The corresponding magnetic attraction characteristic parameter s i , calculate the difference s between it and the threshold τ i -τ, and in s i ∈[s min ,s max ] on the interval of s i -τ is integrated to pass through each vector v′ i The edge characteristics of all vectors v′1,v′2,...,v′ are reflected k The overall system characteristics, that is, the change of a single critical value s i -τ has global universality that is not affected by microscopic factors:

[0059]

[0060] Then, the individual fluctuations are used as the universal correlation of the polynomial roots of the overall fluctuations to modify the local visual feature encoding vector v′ of each animation sketch. i :

[0061]

[0062] That is, while maintaining the local visual feature encoding vector v′ of each animation sketch i As an independent distribution, the universal correlation relationship is determined by algebraic equations, so as to realize the local visual feature encoding vector v′1, v′2, ..., v′ of the animation sketch with the weak correlation stability of the correlation metric. k The neighborhood association between them is universal, thus promoting the extraction of subsequent contextual joint semantics.

[0063] Subsequently, the fast interactive set of animation effect semantic feature vectors and corrected animation sketch local visual feature encoding vectors is input into a multimodal multi-scale joint encoder based on the LSTM-Transformer hybrid architecture to obtain a cross-modal animation effect local context semantic joint encoding vector and a cross-modal animation effect global context semantic joint encoding vector. It should be understood that the corrected sketch local features may contain enhanced associated elements (such as the limb angles of the character's headwind posture and the lines of fluttering clothes), but the traditional single architecture is difficult to capture the temporal coherence of the action sequence (such as the dynamic relationship between the rhythm of the steps and the change in wind force) and the global coordination of the scene elements (such as the physical interaction between the storm particle effect and the character's movement). In the technical solution of the present application, the LSTM-Transformer hybrid architecture breaks through the capability limitations of the traditional model in cross-modal feature fusion by integrating temporal modeling and spatial attention mechanism, and realizes the multi-scale collaborative expression of animation elements.

[0064] During this process, the LSTM module uses a gated recurrent mechanism to parse and correct short-term dependencies in feature sequences (such as the mechanical transmission of a character's foot trajectory in consecutive frames), capturing the microscopic dynamics of motion decomposition. The Transformer's self-attention layer establishes long-range correlations between global features (such as the energy conservation relationship between storm intensity parameters and the character's motion resistance), coordinating the macroscopic layout of scene elements. This approach significantly improves the semantic fidelity and visual realism of complex animation generation.

[0065] In a specific example of the present application, the fast interactive set of the animation effect semantic feature vector and the modified animation sketch local visual feature encoding vector is input into a multimodal multi-scale joint encoder based on the LSTM-Transformer hybrid architecture to obtain a cross-modal animation effect local context semantic joint encoding vector and a cross-modal animation effect global context semantic joint encoding vector according to the following formula; wherein, the formula is:

[0066] h local =LSTM({v;v″1,v″2,...,v″ k})

[0067] h global =MultiHeadAttn({u;v″1,v″2,...,v″ k})

[0068] Among them, v″1,v″2,...,v″ k are the first, second and kth corrected animation sketch local visual feature encoding vectors in the fast interaction set of the corrected animation sketch local visual feature encoding vectors, LSTM(·) represents the LSTM encoder, h localis the local context semantic joint encoding vector of the cross-modal animation effect, h global It is the global context semantic joint encoding vector of the cross-modal animation effect, and MultiHeadAttn represents the multi-head attention mechanism.

[0069] Finally, the cross-modal animation effect local context semantic joint coding vector and the cross-modal animation effect global context semantic joint coding vector are fused to obtain the animation effect intention multimodal joint coding vector. It should be understood that when the user sketches the scene of "character crossing the lava zone", the local context coding may focus on the contact deformation details of the character's feet and the high-temperature surface, while the global context coding models the macroscopic correlation between the lava flow trend and the lens movement trajectory. If any dimension is relied upon alone, the generated animation will cause local physical distortion or overall rhythm disorder. Therefore, in the technical solution of the present application, the information island effect of the traditional hierarchical processing mode is broken through by integrating the features of the mathematical space, and a cross-modal representation system with both detail authenticity and global coordination is constructed.

[0070] During this process, the local context encoding vector carries high-frame-rate details such as character joint motion and particle effect triggering frequency, while the global encoding vector controls long-term animation properties such as scene lighting gradients and camera movement logic. A nonlinear fusion algorithm dynamically weights the two in latent space, enabling the semantic description of "traveling through lava" to accurately reproduce the immediate physical feedback of the foot contacting the lava (such as the lava splash pattern) while maintaining the dynamic matching of the overall motion of the lava flow with the character's displacement speed, forming a self-consistent animation logic that combines microscopic details with macroscopic rhythms. This significantly improves the spatiotemporal consistency of complex animation generation.

[0071] In a specific example of the present application, the cross-modal animation effect local context semantic joint coding vector and the cross-modal animation effect global context semantic joint coding vector are fused using the following formula to obtain the animation effect intention multimodal joint coding vector; wherein, the formula is:

[0072] h fusion =W fuse [h local ;h global ]+b fuse

[0073] Among them, W fuse and b fuse are the fusion weight matrix and fusion bias vector respectively, [·;·] represents vector cascade, h fusion A multimodal joint encoding vector for the animation effect intention.

[0074] Specifically, in S23, the multimodal joint features of the animation effect intention are input into the AIGC module to obtain a low-fidelity preview image of the three-dimensional animation. It should be understood that although the multimodal joint features of the animation effect intention obtained through multimodal understanding and fusion accurately encode the user's creative intention, they themselves are still an abstract, machine-readable numerical representation that the user cannot directly perceive. Therefore, in order to allow users to intuitively see the visualization results of their preliminary ideas and provide a basis for subsequent interactive adjustments, in the technical solution of the present application, the multimodal joint features of the animation effect intention are input into the AIGC module to obtain a low-fidelity preview image of the three-dimensional animation.

[0075] It's worth noting that the AIGC module here specifically refers to one or a group of specially trained deep learning models whose core function is to generate 3D animation content based on input feature instructions. Its working principles are typically based on advanced generative model architectures, such as generative adversarial networks (GANs), variational autoencoders (VAEs), diffusion models, or Transformers optimized for 3D scenes and animation sequences. These models learn from massive amounts of "feature-3D animation" paired data, mastering the mapping relationship between abstract feature descriptions and concrete 3D scenes, character poses, action sequences, camera movements, and even preliminary lighting and shadow effects. This allows users to instantly assess the system's understanding of their sketches and text descriptions, determining whether the current generation direction aligns with their intended vision. Based on this preview, users can then more effectively refine their instructions in the next round, gradually approaching their ultimate creative goal. This effectively avoids the multiple ineffective attempts and inefficiencies caused by initial misunderstandings in traditional AIGC solutions.

[0076] In particular, S3 performs structured scene decoding on the low-fidelity preview image of the three-dimensional animation to obtain a structured scene description, which includes a list of detailed interaction modules. It should be understood that when the system generates a low-fidelity preview image, although the user can intuitively perceive the prototype of the animation, it is difficult to accurately locate the details that need to be adjusted. Therefore, in the technical solution of this application, the animation effect intention multimodal joint coding vector is subjected to structured scene decoding based on the RNN model to obtain a structured scene description, which includes a list of detailed interaction modules. This structured information can clearly list the various components of the scene and their adjustable properties, thereby providing users with a clear interactive interface and control points. In this way, the internal, abstract intent representation of the AIGC model can be converted into a structured data that the user can intuitively understand and perform precise operations on, effectively solving the problem of users having difficulty in accurately controlling details and poor interactive experience in existing AIGC animation generation solutions, making the entire creation process more transparent, efficient, and personalized.

[0077] In specific implementation, the multimodal joint encoding vector of the animation effect intent is first input into a pre-trained recurrent neural network (RNN)-based decoder model. During the training phase, the RNN decoder learns how to map the input joint encoding vector to the corresponding structured scene description text or data structure. During the inference (i.e., decoding) phase, the RNN uses this joint encoding vector as the initial state or continuous input and then gradually generates a series of tokens or symbols. For example, the RNN may first generate a token representing "scene object," followed by the object type (such as "character," "prop," or "environmental element"), followed by the specific attributes of the object (such as the character's action "walking," the prop's material "metal," and the environment's weather "sunny") and the current values ​​of these attributes. More importantly, the decoding process identifies which attributes are suitable for user interaction and generates corresponding refined interaction module entries for these attributes. For example, for the character's "action" attribute, the decoder may generate an interaction module containing a drop-down list of other actions the character can perform (such as "run," "jump," and "stand"); for the color attribute, a color selector module may be generated. These generated token sequences are ultimately combined into a complete, structured scene description (e.g., in JSON or XML format), which contains all the main elements and their attributes, as well as a clear list of detailed interaction modules, each of which points to an editable unit in the scene and its available interaction controls.

[0078] In particular, the S4 obtains the refinement instructions for a specific refinement interaction module in the list of refinement interaction modules input by the user to obtain an updated structured scene description. It should be understood that although the initially generated low-fidelity preview image is based on the user's sketch and text description, it is often only a preliminary and possibly biased prototype of the concept. Therefore, in the technical solution of the present application, the refinement instructions for a specific refinement interaction module in the list of refinement interaction modules input by the user are further obtained to obtain an updated structured scene description. In this way, users can easily find and modify almost any aspect of the animation they want to adjust, whether it is a specific action of a character, the position and color of an object in the scene, the trajectory of the lens movement, or subtle changes in the overall light and shadow atmosphere. Each refinement instruction can be accurately reflected in the updated structured scene description, providing clear and accurate guidance for the subsequent AIGC module to regenerate the updated preview image of the three-dimensional animation. This highly controllable interactive method greatly improves the user's creative experience and efficiency, effectively solves the problem of repeated revisions caused by AIGC's misunderstanding or insufficient expression, and enables users to transform from passively accepting generated results to actively guiding and shaping the creative content.

[0079] In particular, in S5, the updated structured scene description is input into the AIGC module to obtain an updated preview of the 3D animation. In the technical solution of the present application, in order to ensure that each refinement operation of the user is accurately reflected in the final animation generation process and realize true interactive creation, the updated structured scene description is feature-reverse mapped and encoded to obtain an updated animation effect intention multimodal joint encoding vector. By reverse mapping the updated structured scene description to the updated animation effect intention multimodal joint encoding vector, the system can effectively integrate the local and specific modification intentions expressed by the user through the refinement interaction module into the original global creative intention, thereby generating an animation effect that retains the original creative core and incorporates the latest adjustments. This enables users to iteratively optimize based on the preliminary results generated by AIGC, gradually approaching the ideal animation effect they envisioned, greatly improving controllability and user satisfaction. Among them, the updated animation effect intention multimodal joint encoding vector not only retains the macro information of the user's original creative intention (from the initial sketch and text), but also accurately incorporates the specific adjustments made by the user later through the refinement interaction module. Furthermore, the updated animation effect intention multimodal joint encoding vector is input into the AIGC module to generate a preview of the 3D animation update. This allows users to make precise modifications through an intuitive structured editing interface and receive immediate visual feedback of the changes.

[0080] In summary, the interactive 3D animation generation method based on AIGC technology according to the embodiment of the present application is explained. It performs cross-modal joint understanding of the animation sketch and natural language description input by the user, and generates a preliminary low-fidelity preview image of the 3D animation based on this. Then, by introducing a structured scene decoding mechanism, the complex multimodal joint features of the animation intention are parsed into a structured scene description containing a series of refined interaction modules, so as to decompose the macro animation intention into local units that can be precisely intervened by the user. The user can issue specific detailed instructions for these specific refined interaction modules. After receiving the instructions, the system updates the structured scene description and generates an updated 3D animation preview image based on the updated structured scene description. In this way, the user can make targeted and refined adjustments to the animation, which greatly enhances the controllability and flexibility of animation generation and effectively solves the problem of difficult detail adjustment in traditional AIGC solutions.

[0081] Furthermore, an interactive three-dimensional animation generation system based on AIGC technology is also provided.

[0082] Figure 4 FIG is a block diagram of an interactive 3D animation generation system based on AIGC technology according to an embodiment of the present application. Figure 4As shown, according to an embodiment of the present application, an interactive 3D animation generation system 300 based on AIGC technology includes: a user intention acquisition module 310, which is used to obtain an animation sketch and a natural language description of an animation effect input by a user; a 3D animation preliminary generation module 320, which is used to input the animation sketch and the natural language description of the animation effect into the AIGC module to obtain a low-fidelity preview image of the 3D animation; a structured scene decoding module 330, which is used to perform structured scene decoding on the low-fidelity preview image of the 3D animation to obtain a structured scene description, wherein the structured scene description includes a list of refined interaction modules; a scene description update module 340, which is used to obtain a refinement instruction input by a user for a specific refined interaction module in the list of refined interaction modules to obtain an updated structured scene description; and a 3D animation update module 350, which is used to input the updated structured scene description into the AIGC module to obtain an updated preview image of the 3D animation.

[0083] As described above, the interactive 3D animation generation system 300 based on AIGC technology according to an embodiment of the present application can be implemented in various wireless terminals, such as a server equipped with an interactive 3D animation generation algorithm based on AIGC technology. In one possible implementation, the interactive 3D animation generation system 300 based on AIGC technology according to an embodiment of the present application can be integrated into a wireless terminal as a software module and / or a hardware module. For example, the interactive 3D animation generation system 300 based on AIGC technology can be a software module in the operating system of the wireless terminal, or can be an application developed specifically for the wireless terminal. Of course, the interactive 3D animation generation system 300 based on AIGC technology can also be one of the many hardware modules of the wireless terminal.

[0084] Alternatively, in another example, the interactive three-dimensional animation generation system 300 based on AIGC technology and the wireless terminal may also be separate devices, and the interactive three-dimensional animation generation system 300 based on AIGC technology may be connected to the wireless terminal via a wired and / or wireless network and transmit interactive information in accordance with an agreed data format.

[0085] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. An interactive 3D animation generation method based on AIGC technology, characterized in that: include: Obtaining an animation sketch and a natural language description of the animation effect input by a user; Input the animation sketch and natural language description of the animation effect into the AIGC module to obtain a low-fidelity preview of the 3D animation; Performing structured scene decoding on the low-fidelity preview image of the three-dimensional animation to obtain a structured scene description, wherein the structured scene description includes a list of detailed interaction modules; Obtaining a refinement instruction input by a user for a specific refinement interaction module in the list of refinement interaction modules to obtain an updated structured scene description; The updated structured scene description is input into the AIGC module to obtain the updated 3D animation preview.

2. The interactive 3D animation generation method based on AIGC technology according to claim 1, characterized in that: Input the animation sketch and natural language description of the animation effect into the AIGC module to obtain a low-fidelity preview of the 3D animation, including: Parse the natural language descriptions of animation sketches and animation effects respectively to obtain the visual features of animation sketches and the semantic features of animation effects; Perform cross-modal joint understanding of animation intent on the animation effect semantic features and animation sketch visual features to obtain multimodal joint features of animation effect intent; The multimodal joint features of animation effect intention are input into the AIGC module to obtain a low-fidelity preview image of the 3D animation.

3. The interactive 3D animation generation method based on AIGC technology according to claim 2, characterized in that: The natural language descriptions of animation sketches and animation effects are parsed separately to obtain the visual features of animation sketches and the semantic features of animation effects, including: Use the NLU module to parse the natural language description of the animation effect to obtain the animation effect semantic feature vector as the animation effect semantic feature; The animation sketch is parsed using the CV module to obtain an animation sketch visual feature map as the animation sketch visual feature.

4. The interactive 3D animation generation method based on AIGC technology according to claim 2, characterized in that: The animation intent is cross-modally understood by combining the animation effect semantic features and the animation sketch visual features to obtain the multimodal joint features of the animation effect intent, including: Decoupling the visual feature map of the animation sketch to obtain a set of local visual feature encoding vectors of the animation sketch; Based on the feature-simulated magnetic attraction effect between the animation effect semantic feature vector and each animation sketch local visual feature coding vector in the set of animation sketch local visual feature coding vectors, a multimodal and multi-scale joint coding analysis is performed on the set of animation effect semantic feature vectors and animation sketch local visual feature coding vectors to obtain a cross-modal animation effect local context semantic joint coding vector and a cross-modal animation effect global context semantic joint coding vector; The cross-modal animation effect local context semantic joint encoding vector and the cross-modal animation effect global context semantic joint encoding vector are fused to obtain the animation effect intention multimodal joint encoding vector as the animation effect intention multimodal joint feature.

5. The interactive 3D animation generation method based on AIGC technology according to claim 4, characterized in that: Based on the feature-simulated magnetic attraction effect between the animation effect semantic feature vector and each animation sketch local visual feature coding vector in the set of animation sketch local visual feature coding vectors, a multimodal and multi-scale joint coding analysis is performed on the set of animation effect semantic feature vectors and animation sketch local visual feature coding vectors to obtain a cross-modal animation effect local context semantic joint coding vector and a cross-modal animation effect global context semantic joint coding vector, including: Calculating a feature simulated magnetic effect value between the animation effect semantic feature vector and each animation sketch local visual feature encoding vector in a set of animation sketch local visual feature encoding vectors to obtain a set of animation effect visual feature simulated magnetic effect values; Based on the comparison between the set of simulated magnetic effect values ​​of the animation effect visual features and a preset threshold, a fast interactive set of animation sketch local visual feature encoding vectors is extracted from the set of animation sketch local visual feature encoding vectors; The fast interactive set of animation effect semantic feature vectors and animation sketch local visual feature encoding vectors is input into a multimodal multi-scale joint encoder based on the LSTM-Transformer hybrid architecture to obtain a cross-modal animation effect local context semantic joint encoding vector and a cross-modal animation effect global context semantic joint encoding vector.

6. The interactive 3D animation generation method based on AIGC technology according to claim 5, characterized in that: The fast interactive set of animation effect semantic feature vectors and animation sketch local visual feature encoding vectors is input into a multimodal multi-scale joint encoder based on the LSTM-Transformer hybrid architecture to obtain a cross-modal animation effect local context semantic joint encoding vector and a cross-modal animation effect global context semantic joint encoding vector, including: Performing universal correlation correction on each animation sketch local visual feature encoding vector in the fast interaction set of animation sketch local visual feature encoding vectors to obtain a corrected fast interaction set of animation sketch local visual feature encoding vectors; The fast interactive set of animation effect semantic feature vectors and corrected animation sketch local visual feature encoding vectors is input into a multimodal multi-scale joint encoder based on the LSTM-Transformer hybrid architecture to obtain cross-modal animation effect local context semantic joint encoding vectors and cross-modal animation effect global context semantic joint encoding vectors.

7. The interactive 3D animation generation method based on AIGC technology according to claim 1, characterized in that: Performing structured scene decoding on the low-fidelity preview image of the 3D animation to obtain a structured scene description, wherein the structured scene description includes a list of detailed interaction modules, including: The multimodal joint encoding vector of the animation effect intention is subjected to structured scene decoding based on the RNN model to obtain a structured scene description, which includes a list of refined interaction modules.

8. The interactive 3D animation generation method based on AIGC technology according to claim 1, characterized in that: Input the updated structured scene description into the AIGC module to obtain the updated 3D animation preview, including: Perform feature reverse mapping encoding on the updated structured scene description to obtain an updated multimodal joint encoding vector of the animation effect intention; The updated animation effect intention multimodal joint encoding vector is input into the AIGC module to obtain a 3D animation update preview image.

9. An interactive 3D animation generation system based on AIGC technology, characterized in that: include: A user intention acquisition module is used to obtain the animation sketch and natural language description of the animation effect input by the user; The 3D animation preliminary generation module is used to input the animation sketch and natural language description of the animation effect into the AIGC module to obtain a low-fidelity preview image of the 3D animation; A structured scene decoding module is used to perform structured scene decoding on the 3D animation low-fidelity preview image to obtain a structured scene description, wherein the structured scene description includes a list of detailed interaction modules; A scene description updating module, configured to obtain a refinement instruction input by a user for a specific refinement interaction module in the list of refinement interaction modules to obtain an updated structured scene description; The 3D animation update module is used to input the updated structured scene description into the AIGC module to obtain a 3D animation update preview image.

Citation Information

Patent Citations

  • Three-dimensional model editing method and device based on multiple modes, equipment and medium

    CN119540504A

  • Systems and Methods for Language-Based Three-Dimensional Interactive Environment Construction and Interaction

    US20250029348A1

Cited By

  • Animation production system and method

    CN121259133A