Multi-mode driving character animation automatic generation method and system for virtual reality
By using a multimodal information-driven method for automatic generation of character animation, the problems of stiff action rhythm and disconnect between plot and atmosphere in virtual reality are solved. This method achieves natural matching between action and plot and full utilization of multimodal information, thereby enhancing the narrative appeal and immersion of virtual reality animation.
Patent Information
- Application Number
- CN202511182412.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-25
AI Technical Summary
Existing methods and systems for automatically generating multimodal character animations suffer from problems in virtual reality, such as stiff movement rhythm, disconnect between plot and atmosphere, unreasonable action sequence, insufficient matching between character performance and plot, and inadequate utilization of multimodal information, resulting in a weak sense of immersion.
By acquiring multimodal creation information, generating a stream of creation instruction symbols using a contextual intent parsing mechanism, forming a rhythmic event queue by combining plot node mapping and priority drift rules, generating character actions using a master-slave dual-track logic, introducing a camera perspective projection mechanism, automatically retrieving motion resource libraries and performing style transfer rendering, and generating animations that conform to multimodal information.
It achieves a natural match between action and plot, enhances the narrative appeal and immersive experience of virtual reality scenes, makes action performance clear and distinct, and enriches emotional expression driven by multimodal information, thereby improving the production efficiency and immersion of virtual reality animation.
Smart Images

Figure CN121010678A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of animation production technology, and more specifically, to a method and system for automatically generating multimodal driven character animations for virtual reality. Background Technology
[0002] Existing methods and systems for automatically generating multimodal character animations mainly suffer from the following problems:
[0003] With the rapid development of virtual reality (VR) technology, virtual character animation is playing an increasingly important role in film, games, and interactive applications. However, existing methods and systems for automatically generating multimodal character animation still have many shortcomings.
[0004] Traditional methods typically rely solely on timestamps of action triggers or simple linear sequencing to generate event sequences. This results in stiff character movements that fail to reflect the pacing and rhythm of the plot. Furthermore, current technologies often consider only temporal sequence or a single plot weight when generating animations, causing a disconnect between action rhythm and narrative atmosphere, making it difficult to adapt to different scene requirements. When the priority difference between adjacent events is too large, traditional linear mapping methods can cause abrupt changes in event intervals, resulting in an unnatural distribution of action rhythm and thus disrupting the user's immersion in the virtual reality environment.
[0005] Existing technologies largely rely on uniform time sequences or manual adjustments, lacking intelligent judgment of the urgency of events and the pacing of the plot. This results in illogical action sequences in the generated animations and insufficient matching between character performance and the storyline. Furthermore, traditional methods cannot distinguish between core and supplementary actions, easily leading to chaotic and overlapping character movements, making core actions less prominent and hindering the effective expression of emotions and dramatic tension. Existing technologies also have limitations in utilizing multimodal information. Current methods typically do not fully utilize multimodal creative information such as voice, text, music, and sketches for action generation, resulting in generated character movements lacking emotional drive and expressive intent, leading to insufficient immersion. During action generation, the lack of dynamic adjustment mechanisms based on time, emotion, and plot points makes it difficult to achieve automatic adaptation of character movements and the generation of supplementary actions, limiting the naturalness and expressiveness of character animations in virtual reality scenes.
[0006] In view of this, the present invention proposes a method for automatically generating multimodal driven character animations for virtual reality to solve the above problems. Summary of the Invention
[0007] To overcome the aforementioned deficiencies of the prior art and to achieve the above objectives, the present invention provides the following technical solution: a method for automatically generating multimodal driven character animations for virtual reality, comprising:
[0008] S1. Obtain multimodal creation information input by the developer, and transform the multimodal creation information into a creation instruction symbol stream containing plot rhythm markers, emotional color markers, and action tendency markers through the contextual intent parsing mechanism;
[0009] S2. Decompose the creative instruction symbol flow into different independent event units, and through plot node mapping and priority drift rules, assign the independent event units to rhythmic grid points on the timeline to form a rhythmic event queue.
[0010] S3. Based on the event queue, generate character actions through master-slave dual-track logic, introduce a camera perspective projection mechanism, automatically adjust the expressive tension of character actions according to scene changes, and generate a logical frame sequence with a camera feel.
[0011] S4. Automatically retrieve action units from the preset action resource library based on the logical frame sequence, match the action units with the emotion dimension, spatial dimension and scene dimension through a cross-dimensional index table, and output a continuous action fragment stream.
[0012] S5 receives the motion clip stream and performs automatic rendering binding on the character. It adopts a style transfer rendering mechanism to generate animation versions with different art styles and exports them as finished animation files that can be directly called on different production platforms.
[0013] Specifically, the method for obtaining the creation instruction symbol stream includes:
[0014] The system acquires multimodal creation information input by developers, including voice description information, sketch action information, plot script text, character position change information, and background music information; it normalizes the timestamps of each modality and aligns them using a dynamic time warping method to form a multimodal feature sequence on a unified timeline;
[0015] Using an attention-weighted fusion method, the features of each modality are integrated according to their contextual relevance to generate a unified context vector. Based on the music beat, text event interval, and speech rate, music beat features, text event interval features, and speech rate features are extracted from the unified context vector and generated by weighted combination.
[0016] Emotional analysis is performed on text, speech, and music modalities. Emotional features from each modality are fused using adjustable weights to generate standardized emotion vectors, which serve as emotion color markers. Semantic description information, sketch action information, and character position change information are mapped to action probability distributions to form action tendency markers. Plot rhythm markers, emotion color markers, and action tendency markers are integrated to form a unified creative instruction symbol flow.
[0017] Specifically, the method for forming a rhythmic event queue includes:
[0018] The creation instruction symbol flow is decomposed into different independent event units. For each independent event unit, an event priority score is calculated based on the importance of plot nodes, emotional intensity, and time priority. The event priority score is then non-linearly mapped to the plot timeline, with events with higher priority scores automatically mapped to more densely packed rhythm points.
[0019] When the difference in priority scores between adjacent events exceeds a preset priority score difference threshold, a priority drift mechanism is used to adjust the time interval between events, so that the event units form a rhythmic arrangement on the time axis, resulting in a rhythmic event queue.
[0020] Specifically, the method for generating character actions through master-slave dual-track logic includes:
[0021] In the received event queue, each independent event unit is parsed into an action trigger point. Each action trigger point includes the start timestamp of the action, the end timestamp of the action, and the duration of the action. A time priority score is calculated based on how close the trigger time of the action is to the current time.
[0022] Action trigger points are sorted according to time priority scores. Actions with the highest priority scores are classified as main track actions, and actions with other priority scores are classified as secondary track actions. Generation strategies are established for main track actions and secondary track actions respectively. Main track actions generate the character's core actions based on the action time sequence, plot node mapping, and action tendency markers. Secondary track actions generate the character's supplementary actions based on the time offset of the main track actions and the emotional color markers.
[0023] Specifically, the method for generating a logical frame sequence with a cinematic feel includes:
[0024] The system obtains the current camera's spatial position, orientation information, focal length, and motion trajectory in a preset virtual scene, and calculates the projection parameters of the action from the camera's perspective based on the relative spatial relationship between the character and the camera. The projection parameters include the viewing angle, visible area, and depth of field.
[0025] For each main track action and secondary track action, a motion performance tension score is calculated based on the motion amplitude, speed, motion trigger point position, and focal length adaptation to measure the visual impact and emotional expression intensity of the action under the current camera perspective.
[0026] The method for obtaining the motion performance tension score includes quantifying the motion amplitude, motion speed, the position of the motion trigger point on the time axis, the key points of the motion, and the focal length adaptation into independent scores, and then weighting and summing them through preset weights to form a comprehensive score, which is used as the motion performance tension score.
[0027] Based on the calculated action performance tension score, the character's actions are adjusted, including scaling the action amplitude, increasing or decreasing the action speed, and advancing or delaying the action time. The tension-adjusted character actions are then combined with the time information of the master and slave tracks and the action trigger points to generate a logical frame sequence containing character posture information, action amplitude, speed, timestamp, and camera projection parameters.
[0028] Specifically, the method for automatically retrieving action units from a preset action resource library based on a logical frame sequence includes:
[0029] Different action units are stored in the preset action resource library. Each action unit corresponds to action type, action amplitude, speed curve, emotion label, spatial position change and adaptation scene conditions. The action units are mapped to emotion dimension, spatial dimension, scene dimension and rhythm dimension through a cross-dimensional index table.
[0030] For each frame in the logical frame sequence, action units are filtered in the preset action resource library according to the logical frame attributes, and a set of action units that meet the preset requirements in the dimensions of emotion, space, scene and rhythm is retained.
[0031] Specifically, the method for acquiring the continuous action segment stream includes:
[0032] The cosine similarity method is used to quantify and match the attributes of logical frames with the attributes of action units to obtain the corresponding matching scores. According to the time order of the logical frame sequence, the action unit with the highest matching score is selected for each frame and arranged continuously in the time order of the logical frame sequence to form a continuous action segment stream.
[0033] Specifically, the method for performing automatic rendering binding on the character includes:
[0034] Receive a continuous stream of motion clips, each containing key motion points, motion amplitude, speed curve, emotion markers, motion tendency markers, and scene parameters; load a character skeleton model for the virtual character, and map the key motion points in the motion clip stream to the corresponding nodes of the character skeleton joints, thus binding the motion to the character skeleton.
[0035] Based on the timestamp information in the continuous action fragment stream, each action fragment is synchronously bound to the character's skeletal model in chronological order, so that the action trigger point, duration, and action order are consistent with the logical frame sequence.
[0036] Perform interpolation and transition processing on frames between consecutive action segments, including position interpolation, rotation interpolation, and speed smoothing; map the emotional markers and action performance tension information carried in the action segment stream to the character's action parameters;
[0037] The character's motion parameters include hand tension, limb curvature, and facial expression range, so that the character's movements not only conform to physical logic but also express the preset emotions and plot tension, ultimately completing the automatic rendering and binding of continuous motion segments onto the virtual character.
[0038] Specifically, the method for generating animation versions with different art styles includes:
[0039] The virtual character motion frames that have been automatically rendered and bound are input into the preset style transfer rendering module. Based on the preset art style samples, the module extracts the content features of the character motion frames through a deep learning model, and maps and fuses the content features with the preset target style features to generate stylized animation frames that conform to the target art style.
[0040] The stylized animation frames are continuously optimized, including position, rotation, speed, and smoothness processing. Based on the rendering requirements of different platforms, the stylized animation frames are exported to different standard animation file formats, so that the exported animation files can be directly called in various game engines, animation production software, or virtual reality platforms, ultimately resulting in virtual character animation files with different art styles.
[0041] A multimodal-driven character animation automatic generation system for virtual reality includes:
[0042] The creative intent capture module acquires multimodal creative information input by developers and transforms it into a stream of creative instruction symbols containing plot rhythm markers, emotional color markers, and action tendency markers through a contextual intent parsing mechanism.
[0043] The narrative event construction module decomposes the creative instruction symbol flow into different independent event units, and assigns the independent event units to rhythmic grid points on the timeline through plot node mapping and priority drift rules, forming a rhythmic event queue.
[0044] The character animation generation module generates character actions based on the event queue through a master-slave dual-track logic. It introduces a camera perspective projection mechanism to automatically adjust the expressive tension of the character actions according to scene changes, generating a logical frame sequence with a cinematic feel.
[0045] The action resource retrieval module automatically retrieves action units from the preset action resource library based on the logical frame sequence, matches the action units with the emotion dimension, spatial dimension and scene dimension through a cross-dimensional index table, and outputs a continuous action fragment stream.
[0046] The rendering and export module receives motion clip streams, performs automatic rendering and binding on the characters, adopts a style transfer rendering mechanism to generate animation versions with different art styles, and exports them as finished animation files that can be directly used on different production platforms.
[0047] Compared with the prior art, the present invention has the following beneficial effects:
[0048] This invention uses a non-linear mapping between event priority scores and the plot timeline to automatically align high-priority events with dense rhythmic points, thus naturally reflecting the progression and release of plot tension and enhancing the narrative expressiveness of virtual reality animation. Multimodal information fusion improves the accuracy of event priority calculation, avoiding the limitations of single-dimensional driving and making the generated results more in line with user expectations and immersive experience. A priority drift mechanism dynamically fine-tunes the time intervals, preventing abrupt breaks in event intervals and achieving smooth transitions of events on the timeline, thereby ensuring a natural and fluid overall rhythm distribution. It can generate action event queues that conform to plot logic and rhythmic patterns in the virtual reality environment, maintaining a unified rhythmic atmosphere between character actions, music, voice, and plot development, enhancing the narrative appeal and immersive experience of virtual reality scenes.
[0049] By calculating time priority scores for action trigger points, considering the start time, end time, duration, and key rhythmic points of actions, the triggering order of character actions automatically adapts to the rhythm of the plot, ensuring smooth and natural animation and enhancing immersion in virtual reality scenes. Actions with the highest priority scores are designated as primary track actions, generating core character actions; other actions are designated as secondary track actions, generating supplementary actions, achieving a clear hierarchy of action performance, highlighting plot points, and ensuring clear logic in character actions. Primary track actions, combined with action tendency markers, generate core actions; secondary track actions, combined with emotional color markers, generate supplementary actions, enabling actions to respond to multimodal creative information, making character actions more aligned with the developer's intentions and enriching emotional expression. Separate generation strategies are established for primary and secondary track actions, enabling automatic adjustment and dynamic supplementation of actions, reducing manual adjustments by animators, improving the efficiency of virtual reality animation production, and ensuring a high degree of consistency between actions, plot, and emotions. Through time priority scoring, primary / secondary track division, and multimodal information-driven action generation strategies, character actions are highly matched to the plot in time, space, and emotion, enhancing the user's realistic experience and immersion in the virtual reality environment. Attached Figure Description
[0050] Figure 1 This is a schematic diagram of the process for automatically generating multimodal driven character animations for virtual reality according to the present invention.
[0051] Figure 2 This is a schematic diagram of the multimodal-driven character animation automatic generation system for virtual reality according to the present invention.
[0052] Figure 3 This is a schematic diagram of the method for automatically rendering and binding characters provided by the present invention. Detailed Implementation
[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] Example 1
[0055] Please see Figure 1 and Figure 3 As shown, this embodiment provides a method for automatically generating multimodal driven character animations for virtual reality, specifically including the following steps:
[0056] S1. Obtain multimodal creation information input by the developer, and transform the multimodal creation information into a creation instruction symbol stream containing plot rhythm markers, emotional color markers, and action tendency markers through the contextual intent parsing mechanism;
[0057] S2. Decompose the creative instruction symbol flow into different independent event units, and through plot node mapping and priority drift rules, assign the independent event units to rhythmic grid points on the timeline to form a rhythmic event queue.
[0058] S3. Based on the event queue, generate character actions through master-slave dual-track logic, introduce a camera perspective projection mechanism, automatically adjust the expressive tension of character actions according to scene changes, and generate a logical frame sequence with a camera feel.
[0059] S4. Automatically retrieve action units from the preset action resource library based on the logical frame sequence, match the action units with the emotion dimension, spatial dimension and scene dimension through a cross-dimensional index table, and output a continuous action fragment stream.
[0060] S5 receives the motion clip stream and performs automatic rendering binding on the character. It adopts a style transfer rendering mechanism to generate animation versions with different art styles and exports them as finished animation files that can be directly called on different production platforms.
[0061] Methods for obtaining the creation instruction symbol stream include:
[0062] The system acquires multimodal creation information input by developers, including voice description information, sketch action information, plot script text, character position change information, and background music information; it normalizes the timestamps of each modality and aligns them using a dynamic time warping method to form a multimodal feature sequence on a unified timeline;
[0063] Using an attention-weighted fusion method, features from various modalities are weighted and integrated according to their contextual relevance to generate a unified context vector. Based on music beat, text event interval, and speech rate, music beat features, text event interval features, and speech rate features are extracted from the unified context vector and generated by weighted combination to produce plot rhythm markers, so that the generated creative instruction symbol stream can reflect the sense of rhythm under multimodal drive. The attention weights are dynamically adjusted according to the relevance of each modality to the current creative context, so that the contribution of each modality to the generated creative instruction symbol stream can be adaptively changed at different time points.
[0064] Emotional analysis is performed on text, speech, and music modalities. Emotional features from each modality are fused using adjustable weights to generate standardized emotion vectors, which serve as emotion color markers. Semantic description information, sketch action information, and character position change information are mapped to action probability distributions to form action tendency markers. Plot rhythm markers, emotion color markers, and action tendency markers are integrated to form a unified creative instruction symbol flow.
[0065] Methods for creating rhythmic event queues include:
[0066] The creation instruction symbol flow is decomposed into different independent event units. For each independent event unit, an event priority score is calculated based on the importance of plot nodes, emotional intensity, and time priority. The event priority score is then non-linearly mapped to the plot timeline (e.g., using a sigmoid function). Events with higher event priority scores are automatically mapped to more densely packed rhythm points.
[0067] Preset number The rhythmic grid points of each event unit on the timeline are: The time interval between the previous event unit and the previous event unit is ;in, This indicates the time interval between the current event unit and the previous event unit; Indicates the first The plot priority rating of each event unit; This indicates the maximum upper limit of the rhythm interval, used to constrain the stretching range of the time axis; This represents the rhythm sensitivity adjustment coefficient, which determines how sensitive priority changes are to interval adjustments; This represents the rhythm balance factor, used to adjust the center offset of the event interval distribution; Indicates the index of the event unit;
[0068] The importance of plot nodes is determined by core event tags (such as “turning point”, “climax”, and “ending point”) extracted from the plot script text, which can be calculated by plot structure analysis algorithms (such as key event extraction models based on natural language processing). Plot nodes are divided into different levels, for example, ordinary events are recorded as 1, supporting key events as 2, turning or climax nodes as 3, and ending or key plot driving nodes as 4, thus forming a discretized importance score.
[0069] Emotion intensity is determined using an emotion classification model (such as a multimodal emotion recognition network) to output emotion scores such as pleasure, tension, and sadness. The emotion vector is projected into an emotion intensity scalar interval [0,1], where 0 represents no obvious emotion and 1 represents the highest emotion intensity. For example, a high emotion intensity score can be obtained when high-pitched music is accompanied by a fast tempo, the text triggers the "anger" label, and the speech is rapid.
[0070] When the difference in priority scores between adjacent events exceeds a preset priority score difference threshold, a priority drift mechanism is used to adjust the time interval between events, so that the event units form a rhythmic arrangement on the time axis, resulting in a rhythmic event queue.
[0071] This solution addresses the technical problems of existing technologies: Traditional methods often generate event sequences solely based on action trigger timestamps or simple linear ordering, resulting in rigid action trigger rhythms that fail to reflect the pacing and cadence of the plot. Existing technologies mostly consider only temporal sequence or a single plot weight, leading to a disconnect between the generated action rhythm and the narrative atmosphere, making it difficult to adapt to different types of scene requirements. When the priority difference between adjacent events is too large, traditional linear mapping methods can cause abrupt changes in event intervals, resulting in an unnatural rhythm distribution and disrupting the user's immersion in virtual reality.
[0072] Methods for generating character actions using a master-slave dual-track logic include:
[0073] In the received event queue, each independent event unit is parsed into an action trigger point. Each action trigger point includes the start timestamp of the action, the end timestamp of the action, and the duration of the action. A time priority score is calculated based on how close the trigger time of the action is to the current time.
[0074] Time priority rating ;in, The score indicates the time priority of the action, reflecting the urgency of the action in terms of time. The higher the score, the higher the priority. Indicates the start timestamp of the action; Indicates the deadline timestamp for the action; Indicates the duration of the action; This represents the rhythmic key point factor, with a value of 0 or 1. This indicates the urgency weight, reflecting the degree of influence of urgency in the time priority score; This indicates the weight of key rhythm points, used to adjust the degree to which whether an event is at a key rhythm point affects the time priority score. This indicates the duration weight, used to adjust the impact of event duration on time priority scoring;
[0075] Action trigger points are sorted according to time priority scores. Actions with the highest priority scores are classified as main track actions, and actions with other priority scores are classified as secondary track actions. Generation strategies are established for main track actions and secondary track actions respectively. Main track actions generate the character's core actions based on the action time sequence, plot node mapping, and action tendency markers. Secondary track actions generate the character's supplementary actions based on the time offset of the main track actions and the emotional color markers.
[0076] This solution addresses the following issues with existing technologies: Existing technologies often rely on a uniform time sequence or manual adjustments, failing to intelligently determine the action trigger order based on the urgency of events and the pacing of the plot. This results in unnatural animation rhythm and a mismatch between character performance and the storyline. Existing methods often fail to distinguish between core and supplementary actions, leading to chaotic character movement stacking, a lack of emphasis on core actions, and difficulty in conveying emotional expression and plot tension. Existing technologies typically fail to fully utilize multimodal creative information such as voice, text, music, and sketches for action generation, resulting in a lack of multimodal-driven emotional and behavioral expression in character movements, leading to insufficient immersion. Furthermore, the lack of a dynamic adjustment mechanism based on time, emotion, and plot nodes during action generation hinders the automatic adaptation of character movements and the generation of supplementary actions.
[0077] Methods for generating logical frame sequences with a cinematic feel include:
[0078] The system obtains the current camera's spatial position, orientation information, focal length, and motion trajectory in a preset virtual scene, and calculates the projection parameters of the action from the camera's perspective based on the relative spatial relationship between the character and the camera. The projection parameters include the viewing angle, visible area, and depth of field.
[0079] The perspective angle is calculated based on the relative spatial relationship between each key action point and the camera position, as well as the camera orientation vector. The angle between the key action point and the camera's visual axis is calculated, and the smaller the angle, the closer the action is to the center of the camera, and the more prominent the visual performance. The visible area is determined by projecting the key action points onto the camera's projection plane and judging whether the projected points are located within the camera's visible area and their relative position in the frame. This determines the visible coverage of the action in the frame, in order to identify the visual priority of key actions. The depth of field is calculated based on the camera's focal length, aperture, and focus position. The spatial distance difference between the key action point and the camera's focus is calculated and mapped to a depth of field weight to obtain the depth of field range. The higher the depth of field weight, the clearer the action is within the focus range, thus enhancing the visual expressiveness.
[0080] For each main track action and secondary track action, a motion performance tension score is calculated based on the motion amplitude, speed, motion trigger point position, and focal length adaptation to measure the visual impact and emotional expression intensity of the action under the current camera perspective.
[0081] The method for obtaining the motion performance tension score includes quantifying the motion amplitude, motion speed, the position of the motion trigger point on the time axis, the key points of the motion, and the focal length adaptation into independent scores, and then weighting and summing them through preset weights to form a comprehensive score, which is used as the motion performance tension score.
[0082] Action amplitude scoring extracts the 3D spatial displacement amplitude of key action points (such as joint movement distance and rotation angle changes), maps the amplitude to a fixed scoring range using a normalization function (e.g., mapping the maximum amplitude to 1 and the minimum amplitude to 0). Action speed scoring calculates the speed change (displacement / time) between adjacent keyframes, uses a smoothing function or moving average to eliminate noise, and then maps the speed value to a 0-1 range as the score. The position of the action trigger point on the timeline is scored using time priority scoring; the earlier the time or the closer it is to a key plot point, the higher the score. Action key point scoring projects the action key point onto the camera's view plane, calculates the Euclidean distance to the camera center, and maps it to a 0-1 score using a reciprocal or normalization function; the closer the distance, the higher the score. Focal length adaptation scoring calculates the distance difference between the action key point and the camera's focus, maps it to a scoring function (such as a Gaussian decay function), and the smaller the distance, the higher the score, reflecting the clarity of the action within the focus area.
[0083] Based on the calculated action performance tension score, the character's actions are adjusted, including scaling the action amplitude, increasing or decreasing the action speed, and advancing or delaying the action time. The tension-adjusted character actions are then combined with the time information of the master and slave tracks and the action trigger points to generate a logical frame sequence containing character posture information, action amplitude, speed, timestamp, and camera projection parameters.
[0084] The method for automatically retrieving action units from a preset action resource library based on logical frame sequences includes:
[0085] Different action units are stored in the preset action resource library. Each action unit corresponds to action type, action amplitude, speed curve, emotion label, spatial position change and adaptation scene conditions. The action units are mapped to emotion dimension, spatial dimension, scene dimension and rhythm dimension through a cross-dimensional index table.
[0086] For each frame in the logical frame sequence, action units are filtered in the preset action resource library according to the logical frame attributes, and a set of action units that meet the preset requirements in the dimensions of emotion, space, scene and rhythm is retained.
[0087] Methods for obtaining continuous action fragment streams include:
[0088] The cosine similarity method is used to quantify and match the attributes of logical frames with the attributes of action units to obtain the corresponding matching scores. According to the time order of the logical frame sequence, the action unit with the highest matching score is selected for each frame and arranged continuously in the time order of the logical frame sequence to form a continuous action segment stream.
[0089] Methods for performing automatic rendering binding on characters include:
[0090] Receive a continuous stream of motion clips, each containing key motion points, motion amplitude, speed curve, emotion markers, motion tendency markers, and scene parameters; load a character skeleton model for the virtual character, and map the key motion points in the motion clip stream to the corresponding nodes of the character skeleton joints, thus binding the motion to the character skeleton.
[0091] Based on the timestamp information in the continuous action fragment stream, each action fragment is synchronously bound to the character's skeletal model in chronological order, so that the action trigger point, duration, and action order are consistent with the logical frame sequence.
[0092] Interpolation and transition processing are performed on frames between consecutive action segments, including position interpolation, rotation interpolation, and speed smoothing, to eliminate action jumps and make the character's movements natural and smooth after rendering and binding; emotional markers and action expression tension information carried in the action segment stream are mapped to the character's action parameters.
[0093] The character's motion parameters include hand tension, limb curvature, and facial expression range, so that the character's movements not only conform to physical logic but also express the preset emotions and plot tension, ultimately completing the automatic rendering and binding of continuous motion segments onto the virtual character.
[0094] Methods for generating animation versions with different art styles include:
[0095] The virtual character motion frames that have been automatically rendered and bound are input into the preset style transfer rendering module. Based on the preset art style samples, the module extracts the content features of the character motion frames through a deep learning model, and maps and fuses the content features with the preset target style features to generate stylized animation frames that conform to the target art style.
[0096] Continuity optimization is performed on stylized animation frames, including position, rotation, speed, and smoothness processing, to eliminate motion frame skipping, flickering, or breakage caused by style transfer and ensure smooth animation.
[0097] Based on the rendering requirements of different platforms, stylized animation frames are exported into different standard animation file formats, so that the exported animation files can be directly called in various game engines, animation production software or virtual reality platforms, ultimately resulting in virtual character animation files with different art styles.
[0098] The preset priority score difference threshold is set by staff based on historical data analysis results. This historical analysis process includes the system collecting multiple priority score differences and calculating their average value as a reference to obtain the preset priority score difference threshold.
[0099] In this embodiment, by non-linearly mapping event priority scores to the plot timeline, high-priority events automatically align with dense rhythmic points, thus naturally reflecting the progression and release of plot tension and enhancing the narrative expressiveness of virtual reality animation. Multimodal information fusion improves the accuracy of event priority calculation, avoiding the limitations of single-dimensional driving, and making the generated results more in line with users' psychological expectations and immersive experience. Through a priority drift mechanism, the time interval is dynamically fine-tuned to avoid abrupt breaks in event intervals, achieving a smooth transition of events on the timeline, thereby ensuring the naturalness and fluency of the overall rhythm distribution. It can generate action event queues that conform to the plot logic and rhythmic rules in the virtual reality environment, ensuring that character actions, music, voice, and plot development maintain a unified rhythmic atmosphere, enhancing the narrative appeal and immersive experience of the virtual reality scene.
[0100] By calculating time priority scores for action trigger points, considering the start time, end time, duration, and key rhythmic points of actions, the triggering order of character actions automatically adapts to the rhythm of the plot, ensuring smooth and natural animation and enhancing immersion in virtual reality scenes. Actions with the highest priority scores are designated as primary track actions, generating core character actions; other actions are designated as secondary track actions, generating supplementary actions, achieving a clear hierarchy of action performance, highlighting plot points, and ensuring clear logic in character actions. Primary track actions, combined with action tendency markers, generate core actions; secondary track actions, combined with emotional color markers, generate supplementary actions, enabling actions to respond to multimodal creative information, making character actions more aligned with the developer's intentions and enriching emotional expression. Separate generation strategies are established for primary and secondary track actions, enabling automatic adjustment and dynamic supplementation of actions, reducing manual adjustments by animators, improving the efficiency of virtual reality animation production, and ensuring a high degree of consistency between actions, plot, and emotions. Through time priority scoring, primary / secondary track division, and multimodal information-driven action generation strategies, character actions are highly matched to the plot in time, space, and emotion, enhancing the user's realistic experience and immersion in the virtual reality environment.
[0101] Example 2
[0102] Please see Figure 2 As shown, parts not described in detail in this embodiment are described in Embodiment 1. A multimodal driven character animation automatic generation system for virtual reality is provided, including:
[0103] The creative intent capture module acquires multimodal creative information input by developers and transforms it into a stream of creative instruction symbols containing plot rhythm markers, emotional color markers, and action tendency markers through a contextual intent parsing mechanism.
[0104] The narrative event construction module decomposes the creative instruction symbol flow into different independent event units, and assigns the independent event units to rhythmic grid points on the timeline through plot node mapping and priority drift rules, forming a rhythmic event queue.
[0105] The character animation generation module generates character actions based on the event queue through a master-slave dual-track logic. It introduces a camera perspective projection mechanism to automatically adjust the expressive tension of the character actions according to scene changes, generating a logical frame sequence with a cinematic feel.
[0106] The action resource retrieval module automatically retrieves action units from the preset action resource library based on the logical frame sequence, matches the action units with the emotion dimension, spatial dimension and scene dimension through a cross-dimensional index table, and outputs a continuous action fragment stream.
[0107] The rendering and export module receives motion clip streams, performs automatic rendering and binding on the characters, adopts a style transfer rendering mechanism to generate animation versions with different art styles, and exports them as finished animation files that can be directly used on different production platforms.
[0108] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.
[0109] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A method for automatically generating multimodal driven character animations for virtual reality, characterized in that, include: S1. Obtain multimodal creation information input by the developer, and transform the multimodal creation information into a creation instruction symbol stream containing plot rhythm markers, emotional color markers, and action tendency markers through the contextual intent parsing mechanism; S2. Decompose the creative instruction symbol flow into different independent event units, and through plot node mapping and priority drift rules, assign the independent event units to rhythmic grid points on the timeline to form a rhythmic event queue. S3. Based on the event queue, generate character actions through master-slave dual-track logic, introduce a camera perspective projection mechanism, automatically adjust the expressive tension of character actions according to scene changes, and generate a logical frame sequence with a camera feel. S4. Automatically retrieve action units from the preset action resource library based on the logical frame sequence, match the action units with the emotion dimension, spatial dimension and scene dimension through a cross-dimensional index table, and output a continuous action fragment stream. S5 receives the motion clip stream and performs automatic rendering binding on the character. It adopts a style transfer rendering mechanism to generate animation versions with different art styles and exports them as finished animation files that can be directly called on different production platforms.
2. The method for automatically generating multimodal driven character animations for virtual reality according to claim 1, characterized in that, The method for obtaining the creation instruction symbol stream includes: The system acquires multimodal creation information input by developers, including voice description information, sketch action information, plot script text, character position change information, and background music information; it normalizes the timestamps of each modality and aligns them using a dynamic time warping method to form a multimodal feature sequence on a unified timeline; Using an attention-weighted fusion method, the features of each modality are integrated according to their contextual relevance to generate a unified context vector. Based on the music beat, text event interval, and speech rate, music beat features, text event interval features, and speech rate features are extracted from the unified context vector and generated by weighted combination. Emotional analysis is performed on text, speech, and music modalities. Emotional features from each modality are fused using adjustable weights to generate standardized emotion vectors, which serve as emotion color markers. Semantic description information, sketch action information, and character position change information are mapped to action probability distributions to form action tendency markers. Plot rhythm markers, emotion color markers, and action tendency markers are integrated to form a unified creative instruction symbol flow.
3. The method for automatically generating multimodal driven character animations for virtual reality according to claim 2, characterized in that, The method for forming a rhythmic event queue includes: The creation instruction symbol flow is decomposed into different independent event units. For each independent event unit, an event priority score is calculated based on the importance of plot nodes, emotional intensity, and time priority. The event priority score is then non-linearly mapped to the plot timeline, with events with higher priority scores automatically mapped to more densely packed rhythm points. When the difference in priority scores between adjacent events exceeds a preset priority score difference threshold, a priority drift mechanism is used to adjust the time interval between events, so that the event units form a rhythmic arrangement on the time axis, resulting in a rhythmic event queue.
4. The method for automatically generating multimodal driven character animations for virtual reality according to claim 3, characterized in that, The method for generating character actions through master-slave dual-track logic includes: In the received event queue, each independent event unit is parsed into an action trigger point. Each action trigger point includes the start timestamp of the action, the end timestamp of the action, and the duration of the action. A time priority score is calculated based on how close the trigger time of the action is to the current time. Action trigger points are sorted according to time priority scores. Actions with the highest priority scores are classified as main track actions, and actions with other priority scores are classified as secondary track actions. Generation strategies are established for main track actions and secondary track actions respectively. Main track actions generate the character's core actions based on the action time sequence, plot node mapping, and action tendency markers. Secondary track actions generate the character's supplementary actions based on the time offset of the main track actions and the emotional color markers.
5. The method for automatically generating multimodal driven character animations for virtual reality according to claim 4, characterized in that, The method for generating a logical frame sequence with a cinematic feel includes: The system obtains the current camera's spatial position, orientation information, focal length, and motion trajectory in a preset virtual scene, and calculates the projection parameters of the action from the camera's perspective based on the relative spatial relationship between the character and the camera. The projection parameters include the viewing angle, visible area, and depth of field. For each main track action and secondary track action, a motion performance tension score is calculated based on the motion amplitude, speed, position of the motion trigger point, and focal length adaptation. This score measures the visual impact and emotional expression intensity of the action from the current camera perspective. The motion performance tension score is obtained by quantifying the motion amplitude, motion speed, position of the motion trigger point on the timeline, key points of the motion, and focal length adaptation into independent scores, and then weighting and summing them using preset weights to form a comprehensive score. This comprehensive score is used as the motion performance tension score. Based on the calculated action performance tension score, the character's actions are adjusted, including scaling the action amplitude, increasing or decreasing the action speed, and advancing or delaying the action time. The tension-adjusted character actions are then combined with the time information of the master and slave tracks and the action trigger points to generate a logical frame sequence containing character posture information, action amplitude, speed, timestamp, and camera projection parameters.
6. The method for automatically generating multimodal driven character animations for virtual reality according to claim 5, characterized in that, The method for automatically retrieving action units from a preset action resource library based on logical frame sequences includes: Different action units are stored in the preset action resource library. Each action unit corresponds to action type, action amplitude, speed curve, emotion label, spatial position change and adaptation scene conditions. The action units are mapped to emotion dimension, spatial dimension, scene dimension and rhythm dimension through a cross-dimensional index table. For each frame in the logical frame sequence, action units are filtered in the preset action resource library according to the logical frame attributes, and a set of action units that meet the preset requirements in the dimensions of emotion, space, scene and rhythm is retained.
7. The method for automatically generating multimodal driven character animations for virtual reality according to claim 6, characterized in that, The method for obtaining the continuous action segment stream includes: The cosine similarity method is used to quantify and match the attributes of logical frames with the attributes of action units to obtain the corresponding matching scores. According to the time order of the logical frame sequence, the action unit with the highest matching score is selected for each frame and arranged continuously in the time order of the logical frame sequence to form a continuous action segment stream.
8. The method for automatically generating multimodal driven character animations for virtual reality according to claim 7, characterized in that, The method for performing automatic rendering binding on the character includes: Receive a continuous stream of motion clips, each containing key motion points, motion amplitude, speed curve, emotion markers, motion tendency markers, and scene parameters; load a character skeleton model for the virtual character, and map the key motion points in the motion clip stream to the corresponding nodes of the character skeleton joints, thus binding the motion to the character skeleton. Based on the timestamp information in the continuous action segment stream, each action segment is synchronously bound to the character's skeletal model in chronological order, ensuring that the action trigger point, duration, and action order are consistent with the logical frame sequence; interpolation and transition processing are performed on the frames between continuous action segments, including position interpolation, rotation interpolation, and speed smoothing; and the emotion markers and action performance tension information carried in the action segment stream are mapped to the character's action parameters. The character's motion parameters include hand tension, limb curvature, and facial expression range, so that the character's movements not only conform to physical logic but also express the preset emotions and plot tension, ultimately completing the automatic rendering and binding of continuous motion segments onto the virtual character.
9. The method for automatically generating multimodal driven character animations for virtual reality according to claim 8, characterized in that, The method for generating animation versions with different art styles includes: The virtual character motion frames that have been automatically rendered and bound are input into the preset style transfer rendering module. Based on the preset art style samples, the module extracts the content features of the character motion frames through a deep learning model, and maps and fuses the content features with the preset target style features to generate stylized animation frames that conform to the target art style. The stylized animation frames are continuously optimized, including position, rotation, speed, and smoothness processing. Based on the rendering requirements of different platforms, the stylized animation frames are exported to different standard animation file formats, so that the exported animation files can be directly called in various game engines, animation production software, or virtual reality platforms, ultimately resulting in virtual character animation files with different art styles.
10. A multimodal driven character animation automatic generation system for virtual reality, used to implement the multimodal driven character animation automatic generation method for virtual reality as described in any one of claims 1 to 9, characterized in that, include: The creative intent capture module acquires multimodal creative information input by developers and transforms it into a stream of creative instruction symbols containing plot rhythm markers, emotional color markers, and action tendency markers through a contextual intent parsing mechanism. The narrative event construction module decomposes the creative instruction symbol flow into different independent event units, and assigns the independent event units to rhythmic grid points on the timeline through plot node mapping and priority drift rules, forming a rhythmic event queue. The character animation generation module generates character actions based on the event queue through a master-slave dual-track logic. It introduces a camera perspective projection mechanism to automatically adjust the expressive tension of the character actions according to scene changes, generating a logical frame sequence with a cinematic feel. The action resource retrieval module automatically retrieves action units from the preset action resource library based on the logical frame sequence, matches the action units with the emotion dimension, spatial dimension and scene dimension through a cross-dimensional index table, and outputs a continuous action fragment stream. The rendering and export module receives motion clip streams, performs automatic rendering and binding on the characters, adopts a style transfer rendering mechanism to generate animation versions with different art styles, and exports them as finished animation files that can be directly used on different production platforms.