Digital figure action dynamic adjustment method based on reinforcement learning

By integrating improved hierarchical reinforcement learning with the Koopman linearized subspace control model, the problems of semantic sub-target modeling and hierarchical strategy in digital avatar action control are solved, realizing personalized, natural and smooth action control and improving the interactive capabilities of digital avatars in multiple scenarios.

CN121168569BActive Publication Date: 2026-02-10JIANGSU ELECTRIC POWER INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511704829.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-20
Publication Date
2026-02-10
Estimated Expiration
2045-11-20

AI Technical Summary

Technical Problem

Existing technologies for digital avatar motion control suffer from problems such as insufficient semantic sub-target modeling capabilities, difficulty in expressing differentiated user intentions through hierarchical strategy structures, lack of controllability and interpretability in the control phase, and loose strategy update mechanisms. These issues fail to meet the demands for high-quality, personalized, and natural motion control in human-computer interaction.

Method used

By integrating an improved hierarchical reinforcement learning structure with the Koopman linearized subspace control model, semantic sub-target vectors are generated through semantic parsing. High- and low-level policy networks, style regulation mechanisms, graph-structured action skill libraries, and controllable action trajectory optimization control models are constructed to achieve personalized action generation and dynamic adjustment.

Benefits of technology

It enhances the digital avatar's ability to express actions and adapt in multiple scenarios and user interactions. It has strong semantic understanding capabilities, high style adaptability, natural action transitions, high trajectory control precision, and strong strategy adaptive update capabilities, making it suitable for interactive applications such as virtual human interaction, virtual customer service, and immersive digital scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121168569B_ABST
    Figure CN121168569B_ABST
Patent Text Reader

Abstract

The application discloses a digital image action dynamic adjustment method based on reinforcement learning, which comprises the following steps: collecting multi-source interaction data between a user and a digital image, performing semantic analysis, and generating a semantic sub-target vector; constructing an improved hierarchical reinforcement learning structure according to a score value, generating a sub-target intention embedding vector, inputting the vector into a style regulator and a low-layer action strategy network, and outputting an action trajectory reference path; dynamically adjusting a graph structure action skill library by using a graph attention mechanism, and generating an action combination candidate graph; constructing a Koopman linearization subspace control model, performing subspace trajectory optimization, and outputting a continuous action execution instruction; updating the control model and the strategy network according to environment feedback data, and realizing strategy iteration optimization; and finally packaging key modules into a digital image behavior package. The application improves the real-time performance, continuity and semantic consistency of digital image action adjustment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, and in particular relates to a method for dynamically adjusting the actions of digital images based on reinforcement learning. Background Technology

[0002] With the widespread application of virtual humans and digital avatars in emerging scenarios such as social media platforms, online education, virtual customer service, digital marketing, and the metaverse, the intelligent behavior and natural expression of digital avatars are gradually becoming core factors influencing user interaction experience. Traditional digital avatar action control mainly relies on preset action libraries and fixed behavior scheduling logic, usually adopting rule-driven or template matching methods. This makes it difficult to effectively adapt to the semantic differences in user commands, individual preferences, and changing interaction contexts, resulting in repetitive actions, rigid performance, and a lack of flexibility, failing to meet the needs of humanized and emotional interaction.

[0003] In recent years, reinforcement learning has achieved a series of breakthroughs in robot control and intelligent policy generation, providing new ideas for the generation and dynamic adjustment of digital avatars' actions. By constructing policy networks based on interactive rewards, reinforcement learning methods enable systems to continuously optimize their behavioral policies through environmental feedback, forming strong adaptability and generalization capabilities. However, directly applying traditional reinforcement learning structures to the action adjustment of digital avatars still faces multiple challenges. First, the action adjustment tasks of digital avatars typically have goal decomposition and hierarchical decision-making characteristics. Relying solely on a single-layer policy network is insufficient to model complex task structures, easily leading to slow policy convergence or getting trapped in local optima. Second, user input often manifests as multimodal data such as natural language, emotional states, and behavioral responses. There are granularity differences and semantic ambiguities between different modalities, making it difficult for traditional policy modeling methods to achieve semantic consistency parsing and action mapping. Furthermore, the execution of actions by digital avatars inherently possesses continuous control characteristics and physical consistency requirements. Reinforcement learning policies often lack controllable state evolution mechanisms, resulting in inconsistent, unstable, or distorted generated actions.

[0004] Against this backdrop, some studies have attempted to apply hierarchical reinforcement learning to digital human motion control scenarios. By introducing a hierarchical structure of high-level task scheduling and low-level action execution, they aim to improve policy expressiveness and task adaptability. High-level policies are responsible for generating sub-goals at the intermediate or semantic levels, while low-level policies execute specific action sequences based on these sub-goals. Although this method alleviates the task complexity problem to some extent, in existing structures, the generation of sub-goals often relies on manual setting or fixed decomposition logic, lacking the ability to dynamically model user intent and failing to effectively address the semantic ambiguity and polysemy inherent in natural language. Furthermore, the collaborative policy update mechanism between levels is still immature, and the training process is easily affected by high-dimensional action spaces, resulting in instability and ultimately failing to meet the requirements of realistic natural expression in terms of control performance.

[0005] On the other hand, trajectory control is also a core issue restricting the natural expression of digital avatars. During the execution of digital avatar actions, the system must not only consider the mapping between semantic instructions and action intentions, but also ensure that the actual execution trajectory conforms to human cognition in terms of temporal continuity, posture smoothness, and physical accessibility. Traditional reinforcement learning faces modeling difficulties in nonlinear system control, struggling to effectively handle potential structural constraints and dynamic coupling relationships in the action state space. Therefore, how to combine highly interpretable and robust modeling control methods to compensate for the instability in the control phase of traditional reinforcement learning has become an important research direction.

[0006] Koopman operator theory provides strong support for solving control problems of such nonlinear systems. By constructing a state mapping function to elevate the original nonlinear system to a high-dimensional subspace, linear modeling and control are achieved. The Koopman method can provide a predictable and optimizable control structure while preserving the dynamic characteristics of the original system. Existing research has applied the Koopman model to fields such as robot control and traffic system optimization, achieving certain results. However, a systematic modeling framework is still lacking in the dynamic motion control of digital avatars. Especially in interactive systems driven by multiple objectives and involving both semantic commands and user behavior, how to construct a goal-oriented Koopman control model and deeply integrate it with reinforcement learning structures to form a closed-loop feedback mechanism between action intent parsing, trajectory optimization, and policy update remains a challenging and crucial issue.

[0007] In summary, existing technologies for dynamically adjusting the actions of digital avatars suffer from several shortcomings. These include insufficient semantic sub-target modeling capabilities, difficulty in expressing differentiated user intentions through hierarchical policy structures, a lack of controllability and interpretability in the control phase, and a loose policy update mechanism. Consequently, these technologies fail to meet the demands for high-quality, personalized, and natural-smooth action control in human-computer interaction. Therefore, there is an urgent need to propose a novel approach that integrates improved hierarchical reinforcement learning with Koopman linearized subspace control. This approach would establish a complete technical path encompassing semantic target-driven mechanisms, personalized style adjustment, motion graph generation, and trajectory optimization control, thereby enhancing the dynamic action expression and adaptability of digital avatars across multiple scenarios, users, and tasks. Summary of the Invention

[0008] To address the problems of existing technologies, this invention proposes a dynamic adjustment method for digital avatar actions based on reinforcement learning. This method integrates an improved hierarchical reinforcement learning structure with a Koopman linearized subspace control model. Addressing the issues of lack of semantic understanding, style control, and poor trajectory continuity in existing digital avatar actions, it constructs a high- and low-level policy network driven by semantic sub-goals, a style control mechanism, a graph-structured action skill library, and a controllable action trajectory optimization model. This systematically achieves personalized action generation and dynamic adjustment based on user intent. This method possesses advantages such as strong semantic understanding, high style adaptability, natural action transitions, high trajectory control accuracy, and strong policy adaptive update capabilities, making it suitable for various interactive application scenarios such as virtual human interaction, virtual customer service, and immersive digital scenes.

[0009] The technical solution of the present invention is as follows:

[0010] A method for dynamic adjustment of digital avatar actions based on reinforcement learning includes:

[0011] Collect multi-source interaction data between users and digital avatars, perform semantic parsing on the multi-source interaction data, generate semantic sub-target vectors, and calculate target granularity scores based on the semantic sub-target vectors;

[0012] An improved hierarchical reinforcement learning structure is constructed based on the target granularity score. The improved hierarchical reinforcement learning structure includes a high-level policy network, a style regulator, and a low-level action policy network. The high-level policy network generates a sub-target intent embedding vector, and the style regulator generates a style regulation vector. The sub-target intent embedding vector and the style regulation vector are input into the low-level action policy network to output an action trajectory reference path.

[0013] A graph-structured action skill library is constructed. Based on the semantic sub-target vector and the action trajectory reference path, the adjacency structure of the graph-structured action skill library is dynamically adjusted through a graph attention mechanism to generate a candidate graph of action combinations.

[0014] Construct a Koopman linearized subspace control model, establish an observable mapping function using training data, and obtain a linear system model in the Koopman subspace based on the subspace state vector and control input.

[0015] Based on the motion trajectory reference path and motion combination candidate map, trajectory optimization is performed in the Koopman subspace to generate the optimal subspace motion control sequence. The sequence is then input into the decoder and mapped back to the original motion state space, outputting a digital image of continuous motion execution instructions.

[0016] Collect environmental feedback data after the execution of continuous action execution instructions, update parameters based on environmental feedback data, and simultaneously update the graph structure action skill library and linear system model to achieve iterative optimization of the strategy of the improved hierarchical reinforcement learning structure;

[0017] Save the updated high-level strategy network, style regulator, low-level action strategy network, graph-structured action skill library, and Koopman linearized subspace control model as a digital character behavior package.

[0018] Furthermore, the process of collecting multi-source interaction data between users and digital avatars, performing semantic parsing on the multi-source interaction data to generate semantic sub-target vectors, and calculating target granularity scores based on the semantic sub-target vectors, specifically involves:

[0019] Collect multi-source interaction data between users and digital avatars, including natural language commands, scene tags, user behavior data, and historical response records;

[0020] Perform unified formatting on multi-source interactive data to construct semantic parsing input vectors;

[0021] The semantic parsing input vector is fed into the semantic parsing network to extract contextual semantic features and generate semantic sub-target vectors;

[0022] The target granularity score is calculated based on the semantic complexity, instruction abstraction, and behavioral expectation clarity in the semantic sub-target vector.

[0023] Furthermore, the high-level policy network in the improved hierarchical reinforcement learning structure receives semantic sub-target vectors and target granularity scores as inputs, and generates sub-target intent embedding vectors using a dynamic hierarchical mechanism. The dynamic hierarchical mechanism adaptively adjusts the number of sub-targets and the abstraction level of the high-level policy output according to the target granularity scores.

[0024] The style modifier receives the sub-target intent embedding vector, user personality parameters, and contextual emotion labels. It uses a fusion mapping network to generate a style control vector. The fusion mapping network uses an attention mechanism to weight and combine personality and emotion factors to achieve personalized adjustment of action style and behavioral trends.

[0025] The low-level action policy network receives the concatenated features of the sub-target intent embedding vector and style control vector as input, and outputs a continuous action trajectory reference path based on the conditional policy generation structure. The action trajectory reference path is composed of target state points at multiple time steps, which guides the generation of action combinations.

[0026] Furthermore, the improved hierarchical reinforcement learning structure is constructed based on the target granularity score. This improved hierarchical reinforcement learning structure includes a high-level policy network, a style regulator, and a low-level action policy network. The high-level policy network generates a sub-target intent embedding vector, and the style regulator generates a style regulation vector. The sub-target intent embedding vector and the style regulation vector are input together into the low-level action policy network to output a motion trajectory reference path. Specifically:

[0027] Input the semantic sub-target vector into the high-level policy network, and output the sub-target intent embedding vector;

[0028] User personality parameters are generated based on the user behavior data and historical response records. These user personality parameters include action style preferences, response rhythm tendencies, and risk tolerance characteristics.

[0029] By combining natural language instructions and user behavior feedback in the current interaction round, a contextual emotion label is constructed, which reflects the emotional orientation and intensity of the current interaction state.

[0030] The style modifier is input by embedding sub-target intent into a vector, user personality parameters, and upper and lower emotion tags to generate a style modifier vector.

[0031] The sub-target intent embedding vector and style control vector are concatenated to form a fused feature representation, which is then input into the low-level action policy network and outputs an action trajectory reference path.

[0032] Furthermore, the construction of the graph-structured action skill library involves dynamically adjusting the adjacency structure of the graph-structured action skill library based on the semantic sub-target vector and the action trajectory reference path using a graph attention mechanism to generate a candidate graph of action combinations, specifically as follows:

[0033] A graph-structured action skill library is constructed based on a predefined set of atomic action skills. Each atomic action skill is represented as a node in the graph, and the state transition probability between atomic actions is calculated based on historical behavior data to construct a set of edges between nodes.

[0034] Initialize the graph-structured action skill library to form an adjacency matrix with associated weights, which describes the transition strength and reachability relationship between action skills;

[0035] The semantic sub-target vector and the action trajectory reference path are used as joint control information and input into the graph attention mechanism module to dynamically update the adjacency matrix and adjust the edge weights to highlight the action transfer path related to the current semantic target.

[0036] Based on the updated graph-structured action skill library, a combination search is performed on candidate action nodes to generate an action combination candidate graph, providing optional skill sequences.

[0037] Furthermore, the construction of the Koopman linearized subspace control model involves establishing an observable mapping function using training data, and training a linear system model in the Koopman subspace based on the subspace state vector and control input. Specifically:

[0038] A goal-oriented Koopman linearized subspace control model is constructed, which includes an observable mapping function module and a goal-enhanced subspace linear system modeling module.

[0039] The original action state vector sequence, control input sequence and corresponding semantic sub-target vector sequence are collected. The original action state vector is input into the observable mapping function module composed of a multi-layer neural network to generate subspace state vectors.

[0040] The subspace state vector, control input sequence, and semantic subtarget vector are input into the target-enhanced subspace linear modeling module to fit and construct a linear system including a state transition matrix, a control input mapping matrix, and a target guidance mapping matrix;

[0041] The model is trained and evaluated based on trajectory reconstruction error, semantic response consistency index and system stability constraints. When each performance index meets the preset threshold, the parameters of the final goal-oriented Koopman linearized subspace control model are determined and solidified into the subspace prediction model for the motion trajectory optimization stage.

[0042] Furthermore, the process of performing trajectory optimization in the Koopman subspace based on the action trajectory reference path and action combination candidate map to generate the optimal subspace action control sequence, inputting it into the decoder for inverse mapping to the original action state space, and outputting a digitally visualized continuous action execution instruction, specifically includes:

[0043] Based on the Koopman linearized subspace control model, the motion trajectory reference path and motion combination candidate map are mapped to the corresponding subspace state vector sequence and control input sequence.

[0044] In the Koopman linearized subspace control model, the optimal control trajectory is solved to generate the optimal subspace motion control sequence with the goal of minimizing the trajectory offset loss and control cost function.

[0045] The optimal subspace control sequence and the corresponding sub-objective vector are input into the Koopman linearized subspace control model to predict the optimal subspace state sequence.

[0046] The trained decoder function is used to reverse map the optimal subspace state sequence to restore the action state sequence in the original action state space, and generate a digital image of continuous action execution instructions.

[0047] Furthermore, the step of collecting environmental feedback data after the execution of continuous action execution instructions, updating parameters based on the environmental feedback data, and simultaneously updating the graph-structured action skill base and the linear system model to achieve iterative optimization of the improved hierarchical reinforcement learning structure specifically involves:

[0048] Collect environmental feedback data after the execution of the continuous action execution instructions. The environmental feedback data includes user response characteristics, behavior result status and action naturalness index.

[0049] An experience playback dataset is jointly constructed by environmental feedback data and multi-source interaction data. The network parameters of the high-level policy network, style regulator and low-level action policy network are updated respectively. Reinforcement learning policy iterative training is performed by policy gradient optimization method.

[0050] Based on the motion deviation and trajectory response error observed in the environmental feedback data, the node transfer weights and adjacency structure in the graph-structured motion skill library are adaptively adjusted by adopting the minimum reconstruction error criterion.

[0051] By simultaneously utilizing the triplet samples of state-control input-target vector in the latest environmental feedback samples, the state transition matrix, control mapping matrix, and target guidance matrix of the Koopman linearized subspace control model are incrementally updated, thereby improving the control model's adaptability to changes in user intent and achieving dynamic optimization of the improved hierarchical reinforcement learning structure.

[0052] Furthermore, the step of saving the updated high-level policy network, style regulator, low-level action policy network, graph-structured action skill library, and Koopman linearized subspace control model as a digital avatar behavior package specifically involves:

[0053] The updated high-level strategy network, style regulator, low-level action strategy network, graph structure action skill library and Koopman linearized subspace control model are uniformly encapsulated to construct a digital image behavior package.

[0054] The configuration parameters, training weights, structural topology, and input / output interfaces of each module in the digital image behavior package are archived and encoded to form a standardized model file that can be loaded and deployed.

[0055] The digital avatar behavior package is loaded into the new interactive scenario, and new multi-source interactive data input is received. The action strategy reasoning and control execution are completed through an improved hierarchical reinforcement learning structure, supporting the dynamic action adjustment task of the digital avatar in cross-scenario scenarios.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] This invention constructs a method for dynamically adjusting the actions of digital images by integrating an improved hierarchical reinforcement learning structure and a Koopman linearized subspace control model. This method overcomes multiple bottlenecks in traditional digital interactive systems in terms of action generation continuity, accuracy of personalized expression, semantic intent adaptability, and execution control stability. It achieves dynamic action planning, optimization, and adaptive execution that caters to users' personalized needs and contextual changes, and possesses high intelligence and wide adaptability.

[0058] First, this invention makes key innovations in semantic understanding and action generation driving mechanisms. By collecting multi-source interaction data between users and digital avatars, including natural language commands, scene tags, user behavior data, and historical response records, and integrating semantic parsing and target-granularity scoring mechanisms, it effectively transforms complex, multimodal user inputs into structured semantic sub-target vectors. These semantic sub-target vectors not only clarify the target intent of digital avatar action planning but also provide a refined and continuous target-driven basis for subsequent hierarchical strategy modeling, solving the problems of coarse target parsing and unclear intent communication in traditional methods.

[0059] Secondly, this invention proposes an improved hierarchical reinforcement learning method with an innovative structure. The high-level policy network generates sub-target intent embeddings based on semantic sub-target vectors, and combines this with a style modifier constructed from user history, personality parameters, and contextual emotion labels to achieve individualized style control. The low-level action policy network integrates intent embeddings and style modulation vectors, outputting continuous action trajectory reference paths, achieving coordinated adaptation between the policy execution level and the expressive style. This hierarchical structure not only decouples the target levels but also enhances the controllability and interpretability of policy generation, effectively addressing the problems of style mutations and policy confusion in dynamic interactions across multiple scenarios.

[0060] Furthermore, this invention employs a graph-structured action skill library design for action skill modeling and combination mechanisms. Atomic action skills are modeled as graph nodes, and transition probabilities between edges are constructed using historical behavior data. Adjacency structures are adjusted in real-time based on a graph attention mechanism, thereby building a dynamic action combination candidate graph. This structure fully leverages the accumulated experience from historical execution data, achieving structured, modular, and flexible action generation. This allows digital images to quickly combine semantically consistent and naturally connected action sequences driven by different semantic sub-goals.

[0061] Furthermore, this invention introduces a Koopman linearized subspace control model to address the problems of high computational complexity and low trajectory control accuracy in traditional motion control models during nonlinear dynamic process modeling. By constructing an observable mapping function, the original motion state is embedded into the Koopman subspace. Within this subspace, a linear dynamic system is used to model the motion control strategy, perform trajectory optimization and control input inference, and finally, a decoder maps the result back to the original motion space, achieving a precise mapping from control input to physical motion. This method significantly improves the computational efficiency and execution stability of the system in complex nonlinear control scenarios and possesses good engineering deployability.

[0062] Finally, regarding strategy optimization, this invention collects environmental feedback data after the execution of continuous action commands, including user response characteristics, behavioral outcome status, and action naturalness indicators, and dynamically updates the parameters of the high-level policy network, style regulator, low-level action policy network, and Koopman control model to form an adaptive closed-loop policy update mechanism. This mechanism achieves the co-evolution of the policy and control models, not only continuously improving interaction quality based on user feedback but also possessing continuous learning and transfer capabilities, significantly enhancing its ability to adapt to new scenarios and new users. Attached Figure Description

[0063] Figure 1 This is a flowchart illustrating a method for dynamically adjusting the actions of a digital avatar based on reinforcement learning.

[0064] Figure 2 This is a schematic diagram of an improved hierarchical reinforcement learning structure. Detailed Implementation

[0065] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading this invention, any modifications of the invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.

[0066] Example 1:

[0067] This invention provides a method for dynamically adjusting the actions of digital avatars based on reinforcement learning, such as... Figure 1 and Figure 2 As shown, it includes:

[0068] Step 1) Collect multi-source interaction data between users and digital avatars, perform semantic parsing on the multi-source interaction data, generate semantic sub-target vectors, and calculate target granularity scores based on the semantic sub-target vectors;

[0069] Step 2) Construct an improved hierarchical reinforcement learning structure based on the target granularity score. The improved hierarchical reinforcement learning structure includes a high-level policy network, a style regulator, and a low-level action policy network. The high-level policy network generates a sub-target intent embedding vector, and the style regulator generates a style regulation vector. The sub-target intent embedding vector and the style regulation vector are input into the low-level action policy network to output the action trajectory reference path.

[0070] Step 3) Construct a graph-structured action skill library. Based on the semantic sub-target vector and the action trajectory reference path, dynamically adjust the adjacency structure of the graph-structured action skill library through a graph attention mechanism to generate a candidate graph of action combinations.

[0071] Step 4) Construct a Koopman linearized subspace control model, establish an observable mapping function using training data, and train a linear system model in the Koopman subspace based on the subspace state vector and control input.

[0072] Step 5) Perform trajectory optimization in the Koopman subspace based on the motion trajectory reference path and motion combination candidate map to generate the optimal subspace motion control sequence, input it into the decoder to back-map to the original motion state space, and output a digital image of continuous motion execution instructions.

[0073] Step 6) Collect environmental feedback data after the execution of continuous action execution instructions, update parameters based on environmental feedback data, and simultaneously update the graph structure action skill library and linear system model to achieve iterative optimization of the improved hierarchical reinforcement learning structure;

[0074] Step 7) Save the updated high-level policy network, style regulator, low-level action policy network, graph structure action skill library, and Koopman linearized subspace control model as a digital image behavior package.

[0075] This invention collects multi-source interaction data between users and digital avatars and performs semantic parsing to accurately extract semantic sub-target vectors. It dynamically adjusts the target abstraction level based on target granularity scores, achieving refined modeling and context awareness of user intent. An improved hierarchical reinforcement learning structure, combined with a style regulator, models user personality and emotional state, enabling personalized adjustment and natural expression of action styles, improving the coordination and emotional fit of interactive responses. By dynamically updating the adjacency structure of the graph-structured action skill base through a graph attention mechanism, it flexibly adjusts action combination paths based on the current semantic target and trajectory reference path, enhancing the diversity and adaptability of complex action planning. The Koopman linearized subspace control model introduces observable mapping and linear modeling mechanisms, effectively improving computational efficiency and trajectory generation accuracy in nonlinear behavior control tasks. Through trajectory optimization and inverse mapping, it outputs highly continuous and consistent digital avatar action commands. Iterative optimization of the policy network and control model using environmental feedback data constructs an adaptive reinforcement learning closed-loop system, enabling dynamic adjustment and generalization enhancement of the policy in long-term interactions. Ultimately, by encapsulating the trained core model modules into a digital avatar behavior package, cross-scenario deployment and strategy migration are supported, significantly improving the practicality, scalability, and intelligence of the digital avatar system.

[0076] In one embodiment, multi-source interaction data between users and digital avatars is collected, semantic parsing is performed on the multi-source interaction data to generate semantic sub-target vectors, and target granularity scores are calculated based on the semantic sub-target vectors, specifically as follows:

[0077] Collect multi-source interaction data between users and digital avatars. The multi-source interaction data includes natural language commands, scene tags, user behavior data, and historical response records.

[0078] Perform unified formatting on multi-source interactive data to construct semantic parsing input vectors;

[0079] The semantic parsing input vector is fed into the semantic parsing network to extract contextual semantic features and generate semantic sub-target vectors;

[0080] The target granularity score is calculated based on the semantic complexity, instruction abstraction, and behavioral expectation clarity in the semantic sub-target vector.

[0081] By collecting multi-source interaction data, including natural language commands, scene labels, user behavior data, and historical response records, the system can comprehensively capture the interaction context between users and digital avatars, enhancing the breadth and depth of semantic understanding. Unified formatting of multimodal data and construction of semantic parsing input vectors enable the semantic parsing network to fully integrate various data types, improving the consistency and accuracy of semantic modeling. The contextual semantic features extracted by the semantic parsing network possess stronger command comprehension and scene association capabilities, helping to generate semantic sub-target vectors that better reflect actual intentions. By introducing multi-dimensional indicators such as semantic complexity, command abstraction, and behavioral expectation clarity to evaluate semantic sub-target vectors, quantitative modeling of target granularity can be achieved. This allows the system to adaptively adjust target splitting strategies, providing more precise target guidance and abstraction level support for subsequent hierarchical strategy generation and action control in reinforcement learning structures. This significantly improves the digital avatar's response flexibility and task completion rate in complex and ever-changing interaction scenarios.

[0082] In one embodiment, the high-level policy network in the improved hierarchical reinforcement learning structure receives semantic sub-target vectors and target granularity scores as inputs, and generates sub-target intent embedding vectors using a dynamic hierarchical mechanism. The dynamic hierarchical mechanism adaptively adjusts the number of sub-targets and the abstraction level of the high-level policy output according to the target granularity scores.

[0083] The style modifier receives the sub-target intent embedding vector, user personality parameters, and contextual emotion labels. It uses a fusion mapping network to generate style control vectors. The fusion mapping network uses an attention mechanism to weight and combine personality and emotion factors to achieve personalized adjustment of action style and behavioral trends.

[0084] The low-level action policy network receives the concatenated features of the sub-target intent embedding vector and style modulation vector as input. Based on the conditional policy, it generates a continuous action trajectory reference path, which is composed of target state points at multiple time steps, guiding the generation of action combinations.

[0085] The introduction of an improved hierarchical reinforcement learning structure plays a crucial role in enhancing the dynamic adjustment process of digital avatar actions. First, the high-level policy network, by combining semantic sub-target vectors with target granularity scores, introduces a dynamic hierarchical mechanism. This allows the policy output to adaptively adjust the abstraction level and number of sub-targets, thereby enhancing the understanding of targets with varying semantic complexity and significantly improving the system's response accuracy and understanding precision when handling ambiguous expressions or nested multi-target instructions. Second, the style regulator, through a fusion mapping network, weights the sub-target intent embedding vector with user personality parameters and contextual emotion labels. This introduces personalization and emotion adjustment mechanisms before action generation, enabling the system-generated actions to reflect more natural, vivid, and user-preferred style characteristics, enhancing the digital avatar's personalized expressiveness and interactive affinity. Furthermore, the low-level action policy network, after receiving the fused feature input, generates a structured action trajectory reference path based on the conditional policy structure, ensuring the logical continuity and dynamic adaptability of the action generation process and effectively guiding the subsequent combination generation operation of the graph-structured action skill library. In summary, this improved structure significantly enhances the flexibility, personalization, and contextual adaptability of digital avatar action planning through high-low level strategy synergy, sub-goal level adaptive adjustment, and emotional style personality fusion control, thereby improving the accuracy, coherence, and user satisfaction of overall action expression.

[0086] In one embodiment, an improved hierarchical reinforcement learning structure is constructed based on the target granularity score. This structure includes a high-level policy network, a style regulator, and a low-level action policy network. The high-level policy network generates a sub-target intent embedding vector, and the style regulator generates a style regulation vector. The sub-target intent embedding vector and the style regulation vector are then input into the low-level action policy network to output a motion trajectory reference path. Specifically:

[0087] Input the semantic sub-target vector into the high-level policy network, and output the sub-target intent embedding vector;

[0088] User personality parameters are generated based on user behavior data and historical response records. These parameters include action style preferences, response rhythm tendencies, and risk tolerance characteristics.

[0089] By combining natural language commands and user behavior feedback in the current interaction round, a contextual sentiment label is constructed, which reflects the emotional orientation and intensity of the current interaction state.

[0090] The style modifier is input by embedding sub-target intent into a vector, user personality parameters, and upper and lower emotion tags to generate a style modifier vector.

[0091] The sub-target intent embedding vector and style control vector are concatenated to form a fused feature representation, which is then input into the low-level action policy network and outputs an action trajectory reference path.

[0092] This invention achieves a closed-loop control mechanism from high-level semantic understanding to personalized action generation by introducing an improved hierarchical reinforcement learning structure guided by target granularity scoring. First, the semantic sub-target vector is input into the high-level policy network, outputting a sub-target intent embedding vector. This effectively abstracts and structures the semantic target, enabling the system to accurately extract the core intent when faced with complex natural language expressions from users, enhancing its task-driven recognition capabilities. Second, through statistical analysis of user behavior data and historical response records, action style preferences, response rhythm tendencies, and risk tolerance features are extracted to form user personality parameters. This allows the system to identify and remember user interaction characteristics in different scenarios, establishing a long-term user profile. Furthermore, contextual emotion tags are constructed by combining natural language instructions and user behavior feedback in the current round, thereby capturing the user's emotional fluctuations and intensity during the current interaction process, achieving emotion perception and real-time adaptation. Subsequently, the sub-target intent embedding vector, user personality parameters, and contextual emotion tags are input into a style regulator. The generated style regulation vector, through weighted fusion adjustment of multi-source features, makes the action style more consistent with the user's individual characteristics and current emotional state in terms of emotional expression and rhythm control. Ultimately, by fusing sub-target intent embedding vectors and style control vectors to form low-level policy inputs, precise output of action trajectory reference paths is achieved, enhancing the continuity, flexibility, and personalized expression of action generation. Overall, this implementation strengthens the coupling between semantic understanding and personality modeling, enabling the action decision-making process to maintain semantic accuracy while taking into account individual preferences and emotional adaptation. This significantly improves the naturalness of digital avatar behavior, user satisfaction, and interactive immersion in multi-round human-computer interactions.

[0093] In one embodiment, a graph-structured action skill library is constructed. Based on the semantic sub-target vectors and action trajectory reference paths, the adjacency structure of the graph-structured action skill library is dynamically adjusted using a graph attention mechanism to generate a candidate graph of action combinations. Specifically:

[0094] A graph-structured action skill library is constructed based on a predefined set of atomic action skills. Each atomic action skill is represented as a node in the graph, and the state transition probability between atomic actions is calculated based on historical behavior data to construct a set of edges between nodes.

[0095] Initialize the graph-structured action skill library to form an adjacency matrix with associated weights, which describes the transition strength and reachability relationship between action skills;

[0096] The semantic sub-target vector and the action trajectory reference path are used as joint control information and input into the graph attention mechanism module to dynamically update the adjacency matrix and adjust the edge weights to highlight the action transfer path related to the current semantic target.

[0097] Based on the updated graph-structured action skill library, a combination search is performed on candidate action nodes to generate an action combination candidate graph, providing optional skill sequences.

[0098] This invention utilizes a graph attention mechanism jointly driven by semantic sub-target vectors and action trajectory reference paths to dynamically adjust the adjacency structure of a graph-structured action skill library, significantly improving the adaptability and contextual relevance of digital avatar action generation. First, a graph-structured action skill library is constructed based on atomic action skill sets, and state transition probabilities are calculated using historical behavioral data, forming a graph representation with semantic and temporal dependencies. This allows for structured modeling of the intrinsic relationships between action skills, providing a scalable expressive foundation for subsequent combination and decision-making. Second, an adjacency matrix is ​​constructed by introducing association weights during graph initialization, clearly expressing the transition strength and behavioral path preferences between action skills, exhibiting good controllability. Furthermore, a graph attention mechanism is introduced to dynamically update the adjacency structure, fully combining the current semantic sub-target vector and action trajectory reference path as joint control information. Edge weights are adjusted in real-time during model execution, dynamically emphasizing task-related skill transition channels, suppressing irrelevant paths, and improving the semantic fit and goal orientation of action combination generation. In addition, by searching for candidate action node combinations in the updated graph structure, not only is the efficiency of action combination generation improved, but the behavioral continuity and contextual consistency of skill sequences are also guaranteed. Overall, this method breaks the static nature of action selection in traditional fixed action libraries, realizes the reconstruction and dynamic adaptation of skill libraries based on semantic and behavioral guidance, significantly enhances the action organization ability and response intelligence level of digital avatars in complex, multi-turn interactive tasks, effectively alleviates the problems of inconsistent action selection and abrupt style when facing ambiguous semantic targets in real scenarios, and further improves the naturalness and credibility of virtual interactive experience.

[0099] In one embodiment, a Koopman linearized subspace control model is constructed, an observable mapping function is established using training data, and a linear system model in the Koopman subspace is obtained by training based on the subspace state vector and control input. Specifically:

[0100] A goal-oriented Koopman linearized subspace control model is constructed, which includes an observable mapping function module and a goal-enhanced subspace linear system modeling module.

[0101] The original action state vector sequence, control input sequence and corresponding semantic sub-target vector sequence are collected. The original action state vector is input into the observable mapping function module composed of a multi-layer neural network to generate subspace state vectors.

[0102] The subspace state vector, control input sequence, and semantic sub-target vector are input into the target-enhanced subspace linear modeling module to fit and construct a linear system containing a state transition matrix, a control input mapping matrix, and a target guidance mapping matrix;

[0103] The model is trained and evaluated based on trajectory reconstruction error, semantic response consistency index and system stability constraints. When each performance index meets the preset threshold, the parameters of the final goal-oriented Koopman linearized subspace control model are determined and solidified into the subspace prediction model for the motion trajectory optimization stage.

[0104] This invention achieves accurate modeling and efficient prediction of digital avatar action states by introducing a goal-oriented Koopman linearized subspace control model. Based on the original action state vector, control input sequence, and semantic sub-objective vector, this method transforms the nonlinear original state into a subspace state with linear characteristics by constructing an observable mapping function. This linear mapping not only improves the stability of the model in the state representation process but also provides clear structural support for the subsequent design and optimization of the controller. During modeling, the system integrates semantic sub-objective information, constructing a goal-enhanced linear subspace system. This allows the control model to not only express the relationship between the state and control input but also capture the guiding effect of user intent on action evolution. This fusion approach enables the digital avatar to exhibit higher semantic matching degree and action response flexibility when executing complex instructions. Simultaneously, this method introduces multi-dimensional performance indicators such as trajectory reconstruction error, semantic response consistency, and system stability to train and evaluate the model, thereby ensuring the model's generalization ability and control accuracy under different semantic objectives. The final trained Koopman linearized subspace control model can be embedded in the motion trajectory optimization module, enabling the system to stably generate highly adaptive and coherent control sequences to the target during motion execution. This avoids control drift and motion breakage caused by the high difficulty of nonlinear modeling in traditional methods, significantly improving the intelligent performance of virtual characters and the user interaction experience in dynamic interactive scenarios.

[0105] In one embodiment, trajectory optimization is performed in the Koopman subspace based on the motion trajectory reference path and the motion combination candidate map to generate the optimal subspace motion control sequence. This sequence is then input to the decoder and mapped back to the original motion state space, outputting a digitally visualized continuous motion execution instruction. Specifically:

[0106] Based on the Koopman linearized subspace control model, the motion trajectory reference path and motion combination candidate map are mapped to the corresponding subspace state vector sequence and control input sequence.

[0107] In the Koopman linearized subspace control model, the optimal control trajectory is solved to generate the optimal subspace motion control sequence with the goal of minimizing the trajectory offset loss and control cost function.

[0108] The optimal subspace control sequence and the corresponding sub-objective vector are input into the Koopman linearized subspace control model to predict the optimal subspace state sequence.

[0109] The trained decoder function is used to reverse map the optimal subspace state sequence to restore the action state sequence in the original action state space, and generate a digital image of continuous action execution instructions.

[0110] This invention leverages the Koopman linearized subspace control model to jointly optimize the motion trajectory reference path and the action combination candidate graph, effectively improving the coherence and response accuracy of digital avatars in dynamic scenes. By mapping the motion trajectory reference path and the action combination candidate graph to state vector sequences and control input sequences in the Koopman subspace, respectively, the system can perform optimization control solutions within a subspace with linear dynamic characteristics. Compared to traditional nonlinear control methods, this significantly reduces solution complexity and computational overhead, while enhancing the interpretability and convergence stability of the optimization results. During trajectory optimization, an optimization strategy focused on minimizing trajectory offset loss and control cost ensures that the generated action control sequence accurately matches the original intent path, while avoiding redundant control and enhancing system execution efficiency. Incorporating sub-target vectors as auxiliary input further enhances the response matching capability of the optimization results to semantic targets, ensuring that the digital avatar exhibits highly consistent semantic behavior trends during task execution. After trajectory optimization, the system uses a trained decoder to inversely map the optimal subspace state sequence, restoring it to a continuous action sequence in the original action space, ultimately outputting continuous action execution instructions for the digital avatar. This approach not only ensures the physical executability and semantic coherence of action sequences, but also significantly improves the naturalness and coordination of action generation, avoiding problems such as action breakage, response delay and semantic deviation in traditional methods, and demonstrating stronger action intelligence and user adaptability in multi-round complex interactions.

[0111] In one embodiment, environmental feedback data is collected after the execution of continuous action execution instructions. Parameters are updated based on the environmental feedback data, and the graph-structured action skill base and linear system model are updated simultaneously to achieve iterative optimization of the improved hierarchical reinforcement learning structure. Specifically:

[0112] Collect environmental feedback data after the execution of continuous action execution instructions. The environmental feedback data includes user response characteristics, behavior result status and action naturalness index.

[0113] An experience playback dataset is jointly constructed by environmental feedback data and multi-source interaction data. The network parameters of the high-level policy network, style regulator and low-level action policy network are updated respectively. Reinforcement learning policy iterative training is performed by policy gradient optimization method.

[0114] Based on the motion deviation and trajectory response error observed in the environmental feedback data, the node transfer weights and adjacency structure in the graph-structured motion skill library are adaptively adjusted by adopting the minimum reconstruction error criterion.

[0115] By simultaneously utilizing the triplet samples of state-control input-target vector in the latest environmental feedback samples, the state transition matrix, control mapping matrix, and target guidance matrix of the Koopman linearized subspace control model are incrementally updated, thereby improving the control model's adaptability to changes in user intent and achieving dynamic optimization of the improved hierarchical reinforcement learning structure.

[0116] By collecting environmental feedback data after executing continuous action commands, the system dynamically updates the policy network and control model, significantly improving the adaptability and robustness of the digital avatar in complex interaction processes. The collected environmental feedback data covers user response characteristics, behavioral outcome states, and action naturalness indicators, comprehensively reflecting the matching effect of action execution with user intent and behavioral deviations during the interaction process. This data not only provides realistic and usable supervision signals for subsequent policy optimization but also constitutes a high-quality experience playback dataset, contributing to the robust convergence of the reinforcement learning model in multi-round interactions.

[0117] This invention constructs an experience playback set by jointly integrating feedback data with original multi-source interaction data. The system can simultaneously update the parameters of the high-level policy network, style regulator, and low-level action policy network, thereby enhancing the personalized expression and target matching capabilities of action strategies at different levels of abstraction. Iterative training using a policy gradient optimization method allows the system to continuously adjust action generation strategies to maximize user response scores and interaction satisfaction, effectively avoiding the generalization difficulties of traditional fixed strategies. Furthermore, by analyzing action deviations and trajectory response errors in environmental feedback, the adjacency structure of the graph-structured action skill base is adaptively adjusted. The connection weights between nodes dynamically reflect the current interaction intent and action transfer trends, making the generated action combinations more relevant and stable, significantly improving the skill base's adaptability to new targets and scenarios. Simultaneously, the system also uses the state-control input-target vector triples in the feedback samples to incrementally update the key matrix of the linearized subspace control model, effectively enhancing the predictive ability and sensitivity to changes in user behavior. This strategy optimization mechanism not only enables the improved hierarchical reinforcement learning structure to maintain accuracy and consistency under constantly changing user intentions and environmental conditions, but also provides technical support for the long-term interactive learning capabilities of digital avatars, making them perform well in multiple dimensions such as personalization, coherence and real-time performance.

[0118] In one embodiment, the updated high-level policy network, style regulator, low-level action policy network, graph-structured action skill library, and Koopman linearized subspace control model are saved as a digital avatar behavior package, specifically:

[0119] The updated high-level strategy network, style regulator, low-level action strategy network, graph structure action skill library and Koopman linearized subspace control model are uniformly encapsulated to construct a digital image behavior package.

[0120] The configuration parameters, training weights, structural topology, and input / output interfaces of each module in the digital image behavior package are archived and encoded to form a standardized model file that can be loaded and deployed.

[0121] The system loads digital avatar behavior packages into new interactive scenarios, receives new multi-source interactive data inputs, and completes action strategy reasoning and control execution through an improved hierarchical reinforcement learning structure, supporting dynamic action adjustment tasks of digital avatars across scenarios.

[0122] This invention encapsulates the updated high-level policy network, style regulator, low-level action policy network, graph-structured action skill library, and Koopman linearized subspace control model into a unified digital avatar behavior package. This not only achieves modular management and efficient invocation of each module but also significantly improves the system's portability and deployment efficiency across different interaction scenarios. By uniformly archiving and encoding the configuration parameters, training weights, structural topology, and input / output interfaces of each component module within the behavior package, the system can generate standardized, portable model files, facilitating rapid loading and version control, and providing fundamental support for synchronous deployment and large-scale promotion across multiple devices. As a scalable intelligent agent capability carrier, this digital avatar behavior package allows the system to perform action policy inference and control output without retraining when facing new users, new semantic targets, or new interaction scenarios. This significantly reduces response latency and computational overhead, improving the real-time performance and stability of the user experience. Furthermore, this structure possesses excellent sustainable iteration capabilities. By continuously collecting feedback data during the interaction process and updating the policy, the behavior package can be periodically reconstructed and its performance optimized, ensuring the continuous intelligent evolution of the digital avatar during long-term use. More importantly, the behavior package storage mechanism endows digital avatars with "memory" and "habit" capabilities. In future scenarios, differentiated transfer learning can be performed based on historical behavior packages to generate action responses with individual style, target sensitivity, and dynamic adaptability, significantly enhancing the system's personalized expressiveness and generalization ability. Therefore, this behavior package mechanism provides crucial foundational support for the dynamic adjustment of digital avatar actions in multi-scenario, multi-user, and multi-task environments, possessing significant engineering promotion value and application prospects.

[0123] To verify the feasibility of this invention in practice, this example applies it to a children's voice-interactive education scenario within a virtual character-driven platform. The platform targets children aged 5 to 10, providing functional modules for multi-round voice-interactive teaching using a "digital avatar teacher," including English speaking practice, picture book explanations, behavioral guidance, and interactive mini-games. In practical applications, traditional preset animation templates and fixed action scripts suffer from stiff responses, unnatural emotional expression, and difficulty adapting to different semantic goals. Especially when faced with complex or emotional commands, the digital avatar's response is significantly incompatible, impacting children's immersion and engagement. This invention, through a reinforcement learning-driven dynamic action adjustment mechanism, achieves adaptive action generation oriented towards semantic sub-goals in this scenario, significantly improving the naturalness and accuracy of the interaction.

[0124] During the testing phase, 320 students from 8 primary schools in a certain region were selected and divided into two groups for the experiment. The control group (160 students) used an older version of the system that did not include the technology of this invention, while the experimental group (160 students) used an improved platform system that integrates this invention. The interactive content included three types of tasks: "English dialogue and Q&A," "rhythm game interaction," and "safety education simulation." During multiple rounds of interaction, the system collected various data, including user voice input, response semantics, number of interaction rounds, action execution feedback, children's facial expression and emotion annotations, and teacher scores.

[0125] For example, in the "English Conversation Q&A" task, when a child inputs "Can you show me how to say thank you in English?", the traditional system simply plays the recording and mechanically nods. The system using this invention, however, analyzes the emotional intent in the instruction, recognizing it as "request for demonstration + emotional expression," dynamically generating a complex behavioral chain that includes voice imitation and semantic emphasis. Furthermore, through a style adjustment mechanism, it generates a softer tone and smiling expression based on the user's past preferences, making the overall performance more approachable. In the "Rhythm Game Interaction," after a child inputs "Let's dance faster!", the experimental group system dynamically recognizes the "speed enhancement" intent based on the target granularity score and adjusts the edge weights between nodes in the graph-structured action skill library, making the generated actions more rhythmic and synchronized. Using the Koopman subspace control model improves the continuity score of action transitions.

[0126] The specific experimental results are summarized in the table below:

[0127] Table 1. Performance Comparison of Experimental and Control Group Systems under Different Task Scenarios

[0128] project control group performance experimental group performance Increase Average action response time (seconds) 2.1 1.4 -33.3% Action semantic adaptation rate 61.2% 87.6% +43.2% motion trajectory continuity score 6.7 (out of 10) 8.2 (out of 10) +22.4% Average trajectory offset 1.15 (Unit: Action Coding Spatial Distance) 0.75 (unit: action coding space distance) -35.2% Environmental feedback convergence rounds 17.8 wheels 10.5 wheels -41.0% Interaction Satisfaction Survey (Children) The satisfaction rate was 65.3%. The satisfaction rate was 93.1%. +42.5% Teacher evaluation approval 68.6% 89.4% +30.3%

[0129] As shown in Table 1, the experimental group using the method of this invention significantly outperformed the control group in several key performance indicators. First, regarding action response time, the experimental group averaged 1.4 seconds, a reduction of approximately 33.3% compared to the control group. This indicates that the improved hierarchical reinforcement learning structure, combined with sub-target parsing and style control mechanisms, can quickly generate personalized action response paths, improving the system's real-time performance. Second, the action semantic adaptation rate increased from 61.2% to 87.6%, demonstrating that the action generation method driven by semantic sub-target vectors more accurately corresponds to the user's deeper intent, significantly enhancing the naturalness of the interaction.

[0130] Furthermore, the continuity score of the action trajectory improved from 6.7 to 8.2, and the average trajectory offset decreased by 35.2%, verifying the effectiveness of the Koopman linearized subspace control model in terms of the smoothness and stability of action generation. Simultaneously, the number of convergence rounds of the policy network decreased from 17.8 rounds to 10.5 rounds, a reduction of 41%, reflecting the faster policy optimization capability of the reinforcement learning structure of this invention under environmental feedback-driven conditions. The simultaneous increase in user subjective satisfaction and teacher approval (+42.5% and +30.3%, respectively) further validates the practicality and innovation of this invention from a practical experience perspective.

[0131] This invention not only systematically enhances the traditional digital avatar control system in terms of theoretical structure, but also demonstrates superior dynamic response capabilities, personalized matching capabilities, and system robustness in practical applications, fully meeting the needs for intelligent control of digital avatars with high naturalness and strong interactivity. This method has broad application prospects, especially suitable for semantic-driven interaction scenarios such as educational companionship, virtual assistants, and rehabilitation training.

[0132] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for dynamically adjusting the actions of a digital avatar based on reinforcement learning, characterized in that, include: Collect multi-source interaction data between users and digital avatars, perform semantic parsing on the multi-source interaction data, generate semantic sub-target vectors, and calculate target granularity scores based on the semantic sub-target vectors; An improved hierarchical reinforcement learning structure is constructed based on the target granularity score. The improved hierarchical reinforcement learning structure includes a high-level policy network, a style regulator, and a low-level action policy network. The high-level policy network generates a sub-target intent embedding vector, and the style regulator generates a style regulation vector. The sub-target intent embedding vector and the style regulation vector are input into the low-level action policy network to output an action trajectory reference path. A graph-structured action skill library is constructed. Based on the semantic sub-target vector and the action trajectory reference path, the adjacency structure of the graph-structured action skill library is dynamically adjusted through a graph attention mechanism to generate a candidate graph of action combinations. Construct a Koopman linearized subspace control model, establish an observable mapping function using training data, and obtain a linear system model in the Koopman subspace based on the subspace state vector and control input. Based on the motion trajectory reference path and motion combination candidate map, trajectory optimization is performed in the Koopman subspace to generate the optimal subspace motion control sequence. The sequence is then input into the decoder and mapped back to the original motion state space, outputting a digital image of continuous motion execution instructions. Collect environmental feedback data after the execution of continuous action execution instructions, update parameters based on environmental feedback data, and simultaneously update the graph structure action skill library and linear system model to achieve iterative optimization of the strategy of the improved hierarchical reinforcement learning structure; Save the updated high-level strategy network, style regulator, low-level action strategy network, graph structure action skill library and Koopman linearized subspace control model as a digital image behavior package. The high-level policy network in the improved hierarchical reinforcement learning structure receives semantic sub-target vectors and target granularity scores as inputs, and generates sub-target intent embedding vectors using a dynamic hierarchical mechanism. The dynamic hierarchical mechanism adaptively adjusts the number of sub-targets and the abstraction level of the high-level policy output according to the target granularity scores. The style modifier receives the sub-target intent embedding vector, user personality parameters, and contextual emotion labels. It uses a fusion mapping network to generate a style control vector. The fusion mapping network uses an attention mechanism to weight and combine personality and emotion factors to achieve personalized adjustment of action style and behavioral trends. The low-level action policy network receives the concatenated features of the sub-target intent embedding vector and style control vector as input, and outputs a continuous action trajectory reference path based on the conditional policy generation structure. The action trajectory reference path is composed of target state points at multiple time steps, which guides the generation of action combinations.

2. The method for dynamic adjustment of digital image actions based on reinforcement learning according to claim 1, characterized in that, The process involves collecting multi-source interaction data between users and digital avatars, performing semantic parsing on the multi-source interaction data to generate semantic sub-target vectors, and calculating target granularity scores based on the semantic sub-target vectors. Specifically: Collect multi-source interaction data between users and digital avatars, including natural language commands, scene tags, user behavior data, and historical response records; Perform unified formatting on multi-source interactive data to construct semantic parsing input vectors; The semantic parsing input vector is fed into the semantic parsing network to extract contextual semantic features and generate semantic sub-target vectors; The target granularity score is calculated based on the semantic complexity, instruction abstraction, and behavioral expectation clarity in the semantic sub-target vector.

3. The method for dynamic adjustment of digital image actions based on reinforcement learning according to claim 2, characterized in that, The improved hierarchical reinforcement learning structure is constructed based on the target granularity score. This structure includes a high-level policy network, a style regulator, and a low-level action policy network. The high-level policy network generates a sub-target intent embedding vector, and the style regulator generates a style regulation vector. The sub-target intent embedding vector and the style regulation vector are input together into the low-level action policy network, which outputs a motion trajectory reference path. Specifically: Input the semantic sub-target vector into the high-level policy network, and output the sub-target intent embedding vector; User personality parameters are generated based on the user behavior data and historical response records. These user personality parameters include action style preferences, response rhythm tendencies, and risk tolerance characteristics. By combining natural language instructions and user behavior feedback in the current interaction round, a contextual emotion label is constructed, which reflects the emotional orientation and intensity of the current interaction state. The style modifier is input by embedding sub-target intent into a vector, user personality parameters, and upper and lower emotion tags to generate a style modifier vector. The sub-target intent embedding vector and style control vector are concatenated to form a fused feature representation, which is then input into the low-level action policy network and outputs an action trajectory reference path.

4. The method for dynamic adjustment of digital image actions based on reinforcement learning according to claim 3, characterized in that, The construction of the graph-structured action skill library involves dynamically adjusting the adjacency structure of the graph-structured action skill library based on semantic sub-target vectors and action trajectory reference paths using a graph attention mechanism to generate a candidate graph of action combinations. Specifically: A graph-structured action skill library is constructed based on a predefined set of atomic action skills. Each atomic action skill is represented as a node in the graph, and the state transition probability between atomic actions is calculated based on historical behavior data to construct a set of edges between nodes. Initialize the graph-structured action skill library to form an adjacency matrix with associated weights, which describes the transition strength and reachability relationship between action skills; The semantic sub-target vector and the action trajectory reference path are used as joint control information and input into the graph attention mechanism module to dynamically update the adjacency matrix and adjust the edge weights to highlight the action transfer path related to the current semantic target. Based on the updated graph-structured action skill library, a combination search is performed on candidate action nodes to generate an action combination candidate graph, providing optional skill sequences.

5. The method for dynamic adjustment of digital image actions based on reinforcement learning according to claim 4, characterized in that, The construction of the Koopman linearized subspace control model involves establishing an observable mapping function using training data, and training a linear system model in the Koopman subspace based on the subspace state vector and control input. Specifically: A goal-oriented Koopman linearized subspace control model is constructed, which includes an observable mapping function module and a goal-enhanced subspace linear system modeling module. The original action state vector sequence, control input sequence and corresponding semantic sub-target vector sequence are collected. The original action state vector is input into the observable mapping function module composed of a multi-layer neural network to generate subspace state vectors. The subspace state vector, control input sequence, and semantic subtarget vector are input into the target-enhanced subspace linear modeling module to fit and construct a linear system including a state transition matrix, a control input mapping matrix, and a target guidance mapping matrix; The model is trained and evaluated based on trajectory reconstruction error, semantic response consistency index and system stability constraints. When each performance index meets the preset threshold, the parameters of the final goal-oriented Koopman linearized subspace control model are determined and solidified into the subspace prediction model for the motion trajectory optimization stage.

6. The method for dynamic adjustment of digital image actions based on reinforcement learning according to claim 5, characterized in that, The process involves performing trajectory optimization in the Koopman subspace based on the action trajectory reference path and action combination candidate map, generating the optimal subspace action control sequence, inputting it into the decoder for inverse mapping to the original action state space, and outputting a digitally visualized continuous action execution instruction. Specifically: Based on the Koopman linearized subspace control model, the motion trajectory reference path and motion combination candidate map are mapped to the corresponding subspace state vector sequence and control input sequence. In the Koopman linearized subspace control model, the optimal control trajectory is solved to generate the optimal subspace motion control sequence with the goal of minimizing the trajectory offset loss and control cost function. The optimal subspace control sequence and the corresponding sub-objective vector are input into the Koopman linearized subspace control model to predict the optimal subspace state sequence. The trained decoder function is used to reverse map the optimal subspace state sequence to restore the action state sequence in the original action state space, and generate a digital image of continuous action execution instructions.

7. The method for dynamic adjustment of digital image actions based on reinforcement learning according to claim 6, characterized in that, The process involves collecting environmental feedback data after the execution of continuous action execution instructions, updating parameters based on this data, and simultaneously updating the graph-structured action skill base and the linear system model. This enables iterative optimization of the strategy within the improved hierarchical reinforcement learning structure. Specifically: Collect environmental feedback data after the execution of the continuous action execution instructions. The environmental feedback data includes user response characteristics, behavior result status and action naturalness index. An experience playback dataset is jointly constructed by environmental feedback data and multi-source interaction data. The network parameters of the high-level policy network, style regulator and low-level action policy network are updated respectively. Reinforcement learning policy iterative training is performed by policy gradient optimization method. Based on the motion deviation and trajectory response error observed in the environmental feedback data, the node transfer weights and adjacency structure in the graph-structured motion skill library are adaptively adjusted by adopting the minimum reconstruction error criterion. By simultaneously utilizing the triplet samples of state-control input-target vector in the latest environmental feedback samples, the state transition matrix, control mapping matrix, and target guidance matrix of the Koopman linearized subspace control model are incrementally updated, thereby improving the control model's adaptability to changes in user intent and achieving dynamic optimization of the improved hierarchical reinforcement learning structure.

8. The method for dynamic adjustment of digital image actions based on reinforcement learning according to claim 7, characterized in that, The process of saving the updated high-level policy network, style regulator, low-level action policy network, graph-structured action skill library, and Koopman linearized subspace control model as a digital avatar behavior package specifically involves: The updated high-level strategy network, style regulator, low-level action strategy network, graph structure action skill library and Koopman linearized subspace control model are uniformly encapsulated to construct a digital image behavior package. The configuration parameters, training weights, structural topology, and input / output interfaces of each module in the digital image behavior package are archived and encoded to form a standardized model file that can be loaded and deployed. The digital avatar behavior package is loaded into the new interactive scenario, and new multi-source interactive data input is received. The action strategy reasoning and control execution are completed through an improved hierarchical reinforcement learning structure, supporting the dynamic action adjustment task of the digital avatar in cross-scenario scenarios.

Citation Information

Patent Citations

  • Holographic display digital human speech recognition enhancement method based on multi-mode interactive learning

    CN119541459A

  • Health management service system and method based on AI optimization

    CN120878209A