Policy generation methods, devices, equipment, and media based on hierarchical reinforcement learning

CN121168515BActive Publication Date: 2026-08-14PING AN TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0004]本发明的主要目的在于提供一种基于分层强化学习的策略生成方法、装置、设备及存储介质,旨在解决现有技术在复杂动态环境下难以同时建模状态演变的时序因果关系与高维动作策略的分层生成与优化,导致策略泛化能力差、响应不及时且难以收敛的技术问题

Benefits of technology

[0021] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a policy generation method, apparatus, device, and medium based on hierarchical reinforcement learning, comprising: acquiring multi-source raw scene data and preprocessing it to generate a structured scene dataset; constructing a dynamic causal graph based on the structured scene dataset; receiving environmental state information and generating sub-goals and specific actions through a hierarchical reinforcement learning architecture; acquiring state change information after the execution of specific actions and combining the environmental state information, state change information, specific actions, and sub-goals to generate a policy reward signal; jointly training and updating the dynamic causal graph and hierarchical reinforcement learning model parameters based on the policy reward signal; and using the updated hierarchical reinforcement learning model to generate optimized action policies. This invention models the state evolution relationship by constructing a dynamic causal graph, enabling reinforcement learning to gain a causal understanding of state change trends. Combined with a hierarchical reinforcement learning architecture, it achieves the decomposition and optimization of sub-goals and actions, improving the response accuracy and generalization ability of action policies in complex environments, thereby improving the stability of task completion and the convergence efficiency of policy training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121168515B_ABST
    Figure CN121168515B_ABST
Patent Text Reader

Abstract

This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a policy generation method, apparatus, device, and medium based on hierarchical reinforcement learning, comprising: acquiring multi-source raw scene data and generating a structured scene dataset; constructing a dynamic causal graph; processing environmental state information to generate sub-objectives and specific actions; generating a policy reward signal based on state changes; jointly training and updating the dynamic causal graph and the hierarchical reinforcement learning model based on the policy reward signal; and generating an optimized action policy. This invention models the state evolution relationship by constructing a dynamic causal graph, enabling reinforcement learning to gain a causal understanding of state change trends. Combined with a hierarchical reinforcement learning architecture, it achieves the decomposition and optimization of sub-objectives and actions, improving the response accuracy and generalization ability of action policies in complex environments, thereby improving the stability of task completion and the convergence efficiency of policy training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a policy generation method, apparatus, device, and storage medium based on hierarchical reinforcement learning. Background Technology

[0002] In the healthcare field, with the widespread application of intelligent devices such as surgical robots and rehabilitation robots, the generation and control of motion strategies have become crucial for ensuring operational accuracy and patient safety. Current technologies largely rely on pre-set rule systems or experience-based fixed strategies for motion planning, lacking the ability to respond to real-time environmental changes and individual differences. In real-world medical scenarios, patients' physiological states are highly dynamic, with significant differences in tissue structure and physiological parameters among individuals. Fixed motion sequences and static strategies are difficult to flexibly adjust to minor fluctuations in environmental conditions, easily leading to operational errors or disrupting normal treatment procedures. Furthermore, surgical procedures involve long causal delays; the impact of a particular action on the final treatment outcome often only gradually becomes apparent after multiple steps. Traditional reinforcement learning methods struggle to accurately establish the correlation between actions and long-term effects, resulting in slow strategy training and difficulty in convergence, severely limiting the effectiveness of reinforcement learning in medical robotics.

[0003] In the fintech sector, financial decision-making systems face complex external environmental changes, such as market fluctuations, policy adjustments, and dynamic user behavior, placing higher demands on system responsiveness and strategy adaptability. Traditional risk control and resource allocation strategies often employ static modeling methods, which struggle to capture the potential causal relationships between multidimensional data in real time, leading to policy lags or misjudgments. In high-dimensional financial behavior modeling, single-layer strategy structures are insufficient for decomposing and optimizing complex operational sequences, making it difficult for the system to meet the real-time and accuracy requirements of tasks such as anomaly detection and behavior sequence deduction, thus impacting overall business response efficiency and stability. Summary of the Invention

[0004] The main objective of this invention is to provide a policy generation method, apparatus, device, and storage medium based on hierarchical reinforcement learning. This invention aims to solve the technical problems of existing technologies, which struggle to simultaneously model the temporal causal relationships of state evolution and the hierarchical generation and optimization of high-dimensional action policies in complex dynamic environments, resulting in poor policy generalization ability, untimely response, and difficulty in convergence.

[0005] To achieve the above objectives, this invention provides a policy generation method based on hierarchical reinforcement learning, comprising:

[0006] Acquire multi-source raw scene data and preprocess the multi-source raw scene data to generate a structured scene dataset;

[0007] Based on the structured scene dataset, a dynamic causal graph is constructed to describe the causal relationships of state evolution in the scene;

[0008] Receive environmental state information, process the environmental state information through a hierarchical reinforcement learning architecture that includes high-level policies and low-level policies, and generate sub-goals and specific actions;

[0009] Obtain the state change information after the specific action is executed, and use the dynamic cause-effect graph to generate a policy reward signal based on the environmental state information, state change information, specific action and sub-goal;

[0010] Based on the policy reward signal, the parameters of the dynamic causal graph and the hierarchical reinforcement learning model are updated through joint training;

[0011] Optimized action strategies are generated using the updated hierarchical reinforcement learning model.

[0012] Furthermore, to achieve the above objectives, the present invention provides a policy generation apparatus based on hierarchical reinforcement learning, comprising:

[0013] The data preprocessing module is used to acquire multi-source raw scene data and preprocess the multi-source raw scene data to generate a structured scene dataset.

[0014] The causal modeling module is used to construct a dynamic causal graph based on the structured scene dataset to describe the causal relationships of state evolution in the scene.

[0015] The hierarchical policy generation module is used to receive environmental state information, process the environmental state information through a hierarchical reinforcement learning architecture that includes high-level policies and low-level policies, and generate sub-goals and specific actions.

[0016] The state evolution analysis module is used to obtain state change information after the specific action is executed, and to generate a policy reward signal based on the environmental state information, state change information, specific action and sub-goal using the dynamic cause-effect graph;

[0017] The joint optimization training module is used to update the parameters of the dynamic causal graph and the hierarchical reinforcement learning model through joint training based on the policy reward signal.

[0018] The action policy output module is used to generate optimized action policies through the updated hierarchical reinforcement learning model.

[0019] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a policy generation program based on hierarchical reinforcement learning stored in the memory and executable on the processor, wherein when the policy generation program based on hierarchical reinforcement learning is executed by the processor, it implements the steps of the policy generation method based on hierarchical reinforcement learning as described above.

[0020] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a policy generation program based on hierarchical reinforcement learning, wherein when the policy generation program based on hierarchical reinforcement learning is executed by a processor, it implements the steps of the policy generation method based on hierarchical reinforcement learning as described above.

[0021] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a policy generation method, apparatus, device, and medium based on hierarchical reinforcement learning, comprising: acquiring multi-source raw scene data and preprocessing it to generate a structured scene dataset; constructing a dynamic causal graph based on the structured scene dataset; receiving environmental state information and generating sub-goals and specific actions through a hierarchical reinforcement learning architecture; acquiring state change information after the execution of specific actions and combining the environmental state information, state change information, specific actions, and sub-goals to generate a policy reward signal; jointly training and updating the dynamic causal graph and hierarchical reinforcement learning model parameters based on the policy reward signal; and using the updated hierarchical reinforcement learning model to generate optimized action policies. This invention models the state evolution relationship by constructing a dynamic causal graph, enabling reinforcement learning to gain a causal understanding of state change trends. Combined with a hierarchical reinforcement learning architecture, it achieves the decomposition and optimization of sub-goals and actions, improving the response accuracy and generalization ability of action policies in complex environments, thereby improving the stability of task completion and the convergence efficiency of policy training. Attached Figure Description

[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:

[0023] Figure 1 This is a schematic diagram of an application environment for a policy generation method based on hierarchical reinforcement learning in one embodiment of the present invention;

[0024] Figure 2 This is a flowchart illustrating an embodiment of the policy generation method based on hierarchical reinforcement learning according to the present invention.

[0025] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the policy generation device based on hierarchical reinforcement learning of the present invention.

[0026] Figure 4This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0027] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0028] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0029] The policy generation method based on hierarchical reinforcement learning provided in this invention can be applied to, for example... Figure 1 In this application environment, the user terminal communicates with the server via a network. The server can obtain multi-source raw scene data from the user terminal and preprocess it to generate a structured scene dataset. Based on the structured scene dataset, it constructs a dynamic causal graph, receives environmental state information, generates sub-goals and specific actions through a hierarchical reinforcement learning architecture, obtains state change information after the execution of specific actions, and combines environmental state information, state change information, specific actions, and sub-goals to generate a policy reward signal. Based on the policy reward signal, it jointly trains and updates the parameters of the dynamic causal graph and the hierarchical reinforcement learning model, and uses the updated hierarchical reinforcement learning model to generate optimized action policies. This invention models the state evolution relationship by constructing a dynamic causal graph, enabling reinforcement learning to gain a causal understanding of state change trends. Combined with a hierarchical reinforcement learning architecture, it achieves the decomposition and optimization of sub-goals and actions, improving the response accuracy and generalization ability of action policies in complex environments, thereby improving the stability of task completion and the convergence efficiency of policy training. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster composed of multiple servers. The invention will be described in detail below through specific embodiments.

[0030] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the policy generation method based on hierarchical reinforcement learning provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0031] like Figure 2 As shown, the policy generation method based on hierarchical reinforcement learning proposed in this invention includes the following steps:

[0032] S10, acquire multi-source raw scene data, and preprocess the multi-source raw scene data to generate a structured scene dataset;

[0033] In this embodiment, to achieve an action generation process that can adapt to dynamic changes in complex environments, it is necessary to construct a data structure that supports subsequent causal modeling and strategy reasoning, starting from multi-source heterogeneous raw scene data. First, it is necessary to acquire multi-source raw scene data, including image acquisition devices, various environmental sensors, and structured database systems. Image acquisition devices may include RGB cameras, depth cameras, endoscope camera modules, etc., used to acquire visual content related to the target task; environmental sensors include temperature, pressure, acceleration, electrophysiological signals, etc., suitable for capturing physical or physiological change signals; structured database systems can provide tagged medical records, device status logs, or historical transaction data. These various types of data often exhibit heterogeneity in terms of time, resolution, structural format, and sampling accuracy, therefore preprocessing is required to build a unified data foundation.

[0034] After data acquisition, noise removal and outlier removal are necessary. Outliers are identified using density-based methods, or abrupt changes in the sequence data are processed using moving averages and sliding window filtering to obtain stable and reliable basic observation data. Alignment of temporal information is a crucial prerequisite for multi-source fusion. Global timestamp synchronization mechanisms are commonly used to align asynchronously sampled data streams to a unified timeline, or linear interpolation and dynamic time warping algorithms are employed for time-dimension alignment. To address the inconsistency in data dimensions between modalities, standardization of each data channel is required. This standardization includes min-max scaling, z-score normalization, or nonlinear function transformations to unify the representation range of different modalities and improve the effectiveness of subsequent feature fusion stages.

[0035] Based on standardization, feature extraction is performed on various types of input data. Image data can be extracted using convolutional neural networks to obtain spatial structure information; sensor data can be transformed using frequency domain to obtain dynamic change patterns; and structured data uses field filtering and statistical coding methods to extract key indicators. After feature extraction, the features from different modalities need to be fused. Fusion methods include early stitching, mid-stage alignment fusion, and late-stage decision-level fusion. To adapt to the variable requirements in causal modeling, the fused feature representations need further dimensional reconstruction or context expansion through linear mapping or self-attention mechanisms, enabling different types of variables to express clear state characterization information in a unified feature space. The final output is a structured scene dataset, which should include fields such as temporal index, variable name, state value, and data source information to support subsequent causal modeling, policy evaluation, and action generation operations.

[0036] In one implementation, multiple cameras are deployed at different spatial angles and simultaneously connected to physiological monitoring sensors and electronic medical record systems to construct three types of data sources: images, sensors, and structured data. A unified data acquisition gateway acquires raw data from different modalities, and a high-precision timestamp synchronization mechanism is used to normalize the sampling time of each data source. For image data, a ResNet network is used to extract backbone features; for sensor data, wavelet transform is used to extract spectral features; and for structured data, field normalization and value range encoding are performed to obtain structural index features. In the fusion stage, a Transformer is used to perform cross-attention fusion of multimodal features, outputting a unified structured scene state representation. Finally, a structured scene dataset suitable for modeling is constructed in JSON format.

[0037] In another implementation approach, applied to financial risk analysis tasks, a time-labeled data structure is constructed by combining user interaction logs, historical transaction records, and external credit scoring data. Transaction data undergoes frequency analysis, log data uses a BERT model to extract contextual features, and credit data uses rules to extract key fields and convert them into vector representations. After fusion, a set of state variables is generated through MLP mapping, further generating a structured scenario dataset.

[0038] Example Description: In the healthcare field, this can be applied to minimally invasive surgical assistance systems. By combining endoscopic images, intraoperative physiological signals, and preoperative medical information to construct scene data, the system can achieve state recognition and risk prediction. During surgery, the system can adjust its operational strategies based on continuous visual changes and the patient's physiological state. Through structured scene dataset input, the system enhances the surgical robot's understanding of the surgical scenario.

[0039] In the fintech business field, it can be used for credit behavior modeling tasks. By integrating user behavior logs, payment transaction data and credit information, it can build a structured scenario dataset that can characterize the evolution of user risk, support dynamic prediction of future risk trends and generation of action strategies, and improve the dynamic adaptability of credit strategies and the effectiveness of risk control.

[0040] This embodiment integrates multi-source heterogeneous data, completes time alignment and standardization processing, and constructs a unified structured scene dataset. This enhances the collaborative expression capabilities between different modalities while maintaining information integrity, providing sufficient data support for subsequent causal modeling and strategy generation, thereby improving the adaptability and accuracy of the action generation process.

[0041] S20, Based on the structured scene dataset, construct a dynamic causal graph to describe the causal relationships of state evolution in the scene;

[0042] In this embodiment, constructing a dynamic causal graph to characterize the causal relationships of state evolution requires a pre-processed and feature-fused structured scene dataset as input. First, it is necessary to identify a set of state variables from this dataset that can participate in causal analysis. These state variables can originate from device operating states, sensor values, user behavior parameters, environmental indicators, etc., and typically have clear naming conventions, data types, time-series distributions, and sampling periods. During variable selection, the interactivity and temporal dependencies between variables need to be considered to avoid introducing redundant or weakly correlated variables that could interfere with the causal modeling process.

[0043] After obtaining the set of state variables, a preliminary causal structure model needs to be generated based on time series causal discovery algorithms. Commonly used methods include Granger causality tests, structure learning algorithms based on mutual information or conditional independence, and multivariate time series causal inference methods such as PCMCI (Peter-Clark Momentary Conditional Independence). These algorithms can output directed edges between variables and their causal strength scores, thereby constructing a directed acyclic graph (DAG) or directed graph model to represent the preliminary causal structure.

[0044] Because time-series data often involves complex factors such as asynchronous changes, time delay effects, and state drift, time-delay embedding processing of structured data is necessary to enhance the model's ability to model potential time-delay relationships. This processing constructs a joint representation of state variables at multiple time points using a sliding window approach, enabling the model to capture the multi-scale influence paths of the current state on future states. By introducing data with delayed embedding into causal modeling, the explanatory power and stability of causal edges can be significantly improved.

[0045] After the initial causal graph is established, structural optimization is required to improve its ability to predict state evolution under different scenarios. Optimization strategies include regularization constraints on edge weights, structural pruning based on prediction errors, and robust adjustment of the causal graph through multi-model fusion. Some optimization strategies can integrate dynamic Bayesian networks, graph neural networks, or sparse structure learning methods to enable the causal graph to dynamically evolve with data updates.

[0046] Finally, to ensure the reliability of the causal graph, consistency verification is required. This process is accomplished by comparing the consistency measures between the observed state evolution trajectories in the structured dataset and the evolution paths predicted by the causal graph. Verification methods may include distribution matching based on Kullback-Leibler divergence, reconstruction accuracy evaluation based on path matching rate, or behavioral simulation testing based on replay mechanisms.

[0047] The variable selection rules in the state variable identification process can be adjusted to adapt to different data source types. For example, for visual feature variables extracted from image streams, a spatial region attention mechanism can be used to filter out low-relevance regions; for data collected by sensors, the information content of variables can be evaluated based on noise variance and information entropy, and variables with low information content can be removed.

[0048] In the causal structure learning stage, a graph learning method combining neural network embedding structures and self-attention mechanisms can be used instead of traditional Granger tests or conditionally independent algorithms to adapt to nonlinear or high-dimensional variable scenarios. Regarding time-delay embedding, the window width and step size can be dynamically set according to the system response delay to avoid error propagation caused by uniform delay embedding.

[0049] To optimize causal graphs, reinforcement learning feedback mechanisms or joint modeling strategies can be introduced to achieve synchronous iteration of edge structure updates and state predictions, so that the causal graph gradually converges to a topological structure that can accurately reflect the state evolution mechanism.

[0050] Example: In the healthcare business, for the motion generation requirements during the operation of surgical robots, dynamic cause-effect graphs can be used to model the causal paths between different operation steps and the patient's physiological feedback. This supports the prediction of subsequent state changes based on current physiological parameters, thereby allowing for advance adjustment of operation strategies to avoid risks.

[0051] In the field of fintech business, when faced with the complex relationship between behavioral variables and the final credit granting result in credit approval or risk assessment, a causal structure between multidimensional features can be constructed through dynamic causal graphs. This can support the explanation of the impact path of individual behavior on the evolution of credit scores and assist decision-making systems in making rapid strategy adjustments in a highly uncertain environment.

[0052] This embodiment systematically filters, embeds, and models causal relationships among variables in a structured scene dataset. This enables the identification of predictive state-dependent paths in complex environments and the construction of a causal graph model with dynamic evolution capabilities. This causal graph accurately depicts the logical structure of multiple variables changing over time, providing causal explanations for subsequent policy generation or behavioral decisions, thereby improving the response speed and policy stability of the action generation system.

[0053] S30, Receive environmental state information, process the environmental state information through a hierarchical reinforcement learning architecture that includes high-level policies and low-level policies, and generate sub-goals and specific actions;

[0054] In this embodiment, environmental state information refers to observable data reflecting the current system operating environment, object status, or task progress. This data may include time-series sensor data, image recognition results, user operation feedback, task execution logs, and physical parameter sampling values. This information needs to be represented in vectorized or tensor structure form to adapt to the spatial structure and temporal dynamics requirements of subsequent processing.

[0055] The received environmental state information is input into a hierarchical reinforcement learning architecture for processing. This architecture includes at least two structural modules: a high-level policy and a low-level policy. The high-level policy is responsible for abstracting information such as task stages, contextual intentions, or stage goals from the environmental state, and the output is a stage intention signal expressed in the form of intermediate goals or sub-goals. These sub-goals can take the form of discrete labels, continuous goal vectors, or multi-dimensional goal parameters, depending on the type and complexity of the task representation space.

[0056] The high-level policy module can be implemented as a policy network based on policy gradients, Actor-Critic, value-selective policy tables, or a Transformer structure. Its input is the current state representation, and its output must have clear meaning and decomposability in the policy space for easy use by subsequent modules. The sub-objectives extracted by the high-level policy are input into the low-level policy module along with the current environment state information.

[0057] The underlying strategy module receives sub-target signals and generates executable action instructions based on the current environmental state. Its main tasks are to solve continuous control, action decoding, and execution parameter generation problems. The underlying strategy typically uses convolutional neural networks, recurrent neural networks, or structures with attention mechanisms to jointly model the state and sub-targets, and outputs corresponding action space parameters, such as multi-joint control instructions, click selection positions, and motion path control points. The generated actions must be parsable into action vectors, behavior labels, or operation commands that the system execution unit can recognize.

[0058] A hierarchical structure with a pluggable sub-target representation space can be adopted, defining the output structure of the high-level policy module in a configurable format. For example, in an execution scenario, sub-targets can be represented as spatial coordinate points, behavior stage labels, or subtask identifiers. The depth of the high-level policy structure, the width of the induction window, or the diversity of policy outputs can be adjusted to suit different task complexities.

[0059] The underlying policy can be constructed as a conditional policy network, with conditions including the current state and sub-goals. A dual-input attention mechanism can be incorporated to enhance the joint modeling capability between sub-goals and the environmental state. In the discrete action space, a Q-learning structure is used to handle sub-goal-driven behavior selection, while in the continuous space, DDPG (Deep Deterministic Policy Gradient) or SAC (Soft Actor-Critic) is used for high-precision action control.

[0060] An intermediate reward signal evaluation mechanism can also be introduced between high and low levels to make bidirectional reinforcement adjustments to the reachability of high-level sub-goals and low-level execution feedback, thereby improving learning efficiency and action stability.

[0061] Example: In the field of healthcare, surgical robots receive environmental status information such as patient vital signs, endoscopic images, and instrument postures. The high-level strategy identifies the current surgical stage and generates sub-goals such as "clamp position alignment" or "suture start point positioning". The low-level strategy controls the multi-joint arm to perform precise operations based on the sub-goals, thereby completing the execution of a high-precision operation path.

[0062] In the fintech business, risk assessment robots form an environmental state based on users' real-time interaction behavior, input data and feedback results. High-level strategies output sub-goals such as "completing key information" or "verifying financial stability". The underlying strategies dynamically select questioning methods, prompt structures and guidance paths based on the goals to generate strategic interaction behaviors.

[0063] This embodiment uses a hierarchical reinforcement learning architecture to process the environment state, introducing structured decoupling when dealing with high-dimensional states and complex action spaces. It divides the policy learning task into two complementary tasks: sub-goal abstraction and concrete action execution, improving policy convergence speed and generalization ability. The sub-goal guidance provided by the high-level policy enables phased planning in action generation, while the low-level policy optimizes the control response based on sub-goals, thereby enhancing the flexibility and stability of task execution.

[0064] S40, obtain the state change information after the specific action is executed, and use the dynamic cause-effect graph to generate a policy reward signal based on the environmental state information, state change information, specific action and sub-goal;

[0065] In this embodiment, the state change information after a specific action is performed refers to the changes in the environmental state after the system performs a specific action, which is reflected in the update of state variable values. This change information is usually represented in the form of a time series or state pairs, and the data structure includes state transition tensors, event logs, parameter change vectors, etc., which can be obtained through system monitoring modules, sensor feedback, or log feedback mechanisms.

[0066] Dynamic causal graphs are used to model the causal relationships between different state variables that evolve over time. Their structure includes causal nodes, edge directions, and condition strengths, used to infer possible paths for state changes. In this process, the dynamic causal graph receives current environmental state information, specific actions, and sub-goals as input conditions, and derives the expected state path after the action is executed. This path is not a static prediction result, but rather a reasoning path generated based on the graph structure under the current context, used as a reference for reward evaluation.

[0067] By comparing the actual state change information with the expected state evolution path, a state change difference index can be calculated using methods such as path matching degree, node similarity, and state offset. This index quantifies the degree of difference between the current action and the system's expectation and is a fundamental component in constructing the reward signal.

[0068] The reward signal comprises intrinsic and extrinsic reward components. Intrinsic rewards typically rely on state change difference indicators to measure the consistency of actions with the expected evolutionary path, used during the training phase to enhance the learning ability of the strategy structure. Extrinsic rewards are usually calculated based on task completion progress or business goal achievement rate, such as whether a certain intermediate state has been reached, whether the operation time is within the expected range, or whether a sub-task has been completed. The weight ratio of intrinsic and extrinsic rewards is not fixed but dynamically adjusted according to the current training phase. For example, the proportion of intrinsic rewards is increased in the early stages when exploration is emphasized, while the weight of extrinsic goals is increased during the convergence phase.

[0069] Finally, the two reward components are merged and a policy reward signal is generated through weighted summation or nonlinear mapping. This signal serves as the objective function input during the hierarchical policy network update process, driving the policy optimization direction.

[0070] Attention-based difference detection models can be used to align state change information with expected state paths, extracting the most influential state offsets to construct state change difference indices. For example, introducing conditionally adjustable jump edges or memory mechanisms into graph structures can enhance the temporal representation of causal paths. The expected path can be generated by combining the current action with finite-step reasoning on the conditionally activated edges of the causal graph to construct an evolution sequence within a time window.

[0071] Internal rewards can be designed as path consistency scores, such as the percentage of matching state nodes or the achievement rate of key states. External rewards can refer to task completion indicators, user feedback scores, or the convergence value of the system's objective function. Reward weights can be adaptively adjusted using polynomial scheduling functions, warm-start strategies, or learning regulators to achieve dynamic adaptation.

[0072] Furthermore, a reward attribution mechanism can be introduced to distribute global rewards backward to each intermediate action node, improving the fine-grained control over policy updates. In a layered architecture, reward signals of different dimensions can be calculated for high-level and low-level policies respectively, and then fused based on the dependencies between the main task and sub-goals.

[0073] Example description: In the field of healthcare, after performing a suturing action, the surgical robot collects state data on changes in tissue tension and the position of the suture line. The system compares this state change with the expected healing path derived from the dynamic cause-effect graph. If the deviation is too large, a negative intrinsic reward is generated. At the same time, an extrinsic reward is generated by combining the task progress assessment with the suture integrity assessment. The comprehensive generation of reward signals is used for strategy optimization.

[0074] In the fintech business, the automated review system uses changes in system parameters after a user submits an operation, such as adjustments to credit scores or increases in data completeness, to constitute a state change. This state is then compared with the target state in the cause-and-effect graph. If the direction of the change aligns with the expected process, a positive internal reward is given. At the same time, external rewards are evaluated in conjunction with the user's task completion progress, such as whether the review step has entered the next stage. Ultimately, these are used to update the decision-making strategy generation logic.

[0075] This embodiment compares and quantifies the differences between state changes and expected causal paths, combined with a structured internal and external reward component design, so that policy learning no longer relies on a single delayed feedback signal, improving the timeliness and fine granularity of training feedback. Introducing a dynamic causal graph as a reference for expected generation enhances the policy's ability to perceive and constrain causal relationships, improves the consistency and stability in the generation of complex action sequences, and simultaneously achieves a unified integration of structure guidance and goal guidance in the learning process.

[0076] S50, based on the policy reward signal, update the parameters of the dynamic causal graph and the hierarchical reinforcement learning model through joint training;

[0077] In this embodiment, the joint training process is driven by the policy reward signal. Based on previously collected environmental state information, specific actions, sub-objectives, and state change data, an experience sample set is constructed as the training data source for updating model parameters. The experience data includes multi-dimensional temporal information and policy behavior feedback. Its storage structure should support temporal backtracking, state transition queries, and associated reward extraction, and commonly employs queue structures, buffer pool structures, or graph database structures.

[0078] The goal of updating parameters in dynamic causal graphs is to improve the accuracy and stability of causal paths between states, making the expected states generated based on these paths more closely resemble the evolutionary trajectories in the real environment. The parameters of a causal graph mainly include node state transition probabilities, edge weights, and time delay functions. These parameters are updated by minimizing the deviation between the actual state evolution and the predicted graph path.

[0079] The hierarchical reinforcement learning model consists of a high-level policy network and a low-level policy network, which respectively learn sub-objective generation and action parameter optimization. The high-level policy network learns the effectiveness of sub-objective setting and is responsible for achieving the global goal; the low-level policy network focuses on executing fine-grained actions in specific contexts and is responsible for the efficiency of policy execution. The parameters of both networks need to be optimized through gradient backpropagation or policy update based on the policy reward signal.

[0080] Model performance changes are a key factor guiding the joint training process and need to be evaluated in real time using signals such as structural stability metrics, policy convergence metrics, or reward volatility. When phenomena such as slow performance improvement, unstable rewards, or significant causal path deviations are detected, the training optimization direction should be adjusted. For example, increasing the learning rate of higher-level systems to improve the ability to adjust the objective, or converging the lower-level policies to improve execution stability.

[0081] Parameter updates in joint training can be implemented using phased, alternating, or fusion loss functions. In phased mode, the causal graph is fixed first, and only the hierarchical policy is optimized, then the causal graph is optimized again with the policy fixed. In alternating update mode, both models are updated once per round. In fusion loss mode, the prediction error loss and the policy optimization objective are combined into a unified training objective.

[0082] Empirical data can be stored as triples (state, action, reward) and expanded to quintuples (state, action, sub-goal, state change, reward) to simultaneously drive the optimization process of two models. During causal graph optimization, variational inference algorithms or graph attention network-based methods can be used to probabilistically model the paths between states, thereby updating edge connection probabilities and path selection weights. A structural penalty mechanism can also be introduced to update the causal graph structure to prevent overfitting or graph structure oscillations.

[0083] In the reinforcement learning part, reinforcement learning algorithms such as PPO, DDPG, or TD3 can be used to model the high-level and low-level policies separately. The high-level network outputs sub-objective embeddings, while the low-level network receives the joint encoding of the environment state and the sub-objective as input and outputs continuous action parameters. The reward signal is used in the objective function of the two-layer network, for example, to construct the policy gradient in the form of a weighted advantage function.

[0084] In terms of performance tuning mechanisms, metrics such as reward stability (e.g., moving average difference rate) and model performance growth rate can be introduced. By controlling training hyperparameters, adaptive sampling frequency, or adjusting model structural complexity, a balance between training stability and policy interpretability can be achieved.

[0085] Example Description: In the healthcare field, intelligent guidance devices use recorded status changes and operational feedback during surgery to create training samples. Key operations such as resection and suturing, along with their subsequent effects, are used as data to optimize causal models and control strategies. When strategy performance is unstable, the system automatically adjusts the causal path modeling method and strategy network hyperparameters to improve operational accuracy and adaptability.

[0086] In the fintech business, credit approval systems continuously optimize risk causal graphs and approval strategies through joint training. The system records customer behavior, approval decisions, and subsequent risk outcomes, forming an experience set that includes status, actions, and rewards, driving updates to the risk causal structure and approval operation strategies. When data feedback is insufficient or risk predictions deviate, the system can dynamically adjust the update ratio and model complexity to enhance the robustness and interpretability of the strategy.

[0087] This embodiment introduces a joint training mechanism to achieve a high degree of synergy between causal modeling and policy optimization at the data usage and learning objective levels, avoiding the accumulation of biases and information waste caused by isolated model training. Updates to the causal graph parameters enhance the structural expressiveness of state evolution prediction, providing more structurally constrained support for reward signal generation; updates to the hierarchical policy network parameters improve the efficiency and accuracy of action generation. Joint optimization further maximizes data utilization efficiency, contributing to rapid policy convergence and improved generalization capabilities.

[0088] S60 generates optimized action policies through an updated hierarchical reinforcement learning model.

[0089] In this embodiment, the process of generating the optimized action policy relies on the updated hierarchical reinforcement learning model, and is divided into three parts: high-level policy generates goal-oriented instructions, goal-oriented instructions generate optimized sub-goals, low-level policy generates action parameters based on optimized sub-goals, constructs action sequences, verifies compatibility, and outputs the policy. Current environment state information includes continuous or discrete state variables describing the current system environment, originating from the perception system, operational feedback, or scene data updates. This state information maintains consistency in its representational dimension with the environment seen during the model's previous training phases to ensure policy transfer and generalization capabilities.

[0090] The high-level policy model takes the current environmental state as input and outputs a goal-oriented instruction. This instruction can be an embedded vector, task symbol, or operation cue, used to guide subsequent goal decomposition. This instruction possesses directionality, abstraction, and task-driven characteristics, and can represent the global goal intent. Optimization sub-goals are intermediate task decompositions performed under the guidance of the high-level policy, combined with the environmental state. They are typically represented as secondary goal vectors or expected local state changes, used to constrain the range of low-level policy generation.

[0091] The underlying strategy model receives joint inputs of environmental state information and optimization sub-objectives to generate action execution parameters. These parameters cover dimensions such as position offset, velocity control, angle change, and mechanical adjustment, possessing continuity and high precision characteristics, and are used to drive the execution module to perform micro-actions. The action execution parameters are then transformed into a sequence of executable actions with temporal consistency through action splicing and time scheduling processes.

[0092] To ensure the rationality and security of the generated action strategy, a compatibility verification operation is introduced. This verification, based on the established dynamic causal graph structure, performs path evolution reasoning on the expected results of the action sequence to determine whether the sequence is consistent with the state evolution logic. Compatibility verification includes path similarity calculation, causal conflict detection, and state offset evaluation to ensure that the generated strategy will not lead to uncontrollable behavior during future state evolution. Finally, the verified action sequence is used as the output of the optimized action strategy to control the execution module to complete the target task.

[0093] A high-level policy model constructed using a graph neural network can be used in the implementation. The environment state is encoded as node embeddings, and the goal-oriented instructions are propagated as global control signals in the graph structure. Then, an attention mechanism aggregates the node information to form sub-goals. During the sub-goal generation process, a task memory mechanism or external knowledge guidance can be introduced to stabilize goal setting.

[0094] The underlying strategy model can employ a multi-input neural control architecture, embedding sub-goals and environmental states separately and then concatenating them into an input control network. This network outputs multi-dimensional continuous parameters that control the position, velocity, or stiffness of the end effector. The construction process of the executable action sequence can integrate a time window sliding mechanism, ensuring the output actions possess both execution continuity and temporal rationality.

[0095] Compatibility verification can employ a graph-based prediction system, using the structure in a dynamic causal graph to predict subsequent state sequences and comparing the differences with the predicted states after action execution. Constraint mechanisms such as maximum offset thresholds and structural consistency measures can be set; if these are violated, the use of the policy sequence is rejected. Finally, the action policy is input into the execution unit through a scheduling system to implement instruction issuance.

[0096] Example: In the healthcare field, after completing reinforcement learning model training, the surgical robot generates optimized action strategies based on the current real-time patient status (such as endoscopic images and tissue stiffness feedback). The high-level strategy instructs "complete vascular avoidance suturing," with the sub-objective being "locating the suture start and end points." The low-level strategy generates specific surgical operation parameters. The generated action sequence is then verified through a causal graph to check for any risk of tissue damage. After ensuring the safety of the execution path, it is sent to the robot control unit.

[0097] In the fintech business, after acquiring customer transaction behavior status information, intelligent auditing robots generate target-oriented instructions such as "identifying potential fraud patterns" through high-level strategies, and then generate sub-objectives such as "locating high-frequency trading accounts" and "confirming unauthorized payment paths." The underlying strategies generate monitoring indicators and behavioral analysis paths. The resulting sequence of analysis operations is verified through a dynamic causal graph to determine whether it may trigger systemic risk assessment biases. Only strategies that pass verification are used for automated decision-making and report generation.

[0098] This embodiment introduces an updated hierarchical reinforcement learning structure, which aligns the policy execution process with the global objective. The higher-level layer controls the task direction, while the lower-level layer provides fine-grained control; the two are decoupled, optimized, and collaborative. Before action policies are generated, causal graph compatibility verification is introduced, enhancing the stability and safety of the policy in the state evolution path and preventing overfitting, policy jumps, or unexpected state transitions. The policy sequence possesses clear goal guidance, a hierarchical structure, and structural consistency guarantees, significantly improving task execution efficiency, success rate, and interpretability.

[0099] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a policy generation method, apparatus, device, and medium based on hierarchical reinforcement learning, comprising: acquiring multi-source raw scene data and preprocessing it to generate a structured scene dataset; constructing a dynamic causal graph based on the structured scene dataset; receiving environmental state information and generating sub-goals and specific actions through a hierarchical reinforcement learning architecture; acquiring state change information after the execution of specific actions and combining the environmental state information, state change information, specific actions, and sub-goals to generate a policy reward signal; jointly training and updating the dynamic causal graph and hierarchical reinforcement learning model parameters based on the policy reward signal; and using the updated hierarchical reinforcement learning model to generate an optimized action policy. This invention models the state evolution relationship by constructing a dynamic causal graph, enabling reinforcement learning to gain a causal understanding of state change trends. Combined with a hierarchical reinforcement learning architecture, it achieves the decomposition and optimization of sub-goals and actions, improving the response accuracy and generalization ability of action policies in complex environments, thereby improving the stability of task completion and the convergence efficiency of policy training.

[0100] In one embodiment, step S10 above includes:

[0101] S101 collects multi-source raw scene data, including visual data, sensor data, and structured table data;

[0102] S102, Clean the multi-source original scene data to remove noise and outliers, and generate cleaned multi-source original scene data;

[0103] S103, Align the time series dimension in the cleaned multi-source original scene data to generate time-aligned multi-source original scene data;

[0104] S104, standardize the time-aligned multi-source original scene data of different modalities to generate standardized multi-source original scene data;

[0105] S105, extract features from the standardized multi-source original scene data to generate a feature set;

[0106] S106, fuse the feature set to generate a fused feature representation;

[0107] S107, Generate a structured scene dataset based on the fused feature representation.

[0108] In this embodiment, the process of generating a structured scene dataset begins with the acquisition of multi-source raw scene data, encompassing three types of data sources: visual data, sensor data, and structured tabular data. Visual data mainly refers to two-dimensional or three-dimensional perceptual inputs such as images, video sequences, and depth maps, typically acquired through cameras, infrared imaging devices, etc., possessing unstructured characteristics and spatial semantic density. Sensor data includes acceleration and gyroscope signals acquired by inertial measurement units (IMUs), tactile feedback and electrophysiological signals output by pressure sensors, and other types of continuous time-series signals; this type of data features high temporal resolution and continuity. Structured tabular data originates from status indicators and event logs recorded by environmental management systems or external interfaces, featuring predefined fields and data constraints, suitable for supplementing state variables and labeling event boundaries. These three types of data sources differ significantly in acquisition dimensions, sampling frequency, and semantic structure, requiring unified representation through downstream processing steps.

[0109] After data acquisition, the original multi-source scene data needs to be cleaned. This step aims to remove noise and outliers introduced by factors such as hardware errors, environmental interference, and data loss. Visual data cleaning can incorporate image quality evaluation metrics such as sharpness, brightness distribution, and compression artifact levels for filtering, and incomplete areas can be repaired using image reconstruction networks. For sensor data, extreme values ​​that do not conform to physical laws are detected using statistical filtering and sliding window outlier detection methods, and missing intervals are repaired based on Kalman filtering or time series interpolation methods. For structured tabular data, data consistency and integrity are restored through field validity checks, business rule verification, and missing value imputation. The cleaned data maintains the original data source structure but possesses higher usability and robustness, supporting subsequent analysis.

[0110] Even after cleaning, the multi-source data still exhibits inconsistencies in the time dimension, requiring unification through a time series alignment mechanism. First, the timestamp format and sampling frequency of all data sources are identified, and a reference time base is selected as the alignment target, such as a millisecond-level event timeline or a unified starting point offset. Low-frequency data is aligned using interpolation, such as spline interpolation or local linear regression to predict intermediate values; high-frequency data employs aggregated downsampling methods to compress redundant samples, such as sliding window averaging or max pooling. The alignment result ensures that multi-source data form corresponding sample points on the same time index, achieving synchronous fusion of asynchronous observations and providing time consistency support for intermodal interaction modeling.

[0111] After time alignment, data from different modalities needs to be standardized to address the impact of inconsistent dimensions and numerical skewness on subsequent modeling. For visual modalities, pixel values ​​can be normalized to the [0,1] range to enhance network input stability. For sensor modalities, Z-score normalization or Min-Max scaling is used to eliminate dimensional differences and improve gradient propagation efficiency. Categorical fields in structured tabular data are represented using one-hot encoding or embedding encoding, while numerical fields are uniformly converted to standardized floating-point numbers. The standardization process should preserve the semantic structure and relative trends of the original data to avoid information loss.

[0112] The feature extraction process is based on standardized multi-source inputs, with extraction mechanisms designed separately for each modality. Visual modalities typically use convolutional neural networks or visual Transformer structures to extract local spatial and global semantic features, outputting fixed-length vectors or tensors. Sensor modalities employ one-dimensional convolutional networks, recurrent neural networks, or transformer structures to extract dynamic patterns from time series data. Structured data can be extracted using fully connected networks or graph embedding mechanisms to extract inter-field dependencies. The extraction results are summarized into a multimodal feature set, preserving independent semantics between modalities.

[0113] The generation of fused feature representations is based on the above feature set through intermodal integration. Common fusion strategies include feature concatenation, attention-weighted fusion, and cross-transformation modules. The fusion mechanism needs to retain the important information of each modality and enhance their collaborative ability in a unified representation space. For example, gating mechanisms can be introduced to control the contribution ratio of each modality to the final representation, or modality consistency loss can be used to optimize the collaborative representation ability of modalities. As a high-dimensional vector structure, the fused feature representation contains multi-dimensional contextual information such as time, space, and category, and also possesses a unified embedding structure, making it suitable for downstream modeling tasks.

[0114] The structured scene dataset is generated based on the extraction and format reconstruction of fused feature representations. The fused features at each time step are combined with their corresponding labels, identifiers, task states, etc., to form structured samples with temporal order, multimodal expression, and task semantics. This dataset is scalable and generalizable, supporting multi-task joint modeling and transfer learning applications.

[0115] This embodiment collects raw data covering three heterogeneous modalities: visual, sensor, and structured tables. Preprocessing operations such as cleaning, time alignment, and standardization are then performed sequentially. This effectively addresses issues related to sampling frequency, noise interference, format differences, and timestamp asynchrony in multi-source data, enhancing the integration capability of the raw data. Furthermore, through feature extraction and modality fusion, the multimodal raw observation information is transformed into a unified semantic representation, achieving efficient mapping from raw data to modeling data. The resulting structured scene dataset possesses high-dimensional semantic information, temporal continuity, and task context relevance, providing stable, unified, and high-quality data support for subsequent causal modeling and policy optimization, significantly improving overall modeling accuracy and policy generalization ability.

[0116] In one embodiment, step S20 above includes:

[0117] S201, Analyze the structured scene dataset to identify the set of state variables;

[0118] S202, perform causal discovery based on the set of state variables, and generate a causal structure model;

[0119] S203, Perform time delay embedding processing on the time series data in the structured scene dataset to generate time delay embedded data;

[0120] S204, Integrate the causal structure model and time delay embedded data to construct an initial causal graph;

[0121] S205, Optimize the initial causal graph to improve the state evolution prediction accuracy, and generate an optimized causal graph;

[0122] S206, verify the consistency between the optimized causal graph and the structured scene dataset, and generate a dynamic causal graph.

[0123] In this embodiment, the process of constructing a dynamic causal graph to describe the causal relationships of state evolution relies on the systematic analysis and layer-by-layer modeling of the structured scene dataset. The aim is to extract the stable causal dependencies between state variables in the time series and form a graph structure representation that can dynamically adapt to state evolution.

[0124] First, the structured scenario dataset needs to be analyzed to identify the set of state variables. Structured scenario datasets typically contain multi-dimensional feature representations across multiple time steps, with each feature potentially corresponding to a system state or external intervention variable. The identification process for state variables includes field parsing, semantic clustering, and correlation filtering. Field parsing uses field names, units, types, and source identifiers to initially filter out non-state fields such as labels and IDs that lack dynamic features. Semantic clustering combines field meaning and change patterns for grouping; for example, word vector embedding can be used to merge multiple data fields representing the same physical quantity (e.g., temperature, pressure) but expressing different values. Correlation filtering uses metrics such as mutual information, correlation coefficient matrices, or minimum description length to remove variables dominated by noise or irrelevant fluctuations, retaining state variables that are highly correlated with behavioral outcomes and have dynamic system significance. The resulting set of state variables provides the basic set of nodes for causal modeling, constituting the node space of the graph.

[0125] Causal discovery is performed based on the aforementioned set of state variables to generate a causal structure model. The goal of causal discovery is to identify directional dependency paths between state variables, i.e., the temporal and logical "cause-effect" relationships between variables. Causal discovery methods may include structural equation modeling (SEM), constraint-based PC algorithms, scoring-based GIES algorithms, and neural causal discovery methods based on adversarial perturbations and interpretability mechanisms. For time-series state variables, Granger causality tests, attention weight alignment, or deep residual modeling can be combined to extract the time lag and directional strength of causal relationships. Stability constraints and reconfigurability loss terms are introduced during model training to ensure consistency of the generated structure across multiple subsamples. The final causal structure model represents the static causal topology between state variables in the form of a directed acyclic graph (DAG).

[0126] Considering that causal relationships between state variables often have time dependencies, time-delay embedding is required for time-series data in structured scenario datasets. The purpose of time-delay embedding is to explicitly introduce the historical state of a single variable into the modeling process to construct the lag structure in the state transition path. This process can be achieved by using Takens embedding theory to merge the current value of each variable with data from multiple historical time steps into a high-dimensional vector. The embedding dimension and lag step size can be determined by the descent point of the mutual information curve or a pseudo-autocorrelation function. The embedded data retains the dynamic trajectory information of the variables themselves, providing an informational basis for capturing delayed causal effects between variables. The embedding vectors of multiple variables can be further concatenated or stacked for joint modeling.

[0127] By integrating the aforementioned causal structure model with time-delayed embedded data, an initial causal graph is constructed. The integration process requires expanding the embedded data into a causal graph, incorporating the historical state of each variable as an independent node into the graph structure and connecting it to the current state of the target variable via lagged causal edges. This expanded graph structure simultaneously contains the directional connections between variables in a static structure and the temporal evolution relationships in a dynamic path, possessing the ability to capture both lateral causal and longitudinal temporal paths. The graph structure can be parameterized within a unified graph neural network framework, and a graph attention mechanism can be introduced to model the importance of causal edges.

[0128] To improve the accuracy of state evolution prediction, joint optimization of the structure and parameters of the initial causal graph is required. Structure optimization can remove redundant connections through edge weight sparsification, path pruning, and residual graph construction to avoid overfitting of causal paths. Parameter optimization includes end-to-end learning of node state propagation weights, causal edge transformation matrices, and state aggregation functions. The training objective is to minimize the deviation between the current state prediction and the actual observation or to maximize the consistency score of the structure graph. The optimization process can incorporate supervisory signals, such as indirect supervision through comparative prediction loss of state evolution, or generative adversarial strategies can be used to verify the distribution consistency between the graph output and the actual evolution.

[0129] After structural optimization, the optimized causal graph needs to be validated to ensure consistency with the structured scenario dataset. The validation process includes graph structure interpretability assessment, in-sample and out-of-sample prediction consistency testing, and recoverability-based inversion testing. If the predicted result of a certain state variable deviates significantly from its actual observed value, it is necessary to trace back to identify any weak edges or erroneous dependencies in its causal path and adjust the graph structure using Bayesian optimization or graph structure reinforcement learning mechanisms. The final validated causal graph, possessing long-term predictive power, causal interpretability, and dynamic adaptability, is defined as a dynamic causal graph, suitable for subsequent reward generation and policy optimization processes.

[0130] This embodiment identifies a set of state variables and constructs a causal structure model. It introduces time delay embedding to capture temporal lag relationships and generates a dynamic causal graph with the support of structural optimization and consistency verification mechanisms. This effectively solves problems such as sparse causal relationships, uninterpretable structure, and implicit temporal dependencies in high-dimensional time-series data. The dynamic causal graph possesses structural representation and generalization capabilities in modeling the evolutionary paths and behavioral influences between variables, providing a state transition reference with prior structure for subsequent policy learning, and significantly improving causal consistency and policy inference efficiency.

[0131] In one embodiment, step S30 above includes:

[0132] S301, receives environmental status information;

[0133] S302, The environmental state information is processed by a high-level strategy of a hierarchical reinforcement learning architecture to generate high-level decision instructions;

[0134] S303, Generate sub-targets based on the high-level decision-making instructions;

[0135] S304, The environmental state information and sub-objectives are processed through the underlying strategy of the hierarchical reinforcement learning architecture to generate action execution parameters;

[0136] S305, Generate a specific action based on the action execution parameters.

[0137] In this embodiment, to adapt to complex scenarios in dynamic environments where task objectives change frequently, state response times are sensitive, and action output accuracy requirements are extremely high, a reinforcement learning architecture with hierarchical decision-making capabilities needs to be constructed. This architecture effectively decomposes the coupling relationship between long-term tasks and immediate responses by distinguishing between high-level abstract goal setting and low-level specific action output, thereby improving the generalization and response stability of the policy.

[0138] First, environmental state information is received. This information is typically collected in real-time or periodically by multiple sensors, vision devices, scene recognition modules, or historical state trackers. The information can take the form of image frames, point clouds, pose matrices, structured index vectors, or state embeddings. To accommodate different modal input formats, the environmental state information is uniformly mapped to a continuous vector space before input, such as by extracting image state embeddings through convolutional networks or aggregating multimodal state summaries through transformer structures. During the input process, normalization mechanisms and historical state caching structures can also be introduced to enhance the consistency and temporal coherence of state representation.

[0139] Upon receiving environmental state information, it is processed through high-level policies within the hierarchical reinforcement learning architecture. The high-level policy module aims to analyze the macroscopic characteristics of the environmental state from a global task perspective and formulate long-term goal-oriented policy instructions. High-level policies typically employ policy gradient methods, model predictive control (MPC), or value function approximation mechanisms, outputting abstract policy commands with planning attributes, such as target displacement regions, operation target sequences, or task priority ranking. These policy instructions do not directly act on the execution system but guide lower-level policies to concretize actions. The input to the high-level policy is not limited to the current state information but can also incorporate historical observation trajectories, completed sub-task states, and contextual background provided by external knowledge bases.

[0140] Sub-objectives are generated based on decision instructions output by high-level policies. A sub-objective refers to a sub-task expression generated during the high-level task decomposition process, possessing executability boundaries and applicable to a single policy optimization unit. Sub-objective generation methods include rule-driven instruction translation, structured semantic decoder transformation, or mapping high-level instructions to low-dimensional task target vectors through parameterized transformation. For example, in navigation tasks, the "go to a specified area" instruction given by the high-level policy can be converted into a two-dimensional target point in the current coordinate system. In interactive tasks, "perform a fine-grasping operation" can be converted into gripper posture and force targets.

[0141] After setting the sub-objectives, the environmental state information and the sub-objectives are input into the underlying policy of the hierarchical reinforcement learning architecture. The underlying policy focuses on the generation of specific actions, and its input state space is more granular and constrained by the current task stage. The underlying policy uses an end-to-end policy network, a value function-driven action distribution sampler, or a hybrid action graph optimization model to generate action execution parameters by combining the current perception state and the target vector. The action execution parameters here may include multi-dimensional continuous control commands (such as changes in the joint angles of a multi-degree-of-freedom robotic arm), discrete action selection probabilities (such as menu item switching), or hybrid control sequences (such as parallel output of decision tree paths and PID control quantities). To improve the sensitivity of the underlying policy to sub-objectives, soft constraint regularization terms can be introduced or a target conditional encoding mechanism can be used to embed the target representation within the policy network.

[0142] The specific actions are ultimately generated based on the action execution parameters output by the underlying strategy. Generating specific actions refers to converting the control signals output by the strategy into action instruction formats acceptable to the physical system, including but not limited to: converting continuous action vectors into drive voltages or servo control signals for the mechanical execution system; mapping discrete actions into command byte sequences and transmitting them to the control interface; or mapping a graphical action plan table into a task instruction stream and pushing it to a multi-agent system. This process requires consideration of physical constraints on the equipment (such as maximum load, response time, and action boundaries) and safety mechanisms (such as rollback for erroneous operations, limit detection, and soft stop).

[0143] Through the tight coupling between high-level strategies, sub-goal generation, low-level strategies, and action generation, the entire system realizes a continuous mapping process from macro-level task planning to micro-level action control. It effectively compresses the state space, decomposes the strategy and goal, and decouples the execution mechanism, providing a highly adaptable and robust action generation process for environments with high complexity, multi-dimensional states, and frequent multi-goal switching.

[0144] This embodiment introduces a layered architecture of high-level and low-level strategies, enabling the system to establish a dynamic mapping channel between global task planning and local fine-grained action control. High-level strategies are responsible for long-term goal planning, enhancing the global consistency and foresight of task completion, while low-level strategies focus on short-term response and fine-grained control, improving the accuracy and stability of action execution. Sub-goals, acting as an intermediary structure, establish a transition layer between abstract task instructions and executed actions, effectively reducing the dimensionality of the strategy search space and improving strategy learning efficiency.

[0145] In one embodiment, step S40 above includes:

[0146] S401, Obtain the state change information after the specific action is executed;

[0147] S402, using the dynamic cause-effect graph, based on the environmental state information, specific actions, and sub-goals, determine the expected state evolution path;

[0148] S403, detect the degree of conformity between the state change information and the expected state evolution path, and generate a state change difference index;

[0149] S404, Generate an intrinsic reward component based on the state change difference index;

[0150] S405, generates external reward components based on task completion progress;

[0151] S406, determine the weight ratio of the intrinsic reward component and the extrinsic reward component, and dynamically adjust the weight ratio according to the current training stage;

[0152] S407, Based on the adjusted weight ratio, the intrinsic reward component and the extrinsic reward component are weighted and fused to generate a strategy reward signal.

[0153] In this embodiment, after a specific action is output by the system and acts on the environment, the state of the external environment will change. To achieve accurate perception and strategy evaluation of the impact after the action is executed, it is first necessary to acquire information on the state change after the action is executed. This information can be acquired in real time through a continuous state acquisition module or indirectly inferred through a state estimation module. The acquired data includes, but is not limited to, sensor outputs, image sequence changes, and environmental indicator values. To improve the consistency of data comparison, the state change information and the environmental state information before the action is executed can be timestamped and modally unified, making the state change performance comparable.

[0154] Next, the system needs to utilize dynamic causal graphs to deduce the expected state evolution path that the action may cause in the current context. A dynamic causal graph is a structured causal modeling mechanism where nodes represent state variables, edges represent causal relationships, and it possesses the property of temporal evolution. By inputting current environmental state information, specific actions, and sub-goals, the system can search for paths from the initial state to the target state in the causal graph, or predict intermediate state change trends through graph embedding models and causal propagation mechanisms, thereby constructing the expected state evolution path. This process emphasizes the consistency between causal logic and the evolutionary structure of state variables, considering not only direct effects but also the changing trends of intermediate state variables involved in indirect links.

[0155] After constructing the expected state evolution path, it is compared with the actual state change information collected after the action is executed. The comparison can be based on Euclidean distance, KL divergence, distribution overlap, or structural similarity indices of state variable values ​​to obtain the degree of difference in state changes. The difference calculation process is not limited to the final state; trajectory alignment techniques can also be introduced to measure the consistency of intermediate states. The comparison results are expressed in the form of a state change difference index, providing accuracy support for subsequent reward signals.

[0156] The system then generates an intrinsic reward component based on the difference index of state changes. This reward component measures the degree to which the policy itself conforms to the law of state evolution, representing the accuracy of the model's intrinsic causal modeling and the rationality of the action. The intrinsic reward often adopts a negative correlation structure, with higher rewards for smaller differences, to encourage the policy to minimize errors in its modeling consistency with the system. Its construction methods may include backpropagation error, causal path entropy reduction, and state mapping consistency evaluation.

[0157] On the other hand, the system also generates extrinsic reward components based on task completion progress. Extrinsic rewards focus more on task-level performance feedback, such as the degree of achievement of sub-goals, the distance between the final state and the target state, task duration, and resource consumption. Progress metrics can be calculated through an objective function or rely on an external evaluation index system. Their output forms the extrinsic reward, used to guide the strategy towards convergence towards the task objective.

[0158] The system then determines the weight ratio between intrinsic and extrinsic reward components. Due to different training phases, the emphasis on model structure fitting and behavioral goal achievement varies. For example, in the early stages of policy implementation, it is necessary to strengthen causal consistency training, which can assign higher weights to intrinsic rewards; in the later stages of policy stabilization, the weight of extrinsic rewards can be gradually increased to optimize behavioral performance. Therefore, the system needs to design a dynamic weight adjustment mechanism to update the weight ratio function based on indicators such as training epochs, loss convergence speed, and task completion frequency.

[0159] Finally, the system weights and fuses the adjusted intrinsic and extrinsic reward components to generate a policy reward signal. This signal serves as the feedback target for the joint optimization of the policy network and causal graph, participating in subsequent gradient backpropagation and parameter update processes, thereby achieving synergistic driving of adaptive optimization of model behavior and structure fitting.

[0160] This embodiment introduces a reward signal generation mechanism based on causal path modeling, enabling a comprehensive evaluation of policy behavior from two dimensions: structural prediction accuracy and task completion. This overcomes the slow convergence and policy fluctuation problems caused by relying solely on task feedback in traditional reinforcement learning. Intrinsic rewards provide structural consistency supervision, effectively guiding the policy to maintain the accuracy of its modeling of state evolution patterns. Extrinsic rewards ensure that behavior is directed towards the true target task, thereby achieving efficient convergence in the policy optimization process and accurate correction in the causal modeling process. The dynamic weight mechanism endows the system with the ability to adapt to different training stages, improving generalization ability and the stability of the final policy.

[0161] In one embodiment, step S50 above includes:

[0162] S501 stores empirical data including environmental state information, specific actions, policy reward signals, state change information, and sub-objectives.

[0163] S502, Optimize the parameters of the dynamic causal graph based on the empirical data to improve prediction consistency;

[0164] S503, Optimize the parameters of the hierarchical reinforcement learning model based on the empirical data to maximize the cumulative reward;

[0165] S504, detect the performance changes of the optimized hierarchical reinforcement learning model, and adjust the model optimization direction based on the performance changes;

[0166] S505, based on the adjusted optimization direction, iteratively update the parameters of the dynamic causal graph and the hierarchical reinforcement learning model.

[0167] In this embodiment, to achieve a dual improvement in both environmental adaptability and policy stability, the feedback information generated during execution needs to be transformed into training data that can be used for optimization. First, empirical data is constructed by uniformly collecting key variables from each round of interaction. This empirical data includes environmental state information, specific actions, policy reward signals, state change information, and sub-objectives, covering multiple dimensions such as input, execution, feedback, and goal setting. To ensure data quality, an experience replay mechanism can be used to filter and denoise the collected samples, or a priority caching mechanism can be constructed to strengthen the weight of representative samples, thereby improving the stability and representativeness of the optimization process.

[0168] For updating dynamic causal graphs, the system needs to optimize structural edge weights, time delay parameters, and causal path directions based on collected empirical data. To improve the consistency between predicted state evolution paths and actual paths, a structural regression approach based on minimum structural entropy can be used, or a graph neural network can be used to calculate the joint distribution changes of state pairs and feed them back to the structural layer for adjustment. Incorporating the intrinsic components of the policy reward signal during optimization helps identify which causal paths play a key role in target transition, thereby improving the alignment accuracy between causal paths and actual evolution paths.

[0169] For updating the hierarchical reinforcement learning model, the system can optimize the parameters of both the high-level and low-level policy networks separately. Based on the policy reward signals in empirical data, and with the objective function of maximizing long-term cumulative reward, reinforcement learning algorithms such as policy gradient, proximal policy optimization (PPO), or distributed deep Q-networks are used to perform parameter updates. The high-level policy mainly adjusts the goal-setting logic, such as the sub-goal decomposition method and abstraction granularity, while the low-level policy mainly adjusts the action output accuracy and response speed. During training, the model needs to continuously sample from empirical data and perform policy replay, combining the reward signals and policy outputs to calculate gradients and iterate parameters.

[0170] To ensure the effectiveness of the training convergence direction, the system needs to monitor the performance changes of the hierarchical reinforcement learning model in real time after optimization. Performance monitoring can be based on evaluation metrics such as task completion rate, average reward, state transition error, and sub-objective achievement rate. By comparing the performance differences between the current training cycle and the historical baseline model, it can identify whether there are problems such as overfitting, vanishing gradients, or policy oscillations in the current parameters.

[0171] Based on performance testing results, the system adjusts its optimization direction. If the performance of the current strategy declines, the parameters need to be rolled back to the previous stable state or switched to a lower learning rate for fine-tuning; if the prediction error of the causal graph increases, the causal edges need to be regularized or dynamically pruned. This adjustment process is driven by performance feedback and constructs an adaptive training path control mechanism.

[0172] Finally, based on a clear optimization direction, the parameters of the dynamic causal graph and the hierarchical reinforcement learning model are iteratively updated to simultaneously improve the accuracy of causal modeling and the quality of policy output. The joint training of the two models is not conducted in isolation, but rather through shared policy reward signals and a unified empirical data stream, achieving information synergy and forming an end-to-end differentiable structure-behavior coupled optimization framework.

[0173] This embodiment constructs empirical data encompassing multidimensional inputs and feedback, and synchronously updates the causal structure and behavioral strategies based on this data. This enables the model to not only improve the real-time performance and accuracy of action decisions but also continuously adjust its cognitive expression of the evolutionary patterns of the environment. Causal graph optimization enhances the modeling ability of the dynamic structure of the environment, while hierarchical policy updates improve the flexibility and adaptability of task execution. Joint training of both eliminates the disconnect between modeling and behavior, significantly improving the system's convergence speed, stability, and generalization ability. The model performance detection and optimization path adjustment mechanism introduces dynamic feedback control into the training process, improving training efficiency and effectively preventing policy degradation or causal structure shift.

[0174] In one embodiment, step S60 above includes:

[0175] S601, obtain current environmental status information;

[0176] S602, The current environment state information is processed by the high-level policy of the updated hierarchical reinforcement learning model to generate a target-oriented instruction;

[0177] S603, Generate optimized sub-targets based on the target-oriented instructions;

[0178] S604, The current environment state information and optimization sub-objective are processed through the underlying policy of the updated hierarchical reinforcement learning model to generate optimized action parameters;

[0179] S605, Generate an executable action sequence based on the optimized action parameters;

[0180] S606, Verify the compatibility between the executable action sequence and the dynamic cause-effect graph;

[0181] S607 outputs the validated executable action sequence as an optimized action strategy.

[0182] In this embodiment, during the action decision generation phase, the system receives and parses the current environmental state information based on the trained hierarchical reinforcement learning structure. The environmental state information indicates the observed values ​​of various key variables in the current scene, covering visual input, sensor readings, task status indicators, and external constraint parameters. This state information can constitute a multidimensional tensor or a structured feature map. To adapt to the input formats of different policy modules, the system needs to perform unified vectorization or embedding mapping processing on this state information, while maintaining consistency with the feature encoding structure from the training phase.

[0183] The embedded environmental state information is first input to the high-level policy module. This high-level policy module is a policy decision network with abstraction capabilities, whose role is to identify behavioral intentions with global guiding significance from the current environmental state. Through this policy network, a set of goal-oriented instructions can be generated, representing the abstract goals or intentional states to be achieved in the current task phase. Goal-oriented instructions can be logical tags, semantic codes, or structured goal templates, and their form is strongly correlated with the task type. In complex tasks, graph attention networks can be used to extract semantic hierarchical information or combined with behavioral graphs to generate medium- and long-term goals.

[0184] Subsequently, the goal-oriented instructions are transformed into a sub-goal structure adapted to the underlying execution module through a sub-goal generation mechanism. This transformation process can be implemented through analytical mapping functions or shallow neural networks, ensuring that the high-level abstract goal has a clear execution decomposition method in the spatial, temporal, and operational dimensions, and finally outputs a target variable vector or specific sub-task label for lower-level action decisions.

[0185] The generated optimized sub-objective and the current environmental state information are input into the underlying policy module. This module focuses on specific action generation, and its input dimensions typically include continuous state variables and discrete target variables, while the output is a vector of action parameters. The underlying policy network can be implemented using algorithms such as Deep Deterministic Policy Gradient (DDPG), Proximal Policy Optimization (PPO), or Soft Behavior Policy (SAC), performing policy retrieval and action construction in the input space. The optimized action parameters can be structurally represented as a range of values ​​for continuous actions (such as the angular velocity and position coordinates of a robotic arm) or as a combination of multiple discrete actions (such as multi-arm collaboration and tool switching).

[0186] Subsequently, the system constructs an executable action sequence based on optimized action parameters. The construction methods can be divided into two categories: parallel generation and temporal concatenation. The former is suitable for generating all action elements at once, while the latter is suitable for execution strategies based on predictive feedback and gradual generation. Each action sequence needs to contain complete control instruction formats and timing information for easy deployment and execution.

[0187] After action generation, to ensure consistency between the behavioral output and causal modeling, the executable action sequence needs to be input into the compatibility verification module. This module matches the expected state change trajectory under the generated action with the state evolution path in the current dynamic causal graph, evaluating its reachability and rationality on the causal path. The verification mechanism can be implemented through path alignment, state transition scoring, or a structured path matching network. If significant deviations exist, the higher-level policy can be guided to regenerate goal-oriented instructions for iterative correction.

[0188] Finally, the system outputs the action sequence that has passed causal consistency verification as the optimized action strategy. This strategy contains a complete set of action instructions, target location information, and their context state dependencies. It is deployable and highly adaptable, and can achieve adaptive, efficient, and robust task completion in actual operation.

[0189] Example Description: In an intelligent surgical assistance system, the robot needs to perform complex soft tissue suturing operations. The task requires it to autonomously plan needle entry points and suturing paths on irregular soft tissue surfaces, responding in real-time to changes in tissue stretching, vascular interference, and other conditions, avoiding accidental injury and ensuring suturing strength and precision. To achieve this task, the system first collects multi-source raw scene data, including: surgical visual image streams (such as endoscopic image sequences), multimodal tactile sensor data (such as tissue tension and pressure sensing), and structured preoperative patient data (such as 3D reconstruction models of the surgical site and individual anatomical annotations). The system cleans these raw data to remove signal drift and occlusion distortion, and then synchronizes the multimodal data based on timestamps to ensure complete alignment of visual actions and force feedback. Subsequently, standardization operations (such as grayscale normalization and tension unit unification) are performed to extract key feature points and surface state descriptions, forming a fusion feature representation that expresses tissue deformation and force feedback response. Finally, a structured scene dataset is constructed to characterize the suturing operation space.

[0190] The system further analyzes the structured scene dataset, extracting state variables such as tissue edge states, suture tension change curves, and key point tension propagation paths. It uses a graph-based causal discovery method to identify the dependencies between states (e.g., "tension change" affects "suture point displacement"), generating an initial causal structure graph. To supplement the temporal influence of tissue states after the action is applied, the system performs delayed embedding on the time series (e.g., introducing historical tension from t-1 to t-3 as the current causal edge input), enhancing the dynamic description capability of causal paths. After combining the causal structure graph with temporal embedding information to generate the initial causal graph, the system optimizes node connections and edge weights through Bayesian structure scoring and prediction error feedback. The fit between the reconstructed path and the actual suture action feedback is used to verify the graph structure's alignment with reality. Finally, a dynamic causal graph is constructed, representing the temporal causal network between "action—state evolution—safety boundary" during surgery.

[0191] In each operation cycle, the robot extracts the environmental state from the current sensor and visual information, such as the current position of the suture needle at the tissue interface, the degree of tissue surface deformation, and force feedback anomalies. This state is input into a hierarchical reinforcement learning model. The high-level policy module generates target-oriented instructions based on the suture pattern and the tissue target tension structure, such as "complete the closure of the lower left tension zone" or "prioritize the area near blood vessels." This target is further converted into a sub-target vector, representing the suture task segment and acceptable tension threshold in specific spatial coordinates. The low-level policy module then generates motion parameters based on this sub-target and the current perception state, such as the suture needle incident angle, insertion force, and gripper opening and closing amplitude. The specific motion combinations constructed in this way form a set of deployable mechanical control commands.

[0192] After executing an action, the system collects information on state changes, such as suture point position offset, changes in surgical area tension, and suture completion indicators. It then invokes a pre-constructed dynamic causal graph to predict the expected state change path based on the pre-execution state, selected action, and sub-objectives. The system then compares the actual collected change data with the predicted path to generate a state change difference index. This difference is used to construct intrinsic reward components; for example, "maintaining suture tension within the ideal range" earns a high score, while a risk of tissue tearing results in a negative score. Simultaneously, the system determines extrinsic reward components through a task progress management module, such as "completing a 1cm suture path" or "reducing one incorrect insertion attempt." In the early stages of training, the system assigns higher weight to extrinsic rewards to promote learning progress. Once the system enters a stable phase, it gradually strengthens intrinsic rewards to improve the stability and precision of the actions, ultimately merging them into a complete policy reward signal.

[0193] The reward signal, along with information such as the current environmental state, action selection, and sub-objectives, forms an empirical dataset, which is then fed into the joint training module. First, the dynamic causal graph structure continues to refine paths and update edge weights in the new data, ensuring it always reflects the latest organizational state evolution. Simultaneously, the hierarchical reinforcement learning model adjusts network parameters based on policy rewards, improving the overall policy performance through policy gradient optimization. The system continuously evaluates model performance, such as action execution time, stitching error, and policy convergence speed, to dynamically adjust the learning rate, update direction, and training plan, gradually improving the system's policy quality and convergence efficiency in high-dimensional task spaces.

[0194] Ultimately, after complete training, the robot can directly receive current environmental state information in new suturing tasks. It automatically generates corresponding target-oriented instructions (e.g., "suture the vascular side tissue at the end") through high-level strategies and further generates optimized sub-objectives (e.g., "target position P, tension target of 3.5N"). The low-level strategy calculates refined control actions, outputting the optimal combination of needle insertion point and path curvature to generate a complete suturing action sequence. The system then verifies the path using a causal graph to ensure that the action sequence does not lead to excessive tissue stress or uncontrollable deformation, ensuring that each operation meets physiological limitations and medical safety requirements. This ultimately forms an optimized action strategy that balances accuracy, safety, and efficiency, which is deployed to the robot's control interface to complete the suturing task.

[0195] In the intelligent lending system, to achieve autonomous approval and dynamic risk control of customer credit applications, the system first collects data from multiple sources, including structured documents submitted by customers (such as income statements, credit information, and asset certificates), real-time business interaction data (such as form filling paths, page dwell time, and operation sequence), and unstructured text data provided by external system sensor interfaces (such as customer service interaction text and open banking interface data). All data undergoes a cleaning process to remove outliers and format errors, such as identifying invalid ID numbers, duplicate submissions, and automated script noise in the behavior trajectory. The system then performs time alignment processing to ensure strict synchronization of time annotations such as "income statement changes" and "abnormal interaction behavior," forming a unified time-series scenario data oriented towards the customer's behavioral lifecycle. During standardization, normalization processing is performed on fields from different sources, including credit score vector normalization, form behavior encoding conversion, and text sentiment label discretization, ensuring dimensional comparability between features. The system further extracts structured features representing the evolution of customer status, such as "income trend change rate," "interaction interruption frequency," and "multi-device switching behavior intensity," and merges them to generate a structured scenario dataset for credit approval scenarios.

[0196] Next, the system analyzes customer state variables in the structured scenario dataset, such as "asset fluctuations," "number of risk warnings," "historical approval paths," and "degree of deviation from current behavior," to perform causal discovery operations and identify causal dependencies and evolutionary directions. For example, the system identifies that "frequent application process jumps" may cause "approval delays," or that "multiple changes to loan amounts" may increase the impact on "pre-approval rejection rates." Furthermore, time-delay embedding technology is applied to the time-series data, introducing time-series features such as changes in fund inflows over the past three days and the type of the previous approval behavior, and integrating them into the initial causal graph structure. Subsequently, based on the deviation between the actual approval results and the predicted results in the graph, the system optimizes the connection relationships and strength of the causal structure graph, improving the modeling accuracy and behavioral interpretability of customer behavior trends. The final constructed dynamic causal graph clearly describes the causal link from "multiple rounds of customer interaction behavior" to "changes in risk control scores" and then to "approval results."

[0197] In the actual approval process, the system first receives real-time information about the customer's current environment, including the current filling behavior, device environment, and behavior tags. This state input is processed by a hierarchical reinforcement learning architecture, where higher-level policies generate abstract target instructions, such as "identifying potential fraudulent tendencies" or "judging the reasonableness of the credit limit." These instructions are then transformed into sub-objectives, such as "verifying the authenticity of salary slips" or "inferring whether the target repayment period matches past credit history." Lower-level policies then combine these sub-objectives with the details of the current state to generate specific operational parameters, such as "calling a third-party API to verify the bank card's location," "adjusting the initial credit limit to 150,000," or "postponing the submission of the request and waiting for supplementary materials," thus forming an executable action chain.

[0198] After an action is executed, the system obtains information on post-action status changes through business process records and customer response data, such as indicators like "customer acceptance rate of system decisions," "success rate of supplementary material uploads," and "customer application interruption." Based on a dynamic causal graph, the system predicts the ideal state evolution path by combining the current environmental state, sub-goals, and executed actions. It then compares the actual state changes with this path to generate state offset indicators, such as "risk label increase exceeding prediction by more than 5%" or "number of failed fund verifications lower than expected." The system uses this to construct an intrinsic reward signal reflecting the contribution of the current action to the consistency of the system's modeling predictions. Simultaneously, it combines extrinsic rewards, such as key business indicators like "successful approval rate," "customer retention behavior," and "credit limit hit rate," to generate external rewards. The weights of internal and external rewards are automatically adjusted based on the system's current training round or deployment scenario. For example, in the early stages of model deployment, external rewards are strengthened to drive business goals, while the weight of intrinsic rewards is increased after stable operation to enhance risk control capabilities, thus merging to generate a complete strategy reward signal.

[0199] The system stores current experience data, including environmental state information, actions performed, policy rewards, state changes, and sub-objectives, as a basis for subsequent training. First, it fine-tunes the parameters of the dynamic causal graph, such as updating edge weight strength and node transition probabilities, to improve the consistency of predictions regarding customer behavior evolution. Then, it optimizes the parameters of the hierarchical reinforcement learning model using policy gradient or value iteration methods, continuously increasing the cumulative reward in the high-dimensional state and action space. Throughout this process, the system monitors model performance changes in real time, such as indicators like "approval efficiency improvement" and "rejection rate trend," automatically adjusting the learning rate, objective function, or optimization path to ensure that model training always aligns with the system's goal evolution. Finally, it performs a joint adaptive update of the causal graph and the policy network.

[0200] After the update, the system redeploys the intelligent approval decision-making process using the latest hierarchical reinforcement learning model. When a new customer enters the approval process, the system immediately obtains their environmental status information, generates goal-oriented instructions (such as "avoiding potential high delinquency risks") through high-level strategies, and then decomposes them into sub-goals (such as "score abnormalities in large-amount transfers in the past three months"). These sub-goals are then linked with data analysis, model invocation, and approval path generation through low-level strategies, outputting a set of executable action parameters. These parameters are transformed into action sequences, such as "suspend approval and wait for supplementary bank statements," "recommend the credit limit to 30,000 yuan," or "suggest entering the manual review channel." The system then verifies whether the action sequence maintains structural consistency with the dynamic causal graph, ensuring that the strategy does not cause system-level risk shifts or logical loops. After confirmation, the final approval strategy is output and automatically applied to the approval process or submitted to the manual intervention node.

[0201] This embodiment introduces an updated hierarchical reinforcement learning structure and combines it with dynamic causal graphs for action verification. The system achieves a closed-loop process from abstract goal setting, sub-goal construction, specific action generation to causal consistency verification. High-level policies ensure goal-oriented and globally consistent behavior generation, while low-level policies improve the execution granularity and response speed of action generation. The causal graph verification stage strengthens the structural matching between the policy and the environment. This collaborative process results in a final action policy that significantly outperforms traditional reinforcement learning methods in terms of accuracy, stability, interpretability, and task adaptability, making it particularly suitable for complex environments requiring dynamic control and high-precision execution.

[0202] In one embodiment, a policy generation apparatus based on hierarchical reinforcement learning is provided, which corresponds one-to-one with the policy generation method based on hierarchical reinforcement learning in the above embodiments. (Refer to...) Figure 3 , Figure 3This is a schematic diagram of the functional modules of a preferred embodiment of the policy generation device based on hierarchical reinforcement learning of the present invention. The module includes a data preprocessing module 10, a causal modeling module 20, a hierarchical policy generation module 30, a state evolution analysis module 40, a joint optimization training module 50, and an action policy output module 60. Detailed descriptions of each functional module are as follows:

[0203] The data preprocessing module 10 is used to acquire multi-source raw scene data and preprocess the multi-source raw scene data to generate a structured scene dataset.

[0204] The causal modeling module 20 is used to construct a dynamic causal graph based on the structured scene dataset to describe the causal relationships of state evolution in the scene.

[0205] The hierarchical policy generation module 30 is used to receive environmental state information, process the environmental state information through a hierarchical reinforcement learning architecture that includes high-level policies and low-level policies, and generate sub-goals and specific actions.

[0206] The state evolution analysis module 40 is used to obtain state change information after the specific action is executed, and generate a policy reward signal based on the environmental state information, state change information, specific action and sub-goal using the dynamic cause-effect graph;

[0207] The joint optimization training module 50 is used to update the parameters of the dynamic causal graph and the hierarchical reinforcement learning model through joint training based on the policy reward signal.

[0208] Action policy output module 60 is used to generate optimized action policies through the updated hierarchical reinforcement learning model.

[0209] In one embodiment, the data preprocessing module 10 is specifically used for:

[0210] Collect raw scene data from multiple sources, including visual data, sensor data, and structured tabular data;

[0211] The multi-source raw scene data is cleaned to remove noise and outliers, generating cleaned multi-source raw scene data.

[0212] Align the time series dimension in the cleaned multi-source original scene data to generate time-aligned multi-source original scene data;

[0213] The time-aligned multi-source raw scene data of different modalities are standardized to generate standardized multi-source raw scene data;

[0214] Features are extracted from the standardized multi-source original scene data to generate a feature set;

[0215] The feature sets are fused to generate a fused feature representation;

[0216] A structured scene dataset is generated based on the fused feature representation.

[0217] In one embodiment, the causal modeling module 20 is specifically used for:

[0218] Analyze the structured scene dataset to identify the set of state variables;

[0219] Causal discovery is performed based on the set of state variables to generate a causal structure model;

[0220] Time-delay embedding processing is performed on the time-series data in the structured scene dataset to generate time-delay embedded data;

[0221] Integrate the causal structure model and time-delay embedded data to construct an initial causal graph;

[0222] The initial causal graph is optimized to improve the accuracy of state evolution prediction, and an optimized causal graph is generated.

[0223] Verify the consistency between the optimized causal graph and the structured scene dataset, and generate a dynamic causal graph.

[0224] In one embodiment, the hierarchical strategy generation module 30 is specifically used for:

[0225] Receive environmental status information;

[0226] The environment state information is processed by a high-level policy of a hierarchical reinforcement learning architecture to generate high-level decision instructions;

[0227] Sub-objectives are generated based on the aforementioned high-level decision-making instructions;

[0228] The underlying strategy of the hierarchical reinforcement learning architecture processes the environmental state information and sub-objectives to generate action execution parameters;

[0229] A specific action is generated based on the action execution parameters.

[0230] In one embodiment, the state evolution analysis module 40 is specifically used for:

[0231] Obtain the state change information after the specific action is executed;

[0232] Using the dynamic cause-effect graph, the expected state evolution path is determined based on the environmental state information, specific actions, and sub-goals;

[0233] The degree of conformity between the state change information and the expected state evolution path is detected, and a state change difference index is generated.

[0234] An intrinsic reward component is generated based on the aforementioned state change difference index;

[0235] External reward components are generated based on the task completion progress;

[0236] Determine the weight ratio of the intrinsic reward component and the extrinsic reward component, and dynamically adjust the weight ratio according to the current training stage;

[0237] The strategy reward signal is generated by weighting and fusing the intrinsic reward component and the extrinsic reward component based on the adjusted weight ratio.

[0238] In one embodiment, the joint optimization training module 50 is specifically used for:

[0239] The storage contains empirical data including environmental state information, specific actions, policy reward signals, state change information, and sub-objectives.

[0240] The parameters of the dynamic causal graph are optimized based on the empirical data to improve prediction consistency;

[0241] The parameters of the hierarchical reinforcement learning model are optimized based on the empirical data to maximize the cumulative reward;

[0242] The performance changes of the optimized hierarchical reinforcement learning model are detected, and the optimization direction of the model is adjusted based on the performance changes;

[0243] The parameters of the dynamic causal graph and the hierarchical reinforcement learning model are iteratively updated based on the adjusted optimization direction.

[0244] In one embodiment, the action strategy output module 60 is specifically used for:

[0245] Obtain current environment status information;

[0246] The updated hierarchical reinforcement learning model processes the current environment state information through a high-level strategy to generate goal-oriented instructions.

[0247] Optimized sub-objectives are generated based on the target-oriented instructions;

[0248] The updated hierarchical reinforcement learning model processes the current environment state information and optimizes the sub-objectives through its underlying policy, generating optimized action parameters.

[0249] Generate an executable action sequence based on the optimized action parameters;

[0250] Verify the compatibility between the executable action sequence and the dynamic cause-effect graph;

[0251] The validated executable action sequence is output as an optimized action strategy.

[0252] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the server-side functions or steps of a hierarchical reinforcement learning-based policy generation method.

[0253] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the user-side functions or steps of a hierarchical reinforcement learning-based policy generation method.

[0254] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0255] Acquire multi-source raw scene data and preprocess the multi-source raw scene data to generate a structured scene dataset;

[0256] Based on the structured scene dataset, a dynamic causal graph is constructed to describe the causal relationships of state evolution in the scene;

[0257] Receive environmental state information, process the environmental state information through a hierarchical reinforcement learning architecture that includes high-level policies and low-level policies, and generate sub-goals and specific actions;

[0258] Obtain the state change information after the specific action is executed, and use the dynamic cause-effect graph to generate a policy reward signal based on the environmental state information, state change information, specific action and sub-goal;

[0259] Based on the policy reward signal, the parameters of the dynamic causal graph and the hierarchical reinforcement learning model are updated through joint training;

[0260] Optimized action strategies are generated using the updated hierarchical reinforcement learning model.

[0261] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0262] Acquire multi-source raw scene data and preprocess the multi-source raw scene data to generate a structured scene dataset;

[0263] Based on the structured scene dataset, a dynamic causal graph is constructed to describe the causal relationships of state evolution in the scene;

[0264] Receive environmental state information, process the environmental state information through a hierarchical reinforcement learning architecture that includes high-level policies and low-level policies, and generate sub-goals and specific actions;

[0265] Obtain the state change information after the specific action is executed, and use the dynamic cause-effect graph to generate a policy reward signal based on the environmental state information, state change information, specific action and sub-goal;

[0266] Based on the policy reward signal, the parameters of the dynamic causal graph and the hierarchical reinforcement learning model are updated through joint training;

[0267] Optimized action strategies are generated using the updated hierarchical reinforcement learning model.

[0268] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0269] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0270] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0271] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A policy generation method based on hierarchical reinforcement learning, characterized in that, Includes the following steps: Acquire multi-source raw scene data and preprocess the multi-source raw scene data to generate a structured scene dataset. The multi-source raw scene data includes visual data, sensor data, and structured tabular data. Based on the structured scene dataset, a dynamic causal graph is constructed to describe the causal relationships of state evolution in the scene; Receive environmental state information, process the environmental state information through a hierarchical reinforcement learning architecture that includes high-level policies and low-level policies, and generate sub-goals and specific actions; The process involves: acquiring state change information after the execution of a specific action; using the dynamic causal graph to generate a policy reward signal based on the environmental state information, state change information, specific action, and sub-objective; obtaining state change information after the execution of the specific action; determining an expected state evolution path based on the environmental state information, specific action, and sub-objective using the dynamic causal graph; detecting the conformity between the state change information and the expected state evolution path to generate a state change difference index; generating an intrinsic reward component based on the state change difference index; generating an extrinsic reward component based on the task completion progress; determining the weight ratio between the intrinsic and extrinsic reward components and dynamically adjusting the weight ratio according to the current training stage; and weightedly fusing the intrinsic and extrinsic reward components based on the adjusted weight ratio to generate a policy reward signal. Based on the policy reward signal, the parameters of the dynamic causal graph and the hierarchical reinforcement learning model are updated through joint training, including: storing empirical data containing environmental state information, specific actions, policy reward signals, state change information, and sub-objectives; optimizing the parameters of the dynamic causal graph based on the empirical data to improve prediction consistency; optimizing the parameters of the hierarchical reinforcement learning model based on the empirical data to maximize cumulative reward; detecting performance changes in the optimized hierarchical reinforcement learning model and adjusting the model optimization direction based on the performance changes; and iteratively updating the parameters of the dynamic causal graph and the hierarchical reinforcement learning model based on the adjusted optimization direction. Optimized action strategies are generated using the updated hierarchical reinforcement learning model.

2. The policy generation method based on hierarchical reinforcement learning as described in claim 1, characterized in that, Acquire multi-source raw scene data and preprocess the multi-source raw scene data to generate a structured scene dataset, including: Collect raw scene data from multiple sources, including visual data, sensor data, and structured tabular data; The multi-source raw scene data is cleaned to remove noise and outliers, generating cleaned multi-source raw scene data. Align the time series dimension in the cleaned multi-source original scene data to generate time-aligned multi-source original scene data; The time-aligned multi-source raw scene data of different modalities are standardized to generate standardized multi-source raw scene data; Features are extracted from the standardized multi-source original scene data to generate a feature set; The feature sets are fused to generate a fused feature representation; A structured scene dataset is generated based on the fused feature representation.

3. The policy generation method based on hierarchical reinforcement learning as described in claim 1, characterized in that, Based on the structured scene dataset, a dynamic causal graph is constructed to describe the causal relationships of state evolution in the scene, including: Analyze the structured scene dataset to identify the set of state variables; Causal discovery is performed based on the set of state variables to generate a causal structure model; Time-delay embedding processing is performed on the time-series data in the structured scene dataset to generate time-delay embedded data; Integrate the causal structure model and time-delay embedded data to construct an initial causal graph; The initial causal graph is optimized to improve the accuracy of state evolution prediction, and an optimized causal graph is generated. Verify the consistency between the optimized causal graph and the structured scene dataset, and generate a dynamic causal graph.

4. The policy generation method based on hierarchical reinforcement learning as described in claim 1, characterized in that, Receive environmental state information, process the environmental state information through a hierarchical reinforcement learning architecture that includes high-level and low-level policies, and generate sub-goals and specific actions, including: Receive environmental status information; The environment state information is processed by a high-level policy of a hierarchical reinforcement learning architecture to generate high-level decision instructions; Sub-objectives are generated based on the aforementioned high-level decision-making instructions; The underlying strategy of the hierarchical reinforcement learning architecture processes the environmental state information and sub-objectives to generate action execution parameters; A specific action is generated based on the action execution parameters.

5. The policy generation method based on hierarchical reinforcement learning as described in claim 1, characterized in that, The updated hierarchical reinforcement learning model generates optimized action policies, including: Obtain current environment status information; The updated hierarchical reinforcement learning model processes the current environment state information through a high-level strategy to generate goal-oriented instructions. Optimized sub-objectives are generated based on the target-oriented instructions; The updated hierarchical reinforcement learning model processes the current environment state information and optimizes the sub-objectives through its underlying policy, generating optimized action parameters. Generate an executable action sequence based on the optimized action parameters; Verify the compatibility between the executable action sequence and the dynamic cause-effect graph; The validated executable action sequence is output as an optimized action strategy.

6. A policy generation device based on hierarchical reinforcement learning, characterized in that, The policy generation device based on hierarchical reinforcement learning includes: The data preprocessing module is used to acquire multi-source raw scene data and preprocess the multi-source raw scene data to generate a structured scene dataset. The multi-source raw scene data includes visual data, sensor data, and structured table data. The causal modeling module is used to construct a dynamic causal graph that describes the causal relationships of state evolution in the scenario based on the structured scenario dataset. The hierarchical policy generation module is used to receive environmental state information, process the environmental state information through a hierarchical reinforcement learning architecture that includes high-level policies and low-level policies, and generate sub-goals and specific actions. The state evolution analysis module is used to acquire state change information after the execution of the specific action, and generate a policy reward signal based on the environmental state information, state change information, specific action, and sub-objective using the dynamic causal graph. This includes: acquiring state change information after the execution of the specific action; determining the expected state evolution path based on the environmental state information, specific action, and sub-objective using the dynamic causal graph; detecting the conformity between the state change information and the expected state evolution path, and generating a state change difference index; generating an intrinsic reward component based on the state change difference index; generating an extrinsic reward component based on the task completion progress; determining the weight ratio of the intrinsic reward component and the extrinsic reward component, and dynamically adjusting the weight ratio according to the current training stage; and weightedly fusing the intrinsic reward component and the extrinsic reward component based on the adjusted weight ratio to generate a policy reward signal. The joint optimization training module is used to update the parameters of the dynamic causal graph and the hierarchical reinforcement learning model through joint training based on the policy reward signal. This includes: storing empirical data containing environmental state information, specific actions, policy reward signals, state change information, and sub-objectives; optimizing the parameters of the dynamic causal graph based on the empirical data to improve prediction consistency; optimizing the parameters of the hierarchical reinforcement learning model based on the empirical data to maximize cumulative reward; detecting performance changes in the optimized hierarchical reinforcement learning model and adjusting the model optimization direction based on the performance changes; and iteratively updating the parameters of the dynamic causal graph and the hierarchical reinforcement learning model based on the adjusted optimization direction. The action policy output module is used to generate optimized action policies through the updated hierarchical reinforcement learning model.

7. A computer device, characterized in that, The computer device includes a memory, a processor, and a policy generation program based on hierarchical reinforcement learning stored in the memory and executable on the processor. When executed by the processor, the policy generation program based on hierarchical reinforcement learning implements the steps of the policy generation method based on hierarchical reinforcement learning as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The storage medium stores a policy generation program based on hierarchical reinforcement learning, which, when executed by a processor, implements the steps of the policy generation method based on hierarchical reinforcement learning as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Private domain live broadcast peak hot spot prediction and content scheduling method based on deep learning

    CN119450099A