Intelligent decision support method based on reinforcement learning

By performing spatiotemporal feature encoding and policy gradient algorithm updates in reinforcement learning, a decision action space mapping table is constructed, which solves the problems of decision delay and error in existing technologies and achieves efficient decision support in complex dynamic scenarios.

CN121009928APending Publication Date: 2025-11-25YANGO UNIV

Patent Information

Application Number
CN202511543046.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing reinforcement learning-based decision-making methods struggle to fully utilize historical interaction data in complex and dynamic scenarios, neglect the correlation between environmental states and decision actions, and fail to update policies in real time, leading to decision delays and errors. They are also unable to quickly process massive amounts of real-time environmental data.

Method used

By acquiring historical interaction data under the target decision-making scenario, spatiotemporal feature encoding is performed to generate state feature vectors, a decision action space mapping table is constructed, and a policy gradient algorithm is used for dynamic updating to build a real-time decision support engine and output the optimal decision action.

Benefits of technology

It improves the accuracy and efficiency of decision-making, enables rapid response to environmental changes, and meets the timeliness requirements of financial transactions, equipment scheduling, and resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009928A_ABST
    Figure CN121009928A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent decision, and discloses an intelligent decision support method based on reinforcement learning. The method comprises the following steps: acquiring a historical interaction data set containing a decision action sequence, an environment state sequence and a corresponding instant reward signal in a target decision scene; the data set is input into a state feature extraction network for spatial-temporal feature coding, and a state feature vector set with time sequence relevance is generated; constructing a decision action space mapping table based on the set, wherein candidate decision actions and expected accumulated rewards corresponding to the state feature vectors are recorded in the table; dynamically updating the mapping table by adopting a strategy gradient algorithm to generate an optimized strategy gradient parameter set; and constructing a real-time decision support engine according to the parameter set, wherein the engine can respond to the environment state change and output the optimal decision action. The method adapts to a dynamic decision-making scene, and assists in efficiently outputting decision-making results meeting requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent decision-making technology, specifically to an intelligent decision support method based on reinforcement learning. Background Technology

[0002] Intelligent decision-making plays a crucial role in numerous fields, including financial risk control, intelligent manufacturing equipment scheduling, and medical resource allocation. These scenarios share the common characteristic of dynamically changing environmental conditions over time, and the decision-making outcomes directly impact operational efficiency and profitability. Traditional decision-making systems often rely on manually preset rules or static models. When unexpected changes occur in the environment—such as sudden policy adjustments in the financial market, temporary abnormal parameters in manufacturing equipment, or significant short-term fluctuations in the supply and demand of medical resources—traditional systems often struggle to adapt quickly, leading to decreased decision-making accuracy and even errors.

[0003] With the development of artificial intelligence technology, reinforcement learning, due to its ability to learn through trial and error in interaction between intelligent agents and the environment, has gradually become an important technological direction in the field of intelligent decision-making. However, existing reinforcement learning-based decision-making methods still have many shortcomings: They do not fully utilize historical interaction data; most methods only extract static features from the data, ignoring the correlation between environmental states and decision actions across different time dimensions. This results in the generated state representations failing to fully reflect the dynamic changes in the scene, thus affecting the rationality of subsequent decisions. The correspondence between decision actions and states lacks a clear structure and mapping mechanism. The decision-making process requires temporary calculation of action rewards, increasing decision delays and potentially leading to biased action selection due to errors in the calculation process. Policy update mechanisms are mostly periodic, making it difficult to respond to sudden environmental changes in real time. When the environmental state of the target decision scenario changes significantly within the update cycle, the updated policy will deviate significantly from the current scenario requirements, failing to effectively guide decision-making. Some methods lack dedicated real-time decision output modules, making it difficult to quickly process information and generate optimal decision actions when faced with massive amounts of real-time environmental data, thus failing to meet the timeliness requirements of decision-making in real-world scenarios. These issues collectively limit the effectiveness of existing reinforcement learning decision-making methods in complex and dynamic scenarios, making it difficult to fully realize the value of intelligent decision-making. Summary of the Invention

[0004] The purpose of this invention is to provide an intelligent decision support method based on reinforcement learning to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides an intelligent decision support method based on reinforcement learning, the method comprising:

[0006] Acquire a set of historical interaction data under the target decision-making scenario, wherein the set of historical interaction data includes a sequence of decision-making actions, a sequence of environmental states, and corresponding instantaneous reward signals;

[0007] The historical interaction data set is input into a state feature extraction network for spatiotemporal feature encoding to generate a set of state feature vectors with temporal correlation.

[0008] A decision action space mapping table is constructed based on the set of state feature vectors. The decision action space mapping table records the candidate decision action and its expected cumulative reward corresponding to each state feature vector.

[0009] The decision action space mapping table is dynamically updated using a policy gradient algorithm to generate an optimized set of policy gradient parameters.

[0010] A real-time decision support engine is constructed based on the optimized set of policy gradient parameters. The real-time decision support engine is used to respond to changes in environmental state and output the optimal decision action.

[0011] Preferably, the step of inputting the historical interaction data set into a state feature extraction network for spatiotemporal feature encoding processing to generate a set of state feature vectors with temporal correlation includes:

[0012] The environmental state sequence is extracted from the historical interaction data set, and the environmental state sequence is segmented by a sliding window to generate multiple local state sub-sequences.

[0013] A pre-trained spatiotemporal convolutional network is invoked to perform feature extraction processing on each of the local state subsequences, generating a local spatiotemporal feature vector;

[0014] The local spatiotemporal feature vector is input into a long short-term memory network to model temporal dependencies and generate a global state feature vector.

[0015] The global state feature vector is subjected to dimensionality compression to generate a standardized set of state feature vectors.

[0016] Preferably, the step of constructing a decision action space mapping table based on the set of state feature vectors includes:

[0017] Based on each state feature vector in the state feature vector set, query the decision actions executed under the same or similar states in the historical interaction data set;

[0018] Calculate the expected cumulative reward for each decision action under the state feature vector, which is obtained by iterative calculation using the Bellman equation;

[0019] The state feature vector, candidate decision actions, and expected cumulative reward are stored in key-value pairs to form an initial decision action space mapping table.

[0020] The initial decision action space mapping table is extended by action exploration, and an unverified but constraint-compliant exploratory decision action is added to each state feature vector.

[0021] Preferably, the step of dynamically updating the decision action space mapping table using a policy gradient algorithm includes:

[0022] Sample the current state feature vector from the real-time decision-making environment and query the decision action space mapping table to obtain the corresponding set of candidate decision actions;

[0023] Based on the action selection probability distribution of the policy gradient algorithm, a target decision action is selected from the set of candidate decision actions and executed;

[0024] Record the transition results of the environmental state after the execution of the target decision action and the obtained immediate reward signal;

[0025] Based on the transition results and the immediate reward signal, the policy gradient update amount is calculated, and the expected cumulative reward in the decision action space mapping table is corrected by backpropagation.

[0026] Preferably, the step of constructing a real-time decision support engine based on the optimized policy gradient parameter set includes:

[0027] The optimized policy gradient parameter set is loaded into the online inference framework to form a policy network that can respond in real time;

[0028] Configure a state feature preprocessing pipeline for the policy network, which is used to convert the original environment state into a state feature vector;

[0029] A post-processing mechanism for action output is established to perform feasibility verification and boundary constraint correction on the decision actions output by the policy network.

[0030] Preferably, the state feature preprocessing pipeline includes:

[0031] Receive raw environmental state data and perform missing value imputation and noise filtering on the raw environmental state data;

[0032] The same spatiotemporal convolutional network and long short-term memory network as those used in the training phase are invoked to extract features from the cleaned environmental state data.

[0033] The extracted feature vectors are matched with the state feature vectors in the decision action space mapping table to ensure consistency of the feature space.

[0034] Preferably, the post-processing mechanism for the action output includes:

[0035] Obtain the original decision action output by the policy network, and query the domain knowledge base to verify the physical feasibility of the original decision action;

[0036] For original decision actions that violate physical constraints, nearest neighbor replacement is performed, replacing them with the candidate decision actions with the highest feasibility under the same state in the decision action space mapping table;

[0037] The corrected decision action is marked as the final output action, and the training sample library of the policy network is updated synchronously.

[0038] Preferably, the method for constructing the domain knowledge base includes:

[0039] Collect physical constraint rules and operational specifications documents for the target decision-making scenario, and convert them into a set of structured constraint conditions;

[0040] Establish a mapping table between constraints and decision actions, and record the range of boundary parameters that each decision action is allowed to execute;

[0041] The system periodically receives abnormal action records from the real-time decision support engine and dynamically expands the constraint entries in the mapping table.

[0042] Preferably, the training sample library of the synchronous update policy network includes:

[0043] The revised decision-making actions, environmental state transition results, and immediate reward signals are combined into new training samples.

[0044] Calculate the feature space distance between the new training sample and the samples in the existing sample library; if it exceeds a preset threshold, add it to the sample library.

[0045] When the sample library capacity reaches its limit, an importance sampling strategy is used to eliminate the historical samples with the highest redundancy.

[0046] Preferably, the importance sampling strategy includes:

[0047] Each training sample is assigned a dynamic weight, which depends on the exploratory value and historical usage frequency of the decision action in the sample;

[0048] Periodically calculate the weight distribution of all samples in the sample library and remove low-weight samples after weight sorting;

[0049] The sample gaps triggered by the removal operation will be used to store newly acquired high-value training samples.

[0050] Compared with the prior art, the beneficial effects of the present invention are:

[0051] In the data acquisition stage, the method focuses on the target decision-making scenario and collects a set of historical interaction data that includes decision action sequences, environmental state sequences, and immediate reward signals. Instead of collecting a single type of data in a scattered manner, it covers the complete "environment-action-feedback" link in the decision-making process, so that subsequent analysis has comprehensive data support, avoids decision bias caused by missing data dimensions, and makes the foundation of decision analysis more solid.

[0052] When historical interaction data is input into a state feature extraction network for spatiotemporal feature encoding, the network model can deeply mine the hidden temporal and spatial correlations in the data. For example, the evolution of environmental states at different time points and the sequential influence of different decision actions in the same state. The generated set of state feature vectors is no longer isolated static data, but a feature carrier that can reflect the dynamic change trend of the scene. This kind of feature vector with temporal correlation can more accurately depict the essence of the environmental state, providing a more realistic basis for subsequent action mapping and making the matching of actions and states more reasonable.

[0053] The decision action space mapping table, constructed based on the set of state feature vectors, directly associates each state feature vector with the corresponding candidate decision action and expected cumulative reward, forming a structured correspondence. During the decision-making process, there is no need to repeatedly calculate the action reward. The mapping table can be used to quickly locate potential action options and their respective expected rewards in different states. This not only reduces the amount of computation in the decision-making process, but also avoids the blindness of selection caused by temporary calculations, making the decision selection more directional and improving decision-making efficiency.

[0054] When the policy gradient algorithm is used to dynamically update the mapping table, the algorithm can adjust the parameters in real time according to the changes in the environmental state, rather than relying on fixed periods or manual triggering of updates. This dynamic update mechanism can ensure that the decision-making strategy is continuously optimized as the scenario changes. Even if the target decision-making scenario experiences sudden state or long-term trend changes, the updated policy gradient parameter set can still maintain its adaptability to the scenario requirements, so that the decision-making strategy always fits the actual situation of the current scenario and avoids decision failure due to policy lag.

[0055] The real-time decision support engine, built on an optimized set of parameters, is specifically designed with a response mechanism for changes in environmental state. It can quickly receive new environmental state information and directly output the optimal decision action based on the optimized strategy parameters, without going through a complex intermediate calculation process. In scenarios with high requirements for decision timeliness, such as financial transactions, equipment scheduling, and resource allocation, it can effectively shorten the interval from state perception to action output, ensuring that decisions can be implemented in a timely manner and applied to the actual scenario, and allowing decision results to be quickly transformed into practical value. Attached Figure Description

[0056] Figure 1 This is a schematic diagram illustrating the working principle of the reinforcement learning-based intelligent decision support method described in this invention.

[0057] Figure 2 A flowchart for constructing a decision action space mapping table;

[0058] Figure 3 Flowchart for the state feature preprocessing pipeline;

[0059] Figure 4 A diagram for optimizing operational status and decision-making. Detailed Implementation

[0060] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0061] Please see Figure 1 This invention provides an intelligent decision support method based on reinforcement learning. The method includes: acquiring a historical interaction data set under a target decision scenario, which includes a sequence of decision actions, a sequence of environmental states, and corresponding immediate reward signals; the historical interaction data set is derived from system logs or simulated environments to ensure data coverage of multiple decision scenarios. The historical interaction data set is input into a state feature extraction network for spatiotemporal feature encoding. This network uses a deep learning architecture to learn features from the state sequences and outputs a set of state feature vectors with temporal correlation; the design of the state feature extraction network focuses on capturing the dynamic evolution of environmental states. Based on the set of state feature vectors, a decision action space mapping table is constructed. This mapping table stores candidate decision actions and their expected cumulative rewards for each state feature vector in key-value pairs; the expected cumulative reward is calculated using a value function approximation method in reinforcement learning. A policy gradient algorithm is used to dynamically update the decision action space mapping table. The policy gradient algorithm calculates the policy gradient by sampling environmental interaction data, thereby adjusting the parameters in the mapping table and generating an optimized set of policy gradient parameters; the update process involves iterative loops of policy evaluation and improvement. A real-time decision support engine is built based on the optimized set of policy gradient parameters. This engine integrates a policy network and a state processing module, which can respond to changes in environmental state in real time and output the optimal decision action based on the current policy. The engine is deployed with low latency and high reliability requirements in mind.

[0062] Example 1: The specific process of spatiotemporal feature encoding using the state feature extraction network begins with extracting environmental state sequences from a historical interaction dataset. These sequences are arranged in strict chronological order, with each time step containing multidimensional environmental variables, such as sensor readings or system metrics. These variables collectively characterize the dynamic changes in the decision-making scenario. The extraction process involves data parsing and timestamp alignment to ensure the continuity and consistency of the sequence, avoiding time breaks caused by missing data or jitter. A sliding window segmentation process is then applied to the environmental state sequences. The size of the sliding window is flexibly determined based on the time scale of the decision-making scenario. For example, a shorter window is used in real-time control scenarios to capture rapid changes, while a longer window is used in long-term planning scenarios to cover macroscopic trends. Multiple local state subsequences are generated, covering continuous time segments. Overlapping or non-overlapping strategies are used to balance computational efficiency and feature integrity. The length of the local state subsequences needs to be optimized to capture sufficient context without introducing excessive noise. The sliding window segmentation uses a configurable step size parameter, which is dynamically adjusted according to the data sampling frequency. For example, a smaller step size is used for high-frequency data to increase sequence density, while a larger step size is used for low-frequency data to reduce redundancy. A boundary processing mechanism is implemented during the generation of local state subsequences to fill or truncate the start and end parts of the sequence, ensuring that all subsequences have a uniform length for easy subsequent processing.

[0063] A pre-trained spatiotemporal convolutional network is invoked to extract features from each local state subsequence. This network comprises multiple layers of convolution and pooling operations. The convolutional layers use three-dimensional kernels to simultaneously capture local patterns in both spatial and temporal dimensions, while the pooling layers reduce feature dimensionality and enhance translation invariance. The convolutional kernel parameters are pre-trained using historical data. Unsupervised learning, such as autoencoders or contrastive learning, is employed during pre-training to enable the network to learn a general spatiotemporal representation without relying on specific task labels. Local spatiotemporal feature vectors are generated, encoding key patterns within the subsequence, such as periodicity and anomalous fluctuations. Multi-scale convolutional kernels are applied during feature extraction; small-scale kernels capture detailed features, while large-scale kernels integrate macroscopic information. The dimensionality of the local spatiotemporal feature vectors is consistent with the number of output channels of the convolutional network, and non-linearity is typically introduced through activation functions such as ReLU. Fine-tuning of the pre-trained network can be performed according to specific scenarios, using labeled data and supervised learning to refine feature representations and improve adaptability to the target decision-making task. Local spatiotemporal feature vectors are input into a Long Short-Term Memory (LSTM) network for temporal dependency modeling. The LTM network has a forget gate, input gate, and output gate structure. The forget gate controls the degree of retention of historical information, the input gate regulates the integration of new information, and the output gate determines the exposure of hidden states. The network selectively retains or discards information through a gating mechanism to handle long-term dependencies in sequence data and avoid gradient vanishing or exploding problems. During temporal dependency modeling, the number of hidden layers is configured according to the sequence complexity; shallow networks are used for simple sequences, while deep networks are used for complex sequences to enhance expressive power. However, overfitting must be prevented through regularization such as dropout. The network outputs a global state feature vector, which integrates the contextual information of the entire sequence, such as trend direction and abrupt change points. The global vector is generated from the final output of the hidden states or through sequence pooling. The LTM network is trained using a backpropagation algorithm over time, and gradient pruning techniques stabilize the learning process. The initial state can be set to zero or learned from the data.

[0064] The global state feature vectors undergo dimensionality compression. Principal component analysis (PCA) or autoencoders are used to reduce feature dimensions. PCA finds the direction of maximum variance through linear transformation, while autoencoders learn a compact representation through an encoder-decoder structure. This process eliminates redundant information while retaining most of the original variance. The dimensionality compression rate is balanced according to application requirements; a high compression rate saves storage but may result in information loss. A standardized set of state feature vectors is generated. The standardization process includes normalization and scaling. Normalization maps feature values ​​to the range of zero to one or negative one to one, while scaling adjusts the feature scale to ensure a balanced contribution from all dimensions, ensuring feature values ​​are of the same magnitude and preventing any single dimension from dominating model training. Common standardization methods include min-max scaling or Z-score standardization. Z-score standardization makes the feature distribution mean zero and variance one, enhancing numerical stability. The output format of the state feature vector set is designed as a tensor structure for easy batch processing and efficient computation, such as using TensorFlow tensors. Specifically, the output format of the state feature vector set is designed using TensorFlow tensor structures to achieve batch processing and efficient computation. A tensor is a multi-dimensional array structure that can uniformly represent batch data. In practice, the state feature vectors are organized into a three-dimensional tensor with a shape defined as [batch_size, sequence_length, feature_dimension], where batch_size represents the number of samples processed in a single batch (e.g., 32 or 64), sequence_length represents the time series step size (determined by the sliding window size, e.g., 10 time steps), and feature_dimension represents the dimension of each state feature vector (e.g., 128 dimensions). In implementation, TensorFlow's tf.Tensor objects are used to store data, and an input pipeline is built using the tf.data.Dataset API to batch load data from historical interactive datasets. This pipeline supports parallel reading and preprocessing operations, such as applying the feature extraction network using the tf.map_fn function. Efficient computation is achieved through TensorFlow's graph execution model: a computation graph is defined, encapsulating spatiotemporal convolutional networks and long short-term memory networks as subclasses of `tf.keras.Model`; the `tf.function` decorator is used to compile forward propagation into a static graph, leveraging the parallel computing capabilities of GPUs (such as CUDA cores) to accelerate convolution and LSTM operations. For example, convolutional layers use `tf.keras.layers.Conv1D` for temporal convolutions, and LSTM layers use `tf.keras.layers.LSTM` to handle sequence dependencies. Tensor structures allow matrix operations to be performed on the entire batch of data at once, reducing Python loop overhead, and gradient computation is supported through TensorFlow's automatic differentiation mechanism.Furthermore, the tensor format ensures compatibility with the input of subsequent decision action space mapping tables, avoiding data transformation delays. The entire design emphasizes memory layout optimization, such as using row-first storage to improve cache hit rate, thereby achieving millisecond-level feature extraction in scenarios such as financial risk control. The state feature extraction network is trained using the backpropagation algorithm, and the loss function is optimized based on reconstruction error or prediction accuracy. Reconstruction error is used for unsupervised learning to measure the difference between the input and the reconstructed output, while prediction accuracy is used for supervised learning to evaluate the utility of features for downstream tasks. The training data is divided into training and validation sets, an early stopping strategy prevents overfitting, and optimizers such as Adam or SGD dynamically adjust the learning rate.

[0065] The entire processing flow achieves efficient transformation from raw state sequences to compact feature vectors, integrating parallel computing techniques such as GPU-accelerated convolution to improve processing speed. Specifically, an environmental state sequence is extracted from a historical interaction dataset. This sequence is arranged chronologically and contains multidimensional environmental variables. Next, a sliding window segmentation process is applied to the environmental state sequence to generate multiple local state subsequences. This decomposition transforms the long sequence into manageable fragments, facilitating subsequent feature learning. A pre-trained spatiotemporal convolutional network is invoked to extract features from each local state subsequence. The spatiotemporal convolutional network employs multi-layer convolution and pooling operations, using a three-dimensional kernel to capture local patterns in both spatial and temporal dimensions, generating local spatiotemporal feature vectors. These local spatiotemporal feature vectors are then input into a Long Short-Term Memory (LSTM) network for temporal dependency modeling. The LTM network uses a gating mechanism to process long-term dependencies in the sequence data, generating a global state feature vector that integrates the contextual information of the entire sequence. Finally, the global state feature vector undergoes dimensionality compression. Principal component analysis or autoencoder techniques are used to reduce feature dimensions, and standardization operations such as normalization or scaling are implemented to generate a standardized set of state feature vectors. The entire process reduces computational complexity through sliding window segmentation and, combined with the efficient feature learning of deep learning models, achieves a smooth transformation from the original state sequence to a compact feature vector. The parallel computing techniques integrated into the process significantly improve processing speed. In implementing the spatiotemporal convolutional network and the long short-term memory network, the system leverages the parallel computing capabilities of GPUs for acceleration, such as optimizing convolution operations and LSTM computations through CUDA cores. The three-dimensional kernel operations of the convolutional network simultaneously handle both spatial and temporal dimensions, and thanks to the parallel architecture of the GPU, it can process multiple local state subsequences in batches. The gating mechanism computation of the long short-term memory network is also optimized through parallelization, utilizing the graph execution mode of frameworks such as TensorFlow to compile forward propagation into a static graph, reducing Python interpretation overhead. This parallel acceleration enables the feature extraction process to efficiently handle large-scale data, meeting the low-latency requirements of real-time decision-making. The training data is partitioned using a stratified sampling method to ensure distribution consistency. When training the state feature extraction network, historical interaction data is divided into training and validation sets, and a stratified sampling strategy is used to group the data according to key features, such as by environmental state type or decision action category. This approach ensures that the training and validation sets have similar distributions under various scenario conditions, avoiding overfitting or underfitting caused by data bias. Stratified sampling combined with training techniques such as early stopping further optimizes the model's generalization ability, enabling the feature extraction network to robustly adapt to dynamic environmental changes. The stratified sampling method used to partition the training data ensures consistent distribution. This design allows the state feature extraction network to quickly adapt to changes in data distribution in dynamic scenarios. The parameters of the feature extraction network can be fine-tuned through online learning, with network weights periodically updated with new data to adapt to data distribution drift.In the implementation of spatiotemporal feature encoding, sliding window segmentation needs to consider computational resource limitations, large sequences may be processed in blocks, and memory management strategies such as streaming processing should be used to avoid overflow. The selection of pre-trained spatiotemporal convolutional networks can be based on transfer learning, and public pre-trained models can be reused to reduce training time.

[0066] In the feature extraction stage of local state subsequences, the number of input channels of the convolutional network matches the dimension of the state variables, the padding strategy keeps the spatial size constant, and the stride controls the downsampling rate. Multi-scale feature fusion is achieved through skip connections or feature pyramids to enhance the utilization of information at different levels of abstraction. In the configuration of the Long Short-Term Memory (LSTM) network, the number of hidden units affects the model capacity; too many units lead to overfitting, while too few lead to underfitting. Hyperparameter tuning uses grid search or Bayesian optimization. Specifically, the number of hidden units in the LSTM network is tuned using a grid search method to balance the model capacity. In the implementation, the search space is defined as follows: candidate values ​​for the number of hidden units are [32, 64, 128, 256], corresponding to 1 or 2 layers in the network (to avoid excessive depth leading to gradient vanishing). The grid search uses 5-fold cross-validation: the training set is evenly divided into 5 subsets, and 4 subsets are used for training and 1 subset for validation, repeated 5 times, and the average performance metric (such as validation set accuracy) is taken. During training, the LSTM network uses `tf.keras.layers.LSTM` layers with tanh activation function and sigmoid activation function for recurrent activation. The initializer uses Glorot uniform initialization. The optimizer is fixed at Adam (learning rate 0.001), with a maximum of 100 training epochs and an early stopping patience value of 5. Evaluation metrics also include training time to select the most efficient configuration. For example, in resource allocation scenarios, 128 hidden units and a single-layer structure achieve the best bias-variance tradeoff in most cases. The grid search process is implemented using Python's `GridSearchCV` class, with parallel processing to accelerate the search. Furthermore, hyperparameter tuning incorporates model complexity penalties: increasing L2 regularization strength when the number of hidden units is too high (e.g., 256), and reducing the dropout ratio when it is too low (e.g., 32). After tuning, the optimal parameters are persisted to a configuration file for direct loading during the inference phase. This approach ensures the LSTM network maintains high robustness in temporal data modeling. Sequence modeling handles variable-length sequences through padding and masking, ignoring invalid time steps. In dimensionality compression, feature selection in principal component analysis is based on eigenvalue magnitude, retaining principal components. The bottleneck layer size of the autoencoder is determined experimentally. Standardization may introduce bias and must be consistently applied during training and inference. The storage of the state feature vector set utilizes efficient data structures, such as HDF5 format, to support large-scale data. Feature retrieval is optimized through indexing. Regularization techniques in network training include weight decay and batch normalization to improve generalization ability. The loss function can combine multiple objectives, such as optimization of reconstruction loss and contrastive loss. A monitoring and logging system is integrated into the feature extraction process to track feature quality and computation time, and invalid input is handled through anomaly detection.The overall implementation emphasizes modularity, with each component capable of independent replacement or upgrades. For example, a convolutional network can be replaced with a graph convolutional network to handle non-Euclidean data, and an LSTM can be replaced with a Transformer to handle long sequences. Real-time requirements are considered during system deployment, and the feature extraction latency must meet the response time constraints of the decision engine. Through these steps, the state feature extraction network can robustly encode complex spatiotemporal dynamics, providing high-quality feature representations for subsequent decision-making.

[0067] Parameter tuning for spatiotemporal feature encoding is an iterative process. Window size and network structure are selected through cross-validation, with validation metrics including feature discrimination or downstream task performance. A caching mechanism is integrated into the data processing pipeline to avoid redundant computation of historical data. Data augmentation techniques during the pre-training phase, such as time warping or adding noise, improve model robustness. Visualization tools for local spatiotemporal feature vectors aid in understanding feature distribution and facilitate debugging. The initialization strategy of Long Short-Term Memory (LSTM) networks affects convergence speed; for example, Xavier initialization maintains gradient stability. Features after dimensionality compression can be used for cluster analysis to explore natural grouping of state patterns. Standardization steps may involve outlier handling, such as pruning extreme values ​​to prevent scaling distortion. Version management of feature vector sets supports experimental reproducibility and records parameter and data processing history. Gradient checking techniques during training numerically verify the correctness of backpropagation. Distributed training accelerates learning on large-scale data, with parameter server architecture coordinating multiple nodes. Interpretable methods for feature extraction networks, such as attention mechanisms, highlight important time steps or variables. Error handling mechanisms in implementation address data corruption or network failures, and rollback strategies ensure system availability. The entire encoding process is seamlessly integrated with the decision action space mapping table construction, and feature consistency checks ensure that the inputs of downstream modules are compatible; resource allocation is dynamically adjusted, and computationally intensive operations are prioritized for scheduling to high-performance hardware.

[0068] Example 2: See Figure 2The construction of the decision action space mapping table begins by querying the historical interaction data set for decision actions performed under the same or similar states based on each state feature vector in the state feature vector set. This step relies on an efficient feature matching algorithm to establish the association between states and historical behaviors. The feature matching process employs a tree-based fast nearest neighbor search algorithm, such as a ball tree or kd-tree. These data structures can quickly locate neighboring points in a high-dimensional feature space. The threshold setting for similarity measurement needs to undergo sensitivity analysis to avoid data noise or information loss caused by overly lenient or strict matching. Specifically, the feature matching process uses the kd-tree algorithm to achieve fast nearest neighbor search. A kd-tree is a binary tree structure used to segment high-dimensional feature spaces. The specific construction process is as follows: a sample is randomly selected from the state feature vector set as the root node, and feature dimensions are alternately selected (e.g., prioritizing the dimension with the largest variance), recursively constructing left and right subtrees using the median as the split point. During the search, given the query state feature vector, the nearest neighbor is found through tree traversal, and the distance metric used is Euclidean distance. The similarity metric threshold θ is set through sensitivity analysis: θ is initially set to 0.1, and the recall-precision curve is calculated based on historical data; θ is searched in a grid within the range [0.05, 0.2] with increments of 0.01, and the threshold that maximizes the F1 score is selected (e.g., θ = 0.12). Sensitivity analysis uses cross-validation to avoid overfitting. During matching, if the nearest neighbor distance is greater than θ, the state is considered unseen, triggering an exploration mechanism; otherwise, the corresponding candidate action is returned. The kd-tree implementation uses the KDTree class from scikit-learn, supporting batch queries and reducing the time complexity from O(n) to O(logn). Furthermore, the tree structure is periodically rebuilt (e.g., every 1000 queries) to adapt to data distribution drift. This design achieves millisecond-level state matching in intelligent manufacturing scheduling, reducing decision latency. When calculating the expected cumulative reward for each decision action under the corresponding state feature vector, the system uses the value function approximation theory of reinforcement learning to solve for the long-term value through iterative calculation of the Bellman equation. The iterative process of the Bellman equation adopts a dynamic programming method, starting from the terminal state and backpropagating the value estimate. In each iteration, the Q-value of all state-action pairs is updated until the numerical change converges within a preset tolerance threshold. The calculation of the expected cumulative reward introduces a discount factor to balance the weight of immediate reward and future reward. The value of the discount factor affects the foresight of the agent and is usually determined through grid search or empirical adjustment.

[0069] The initial decision action space mapping table is formed by storing state feature vectors, candidate decision actions, and calculated expected cumulative rewards in key-value pairs. In the design of these key-value pairs, the key typically uses a hash digest of the state feature vector or a unique identifier to improve retrieval efficiency, while the value is encapsulated as a structured data object containing metadata such as action identifier, expected reward value, and access count. The initial mapping table is often constructed using a batch processing model, fully utilizing distributed computing frameworks such as Spark for parallel construction to handle large-scale historical datasets. The initial decision action space mapping table undergoes action exploration extension processing. This processing aims to address the limitations of purely historical data-based approaches by introducing unverified but constraint-compliant exploratory decision actions to enhance the strategy's exploration capabilities. The generation strategies for exploratory actions can include model-based predictive actions, action variants generated by random perturbations, or heuristic actions based on domain knowledge. All these new actions must pass pre-defined constraint filters, such as physical feasibility checks or business rule verification, to ensure their safety.

[0070] The process of dynamically updating the decision action space mapping table using a policy gradient algorithm involves sampling the current state feature vector from the real-time decision environment. This sampling mechanism is typically connected to an environment simulator or the sensor interface of the actual system, enabling real-time acquisition of the latest state representation. The decision action space mapping table is then queried to obtain a set of candidate decision actions corresponding to the current state feature vector. This query operation needs to handle situations where states may not perfectly match; in such cases, it returns the union of action sets corresponding to similar states or a distance-weighted combination of actions. Based on the action selection probability distribution of the policy gradient algorithm, a target decision action is selected from the candidate set and executed. The policy gradient algorithm typically maintains a parameterized policy function, such as the Softmax policy, which maps state features to a probability distribution in the action space. The action selection process balances exploration and exploitation, employing an ε-greedy policy method. In specific implementations, the initial ε value is set to 0.1, representing a 10% probability of exploration (randomly selecting an action) and a 90% probability of exploitation (selecting the action with the highest expected cumulative reward). ε decays over time to balance the learning phase: every 1000 actions, ε is multiplied by a decay factor of 0.99, with a lower bound of 0.01 to ensure continuous exploration. During strategy implementation, the candidate action set of the current state feature vector is queried from the decision action space mapping table, and a random number r∈[0,1] is generated. If r<ε, an action is uniformly and randomly selected from the candidate actions; otherwise, the action with the largest expected cumulative reward is selected. The reward value is pre-calculated using the Bellman equation and stored in the mapping table. Exploratory actions must pass feasibility verification through the domain knowledge base to avoid unsafe operations. In addition, the ε-greedy strategy is combined with simulated annealing: when the environment changes abruptly (such as an increase in reward variance), ε is temporarily reset to 0.2 to enhance exploration. The entire process is implemented in Python, using numpy.random.choice for random selection to ensure repeatability. This design effectively avoids local optima and improves the robustness of the strategy in medical resource allocation scenarios.

[0071] Recording the state transitions of the environment after executing the selected target decision action and the resulting immediate reward signal is a crucial step in the update loop. The state transition results record the new state feature vector of the environment and whether it has entered the termination state. The immediate reward signal is generated by a predefined reward function, the design of which directly affects the agent's learning objective. Based on the recorded state transitions and reward signals, the policy gradient update is calculated. The policy gradient theorem provides a theoretical basis for the parameter update direction. The calculation of the update usually requires estimating the state-action value function, which can be achieved through Monte Carlo methods, temporal difference learning, or actor-commentator architectures. The calculated policy gradient update is then used to backpropagate and correct the expected cumulative reward value in the decision action space mapping table. The correction process essentially incorporates new empirical evidence into the value estimation. It may involve directly updating the reward value of the queried state-action pair or adjusting the model parameters that generate these values ​​through a function approximator. Backpropagation may be synchronous or asynchronous and is often combined with an experience replay buffer to reuse past experience, improving data utilization efficiency and learning stability.

[0072] In the similarity query stage of mapping table construction, the system builds an index to accelerate retrieval speed under large-scale data. The choice of index structure needs to consider the dimensionality of feature vectors and the sparsity of data distribution. The iterative calculation process of expected cumulative reward needs to deal with the curse of dimensionality that may be caused by large-scale state spaces. At this time, function approximation techniques, such as linear functions, neural networks, or decision trees, are often used to compactly represent value functions. In the action exploration extension processing, the definition and verification of constraints need to be integrated with an independent constraint management module, which can dynamically load and apply domain constraint rules. The injection strategy of exploratory actions needs to be carefully designed to avoid destroying the learned excellent strategies. Usually, its injection probability will gradually decay as the learning process progresses.

[0073] During the policy gradient update phase, the state feature vectors sampled from the real-time environment need to undergo a feature preprocessing pipeline consistent with the training phase to ensure the consistency of the feature space. Obtaining the candidate decision action set may involve complex set operations, such as merging, sorting, and deduplicating action lists corresponding to multiple similar states. The generation of the probability distribution for action selection depends on the current parameters of the policy network, which needs to periodically extract information from the mapping table for updates or co-evolve with the mapping table. The environmental feedback obtained after executing an action needs to be accurately and without delay recorded, and timestamped and labeled with context labels for traceability. The calculation of the policy gradient update involves variance control techniques for gradient estimation, such as baseline subtraction or qualification traces, to ensure the stability and convergence speed of the learning process. When performing backpropagation correction on the mapping table, an appropriate learning rate scheduling strategy needs to be adopted, such as an adaptive learning rate algorithm or annealing strategy, to balance convergence speed and accuracy. The entire dynamic update loop is designed as a continuously running process, capable of adapting to the non-stationarity of the environment, and by continuously incorporating new interaction experiences, the decision action space mapping table gradually approaches the optimal decision policy.

[0074] The storage backend of the mapping table needs to be highly available and scalable to support frequent read / write operations and high concurrency access, potentially employing a distributed key-value store or in-memory database. A version management mechanism for historical interaction data ensures that updates can be rolled back to a previous stable state to address potential learning degradation. The generation and filtering mechanism for exploratory actions needs to be tightly coupled with security checks; any new action must undergo multi-layered verification before being added to the mapping table, including static rule checks and dynamic simulation verification. The implementation details of the policy gradient algorithm, such as the choice of optimizer and support for parallel environment sampling, significantly impact update efficiency. Latency and throughput during real-time sampling and updating need to be monitored and optimized to meet the real-time requirements of the decision support system. Mapping table parameter correction operations need to possess transactional characteristics to ensure data consistency in the event of system failure.

[0075] Example 3: See Figure 3The construction of a real-time decision support engine begins with loading the optimized policy gradient parameter set into an online inference framework. This framework typically employs a high-performance machine learning service architecture, capable of supporting low-latency loading and hot updates of the model. The parameter set is stored in a serialized format, such as Protocol Buffers. The loading process includes integrity checks and version compatibility checks to prevent inference errors caused by data corruption or version inconsistencies. Specifically, the parameter set is stored and loaded using the Protocol Buffers serialization format. In practice, the parameter structure is defined using a Protocol Buffers .proto file: the message type PolicyParams contains repeating fields for storing policy gradient parameters (such as weight matrices and bias vectors), and each parameter includes a name, dimension, a numerical array, and a version number (e.g., v1.0.0). During serialization, the parameter object is encoded into a binary stream using the Protocol Buffers Python API, and a CRC32 checksum is calculated and appended to the end of the file. The loading process includes: reading the binary file and verifying checksum matching; checking if the version number matches the current engine version, otherwise throwing an exception; after deserialization, using assertions to check if the parameter dimensions match the network structure (e.g., the weight shape should conform to the input size of the LSTM layer). Integrity verification is strengthened using a hash function (e.g., SHA-256): the parameter hash value is calculated during storage and recalculated and compared during loading. Version management uses semantic versioning, terminating loading if the major version number is incompatible. The binary file is stored in a distributed file system (e.g., HDFS), supporting incremental updates. This design ensures safe parameter loading and avoids inference errors in the financial trading engine. The resulting real-time responsive policy network is the core computing unit of the engine. Its network structure is usually optimized to balance accuracy and speed, possibly using techniques such as pruning, quantization, or knowledge distillation to reduce model size. The network's forward propagation process is optimized for hardware characteristics, such as using the instruction set of the GPU's TensorCorePU to accelerate matrix operations. Specifically, the policy network uses quantization techniques to optimize the network structure to balance accuracy and speed. In practice, a post-training quantization method is used: FP32 weights are converted to INT8 format, reducing memory usage by 75%. The quantization process is implemented using TensorFlow's TFLite converter: the trained FP32 model is loaded, and the quantization range is calibrated using a representative dataset (such as a validation set); symmetric quantization is applied, mapping weights and activation values ​​to the [-128, 127] interval. During inference, INT8 computation is accelerated using GPU TensorCore hardware, supporting 4x4 matrix multiplication and addition operations.Forward propagation optimizations include: fusing network operations into a single kernel function (e.g., fusing convolution, bias, and ReLU); accelerating matrix multiplication using CUDA's cuBLAS library; and parallelizing 8-bit integer operations using the AVX2SIMD instruction set in a CPU environment. The network structure is simplified to a single hidden layer fully connected network, with the input dimension matching the state feature vector and the output dimension consistent with the action space. The quantized model is deployed on a Triton inference server, supporting dynamic batch processing with latency below 10 milliseconds. This design enables high-throughput decision-making in device scheduling scenarios. The state feature preprocessing pipeline configured for the policy network is responsible for converting the raw environment state into a normalized state feature vector. This pipeline is designed with a modular, pluggable architecture, allowing flexible adjustments to processing steps based on the characteristics of different data sources. The pipeline's input interfaces with a data acquisition system, capable of receiving real-time data streams from sensors, databases, or message queues. After receiving the raw environmental state data, the pipeline performs missing value imputation and noise filtering. The missing value handling strategy is selected based on the data characteristics; for time series data, forward imputation, linear interpolation, or prediction methods based on neighboring data may be used. Noise filtering may employ moving average filters, median filters, or more complex wavelet transform-based denoising algorithms to eliminate interference from random fluctuations and measurement errors. Calling the exact same spatiotemporal convolutional network and long short-term memory network used in the training phase to extract features from the cleaned environmental state data is crucial to ensuring feature consistency. The network weights and structure must remain absolutely consistent with those in the training phase. The feature extraction process is executed in a dedicated inference engine, utilizing the TensorRT framework for efficient cross-platform inference. Specifically, the feature extraction process uses the TensorRT inference engine to ensure consistency. In practice, the spatiotemporal convolutional network and LSTM network from the training phase are converted into an optimization engine via TensorRT's Python API: the TensorFlow model is exported to ONNX format; an optimization computation graph is built using TensorRT's Builder, applying layer fusion, accuracy calibration, and automatic kernel tuning. Network weights are loaded from training checkpoints to ensure absolute consistency with the training phase; structural parameters such as convolutional kernel size and the number of LSTM hidden units are hard-coded in configuration files. During inference, the cleaned environment state data is converted into a tensor format supported by TensorRT, and synchronous inference is performed using the execute method. The engine is deployed in Docker containers, supporting GPU inference (such as NVIDIA T4), and utilizes TensorRT's dynamic shape handling for variable-length sequences. Consistency checks include: verifying the shape of the input tensor matches the training data before inference, and performing similarity matching between the output feature vector and the decision action space mapping table.A key step is to perform similarity matching between the extracted feature vectors and the state feature vectors in the decision action space mapping table. The purpose is to locate real-time features into the existing knowledge structure. The matching algorithm needs to be efficient and accurate. For example, it can use approximate nearest neighbor search techniques such as locality-sensitive hashing or product quantization to complete the fast retrieval of high-dimensional vectors in milliseconds. The threshold for successful matching needs to be set carefully. Too low a threshold may lead to matching irrelevant states, while too high a threshold may prevent the system from responding to unseen states.

[0076] The established post-processing mechanism for action output performs multi-level verification and correction of the original decision actions output by the policy network. This mechanism is an important bridge connecting the learned policy and actual physical constraints. The post-processing mechanism obtains the original decision actions output by the policy network. These actions are directly generated based on the learned policy function and may not consider the hard constraints of the instantaneous environment. The first step of post-processing is to query the domain knowledge base to verify the physical feasibility of the original decision actions. The domain knowledge base stores constraints such as system dynamics, operational safety range, and business rules in a structured form. The verification process is usually transformed into a constraint satisfaction problem, checking whether all action parameters fall within the feasible domain defined by each variable. Actions that pass the verification can be directly marked as the final output actions; while for original decision actions that violate physical constraints, a correction process is initiated, replacing them with the most feasible candidate decision actions that exist in the same or most similar state in the decision action space mapping table. The correction logic needs to weigh the similarity of actions with the reward value, prioritizing the action with the highest expected cumulative reward under the constraint conditions. The state feature preprocessing pipeline deeply integrates data quality monitoring, calculating the statistical characteristics of input data in real time, such as mean, variance, and outlier ratio, and dynamically adjusting preprocessing parameters. During feature extraction, the batch size for network inference is optimized to fully utilize computing resources while meeting the latency cap required for real-time performance. A priority mechanism is introduced in the similarity matching process, assigning higher matching priority to frequently occurring or high-value state features to accelerate the response speed of commonly used states. The uncertainty of matching results is quantified by calculating matching scores or confidence levels; low-confidence matches trigger the system's exploration mechanism or degradation strategy. The implementation of the action output post-processing mechanism includes an asynchronous simulation verification step. For certain complex or high-risk actions, a lightweight simulation environment is used for instantaneous extrapolation before final output to predict the short-term consequences of the action, further reducing risk. The entire real-time decision support engine's operational status is comprehensively monitored, including inference latency, decision quality metrics, and system load. Monitoring data is used for the engine's elastic scaling and performance tuning. The engine is designed with a fault-tolerant architecture, possessing degradation strategies in case of failure at any stage of preprocessing, inference, or post-processing, such as switching to a rule-based backup decision generator to ensure continuous system availability. Through this meticulous design and implementation, the real-time decision support engine can safely, reliably, and efficiently apply optimized strategies trained offline to dynamically changing real-world environments.

[0077] Integration of the engine with external systems is achieved through well-defined API interfaces, typically using standard protocols such as gRPC, supporting high-concurrency requests and asynchronous response modes. Specifically, the service contract is defined using Protocol Buffers: the DecisionService service contains the RPC method GetOptimalAction, with the input message being StateRequest (containing environment state data) and the output message being ActionResponse (containing the decision action). The gRPC server is implemented in C++, bound to port 50051, and supports HTTP / 2 multiplexing to handle high-concurrency requests. Asynchronous response mode is handled through CompletionQueue: client requests are placed in the queue, and the server uses a thread pool (e.g., 4 threads) to process feature extraction and inference in parallel, returning the result via callback upon completion. Concurrency control is set to a maximum of 1000 connections and a timeout of 5 seconds. Load balancing is implemented through a gRPC client-side load balancer, supporting multiple engine instances. The API interface integrates authentication mechanisms (such as TLS / SSL encryption) to ensure data security. This design supports thousands of decision requests per second in real-time risk control systems. The structure of request / response messages is carefully designed, including necessary metadata such as timestamps, session IDs, and data sources to facilitate issue tracking and system debugging. The online learning capability of the policy network allows the engine to continuously fine-tune its parameters during runtime. This continuous learning function is achieved through an incremental learning algorithm. Newly arrived interaction data is filtered and weighted before being used in small batches for fine-tuning network parameters, with the learning rate set extremely low to avoid disrupting already learned effective policies. Data transformation operations in the preprocessing pipeline, such as normalization and standardization, use parameters (e.g., mean, standard deviation) that are completely consistent with those used in the model training phase. These parameters are persisted in configuration files and loaded during pipeline initialization to prevent feature distortion due to data distribution shifts. Engine performance optimization involves multiple levels. At the code level, computations on critical paths are implemented using efficient base libraries, such as the BLAS library for matrix operations. At the system level, caching mechanisms are used to store frequently used feature vectors or matching results, reducing redundant computations. The resource management module dynamically allocates CPU, memory, and GPU resources to ensure service quality is maintained even under high load. The engine's deployment scheme considers high availability requirements, potentially employing multi-instance deployment and load balancing strategies, ensuring that a single instance failure will not cause service interruption. Security considerations are integrated throughout the engine's design, including validating input data to prevent injection attacks, validating the reasonableness of output actions to avoid dangerous operations, and implementing strict access control to ensure that only authorized systems can trigger the decision-making process.The logging system records every decision cycle of the engine in detail, including input status, candidate actions, final output actions, and post-processing reasons. These logs are used for offline analysis, model evaluation, and system auditing.

[0078] Example 4: The implementation process of the action output post-processing mechanism is illustrated using an industrial boiler temperature control scenario as an example. In this scenario, the decision action is to adjust the opening of the fuel valve, and the environmental state includes multi-dimensional variables such as water temperature, pressure, and flow rate. The system obtains the original decision action output by the strategy network. For example, the network might directly output an instruction to instantly increase the valve opening by 40%. Mathematically, this instruction is the optimal solution for the reward function, but it may exceed the safe operating range of the equipment. Verifying the physical feasibility of the original decision action by querying the domain knowledge base is the foundation for subsequent processing. The domain knowledge base stores the physical constraint rules of the boiler system, such as structured conditions like "a single valve opening change must not exceed 15%" and "it is forbidden to increase the valve opening when the pressure exceeds 2MPa." The verification process compares the action parameters with the constraints in the knowledge base one by one, generating a feasibility report. When the verification finds that the original action violates the "single change range" constraint, the system triggers the nearest neighbor replacement processing logic. This logic queries the decision action space mapping table for the K historical states most similar to the current state feature vector. From the candidate decision actions corresponding to these states, the action that satisfies all constraints and has the highest expected cumulative reward is selected as the alternative.

[0079] Referring to Table 1, the construction of the domain knowledge base begins with collecting equipment operation manuals, safety procedures, and technical documents related to boiler temperature control. These unstructured documents are transformed into a machine-readable set of structured constraints through entity recognition and relation extraction using natural language processing technology. For example, key parameters such as "safety threshold" and "maximum adjustment rate" are extracted from the documents, and logical relationships between parameters are established. Establishing a mapping relationship table between constraints and decision actions requires clarifying the legal value range and associated conditions of each action variable. The mapping relationship table is stored using a relational database table structure, facilitating efficient joint queries and dynamic updates. The table design includes fields such as action type, parameter boundaries, and effective conditions, forming a complete constraint rule network. Regularly receiving abnormal action records from the real-time decision support engine is an important way for the knowledge base to evolve. Abnormal records contain action parameters, violated constraint types, and environmental context. The system discovers new potential constraint rules from these records using pattern mining algorithms, which are then added to the mapping relationship table after manual review.

[0080] Table 1: Boiler Control Action Constraint Mapping Table

[0081] The physical feasibility verification in the post-processing mechanism for action output employs a multi-level verification strategy. The first level performs a static range check to quickly filter out parameter values ​​that are clearly out of bounds. The second level performs a dynamic correlation verification, combining the current environmental state to determine whether the action is allowed to execute. Nearest neighbor replacement processing considers not only the numerical proximity of actions but also introduces a safety weight factor, prioritizing safe actions that have been executed frequently in the historical record and have never caused anomalies. The replacement algorithm uses a weighted scoring mechanism to comprehensively evaluate the action's reward value, safety record, and state matching degree. When the corrected decision action is marked as the final output action, the system generates a detailed correction log, recording the original action, the reason for the correction, the replacement action, and its expected reward. These logs are used for subsequent analysis and strategy optimization. Version management of the domain knowledge base uses semantic version control. Each addition, deletion, or modification of constraint entries generates a new version number, and the modification history and basis are retained. The knowledge base access interface provides version query and rollback functions to ensure the traceability of constraint updates. The process of dynamically expanding constraint entries includes a conflict detection mechanism. When a newly added constraint logically contradicts an existing constraint, the system will prompt the administrator for manual resolution. Expansion operations are typically scheduled during periods of low system load, and transaction processing is used to ensure data consistency. Anomaly logs are collected using a sliding time window, focusing on frequently occurring anomaly patterns to promptly identify changes in the system's operating environment.

[0082] The integration of the post-processing mechanism and the real-time control system is achieved through an event-driven architecture. When the policy network outputs an action, an asynchronous verification request is triggered, and the verification result is returned to the decision engine via a message queue. This design avoids blocking the main decision-making process and ensures the system's real-time responsiveness. The execution effect of the corrected action is continuously monitored. If multiple corrections occur consecutively and the reward value of the corrected action is significantly lower than the original action, it may indicate that the policy network needs retraining or that the constraints are too strict. The domain knowledge base maintenance interface provides visualization tools, allowing administrators to graphically view and adjust constraints. The system automatically checks the compatibility of the modified constraints with existing knowledge. When synchronously updating the training sample library of the policy network, correction records are assigned data labels, indicating that these are samples that have undergone post-processing correction. These samples can be used to strengthen constraint awareness learning in subsequent training. The sample library update strategy considers data balance to avoid post-processed samples excessively affecting the exploratory nature of the original policy. The selection criteria for important samples include the magnitude of the correction, the degree of reward difference, and the scarcity of the action. The implementation of the entire post-processing mechanism enables the reinforcement learning system to gradually improve decision-making quality while strictly adhering to domain knowledge constraints, effectively integrating data-driven intelligence with human experience and knowledge. In actual operation, the post-processing mechanism may encounter boundary conditions. For example, when there are no fully satisfying alternative actions in the decision action space mapping table, the system will initiate a degradation strategy, selecting partially satisfying actions according to preset priorities, or directly outputting a safe and conservative default action. Performance optimization of the post-processing component includes establishing a cache index for constraint condition queries and pre-compiling frequently used verification rules; the component's health status is monitored through heartbeat detection and timeout mechanisms, automatically switching to bypass mode in case of anomalies.

[0083] See Figure 4The diagram above illustrates the dynamic changes of four key state variables in the boiler system. The water temperature curve reflects the stability of the heating control, aiming to maintain it within the ideal range of 65-75°C. The pressure curve shows the system pressure fluctuating within safety constraints; when it approaches the 2.3MPa warning value, the system automatically adjusts its control strategy. The flow rate curve reflects the effect of pump speed regulation, maintaining it within a reasonable range of 120-180. The water level curve shows the balanced state of the water supply system, maintaining an optimal range of 40-60% through intelligent regulation. The coordinated changes of the curves demonstrate the system's ability to coordinate multi-variable control. The diagram below compares the original decision actions with the action sequences corrected by post-processing. The dashed line represents the original action directly output by the neural network, while the solid line represents the safe action after physical constraint verification and nearest neighbor replacement processing. Several key points can be observed: when the original valve adjustment exceeds the safety limit of ±15% / s, the system automatically corrects it to within the constraint range; under high pressure, the system prohibits dangerous operations that significantly increase valve pressure; the constraint boundaries represented by the dotted lines provide a safe operating range for the actions. This post-processing mechanism effectively prevents equipment damage and safety accidents while maintaining control performance.

[0084] Example 5: The process of synchronously updating the policy network begins by combining the revised decision action with the environmental state transition result and the immediate reward signal into a new training sample. This combination operation ensures that each sample contains a complete four-tuple of information: state, action, reward, and new state, forming a standard reinforcement learning transition sample. The format of the new sample is consistent with historical data for easy unified processing and management. A timestamp and source label are attached when the sample is generated to trace the data origin. Calculating the feature space distance between the new training sample and the samples in the existing sample library is a key step in determining whether to add a sample. The feature space distance is calculated using the Mahalanobis distance method. This distance metric considers the covariance structure of the feature vectors and reflects the data distribution characteristics better than Euclidean distance. Before distance calculation, the feature vectors need to be standardized to eliminate the influence of dimensions and ensure the fairness of the distance calculation. The distance threshold θ is set based on the diversity requirements of the sample library and is dynamically adjusted by analyzing the distribution density of historical samples. Too high a threshold may lead to the addition of redundant samples, while too low a threshold may reject valuable new samples. The distance calculation process utilizes batch matrix operations to optimize efficiency and support rapid comparison of large-scale sample libraries.

[0085] When the sample library capacity reaches its limit, the system automatically triggers an importance sampling strategy to eliminate the most redundant historical samples. The sample library capacity limit is preset based on available memory and performance requirements, typically ranging from tens of thousands to millions of samples. The elimination strategy aims to retain samples with high information content while removing duplicate or low-value samples, thereby maintaining the vitality and representativeness of the sample library. The core of the importance sampling strategy is to assign dynamic weights to each training sample. The calculation is based on the exploratory value of decision-making actions in the sample. and historical usage frequency Value is explored through prediction error or uncertainty estimation of the value function, and historical frequency records show the number of times samples were used for training; the weight allocation formula is:

[0086] ;

[0087] in: Indicates sample Dynamic weights, Indicates sample The value of exploration Indicates sample Historical usage frequency It is a small positive constant used to prevent division by zero errors and invalid logarithmic parameters. This formula assigns higher weights to samples with high exploration value but low usage frequency, encouraging the retention of novel and informative samples.

[0088] Regularly calculating the weight distribution of all samples in the sample library is fundamental to management operations. This calculation generates a cumulative distribution function by statistically analyzing the weight values ​​of all samples, which is used to determine the elimination boundary. The calculation cycle is set according to the sample update frequency; shorter cycles (e.g., once per minute) are used in high-update environments, while longer cycles (e.g., once per hour) are used in low-update environments. When removing low-weight samples after weight sorting, the system arranges samples in ascending weight order. The removal ratio is dynamically determined based on the number of new samples; for example, if N new samples are planned to be added, the N oldest samples with the lowest weights are removed. Removal operations use transaction processing to ensure data consistency and avoid errors caused by concurrent access. Sample vacancies triggered by removal operations are used to store newly collected high-value training samples. Before adding a new sample, its distance to the remaining samples is verified again to prevent immediate redundancy. Metadata such as the addition time and initial weight values ​​are recorded during the sample addition process for easy subsequent management. In the implementation of synchronously updating the training sample library, the combination operation of new samples integrates a data verification step, checking the completeness of sample fields and the rationality of their value ranges. Invalid samples are discarded and logged. The calculation of feature space distance uses an approximate nearest neighbor algorithm for acceleration; for example, locality-sensitive hashing maps high-dimensional vectors to a low-dimensional space for fast similarity searching. The dynamic adjustment mechanism of the distance threshold θ monitors the coverage metric of the sample library. When the average distance between a new sample and its nearest neighbor remains consistently low, the threshold is automatically increased; when the average distance is high, the threshold is decreased to attract more samples. Sample library capacity management employs a hierarchical storage strategy: frequently used samples are stored in memory, while infrequently used samples are archived to disk, balancing access speed and storage costs.

[0089] The weight updates for the importance sampling strategy are performed simultaneously with sample usage, and each time a sample is sampled for training, its historical usage frequency is considered. Incremental growth, exploring value The weights are updated based on the sample's contribution to training, for example, by measuring temporal difference error. The weight calculation process avoids frequent global recalculations by using an incremental update mechanism, adjusting weights only when a sample is accessed or modified. Weight distribution calculation is optimized by maintaining a weight heap data structure to track minimum and maximum weights in real time, quickly obtaining quantile information. Sample removal operations consider the correlation between samples to avoid removing too many samples from the same scenario at once, which could lead to distribution bias. The priority of adding new samples is based on their weight and diversity contribution; the system may temporarily postpone adding samples with too low a weight even if there are vacancies, waiting for higher-quality samples. The persistent storage of the sample library uses a compact columnar format to reduce disk usage, and periodic snapshots prevent data loss. The sample retrieval interface supports conditional queries and random sampling to meet the needs of different training algorithms. The entire synchronous update mechanism is asynchronous with the training process, decoupled through a producer-consumer model to avoid blocking the training process. The monitoring system tracks the sample library size, weight distribution, and diversity indicators, triggering alarms when indicators are abnormal. Through the above design, the training sample library can continuously evolve, providing high-quality and diverse training data for the policy network and promoting a steady improvement in intelligent decision-making capabilities.

[0090] Exploring the value in the specific calculation of dynamic weights The estimation may employ various methods, such as prediction variance based on ensemble models or exploration rewards based on state access frequency. The estimation process utilizes online learning algorithms to adapt to environmental changes; historical usage frequency; The records take into account time decay, and recent data is weighted more heavily to reflect current needs. The weighting formula... Values ​​are typically set to 1 or less, adjusted according to numerical stability requirements; variations in the formula may introduce adjustments to balance exploration and utilization, but the core idea remains consistent. Weight distribution calculations may employ random sampling estimation to reduce computation, especially when the sample library is extremely large; low-weight samples removed may be archived to secondary storage instead of being immediately deleted, for later analysis or recovery. Distance verification before sample addition may introduce a secondary threshold, prioritizing the addition of new samples when their distances to multiple nearest neighbors exceed the threshold, enhancing diversity; the sample library's capacity limit may be flexibly adjusted, dynamically scaling based on system load and data value. The implementation of the importance sampling strategy considers fairness, avoiding the systematic removal of samples in certain scenarios, and protecting rare state samples through the design of the weight formula; the sample library's update log is used for auditing and debugging, recording the history of each sample's addition, removal, and weight changes. Interaction with the policy network occurs through a standard data interface; samples are batch-sampled from the sample library during training, and feedback is used for information-based closed-loop optimization after updates. The implementation of this entire embodiment makes the training sample library an intelligent data management component, capable of adaptively maintaining data quality and supporting the long-term stable operation of the reinforcement learning system in complex environments.

[0091] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0092] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. An intelligent decision support method based on reinforcement learning, characterized in that, The method includes: Acquire a set of historical interaction data under the target decision-making scenario, wherein the set of historical interaction data includes a sequence of decision-making actions, a sequence of environmental states, and corresponding instantaneous reward signals; The historical interaction data set is input into a state feature extraction network for spatiotemporal feature encoding to generate a set of state feature vectors with temporal correlation. A decision action space mapping table is constructed based on the set of state feature vectors. The decision action space mapping table records the candidate decision action and its expected cumulative reward corresponding to each state feature vector. The decision action space mapping table is dynamically updated using a policy gradient algorithm to generate an optimized set of policy gradient parameters. A real-time decision support engine is constructed based on the optimized set of policy gradient parameters. The real-time decision support engine is used to respond to changes in environmental state and output the optimal decision action.

2. The intelligent decision support method based on reinforcement learning according to claim 1, characterized in that, The step of inputting the historical interaction data set into a state feature extraction network for spatiotemporal feature encoding processing to generate a set of state feature vectors with temporal correlation includes: The environmental state sequence is extracted from the historical interaction data set, and the environmental state sequence is segmented by a sliding window to generate multiple local state sub-sequences. A pre-trained spatiotemporal convolutional network is invoked to perform feature extraction processing on each of the local state subsequences, generating a local spatiotemporal feature vector; The local spatiotemporal feature vector is input into a long short-term memory network to model temporal dependencies and generate a global state feature vector. The global state feature vector is subjected to dimensionality compression to generate a standardized set of state feature vectors.

3. The intelligent decision support method based on reinforcement learning according to claim 2, characterized in that, The construction of the decision action space mapping table based on the set of state feature vectors includes: Based on each state feature vector in the state feature vector set, query the decision actions executed under the same or similar states in the historical interaction data set; Calculate the expected cumulative reward for each decision action under the state feature vector, which is obtained by iterative calculation using the Bellman equation; The state feature vector, candidate decision actions, and expected cumulative reward are stored in key-value pairs to form an initial decision action space mapping table. The initial decision action space mapping table is extended by action exploration, and an unverified but constraint-compliant exploratory decision action is added to each state feature vector.

4. The intelligent decision support method based on reinforcement learning according to claim 3, characterized in that, The step of dynamically updating the decision action space mapping table using the policy gradient algorithm includes: Sample the current state feature vector from the real-time decision-making environment and query the decision action space mapping table to obtain the corresponding set of candidate decision actions; Based on the action selection probability distribution of the policy gradient algorithm, a target decision action is selected from the set of candidate decision actions and executed; Record the transition results of the environmental state after the execution of the target decision action and the obtained immediate reward signal; Based on the transition results and the immediate reward signal, the policy gradient update amount is calculated, and the expected cumulative reward in the decision action space mapping table is corrected by backpropagation.

5. The intelligent decision support method based on reinforcement learning according to claim 4, characterized in that, The step of constructing a real-time decision support engine based on the optimized policy gradient parameter set includes: The optimized policy gradient parameter set is loaded into the online inference framework to form a policy network that can respond in real time; Configure a state feature preprocessing pipeline for the policy network, which is used to convert the original environment state into a state feature vector; A post-processing mechanism for action output is established to perform feasibility verification and boundary constraint correction on the decision actions output by the policy network.

6. The intelligent decision support method based on reinforcement learning according to claim 5, characterized in that, The state feature preprocessing pipeline includes: Receive raw environmental state data and perform missing value imputation and noise filtering on the raw environmental state data; The same spatiotemporal convolutional network and long short-term memory network as those used in the training phase are invoked to extract features from the cleaned environmental state data. The extracted feature vectors are matched with the state feature vectors in the decision action space mapping table to ensure consistency of the feature space.

7. The intelligent decision support method based on reinforcement learning according to claim 6, characterized in that, The post-processing mechanism for the action output includes: Obtain the original decision action output by the policy network, and query the domain knowledge base to verify the physical feasibility of the original decision action; For original decision actions that violate physical constraints, nearest neighbor replacement is performed, replacing them with the candidate decision actions with the highest feasibility under the same state in the decision action space mapping table; The corrected decision action is marked as the final output action, and the training sample library of the policy network is updated synchronously.

8. The intelligent decision support method based on reinforcement learning according to claim 7, characterized in that, The methods for constructing the domain knowledge base include: Collect physical constraint rules and operational specifications documents for the target decision-making scenario, and convert them into a set of structured constraint conditions; Establish a mapping table between constraints and decision actions, and record the range of boundary parameters that each decision action is allowed to execute; The system periodically receives abnormal action records from the real-time decision support engine and dynamically expands the constraint entries in the mapping table.

9. The intelligent decision support method based on reinforcement learning according to claim 8, characterized in that, The training sample library of the synchronous update policy network includes: The revised decision-making actions, environmental state transition results, and immediate reward signals are combined into new training samples. Calculate the feature space distance between the new training sample and the samples in the existing sample library; if it exceeds a preset threshold, add it to the sample library. When the sample library capacity reaches its limit, an importance sampling strategy is used to eliminate the historical samples with the highest redundancy.

10. The intelligent decision support method based on reinforcement learning according to claim 9, characterized in that, The importance sampling strategy includes: Each training sample is assigned a dynamic weight, which depends on the exploratory value and historical usage frequency of the decision action in the sample; Periodically calculate the weight distribution of all samples in the sample library and remove low-weight samples after weight sorting; The sample gaps triggered by the removal operation will be used to store newly acquired high-value training samples.

Citation Information

Patent Citations

  • Table processing method and device driven by natural language, equipment and medium

    CN120542396A

  • Virtual power plant response optimization scheduling system and method based on reinforcement learning

    CN120728750A

  • Unmanned vehicle dynamic path planning method based on multi-source sensor fusion

    CN120778136A

  • Intelligent traffic management using reinforcement learning

    DE202025102042U1

Cited By

  • System based on AI multi-dimensional model and intelligent decision-making method

    CN121880912A

  • Systems and Intelligent Decision-Making Methods Based on AI Multi-Dimensional Models

    CN121880912B