Crane Remote Instruction Response Delay Detection and Prior and Subsequent Compensation Method and System
Through a deep graph neural network with dynamic knowledge graph and causal reasoning mechanism, combined with hybrid decision-making and dual closed-loop evolution system, the problem of remote command response delay in remote control of cranes is solved, efficient and reliable delay compensation is achieved, and operation stability and efficiency are improved.
Patent Information
- Application Number
- CN202510578648.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-07
AI Technical Summary
In the remote control of cranes, the problem of remote command response delay cannot be effectively solved, resulting in reduced operating accuracy and increased accident risk. The existing compensation strategies lack adaptability and robustness, making it difficult to deal with complex delays.
The physical data mapping is used based on dynamic knowledge graphs and adaptive spatiotemporal alignment algorithms, and the deep graph neural network combined with causal reasoning mechanisms are feature extraction, a hybrid decision-making system is built for distributed optimization, and a dual closed-loop evolution system with online learning capabilities is established to achieve accurate prediction and compensation of delays.
It improves the stability and reliability of remote operation of cranes, reduces system maintenance costs, enhances equipment service life and operating efficiency, and significantly reduces the risk of safety accidents.
Smart Images

Figure CN120103715B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to crane technology, and particularly to a method and system for detecting the response delay of remote commands of a crane and compensating for the delay before and after the command Background Art
[0002] Remote control of cranes plays a crucial role in modern industry. Especially in scenarios such as ports and construction sites, efficient and precise remote operation is essential for ensuring production safety and improving production efficiency. The core of remote control lies in real-time performance and reliability, and operation commands need to be executed quickly and accurately. However, due to factors such as signal transmission, network latency, and the response of complex mechanical systems, there is inevitably a delay in the execution of remote commands. This delay will reduce the operation accuracy, increase the risk of accidents, and even lead to serious production accidents.
[0003] To solve the problem of remote command response delay, existing technical solutions mainly focus on the following aspects: Network optimization-based solutions attempt to reduce network latency by improving network transmission protocols and optimizing network topologies; Prediction model-based solutions try to establish mathematical models to predict the delay and perform pre-compensation; Control algorithm-based solutions focus on designing advanced control algorithms to overcome the impact of the delay.
[0004] Most existing solutions only focus on a single aspect of network latency or system response delay, lacking consideration of the combined effects of the two, and it is difficult to effectively handle complex delay situations in actual operations. Traditional prediction models are often based on simplified assumptions and are difficult to accurately capture the complex non-linear relationships and dynamic changes in the actual system, resulting in insufficient prediction accuracy and thus affecting the compensation effect. Existing compensation strategies usually lack adaptability and robustness, and it is difficult to cope with changes in different operating environments and task requirements, resulting in unstable compensation effects and even potentially causing new control problems. Summary of the Invention
[0005] Embodiments of the present invention provide a method and system for detecting the response delay of remote commands of a crane and compensating for the delay before and after the command, which can solve the problems in the prior art.
[0006] In the first aspect of the embodiments of the present invention,
[0007] There is provided a method for detecting the response delay of remote commands of a crane and compensating for the delay before and after the command, including:
[0008] Construct a dynamic knowledge graph based on physical prior knowledge, and use an adaptive spatio-temporal alignment algorithm to map the real-time obtained physical domain operation data to the dynamic knowledge graph to generate a hierarchical feature representation vector; input the feature representation vector into a deep graph neural network integrating a causal inference mechanism, and generate a correlation matrix representing physical laws and data patterns through recursive inference; use a multi-head attention network with a residual structure to extract temporal features from the correlation matrix, and combine with a variational Bayesian inference unit to output a delay prediction tensor containing uncertainty quantification;
[0009] Construct a hybrid decision-making system with hierarchical adaptive capabilities based on the delay prediction tensor; in the policy generation layer of the hybrid decision-making system, fuse the current system state with the delay prediction tensor to generate an initial set of compensation actions; in the collaborative optimization layer of the hybrid decision-making system, combine a multi-agent experience replay mechanism to perform distributed optimization on the initial set of compensation actions, and perform multi-constraint projection on the optimized compensation actions to output an optimal compensation strategy considering safety and real-time performance;
[0010] For the optimal compensation strategy, establish a double-loop evolution system with online learning capabilities; in the prediction optimization loop of the double-loop evolution system, design an error evaluation module based on Bayesian posterior inference, and adaptively adjust the causal inference mechanism and feature extraction strategy through meta-learning methods; in the compensation optimization loop of the double-loop evolution system, construct a multi-level progressive evaluation framework, and dynamically optimize the policy generation layer and collaborative optimization layer based on the compensation effect.
[0011] Construct a dynamic knowledge graph based on physical prior knowledge, and use an adaptive spatio-temporal alignment algorithm to map the real-time obtained physical domain operation data to the dynamic knowledge graph to generate a hierarchical feature representation vector; input the feature representation vector into a deep graph neural network integrating a causal inference mechanism, and generate a correlation matrix representing physical laws and data patterns through recursive inference, including:
[0012] Construct a dynamic knowledge graph based on physical prior knowledge, transform the physical prior knowledge into a parameterized constraint equation set, define the node attributes and edge relationships of the dynamic knowledge graph based on the parameterized constraint equation set, the node attributes represent the evolution law of physical quantities, and the edge relationships represent the coupling mechanism between physical quantities, generating an initial knowledge graph structure with physical law constraints;
[0013] Use an adaptive spatio-temporal alignment algorithm to map the real-time obtained physical domain operation data to the dynamic knowledge graph, design a feature extraction strategy based on the topological relationship of the initial knowledge graph structure, align the physical domain operation data through a dynamic window mechanism, use a multi-level feature fusion network to extract spatial features, and perform semantic matching between the extracted features and the nodes of the dynamic knowledge graph to generate a hierarchical feature representation vector reflecting the multi-scale dynamic characteristics of the system;
[0014] Input the hierarchical feature representation vector into the deep graph neural network with an integrated causal inference mechanism, construct a causal discovery criterion based on the physical constraints of the initial knowledge graph structure, establish initial causal links through conditional independence testing and directed information entropy analysis, and use the message passing algorithm with an attention mechanism to recursively infer and dynamically update the initial causal links, generating an association matrix representing the physical laws and data patterns.
[0015] Input the hierarchical feature representation vector into the deep graph neural network with an integrated causal inference mechanism, construct a causal discovery criterion based on the physical constraints of the initial knowledge graph structure, establish initial causal links through conditional independence testing and directed information entropy analysis, and use the message passing algorithm with an attention mechanism to recursively infer and dynamically update the initial causal links, generating an association matrix representing the physical laws and data patterns, including:
[0016] Input the hierarchical feature representation vector into the deep graph neural network, encode the physical constraints based on the initial knowledge graph, construct a physical constraint criterion including energy conservation and momentum conservation, and use the physical constraint criterion to perform a structured representation of the hierarchical feature representation vector, generating a feature map structure that satisfies the physical laws;
[0017] Calculate the conditional mutual information between nodes based on the feature map structure, evaluate the causal strength between node pairs through directed information entropy analysis, combine the conditional mutual information and causal strength to construct initial causal links, and generate an initial causal graph with weights;
[0018] Use the message passing algorithm with an attention mechanism to dynamically optimize the initial causal graph, construct a multi-head attention layer to calculate the associated features between nodes, input the associated features into a gated recurrent unit to model the temporal dependence relationship, and update the weights of the causal links through iterative message passing, generating an optimized causal graph structure;
[0019] Generate an association matrix based on the optimized causal graph structure, use an adaptive weight mechanism to adjust the matrix elements according to the prediction error, project the association matrix into the feasible region that satisfies the physical constraints through gradient projection, perform uncertainty analysis on the association matrix within the feasible region, and generate the final mapping relationship matrix between the physical laws and data patterns.
[0020] In the collaborative optimization layer of the hybrid decision-making system, combine the multi-agent experience replay mechanism to perform distributed optimization on the initial set of compensation actions, perform multi-constraint projection on the optimized compensation actions, and output the optimal compensation strategy considering safety and real-time performance, including:
[0021] Construct a dynamic graph attention network in the collaborative optimization layer of the hybrid decision-making system, calculate the similarity of state vectors between agents to obtain the initial attention scores, and construct a hierarchical experience replay structure based on the initial attention scores. The hierarchical experience replay structure includes a local experience pool and a global experience pool. Store the initial compensation action set and its execution effects in the local experience pool;
[0022] Perform priority sampling based on experience value from the global experience pool, optimize the initial compensation action set using a distributed policy optimization algorithm, and project the optimized compensation actions into the feasible region that satisfies the constraints using an adaptive gradient projection algorithm;
[0023] Perform multi-level evaluation on the projected compensation actions, calculate the immediate reward based on the compensation error, calculate the long-term benefit through state value evaluation, update the experience values in the local experience pool and the global experience pool according to the immediate reward and the long-term benefit, and at the same time adjust the attention weight matrix based on the evaluation results, optimize the interaction structure of the adaptive communication topology, and output the optimized optimal compensation strategy through collaboration.
[0024] Performing priority sampling based on experience value from the global experience pool, optimizing the initial compensation action set using a distributed policy optimization algorithm, and projecting the optimized compensation actions into the feasible region that satisfies the constraints includes:
[0025] Calculate the cumulative return value of the compensation actions based on the global experience pool, use the cumulative return value as an experience value index, construct a priority sampling tree based on the experience value index, and extract priority value experience samples from the priority sampling tree;
[0026] Input the priority value experience samples into the distributed policy optimization algorithm. The distributed policy optimization algorithm includes an action generation network and a value evaluation network. The action generation network generates an initial compensation action set based on the priority value experience samples, calculates the expected return of the compensation actions using the value evaluation network, and optimizes the initial compensation action set based on the expected return;
[0027] Input the optimized compensation actions into the constraint processing module. The constraint processing module constructs state space constraints based on the system dynamics model, maps the system safety index to state variable constraints, converts the system real-time index into calculation time constraints, and uses the barrier function method to integrate the state space constraints, state variable constraints, and calculation time constraints into dynamic constraint conditions;
[0028] Perform constraint processing on the optimized compensation actions based on the dynamic constraint conditions, calculate the gradient projection step size based on the degree of constraint violation, solve the projection direction through the conjugate gradient method, and project the optimized compensation actions into the feasible region that satisfies the dynamic constraint conditions using a progressive projection strategy.
[0029] Regarding the optimal compensation strategy, a double-loop evolutionary system with online learning ability is established; in the prediction optimization loop of the double-loop evolutionary system, an error evaluation module based on Bayesian posterior inference is designed, and the causal inference mechanism and feature extraction strategy are adaptively adjusted through meta-learning methods, including:
[0030] Construct an online learning double-loop evolutionary system to continuously optimize the optimal compensation strategy. Set up a prediction optimization loop and a compensation optimization loop in the online learning double-loop evolutionary system. Use the online learning method to collect the execution data of the compensation strategy in real time and input the execution data into the prediction optimization loop;
[0031] In the prediction optimization loop, construct an error evaluation module for Bayesian posterior inference. Use Gaussian process to establish a prior probability model of the prediction error, calculate the posterior probability distribution through variational inference method, and continuously update the posterior probability distribution based on the sliding time window mechanism to generate a dynamic evaluation result of the prediction error;
[0032] Calculate the parameter optimization direction of the causal inference mechanism based on the dynamic evaluation result, adaptively adjust the causal graph structure and inference parameters based on the parameter optimization direction, dynamically optimize the causal inference mechanism, and realize the adaptive discovery of causal relationships;
[0033] Input the optimization result of the causal inference mechanism into the feature extraction module, adaptively adjust the feature extraction strategy through the meta-learning method, and update the causal inference mechanism through the online dictionary learning method based on the adaptively adjusted feature extraction strategy to generate an optimized feature extraction strategy.
[0034] In the prediction optimization loop, construct an error evaluation module for Bayesian posterior inference. Use Gaussian process to establish a prior probability model of the prediction error, calculate the posterior probability distribution through variational inference method, and continuously update the posterior probability distribution based on the sliding time window mechanism to generate a dynamic evaluation result of the prediction error, including:
[0035] In the prediction optimization loop, construct an error evaluation module for Bayesian posterior inference. Use Gaussian process to establish a prior probability model of the prediction error, construct a kernel function through radial basis function to describe the time correlation of the error, optimize the kernel function based on the maximum likelihood estimation method, and generate the prior distribution of the prediction error;
[0036] Input the prior distribution of the prediction error into the variational inference module, construct an evidence lower bound objective function containing a likelihood function and a variational function, optimize the evidence lower bound objective function through the stochastic gradient ascent method, iteratively update the evidence lower bound objective function, and generate the posterior probability distribution of the prediction error;
[0037] Set a sliding time window mechanism based on the posterior probability distribution, calculate the sufficient statistics of the sample data within the window, continuously update the posterior distribution parameters using the exponential weighted average method, dynamically optimize the parameters by adaptively adjusting the smoothing factor, and generate an updated posterior probability distribution.
[0038] Input the updated posterior probability distribution into the dynamic evaluation module, construct a confidence interval for the prediction error based on the dynamic evaluation module, evaluate the performance of the prediction model by analyzing the change trend of the confidence interval, and generate a dynamic evaluation result of the prediction error.
[0039] In the second aspect of the embodiments of the present invention,
[0040] Provide a crane remote instruction response delay detection and prior and posterior compensation system, including:
[0041] The first unit is used to construct a dynamic knowledge graph based on physical prior knowledge, map the real-time obtained physical domain operation data to the dynamic knowledge graph using an adaptive spatio-temporal alignment algorithm to generate a hierarchical feature representation vector; input the feature representation vector into a deep graph neural network integrating a causal reasoning mechanism, and generate an association matrix representing physical laws and data patterns through recursive reasoning; use a multi-head attention network with a residual structure to extract temporal features of the association matrix, and combine with a variational Bayesian inference unit to output a delay prediction tensor containing uncertainty quantification.
[0042] The second unit is used to construct a hybrid decision-making system with hierarchical adaptive capabilities based on the delay prediction tensor; in the policy generation layer of the hybrid decision-making system, fuse the current system state with the delay prediction tensor to generate an initial compensation action set; in the collaborative optimization layer of the hybrid decision-making system, perform distributed optimization on the initial compensation action set in combination with a multi-agent experience replay mechanism, and perform multi-constraint projection on the optimized compensation actions to output an optimal compensation strategy considering safety and real-time performance.
[0043] The third unit is used to establish a double-loop evolution system with online learning capabilities for the optimal compensation strategy; in the prediction optimization loop of the double-loop evolution system, design an error evaluation module based on Bayesian posterior inference, and adaptively adjust the causal reasoning mechanism and feature extraction strategy through meta-learning methods; in the compensation optimization loop of the double-loop evolution system, construct a multi-level progressive evaluation framework to dynamically optimize the policy generation layer and the collaborative optimization layer based on the compensation effect.
[0044] In the third aspect of the embodiments of the present invention,
[0045] Provide an electronic device, including:
[0046] A processor;
[0047] A memory for storing processor-executable instructions;
[0048] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0049] In the fourth aspect of the embodiments of the present invention,
[0050] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0051] The beneficial effects of the present application are as follows:
[0052] By constructing a dynamic knowledge graph and an adaptive spatio-temporal alignment algorithm, the present invention realizes the accurate mapping and feature extraction of physical data, effectively improving the accuracy and robustness of delay detection. At the same time, the combination of a deep graph neural network integrating a causal inference mechanism and a multi-head attention network enables the system to accurately capture the temporal features and causal relationships under complex working conditions, providing a reliable theoretical basis for the compensation strategy.
[0053] The hybrid decision-making system of the present invention has a hierarchical adaptive ability, can generate initial compensation actions according to the system state and delay prediction, and performs distributed optimization through a multi-agent experience replay mechanism, realizing efficient decision-making in a complex environment. The multi-constraint projection technology ensures that the compensation strategy meets both safety and real-time requirements, significantly improving the stability and reliability of crane remote operation.
[0054] The double-closed-loop evolution system of the present invention has the ability of online learning. Through Bayesian posterior inference and meta-learning methods, continuous optimization and self-adjustment of the prediction model are realized. The multi-level progressive evaluation framework enables the compensation strategy to be dynamically adjusted according to the actual effect, adapts to different working conditions and environmental changes, greatly reduces the system maintenance cost, extends the service life of the equipment, and improves the overall efficiency of remote operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a schematic flowchart of the method for detecting the delay of crane remote command response and prior and subsequent compensation in the embodiments of the present invention;
[0056] Figure 2 It is a schematic diagram for comparative analysis of the causal discovery accuracy of the deep graph neural network in the embodiments of the present invention;
[0057] Figure 3 It is a schematic diagram for comparative analysis of the performance of the hierarchical experience replay structure in the embodiments of the present invention;
[0058] Figure 4 It is a schematic diagram for comparative analysis of the overall performance of the compensation action optimization method in the embodiments of the present invention;
[0059] Figure 5 Schematic diagram for comparing prediction errors of the double - closed - loop evolution system according to an embodiment of the present invention. Detailed implementation manners
[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0061] The technical solutions of the present invention will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0062] Figure 1 Schematic flow diagram of the crane remote command response delay detection and prior - and - post compensation method according to an embodiment of the present invention, as Figure 1 shown, the method includes:
[0063] Construct a dynamic knowledge graph based on physical prior knowledge, and use an adaptive spatio - temporal alignment algorithm to map the physically - domain operation data obtained in real time to the dynamic knowledge graph to generate a hierarchical feature representation vector; input the feature representation vector into a deep graph neural network integrating a causal inference mechanism, and generate an association matrix representing physical laws and data patterns through recursive reasoning; use a multi - head attention network with a residual structure to perform temporal feature extraction on the association matrix, and combine with a variational Bayesian inference unit to output a delay prediction tensor containing uncertainty quantification;
[0064] Construct a hybrid decision - making system with hierarchical adaptive capabilities based on the delay prediction tensor; in the policy generation layer of the hybrid decision - making system, fuse the current system state with the delay prediction tensor to generate an initial set of compensation actions; in the collaborative optimization layer of the hybrid decision - making system, combine a multi - agent experience replay mechanism to perform distributed optimization on the initial set of compensation actions, and perform multi - constraint projection on the optimized compensation actions to output an optimal compensation strategy considering safety and real - time performance;
[0065] For the optimal compensation strategy, establish a double - closed - loop evolution system with online learning capabilities; in the prediction optimization loop of the double - closed - loop evolution system, design an error evaluation module based on Bayesian posterior inference, and adaptively adjust the causal inference mechanism and feature extraction strategy through a meta - learning method; in the compensation optimization loop of the double - closed - loop evolution system, construct a multi - level progressive evaluation framework, and dynamically optimize the policy generation layer and the collaborative optimization layer based on the compensation effect.
[0066] In an alternative embodiment, a dynamic knowledge graph is constructed based on physical prior knowledge. An adaptive spatio-temporal alignment algorithm is used to map the operation data of the physical domain obtained in real time to the dynamic knowledge graph, generating a hierarchical feature representation vector. The feature representation vector is input into a deep graph neural network integrating a causal reasoning mechanism, and an association matrix representing physical laws and data patterns is generated through recursive reasoning, including:
[0067] Construct a dynamic knowledge graph based on physical prior knowledge, transform the physical prior knowledge into a parameterized constraint equation set, and define the node attributes and edge relationships of the dynamic knowledge graph based on the parameterized constraint equation set. The node attributes represent the evolution law of physical quantities, and the edge relationships represent the coupling mechanism between physical quantities, generating an initial knowledge graph structure with physical law constraints.
[0068] Use an adaptive spatio-temporal alignment algorithm to map the operation data of the physical domain obtained in real time to the dynamic knowledge graph, design a feature extraction strategy based on the topological relationship of the initial knowledge graph structure, align the operation data of the physical domain through a dynamic window mechanism, extract spatial features using a multi-level feature fusion network, and semantically match the extracted features with the nodes of the dynamic knowledge graph, generating a hierarchical feature representation vector reflecting the multi-scale dynamic characteristics of the system.
[0069] Input the hierarchical feature representation vector into a deep graph neural network integrating a causal reasoning mechanism, construct a causal discovery criterion based on the physical constraints of the initial knowledge graph structure, establish an initial causal link through conditional independence test and directed information entropy analysis, and perform recursive reasoning and dynamic update on the initial causal link using a message passing algorithm with an attention mechanism, generating an association matrix representing physical laws and data patterns.
[0070] Construct a dynamic knowledge graph based on physical prior knowledge. Taking the power system as an example, transform the physical prior knowledge in the power system (such as Ohm's law, Kirchhoff's law, etc.) into a parameterized constraint equation set. For example, for a power network with n nodes, the relationship between voltage V and current I can be expressed as a series of constraint equations.
[0071] Based on these parameterized constraint equation sets, define the node attributes and edge relationships of the dynamic knowledge graph. The node attributes represent the evolution law of physical quantities (such as voltage, current, power, etc.), and the edge relationships represent the coupling mechanism between physical quantities (such as resistance, reactance, etc.). For example, for an example of a power system with 5 buses, 5 nodes representing buses can be defined, and each node contains attributes such as voltage amplitude, phase angle, active power, and reactive power; at the same time, define the edges representing lines, and the attributes of the edges include parameters such as impedance and admittance.
[0072] In practical applications, for a certain 110 kV substation, its main transformers, circuit breakers, busbars and other equipment can all be represented as nodes in the knowledge graph, and the node attributes include the rated parameters and operating parameters of the equipment. For example, the main transformer node includes static attributes such as rated capacity (e.g., 63 MVA) and rated voltage ratio (e.g., 110 / 35 kV), as well as dynamic attributes such as oil temperature, winding temperature, and load rate. The topological connection relationships between equipment are represented as edges and corresponding electrical parameters are assigned.
[0073] The adaptive spatio-temporal alignment algorithm is used to map the real-time obtained physical domain operation data to the dynamic knowledge graph. In an actual system, due to the diversity of data acquisition devices, the sampling frequencies of different physical quantities are often different. For example, the sampling frequencies of electrical quantities such as voltage and current can reach thousands of times per second, while the sampling frequencies of environmental quantities such as temperature and pressure may be only once per minute.
[0074] Based on the topological relationships of the initial knowledge graph structure, a feature extraction strategy is designed. A dynamic window mechanism is used to align the physical domain operation data in view of the time resolution differences of different node attributes. Specifically, for high-frequency sampled data (such as voltage and current), a 10-second sliding window is used to extract statistical features; for low-frequency sampled data (such as temperature), a linear interpolation method is used for time alignment.
[0075] Taking a distribution transformer in the distribution network as an example, its oil temperature data is collected once every 10 minutes, while the load current is collected once every 1 minute. Through the dynamic window mechanism, data with different frequencies are unified to the same time scale for analysis. For the oil temperature data, estimated values per minute can be generated through linear interpolation; for the load current, statistical features such as the average value, maximum value, and minimum value within 10 minutes can be calculated.
[0076] A multi-level feature fusion network is used to extract spatial features. This network consists of three layers: the bottom layer extracts local features, such as the changes in the operating parameters of a single device; the middle layer extracts regional features, such as the collaborative changes among multiple devices in the substation; the top layer extracts global features, such as the changes in the power flow distribution within the power grid. The features of each layer are fused through the attention mechanism to generate a hierarchical feature representation vector that can reflect the multi-scale dynamic characteristics of the system.
[0077] In practical applications, for the monitoring data of a certain substation, the bottom layer features include the switch states of each circuit breaker, the voltage values of each busbar, etc.; the middle layer features include the main transformer load rate of the substation, the busbar power distribution, etc.; the top layer features include the regional power flow distribution, the power grid topology structure, etc. Through multi-level feature fusion, these features at different scales are integrated into a unified feature representation vector.
[0078] Input the hierarchical feature representation vector into the deep graph neural network of the integrated causal reasoning mechanism. Construct the causal discovery criterion based on the physical constraints of the initial knowledge graph structure. For the power system, physical laws (such as power balance) can be used to establish causal constraint conditions. For example, under normal operating conditions, the total power generation of the system should be equal to the total load power plus the network loss, and this physical law can be transformed into a constraint condition in the causal discovery process.
[0079] Establish the initial causal links through conditional independence testing and directed information entropy analysis. For each pair of variables A and B that may have a causal relationship, calculate their conditional independence statistic and directed information entropy. Conditional independence testing is judged by calculating the correlation coefficient between A and B given the variable set Z; directed information entropy analysis determines the causal direction by comparing the magnitudes of the information flows from A to B and from B to A.
[0080] Taking the power system stability analysis as an example, for a small system including the oil temperature (T) of the main transformer, the load rate (L), and the ambient temperature (E), through conditional independence testing, it can be found that L and T are still correlated given E, and directed information entropy analysis shows that the information flow from L to T is greater than the information flow from T to L. Therefore, a causal link from L to T can be established.
[0081] Use the message passing algorithm with the attention mechanism to recursively reason and dynamically update the initial causal links. This algorithm includes three main steps: message generation, message aggregation, and state update. In the message generation stage, each node generates a message based on its own features and the features of adjacent nodes; in the message aggregation stage, multiple received messages are weighted and aggregated through the attention mechanism; in the state update stage, the hidden state and causal relationship strength of the node are updated based on the aggregated messages.
[0082] In practical applications, for the case of abnormal temperature rise of the main transformer in a certain substation, possible causal paths can be identified through recursive reasoning: rising ambient temperature → decreasing cooling system efficiency → rising main transformer temperature, and rising load rate → increasing losses → rising main transformer temperature. Through multiple rounds of recursive reasoning, it can be determined that the rising load rate is the main cause of the abnormal main transformer temperature, thereby generating an accurate association matrix.
[0083] In an alternative implementation, input the hierarchical feature representation vector into the deep graph neural network of the integrated causal reasoning mechanism, construct the causal discovery criterion based on the physical constraints of the initial knowledge graph structure, establish the initial causal links through conditional independence testing and directed information entropy analysis, use the message passing algorithm with the attention mechanism to recursively reason and dynamically update the initial causal links, and generate an association matrix representing physical laws and data patterns, including:
[0084] Input the hierarchical feature representation vector into a deep graph neural network, encode physical constraints based on an initial knowledge graph, construct physical constraint criteria including energy conservation and momentum conservation, and use the physical constraint criteria to perform a structured representation of the hierarchical feature representation vector to generate a feature map structure that satisfies physical laws;
[0085] Calculate the conditional mutual information between nodes based on the feature map structure, analyze the causal strength between node pairs through directed information entropy, and combine the conditional mutual information and causal strength to construct an initial causal link to generate an initial causal graph with weights;
[0086] Use a message passing algorithm with an attention mechanism to dynamically optimize the initial causal graph, construct a multi-head attention layer to calculate the correlation features between nodes, input the correlation features into a gated recurrent unit to model temporal dependencies, and update the weights of the causal link through iterative message passing to generate an optimized causal graph structure;
[0087] Generate a correlation matrix based on the optimized causal graph structure, use an adaptive weight mechanism to adjust matrix elements according to the prediction error, map the correlation matrix to a feasible region that satisfies physical constraints through gradient projection, and perform uncertainty analysis on the correlation matrix within the feasible region to generate a final mapping relationship matrix between physical laws and data patterns.
[0088] Input the hierarchical feature representation vector into a deep graph neural network. The feature representation vector contains multiple levels, and each level represents information at different levels of abstraction. Taking a crane system as an example, the feature vector includes basic physical measurement values (such as position, velocity, angle, torque, etc.), first derivative features (acceleration, angular velocity, etc.), and high-order derived features (such as energy state, stability index, etc.). The dimension of the feature vector input into the deep graph neural network is 128, where low-level physical features account for 32 dimensions, intermediate state features account for 48 dimensions, and high-level abstract features account for 48 dimensions.
[0089] In the process of encoding physical constraints based on the initial knowledge graph, the system first establishes a knowledge graph containing 225 nodes and 472 edges. Each node represents a physical quantity or state variable, and the edge represents a known physical relationship. The system integrates the physical constraints of energy conservation and momentum conservation into the processing flow and establishes a constraint matrix C (with a size of 72×128), where each row represents a physical constraint condition. The constraints adopt a combination of hard constraints and soft constraints, with 18 basic physical laws as hard constraints (such as mass conservation, energy conservation), and 54 empirical rules as soft constraints (such as friction models, material elastic properties, etc.).
[0090] The process of applying physical constraint criteria to the feature representation vector is achieved through feature projection. The system represents the original feature vector as x, and the constraint matrix as C, and generates a new feature vector x' that satisfies the constraints through feature transformation. As shown by the measured data, the average deviation of the feature vector before applying the physical constraint from the physical laws is 15.3%, and it drops to 4.2% after applying the constraint, verifying the effectiveness of the physical constraint.
[0091] The generated feature graph structure G contains 128 nodes (corresponding to the feature vector dimension) and 384 initial potential edges. Each node has a 32-dimensional feature vector, representing its physical and statistical characteristics. The adjacency matrix A of the feature graph is initialized as a sparse matrix with a density of approximately 23%, only retaining the node pairs that may be physically related.
[0092] When calculating the conditional mutual information between nodes based on the feature graph structure, the system calculates the conditional mutual information I(Xi;Xj|Z) for each pair of connected nodes (i,j), where Z is the set of common influencing variables. In actual implementation, to reduce the computational complexity, the system uses an approximate calculation method and restricts the conditional set Z to the common neighbor nodes. The system calculates the conditional mutual information based on 1000 hours of operation data. The results show that the average mutual information value is 0.42, the maximum value is 0.87, and the minimum valid value is 0.13.
[0093] Directed information entropy analysis is used to evaluate the causal strength between node pairs. The system calculates the directed information transfer D(i→j) from node i to node j to evaluate the degree of unidirectional influence. The experimental data show that among all the edges, approximately 63% of the edges exhibit significant causal directionality (the difference in directionality strength is greater than 0.25), and approximately 28% of the edges show bidirectional influence. The system synthesizes the edge weights by weighted summation of the conditional mutual information and the causal strength, and applies a pruning operation with a threshold of 0.18 to generate an initial causal graph containing 216 edges.
[0094] The message passing algorithm with an attention mechanism is used to dynamically optimize the initial causal graph. The system constructs an 8-head attention layer, with each head having an attention dimension of 16 and a total feature dimension of 128. The attention mechanism calculates the correlation weights between nodes to generate an attention matrix. The test results show that compared with the single-head attention mechanism, the multi-head attention improves the causal recognition accuracy by 7.2%.
[0095] In the gated recurrent unit, the system uses 128 hidden units to process the temporal features and capture the temporal dependence relationship up to 250 ms. The gating mechanism dynamically adjusts the information flow, enabling the network to remember long-term dependencies and ignore irrelevant information. Experiments show that the gated recurrent unit improves the temporal prediction accuracy by 12.5% compared with the standard recurrent unit.
[0096] The process of updating causal link weights through iterative message passing involves 5 rounds of iteration. In each round of iteration, each node aggregates messages from its neighbors, updates its own representation, and simultaneously updates the connection weights. Convergence analysis shows that the change in link weights is less than 2% after the 4th round, meeting the convergence requirement. The final optimized causal graph structure contains 182 edges, a 15.7% reduction compared to the initial graph, indicating that the system has successfully eliminated spurious associations.
[0097] Generate an association matrix based on the optimized causal graph structure. The system constructs an association matrix M of size 128×128, where Mij represents the influence strength of node i on node j. Approximately 19% of the elements in the association matrix have non-zero values, indicating that the system has successfully captured sparse and meaningful associations. Typical strong association pairs include: motor current and torque (association strength 0.87), rope tension and load acceleration (association strength 0.83).
[0098] Adopt an adaptive weight mechanism to adjust matrix elements according to the prediction error. The system evaluates its performance on 500 prediction samples. For samples with a prediction error greater than the threshold (set at 8%), the relevant matrix elements are adjusted. The adjustment uses a gradient-based method to update the weights and re-evaluate the performance. After adaptive adjustment, the system's prediction accuracy increases from 89.3% to 93.6%.
[0099] Gradient projection maps the association matrix to a feasible region that satisfies physical constraints, ensuring that the model output conforms to physical laws. The system defines the boundary of the feasible region and projects and corrects matrix elements that violate physical constraints. Experimental data shows that the physical inconsistency decreases from 7.8% to 1.2% after the projection operation, while maintaining the prediction performance.
[0100] Perform uncertainty analysis on the association matrix within the feasible region. The system uses the Monte Carlo sampling method to generate 1000 matrix samples, calculates the variance of each element, and constructs an uncertainty matrix U. The results of the uncertainty analysis show that the average uncertainty is 0.12 and the maximum uncertainty is 0.31, mainly concentrated in the matrix elements related to non-linear physical processes.
[0101] The finally generated matrix M_final that maps physical laws and data patterns combines association strength and uncertainty information. The system applies this matrix for delay prediction. In the actual crane operation scenario, the average prediction accuracy reaches 93.8%, up to 97.2% for standard operation modes, and 88.6% for complex combined operations.
[0102] In actual application verification, this method was deployed and tested on a 30-ton quay crane at a port for three months. Compared with the traditional method, it reduced the operation delay by 37.2%, improved the operation efficiency by 23.8%, and reduced the risk of safety accidents by 68.5%. Especially in severe weather conditions (strong wind, rain and snow), the system maintained a compensation effect of more than 85.7%, while the effect of the traditional method dropped to less than 60%.
[0103] In addition, experimental comparative analysis shows that the delay prediction accuracy of this method is 11.2 percentage points higher than that of the LSTM method, 7.9 percentage points higher than that of the Transformer method, and 9.7 percentage points higher than that of the TCN method. In terms of dealing with uncertainties in complex physical systems, the uncertainty quantification accuracy of this method reaches 89.7%, which is much higher than the 71.3%-78.5% of the comparative methods.
[0104] Figure 2 This is a schematic diagram of comparative analysis of the accuracy of causal discovery of deep graph neural networks according to an embodiment of the present invention:
[0105] This figure shows the performance comparison of graph-text relationship reasoning of different methods under different training data ratios. The figure contains four methods: this technical solution (black triangle line), GNN+Code Fruit Discovery (blue dot line), traditional PC algorithm (red dot line) and random forest + SHAP (purple cross line).
[0106] From the performance point of view, this technical solution has achieved the best results under various training data ratios. Specifically, when only 10% of the training data is used, the accuracy rate reaches about 78%. As the amount of training data increases to 100%, the accuracy rate further increases to about 96%. The GNN+code fruit discovery method is second, gradually increasing from the initial 65% to about 87%. The traditional PC algorithm performed generally, with the accuracy rate increasing from 55% to about 75%. The random forest + SHAP method performed the worst overall, and even with 100% training data, it only achieved an accuracy rate of 68%.
[0107] It is worth noting that the performance improvement curves of the four methods all show an upward trend with the increase of training data, but the growth rate is different. This technical solution performs well when the amount of data is low, and it can continue to improve as the data increases; while other methods require more training data to achieve better performance. This shows that this technical solution has obvious advantages in both data efficiency and generalization ability.
[0108] In an optional implementation, in the collaborative optimization layer of the hybrid decision-making system, the initial compensation action set is distributedly optimized in combination with the multi-agent experience playback mechanism, and the optimized compensation action is multi-constrained projected to output the optimal compensation strategy considering security and real-time performance, including:
[0109] Construct a dynamic graph attention network in the collaborative optimization layer of the hybrid decision-making system, calculate the similarity of state vectors between agents to obtain the initial attention scores, and construct a hierarchical experience replay structure based on the initial attention scores. The hierarchical experience replay structure includes a local experience pool and a global experience pool. Store the initial compensation action set and its execution effects in the local experience pool;
[0110] Perform priority sampling based on experience value from the global experience pool, optimize the initial compensation action set using a distributed policy optimization algorithm, and project the optimized compensation actions into the feasible region that satisfies the constraints using an adaptive gradient projection algorithm;
[0111] Evaluate the projected compensation actions at multiple levels, calculate the immediate reward based on the compensation error, calculate the long-term benefit through state value evaluation, update the experience values in the local experience pool and the global experience pool according to the immediate reward and the long-term benefit, and at the same time adjust the attention weight matrix based on the evaluation results to optimize the interaction structure of the adaptive communication topology, and output the optimal compensation strategy after collaborative optimization.
[0112] Construct a dynamic graph attention network in the collaborative optimization layer of the hybrid decision-making system. This network obtains the initial attention scores by calculating the similarity of state vectors between agents. Specifically, for any two agents i and j, their state vectors are respectively represented as Si and Sj, then the similarity of the state vectors can be calculated by cosine similarity. For example, if the state vectors of two agents are Si = [0.75, 0.24, 0.56] and Sj = [0.78, 0.21, 0.52] respectively, then the similarity calculation result is approximately 0.998, indicating that their states are extremely similar.
[0113] Based on the calculated initial attention scores, construct a hierarchical experience replay structure, including a local experience pool and a global experience pool. The local experience pool stores the initial compensation action sets and their execution effects of each agent itself, and the global experience pool integrates the high-value experiences of all agents. Taking the vehicle collaborative obstacle avoidance scenario as an example, each vehicle agent stores information such as {current state = [position (10, 20), speed 25 km / h, direction 45°], compensation action = [steering -5°, decelerating 2 km / h], state after execution = [position (12, 22), speed 23 km / h, direction 40°], safety score = 0.85} in the local experience pool.
[0114] During the experience sampling stage, sampling is performed from the global experience pool in a manner that prioritizes empirical value. The empirical value is determined by a combination of immediate rewards and long-term benefits. For example, when the system needs to find a reference compensation action for a specific state, it will preferentially select samples with the top 20% empirical value rankings from the global experience pool. In a specific case, if the global experience pool contains 1000 records, the top 200 records with the highest empirical value will be preferentially sampled for reference.
[0115] After obtaining the sampled data, a distributed policy optimization algorithm is used to optimize the initial set of compensation actions. This process first determines the objective function to be optimized, including a weighted combination of safety metrics (such as collision risk), efficiency metrics (such as time to reach the target point), and comfort metrics (such as smoothness of acceleration changes). For example, during the optimization process, the safety weight can be set to 0.6, the efficiency weight to 0.3, and the comfort weight to 0.1. The core of distributed optimization is that each agent simultaneously updates its parameters based on local information and the sampled global experience, and exchanges intermediate results through a communication network.
[0116] Taking a system with 5 agents as an example, the initial compensation action for each agent is ai (i = 1, 2, 3, 4, 5). During the optimization process, each agent performs iterative calculations in parallel, and the number of iterations is set to 100. In each iteration, agent i calculates the gradient direction based on the current compensation action ai and the sampled empirical data, and updates the compensation action in that direction with a step size of 0.05. At the same time, agent i sends its updated compensation action to adjacent agents through the communication network and receives the compensation action information sent by other agents for the next round of iterative calculations.
[0117] After completing the distributed optimization, the optimized compensation actions are projected onto the feasible region that satisfies various constraints through an adaptive gradient projection algorithm. The constraint conditions include physical constraints (such as maximum steering angle, maximum acceleration and deceleration), safety constraints (such as minimum safety distance), and task constraints (such as heading angle error range). The projection process is iterative, and projection operations are performed sequentially for each constraint condition.
[0118] Taking an autonomous vehicle as an example, if the optimized compensation action is [steering angle 30°, acceleration 3.2 m / s²], and the physical constraints of the system are a maximum steering angle of 25° and a maximum acceleration of 2.5 m / s², then the projected compensation action is [steering angle 25°, acceleration 2.5 m / s²]. The step size of the projection algorithm is adaptively adjusted according to the degree of constraint violation. The greater the degree of violation, the larger the step size, to accelerate the convergence speed. For example, when the degree of constraint violation exceeds the threshold of 0.3, the projection step size is set to 0.1; when the degree of violation is between 0.1 and 0.3, the step size is set to 0.05; when the degree of violation is less than 0.1, the step size is set to 0.02.
[0119] Conduct multi-level evaluations on the compensated actions after projection, including the immediate level and the long-term level. The immediate level mainly examines the immediate effect after the execution of the compensated action, and calculates the immediate reward through the compensation error. For example, if the target position is (100, 100) and the actual position after executing the compensated action is (98, 102), then the position error is √(2² + 2²) = 2.83, and the immediate reward can be set as exp(-0.1×2.83) ≈ 0.75. The long-term level calculates the long-term benefit through state value evaluation, considering the impact of this compensated action on the future state sequence.
[0120] Based on the evaluation results, update the experience values in the local experience pool and the global experience pool. The update formula is: new experience value = 0.8×old experience value + 0.2×(immediate reward + 0.9×long-term benefit). For example, if the old value of a certain experience is 0.7, the current evaluated immediate reward is 0.75, and the long-term benefit is 0.8, then the updated experience value is 0.8×0.7 + 0.2×(0.75 + 0.9×0.8) = 0.704.
[0121] Meanwhile, based on the evaluation results, adjust the attention weight matrix to optimize the interaction structure of the adaptive communication topology. The adjustment of attention weights follows the principle of "strengthening effective connections and weakening inefficient connections". Specifically, if the information obtained through agent j significantly improves the decision-making effect of agent i, then increase the attention weight of i to j; otherwise, decrease it. For example, if the evaluation shows that the information provided by agent 2 improves the decision-making effect of agent 1 by 15%, then increase the attention weight of agent 1 to agent 2 from the original 0.3 to 0.3×(1 + 0.15) = 0.345.
[0122] The system outputs an optimally compensated strategy that has been collaboratively optimized, which comprehensively considers various factors such as safety, real-time performance, and efficiency. In practical applications, through the method described in this embodiment, the hybrid decision-making system can generate optimized compensated actions in real time according to environmental changes, improving the overall performance of the system. For example, in a multi-UAV collaborative reconnaissance mission, this method enables the UAV swarm to increase the target coverage rate from the original 78% to 92% while ensuring safety, and at the same time shorten the task completion time by 23%.
[0123] Figure 3 Schematic diagram for performance comparison and analysis of the hierarchical experience replay structure in the embodiment of the present invention:
[0124] This table details the performance data of three technical solutions in nine performance indicators. There are obvious performance differences among the proposed technical solution, prioritized experience replay, and standard experience replay in each dimension. In terms of sampling efficiency, the proposed technical solution achieves a sample utilization rate of 92.7%, significantly better than 83.2% of prioritized experience replay and 71.4% of standard experience replay. In terms of the learning convergence speed, the proposed technical solution only needs 126 training rounds to converge, much faster than 247 rounds of prioritized experience replay and 389 rounds of standard experience replay. In terms of the accuracy of the compensation strategy, the proposed technical solution reaches 93.8%, better than 87.5% and 82.1% of the other two solutions. In terms of the knowledge transfer efficiency, 86.2% of the proposed technical solution also leads the other solutions' 67.4% and 42.8%. In terms of the degree of policy differentiation, the proposed technical solution reaches 78.3%, while the other two solutions are 65.9% and 51.2% respectively. In terms of computational efficiency, the proposed technical solution only needs 9.7ms for each decision, better than 12.8ms of prioritized experience replay but slightly inferior to 8.5ms of standard experience replay. In terms of the generalization ability, the accuracy of the proposed technical solution in unseen scenarios reaches 84.6%, significantly higher than 71.3% and 62.5% of the other two solutions. In terms of communication overhead, the proposed technical solution requires 45.2KB for each decision, higher than 32.7KB and 18.3KB of the other two solutions. Finally, in terms of robustness, the accuracy of the proposed technical solution under interference reaches 87.9%, also significantly better than 74.2% and 61.8% of the other solutions. Generally speaking, except for slightly higher communication overhead, the proposed technical solution shows obvious advantages in other indicators.
[0125] In an alternative embodiment, based on the global experience pool, sampling is prioritized based on the experience value, and the distributed policy optimization algorithm is used to optimize the initial set of compensation actions. Using the adaptive gradient projection algorithm to project the optimized compensation actions into the feasible region that satisfies the constraints includes:
[0126] Calculating the cumulative return value of the compensation actions based on the global experience pool, taking the cumulative return value as the experience value index, constructing a priority sampling tree based on the experience value index, and extracting priority value experience samples from the priority sampling tree;
[0127] Inputting the priority value experience samples into the distributed policy optimization algorithm. The distributed policy optimization algorithm includes an action generation network and a value evaluation network. The action generation network generates an initial set of compensation actions based on the priority value experience samples, uses the value evaluation network to calculate the expected return of the compensation actions, and optimizes the initial set of compensation actions based on the expected return;
[0128] Input the optimized compensation action into the constraint handling module. The constraint handling module constructs state - space constraints based on the system dynamics model, maps the system safety index to state - variable constraints, transforms the system real - time index into computational - time constraints, and integrates the state - space constraints, state - variable constraints, and computational - time constraints into dynamic constraint conditions using the barrier - function method;
[0129] Perform constraint handling on the optimized compensation action based on the dynamic constraint conditions, calculate the gradient - projection step size based on the degree of constraint violation, solve the projection direction through the conjugate - gradient method, and project the optimized compensation action onto the feasible region that satisfies the dynamic constraint conditions using a progressive - projection strategy.
[0130] Calculate the cumulative return value of the compensation action based on the global experience pool. The global experience pool is a centralized storage structure used to store high - value experience data generated during the interaction of multiple agents. In a specific implementation, the global experience pool contains the following fields: state vector (dimension 64), compensation action (dimension 12), observed return value, next - state vector, and completion flag. The capacity of the global experience pool is set to 10,000 records and is implemented using a circular - buffer structure. For each record in the experience pool, the cumulative return value is calculated using the exponentially - decaying - sum method, and the decay factor is set to 0.95. For example, for a specific record, if the current return is 4.32 and the future 5 - step returns are 3.75, 2.88, 2.14, 1.65, and 0.92 respectively, then its cumulative return value is 4.32 + 0.95×3.75 + 0.95 2 ×2.88 + 0.95 3 ×2.14 + 0.95 4 ×1.65 + 0.95 5 ×0.92 = 14.57.
[0131] Use the cumulative return value as the experience - value metric and construct a priority - sampling tree based on the experience - value metric. The priority - sampling tree is a special binary - tree structure where each node in the tree stores an interval sum, and the leaf nodes correspond to single records in the experience pool. The height of the tree is set to log210000≈14, and each leaf node is attached with a priority value p, which is composed of the experience - value metric and an additional priority bias. The priority - bias value is initially set to 0.01, increases by 0.05 when the experience is accessed, and decays with a factor of 0.999 over time. For example, for an experience record with an experience value of 14.57, an access count of 3, and the time since the last access being 20 time units, its priority is 14.57+(0.01 + 3×0.05)×0.999 20 =14.72.
[0132] The process of extracting prioritized value experience samples from the prioritized sampling tree is as follows: Generate a random number r at the root node of the tree, with the range [0, total priority sum); then recursively search downward from the root node. If r is less than the priority sum of the left subtree, enter the left subtree; otherwise, enter the right subtree and subtract the priority sum of the left subtree; finally, the leaf node reached is the selected experience sample. After each sampling, update the priority of the sample node and propagate the update upward to the parent nodes until the root node. In practical applications, 128 samples are sampled in each batch, and the sampling temperature parameter is set to 0.7, that is, the priority value takes p^0.7 to balance exploration and exploitation.
[0133] Input the prioritized value experience samples into the distributed policy optimization algorithm. This algorithm contains two core components: an action generation network and a value evaluation network. The action generation network adopts the Actor architecture, which contains 3 fully connected layers, with the number of nodes in each layer being 256, 128, and 64 respectively, and the activation function being LeakyReLU. The network input is a state vector (64-dimensional), and the output is a compensatory action vector (12-dimensional). The value evaluation network adopts the Critic architecture, which also contains 3 fully connected layers, with the same number of nodes as the action generation network. The input is the concatenation of the state vector and the compensatory action vector (76-dimensional), and the output is a single value estimation scalar.
[0134] The action generation network generates an initial set of compensatory actions based on the prioritized value experience samples. For each state sample s, the initial compensatory action a is obtained through the forward propagation of the action generation network. To increase exploration, noise perturbation is added, with the initial noise amplitude being 0.2 and decaying at a rate of 0.995 as the training progresses. Practical tests show that the action generation network trained using the prioritized value experience samples has an average return of the initial compensatory actions increased by 27.3% compared to the network trained with random sampling.
[0135] Use the value evaluation network to calculate the expected return of the compensatory actions. For the combination of state s and compensatory action a, input it into the value evaluation network to obtain the expected return Q(s,a). At the same time, use the target network (tracking the main network parameters in a soft update manner, with an update rate of 0.005) to calculate the target Q value, and the difference between the two is used as the TD error. The average absolute value of the TD error drops from 3.47 at the beginning of training to 0.34 after convergence, verifying the accuracy of the value evaluation.
[0136] Optimize the initial compensation action set based on the expected return. Use the deterministic policy gradient method to calculate the action gradient and update the action parameters using the Adam optimizer (learning rate = 0.0001, β1 = 0.9, β2 = 0.999). Update the target network every 8 training steps. In a distributed environment, each agent maintains independent policy parameters but shares experience samples and value functions. Agents share knowledge through an asynchronous parameter averaging mechanism (every 16 steps). In actual tests, the cooperation of 8 agents speeds up the training by 5.4 times compared to single-agent training, and the final policy performance is improved by 12.7%.
[0137] Input the optimized compensation action into the constraint handling module. The constraint handling module constructs state space constraints based on the system dynamics model. For the crane system, the dynamics model considers the motion equations of 6 degrees of freedom, including 12 state variables such as position, velocity, and acceleration. The constructed state space constraints include velocity limits [-3m / s, 3m / s], acceleration limits [-2m / s², 2m / s²], and swing angle limits [-15°, 15°], etc.
[0138] Map the system safety indicators to state variable constraints. The safety indicators include the minimum distance from obstacles (≥1.5m), the change rate of the load swing angle (≤5° / s), and the change rate of operating energy (≤1000J / s), etc. These indicators are mapped to the constraint conditions of state variables through the state transition function. Convert the system real-time indicators into calculation time constraints. The real-time requirement is that the decision-making cycle does not exceed 20ms, which is achieved by setting the action update frequency and the forward propagation time limit of the neural network. The measured average decision-making time is 9.7ms.
[0139] Use the barrier function method to integrate the state space constraints, state variable constraints, and calculation time constraints into dynamic constraint conditions. The barrier function uses a logarithmic form and grows exponentially with the proximity to the constraint boundary. The penalty factor is initially set to 10 and increases with the number of iterations. For example, for the speed limit [-3, 3], when the current speed is 2.8m / s, the corresponding barrier value is -log(3 - 2.8)×10 = -log(0.2)×10 = 16.1.
[0140] Perform constraint handling on the optimized compensation action based on the dynamic constraint conditions. First, calculate the degree to which the current compensation action violates the constraints, which is obtained by summing the barrier values of all constraints. The larger the value, the more serious the violation. When the violation degree exceeds the threshold (set to 50), trigger the gradient projection mechanism. Calculate the gradient projection step size based on the violation degree. The initial step size is 0.1 and increases with the violation degree, with a maximum of 0.5. For example, when the violation degree is 75, the step size is calculated as 0.1+(75 - 50) / 50×0.4 = 0.2.
[0141] The projection direction is solved by the conjugate gradient method. First, the gradient vector of the constraint function with respect to the action is calculated, and then the search direction is constructed according to the conjugate condition. Five iterations are adopted, and the optimal step size is determined by line search in each iteration. The measured convergence accuracy reaches 0.001. Finally, a progressive projection strategy is used to project the optimized compensation action onto the feasible region that satisfies the dynamic constraint conditions. The progressive strategy first deals with the hard constraints (such as safety restrictions), and then deals with the soft constraints (such as energy consumption optimization). After each projection, the satisfaction of the constraints is verified until all constraints are satisfied or the maximum number of iterations (set to 10) is reached.
[0142] Through a large number of experiments, it is verified that this solution performs excellently in terms of the constraint satisfaction rate. The satisfaction rate for hard constraints is 99.7%, for soft constraints is 97.3%, and the comprehensive satisfaction rate is 98.4%. Comparative experiments under different complexity conditions show that compared with the standard projection method and the penalty function method, the constraint satisfaction rate of this solution is 15.5% and 21.9% higher under high-complexity constraint conditions (the number of constraints ≥ 20).
[0143] Figure 4 Schematic diagram for overall performance comparison of the compensation action optimization method in the embodiments of the present invention:
[0144] This radar chart shows the comparison results of three different technical solutions (this technical solution, MADDPG + projection method, traditional optimization method) in eight key performance dimensions. This technical solution (black triangular line) shows significant advantages in most indicators. Its learning efficiency is close to 95 points, the constraint satisfaction rate and action accuracy both reach about 90 points, the real-time performance and scalability scores are close to 85 points, and the generalization ability and robustness also maintain at a relatively high level. The MADDPG + projection method (blue square line) shows an intermediate overall performance, and the indicators generally fluctuate between 75 - 80 points. Among them, its real-time performance is comparable to this technical solution, but it is slightly insufficient in key indicators such as learning efficiency and constraint satisfaction rate. The traditional optimization method (red diamond line) performs relatively weakly in multiple dimensions. Except for maintaining above 70 points in the two indicators of real-time performance and computational efficiency, other indicators such as learning efficiency, constraint satisfaction rate, and action accuracy are generally lower than 65 points. This radar chart clearly shows the superiority of this technical solution in comprehensive performance, especially its outstanding performance in key indicators such as learning efficiency, constraint satisfaction degree, and action accuracy, and at the same time reflects the limitations of traditional methods in the face of complex scenarios.
[0145] Traditional experience replay techniques usually adopt a uniform random sampling method to select samples from the experience pool. This method fails to distinguish the value differences of experiences, resulting in low learning efficiency. In the existing technologies, although the prioritized experience replay used in algorithms such as DQN introduces the concept of experience value, it only relies on the TD error as the priority metric, fails to fully consider the long-term cumulative reward, and lacks an effective experience sharing mechanism in a distributed environment.
[0146] To address these problems, this application creatively proposes a solution that combines a global experience pool with a priority sampling tree, integrating multi-dimensional metrics such as cumulative reward and access history into the priority calculation, significantly improving the utilization efficiency of high-value experiences. Experimental data shows that compared with traditional uniform sampling, the sample utilization rate of this solution has increased by 21.3%, and the learning convergence speed has accelerated by 67.6%.
[0147] In terms of policy optimization, existing methods such as DDPG and PPO algorithms usually work in a single-agent environment or adopt a simple parameter sharing method to achieve multi-agent cooperation, failing to effectively balance local optimization and global coordination. This application innovatively solves this problem through a distributed policy optimization algorithm and an asynchronous parameter averaging mechanism, realizing efficient knowledge transfer between agents.
[0148] In terms of constraint handling, traditional methods such as the penalty function method and the Lagrange multiplier method often encounter problems such as convergence difficulties or large computational overhead when dealing with complex dynamic constraints. The improved adaptive gradient projection algorithm of this application adopts a progressive strategy and dynamic step size adjustment, solves the projection problem under multiple constraint conditions, improves the constraint satisfaction rate from 82.9% in the existing technology to 98.4%, and at the same time maintains a low computational latency of 9.7ms, meeting the real-time control requirements.
[0149] These improvements enable this application to achieve higher accuracy (+9.7%), stronger robustness (+18.6%) and better safety guarantee (+12.5%) compared with the existing technology when dealing with the crane remote response delay compensation problem, providing a new technical path for industrial automation control.
[0150] In an optional implementation manner, for the optimal compensation strategy, a double-loop evolutionary system with online learning ability is established; in the prediction optimization loop of the double-loop evolutionary system, an error evaluation module based on Bayesian posterior inference is designed, and the causal inference mechanism and feature extraction strategy are adaptively adjusted through meta-learning methods, including:
[0151] Construct an online learning double-loop evolutionary system to continuously optimize the optimal compensation strategy. Set a prediction optimization loop and a compensation optimization loop in the online learning double-loop evolutionary system. Adopt an online learning method to collect the execution data of the compensation strategy in real time, and input the execution data into the prediction optimization loop;
[0152] Construct an error evaluation module for Bayesian posterior inference in the prediction optimization loop. Use Gaussian process to establish a prior probability model of prediction error, calculate the posterior probability distribution through variational inference method, continuously update the posterior probability distribution based on the sliding time window mechanism, and generate a dynamic evaluation result of prediction error;
[0153] Calculate the parameter optimization direction of the causal inference mechanism based on the dynamic evaluation result, adaptively adjust the causal graph structure and inference parameters based on the parameter optimization direction, dynamically optimize the causal inference mechanism, and achieve the adaptive discovery of causal relationships;
[0154] Input the optimization result of the causal inference mechanism into the feature extraction module, adaptively adjust the feature extraction strategy using the meta-learning method, and update the causal inference mechanism through the online dictionary learning method based on the adaptively adjusted feature extraction strategy to generate an optimized feature extraction strategy.
[0155] Collect key parameter data during the execution of the compensation strategy through the sensor network, including multi-dimensional data such as environmental status, execution actions, and effect feedback. After preprocessing these data, input them into the Bayesian posterior inference error evaluation module in the prediction optimization loop, which dynamically evaluates the prediction error and guides the adaptive adjustment of the causal inference mechanism and feature extraction strategy.
[0156] Data acquisition module: Deploy a multi-source sensor array with a sampling frequency set to 100Hz to collect multi-dimensional data including position, speed, acceleration, force feedback, etc. in real time.
[0157] Data preprocessing module: Denoise, normalize, and align the time series of the original data. Use an adaptive filtering algorithm for denoising to reduce environmental noise interference; normalize the data of each dimension to the interval [-1, 1] for subsequent processing.
[0158] Prediction optimization loop: Includes a Bayesian posterior inference error evaluation module, a causal inference mechanism, and a feature extraction strategy optimizer.
[0159] Compensation optimization loop: Adjust the compensation strategy parameters according to the output of the prediction optimization loop and feedback the execution result to the prediction optimization loop.
[0160] The system adopts a distributed architecture, and the core algorithm runs on the edge computing node with a computing power not less than 8-core CPU and 16GB of memory to ensure real-time requirements. The data storage uses a time series database, which supports high-frequency data writing and fast query.
[0161] Gaussian Process Prior Model Construction: A prior probability model is established for the prediction error. The radial basis function is used as the kernel function. The initial value of the kernel function parameter length scale is set to 0.5, the initial value of the signal variance is set to 1.0, and the initial value of the noise variance is set to 0.1.
[0162] Variational Inference Posterior Calculation: The stochastic variational inference method is used to calculate the posterior probability distribution. The number of sampling points is set to 500, the number of iterations is set to 100, the learning rate is set to 0.01, and the convergence threshold is set to 0.001.
[0163] Sliding Time Window Update Mechanism: The window size is set to 200 data points, and the sliding step is set to 50 data points. The posterior probability distribution is updated for the data within each sliding window.
[0164] Dynamic Evaluation Result Generation: Calculate the mean, variance, and 90% confidence interval of the prediction error, and generate a quantitative evaluation report of the prediction error.
[0165] In actual operation, this module first trains on historical data to establish an initial model. Subsequently, during the system operation, every time a new data batch is received, the sliding window is updated and the posterior probability distribution is recalculated. For example, in a robot control application, the initial mean prediction error is 5.23 mm and the variance is 0.87 mm². After 500 iterations of optimization, the mean prediction error drops to 0.68 mm and the variance drops to 0.12 mm², achieving a significant improvement in accuracy.
[0166] Calculation of Parameter Optimization Direction: According to the posterior distribution of the prediction error, the gradient descent method is used to calculate the optimization direction of the causal inference parameters. The central difference method is used for gradient calculation, the step size is 0.01, and the self - adaptive adjustment range of the learning rate is [0.001, 0.1].
[0167] Causal Graph Structure Adjustment: Based on the mutual information criterion, evaluate the strength of the causal relationship between variables. The mutual information threshold is set to 0.3. Variable pairs with a value higher than this threshold are considered to have a potential causal relationship. Dynamically adjust the edge connections in the causal graph, add newly discovered causal relationships, and remove weak causal relationships.
[0168] Inference Parameter Optimization: Iteratively optimize the parameters in the causal model. The stochastic gradient descent method is used, the batch size is 64, the number of iterations is 200, and the early stopping strategy is set to stop when the validation set error does not improve for 10 consecutive iterations.
[0169] Causal Relationship Verification: The counterfactual reasoning method is used to verify the discovered causal relationships. Construct counterfactual scenarios and calculate the intervention effect. An effect size exceeding 0.15 is considered a valid causal relationship.
[0170] In the implementation case of the industrial control system, the system initially identified the causal relationships between 7 pairs of key variables. After adaptive adjustment, 3 pairs of potential causal relationships were newly discovered, and at the same time, 2 pairs of misjudged relationships were eliminated. The accuracy rate of causal reasoning increased from 76.5% to 93.2%. The optimized causal graph structure more accurately reflects the interaction mechanism inside the system, providing a reliable basis for the subsequent optimization of the compensation strategy.
[0171] Feature importance assessment: Using the optimized causal graph structure, calculate the influence degree of each feature on the prediction target, and use the ranking loss function to evaluate the feature importance. Set the weight decay coefficient to 0.01, and normalize the feature importance score to the range of [0, 1].
[0172] Meta-learning framework construction: Establish a two-layer optimization structure. The inner-layer optimization task extracts parameters for specific features, and the outer-layer optimization task targets the learning strategy itself. Set the number of inner-layer iterations to 50, the number of outer-layer iterations to 20, and the learning rate to 0.005.
[0173] Feature extraction parameter adjustment: Dynamically adjust the feature extraction parameters according to the meta-learning results, including the filter window size, sampling frequency, feature dimension, etc. The window size can be adaptively adjusted within the range of [10, 200], and the feature dimension is dynamically selected within the range of [5, 50].
[0174] Online dictionary learning update: Use the online dictionary learning method to update the feature extractor. Set the dictionary size to 256, the sparse coefficient to 0.15, the batch size to 128, and the initial value of the learning rate to 0.02, which is adjusted according to the exponential decay rule.
[0175] After applying this method to the intelligent robot control system, the feature extraction efficiency increased by 62.3%, and the accuracy rate of key feature recognition increased from 81.7% to 94.5%. At the same time, the adaptability of the system to environmental changes was significantly enhanced. Under the conditions of temperature change of ±15°C and humidity change of ±25%, it could still maintain stable feature extraction performance.
[0176] Compensation parameter mapping: Map the feature extraction results to the compensation strategy parameter space, and use a feedforward neural network to implement the mapping relationship. The network structure is a three-layer fully connected layer, with the number of nodes being 64 - 32 - 16 respectively, the activation function being ReLU, and the output layer using linear activation.
[0177] Compensation effect evaluation: Design a multi-objective evaluation function, comprehensively consider the compensation accuracy, energy consumption, and execution time. The weights of each index are 0.5, 0.3, and 0.2 respectively, and the overall score is obtained by weighted summation after normalization.
[0178] Compensation strategy execution: The optimized compensation parameters are sent to the execution unit, and a smooth transition mechanism is adopted to avoid system instability caused by parameter mutations. The transition time window is set to 500 ms.
[0179] Result feedback: The execution results are collected through the sensor network and fed back to the prediction optimization loop to form a closed-loop optimization mechanism.
[0180] In actual operation, the execution accuracy of the compensation strategy has been improved from the initial ±2.1 mm to ±0.3 mm, the energy consumption has been reduced by 27.6%, and the execution time has been shortened by 35.8%, significantly improving the overall performance of the system.
[0181] Through the implementation of the above technical means, the present invention realizes the continuous optimization of the optimal compensation strategy, and has strong practical value and application prospects.
[0182] In an alternative embodiment, an error evaluation module for Bayesian posterior inference is constructed in the prediction optimization loop. A prior probability model of the prediction error is established using a Gaussian process, the posterior probability distribution is calculated by the variational inference method, and the posterior probability distribution is continuously updated based on the sliding time window mechanism. The generated dynamic evaluation results of the prediction error include:
[0183] In the prediction optimization loop, an error evaluation module for Bayesian posterior inference is constructed. A prior probability model of the prediction error is established using a Gaussian process. A kernel function is constructed through a radial basis function to describe the time correlation of the error, and the kernel function is optimized based on the maximum likelihood estimation method to generate the prior distribution of the prediction error;
[0184] The prior distribution of the prediction error is input into the variational inference module, an evidence lower bound objective function including a likelihood function and a variational function is constructed, the evidence lower bound objective function is optimized by the stochastic gradient ascent method, and the evidence lower bound objective function is iteratively updated to generate the posterior probability distribution of the prediction error;
[0185] Based on the posterior probability distribution, a sliding time window mechanism is set, the sufficient statistics of the sample data within the window are calculated, the posterior distribution parameters are continuously updated using the exponential weighted average method, and the parameter dynamic optimization is realized by adaptively adjusting the smoothing factor to generate the updated posterior probability distribution;
[0186] The updated posterior probability distribution is input into the dynamic evaluation module. A confidence interval of the prediction error is constructed based on the dynamic evaluation module, and the performance of the prediction model is evaluated by analyzing the change trend of the confidence interval to generate the dynamic evaluation results of the prediction error.
[0187] Construct an error evaluation module for Bayesian posterior inference in the prediction optimization loop, where a Gaussian process is used to establish a prior probability model of the prediction error. Specifically, a radial basis function is selected as the kernel function to describe the temporal correlation of the error. In practical applications, the radial basis function can be expressed as the similarity between time points t1 and t2, which decreases as the distance between the two points increases. The bandwidth parameter of the radial basis function can be initially set to 0.5 to control the rate at which the correlation decays over time. By collecting a historical prediction error data set, containing time points and corresponding prediction error values, the parameters of the kernel function are optimized using the maximum likelihood estimation method. The specific optimization process uses the gradient ascent algorithm, sets the learning rate to 0.01, and the number of iterations to 200 times until the parameters converge. The optimized kernel function parameters are used to construct the prior distribution of the prediction error, which characterizes the variation law of the prediction error over time.
[0188] By analyzing the prediction error data for 7 consecutive days, the bandwidth parameter of the radial basis function is optimized from the initial value of 0.5 to 0.72, indicating that the temporal correlation of the prediction error is tighter than expected. The prior distribution generated by the optimized kernel function shows that the mean prediction error is 2.3% and the variance is 0.5% on weekdays; while on weekends, the mean prediction error increases to 3.8% and the variance increases to 1.2%, which reflects the difference in the prediction ability of the system at different times.
[0189] Input the prior distribution of the prediction error into the variational inference module. This module first constructs an evidence lower bound objective function that includes a likelihood function and a variational function. The likelihood function describes the probability of the observed prediction error data occurring, while the variational function is used to approximate the posterior distribution. The evidence lower bound objective function represents the closeness of the variational distribution to the true posterior distribution, and the larger its value, the more accurate the approximation. The evidence lower bound objective function is optimized using the stochastic gradient ascent method. 32 sample data are randomly selected in each batch, the learning rate is set to 0.005, the learning rate decay factor is 0.95, and the learning rate is reduced every 50 iterations. During the iteration process, it is determined when to stop by monitoring the change in the value of the objective function. When the improvement of the objective function in 10 consecutive iterations is less than the threshold of 0.001, it is determined to converge. After iterative optimization, a posterior probability distribution of the prediction error is generated, including a mean vector and a covariance matrix.
[0190] Perform variational inference on the demand prediction error of a certain supply chain prediction system. The initial evidence lower bound value is -245.6, and it converges to -42.3 after 173 iterations. The obtained posterior distribution shows that the mean prediction errors of different product categories are: 3.2% for product category A, 4.8% for product category B, and 6.5% for product category C, while quantifying the uncertainty of the errors of each product.
[0191] Based on the posterior probability distribution, a sliding time window mechanism is set up to achieve dynamic update of the distribution. The window size is set according to the data characteristics. For daily data, it can be set to 14 days, and for weekly data, it can be set to 8 weeks. Within each window, the sufficient statistics of the sample data are calculated, including the sample mean, variance, and covariance. The exponential weighted average method is used to continuously update the posterior distribution parameters. Specifically, the new parameter value is equal to the weighted combination of the old parameter value and the new observed statistics, and the weight coefficient is initially set to 0.3. To adapt to the change rates of different scenarios and achieve dynamic optimization of the parameters, an adaptive adjustment mechanism is introduced: when the change trends of the prediction errors in three consecutive windows are the same, the weight coefficient is increased to 1.5 times; when the prediction errors fluctuate without obvious rules, the weight coefficient is decreased to 0.8 times, and the weight coefficient range is limited within the interval [0.1, 0.6]. Through the above mechanism, an updated posterior probability distribution is generated.
[0192] This mechanism is applied in a retail forecasting system. The initial window size is set to 14 days, and the weight coefficient is 0.3. The system detects that the prediction errors of seasonal products in three consecutive windows continue to increase, and automatically adjusts the weight coefficient to 0.45, enabling the posterior distribution to adapt to the change trend faster. After the update, the mean of the prediction errors of seasonal products is adjusted from 4.2% to 5.7%, and the variance increases from 0.8% to 1.3%, more accurately reflecting the actual prediction situation.
[0193] The updated posterior probability distribution is input into the dynamic evaluation module. This module first constructs a confidence interval for the prediction error based on the posterior distribution, usually choosing a 95% confidence level. The upper and lower bounds of the confidence interval correspond to the 2.5th percentile and the 97.5th percentile of the posterior distribution respectively. By analyzing the change trend of the confidence interval, the performance of the prediction model is evaluated. The specific analysis methods include: monitoring the change in the width of the confidence interval, an increase in width indicates an increase in prediction uncertainty; tracking the displacement of the center point of the confidence interval, an upward shift indicates that the prediction model tends to underestimate, and a downward shift indicates a tendency to overestimate; recording the frequency of the actual error falling within the confidence interval, which should ideally be close to 95% under ideal circumstances. Based on the above analysis, a dynamic evaluation result of the prediction error containing multi-dimensional indicators is generated.
[0194] In the practical application of an energy demand forecasting system, the dynamic evaluation module finds that the 95% confidence interval for weekday forecasts is [1.2%, 3.4%], while it expands to [2.5%, 5.1%] on weekends. After continuous monitoring for one month, the system finds that the proportion of actual errors falling within the confidence interval is 92.3%, close to the ideal value; the center of the confidence interval gradually moves down from 2.3% to 1.8%, indicating an improvement in the performance of the prediction model; after extreme weather events, the width of the confidence interval temporarily expands to 2.3 times the original, accurately reflecting the increase in prediction uncertainty under special circumstances.
[0195] Figure 5Schematic diagram for comparing prediction errors of the double-closed-loop evolutionary system according to the embodiments of the present invention:
[0196] This line graph shows the changing trends of the average prediction errors of three different prediction schemes (the technical scheme of the present application, single-loop prediction, and static prediction model) within 48 hours of system operation time. It can be seen from the graph that the prediction errors of all three methods are between 4.5 and 5 at the initial moment (0 hour). As the system operation time progresses, the technical scheme of the present application (black triangular line) shows the best performance, with the prediction error showing a continuous downward trend, gradually decreasing from approximately 4.8 initially to around 2.2 at 16 hours, and dropping below 1.2 at 48 hours. The single-loop prediction method (blue square line) ranks second, with the error showing a downward trend within the first 16 hours to around 3.0, but experiencing a significant fluctuation and rising to 4.5 at 20 hours, and then basically fluctuating between 4.0 and 5.0. The static prediction model (red diamond line) shows the worst performance, remaining relatively stable between 4.5 and 5.0 in the first 16 hours, but showing a significant increase after 20 hours, with the error increasing to nearly 7.0 and continuing to rise slowly thereafter, reaching approximately 8.0 at 48 hours. This comparison clearly demonstrates that the technical scheme of the present application has better prediction accuracy and stability during long-term operation, especially after the system has been running for more than 20 hours, when its advantages become more obvious.
[0197] The technical scheme proposed in this application effectively solves the technical problems of poor stability and decreasing accuracy of the prediction model in the prior art during long-term operation by introducing an adaptive dynamic prediction mechanism, combining multi-level data analysis and real-time feedback correction strategies.
[0198] In the prior art, the single-loop prediction method mainly relies on a single data source for prediction, lacking the comprehensive utilization of multi-dimensional information, resulting in the prediction results being easily affected by external interference and fluctuating. The static prediction model, on the other hand, uses fixed prediction parameters and cannot be dynamically adjusted according to the system operation state, causing the prediction accuracy to decrease significantly over time, especially showing obvious performance degradation during long-term operation.
[0199] To address the above problems, this application designs an adaptive prediction framework based on deep learning and improves the performance through the following technical means: First, a multi-source data fusion mechanism is constructed to fully utilize multi-dimensional information such as historical data, real-time monitoring data, and environmental factors; second, a dynamic weight adjustment algorithm is adopted to optimize the prediction model parameters in real time according to the system operation state; finally, an error feedback compensation mechanism is introduced to continuously optimize the prediction accuracy.
[0200] The implementation effect shows that this technical solution not only significantly improves the prediction accuracy, but also has better long-term stability. Especially during the continuous operation stage of the system, compared with the existing technical solutions, the prediction error of this application continuously decreases and remains at a low level, fully demonstrating the superiority of this technical solution in practical applications. At the same time, the adaptive characteristics of this solution enable it to better cope with various uncertain factors during system operation, providing more reliable prediction support for the industrial production process.
[0201] In the second aspect of the embodiments of the present invention,
[0202] a crane remote instruction response delay detection and prior and post compensation system is provided, including:
[0203] A first unit for constructing a dynamic knowledge graph based on physical prior knowledge, mapping the real-time obtained physical domain operation data to the dynamic knowledge graph by using an adaptive spatio-temporal alignment algorithm to generate a hierarchical feature representation vector; inputting the feature representation vector into a deep graph neural network integrating a causal reasoning mechanism, and generating an association matrix representing physical laws and data patterns through recursive reasoning; using a multi-head attention network with a residual structure to extract temporal features of the association matrix, and combining with a variational Bayesian inference unit to output a delay prediction tensor containing uncertainty quantification;
[0204] A second unit for constructing a hybrid decision-making system with hierarchical adaptive capabilities based on the delay prediction tensor; in the policy generation layer of the hybrid decision-making system, fusing the current system state with the delay prediction tensor to generate an initial compensation action set; in the collaborative optimization layer of the hybrid decision-making system, combining a multi-agent experience replay mechanism to perform distributed optimization on the initial compensation action set, and performing multi-constraint projection on the optimized compensation actions to output an optimal compensation strategy considering safety and real-time performance;
[0205] A third unit for establishing a double-loop evolutionary system with online learning capabilities for the optimal compensation strategy; in the prediction optimization loop of the double-loop evolutionary system, designing an error evaluation module based on Bayesian posterior inference, and adaptively adjusting the causal reasoning mechanism and feature extraction strategy through meta-learning methods; in the compensation optimization loop of the double-loop evolutionary system, constructing a multi-level progressive evaluation framework to dynamically optimize the policy generation layer and the collaborative optimization layer based on the compensation effect.
[0206] In the third aspect of the embodiments of the present invention,
[0207] an electronic device is provided, including:
[0208] a processor;
[0209] a memory for storing processor-executable instructions;
[0210] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0211] In the fourth aspect of the embodiments of the present invention,
[0212] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0213] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are loaded.
[0214] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting the response delay of remote instructions of a crane and compensating for the prior and subsequent delays, characterized in that Including: Construct a dynamic knowledge graph based on physical prior knowledge, and use an adaptive spatio-temporal alignment algorithm to map the real-time obtained physical domain operation data to the dynamic knowledge graph, generating a hierarchical feature representation vector; Input the feature representation vector into a deep graph neural network integrated with a causal inference mechanism, and generate a correlation matrix representing physical laws and data patterns through recursive inference; Use a multi-head attention network with a residual structure to extract temporal features from the correlation matrix, and combine with a variational Bayesian inference unit to output a delay prediction tensor containing uncertainty quantification; Construct a hybrid decision-making system with hierarchical adaptive capabilities based on the delay prediction tensor; In the policy generation layer of the hybrid decision-making system, fuse the current system state with the delay prediction tensor to generate an initial set of compensation actions; In the collaborative optimization layer of the hybrid decision-making system, combine the multi-agent experience replay mechanism to perform distributed optimization on the initial set of compensation actions, and perform multi-constraint projection on the optimized compensation actions to output an optimal compensation strategy considering safety and real-time performance; For the optimal compensation strategy, establish a dual-loop evolutionary system with online learning capabilities; In the prediction optimization loop of the dual-loop evolutionary system, design an error evaluation module based on Bayesian posterior inference, and adaptively adjust the causal inference mechanism and feature extraction strategy through meta-learning methods; In the compensation optimization loop of the dual-loop evolutionary system, construct a multi-level progressive evaluation framework, and dynamically optimize the policy generation layer and the collaborative optimization layer based on the compensation effect.
2. The method according to claim 1, characterized in that, Construct a dynamic knowledge graph based on physical prior knowledge, and use an adaptive spatio-temporal alignment algorithm to map the real-time obtained physical domain operation data to the dynamic knowledge graph, generating a hierarchical feature representation vector; Input the feature representation vector into a deep graph neural network integrated with a causal inference mechanism, and generate a correlation matrix representing physical laws and data patterns through recursive inference, including: Construct a dynamic knowledge graph based on physical prior knowledge, transform the physical prior knowledge into a parametric constraint equation set, define the node attributes and edge relationships of the dynamic knowledge graph based on the parametric constraint equation set, the node attributes represent the evolution law of physical quantities, and the edge relationships represent the coupling mechanism between physical quantities, generating an initial knowledge graph structure with physical law constraints; Use an adaptive spatio-temporal alignment algorithm to map the real-time obtained physical domain operation data to the dynamic knowledge graph, design a feature extraction strategy based on the topological relationship of the initial knowledge graph structure, align the physical domain operation data through a dynamic window mechanism, use a multi-level feature fusion network to extract spatial features, and perform semantic matching between the extracted features and the nodes of the dynamic knowledge graph to generate a hierarchical feature representation vector reflecting the multi-scale dynamic characteristics of the system; Input the hierarchical feature representation vector into a deep graph neural network integrated with a causal inference mechanism, construct a causal discovery criterion based on the physical constraints of the initial knowledge graph structure, establish an initial causal link through conditional independence test and directed information entropy analysis, and use the message passing algorithm of the attention mechanism to perform recursive inference and dynamic update on the initial causal link, generating a correlation matrix representing physical laws and data patterns.
3. The method according to claim 2, wherein Input the hierarchical feature representation vector into the deep graph neural network of the integrated causal inference mechanism, construct a causal discovery criterion based on the physical constraints of the initial knowledge graph structure, establish initial causal links through conditional independence testing and directed information entropy analysis, and use the message passing algorithm of the attention mechanism to recursively reason and dynamically update the initial causal links, generating an association matrix representing the physical laws and data patterns, including: Input the hierarchical feature representation vector into the deep graph neural network, encode the physical constraints based on the initial knowledge graph, construct a physical constraint criterion including energy conservation and momentum conservation, and use the physical constraint criterion to perform a structured representation of the hierarchical feature representation vector, generating a feature map structure that satisfies the physical laws; Calculate the conditional mutual information between nodes based on the feature map structure, evaluate the causal strength between node pairs through directed information entropy analysis, combine the conditional mutual information and causal strength to construct initial causal links, and generate an initial causal graph with weights; Use the message passing algorithm of the attention mechanism to dynamically optimize the initial causal graph, construct a multi-head attention layer to calculate the association features between nodes, input the association features into the gated recurrent unit to model the temporal dependence relationship, and update the weights of the causal links through iterative message passing, generating an optimized causal graph structure; Generate an association matrix based on the optimized causal graph structure, use an adaptive weight mechanism to adjust the matrix elements according to the prediction error, project the association matrix into the feasible region that satisfies the physical constraints through gradient projection, perform uncertainty analysis on the association matrix within the feasible region, and generate the final mapping relationship matrix between the physical laws and data patterns.
4. The method according to claim 1, characterized in that, In the collaborative optimization layer of the hybrid decision system, combine the multi-agent experience replay mechanism to perform distributed optimization on the initial set of compensation actions, and perform multi-constraint projection on the optimized compensation actions, outputting the optimal compensation strategy considering safety and real-time performance, including: Construct a dynamic graph attention network in the collaborative optimization layer of the hybrid decision system, calculate the similarity of the state vectors between agents to obtain the initial attention scores, and construct a hierarchical experience replay structure based on the initial attention scores. The hierarchical experience replay structure includes a local experience pool and a global experience pool, and store the initial set of compensation actions and their execution effects in the local experience pool; Perform priority sampling based on the experience value from the global experience pool, use a distributed policy optimization algorithm to optimize the initial set of compensation actions, and use an adaptive gradient projection algorithm to project the optimized compensation actions into the feasible region that satisfies the constraints; Perform multi-level evaluation on the projected compensation actions, calculate the immediate reward based on the compensation error, calculate the long-term benefit through state value evaluation, update the experience values in the local experience pool and the global experience pool according to the immediate reward and the long-term benefit, and at the same time adjust the attention weight matrix based on the evaluation results, optimize the interaction structure of the adaptive communication topology, and output the optimally compensated strategy after collaborative optimization.
5. The method according to claim 4, wherein Perform priority sampling based on the experience value from the global experience pool, use a distributed policy optimization algorithm to optimize the initial set of compensation actions, and use an adaptive gradient projection algorithm to project the optimized compensation actions into the feasible region that satisfies the constraints, including: Calculate the cumulative return value of the compensation action based on the global experience pool, use the cumulative return value as the experience value metric, construct a priority sampling tree based on the experience value metric, and extract priority value experience samples from the priority sampling tree; Input the priority value experience samples into the distributed policy optimization algorithm. The distributed policy optimization algorithm includes an action generation network and a value evaluation network. The action generation network generates an initial compensation action set based on the priority value experience samples, uses the value evaluation network to calculate the expected return of the compensation action, and optimizes the initial compensation action set based on the expected return; Input the optimized compensation action into the constraint processing module. The constraint processing module constructs state space constraints based on the system dynamics model, maps the system safety index to state variable constraints, converts the system real-time index into calculation time constraints, and uses the barrier function method to integrate the state space constraints, state variable constraints, and calculation time constraints into dynamic constraint conditions; Perform constraint processing on the optimized compensation action based on the dynamic constraint conditions, calculate the gradient projection step size based on the degree of constraint violation, solve the projection direction through the conjugate gradient method, and project the optimized compensation action onto the feasible region that satisfies the dynamic constraint conditions using the progressive projection strategy.
6. The method according to claim 1, wherein For the optimal compensation strategy, establish a double-loop evolutionary system with online learning ability; in the prediction optimization loop of the double-loop evolutionary system, design an error evaluation module based on Bayesian posterior inference, and adaptively adjust the causal inference mechanism and feature extraction strategy through meta-learning methods, including: Construct an online learning double-loop evolutionary system to continuously optimize the optimal compensation strategy. Set up a prediction optimization loop and a compensation optimization loop in the online learning double-loop evolutionary system, use the online learning method to collect the execution data of the compensation strategy in real time, and input the execution data into the prediction optimization loop; Construct an error evaluation module of Bayesian posterior inference in the prediction optimization loop, establish a prior probability model of the prediction error using Gaussian process, calculate the posterior probability distribution through variational inference method, continuously update the posterior probability distribution based on the sliding time window mechanism, and generate a dynamic evaluation result of the prediction error; Calculate the parameter optimization direction of the causal inference mechanism based on the dynamic evaluation result, adaptively adjust the causal graph structure and inference parameters based on the parameter optimization direction, dynamically optimize the causal inference mechanism, and realize the adaptive discovery of causal relationships; Input the optimization result of the causal inference mechanism into the feature extraction module, adaptively adjust the feature extraction strategy through meta-learning method, and update the causal inference mechanism through the online dictionary learning method based on the adaptively adjusted feature extraction strategy to generate an optimized feature extraction strategy.
7. The method according to claim 6, wherein Construct an error evaluation module of Bayesian posterior inference in the prediction optimization loop, establish a prior probability model of the prediction error using Gaussian process, calculate the posterior probability distribution through variational inference method, continuously update the posterior probability distribution based on the sliding time window mechanism, and generate a dynamic evaluation result of the prediction error, including: Construct an error evaluation module for Bayesian posterior inference in the prediction optimization loop, establish a prior probability model of the prediction error using a Gaussian process, construct a kernel function through a radial basis function to describe the temporal correlation of the error, and optimize the kernel function based on the maximum likelihood estimation method to generate the prior distribution of the prediction error; Input the prior distribution of the prediction error into the variational inference module, construct an evidence lower bound objective function containing a likelihood function and a variational function, optimize the evidence lower bound objective function by the stochastic gradient ascent method, and iteratively update the evidence lower bound objective function to generate the posterior probability distribution of the prediction error; Based on the posterior probability distribution, set up a sliding time window mechanism, calculate the sufficient statistics of the sample data within the window, continuously update the posterior distribution parameters using the exponential weighted average method, and dynamically optimize the parameters by adaptively adjusting the smoothing factor to generate the updated posterior probability distribution; Input the updated posterior probability distribution into the dynamic evaluation module, construct a confidence interval of the prediction error based on the dynamic evaluation module, evaluate the performance of the prediction model by analyzing the change trend of the confidence interval, and generate the dynamic evaluation result of the prediction error.
8. A remote instruction response delay detection and prior and subsequent compensation system for a crane, which is used to implement the method described in any one of the foregoing claims 1-7, and is characterized in that It includes: The first unit is used to construct a dynamic knowledge graph based on physical prior knowledge, and use an adaptive spatio-temporal alignment algorithm to map the real-time obtained physical domain operation data to the dynamic knowledge graph to generate a hierarchical feature representation vector; Input the feature representation vector into the deep graph neural network integrating the causal inference mechanism, and generate an association matrix representing physical laws and data patterns through recursive inference; use a multi-head attention network with a residual structure to extract the temporal features of the association matrix, and combine the variational Bayesian inference unit to output a delay prediction tensor containing uncertainty quantification; The second unit is used to construct a hybrid decision-making system with hierarchical adaptive capabilities based on the delay prediction tensor; in the policy generation layer of the hybrid decision-making system, fuse the current system state with the delay prediction tensor to generate an initial set of compensation actions; in the collaborative optimization layer of the hybrid decision-making system, combine the multi-agent experience replay mechanism to perform distributed optimization on the initial set of compensation actions, and perform multi-constraint projection on the optimized compensation actions to output the optimal compensation strategy considering safety and real-time performance; The third unit is used to establish a double-closed-loop evolution system with online learning ability for the optimal compensation strategy; In the prediction optimization loop of the double-closed-loop evolution system, design an error evaluation module based on Bayesian posterior inference, and adaptively adjust the causal inference mechanism and feature extraction strategy through meta-learning methods; In the compensation optimization loop of the double-closed-loop evolution system, construct a multi-level progressive evaluation framework, and dynamically optimize the policy generation layer and the collaborative optimization layer based on the compensation effect.
9. An electronic device, characterized in that, It includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Intelligent fault diagnosis method and system for portal crane
CN119441777A
Petrochemical production process anomaly diagnosis and optimization method and system integrated with knowledge graph
CN119668245A