Crane remote instruction response delay detection and prior-prior compensation method and system

By constructing a dynamic knowledge graph and a deep graph neural network for causal reasoning, combining a hybrid decision-making system and a dual closed-loop evolution system, the problem of delay in the crane's remote instruction response is solved, and efficient delay detection and compensation strategy optimization is achieved.

CN120103715AActive Publication Date: 2025-06-06NINGBO SPECIAL EQUIP INSPECTION & RES INST

Patent Information

Application Number
CN202510578648.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-06
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of delay in remote command response of cranes, especially in terms of signal transmission and network delay, resulting in reduced operating accuracy and increased accident risk.

Method used

By constructing dynamic knowledge graphs and adaptive spatiotemporal alignment algorithms, hierarchical feature representation vectors are generated, and input them into the depth graph neural network for causal reasoning, and output delay prediction tensors. Based on this, a hybrid decision-making system is built to generate and optimize compensation strategies, and a dual closed-loop evolution system is established for online learning and strategy adjustment.

Benefits of technology

It significantly improves the accuracy and robustness of delay detection, realizes efficient decision-making and compensation strategies in complex environments, and improves the stability and reliability of remote operation of cranes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120103715A_ABST
    Figure CN120103715A_ABST
Patent Text Reader

Abstract

The invention provides a crane remote instruction response delay detection and in-advance compensation method and system, and relates to the technical field of cranes, and the crane remote instruction response delay detection and in-advance compensation method comprises the following steps: adopting an adaptive space-time alignment algorithm to map real-time operation data to a dynamic knowledge graph, and generating a feature vector; inputting the feature vector into a depth map neural network integrated with a causal reasoning mechanism to generate an incidence matrix; a multi-head attention network with a residual structure is adopted to extract time sequence features; constructing a hybrid decision system based on the delay prediction tensor, and outputting an optimal compensation strategy; and establishing a double-closed-loop evolution system with an online learning capability, and dynamically optimizing a prediction and compensation strategy according to a compensation effect. Through the dynamic knowledge graph, causal reasoning, the multi-head attention network and the double-closed-loop evolution system, the remote instruction response delay can be accurately predicted, effective compensation is carried out, and the real-time performance and safety of remote control of the crane are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to crane technology, and in particular to a crane remote command response delay detection and prior and subsequent compensation method and system. Background Art

[0002] Remote control of cranes plays a vital role in modern industry, especially in ports, construction sites and other scenarios. Efficient and accurate remote operation is essential to ensure production safety and improve production efficiency. The core of remote control lies in real-time and reliability, and operating instructions need to be executed quickly and accurately. However, due to factors such as signal transmission, network delays, and complex mechanical system responses, the execution of remote instructions inevitably has delays. This delay will reduce operating accuracy, increase the risk of accidents, and even lead to serious production accidents.

[0003] In order to solve the problem of remote command response delay, existing technical solutions mainly focus on the following aspects: solutions based on network optimization attempt to reduce network delays by improving network transmission protocols, optimizing network topology, etc.; solutions based on prediction models try to establish mathematical models to predict delays and perform pre-compensation; solutions based on control algorithms focus on designing advanced control algorithms to overcome the impact of delays.

[0004] Most existing solutions only focus on a single aspect of network delay or system response delay, lacking consideration of the combined impact of the two, and are unable to effectively deal with complex delay situations in actual operations. Traditional prediction models are often based on simplified assumptions, and are unable to accurately capture the complex nonlinear relationships and dynamic changes in actual systems, resulting in insufficient prediction accuracy, which in turn affects the compensation effect. Existing compensation strategies usually lack adaptability and robustness, and are unable to cope with changes in different operating environments and task requirements, resulting in unstable compensation effects and may even cause new control problems. Summary of the invention

[0005] The embodiment of the present invention provides a crane remote command response delay detection and prior and subsequent compensation method and system, which can solve the problems in the prior art.

[0006] According to a first aspect of the embodiments of the present invention, Provides crane remote command response delay detection and first-come-first-served compensation methods, including: Based on physical prior knowledge, a dynamic knowledge graph is constructed. The real-time acquired physical domain operation data is mapped to the dynamic knowledge graph using an adaptive spatiotemporal alignment algorithm to generate a hierarchical feature representation vector. The feature representation vector is input into a deep graph neural network with an integrated causal reasoning mechanism, and a correlation matrix representing physical laws and data patterns is generated through recursive reasoning. A multi-head attention network with a residual structure is used to extract temporal features from the correlation matrix, and a variational Bayesian reasoning unit is used to output a delay prediction tensor containing uncertainty quantification. A hybrid decision system with hierarchical adaptive capabilities is constructed based on the delay prediction tensor. In the strategy generation layer of the hybrid decision system, the current system state is fused with the delay prediction tensor to generate an initial compensation action set. In the collaborative optimization layer of the hybrid decision system, the initial compensation action set is distributedly optimized in combination with the multi-agent experience playback mechanism, and the optimized compensation actions are projected with multiple constraints to output the optimal compensation strategy that takes into account safety and real-time performance. Aiming at the optimal compensation strategy, a double closed-loop evolutionary system with online learning capability is established. In the prediction optimization loop of the double closed-loop evolutionary system, an error evaluation module based on Bayesian posterior inference is designed, and the causal reasoning mechanism and feature extraction strategy are adaptively adjusted through meta-learning methods. In the compensation optimization loop of the double closed-loop evolutionary system, a multi-level progressive evaluation framework is constructed, and the strategy generation layer and the collaborative optimization layer are dynamically optimized based on the compensation effect.

[0007] Based on physical prior knowledge, a dynamic knowledge graph is constructed. The real-time acquired physical domain operation data is mapped to the dynamic knowledge graph using an adaptive spatiotemporal alignment algorithm to generate a hierarchical feature representation vector. The feature representation vector is input into a deep graph neural network with an integrated causal reasoning mechanism. The association matrix representing the physical laws and data patterns is generated through recursive reasoning, including: Based on physical prior knowledge, a dynamic knowledge graph is constructed. The physical prior knowledge is converted into a set of parameterized constraint equations. Based on the set of parameterized constraint equations, the node attributes and edge relationships of the dynamic knowledge graph are defined. The node attributes represent the evolution law of physical quantities, and the edge relationships represent the coupling mechanism between physical quantities, thus generating an initial knowledge graph structure with physical law constraints. An adaptive spatiotemporal alignment algorithm is used to map the real-time acquired physical domain operation data to the dynamic knowledge graph. A feature extraction strategy is designed based on the topological relationship of the initial knowledge graph structure. The physical domain operation data is aligned through a dynamic window mechanism. A multi-level feature fusion network is used to extract spatial features. The extracted features are semantically matched with the nodes of the dynamic knowledge graph to generate a hierarchical feature representation vector that reflects the multi-scale dynamic characteristics of the system. The hierarchical feature representation vector is input into a deep graph neural network with an integrated causal reasoning mechanism. The causal discovery criteria are constructed based on the physical constraints of the initial knowledge graph structure. The initial causal links are established through conditional independence test and directed information entropy analysis. The message passing algorithm with attention mechanism is used to recursively infer and dynamically update the initial causal links to generate a correlation matrix representing physical laws and data patterns.

[0008] The hierarchical feature representation vector is input into the deep graph neural network with integrated causal reasoning mechanism. The causal discovery criteria are constructed based on the physical constraints of the initial knowledge graph structure. The initial causal links are established through conditional independence test and directed information entropy analysis. The message passing algorithm of the attention mechanism is used to recursively reason and dynamically update the initial causal links. The association matrix representing the physical laws and data patterns is generated, including: The hierarchical feature representation vector is input into the deep graph neural network, the physical constraints are encoded based on the initial knowledge graph, and the physical constraint criteria including energy conservation and momentum conservation are constructed. The hierarchical feature representation vector is structured using the physical constraint criteria to generate a feature graph structure that satisfies the physical laws. The conditional mutual information between nodes is calculated based on the characteristic graph structure, and the causal strength between node pairs is evaluated through directed information entropy analysis. The conditional mutual information and causal strength are combined to construct the initial causal link and generate an initial causal graph with weights. The initial causal graph is dynamically optimized using the message passing algorithm of the attention mechanism. The multi-head attention layer is constructed to calculate the correlation features between nodes. The correlation features are input into the gated recursive unit to model the temporal dependency. The weights of the causal links are updated through iterative message passing to generate the optimized causal graph structure. The association matrix is ​​generated based on the optimized causal graph structure, and the matrix elements are adjusted according to the prediction error using an adaptive weight mechanism. The association matrix is ​​mapped to the feasible domain that satisfies the physical constraints through gradient projection. The uncertainty analysis of the association matrix in the feasible domain is performed to generate the final physical law and data pattern mapping relationship matrix.

[0009] In the collaborative optimization layer of the hybrid decision-making system, the initial compensation action set is distributedly optimized in combination with the multi-agent experience playback mechanism, and the optimized compensation actions are multi-constrained projected to output the optimal compensation strategy that takes into account security and real-time performance, including: A dynamic graph attention network is constructed in the collaborative optimization layer of the hybrid decision-making system, and the state vector similarity between agents is calculated to obtain the initial attention score. A hierarchical experience playback structure is constructed based on the initial attention score. The hierarchical experience playback structure includes a local experience pool and a global experience pool, and the initial compensation action set and its execution effect are stored in the local experience pool. Based on the experience value, the global experience pool is sampled first, the initial compensation action set is optimized by using the distributed strategy optimization algorithm, and the optimized compensation action is projected to the feasible domain that satisfies the constraints by using the adaptive gradient projection algorithm. A multi-level evaluation is performed on the compensation action after projection, the immediate reward is calculated based on the compensation error, the long-term benefit is calculated through state value evaluation, and the experience value in the local experience pool and the global experience pool is updated according to the immediate reward and long-term benefit. At the same time, the attention weight matrix is ​​adjusted based on the evaluation results, the interaction structure of the adaptive communication topology is optimized, and the optimal compensation strategy after collaborative optimization is output.

[0010] Based on the experience value, the global experience pool is sampled first, and the initial compensation action set is optimized using the distributed strategy optimization algorithm. The optimized compensation action is projected to the feasible domain that satisfies the constraints using the adaptive gradient projection algorithm, including: Calculate the cumulative reward value of the compensation action based on the global experience pool, use the cumulative reward value as the experience value indicator, build a priority sampling tree based on the experience value indicator, and extract priority value experience samples from the priority sampling tree; The priority value experience samples are input into the distributed strategy optimization algorithm, which includes an action generation network and a value evaluation network. The action generation network generates an initial compensation action set based on the priority value experience samples, and the value evaluation network is used to calculate the expected benefits of the compensation actions. The initial compensation action set is optimized based on the expected benefits. The optimized compensation action is input into the constraint processing module. The constraint processing module constructs state space constraints based on the system dynamics model, maps the system safety index into state variable constraints, converts the system real-time index into calculation time constraints, and uses the barrier function method to integrate the state space constraints, state variable constraints and calculation time constraints into dynamic constraints. The optimized compensation action is constrained based on dynamic constraints, the gradient projection step is calculated based on the degree of constraint violation, the projection direction is solved by the conjugate gradient method, and the progressive projection strategy is used to project the optimized compensation action to the feasible domain that meets the dynamic constraints.

[0011] Aiming at the optimal compensation strategy, a double closed-loop evolutionary system with online learning capability is established. In the prediction optimization loop of the double closed-loop evolutionary system, an error evaluation module based on Bayesian posterior inference is designed. The causal reasoning mechanism and feature extraction strategy are adaptively adjusted through the meta-learning method, including: Construct an online learning double closed-loop evolutionary system to continuously optimize the optimal compensation strategy. Set up a prediction optimization loop and a compensation optimization loop in the online learning double closed-loop evolutionary system, use online learning methods to collect the execution data of the compensation strategy in real time, and input the execution data into the prediction optimization loop; In the prediction optimization loop, a Bayesian posterior inference error evaluation module is constructed. The Gaussian process is used to establish a priori probability model of the prediction error. The posterior probability distribution is calculated through the variational inference method. The posterior probability distribution is continuously updated based on the sliding time window mechanism to generate dynamic evaluation results of the prediction error. Calculate the parameter optimization direction of the causal reasoning mechanism based on the dynamic evaluation results, adaptively adjust the causal graph structure and reasoning parameters based on the parameter optimization direction, dynamically optimize the causal reasoning mechanism, and realize adaptive discovery of causal relationships; The optimization results of the causal reasoning mechanism are input into the feature extraction module, and the meta-learning method is used to adaptively adjust the feature extraction strategy. Based on the adaptively adjusted feature extraction strategy, the causal reasoning mechanism is updated through the online dictionary learning method to generate an optimized feature extraction strategy.

[0012] In the prediction optimization loop, a Bayesian posterior inference error evaluation module is constructed. The Gaussian process is used to establish a priori probability model of the prediction error. The posterior probability distribution is calculated by the variational inference method. The posterior probability distribution is continuously updated based on the sliding time window mechanism to generate dynamic evaluation results of the prediction error, including: In the prediction optimization loop, a Bayesian posterior inference error evaluation module is constructed. The Gaussian process is used to establish a prior probability model of the prediction error. The kernel function is constructed through the radial basis function to describe the time correlation of the error. The kernel function is optimized based on the maximum likelihood estimation method to generate the prior distribution of the prediction error. The prior distribution of the prediction error is input into the variational inference module, and the evidence lower bound objective function including the likelihood function and the variational function is constructed. The evidence lower bound objective function is optimized by the stochastic gradient ascent method, and the evidence lower bound objective function is iteratively updated to generate the posterior probability distribution of the prediction error. A sliding time window mechanism is set based on the posterior probability distribution, sufficient statistics of the sample data in the window are calculated, and the posterior distribution parameters are continuously updated using the exponential weighted average method. Dynamic optimization of parameters is achieved by adaptively adjusting the smoothing factor to generate an updated posterior probability distribution. The updated posterior probability distribution is input into the dynamic evaluation module, and the confidence interval of the prediction error is constructed based on the dynamic evaluation module. The performance of the prediction model is evaluated by analyzing the changing trend of the confidence interval to generate a dynamic evaluation result of the prediction error.

[0013] According to a second aspect of the embodiments of the present invention, Provide crane remote command response delay detection and first-come-first-served compensation system, including: The first unit is used to build a dynamic knowledge graph based on physical prior knowledge, and uses an adaptive spatiotemporal alignment algorithm to map the real-time physical domain operation data to the dynamic knowledge graph to generate a hierarchical feature representation vector; the feature representation vector is input into a deep graph neural network with an integrated causal reasoning mechanism, and a correlation matrix representing physical laws and data patterns is generated through recursive reasoning; a multi-head attention network with a residual structure is used to extract time series features from the correlation matrix, and a variational Bayesian reasoning unit is used to output a delay prediction tensor containing uncertainty quantification; The second unit is used to build a hybrid decision system with hierarchical adaptive capabilities based on the delay prediction tensor; in the strategy generation layer of the hybrid decision system, the current system state and the delay prediction tensor are fused to generate an initial compensation action set; in the collaborative optimization layer of the hybrid decision system, the initial compensation action set is distributedly optimized in combination with the multi-agent experience playback mechanism, and the optimized compensation action is multi-constrained projected to output the optimal compensation strategy that takes into account safety and real-time performance; The third unit is used to establish a double closed-loop evolutionary system with online learning capabilities for the optimal compensation strategy; in the prediction optimization loop of the double closed-loop evolutionary system, an error evaluation module based on Bayesian posterior inference is designed, and the causal reasoning mechanism and feature extraction strategy are adaptively adjusted through meta-learning methods; in the compensation optimization loop of the double closed-loop evolutionary system, a multi-level progressive evaluation framework is constructed, and the strategy generation layer and the collaborative optimization layer are dynamically optimized based on the compensation effect.

[0014] According to a third aspect of the embodiments of the present invention, An electronic device is provided, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0015] A fourth aspect of the embodiments of the present invention is: A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0016] The beneficial effects of this application are as follows: The present invention achieves accurate mapping and feature extraction of physical data by constructing a dynamic knowledge graph and an adaptive spatiotemporal alignment algorithm, effectively improving the accuracy and robustness of delay detection. At the same time, the combination of a deep graph neural network with an integrated causal reasoning mechanism and a multi-head attention network enables the system to accurately capture the timing characteristics and causal relationships under complex working conditions, providing a reliable theoretical basis for compensation strategies.

[0017] The hybrid decision-making system of the present invention has hierarchical adaptive capabilities, can generate initial compensation actions based on system status and delay prediction, and perform distributed optimization through a multi-agent experience playback mechanism, thus achieving efficient decision-making in complex environments. Multi-constraint projection technology ensures that the compensation strategy meets both safety and real-time requirements, significantly improving the stability and reliability of crane remote operation.

[0018] The dual closed-loop evolutionary system of the present invention has the ability of online learning, and realizes the continuous optimization and self-adjustment of the prediction model through Bayesian posterior inference and meta-learning methods. The multi-level progressive evaluation framework enables the compensation strategy to be dynamically adjusted according to the actual effect, adapting to different working conditions and environmental changes, greatly reducing the system maintenance cost, extending the service life of the equipment, and improving the overall efficiency of remote operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flow chart of a method for detecting a time delay in response to a remote command of a crane and compensating for the delay in response to a remote command of a crane according to an embodiment of the present invention; Figure 2 This is a schematic diagram of comparative analysis of causal discovery accuracy of deep graph neural networks according to an embodiment of the present invention; Figure 3 A schematic diagram of performance comparison and analysis of the hierarchical experience playback structure according to an embodiment of the present invention; Figure 4 This is a schematic diagram of overall performance comparison of the compensation action optimization method according to an embodiment of the present invention; Figure 5 Schematic diagram of comparison of prediction errors of the double closed-loop evolutionary system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0021] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0022] Figure 1 FIG. 1 is a flow chart of a crane remote command response delay detection and prior-later compensation method according to an embodiment of the present invention, as shown in FIG. Figure 1 As shown, the method includes: Based on physical prior knowledge, a dynamic knowledge graph is constructed. The real-time acquired physical domain operation data is mapped to the dynamic knowledge graph using an adaptive spatiotemporal alignment algorithm to generate a hierarchical feature representation vector. The feature representation vector is input into a deep graph neural network with an integrated causal reasoning mechanism, and a correlation matrix representing physical laws and data patterns is generated through recursive reasoning. A multi-head attention network with a residual structure is used to extract temporal features from the correlation matrix, and a variational Bayesian reasoning unit is used to output a delay prediction tensor containing uncertainty quantification. A hybrid decision system with hierarchical adaptive capabilities is constructed based on the delay prediction tensor. In the strategy generation layer of the hybrid decision system, the current system state is fused with the delay prediction tensor to generate an initial compensation action set. In the collaborative optimization layer of the hybrid decision system, the initial compensation action set is distributedly optimized in combination with the multi-agent experience playback mechanism, and the optimized compensation actions are projected with multiple constraints to output the optimal compensation strategy that takes into account safety and real-time performance. Aiming at the optimal compensation strategy, a double closed-loop evolutionary system with online learning capability is established. In the prediction optimization loop of the double closed-loop evolutionary system, an error evaluation module based on Bayesian posterior inference is designed, and the causal reasoning mechanism and feature extraction strategy are adaptively adjusted through meta-learning methods. In the compensation optimization loop of the double closed-loop evolutionary system, a multi-level progressive evaluation framework is constructed, and the strategy generation layer and the collaborative optimization layer are dynamically optimized based on the compensation effect.

[0023] In an optional implementation, a dynamic knowledge graph is constructed based on physical prior knowledge, and an adaptive spatiotemporal alignment algorithm is used to map the real-time acquired physical domain operation data to the dynamic knowledge graph to generate a hierarchical feature representation vector; the feature representation vector is input into a deep graph neural network with an integrated causal reasoning mechanism, and a correlation matrix representing physical laws and data patterns is generated through recursive reasoning, including: Based on physical prior knowledge, a dynamic knowledge graph is constructed. The physical prior knowledge is converted into a set of parameterized constraint equations. Based on the set of parameterized constraint equations, the node attributes and edge relationships of the dynamic knowledge graph are defined. The node attributes represent the evolution law of physical quantities, and the edge relationships represent the coupling mechanism between physical quantities, thus generating an initial knowledge graph structure with physical law constraints. An adaptive spatiotemporal alignment algorithm is used to map the real-time acquired physical domain operation data to the dynamic knowledge graph. A feature extraction strategy is designed based on the topological relationship of the initial knowledge graph structure. The physical domain operation data is aligned through a dynamic window mechanism. A multi-level feature fusion network is used to extract spatial features. The extracted features are semantically matched with the nodes of the dynamic knowledge graph to generate a hierarchical feature representation vector that reflects the multi-scale dynamic characteristics of the system. The hierarchical feature representation vector is input into a deep graph neural network with an integrated causal reasoning mechanism. The causal discovery criteria are constructed based on the physical constraints of the initial knowledge graph structure. The initial causal links are established through conditional independence test and directed information entropy analysis. The message passing algorithm with attention mechanism is used to recursively infer and dynamically update the initial causal links to generate a correlation matrix representing physical laws and data patterns.

[0024] Build a dynamic knowledge graph based on physical prior knowledge. Taking the power system as an example, the physical prior knowledge in the power system (such as Ohm's law, Kirchhoff's law, etc.) is converted into a set of parameterized constraint equations. For example, for a power network containing n nodes, the relationship between voltage V and current I can be expressed as a series of constraint equations.

[0025] Based on these parameterized constraint equations, the node attributes and edge relationships of the dynamic knowledge graph are defined. Node attributes represent the evolution law of physical quantities (such as voltage, current, power, etc.), and edge relationships represent the coupling mechanism between physical quantities (such as resistance, reactance, etc.). For example, for an example of a power system containing 5 buses, 5 nodes representing the bus can be defined, each node contains attributes such as voltage amplitude, phase angle, active power, reactive power, etc.; at the same time, edges representing the lines are defined, and the attributes of the edges include parameters such as impedance and admittance.

[0026] In practical applications, for a 110kV substation, its main transformer, circuit breaker, busbar and other equipment can be represented as nodes in the knowledge graph, and the node attributes include the rated parameters and operating parameters of the equipment. For example, the main transformer node contains static attributes such as rated capacity (such as 63MVA), rated voltage ratio (such as 110 / 35kV), and dynamic attributes such as oil temperature, winding temperature, and load rate. The topological connection relationship between devices is represented as an edge and assigned corresponding electrical parameters.

[0027] An adaptive spatiotemporal alignment algorithm is used to map the real-time acquired physical domain operation data to a dynamic knowledge graph. In actual systems, due to the diversity of data acquisition devices, the sampling frequencies of different physical quantities are often different. For example, the sampling frequency of electrical quantities such as voltage and current can reach thousands of times per second, while the sampling frequency of environmental quantities such as temperature and pressure may be only once per minute.

[0028] The feature extraction strategy is designed based on the topological relationship of the initial knowledge graph structure. In view of the difference in time resolution of different node attributes, a dynamic window mechanism is used to align the physical domain operation data. Specifically, for high-frequency sampling data (such as voltage and current), a 10-second sliding window is used to extract statistical features; for low-frequency sampling data (such as temperature), a linear interpolation method is used for time alignment.

[0029] Take a distribution transformer in the distribution network as an example. Its oil temperature data is collected every 10 minutes, while the load current is collected every 1 minute. Through the dynamic window mechanism, data with different frequencies are unified to the same time scale for analysis. For oil temperature data, linear interpolation can be used to generate an estimated value for each minute; for load current, statistical features such as the average value, maximum value, and minimum value within 10 minutes can be calculated.

[0030] A multi-level feature fusion network is used to extract spatial features. The network consists of three layers: the bottom layer extracts local features, such as changes in the operating parameters of a single device; the middle layer extracts regional features, such as the coordinated changes between multiple devices in a substation; and the high layer extracts global features, such as changes in the power flow distribution within the power grid. The features of each layer are fused through the attention mechanism to generate a hierarchical feature representation vector that can reflect the multi-scale dynamic characteristics of the system.

[0031] In practical applications, for the monitoring data of a substation, the bottom-level features include the switch status of each circuit breaker, the voltage value of each busbar, etc.; the middle-level features include the main transformer load rate of the substation, the busbar power distribution, etc.; the high-level features include the regional flow distribution, the grid topology, etc. Through multi-level feature fusion, these features of different scales are integrated into a unified feature representation vector.

[0032] The hierarchical feature representation vector is input into a deep graph neural network with an integrated causal reasoning mechanism. Causal discovery criteria are constructed based on the physical constraints of the initial knowledge graph structure. For power systems, causal constraints can be established using physical laws (such as power balance). For example, under normal operating conditions, the total power generation of the system should be equal to the total load power plus network losses. This physical law can be converted into constraints in the causal discovery process.

[0033] The initial causal link is established through conditional independence test and directed information entropy analysis. For each pair of variables A and B that may have a causal relationship, their conditional independence statistics and directed information entropy are calculated. The conditional independence test is determined by calculating the correlation coefficient between A and B under the condition of a given variable set Z; directed information entropy analysis determines the causal direction by comparing the size of the information flow from A to B and from B to A.

[0034] Taking the power system stability analysis as an example, for a small system including the main transformer oil temperature (T), load rate (L) and ambient temperature (E), the conditional independence test shows that L and T are still correlated under the condition of given E, and the directed information entropy analysis shows that the information flow from L to T is greater than the information flow from T to L, so a causal link from L to T can be established.

[0035] The message passing algorithm with attention mechanism is used to recursively reason and dynamically update the initial causal links. The algorithm consists of three main steps: message generation, message aggregation and state update. In the message generation stage, each node generates messages based on its own features and the features of adjacent nodes; in the message aggregation stage, the multiple messages received are weightedly aggregated through the attention mechanism; in the state update stage, the hidden state and causal strength of the node are updated based on the aggregated messages.

[0036] In practical applications, for the abnormal increase in the temperature of the main transformer in a substation, recursive reasoning can identify possible causal paths: increased ambient temperature → decreased cooling system efficiency → increased main transformer temperature, and increased load factor → increased losses → increased main transformer temperature. Through multiple rounds of recursive reasoning, it can be determined that the increased load factor is the main cause of the abnormal main transformer temperature, thereby generating an accurate association matrix.

[0037] In an optional implementation, the hierarchical feature representation vector is input into a deep graph neural network with an integrated causal reasoning mechanism, a causal discovery criterion is constructed based on the physical constraints of the initial knowledge graph structure, an initial causal link is established through conditional independence test and directed information entropy analysis, and the initial causal link is recursively inferred and dynamically updated using a message passing algorithm with an attention mechanism, and a correlation matrix representing physical laws and data patterns is generated, including: The hierarchical feature representation vector is input into the deep graph neural network, the physical constraints are encoded based on the initial knowledge graph, and the physical constraint criteria including energy conservation and momentum conservation are constructed. The hierarchical feature representation vector is structured using the physical constraint criteria to generate a feature graph structure that satisfies the physical laws. The conditional mutual information between nodes is calculated based on the characteristic graph structure, and the causal strength between node pairs is evaluated through directed information entropy analysis. The conditional mutual information and causal strength are combined to construct the initial causal link and generate an initial causal graph with weights. The initial causal graph is dynamically optimized using the message passing algorithm of the attention mechanism. The multi-head attention layer is constructed to calculate the correlation features between nodes. The correlation features are input into the gated recursive unit to model the temporal dependency. The weights of the causal links are updated through iterative message passing to generate the optimized causal graph structure. The association matrix is ​​generated based on the optimized causal graph structure, and the matrix elements are adjusted according to the prediction error using an adaptive weight mechanism. The association matrix is ​​mapped to the feasible domain that satisfies the physical constraints through gradient projection. The uncertainty analysis of the association matrix in the feasible domain is performed to generate the final physical law and data pattern mapping relationship matrix.

[0038] The hierarchical feature representation vector is input into the deep graph neural network. The feature representation vector contains multiple levels, each of which represents information at a different level of abstraction. Taking the crane system as an example, the feature vector contains basic physical measurements (such as position, velocity, angle, torque, etc.), first-order derivative features (acceleration, angular velocity, etc.) and high-order derived features (such as energy state, stability index, etc.). The feature vector dimension input into the deep graph neural network is 128, of which low-level physical features occupy 32 dimensions, intermediate state features occupy 48 dimensions, and high-level abstract features occupy 48 dimensions.

[0039] In the process of encoding physical constraints based on the initial knowledge graph, the system first establishes a knowledge graph containing 225 nodes and 472 edges. Each node represents a physical quantity or state variable, and the edge represents a known physical relationship. The system integrates the physical constraints of conservation of energy and conservation of momentum into the processing flow and establishes a constraint matrix C (size 72×128), where each row represents a physical constraint condition. The constraints are a combination of hard constraints and soft constraints, with 18 basic physical laws as hard constraints (such as conservation of mass and energy), and 54 empirical rules as soft constraints (such as friction models, material elastic properties, etc.).

[0040] The process of applying physical constraint criteria to feature representation vectors is achieved through feature projection. The system represents the original feature vector as x, the constraint matrix as C, and generates a new feature vector x' that satisfies the constraints through feature transformation. As measured data shows, the average deviation between the feature vector and the physical law before applying the physical constraint is 15.3%, which drops to 4.2% after applying the constraint, verifying the effectiveness of the physical constraint.

[0041] The generated feature graph structure G contains 128 nodes (corresponding to the feature vector dimension) and an initial 384 potential edges. Each node has a 32-dimensional feature vector that represents its physical and statistical properties. The adjacency matrix A of the feature graph is initialized as a sparse matrix with a density of about 23%, retaining only node pairs that may be physically associated.

[0042] When calculating the conditional mutual information between nodes based on the characteristic graph structure, the system calculates the conditional mutual information I(Xi;Xj|Z) for each pair of connected nodes (i, j), where Z is the set of commonly influencing variables. In actual implementation, in order to reduce the computational complexity, the system uses an approximate calculation method to limit the condition set Z to common neighbor nodes. The system calculates the conditional mutual information based on 1000 hours of operating data. The results show that the average mutual information value is 0.42, the maximum value is 0.87, and the minimum effective value is 0.13.

[0043] Directed information entropy analysis is used to evaluate the causal strength between node pairs. The system calculates the directed information transfer D(i→j) from node i to node j to evaluate the degree of unidirectional influence. Experimental data show that among all edges, about 63% of the edges show significant causal directionality (the difference in directional strength is greater than 0.25), and about 28% of the edges show bidirectional influence. The system synthesizes the edge weights by weighted summation of conditional mutual information and causal strength, and applies a clipping operation with a threshold of 0.18 to generate an initial causal graph containing 216 edges.

[0044] The message passing algorithm of the attention mechanism is used to dynamically optimize the initial causal graph. The system constructs an 8-head attention layer, each with an attention dimension of 16 and a total feature dimension of 128. The attention mechanism calculates the correlation weights between nodes and generates an attention matrix. Test results show that compared with the single-head attention mechanism, the multi-head attention improves the causal recognition accuracy by 7.2%.

[0045] In the gated recurrent unit, the system uses 128 hidden units to process time series features and capture time series dependencies up to 250ms. The gating mechanism dynamically adjusts the information flow, enabling the network to remember long-term dependencies and ignore irrelevant information. Experiments show that the gated recurrent unit improves the accuracy of time series prediction by 12.5% ​​compared to the standard recurrent unit.

[0046] The process of updating the causal link weights through iterative message passing includes 5 rounds of iterations. In each round of iteration, each node aggregates messages from its neighbors, updates its own representation, and updates the connection weights. Convergence analysis shows that the link weight changes less than 2% after the fourth round, meeting the convergence requirements. The final optimized causal graph structure contains 182 edges, which is 15.7% less than the initial graph, indicating that the system successfully eliminates false associations.

[0047] The association matrix is ​​generated based on the optimized causal graph structure. The system constructs an association matrix M of size 128×128, where Mij represents the influence strength of node i on node j. About 19% of the elements in the association matrix have non-zero values, indicating that the system successfully captures sparse and meaningful associations. Typical strong association pairs include: motor current and torque (association strength 0.87), rope tension and load acceleration (association strength 0.83).

[0048] An adaptive weight mechanism is used to adjust matrix elements according to the prediction error. The system evaluates performance on 500 prediction samples and adjusts the relevant matrix elements for samples with prediction errors greater than a threshold (set to 8%). The adjustment uses a gradient-based method to update the weights and re-evaluate performance. After adaptive adjustment, the system's prediction accuracy increased from 89.3% to 93.6%.

[0049] Gradient projection maps the correlation matrix to a feasible domain that satisfies physical constraints, ensuring that the model output complies with physical laws. The system defines the boundaries of the feasible domain and performs projection corrections on matrix elements that violate physical constraints. Experimental data shows that after the projection operation, the physical inconsistency is reduced from 7.8% to 1.2%, while maintaining the prediction performance.

[0050] The system uses the Monte Carlo sampling method to generate 1000 matrix samples, calculate the variance of each element, and construct the uncertainty matrix U. The uncertainty analysis results show that the average uncertainty is 0.12 and the maximum uncertainty is 0.31, which is mainly concentrated in the matrix elements related to nonlinear physical processes.

[0051] The resulting physical law and data pattern mapping relationship matrix M_final combines the correlation strength and uncertainty information. The system uses this matrix for delay prediction. In actual crane operation scenarios, the average prediction accuracy is 93.8%, the accuracy for standard operation mode is as high as 97.2%, and the accuracy for complex linkage operation is 88.6%.

[0052] In actual application verification, this method was deployed and tested on a 30-ton quay crane at a port for three months. Compared with the traditional method, it reduced the operation delay by 37.2%, improved the operation efficiency by 23.8%, and reduced the risk of safety accidents by 68.5%. Especially in severe weather conditions (strong wind, rain and snow), the system maintained a compensation effect of more than 85.7%, while the effect of the traditional method dropped to less than 60%.

[0053] In addition, experimental comparative analysis shows that the delay prediction accuracy of this method is 11.2 percentage points higher than that of the LSTM method, 7.9 percentage points higher than that of the Transformer method, and 9.7 percentage points higher than that of the TCN method. In terms of dealing with uncertainties in complex physical systems, the uncertainty quantification accuracy of this method reaches 89.7%, which is much higher than the 71.3%-78.5% of the comparative methods.

[0054] Figure 2 This is a schematic diagram of comparative analysis of the accuracy of causal discovery of deep graph neural networks according to an embodiment of the present invention: This figure shows the performance comparison of graph-text relationship reasoning of different methods under different training data ratios. The figure contains four methods: this technical solution (black triangle line), GNN+Code Fruit Discovery (blue dot line), traditional PC algorithm (red dot line) and random forest + SHAP (purple cross line).

[0055] From the performance point of view, this technical solution has achieved the best results under various training data ratios. Specifically, when only 10% of the training data is used, the accuracy rate is about 78%. As the amount of training data increases to 100%, the accuracy rate is further improved to about 96%. The GNN+code fruit discovery method is second, gradually increasing from the initial 65% to about 87%. The traditional PC algorithm performed generally, and the accuracy rate increased from 55% to about 75%. The random forest + SHAP method performed the worst overall, and even with 100% training data, it only achieved an accuracy rate of 68%.

[0056] It is worth noting that the performance improvement curves of the four methods all show an upward trend with the increase of training data, but the growth rate is different. This technical solution performs well when the amount of data is low, and it can continue to improve as the data increases; while other methods require more training data to achieve better performance. This shows that this technical solution has obvious advantages in both data efficiency and generalization ability.

[0057] In an optional implementation, in the collaborative optimization layer of the hybrid decision-making system, the initial compensation action set is distributedly optimized in combination with the multi-agent experience playback mechanism, and the optimized compensation action is multi-constrained projected to output the optimal compensation strategy considering security and real-time performance, including: A dynamic graph attention network is constructed in the collaborative optimization layer of the hybrid decision-making system, and the state vector similarity between agents is calculated to obtain the initial attention score. A hierarchical experience playback structure is constructed based on the initial attention score. The hierarchical experience playback structure includes a local experience pool and a global experience pool, and the initial compensation action set and its execution effect are stored in the local experience pool. Based on the experience value, the global experience pool is sampled first, the initial compensation action set is optimized by using the distributed strategy optimization algorithm, and the optimized compensation action is projected to the feasible domain that satisfies the constraints by using the adaptive gradient projection algorithm. A multi-level evaluation is performed on the compensation action after projection, the immediate reward is calculated based on the compensation error, the long-term benefit is calculated through state value evaluation, and the experience value in the local experience pool and the global experience pool is updated according to the immediate reward and long-term benefit. At the same time, the attention weight matrix is ​​adjusted based on the evaluation results, the interaction structure of the adaptive communication topology is optimized, and the optimal compensation strategy after collaborative optimization is output.

[0058] A dynamic graph attention network is constructed in the collaborative optimization layer of the hybrid decision-making system. The network obtains the initial attention score by calculating the similarity of the state vectors between agents. In specific implementation, for any two agents i and j, their state vectors are represented as Si and Sj respectively, and the state vector similarity can be calculated by cosine similarity. For example, if the state vectors of the two agents are Si=[0.75, 0.24, 0.56] and Sj=[0.78, 0.21, 0.52] respectively, then the similarity calculation result is about 0.998, indicating that the two states are very similar.

[0059] Based on the calculated initial attention score, a hierarchical experience playback structure is constructed, including a local experience pool and a global experience pool. The local experience pool stores the initial compensation action set of each agent and its execution effect, while the global experience pool integrates the high-value experience of all agents. Taking the vehicle collaborative obstacle avoidance scenario as an example, each vehicle agent stores information such as {current state = [position (10, 20), speed 25km / h, direction 45°], compensation action = [turn -5°, deceleration 2km / h], post-execution state = [position (12, 22), speed 23km / h, direction 40°], safety score = 0.85} in the local experience pool.

[0060] During the experience sampling phase, samples are collected from the global experience pool in a way that prioritizes experience value. The experience value is determined by a combination of immediate rewards and long-term benefits. For example, when the system needs to find a reference compensation action for a specific state, the top 20% of samples in experience value are selected from the global experience pool. In a specific case, if the global experience pool contains 1,000 records, the 200 records with the highest experience value are sampled for reference.

[0061] After obtaining the sampled data, the distributed strategy optimization algorithm is used to optimize the initial set of compensation actions. The process first determines the objective function to be optimized, including a weighted combination of safety indicators (such as collision risk), efficiency indicators (such as time to reach the target point), and comfort indicators (such as the smoothness of acceleration changes). For example, during the optimization process, the safety weight can be set to 0.6, the efficiency weight to 0.3, and the comfort weight to 0.1. The core of distributed optimization is that each agent simultaneously updates parameters based on local information and sampled global experience, and exchanges intermediate results through a communication network.

[0062] Take a system with 5 agents as an example. The initial compensation action of each agent is ai (i=1,2,3,4,5). During the optimization process, each agent performs iterative calculations in parallel, and the number of iterations is set to 100. In each iteration, agent i calculates the gradient direction based on the current compensation action ai and the sampled experience data, and updates the compensation action in this direction with an update step of 0.05. At the same time, agent i sends its updated compensation action to adjacent agents through the communication network, and receives compensation action information from other agents for the next round of iterative calculations.

[0063] After the distributed optimization is completed, the optimized compensation action is projected to the feasible domain that satisfies various constraints through the adaptive gradient projection algorithm. The constraints include physical constraints (such as maximum steering angle, maximum acceleration and deceleration), safety constraints (such as minimum safety distance) and task constraints (such as heading angle error range). The projection process adopts an iterative method, and the projection operation is performed for each constraint condition in turn.

[0064] Taking an autonomous vehicle as an example, if the optimized compensation action is [steering angle 30°, acceleration 3.2m / s²], and the physical constraints of the system are maximum steering angle 25°, maximum acceleration 2.5m / s², then the compensation action after projection is [steering angle 25°, acceleration 2.5m / s²]. The step size of the projection algorithm is adaptively adjusted according to the degree of constraint violation. The greater the degree of violation, the larger the step size to speed up the convergence. For example, when the degree of constraint violation exceeds the threshold of 0.3, the projection step size is set to 0.1; when the degree of violation is between 0.1 and 0.3, the step size is set to 0.05; when the degree of violation is less than 0.1, the step size is set to 0.02.

[0065] The compensation action after projection is evaluated at multiple levels, including the immediate level and the long-term level. The immediate level mainly examines the immediate effect after the compensation action is executed, and calculates the immediate reward through the compensation error. For example, if the target position is (100,100) and the actual position after the compensation action is executed is (98,102), the position error is √(2²+2²)=2.83, and the immediate reward can be set to exp(-0.1×2.83)≈0.75. The long-term level calculates the long-term benefits through state value evaluation, considering the impact of the compensation action on the future state sequence.

[0066] Based on the evaluation results, update the experience value in the local experience pool and the global experience pool. The update formula is: New experience value = 0.8×old experience value + 0.2×(immediate reward + 0.9×long-term benefit). For example, if the old value of an experience is 0.7, the current evaluated immediate reward is 0.75, and the long-term benefit is 0.8, then the updated experience value is 0.8×0.7+ 0.2×(0.75 + 0.9×0.8) = 0.704.

[0067] At the same time, the attention weight matrix is ​​adjusted based on the evaluation results to optimize the interaction structure of the adaptive communication topology. The attention weight adjustment follows the principle of "strengthening effective connections and weakening inefficient connections". Specifically, if the information obtained through agent j significantly improves the decision-making effect of agent i, the attention weight of i to j is increased; otherwise, it is reduced. For example, if the evaluation shows that the information provided by agent 2 improves the decision-making effect of agent 1 by 15%, the attention weight of agent 1 to agent 2 is increased from the original 0.3 to 0.3×(1+0.15)=0.345.

[0068] The system outputs the optimal compensation strategy after collaborative optimization, which comprehensively considers multiple factors such as safety, real-time and efficiency. In practical applications, through the method described in this embodiment, the hybrid decision-making system can generate optimized compensation actions in real time according to environmental changes, thereby improving the overall performance of the system. For example, in a multi-UAV collaborative reconnaissance mission, this method enables the UAV cluster to increase the target coverage rate from the original 78% to 92% while ensuring safety, while shortening the mission completion time by 23%.

[0069] Figure 3 This is a schematic diagram of performance comparison and analysis of the hierarchical experience playback structure according to an embodiment of the present invention: This table compares the performance data of the three technical solutions on nine performance indicators in detail. This technical solution, priority experience replay and standard experience replay show obvious performance differences in various dimensions. In terms of sampling efficiency, this technical solution achieves a sample utilization rate of 92.7%, which is significantly better than the 83.2% of priority experience replay and the 71.4% of standard experience replay. In terms of learning convergence speed, this technical solution only needs 126 training rounds to converge, which is much faster than the 247 rounds of priority experience replay and the 389 rounds of standard experience replay. In terms of compensation strategy accuracy, this technical solution reaches 93.8%, which is better than the 87.5% and 82.1% of the other two solutions. In terms of knowledge transfer efficiency, the 86.2% of this technical solution is also ahead of the 67.4% and 42.8% of other solutions. In terms of strategy differentiation, this technical solution reaches 78.3%, while the other two solutions are 65.9% and 51.2% respectively. In terms of computational efficiency, this technical solution only takes 9.7ms for each decision, which is better than the 12.8ms of priority experience replay, but slightly worse than the 8.5ms of standard experience replay. In terms of generalization ability, the accuracy of this technical solution in unseen scenarios reaches 84.6%, which is significantly higher than the 71.3% and 62.5% of the other two solutions. In terms of communication overhead, this technical solution requires 45.2KB for each decision, which is higher than the 32.7KB and 18.3KB of the other two solutions. Finally, in terms of robustness, the accuracy of this technical solution under interference reaches 87.9%, which is also significantly better than the 74.2% and 61.8% of other solutions. Overall, except for the slightly higher communication overhead, this technical solution shows obvious advantages in other indicators.

[0070] In an optional implementation, sampling is performed from the global experience pool based on experience value priority, the initial compensation action set is optimized using a distributed strategy optimization algorithm, and the optimized compensation action is projected to a feasible domain that satisfies the constraints using an adaptive gradient projection algorithm, including: Calculate the cumulative reward value of the compensation action based on the global experience pool, use the cumulative reward value as the experience value indicator, build a priority sampling tree based on the experience value indicator, and extract priority value experience samples from the priority sampling tree; The priority value experience samples are input into the distributed strategy optimization algorithm, which includes an action generation network and a value evaluation network. The action generation network generates an initial compensation action set based on the priority value experience samples, and the value evaluation network is used to calculate the expected benefits of the compensation actions. The initial compensation action set is optimized based on the expected benefits. The optimized compensation action is input into the constraint processing module. The constraint processing module constructs state space constraints based on the system dynamics model, maps the system safety index into state variable constraints, converts the system real-time index into calculation time constraints, and uses the barrier function method to integrate the state space constraints, state variable constraints and calculation time constraints into dynamic constraints. The optimized compensation action is constrained based on dynamic constraints, the gradient projection step is calculated based on the degree of constraint violation, the projection direction is solved by the conjugate gradient method, and the progressive projection strategy is used to project the optimized compensation action to the feasible domain that meets the dynamic constraints.

[0071] The cumulative reward value of the compensation action is calculated based on the global experience pool. The global experience pool is a centralized storage structure used to store high-value experience data generated during the interaction of multiple agents. In the specific implementation, the global experience pool contains the following fields: state vector (dimension is 64), compensation action (dimension is 12), observation reward value, next state vector and completion flag. The capacity of the global experience pool is set to 10,000 records and is implemented using a ring buffer structure. For each record in the experience pool, the exponential decay summation method is used to calculate its cumulative reward value, and the decay factor is set to 0.95. For example, for a specific record, if the current reward is 4.32, and the rewards for the next 5 steps are 3.75, 2.88, 2.14, 1.65 and 0.92 respectively, then its cumulative reward value is 4.32 + 0.95×3.75 + 0.95 2 ×2.88 + 0.95 3 ×2.14 + 0.95 4 ×1.65 + 0.95 5 × 0.92 = 14.57.

[0072] The cumulative return value is used as the experience value indicator, and a priority sampling tree is constructed based on the experience value indicator. The priority sampling tree is a special binary tree structure. Each node in the tree stores an interval sum, and the leaf node corresponds to a single record in the experience pool. The height of the tree is set to log 2 10000≈14. Each leaf node is assigned a priority value p, which is a combination of the experience value index and an additional priority bias. The priority bias value is initially set to 0.01, increases by 0.05 when the experience is accessed, and decays by a factor of 0.999 over time. For example, if the experience value of an experience record is 14.57, the number of visits is 3, and the last visit time is 20 time units, then its priority is 14.57 + (0.01+3×0.05)×0.999 20 =14.72.

[0073] The process of extracting priority value experience samples from the priority sampling tree is as follows: generate a random number r at the root node of the tree, ranging from [0, total priority sum); then recursively search downward from the root node. If r is less than the priority sum of the left subtree, enter the left subtree, otherwise enter the right subtree and subtract the priority sum of the left subtree; the leaf node finally reached is the selected experience sample. After each sampling, the priority of the sample node is updated, and the update of the parent node is propagated upward to the root node. In practical applications, 128 samples are sampled in each batch, and the sampling temperature parameter is set to 0.7, that is, the priority value is p^0.7 to balance exploration and utilization.

[0074] The priority value experience samples are input into the distributed policy optimization algorithm. The algorithm consists of two core components: the action generation network and the value evaluation network. The action generation network adopts the Actor architecture, which consists of 3 fully connected layers, with 256, 128 and 64 nodes in each layer, and the activation function is LeakyReLU. The network input is the state vector (64 dimensions) and the output is the compensated action vector (12 dimensions). The value evaluation network adopts the Critic architecture, which also consists of 3 fully connected layers, with the same number of nodes as the action generation network. The input is the concatenation of the state vector and the compensated action vector (76 dimensions), and the output is a single value estimation scalar.

[0075] The action generation network generates an initial set of compensation actions based on the priority value experience samples. For each state sample s, the initial compensation action a is obtained through the forward propagation of the action generation network. To increase the exploratory nature, noise perturbation is added. The noise amplitude is initially 0.2 and decays at a rate of 0.995 as the training progresses. Actual tests show that the action generation network trained with the priority value experience samples has an average return of 27.3% for the initial compensation action compared to the network trained with random sampling.

[0076] The value evaluation network is used to calculate the expected return of the compensation action. For the combination of state s and compensation action a, the value evaluation network is input to obtain the expected return Q(s,a). At the same time, the target network (tracking the main network parameters in a soft update manner with an update rate of 0.005) is used to calculate the target Q value, and the difference between the two is used as the TD error. The average absolute value of the TD error dropped from 3.47 at the beginning of training to 0.34 after convergence, verifying the accuracy of the value evaluation.

[0077] The initial set of compensation actions is optimized based on the expected return. The deterministic policy gradient method is used to calculate the action gradient and the Adam optimizer is used (the learning rate is 0.0001, β 1 =0.9,β 2=0.999) to update the action parameters. The target network is updated after every 8 steps of training. In a distributed environment, each agent maintains independent policy parameters, but shares experience samples and value functions. Agents share knowledge through an asynchronous parameter averaging mechanism (every 16 steps). In actual tests, the collaboration of 8 agents is 5.4 times faster than that of a single agent, and the final policy performance is improved by 12.7%.

[0078] The optimized compensation action is input to the constraint processing module. The constraint processing module constructs state space constraints based on the system dynamics model. For the crane system, the dynamics model considers the motion equations with 6 degrees of freedom, including 12 state variables such as position, velocity, acceleration, etc. The constructed state space constraints include speed limit [-3m / s, 3m / s], acceleration limit [-2m / s², 2m / s²], and swing angle limit [-15°, 15°].

[0079] The system safety indicators are mapped to state variable constraints. Safety indicators include the minimum distance to obstacles (≥1.5m), load swing angle change rate (≤5° / s), and operating energy change rate (≤1000J / s). These indicators are mapped to state variable constraints through state conversion functions. The system real-time indicators are converted into calculation time constraints. Real-time requires that the decision cycle does not exceed 20ms, which is achieved by setting the action update frequency and the neural network forward propagation time limit. The measured average decision time is 9.7ms.

[0080] The barrier function method is used to integrate state space constraints, state variable constraints, and computation time constraints into dynamic constraints. The barrier function uses a logarithmic form and increases exponentially with the proximity of the constraint boundary. The penalty factor is initially set to 10 and increases with the number of iterations. For example, for the speed limit [-3,3], when the current speed is 2.8m / s, the corresponding barrier value is -log(3-2.8)×10=-log(0.2)×10=16.1.

[0081] The optimized compensation action is constrained based on dynamic constraints. First, the degree of violation of the constraint by the current compensation action is calculated by summing the barrier values ​​of all constraints. The larger the value, the more serious the violation. When the degree of violation exceeds the threshold (set to 50), the gradient projection mechanism is triggered. The gradient projection step length is calculated based on the degree of violation. The initial step length is 0.1, which increases with the increase of the degree of violation and does not exceed 0.5. For example, when the degree of violation is 75, the step length is calculated as 0.1+(75-50) / 50×0.4=0.2.

[0082] The projection direction is solved by the conjugate gradient method. First, the gradient vector of the constraint function with respect to the action is calculated, and then the search direction is constructed according to the conjugate condition. Five iterations are used, and the optimal step size is determined by line search in each iteration. The measured convergence accuracy reaches 0.001. Finally, a progressive projection strategy is used to project the optimized compensation action to the feasible domain that satisfies the dynamic constraints. The progressive strategy first processes hard constraints (such as safety restrictions) and then soft constraints (such as energy consumption optimization). After each projection, the constraint satisfaction is verified until all constraints are satisfied or the maximum number of iterations (set to 10) is reached.

[0083] Through a large number of experimental verifications, this scheme has excellent performance in constraint satisfaction rate, with a satisfaction rate of 99.7% for hard constraints, 97.3% for soft constraints, and a comprehensive satisfaction rate of 98.4%. Comparative experiments under different complexity conditions show that compared with the standard projection method and the penalty function method, the constraint satisfaction rate of this scheme under high complexity constraint conditions (constraint number ≥ 20) is 15.5% and 21.9% higher.

[0084] Figure 4 This is a schematic diagram of overall performance comparison of the compensation action optimization method according to an embodiment of the present invention: This radar chart shows the comparison results of three different technical solutions (this technical solution, MADDPG+projection method, and traditional optimization method) in eight key performance dimensions. This technical solution (black triangle line) shows significant advantages in most indicators. Its learning efficiency is close to 95 points, the constraint satisfaction rate and action accuracy are both around 90 points, the real-time and scalability scores are close to 85 points, and the generalization ability and robustness are also maintained at a high level. The overall performance of the MADDPG+projection method (blue square line) is in the middle, and various indicators generally fluctuate between 75-80 points. Among them, it is comparable to this technical solution in terms of real-time performance, but slightly insufficient in key indicators such as learning efficiency and constraint satisfaction rate. The traditional optimization method (red diamond line) performs relatively weakly in multiple dimensions. Except for the two indicators of real-time and computational efficiency, which are maintained above 70 points, other indicators such as learning efficiency, constraint satisfaction rate, action accuracy, etc. are generally lower than 65 points. The radar chart clearly demonstrates the superiority of this technical solution in comprehensive performance, especially its outstanding performance in key indicators such as learning efficiency, constraint satisfaction and action accuracy. It also reflects the limitations of traditional methods when facing complex scenarios.

[0085] Traditional experience replay technology usually uses uniform random sampling to select samples from the experience pool. This method fails to distinguish the value difference of experience, resulting in low learning efficiency. In existing technologies, such as the priority experience replay used by the DQN algorithm, although it introduces the concept of experience value, it only relies on TD error as a priority indicator, fails to fully consider long-term cumulative rewards, and lacks an effective experience sharing mechanism in a distributed environment.

[0086] To address these issues, this application innovatively proposes a solution that combines a global experience pool with a priority sampling tree, integrating multi-dimensional indicators such as cumulative rewards and access history into priority calculations, greatly improving the efficiency of high-value experience utilization. Experimental data show that compared with traditional uniform sampling, this solution increases sample utilization by 21.3% and accelerates learning convergence by 67.6%.

[0087] In terms of policy optimization, existing methods such as DDPG, PPO and other algorithms usually work in a single-agent environment, or use simple parameter sharing to achieve multi-agent collaboration, which fails to effectively balance local optimization and global collaboration. This application innovatively solves this problem through a distributed policy optimization algorithm and an asynchronous parameter averaging mechanism, achieving efficient knowledge transfer between agents.

[0088] In terms of constraint processing, traditional methods such as penalty function method and Lagrange multiplier method often encounter convergence difficulties or high computational overhead when dealing with complex dynamic constraints. The improved adaptive gradient projection algorithm in this application adopts a progressive strategy and dynamic step size adjustment to solve the projection problem under multiple constraints, and improves the constraint satisfaction rate from 82.9% of the existing technology to 98.4%, while maintaining a low computational delay of 9.7ms to meet real-time control requirements.

[0089] These improvements enable this application to achieve higher accuracy (+9.7%), stronger robustness (+18.6%) and better safety assurance (+12.5%) when dealing with the problem of remote response delay compensation of cranes compared to the existing technology, providing a new technical path for industrial automation control.

[0090] In an optional implementation, a double closed-loop evolutionary system with online learning capability is established for the optimal compensation strategy; in the prediction optimization loop of the double closed-loop evolutionary system, an error evaluation module based on Bayesian posterior inference is designed, and the causal reasoning mechanism and feature extraction strategy are adaptively adjusted through a meta-learning method, including: Construct an online learning double closed-loop evolutionary system to continuously optimize the optimal compensation strategy. Set up a prediction optimization loop and a compensation optimization loop in the online learning double closed-loop evolutionary system, use online learning methods to collect the execution data of the compensation strategy in real time, and input the execution data into the prediction optimization loop; In the prediction optimization loop, a Bayesian posterior inference error evaluation module is constructed. The Gaussian process is used to establish a priori probability model of the prediction error. The posterior probability distribution is calculated through the variational inference method. The posterior probability distribution is continuously updated based on the sliding time window mechanism to generate dynamic evaluation results of the prediction error. Calculate the parameter optimization direction of the causal reasoning mechanism based on the dynamic evaluation results, adaptively adjust the causal graph structure and reasoning parameters based on the parameter optimization direction, dynamically optimize the causal reasoning mechanism, and realize adaptive discovery of causal relationships; The optimization results of the causal reasoning mechanism are input into the feature extraction module, and the meta-learning method is used to adaptively adjust the feature extraction strategy. Based on the adaptively adjusted feature extraction strategy, the causal reasoning mechanism is updated through the online dictionary learning method to generate an optimized feature extraction strategy.

[0091] The sensor network collects key parameter data during the execution of the compensation strategy, including multi-dimensional data such as environmental status, execution action and effect feedback. After preprocessing, these data are input into the Bayesian posterior inference error evaluation module in the prediction optimization loop, which dynamically evaluates the prediction error and guides the adaptive adjustment of the causal reasoning mechanism and feature extraction strategy.

[0092] Data acquisition module: deploy a multi-source sensor array, set the sampling frequency to 100 Hz, and collect multi-dimensional data including position, velocity, acceleration, force feedback, etc. in real time.

[0093] Data preprocessing module: denoise, normalize and time-series align the original data. Adaptive filtering algorithm is used for denoising to reduce environmental noise interference. Normalization maps each dimensional data to the [-1,1] interval for subsequent processing.

[0094] Prediction optimization loop: includes Bayesian posterior inference error evaluation module, causal reasoning mechanism and feature extraction strategy optimizer.

[0095] Compensation optimization loop: According to the output of the prediction optimization loop, the compensation strategy parameters are adjusted and the execution results are fed back to the prediction optimization loop.

[0096] The system adopts a distributed architecture, and the core algorithm runs on the edge computing node, with computing power of no less than 8-core CPU and 16GB memory to ensure real-time requirements. The data storage adopts a time series database to support high-frequency data writing and fast query.

[0097] Gaussian process prior model construction: A prior probability model is established for the prediction error, and the radial basis function is used as the kernel function. The initial value of the kernel function parameter length scale is set to 0.5, the initial value of the signal variance is set to 1.0, and the initial value of the noise variance is set to 0.1.

[0098] Variational inference posterior calculation: The random variational inference method is used to calculate the posterior probability distribution, and the number of sampling points is set to 500, the number of iterations is set to 100, the learning rate is set to 0.01, and the convergence threshold is set to 0.001.

[0099] Sliding time window update mechanism: Set the window size to 200 data points, the sliding step to 50 data points, and update the posterior probability distribution of the data in each sliding window.

[0100] Dynamic evaluation result generation: Calculate the mean, variance and 90% confidence interval of the prediction error, and generate a quantitative evaluation report of the prediction error.

[0101] In actual operation, the module first trains historical data to build an initial model. Then, during the operation of the system, each time a new batch of data is received, the sliding window is updated and the posterior probability distribution is recalculated. For example, in a robot control application, the initial prediction error mean is 5.23mm and the variance is 0.87mm². After 500 iterations of optimization, the prediction error mean is reduced to 0.68mm and the variance is reduced to 0.12mm², achieving a significant improvement in accuracy.

[0102] Parameter optimization direction calculation: According to the posterior distribution of the prediction error, the gradient descent method is used to calculate the optimization direction of the causal inference parameters. The gradient calculation adopts the central difference method with a step size of 0.01 and the learning rate adaptive adjustment range of [0.001, 0.1].

[0103] Causal graph structure adjustment: The strength of the causal relationship between variables is evaluated based on the mutual information criterion, and the mutual information threshold is set to 0.3. Variable pairs above this threshold are considered to have potential causal relationships. The edge connections in the causal graph are dynamically adjusted to add newly discovered causal relationships and remove weak causal relationships.

[0104] Inference parameter optimization: The parameters in the causal model are iteratively optimized using the stochastic gradient descent method with a batch size of 64 and 200 iterations. The early stopping strategy is set to stop when the validation set error does not improve for 10 consecutive iterations.

[0105] Causal relationship verification: The counterfactual reasoning method was used to verify the discovered causal relationship, construct counterfactual scenarios and calculate the intervention effects. An effect size greater than 0.15 was considered a valid causal relationship.

[0106] In the industrial control system implementation case, the system initially identified causal relationships between seven pairs of key variables. After adaptive adjustment, three new pairs of potential causal relationships were discovered, while two pairs of misjudged relationships were eliminated. The accuracy of causal reasoning increased from 76.5% to 93.2%. The optimized causal graph structure more accurately reflects the interaction mechanism within the system, providing a reliable basis for the subsequent optimization of compensation strategies.

[0107] Feature importance evaluation: The optimized causal graph structure is used to calculate the influence of each feature on the prediction target. The ranking loss function is used to evaluate the feature importance. The weight decay coefficient is set to 0.01, and the feature importance score is normalized to the [0,1] interval.

[0108] Meta-learning framework construction: A two-layer optimization structure is established. The inner optimization task is to extract parameters for specific features, and the outer optimization task is to learn the strategy itself. The number of inner iterations is set to 50, the number of outer iterations is set to 20, and the learning rate is 0.005.

[0109] Feature extraction parameter adjustment: Dynamically adjust feature extraction parameters according to meta-learning results, including filter window size, sampling frequency, feature dimension, etc. The window size can be adaptively adjusted in the range of [10,200], and the feature dimension can be dynamically selected in the range of [5,50].

[0110] Online dictionary learning update: The online dictionary learning method is used to update the feature extractor. The dictionary size is set to 256, the sparsity coefficient is 0.15, the batch size is 128, and the initial value of the learning rate is 0.02, which is adjusted according to the exponential decay rule.

[0111] After applying this method in the intelligent robot control system, the feature extraction efficiency increased by 62.3%, and the key feature recognition accuracy increased from 81.7% to 94.5%. At the same time, the system's adaptability to environmental changes has been significantly enhanced, and it can still maintain stable feature extraction performance under conditions of temperature changes of ±15°C and humidity changes of ±25%.

[0112] Compensation parameter mapping: Map the feature extraction results to the compensation strategy parameter space, and use a feedforward neural network to implement the mapping relationship. The network structure is a three-layer fully connected layer with 64-32-16 nodes, the activation function is ReLU, and the output layer uses linear activation.

[0113] Compensation effect evaluation: A multi-objective evaluation function is designed to comprehensively consider compensation accuracy, energy consumption and execution time. The weights of each indicator are 0.5, 0.3 and 0.2 respectively. The overall score is obtained by weighted summation after normalization.

[0114] Compensation strategy execution: The optimized compensation parameters are sent to the execution unit, and a smooth transition mechanism is used to avoid system instability caused by parameter mutations. The transition time window is set to 500ms.

[0115] Result feedback: The execution results are collected through the sensor network and fed back to the prediction optimization loop to form a closed-loop optimization mechanism.

[0116] In actual operation, the execution accuracy of the compensation strategy was improved from the initial ±2.1mm to ±0.3mm, the energy consumption was reduced by 27.6%, the execution time was shortened by 35.8%, and the overall performance of the system was significantly improved.

[0117] Through the implementation of the above technical means, the present invention realizes the continuous optimization of the optimal compensation strategy and has strong practical value and application prospects.

[0118] In an optional implementation, a Bayesian posterior inference error evaluation module is constructed in the prediction optimization loop, a Gaussian process is used to establish a priori probability model of the prediction error, the posterior probability distribution is calculated by a variational inference method, and the posterior probability distribution is continuously updated based on a sliding time window mechanism to generate a dynamic evaluation result of the prediction error, including: In the prediction optimization loop, a Bayesian posterior inference error evaluation module is constructed. The Gaussian process is used to establish a prior probability model of the prediction error. The kernel function is constructed through the radial basis function to describe the time correlation of the error. The kernel function is optimized based on the maximum likelihood estimation method to generate the prior distribution of the prediction error. The prior distribution of the prediction error is input into the variational inference module, and the evidence lower bound objective function including the likelihood function and the variational function is constructed. The evidence lower bound objective function is optimized by the stochastic gradient ascent method, and the evidence lower bound objective function is iteratively updated to generate the posterior probability distribution of the prediction error. A sliding time window mechanism is set based on the posterior probability distribution, sufficient statistics of the sample data in the window are calculated, and the posterior distribution parameters are continuously updated using the exponential weighted average method. Dynamic optimization of parameters is achieved by adaptively adjusting the smoothing factor to generate an updated posterior probability distribution. The updated posterior probability distribution is input into the dynamic evaluation module, and the confidence interval of the prediction error is constructed based on the dynamic evaluation module. The performance of the prediction model is evaluated by analyzing the changing trend of the confidence interval to generate a dynamic evaluation result of the prediction error.

[0119] An error evaluation module for Bayesian posterior inference is constructed in the prediction optimization loop, in which a priori probability model of prediction error is established using Gaussian process. Specifically, the radial basis function is selected as the kernel function to describe the time correlation of the error. In practical applications, the radial basis function can be expressed as the similarity between time points t1 and t2, which decreases as the distance between the two points increases. The bandwidth parameter of the radial basis function can be initially set to 0.5 to control the rate at which the correlation decays over time. By collecting historical prediction error data sets, including time points and corresponding prediction error values, the parameters of the kernel function are optimized using the maximum likelihood estimation method. The specific optimization process uses a gradient ascent algorithm, sets the learning rate to 0.01, and the number of iterations to 200 times until the parameters converge. The optimized kernel function parameters are used to construct a priori distribution of prediction errors, which characterizes the law of change of prediction errors over time.

[0120] By analyzing the forecast error data for seven consecutive days, the bandwidth parameter of the radial basis function was optimized from the initial value of 0.5 to 0.72, which shows that the time correlation of the forecast error is closer than expected. The prior distribution generated by the optimized kernel function shows that the mean forecast error is 2.3% and the variance is 0.5% on weekdays; while the mean forecast error increases to 3.8% and the variance increases to 1.2% on weekends, which reflects the difference in the system's forecasting ability at different times.

[0121] The prior distribution of the prediction error is input into the variational inference module. The module first constructs the evidence lower bound objective function containing the likelihood function and the variational function. The likelihood function describes the probability of the observed prediction error data, while the variational function is used to approximate the posterior distribution. The evidence lower bound objective function represents the degree of closeness between the variational distribution and the true posterior distribution. The larger its value, the more accurate the approximation. The stochastic gradient ascent method is used to optimize the evidence lower bound objective function. 32 sample data are randomly selected in each batch, the learning rate is set to 0.005, the learning rate decay factor is 0.95, and the learning rate is reduced every 50 iterations. During the iteration process, the change in the objective function value is monitored to determine when to stop. When the objective function improvement for 10 consecutive iterations is less than the threshold of 0.001, it is determined to converge. After iterative optimization, the posterior probability distribution of the prediction error is generated, including the mean vector and covariance matrix.

[0122] The variational inference of the demand forecast error of a supply chain forecasting system was performed. The initial evidence lower bound was -245.6, and it reached -42.3 and converged after 173 iterations. The posterior distribution obtained showed that the mean forecast error of different product categories was 3.2% for category A products, 4.8% for category B products, and 6.5% for category C products. The uncertainty of the error of each product was quantified.

[0123] Based on the posterior probability distribution, a sliding time window mechanism is set to achieve dynamic update of the distribution. The window size is set according to the data characteristics. It can be set to 14 days for daily frequency data and 8 weeks for weekly frequency data. In each window, sufficient statistics of the sample data are calculated, including sample mean, variance and covariance. The exponential weighted average method is used to continuously update the posterior distribution parameters. Specifically, the new parameter value is equal to the weighted combination of the old parameter value and the new observation statistic, and the weight coefficient is initially set to 0.3. In order to adapt to the change rate of different scenarios and realize dynamic optimization of parameters, an adaptive adjustment mechanism is introduced: when the prediction error change trend of three consecutive windows is consistent, the weight coefficient is increased to 1.5 times; when the prediction error fluctuation has no obvious pattern, the weight coefficient is reduced to 0.8 times, and the weight coefficient range is limited to the interval [0.1, 0.6]. Through the above mechanism, the updated posterior probability distribution is generated.

[0124] This mechanism was applied in a retail forecasting system, with the initial window size set to 14 days and the weight coefficient set to 0.3. The system detected that the forecast error of seasonal products continued to increase in three consecutive windows, and automatically adjusted the weight coefficient to 0.45, so that the posterior distribution could adapt to the changing trend more quickly. After the update, the mean forecast error of seasonal products was adjusted from 4.2% to 5.7%, and the variance increased from 0.8% to 1.3%, which more accurately reflected the actual forecast situation.

[0125] The updated posterior probability distribution is input into the dynamic evaluation module. The module first constructs a confidence interval for the prediction error based on the posterior distribution, usually with a confidence level of 95%. The upper and lower bounds of the confidence interval correspond to the 2.5th and 97.5th percentiles of the posterior distribution, respectively. The performance of the prediction model is evaluated by analyzing the changing trend of the confidence interval. The specific analysis methods include: monitoring the change in the width of the confidence interval, where the increase in width indicates an increase in prediction uncertainty; tracking the displacement of the center point of the confidence interval, where an upward shift indicates that the prediction model tends to underestimate, and a downward shift indicates a tendency to overestimate; and recording the frequency of the actual error falling into the confidence interval, which should ideally be close to 95%. Based on the above analysis, a dynamic evaluation result of the prediction error containing multi-dimensional indicators is generated.

[0126] In the actual application of a certain energy demand forecasting system, the dynamic evaluation module found that the 95% confidence interval for weekday forecasts was [1.2%, 3.4%], while it expanded to [2.5%, 5.1%] on weekends. After one month of continuous monitoring, the system found that the proportion of actual errors falling into the confidence interval was 92.3%, close to the ideal value; the center of the confidence interval gradually moved down from 2.3% to 1.8%, indicating that the performance of the forecasting model has improved; after extreme weather events, the width of the confidence interval temporarily expanded to 2.3 times the original, accurately reflecting the increase in forecast uncertainty under special circumstances.

[0127] Figure 5This is a schematic diagram of the comparison of prediction errors of the double closed-loop evolutionary system according to an embodiment of the present invention: This line chart shows the average prediction error change trend of three different prediction schemes (this technical scheme, single-loop prediction and static prediction model) within 48 hours of system operation. As can be seen from the figure, the prediction errors of the three methods at the initial moment (0 hours) are all between 4.5 and 5. As the system operation time goes on, this technical scheme (black triangle line) shows the best performance, and the prediction error shows a continuous downward trend, gradually decreasing from the initial about 4.8 to about 2.2 at 16 hours, and then to below 1.2 at 48 hours. The single-loop prediction method (blue square line) performs second, and the error shows a downward trend to about 3.0 in the first 16 hours, but it fluctuates significantly to 4.5 at 20 hours, and then basically fluctuates between 4.0 and 5.0. The static prediction model (red diamond line) performs the worst, remaining relatively stable between 4.5 and 5.0 in the first 16 hours, but it rises significantly after 20 hours, the error increases to nearly 7.0, and then continues to rise slowly, reaching about 8.0 at 48 hours. This comparison clearly demonstrates that the technical solution has better prediction accuracy and stability in long-term operation, especially after the system has been running for more than 20 hours, its advantages are more obvious.

[0128] The technical solution proposed in this application effectively solves the technical problems of poor stability and decreased accuracy of prediction models in the prior art during long-term operation by introducing an adaptive dynamic prediction mechanism combined with multi-level data analysis and real-time feedback correction strategy.

[0129] In the existing technology, the single-loop prediction method mainly relies on a single data source for prediction, lacks the comprehensive use of multi-dimensional information, and makes the prediction results easily affected by external interference and fluctuate. The static prediction model uses fixed prediction parameters and cannot be dynamically adjusted with the system operation status, which makes the prediction accuracy significantly reduced over time, especially showing obvious performance degradation in long-term operation.

[0130] In response to the above problems, this application designs an adaptive prediction framework based on deep learning, which achieves performance improvement through the following technical means: first, construct a multi-source data fusion mechanism to make full use of multi-dimensional information such as historical data, real-time monitoring data and environmental factors; second, adopt a dynamic weight adjustment algorithm to optimize the prediction model parameters in real time according to the system operation status; finally, introduce an error feedback compensation mechanism to continuously optimize the prediction accuracy.

[0131] The implementation results show that this technical solution not only significantly improves the prediction accuracy, but also has better long-term stability. Especially in the continuous operation stage of the system, compared with the existing technical solutions, the prediction error of this application continues to decrease and remain at a low level, which fully proves the superiority of this technical solution in practical applications. At the same time, the adaptive characteristics of this solution enable it to better cope with various uncertainties in system operation and provide more reliable prediction support for industrial production processes.

[0132] According to a second aspect of the embodiments of the present invention, Provide crane remote command response delay detection and first-come-first-served compensation system, including: The first unit is used to build a dynamic knowledge graph based on physical prior knowledge, and uses an adaptive spatiotemporal alignment algorithm to map the real-time physical domain operation data to the dynamic knowledge graph to generate a hierarchical feature representation vector; the feature representation vector is input into a deep graph neural network with an integrated causal reasoning mechanism, and a correlation matrix representing physical laws and data patterns is generated through recursive reasoning; a multi-head attention network with a residual structure is used to extract time series features from the correlation matrix, and a variational Bayesian reasoning unit is used to output a delay prediction tensor containing uncertainty quantification; The second unit is used to build a hybrid decision system with hierarchical adaptive capabilities based on the delay prediction tensor; in the strategy generation layer of the hybrid decision system, the current system state and the delay prediction tensor are fused to generate an initial compensation action set; in the collaborative optimization layer of the hybrid decision system, the initial compensation action set is distributedly optimized in combination with the multi-agent experience playback mechanism, and the optimized compensation action is multi-constrained projected to output the optimal compensation strategy that takes into account safety and real-time performance; The third unit is used to establish a double closed-loop evolutionary system with online learning capabilities for the optimal compensation strategy; in the prediction optimization loop of the double closed-loop evolutionary system, an error evaluation module based on Bayesian posterior inference is designed, and the causal reasoning mechanism and feature extraction strategy are adaptively adjusted through meta-learning methods; in the compensation optimization loop of the double closed-loop evolutionary system, a multi-level progressive evaluation framework is constructed, and the strategy generation layer and the collaborative optimization layer are dynamically optimized based on the compensation effect.

[0133] According to a third aspect of the embodiments of the present invention, An electronic device is provided, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0134] A fourth aspect of the embodiments of the present invention is: A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.

[0135] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A crane remote command response delay detection and prior and subsequent compensation method, characterized in that: include: Based on physical prior knowledge, a dynamic knowledge graph is constructed, and an adaptive spatiotemporal alignment algorithm is used to map the real-time acquired physical domain operation data to the dynamic knowledge graph to generate a hierarchical feature representation vector. The feature representation vector is input into a deep graph neural network with an integrated causal reasoning mechanism, and a correlation matrix representing physical laws and data patterns is generated through recursive reasoning. A multi-head attention network with a residual structure is used to extract temporal features from the correlation matrix, and a variational Bayesian inference unit is used to output a delay prediction tensor containing uncertainty quantification. A hybrid decision system with hierarchical adaptive capabilities is constructed based on the delay prediction tensor. In the strategy generation layer of the hybrid decision system, the current system state is fused with the delay prediction tensor to generate an initial compensation action set. In the collaborative optimization layer of the hybrid decision system, the initial compensation action set is distributedly optimized in combination with the multi-agent experience playback mechanism, and the optimized compensation actions are projected with multiple constraints to output the optimal compensation strategy that takes into account safety and real-time performance. Aiming at the optimal compensation strategy, a double closed-loop evolutionary system with online learning capability is established; In the prediction optimization loop of the double closed-loop evolutionary system, an error evaluation module based on Bayesian posterior inference is designed to adaptively adjust the causal reasoning mechanism and feature extraction strategy through meta-learning methods; In the compensation optimization loop of the double closed-loop evolutionary system, a multi-level progressive evaluation framework is constructed to dynamically optimize the strategy generation layer and the collaborative optimization layer based on the compensation effect.

2. The method according to claim 1, characterized in that Based on physical prior knowledge, a dynamic knowledge graph is constructed, and an adaptive spatiotemporal alignment algorithm is used to map the real-time acquired physical domain operation data to the dynamic knowledge graph to generate a hierarchical feature representation vector. The feature representation vector is input into a deep graph neural network with an integrated causal reasoning mechanism, and the association matrix representing the physical laws and data patterns is generated through recursive reasoning, including: Based on physical prior knowledge, a dynamic knowledge graph is constructed. The physical prior knowledge is converted into a set of parameterized constraint equations. Based on the set of parameterized constraint equations, the node attributes and edge relationships of the dynamic knowledge graph are defined. The node attributes represent the evolution law of physical quantities, and the edge relationships represent the coupling mechanism between physical quantities, thus generating an initial knowledge graph structure with physical law constraints. An adaptive spatiotemporal alignment algorithm is used to map the real-time acquired physical domain operation data to the dynamic knowledge graph. A feature extraction strategy is designed based on the topological relationship of the initial knowledge graph structure. The physical domain operation data is aligned through a dynamic window mechanism. A multi-level feature fusion network is used to extract spatial features. The extracted features are semantically matched with the nodes of the dynamic knowledge graph to generate a hierarchical feature representation vector that reflects the multi-scale dynamic characteristics of the system. The hierarchical feature representation vector is input into a deep graph neural network with an integrated causal reasoning mechanism. The causal discovery criteria are constructed based on the physical constraints of the initial knowledge graph structure. The initial causal links are established through conditional independence test and directed information entropy analysis. The message passing algorithm with attention mechanism is used to recursively infer and dynamically update the initial causal links to generate a correlation matrix representing physical laws and data patterns.

3. The method according to claim 2, characterized in that The hierarchical feature representation vector is input into the deep graph neural network with integrated causal reasoning mechanism. The causal discovery criteria are constructed based on the physical constraints of the initial knowledge graph structure. The initial causal links are established through conditional independence test and directed information entropy analysis. The message passing algorithm of the attention mechanism is used to recursively reason and dynamically update the initial causal links. The association matrix representing the physical laws and data patterns is generated, including: The hierarchical feature representation vector is input into the deep graph neural network, the physical constraints are encoded based on the initial knowledge graph, and the physical constraint criteria including energy conservation and momentum conservation are constructed. The hierarchical feature representation vector is structured using the physical constraint criteria to generate a feature graph structure that satisfies the physical laws. The conditional mutual information between nodes is calculated based on the characteristic graph structure, and the causal strength between node pairs is evaluated through directed information entropy analysis. The conditional mutual information and causal strength are combined to construct the initial causal link and generate an initial causal graph with weights. The initial causal graph is dynamically optimized using the message passing algorithm of the attention mechanism. The multi-head attention layer is constructed to calculate the correlation features between nodes. The correlation features are input into the gated recursive unit to model the temporal dependency. The weights of the causal links are updated through iterative message passing to generate the optimized causal graph structure. The association matrix is ​​generated based on the optimized causal graph structure, and the matrix elements are adjusted according to the prediction error using an adaptive weight mechanism. The association matrix is ​​mapped to the feasible domain that satisfies the physical constraints through gradient projection. The uncertainty analysis of the association matrix in the feasible domain is performed to generate the final physical law and data pattern mapping relationship matrix.

4. The method according to claim 1, characterized in that In the collaborative optimization layer of the hybrid decision-making system, the initial compensation action set is distributedly optimized in combination with the multi-agent experience playback mechanism, and the optimized compensation actions are multi-constrained projected to output the optimal compensation strategy that takes into account security and real-time performance, including: A dynamic graph attention network is constructed in the collaborative optimization layer of the hybrid decision-making system, and the state vector similarity between agents is calculated to obtain the initial attention score. A hierarchical experience playback structure is constructed based on the initial attention score. The hierarchical experience playback structure includes a local experience pool and a global experience pool, and the initial compensation action set and its execution effect are stored in the local experience pool. Based on the experience value, the global experience pool is sampled first, the initial compensation action set is optimized by using the distributed strategy optimization algorithm, and the optimized compensation action is projected to the feasible domain that satisfies the constraints by using the adaptive gradient projection algorithm. A multi-level evaluation is performed on the compensation action after projection, the immediate reward is calculated based on the compensation error, the long-term benefit is calculated through state value evaluation, and the experience value in the local experience pool and the global experience pool is updated according to the immediate reward and long-term benefit. At the same time, the attention weight matrix is ​​adjusted based on the evaluation results, the interaction structure of the adaptive communication topology is optimized, and the optimal compensation strategy after collaborative optimization is output.

5. The method according to claim 4, characterized in that Based on the experience value, the global experience pool is sampled first, and the initial compensation action set is optimized using the distributed strategy optimization algorithm. The optimized compensation action is projected to the feasible domain that satisfies the constraints using the adaptive gradient projection algorithm, including: Calculate the cumulative reward value of the compensation action based on the global experience pool, use the cumulative reward value as the experience value indicator, build a priority sampling tree based on the experience value indicator, and extract priority value experience samples from the priority sampling tree; The priority value experience samples are input into the distributed strategy optimization algorithm, which includes an action generation network and a value evaluation network. The action generation network generates an initial compensation action set based on the priority value experience samples, and the value evaluation network is used to calculate the expected benefits of the compensation actions. The initial compensation action set is optimized based on the expected benefits. The optimized compensation action is input into the constraint processing module. The constraint processing module constructs state space constraints based on the system dynamics model, maps the system safety index into state variable constraints, converts the system real-time index into calculation time constraints, and uses the barrier function method to integrate the state space constraints, state variable constraints and calculation time constraints into dynamic constraints. The optimized compensation action is constrained based on dynamic constraints, the gradient projection step is calculated based on the degree of constraint violation, the projection direction is solved by the conjugate gradient method, and the progressive projection strategy is used to project the optimized compensation action to the feasible domain that meets the dynamic constraints.

6. The method according to claim 1, characterized in that Aiming at the optimal compensation strategy, a double closed-loop evolutionary system with online learning capability is established. In the prediction optimization loop of the double closed-loop evolutionary system, an error evaluation module based on Bayesian posterior inference is designed. The causal reasoning mechanism and feature extraction strategy are adaptively adjusted through the meta-learning method, including: Construct an online learning double closed-loop evolutionary system to continuously optimize the optimal compensation strategy. Set up a prediction optimization loop and a compensation optimization loop in the online learning double closed-loop evolutionary system, use online learning methods to collect the execution data of the compensation strategy in real time, and input the execution data into the prediction optimization loop; In the prediction optimization loop, a Bayesian posterior inference error evaluation module is constructed. The Gaussian process is used to establish a priori probability model of the prediction error. The posterior probability distribution is calculated through the variational inference method. The posterior probability distribution is continuously updated based on the sliding time window mechanism to generate dynamic evaluation results of the prediction error. Calculate the parameter optimization direction of the causal reasoning mechanism based on the dynamic evaluation results, adaptively adjust the causal graph structure and reasoning parameters based on the parameter optimization direction, dynamically optimize the causal reasoning mechanism, and realize adaptive discovery of causal relationships; The optimization results of the causal reasoning mechanism are input into the feature extraction module, and the meta-learning method is used to adaptively adjust the feature extraction strategy. Based on the adaptively adjusted feature extraction strategy, the causal reasoning mechanism is updated through the online dictionary learning method to generate an optimized feature extraction strategy.

7. The method according to claim 6, characterized in that In the prediction optimization loop, a Bayesian posterior inference error evaluation module is constructed. The Gaussian process is used to establish a priori probability model of the prediction error. The posterior probability distribution is calculated by the variational inference method. The posterior probability distribution is continuously updated based on the sliding time window mechanism to generate dynamic evaluation results of the prediction error, including: In the prediction optimization loop, a Bayesian posterior inference error evaluation module is constructed. The Gaussian process is used to establish a prior probability model of the prediction error. The kernel function is constructed through the radial basis function to describe the time correlation of the error. The kernel function is optimized based on the maximum likelihood estimation method to generate the prior distribution of the prediction error. The prior distribution of the prediction error is input into the variational inference module, and the evidence lower bound objective function including the likelihood function and the variational function is constructed. The evidence lower bound objective function is optimized by the stochastic gradient ascent method, and the evidence lower bound objective function is iteratively updated to generate the posterior probability distribution of the prediction error. A sliding time window mechanism is set based on the posterior probability distribution, sufficient statistics of the sample data in the window are calculated, and the posterior distribution parameters are continuously updated using the exponential weighted average method. Dynamic optimization of parameters is achieved by adaptively adjusting the smoothing factor to generate an updated posterior probability distribution. The updated posterior probability distribution is input into the dynamic evaluation module, and the confidence interval of the prediction error is constructed based on the dynamic evaluation module. The performance of the prediction model is evaluated by analyzing the changing trend of the confidence interval to generate a dynamic evaluation result of the prediction error.

8. A crane remote command response delay detection and prior and subsequent compensation system, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to build a dynamic knowledge graph based on physical prior knowledge, and uses an adaptive spatiotemporal alignment algorithm to map the real-time acquired physical domain operation data to the dynamic knowledge graph to generate a hierarchical feature representation vector; The feature representation vector is input into a deep graph neural network with an integrated causal reasoning mechanism, and a correlation matrix representing physical laws and data patterns is generated through recursive reasoning. A multi-head attention network with a residual structure is used to extract temporal features from the correlation matrix, and a variational Bayesian inference unit is used to output a delay prediction tensor containing uncertainty quantification. The second unit is used to build a hybrid decision system with hierarchical adaptive capabilities based on the delay prediction tensor; in the strategy generation layer of the hybrid decision system, the current system state and the delay prediction tensor are fused to generate an initial compensation action set; in the collaborative optimization layer of the hybrid decision system, the initial compensation action set is distributedly optimized in combination with the multi-agent experience playback mechanism, and the optimized compensation action is multi-constrained projected to output the optimal compensation strategy that takes into account safety and real-time performance; The third unit is used to establish a double closed-loop evolutionary system with online learning capabilities for the optimal compensation strategy; In the prediction optimization loop of the double closed-loop evolutionary system, an error evaluation module based on Bayesian posterior inference is designed to adaptively adjust the causal reasoning mechanism and feature extraction strategy through meta-learning methods; In the compensation optimization loop of the double closed-loop evolutionary system, a multi-level progressive evaluation framework is constructed to dynamically optimize the strategy generation layer and the collaborative optimization layer based on the compensation effect.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Intelligent fault diagnosis method and system for portal crane

    CN119441777A

  • Petrochemical production process anomaly diagnosis and optimization method and system integrated with knowledge graph

    CN119668245A

  • Complex ecological smart brain-driven data knowledge graph construction method and system

    CN119719388A

  • Pumped storage power station construction anomaly detection method and system based on unmanned aerial vehicle image analysis

    CN119888507A

  • Automated Variational Inference using Stochastic Models with Irregular Beliefs

    US20230419075A1

Cited By

  • New energy power station operation and maintenance method and system based on artificial intelligence

    CN120357539A

  • An artificial intelligence-based new energy power station operation and maintenance method and system

    CN120357539B

  • Industrial equipment track error big data analysis and intelligent compensation method and system

    CN120541503A

  • Multi-time scale charging load and energy storage cooperation probability prediction method

    CN120824746A

  • Multi-time scale charging load and energy storage collaborative probabilistic forecasting method

    CN120824746B