Multi-agent collaborative management and control method and system based on brain-like fusion architecture
By employing a multi-agent collaborative management and control method based on a brain-like fusion architecture, the problem of spatiotemporal fusion and cross-timescale management of multi-source risk information in complex industrial production has been solved. This enables long-term continuous safe production and emergency management, improves the system's collaborative decision-making capabilities and the adaptability of control strategies, and has significant social and economic benefits.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 塔盾信息技术(上海)有限公司
- Filing Date
- 2026-04-08
- Publication Date
- 2026-05-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing safety management technologies struggle to achieve spatiotemporal fusion modeling of multi-source risk information in complex industrial production. They lack continuous reasoning and management across time scales, and there is a lack of effective collaborative decision-making and linkage among multiple roles and systems. Control strategies are not adaptable enough, making it difficult to achieve long-term continuous safe production and emergency management.
A multi-agent collaborative management and control method based on a brain-like fusion architecture is adopted. By constructing a unified intelligent collaborative architecture for safety production and emergency management, a brain-like fusion cognitive mechanism, a long-task continuous reasoning mechanism, a multi-agent collaborative decision-making mechanism, and an execution-feedback-self-learning closed-loop mechanism are introduced to achieve long-term continuous perception, dynamic reasoning, and adaptive optimization of the safety production operation status, risk evolution process, and emergency response behavior.
It improves the temporal consistency and fusion reliability of multi-source safety production and emergency management information, enhances the stability and robustness of safety situation awareness, realizes long-term continuous reasoning and full-process management of accident risk and emergency response, improves the adaptability of control strategies and the overall consistency of collaborative emergency response, and has good engineering adaptability and promotion and application value.
Smart Images

Figure CN122022189A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial safety production technology, and more specifically to a multi-agent collaborative management and control method and system based on a brain-like fusion architecture. Background Technology
[0002] As industrial production scale continues to expand and production system complexity continues to increase, high-risk industrial sectors such as energy, chemical, metallurgy, mining, manufacturing, and industrial parks face increasingly complex safety risk management and emergency response challenges during production operations, placing higher demands on safety management and risk control technologies. During industrial production operations, multiple factors, including equipment status, process parameters, personnel behavior, environmental conditions, and external disturbances, collectively influence the system's operational state. These factors are coupled in both temporal and spatial topological dimensions, resulting in accident risks exhibiting multi-source, multi-modal, and dynamically evolving characteristics.
[0003] However, existing safety management technologies typically rely on single monitoring indicators or short-term window analysis, lacking the ability to integrate and model multi-source risk information across time and space, and also making it difficult to continuously reason about the entire process of accident incubation, risk evolution, and emergency response. Existing technologies mainly focus on monitoring and early warning, information integration, and decision support, but still have certain technical limitations in complex industrial safety and emergency management scenarios, specifically: (1) Multi-source safety production and emergency management information is scattered across different systems, lacking a unified fusion modeling mechanism, making it difficult to form an overall safety situation awareness; (2) Accident risks and emergency events have long-term evolutionary characteristics, and existing technologies are difficult to support continuous reasoning and management across time scales; (3) There is a lack of effective collaborative decision-making and linkage mechanisms among multiple roles and systems, and complex emergency tasks rely on manual coordination; (4) Control strategies are not adaptable enough, making it difficult to achieve continuous optimization and evolution in long-term operation and multiple event handling processes.
[0004] Therefore, how to achieve long-term continuous perception, dynamic reasoning, collaborative decision-making, and adaptive optimization of the operational status of safe production, risk evolution process, and emergency response behavior, and thus realize long-term continuous scientific control in complex safe production and emergency management scenarios, is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] To overcome or at least partially solve the above problems, this invention provides a multi-agent collaborative management and control method and system based on a brain-like fusion architecture. By constructing a unified intelligent collaborative architecture for safe production and emergency management, and introducing a brain-like fusion cognitive mechanism, a long-task continuous reasoning mechanism, a multi-agent collaborative decision-making mechanism, and an execution-feedback-self-learning closed-loop mechanism under this architecture, it realizes long-term continuous perception, dynamic reasoning, collaborative decision-making, and adaptive optimization of safe production operation status, risk evolution process, and emergency response behavior, thereby achieving long-term continuous collaborative management and adaptive optimization in complex safe production and emergency management scenarios.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] In a first aspect, embodiments of the present invention provide a multi-agent collaborative management method based on a brain-like fusion architecture, comprising: Acquire multi-source data related to safety production and emergency management of the target industrial system and perform time-series alignment to obtain time-aligned multi-source data; Based on the time-aligned multi-source data, evidence representations from each data source are obtained, and a security situation awareness vector is obtained based on evidence conflict assessment and dynamic credibility adjustment mechanisms. Based on the security situation cognition vector, a risk field state model is constructed, and combined with dynamic risk thresholds, the risk evolution results are obtained. Based on the risk evolution results, a multi-agent collaborative decision-making model is constructed, and the optimal collaborative strategy for each agent is generated under safety constraints. The optimal collaborative strategy is transformed and executed to obtain the execution result. The effectiveness of the strategy execution is evaluated based on the execution results, and feedback and self-learning optimization are performed to maintain the optimal collaborative strategy.
[0008] In one embodiment, the method for acquiring time-aligned multi-source data is as follows: Based on the multi-source data, cleaning, outlier removal and format standardization are performed, and data identifiers and timestamp information are uniformly generated to obtain structured multi-source data; Based on the structured multi-source data, extract event anchors with business significance; Based on the event anchor points, obtain the optimal time mapping function for each data source relative to the reference time axis; Based on the optimal time mapping function, the timestamps of each data source are mapped to a unified reference time axis to achieve time axis correction of multi-source data and obtain the time-aligned multi-source data.
[0009] In one embodiment, the multi-source data includes: industrial control system operation data, production management and safety management system data, environmental and risk monitoring data, and personnel, equipment and emergency resource status data; The business-significant event anchors include: equipment start / stop events, threshold exceedance events, alarm triggering events, operation status change events, or combinations thereof.
[0010] In one embodiment, the method for obtaining the security situation awareness vector is as follows: Based on the aforementioned time-aligned multi-source data, a brain-like evidence fusion mechanism is used for feature modeling, which is then mapped to the original evidence representation corresponding to each data source. The dynamic reliability of each data source is obtained based on data quality, degree of evidence conflict, and time residual. The original evidence representation is weighted and modified based on the dynamic reliability to obtain the modified evidence support. The multimodal evidence conflict degree is obtained based on the modified evidence support degree; Based on the multimodal evidence conflict degree, the fused evidence quality function is obtained through evidence fusion operation; The evidence quality function is mapped to the security situation perception vector of the target industrial system at the current moment.
[0011] In one embodiment, the method for obtaining the risk evolution result is as follows: A risk field state model of the target industrial system is constructed based on the security situation awareness vector. Based on the risk field state model, the risk state of each node in the target industrial system is inferred over a long continuous time period to obtain the risk field potential energy. Based on the risk field potential energy, the dynamic risk threshold that the target industrial system can withstand is determined through probabilistic security constraints; The risk field potential energy and the dynamic risk threshold together constitute the risk evolution result.
[0012] In one embodiment, the optimal coordination strategy is obtained as follows: Based on the risk evolution results, a multi-agent collaborative decision-making model is constructed, consisting of multiple agents, a global coordination unit, and a security constraint projection mechanism. Based on the multi-agent collaborative decision-making model, a two-layer collaborative optimization and security constraint projection mechanism is adopted, and the upper-layer coordinator generates a global coordination strategy vector. Under security constraints, each agent in the lower layer obtains its local optimal policy based on the global coordination policy vector. The optimal coordination strategy is obtained by aligning the local optimal strategy with the global coordination strategy vector.
[0013] In one embodiment, the multi-agent collaborative decision-making model includes: a perceptual agent, a reasoning agent, a decision-making agent, an execution agent, and a learning agent; The sensing agent is used to receive information about the current security status of the system and environmental status. The reasoning agent is used to perform risk trend reasoning based on risk evolution results and security situation information; The decision-making agent is used to generate candidate security control strategies based on the reasoning results; The execution agent is used to convert candidate security control policies into specific control instructions and execute them; The learning agent is used to update and optimize the strategy parameters based on the execution results and system feedback; The aforementioned intelligent agents make collaborative decisions under the constraints of the safety constraint projection mechanism to ensure that the generated control strategy meets the system safety constraints.
[0014] In one embodiment, the method for obtaining the execution result is as follows: The optimal collaborative strategy is converted into executable control instructions and sent to the target industrial system for execution. The execution result is obtained by real-time collection of equipment status changes and system operation information during the execution process.
[0015] In one embodiment, maintaining the optimal coordination strategy specifically includes: Based on the execution results, obtain the causal attributable execution deviation index; A learning optimization objective is constructed based on the causal attributability performance bias index and the dynamic risk threshold. Based on the learning optimization objective, it is converted into Lagrange form to obtain the Lagrange objective function; Based on the Lagrange objective function, the penalty coefficients for probabilistic security constraints are adaptively updated, so that the target industrial system increases the penalty intensity for security constraints when approaching the risk boundary, and automatically converges to the updated optimal cooperative strategy that satisfies the security constraints.
[0016] In a second aspect, embodiments of the present invention provide a multi-agent collaborative management and control system based on a brain-like fusion architecture, used to execute a multi-agent collaborative management and control method based on a brain-like fusion architecture as described in any of the first aspects, comprising: a multi-source data processing module, a brain-like fusion cognition module, a task risk reasoning module, a multi-agent collaborative decision-making module, a collaborative execution module, and a feedback optimization module. The multi-source data processing module is used to acquire multi-source data related to the safety production and emergency management of the target industrial system and perform time-series alignment to obtain time-aligned multi-source data. The brain-like fusion cognitive module is used to obtain evidence representations from each data source based on the time-aligned multi-source data, and to obtain a security situation cognitive vector based on evidence conflict assessment and dynamic credibility adjustment mechanism. The task risk reasoning module is used to construct a risk field state model based on the security situation cognition vector, and combine it with dynamic risk thresholds to obtain risk evolution results. The multi-agent collaborative decision-making module is used to construct a multi-agent collaborative decision-making model based on the risk evolution results, and generate the optimal collaborative strategy for each agent under safety constraints. The collaborative execution module is used to transform and execute based on the optimal collaborative strategy to obtain the execution result; The feedback optimization module is used to evaluate the effectiveness of the strategy execution based on the execution results, and to perform feedback and self-learning optimization to maintain the optimal collaborative strategy.
[0017] As can be seen from the above technical solution, compared with the prior art, the present invention discloses a multi-agent collaborative management method and system based on a brain-like fusion architecture, which has the following beneficial effects: 1. Improved the temporal consistency and fusion reliability of multi-source safety production and emergency management information: This invention addresses issues such as inconsistent sampling frequencies of multi-source sensing data, system clock drift, and asynchronous event triggering by introducing an optimal time mapping mechanism based on event anchors. This solution addresses these problems from both system architecture and computer science perspectives, enabling the fusion processing of multi-source safety production and emergency management information under a unified time reference. Compared to traditional methods relying on fixed time windows or simple interpolation alignment, this invention maintains the continuity and physical rationality of the time-series mapping in scenarios with dense multi-source events and strong noise interference, making the correlation analysis of cross-system safety events more stable and reliable. In engineering simulations and multi-scenario operational verification, the method described in this invention shows a significant improvement in the stability and consistency of cross-system event matching, maintaining stable output under high-density event and complex interference conditions, providing a more reliable data foundation for subsequent risk inference.
[0018] 2. Enhance the stability and robustness of security situation awareness in complex scenarios: This invention constructs a brain-like fusion cognitive mechanism with the ability to suppress evidence conflicts and dynamically adjust credibility. This mechanism can automatically reduce the impact of abnormal information on overall security situation assessment when there are inconsistencies, gaps, or quality fluctuations in multi-source sensory information. Compared with existing fusion methods based on simple weighting or single model output, this invention maintains the continuity and consistency of situational awareness results under complex operating conditions, multiple risk superpositions, and unstable perception scenarios, significantly reducing overall judgment bias caused by local information anomalies. In comparative analysis and multi-scenario verification, the system's misjudgment rate and unstable output phenomena show a significant downward trend, and the overall cognitive stability is significantly enhanced.
[0019] 3. Achieve continuous reasoning and full-process management of accident risk and emergency response: This invention, by introducing a long-task-driven risk evolution modeling and continuous reasoning mechanism, elevates the processes of safe production operation, risk prevention and control, and emergency response from traditional short-term monitoring and phased analysis to a continuous management process spanning shifts, phases, and even event cycles. This technology enables the system to continuously characterize and dynamically reason about the entire process of accident incubation, risk accumulation, emergency response, and recovery, avoiding the risk assessment distortion caused by fragmented analysis windows. In practical applications and simulation scenarios, this invention significantly enhances the ability to identify risk evolution trends in advance, allowing the system to trigger safety control or emergency preparedness measures earlier in the risk accumulation phase, providing a more ample time window for safety control and emergency preparedness.
[0020] 4. Construct risk control boundaries that dynamically change with the system's carrying capacity to improve the adaptability of management and control strategies: This invention introduces a resource-constrained dynamic risk boundary mechanism, freeing safety risk thresholds from being fixed or entirely dependent on human experience. Instead, it allows for dynamic adjustment based on available emergency resources and workload levels. This mechanism enables the system to automatically adjust the conservatism of control strategies when resource conditions change, risks overlap, or operating conditions switch, avoiding the problems of over-warning or under-warning in complex scenarios caused by traditional fixed threshold methods. Under conditions of multiple operating condition switching and resource changes, invalid warnings are significantly reduced, while frequent false alarms or missed risks caused by fixed thresholds are avoided, significantly improving the rationality and adaptability of the warning strategy.
[0021] 5. Enhance the overall consistency and reliability of multi-role and multi-system collaborative emergency response: This invention constructs a multi-agent collaborative decision-making mechanism under security constraints, enabling collaborative work among intelligent agents with different functions such as perception, reasoning, decision-making, execution, and learning, based on a unified understanding of the security situation. Compared to traditional centralized decision-making or human-led emergency management methods, this invention effectively reduces the risk of single-point decision failure and maintains consistency and stability of collaborative decision-making under conditions of multiple concurrent events, limited resources, and complex constraints. In simulation scenarios involving multi-department collaboration and multi-resource scheduling, both the overall emergency response efficiency and the degree of strategy coordination show a steady upward trend.
[0022] 6. Construct a closed loop of decision-making, execution, and feedback to achieve continuous self-learning and evolution of safety and emergency response strategies: This invention introduces a self-learning optimization mechanism based on execution feedback and causal attribution bias assessment, enabling continuous correction of brain-like fusion model parameters, long-task inference strategies, and multi-agent collaborative decision-making rules during system operation. This closed-loop mechanism allows the system to gradually converge to a more stable and secure control strategy during long-term operation and multiple emergency responses, avoiding the problems of strategy rigidity and high dependence on human experience in traditional systems. During continuous operation and multiple rounds of verification, the system's execution bias index showed a gradual decreasing trend, indicating that this invention possesses good long-term evolution capabilities and engineering stability.
[0023] 7. Possesses good engineering adaptability and application value, generating significant social and economic benefits: This invention employs a modular, model-driven system architecture design, enabling it to interface with industrial control systems, safety management systems, and emergency command systems without significant modifications to existing systems. It exhibits excellent engineering adaptability and scalability. By enhancing accident prevention capabilities, reducing the probability of major risks, and minimizing human intervention, this invention effectively reduces safety production and emergency management costs, minimizes accident losses, and improves production continuity and public safety assurance levels, demonstrating significant social and economic benefits.
[0024] 8. This invention, through the organic combination of brain-like fusion cognitive mechanism, long-task risk evolution modeling, multi-agent safe collaborative decision-making, and closed-loop self-learning optimization mechanism, significantly overcomes the shortcomings of existing technologies in multi-source information fusion, risk evolution characterization, and collaborative emergency response, and has outstanding technological advancement, engineering feasibility, and promotion and application value. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0026] Figure 1 This is a flowchart of a multi-agent collaborative management method based on a brain-like fusion architecture provided in an embodiment of the present invention.
[0027] Figure 2 This is a flowchart of the security situation awareness vector acquisition method provided in this embodiment of the invention.
[0028] Figure 3 This is a schematic diagram illustrating the synergistic relationship between risk evolution and dynamic risk threshold provided in an embodiment of the present invention.
[0029] Figure 4 This is a schematic diagram of a multi-agent collaborative management and control system based on a brain-like fusion architecture provided in an embodiment of the present invention. Detailed Implementation
[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0031] Example 1 To address the problems of existing safety production and emergency management technologies in complex industrial and park scenarios, such as fragmented multi-source information, difficulty in continuously depicting the evolution of risks and accidents, low efficiency of cross-stage emergency response task coordination, and difficulty in long-term stable optimization of control strategies, this invention constructs an overall technical system that integrates multi-source safety production and emergency management information, brain-like cognitive modeling, long-task-driven approaches, multi-agent collaborative decision-making, and closed-loop execution and self-learning optimization. This system enables long-term continuous intelligent control in complex safety production and emergency management scenarios. The specific solution is as follows: like Figure 1 As shown, this embodiment of the invention discloses a multi-agent collaborative management method based on a brain-like fusion architecture, including the following steps. For ease of description, these steps are numbered S1 to S6, and these numbers are not used to limit the sequential relationship between the various steps of this invention: S1 acquires multi-source data related to safety production and emergency management of the target industrial system and performs time-series alignment to obtain time-aligned multi-source data.
[0032] Furthermore, the method for acquiring time-aligned multi-source data is as follows: The data is cleaned, outliers are removed, and formats are standardized based on multi-source data. Data identifiers and timestamps are generated in a unified manner to obtain structured multi-source data. Extracting business-meaning event anchors from structured multi-source data; Obtain the optimal time mapping function for each data source relative to the reference time axis based on event anchors; By mapping the timestamps of each data source to a unified reference time axis using the optimal time mapping function, time axis correction of multi-source data is achieved, resulting in time-aligned multi-source data.
[0033] Furthermore, in this embodiment, the target industrial system is a system in a high-risk industrial scenario such as energy, chemical, metallurgy, mining, or industrial park.
[0034] Furthermore, the multi-source data includes: industrial control system operation data, production management and safety management system data, environmental and risk monitoring data, and personnel, equipment and emergency resource status data; Business-significant event anchors include: equipment start / stop events, threshold exceedance events, alarm triggering events, operation status change events, or combinations thereof.
[0035] Furthermore, the industrial control system operation data includes process parameters and status signals collected by the distributed control system (DCS) and programmable logic controller (PLC). Production management and safety management system data includes data from the Manufacturing Execution System (MES) and the Environmental Health and Safety Management System (EHS). Environmental and risk monitoring data include: gas concentration, temperature and humidity, pressure, and video surveillance; After performing the above preprocessing operations on the multi-source data, it is ensured that data from different systems and different sampling frequencies can be processed subsequently under a unified data structure and temporal semantic framework.
[0036] Furthermore, after completing unified identification and basic cleaning, to address the time inconsistency issue arising from multi-source data under different sampling frequencies and system clock conditions, this invention does not employ simple time interpolation or maximum-minimum difference alignment. Instead, it constructs a cross-source optimal time warp mapping based on the event anchor set, thereby obtaining the optimal time mapping function: ; in, Indicates the first m The optimal time mapping function of each data source relative to the reference time axis is used to map the timestamp of the data source to a unified reference time axis, thereby achieving time alignment of multi-source data. It is a monotonic continuous function. Represents all possible time mapping functions In the given context, find the optimal time mapping function that minimizes the objective function. K represents the number of event anchors extracted from multi-source data; Indicates the first in the reference timeline k Anchor time for each event; Indicates the first m In the data source and the first k The original timestamps corresponding to each event anchor point; Indicates the first m A time mapping function from a data source to a reference time axis; This represents the robust loss function, used to measure the deviation between the time of event anchor points from different data sources and to reduce the impact of anomalous events on the time alignment results; Represents time mapping function The first derivative; This represents the smoothing regularization weight parameter, used to constrain the continuity and monotonicity of the time mapping function; This indicates the time alignment error between event anchors, used to measure the degree of matching between events from different data sources on the reference timeline. This represents a smoothing constraint term for the time mapping function, used to limit the degree to which the time mapping function deviates from linear change, thereby ensuring the continuity and stability of the time mapping process.
[0037] Furthermore, by minimizing the time deviation between corresponding event anchors in different data sources and applying smoothing constraints to the time mapping function, the optimal time mapping function across data sources is obtained. This enables time axis correction of multi-source data on a unified reference time axis, resulting in time-aligned multi-source data with consistent timing. This avoids the cumulative time offset errors introduced by simple interpolation or fixed window methods under complex operating conditions, providing time-consistent data input for subsequent brain-like fusion cognition. The aligned timestamps are then rewritten into the unified industrial data bus buffer to form a standardized time-series data stream that can be directly called by the subsequent risk inference module.
[0038] S2 obtains evidence representations from various data sources based on time-aligned multi-source data, and obtains a security situation awareness vector based on evidence conflict assessment and dynamic credibility adjustment mechanisms.
[0039] Furthermore, such as Figure 2 As shown, the method for obtaining the security situation awareness vector is as follows: Based on time-aligned multi-source data, a brain-like evidence fusion mechanism is used for feature modeling, which is then mapped to the original evidence representation corresponding to each data source. The dynamic reliability of each data source is obtained based on data quality, degree of evidence conflict, and time residual. The original evidence representation is weighted and modified based on dynamic reliability to obtain the modified evidence support. The multimodal evidence conflict degree is obtained based on the revised evidence support degree; Based on the multimodal evidence conflict degree, the fused evidence quality function is obtained through evidence fusion calculation; The evidence quality function is used to map the target industrial system's security situation perception vector at the current moment.
[0040] Furthermore, dynamic reliability specifically refers to: ; in, Indicates the first m Data sources at time t The dynamic reliability is used to measure the credibility of the data source at the current moment and is used as an evidence correction weight. t Indicates the current system uptime; This represents the activation function, which maps the calculation results to the [0,1] interval to keep the dynamic reliability within a reasonable range; Indicates the first m Data sources at time t Data quality is used to reflect the integrity, stability, or reliability of the data source. Indicates the first m The degree of evidence conflict between a data source and other data sources is used to measure the degree of inconsistency between the evidence from that data source and other evidence. Indicates the first m Data sources at time t The time residual is used to represent the remaining time deviation between the data source timestamp and the reference time axis. , , These represent weighting coefficients, which are used to adjust the impact of data quality, degree of evidence conflict, and time residuals on the dynamic reliability calculation. m Indicates the data source number.
[0041] Furthermore, the level of evidence support is revised as follows: ; in, Indicates the th after dynamic reliability correction m A set of security propositions from various data sources A The degree of evidence support is the revised degree of evidence support. Indicates the first m Each data source, before merging, has a set of propositions. A The original evidence represents the degree of support that each data source provides for the set of security propositions; A It represents a set of security situation propositions, used to characterize a certain type of proposition or risk state in the security state of a system.
[0042] Furthermore, the degree of conflict of multimodal evidence is specifically defined as follows: ; in, Indicates time Multimodal evidence conflict degree is used to measure the degree of inconsistency between evidence from different data sources; Indicates the first A subset of evidence corresponding to each data source; This represents the summation of all combinations of multiple subsets of evidence whose intersection is an empty set. The empty set represents a conflicting set of different combinations of evidence that cannot be simultaneously established. This indicates the number of data sources involved in the fusion.
[0043] The above formula is used to measure the overall degree of conflict among multiple sources of evidence. The greater the degree of conflict, the stronger the inconsistency between different data sources.
[0044] Furthermore, the evidence quality function is specifically as follows: ; in, This represents the quality function of the fused evidence, i.e., the quality of the multi-source evidence after conflict correction and fusion calculation, for the set of security propositions. Overall support level; This indicates evidence fusion computation; This indicates that for all conditions satisfying the intersection of multiple subsets of evidence, the result is... Summing the combinations; This indicates that the evidence support from each data source involved in the fusion is multiplied together. Indicates the first The degree of evidence support for the set of propositions from each data source.
[0045] Furthermore, to facilitate subsequent risk reasoning and security decision-making, the fused evidence quality function will be... Further mapped to the system's security situation awareness vector at time tS ( t This is used to characterize the current overall security status of the system. ; in, Indicates the system in time Security situation awareness vector; This represents the situation mapping function, used to map fused evidence into security situation indicators; For example: ,in Indicates the first Risk indicators at any time t The situation value, This represents the number of dimensions in the security posture vector.
[0046] Furthermore, through the aforementioned brain-like evidence fusion and dynamic reliability correction mechanism, stable and reliable security situation cognition results can be obtained even when there are conflicts, uncertainties, or quality differences in multi-source information. This provides a unified data foundation for risk field evolution modeling, multi-agent collaborative decision-making, and security management in subsequent steps.
[0047] S3 constructs a risk field state model based on the security situation cognition vector and combines it with dynamic risk thresholds to obtain the risk evolution results.
[0048] Furthermore, the method for obtaining the risk evolution results is as follows: Constructing a risk field state model of the target industrial system based on security situation awareness vectors; Based on the risk field state model, the risk state of each node in the target industrial system is inferred over a long continuous time period to obtain the risk field potential energy. Based on the risk field potential energy, the dynamic risk threshold that the target industrial system can withstand is determined through probabilistic security constraints; Risk field potential energy and dynamic risk threshold together constitute the outcome of risk evolution.
[0049] Furthermore, this invention characterizes risk as a risk field potential energy that evolves over time and space, and its dynamic evolution model is the risk field state model: ; in, Indicates the elapsed time step The risk field state of the post-system; Indicates the system in time The distribution of risk field intensity, i.e., risk field potential energy; Indicates the time step, used to describe the time discrete interval for risk evolution calculation; Indicates the system in time t Security situation awareness vector; This represents a risk growth function driven by the security situation; This represents the risk diffusion adjustment coefficient, used to control the intensity of risk propagation in the system space; This represents the gradient operator, used to describe the rate of change of the risk field in space; This represents the divergence operator, used to describe the diffusion process of risk in the spatial structure of a system; This represents the risk diffusion coefficient matrix, used to describe the ability of risk to propagate across different spatial regions. Its value is related to the current security situation. related; Represents the gradient of the risk field; Indicates the first Weights of each disturbance factor; Indicates the first The time-varying function of each disturbance factor is used to describe the degree of impact of changes in the environment or equipment status on risk; Indicates the number of disturbance factors; Indicates the control measure number; Indicates the number of control measures; Indicates the first The control efficiency coefficient of each control measure; Indicates the first Each control measure in time The intensity or status of execution; Among them, situation-driven risk generation items This indicates the driving effect of the current security situation of the system on the growth of risk. When the risk index in the security situation vector increases, this item will increase the intensity of system risk. Risk diffusion item This is used to describe the propagation process of risk in the spatial structure of a system, where the diffusion coefficient matrix... Related to the current security situation; Perturbation Amplification Term It is used to describe the amplifying effect of environmental factors or changes in equipment status on the risk evolution process; Risk control and suppression items This indicates the effect of security control measures or emergency response actions on reducing the intensity of system risk.
[0050] Furthermore, the potential energy of the risk field This corresponds to the node risk intensity distribution matrix on the spatial topology of industrial installations, where each node is mapped one-to-one with an actual physical device or work area, used to characterize the risk level of different devices, areas, or work nodes. Through... The dynamic updates can continuously track the propagation, diffusion and evolution of risks in the system, and provide basic state inputs for subsequent risk boundary calculation, collaborative decision-making and emergency response strategy generation.
[0051] Furthermore, to avoid using fixed or empirical thresholds, this invention further constructs an acceptable risk boundary based on resource constraints: ; in, Indicates a dynamic risk threshold; This represents the candidate risk threshold to be optimized. This represents finding the maximum acceptable risk threshold while satisfying probabilistic safety constraints. This represents a probability function used to describe the probability that the risk exceeds a certain threshold; This represents the predicted risk field value of the system at a future time. This represents the upper bound of the allowed risk probability, used to constrain the probability of a risk exceeding a threshold.
[0052] Furthermore, the upper bound of the allowed risk probability for: ; in, This indicates the upper limit of the basic risk probability. In terms of available emergency resource intensity, To adjust the risk threshold dynamically according to the system's carrying capacity, the workload level is adjusted.
[0053] Furthermore, Used to determine the current security situation vector Potential energy of risk field To predict the future evolution; Based on the predicted risks, the acceptable dynamic risk threshold of the system is determined through probabilistic safety constraints. This provides a risk boundary for subsequent safety control and emergency decision-making.
[0054] Dynamic risk threshold This represents the maximum acceptable risk level of the system under current emergency resource capabilities and operational load conditions. Through dynamic risk boundary calculation, the risk tolerance range can be adjusted in real time according to the system's operating status and resource capabilities, thereby providing safety constraints for subsequent collaborative decision-making and control strategies. Risk field potential energy As inputs to the risk boundary calculation model, these variables represent the system risk evolution state and are used to assess whether future risks may exceed the system's acceptable risk boundary. It also provides risk constraints for subsequent multi-agent collaborative decision-making. It enables continuous modeling and reasoning of the entire process of accident incubation, risk accumulation, and emergency response, avoiding the decision-making fragmentation caused by analysis based solely on short time windows.
[0055] Furthermore, such as Figure 3 As shown, the security situation perception vector As input, it drives the operation of the risk field state model; the risk field state model integrates risk generation terms, risk diffusion terms, disturbance amplification terms, and risk control and suppression terms to continuously evolve the system risk state and obtain the risk field potential energy. In the process of calculating the dynamic risk threshold, the dynamic risk threshold is calculated based on the predicted value of the risk field potential energy, emergency resource capacity, and operational load level. .
[0056] By analyzing the potential energy of the risk field With dynamic risk threshold Comparisons are made to form a risk constraint judgment result; when Not exceeding When the system is in a safe state; when Exceed When the time comes, the collaborative control strategy is triggered, and collaborative control constraints are output for subsequent multi-agent collaborative decision-making and control execution.
[0057] S4 constructs a multi-agent collaborative decision-making model based on risk evolution results, and generates the optimal collaborative strategy for each agent under safety constraints.
[0058] Furthermore, the method for obtaining the optimal collaborative strategy is as follows: Based on the risk evolution results, a multi-agent collaborative decision-making model is constructed, consisting of multiple agents, a global coordination unit, and a safety constraint projection mechanism. Based on the multi-agent collaborative decision-making model, a two-layer collaborative optimization and security constraint projection mechanism is adopted, and the upper-layer coordinator generates a global coordination strategy vector. Under security constraints, each agent in the lower layer obtains its local optimal policy based on the global coordination policy vector. The optimal cooperative strategy is obtained by aligning the local optimal strategy with the global coordination strategy vector.
[0059] Furthermore, a multi-agent collaborative decision-making model is constructed based on the risk evolution results, mapping the tasks of safe production operation, risk prevention and control and emergency response into a collaborative optimization problem among multiple decision-making agents; through global collaborative strategy solution and local strategy optimization of each agent, joint decision-making that meets safety constraints is achieved.
[0060] Furthermore, the multi-agent collaborative decision-making model includes: a perceptual agent, a reasoning agent, a decision-making agent, an execution agent, and a learning agent; The sensing agent is used to receive information about the system's current security status and environmental conditions. Inference agent, used to infer risk trends based on risk evolution results and security situation information; A decision-making agent is used to generate candidate security control strategies based on the reasoning results. An executive agent is used to translate candidate security control policies into specific control instructions and execute them. Learning agents are used to update and optimize policy parameters based on execution results and system feedback; The aforementioned intelligent agents make collaborative decisions under the constraints of the safety constraint projection mechanism to ensure that the generated control strategy meets the system safety constraints.
[0061] Furthermore, to avoid the limitations of traditional consistency constraints, this invention adopts a two-layer collaborative optimization and security constraint projection mechanism: the upper-layer coordinator is responsible for generating the global coordination policy vector, and each agent in the lower layer solves its own local optimal policy under security constraints.
[0062] Furthermore, the global coordination strategy vector is: ; in, Indicates time t The global coordination strategy vector; This represents the global collaborative control variable to be solved; This represents the global control variable that minimizes the objective function through optimization. Indicates the number of agents participating in collaborative decision-making; Indicates the agent's index number; Indicates the first i The weight coefficients of each agent are used to represent the importance of that agent in global collaborative decision-making. Indicates the first i An intelligent agent at time t The generated local decision-making strategy; This represents a safety constraint projection operator, used to project agent policies onto a feasible policy space that satisfies safety constraints, ensuring that all collaborative decisions meet safety and compliance boundary requirements; Represents the set of security constraints; This represents the L2 norm, used to measure the distance between two policy vectors.
[0063] Furthermore, the set of security constraints C It consists of industrial safety operation standard boundaries, equipment operation allowable range, resource scheduling constraint matrix, and emergency response process rules.
[0064] Furthermore, after obtaining the global coordination policy vector, the lower-level agents operate within the set of security constraints. C The internal strategy for finding local optima has the following computational model: ; in, Indicates the first i An intelligent agent at time t The optimal local strategy; Indicates the first i Candidate strategies for each agent; Indicates the first i Local decision cost function of an agent; This represents the coordination and consistency weighting coefficient, used to balance the relationship between local optima and global coordination. This model is used to solve the local optimal policy of each agent under safety constraints, and ensures that it is consistent with the global cooperative policy.
[0065] S5 performs the transformation and execution based on the optimal collaborative strategy, and obtains the execution result.
[0066] Furthermore, the method for obtaining the execution result is as follows: The optimal collaborative strategy is converted into executable control commands and sent to the target industrial system for execution. The actual execution result is obtained by collecting real-time data on changes in device status and system operation information during the execution process.
[0067] Furthermore, in this embodiment, instructions are issued to at least one of the industrial control system, emergency command system, resource scheduling system, or safety management system in the target industrial system.
[0068] S6 evaluates the effectiveness of strategy execution based on the execution results, and performs feedback and self-learning optimization to maintain the optimal collaborative strategy.
[0069] Furthermore, maintaining the optimal collaborative strategy specifically includes: Based on the execution results, obtain causal attributable execution deviation indicators; A learning optimization objective is constructed based on causal attributability performance bias indicators and dynamic risk thresholds; Based on the learning optimization objective being transformed into Lagrange form, the Lagrange objective function is obtained. The optimization is based on the Lagrange objective function, and the penalty coefficient of the probabilistic security constraint is adaptively updated so that the penalty intensity of the security constraint is increased when the target industrial system approaches the risk boundary, and automatically converges to the updated optimal cooperative strategy that satisfies the security constraint.
[0070] Furthermore, to distinguish the true sources of execution deviation, evaluate the effectiveness of the optimal collaborative strategy in actual execution, and provide a basis for evaluating execution performance for subsequent feedback learning, this invention introduces a causal attributable execution deviation index: ; in, The causal attribution deviation index is used to measure the degree of deviation between the results of the implementation of a decision-making strategy and the theoretically optimal results. Indicates the system in time The actual observation status; This indicates that the optimal strategy is being executed. The system under action at time t The counterfactual expected state, that is, the state that the system should theoretically reach under the meaning of causal intervention; This represents the mathematical expectation operator.
[0071] Furthermore, based on the execution results and the causal attributable execution deviation index as inputs, the brain-like fusion model, long-task reasoning strategy, and multi-agent collaborative mechanism are self-learned and optimized.
[0072] Furthermore, the causal attribution performance deviation index As an important feedback signal in the reinforcement learning optimization process, the learning optimization objective adopts a reinforcement optimization form under probabilistic safety constraints: ; in, This is a model parameter vector used to represent the learnable parameters of brain-like fusion models, multi-agent collaborative models, and long-task inference models; This is a resource consumption cost function used to represent the cost of using emergency resources or control resources; This is a resource consumption weighting coefficient used to adjust the trade-off between execution deviation and resource cost; The learning optimization objective is used to update the system model parameters under probabilistic safety constraints in order to improve the safety and execution effectiveness of the decision-making strategy.
[0073] Furthermore, the Lagrange form of this learning optimization objective is: ; in, Represent the Lagrange objective function; The adaptive penalty coefficients (Lagrange multipliers) representing probabilistic safety constraints are used to dynamically adjust the penalty intensity for constraint violations based on how close the system is to the risk boundary. By optimizing the Lagrange function, reinforcement learning policy updates under probabilistic safety constraints can be achieved; by adjusting the adaptive penalty coefficient... The adaptive updates enable the system to increase the penalty intensity of safety constraints when approaching the risk boundary, thereby automatically converging to a more conservative and safer collaborative strategy. Through the above optimization process, the model parameters are updated based on the execution feedback and causal deviation evaluation results, enabling the system to gradually converge to a safer and more efficient collaborative control strategy during long-term operation. This allows the system to gradually optimize decision quality during long-term operation and multiple emergency responses, possessing continuous self-learning and evolution capabilities, and maintaining the optimal collaborative strategy.
[0074] Furthermore, the above steps S1-S6 are connected through an industrial safety data bus to form a closed-loop operation, and are transmitted between each step through a safety situation awareness representation and a risk field state matrix, thereby realizing long-term multi-agent collaborative control in safety production and emergency management scenarios.
[0075] Furthermore, compared with traditional safety production and emergency management technologies based on fixed thresholds, single models, or centralized platforms, this invention solves the problem of multi-source information conflict through a brain-like fusion cognitive mechanism, achieves full-process risk management through a long-task reasoning mechanism, and achieves multi-role collaborative handling through a multi-agent collaborative decision-making mechanism, significantly improving the system's stability, adaptability, and long-term operation capability in complex scenarios.
[0076] Example 2 like Figure 4 As shown, based on the same inventive concept, this embodiment of the invention also provides a multi-agent collaborative management and control system based on a brain-like fusion architecture, including: a multi-source data processing module, a brain-like fusion cognition module, a task risk reasoning module, a multi-agent collaborative decision-making module, a collaborative execution module, and a feedback optimization module; The multi-source data processing module is used to acquire multi-source data related to the safety production and emergency management of the target industrial system and perform time-series alignment to obtain time-aligned multi-source data. The brain-like fusion cognitive module is used to obtain evidence representations from various data sources based on time-aligned multi-source data, and to obtain a security situation cognitive vector based on evidence conflict assessment and dynamic credibility adjustment mechanism. The task risk reasoning module is used to construct a risk field state model based on the security situation cognition vector, and combine it with dynamic risk thresholds to obtain the risk evolution results. The multi-agent collaborative decision-making module is used to construct a multi-agent collaborative decision-making model based on risk evolution results and generate the optimal collaborative strategy for each agent under safety constraints. The collaborative execution module is used to transform and execute based on the optimal collaborative strategy to obtain the execution result. The feedback optimization module is used to evaluate the effectiveness of strategy execution based on the execution results, and to perform feedback and self-learning optimization to maintain the optimal collaborative strategy.
[0077] Furthermore, the brain-like fusion cognitive module can be implemented using different multimodal fusion hierarchical structures, evidence conflict modeling methods, or dynamic credibility adjustment strategies; the multi-agent collaborative decision-making module can adopt a centralized, distributed, or hybrid collaborative architecture, and the number of agents, functional divisions, and collaborative relationships can be adjusted according to the application scale; in the task risk reasoning module, the task time span, risk evolution granularity, reasoning step size, and risk threshold constraint strategies can be configured according to the needs of safe production and emergency management; the collaborative execution module can be interface adapted and protocol extended according to different industrial control systems, emergency command systems, or industry standards. All the above alternative implementation methods do not depart from the core technical concept of this invention, which is based on brain-like fusion, long-task reasoning, and multi-agent collaboration, and should be considered to fall within the protection scope of this invention.
[0078] Furthermore, in this embodiment, the functional implementation methods of each functional module correspond one-to-one with the methods described above, and will not be repeated here.
[0079] Example 3 Based on the same inventive concept, the present invention also provides an electronic device, which includes a processor and a memory. The memory stores instructions, which are loaded and executed by the processor to implement a multi-agent collaborative management method based on a brain-like fusion architecture as described in Embodiment 1.
[0080] Based on the same inventive concept, the present invention also provides a computer device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; When the processor executes the program stored in the memory, it can implement a multi-agent collaborative management method based on a brain-like fusion architecture, as shown in Example 1.
[0081] The electronic device may include a processor, a communications interface, a memory, and a communication bus, wherein the processor, communications interface, and memory communicate with each other via the communication bus. The processor can invoke logical instructions from the memory to execute a multi-agent collaborative management method based on a neuromorphic fusion architecture as described in Embodiment 1.
[0082] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0083] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0084] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multi-agent collaborative management and control method based on a brain-like fusion architecture, characterized in that, include: Acquire multi-source data related to safety production and emergency management of the target industrial system and perform time-series alignment to obtain time-aligned multi-source data; Based on the time-aligned multi-source data, evidence representations from each data source are obtained, and a security situation awareness vector is obtained based on evidence conflict assessment and dynamic credibility adjustment mechanisms. Based on the security situation cognition vector, a risk field state model is constructed, and combined with dynamic risk thresholds, the risk evolution results are obtained. Based on the risk evolution results, a multi-agent collaborative decision-making model is constructed, and the optimal collaborative strategy for each agent is generated under safety constraints. The optimal collaborative strategy is transformed and executed to obtain the execution result. The effectiveness of the strategy execution is evaluated based on the execution results, and feedback and self-learning optimization are performed to maintain the optimal collaborative strategy.
2. The multi-agent collaborative management and control method based on a brain-like fusion architecture as described in claim 1, characterized in that, The method for acquiring time-aligned multi-source data is as follows: Based on the multi-source data, cleaning, outlier removal and format standardization are performed, and data identifiers and timestamp information are uniformly generated to obtain structured multi-source data; Based on the structured multi-source data, extract event anchors with business significance; Based on the event anchor points, obtain the optimal time mapping function for each data source relative to the reference time axis; Based on the optimal time mapping function, the timestamps of each data source are mapped to a unified reference time axis to achieve time axis correction of multi-source data and obtain the time-aligned multi-source data.
3. The multi-agent collaborative management and control method based on a brain-like fusion architecture as described in claim 2, characterized in that, The multi-source data includes: industrial control system operation data, production management and safety management system data, environmental and risk monitoring data, and personnel, equipment and emergency resource status data; The business-significant event anchors include: equipment start / stop events, threshold exceedance events, alarm triggering events, operation status change events, or combinations thereof.
4. The multi-agent collaborative management and control method based on a brain-like fusion architecture as described in claim 2, characterized in that, The method for obtaining the security situation awareness vector is as follows: Based on the aforementioned time-aligned multi-source data, a brain-like evidence fusion mechanism is used for feature modeling, which is then mapped to the original evidence representation corresponding to each data source. The dynamic reliability of each data source is obtained based on data quality, degree of evidence conflict, and time residual. The original evidence representation is weighted and modified based on the dynamic reliability to obtain the modified evidence support. The multimodal evidence conflict degree is obtained based on the modified evidence support degree. Based on the multimodal evidence conflict degree, the fused evidence quality function is obtained through evidence fusion operation; The evidence quality function is mapped to the security situation perception vector of the target industrial system at the current moment.
5. The multi-agent collaborative management and control method based on a brain-like fusion architecture as described in claim 4, characterized in that, The method for obtaining the risk evolution result is as follows: A risk field state model of the target industrial system is constructed based on the security situation awareness vector. Based on the risk field state model, the risk state of each node in the target industrial system is inferred over a long continuous time period to obtain the risk field potential energy. Based on the risk field potential energy, the dynamic risk threshold that the target industrial system can withstand is determined through probabilistic security constraints; The risk field potential energy and the dynamic risk threshold together constitute the risk evolution result.
6. The multi-agent collaborative management method based on a brain-like fusion architecture as described in claim 5, characterized in that, The optimal collaborative strategy is obtained as follows: Based on the risk evolution results, a multi-agent collaborative decision-making model is constructed, consisting of multiple agents, a global coordination unit, and a security constraint projection mechanism. Based on the multi-agent collaborative decision-making model, a two-layer collaborative optimization and security constraint projection mechanism is adopted, and the upper-layer coordinator generates a global coordination strategy vector. Under security constraints, each agent in the lower layer obtains its local optimal policy based on the global coordination policy vector. The optimal coordination strategy is obtained by aligning the local optimal strategy with the global coordination strategy vector.
7. The multi-agent collaborative management and control method based on a brain-like fusion architecture as described in claim 6, characterized in that, The multi-agent collaborative decision-making model includes: a perceptual agent, a reasoning agent, a decision-making agent, an execution agent, and a learning agent; The sensing agent is used to receive information about the current security status of the system and environmental status. The reasoning agent is used to perform risk trend reasoning based on risk evolution results and security situation information; The decision-making agent is used to generate candidate security control strategies based on the reasoning results; The execution agent is used to convert candidate security control policies into specific control instructions and execute them; The learning agent is used to update and optimize the strategy parameters based on the execution results and system feedback; The aforementioned intelligent agents make collaborative decisions under the constraints of the safety constraint projection mechanism to ensure that the generated control strategy meets the system safety constraints.
8. The multi-agent collaborative management and control method based on a brain-like fusion architecture as described in claim 7, characterized in that, The method for obtaining the execution result is as follows: The optimal collaborative strategy is converted into executable control instructions and sent to the target industrial system for execution. The execution result is obtained by real-time collection of equipment status changes and system operation information during the execution process.
9. The multi-agent collaborative management and control method based on a brain-like fusion architecture as described in claim 8, characterized in that, Maintaining the optimal collaborative strategy specifically includes: Based on the execution results, obtain the causal attributable execution deviation index; A learning optimization objective is constructed based on the causal attributability performance bias index and the dynamic risk threshold. Based on the learning optimization objective, it is converted into Lagrange form to obtain the Lagrange objective function; Based on the Lagrange objective function, the penalty coefficients for probabilistic security constraints are adaptively updated, so that the target industrial system increases the penalty intensity for security constraints when approaching the risk boundary, and automatically converges to the updated optimal cooperative strategy that satisfies the security constraints.
10. A multi-agent collaborative management and control system based on a brain-inspired fusion architecture, used to execute a multi-agent collaborative management and control method based on a brain-inspired fusion architecture as described in any one of claims 1-9, characterized in that, include: Multi-source data processing module, brain-like fusion cognition module, task risk reasoning module, multi-agent collaborative decision-making module, collaborative execution module, and feedback optimization module; The multi-source data processing module is used to acquire multi-source data related to the safety production and emergency management of the target industrial system and perform time-series alignment to obtain time-aligned multi-source data. The brain-like fusion cognitive module is used to obtain evidence representations from each data source based on the time-aligned multi-source data, and to obtain a security situation cognitive vector based on evidence conflict assessment and dynamic credibility adjustment mechanism. The task risk reasoning module is used to construct a risk field state model based on the security situation cognition vector, and combine it with dynamic risk thresholds to obtain risk evolution results. The multi-agent collaborative decision-making module is used to construct a multi-agent collaborative decision-making model based on the risk evolution results, and generate the optimal collaborative strategy for each agent under safety constraints. The collaborative execution module is used to transform and execute based on the optimal collaborative strategy to obtain the execution result; The feedback optimization module is used to evaluate the effectiveness of the strategy execution based on the execution results, and to perform feedback and self-learning optimization to maintain the optimal collaborative strategy.