Cooperative behavior decision-making method and system for multiple unmanned devices

By normalizing and encapsulating the heterogeneous perception streams of multiple unmanned devices and reconstructing the situational structure, and combining the credit allocation and policy evaluation of multi-agent decision-making networks, the problems of messy data formats and situational reconstruction in collaborative decision-making of multiple unmanned devices are solved, the scientific nature of collaborative decision-making and the synchronicity of execution are improved, and the efficiency and reliability of overall collaborative operations are enhanced.

CN121682113AInactive Publication Date: 2026-03-17ZHEJIANG ASIA PACIFIC INTELLIGENT NETWORK AUTOMOBILE INNOVATION CENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511883523.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-03-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies have failed to effectively address the issues of normalized encapsulation of heterogeneous perception flows, structural reconstruction of situation, credit allocation and strategy evaluation, and dynamic constraints and task capability description in collaborative decision-making among multiple unmanned devices. This results in insufficient rationality and feasibility of collaborative behavioral intention trajectories, making it difficult to guarantee the overall efficiency and reliability of collaborative operations.

Method used

By normalizing and encapsulating heterogeneous perception flows based on the link characteristics and synchronous signaling mechanism of multiple unmanned devices, and combining prior environmental knowledge for situational structure reconstruction, a credit allocation and policy evaluation mechanism of multi-agent decision network is adopted to encode the intention trajectory of collaborative behavior at the instruction granularity. Furthermore, the multi-agent decision network is optimized by establishing a global execution barrier and dynamically updating multi-dimensional performance indicators.

Benefits of technology

It achieves unified data format for collaborative decision-making among multiple unmanned devices, improves the accuracy of situational awareness, enhances the rationality and adaptability of collaborative behavioral intentions, and ensures the consistency of synchronous execution of equipment and the efficiency and reliability of collaborative operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121682113A_ABST
    Figure CN121682113A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses a multi-unmanned equipment cooperative behavior decision method and system, and the method comprises the steps: carrying out the normalization packaging of a heterogeneous perception stream, and obtaining an original data array; mapping the spatial-temporal characteristics of the original data array to preset environment priori knowledge, and performing situation structured reconstruction on the spatial-temporal characteristics and a characteristic dimension and a spatial topological relation of the environment priori knowledge to obtain a situation map; performing joint strategy optimization solution on the multi-agent behavior intention of the situation map to obtain a collaborative behavior intention track; performing instruction granularity coding on the collaborative behavior intention trajectory to obtain an atomic operation instruction sequence; distributing the atomic operation instruction sequence to multiple unmanned devices, and performing global execution barrier establishment on the multiple unmanned devices to obtain a synchronous execution event; performing gradient updating on strategy generation parameters of the multi-agent decision network to obtain a collaborative decision optimization network; according to the invention, the efficiency of cooperative behavior decision-making of multiple unmanned devices can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and system for collaborative behavior decision-making among multiple unmanned devices. Background Technology

[0002] Existing technologies fail to systematically normalize and encapsulate the heterogeneous sensing flows of multiple unmanned devices. The sensing data formats of different devices are chaotic and lack unified standards, making it impossible to form a structured raw data array. This makes it difficult to accurately map the spatiotemporal characteristics of the raw data array with the preset environmental prior knowledge, and also fails to effectively match the feature dimensions and spatial topological relationships of the environmental prior knowledge. Consequently, the situational structure reconstruction process has obvious defects, making it difficult to construct a situational map that can fully reflect the operating status of the devices, their interrelationships, and the environmental conditions.

[0003] Existing technologies lack a comprehensive credit allocation and strategy evaluation system for collaborative decision-making among multiple unmanned devices. They also lack effective support for joint strategy optimization of the behavioral intentions of multiple agents, and the rationality and feasibility of collaborative behavioral intention trajectories are insufficient. Furthermore, the instruction granularity encoding does not fully integrate the dynamic constraints and task capability descriptions of the devices, resulting in poor adaptability of atomic operation instruction sequences. The lack of scientific design in the construction of global execution barriers leads to insufficient consistency in the synchronous execution of devices. Moreover, the absence of a dynamic update mechanism based on multi-dimensional performance indicators and environmental disturbance signals makes it impossible for the collaborative decision-making network to adjust and optimize in a timely manner, making it difficult to guarantee the efficiency and reliability of the overall collaborative operation. Therefore, how to improve the generation efficiency of collaborative decisions has become an urgent problem to be solved. Summary of the Invention

[0004] This invention provides a method and system for collaborative behavior decision-making among multiple unmanned devices to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides a method for collaborative behavior decision-making of multiple unmanned devices, comprising: S1. Based on the link characteristics and synchronization signaling mechanism of multiple unmanned devices, the heterogeneous sensing streams of the multiple unmanned devices are normalized and encapsulated to obtain the original data array of the multiple unmanned devices. S2. Map the spatiotemporal features of the original data array to preset environmental prior knowledge, and perform situational structure reconstruction with the feature dimensions and spatial topology relationship of the environmental prior knowledge to obtain the situational map of the multi-unmanned equipment. S3. Based on the credit allocation and strategy evaluation mechanism of the multi-agent decision-making network in the multi-unmanned equipment, the joint strategy optimization solution is performed on the multi-agent behavioral intentions of the situation map to obtain the collaborative behavioral intention trajectory of the multi-unmanned equipment. S4. Based on the dynamic constraints and task capability description of the multi-unmanned equipment, the collaborative behavior intention trajectory is encoded at the instruction granularity to obtain the atomic operation instruction sequence of the multi-unmanned equipment; S5. Distribute the atomic operation instruction sequence to the multiple unmanned devices and establish a global execution barrier for the multiple unmanned devices to obtain the synchronous execution events of the multiple unmanned devices; S6. Based on the multi-dimensional performance indicators and environmental disturbance signals generated by the synchronous execution events, the policy generation parameters of the multi-agent decision-making network are updated in gradient to obtain the collaborative decision optimization network of the multi-unmanned equipment.

[0006] In a preferred embodiment, the normalization and encapsulation of the heterogeneous sensing streams of the multiple unmanned devices based on the link characteristics and synchronization signaling mechanism of the multiple unmanned devices to obtain the original data array of the multiple unmanned devices includes: Based on the link characteristics of multiple unmanned devices, the continuous sensing flow and state flow of the multiple unmanned devices are captured to obtain the heterogeneous sensing flow of the multiple unmanned devices. The communication protocol is parsed to obtain the sensing data stream of the multiple unmanned devices; Based on the global clock source of the synchronization signaling mechanism in the multi-unmanned equipment, a synchronization timestamp is applied to the sensing data stream to obtain the original data array of the multi-unmanned equipment.

[0007] In a preferred embodiment, the step of mapping the spatiotemporal features of the original data array to preset environmental prior knowledge, and performing situational structure reconstruction with the feature dimensions and spatial topological relationships of the environmental prior knowledge to obtain the situational map of the multi-unmanned equipment, includes: Data decoupling is performed on the original data array to obtain the time series components and spatial distribution components of the multiple unmanned devices; Based on preset environmental prior knowledge, semantic infusion is performed on the time series component and the spatial distribution component to obtain the instantiated feature vector of the multi-unmanned equipment. Based on the spatial topological relationships in the prior environmental knowledge, spatial association reasoning is performed on the entities of the instantiated feature vector to obtain the topological connections, orientation constraints and relative motion events of the multiple unmanned devices; The topological connections, orientation constraints, and relative motion events are woven and integrated to construct the spatial relationship event network of the multiple unmanned devices, thereby obtaining the situation map of the multiple unmanned devices.

[0008] In a preferred embodiment, the credit allocation and policy evaluation mechanism based on the multi-agent decision-making network of the multi-unmanned equipment performs joint policy optimization on the multi-agent behavioral intentions of the situation map to obtain the collaborative behavioral intention trajectory of the multi-unmanned equipment, including: By performing a graph structure traversal on the situation map, the independent state representations and interaction relationship representations of the multiple unmanned devices are obtained. Based on the credit allocation mechanism of the multi-agent decision-making network in the multi-unmanned equipment, the overall effectiveness of the group collaborative task in the multi-unmanned equipment is traced back to the independent state representation to obtain the differentiated contribution of the multi-unmanned equipment. Based on the strategy evaluation mechanism of the multi-agent decision network, the candidate joint strategies of the situation map are prospectively deduced to obtain the strategy superiority order of the multi-unmanned equipment. Based on the differential contribution degree and the superiority-inferiority relationship, the behavioral intentions of the multiple unmanned devices are targetedly coupled to obtain the optimized joint strategy of the multiple unmanned devices; The optimized joint strategy is smoothly connected to obtain the collaborative behavioral intention trajectory of the multiple unmanned devices.

[0009] In a preferred embodiment, the step of encoding the cooperative behavioral intent trajectory at the instruction granularity based on the dynamic constraints and task capability description of the multiple unmanned devices to obtain the atomic operation instruction sequence of the multiple unmanned devices includes: Based on the temporal ordering and state reachability constraints of the multi-unmanned equipment, key state nodes are extracted from the intention trajectory to obtain the target state sequence and expected behavior pattern of the multi-unmanned equipment. Based on the dynamic constraints of the multi-unmanned equipment, the expected behavior pattern is adjusted for boundary compliance to obtain the executable behavior skeleton of the multi-unmanned equipment. Based on the task capability description of the multi-unmanned equipment, the executable behavior skeleton is mapped to the functional operation set of the multi-unmanned equipment to obtain the specific function call instructions of the multi-unmanned equipment. The scheduling gaps of the specific function call instructions are filled to obtain the atomic operation instruction sequence of the multiple unmanned devices.

[0010] In a preferred embodiment, the step of distributing the atomic operation instruction sequence to the multiple unmanned devices and establishing a global execution barrier for the multiple unmanned devices to obtain the synchronous execution events of the multiple unmanned devices includes: The atomic operation instruction sequence is transmitted to the multi-unmanned equipment, and the instruction reception confirmation signal of the multi-unmanned equipment is received. Based on the completeness requirements of the multi-unmanned equipment, the instruction reception confirmation signal is verified to obtain the instruction distribution completeness of the multi-unmanned equipment; Based on the judgment rules and heartbeat detection protocol in the multi-unmanned equipment, distributed heartbeat detection is performed on the collaborative suspension status identifier of the instruction distribution integrity to obtain the global execution barrier of the multi-unmanned equipment. Based on the global execution barrier, a synchronization trigger signal is injected into the multiple unmanned devices to obtain the synchronization execution events of the multiple unmanned devices.

[0011] In a preferred embodiment, the step of updating the policy generation parameters of the multi-agent decision-making network based on the multi-dimensional performance indicators and environmental disturbance signals generated by the synchronous execution events to obtain the collaborative decision optimization network of the multi-unmanned equipment includes: Data separation is performed on the synchronous execution events to obtain multi-dimensional performance indicators and environmental disturbance signals of the multiple unmanned devices; A dual-path attribution analysis was performed on the multi-dimensional performance indicators and the environmental disturbance signals to obtain the incremental adjustment direction and adjustment range of the multiple unmanned devices. Based on the incremental adjustment direction and the adjustment magnitude, the strategy generation parameters are progressively adjusted to obtain the network parameter configuration of the multiple unmanned devices. The network parameter configuration is mapped to the multi-agent decision network to obtain the collaborative decision optimization network of the multi-unmanned equipment.

[0012] In a preferred embodiment, the step of performing a dual-path attribution analysis on the multi-dimensional performance indicators and the environmental disturbance signal to obtain the incremental adjustment direction and adjustment magnitude of the multiple unmanned devices includes: The multi-dimensional performance indicators are structured and analyzed to obtain the overall performance measurement of the group and the individual performance measurement of the multi-unmanned equipment. In the first path of the dual-path attribution analysis, based on the correlation analysis framework of the multiple unmanned devices, the change in the overall effectiveness metric of the group is attributed to the individual effectiveness metric to obtain the strategy contribution credit value of the multiple unmanned devices. In the second path of the dual-path attribution analysis, the environmental disturbance signal is abstracted to obtain the disturbance type and disturbance intensity of the environmental disturbance signal; The disturbance type and the disturbance intensity are mapped to the influence weight factors of the multi-agent decision network; Based on the policy gradient theory of the multi-unmanned equipment, the update intensity of the multi-unmanned equipment is obtained by combining the policy contribution credit value and the influence weight factor. Based on the historical update trends of the multiple unmanned devices, the update intensity is pruned in a standardized manner to obtain the incremental adjustment direction and adjustment range of the multiple unmanned devices.

[0013] In a preferred embodiment, the formula for calculating the update intensity is as follows: ; In the formula, The first of the multiple unmanned devices The update intensity of each agent The first of the multiple unmanned devices Each agent's strategy contributes a credit value. The aforementioned influencing weighting factor, The total number of intelligent agents in the multi-unmanned equipment. The arithmetic mean of the credit values ​​contributed by all agents to the strategies of the multi-unmanned equipment. The preset convergence adjustment coefficient is used. For a preset minimum positive number, This is for natural exponent calculations. This is for absolute value operations.

[0014] To address the aforementioned problems, the present invention also provides a multi-unmanned equipment collaborative behavior decision-making system, the system comprising: The data encapsulation module is used to normalize and encapsulate the heterogeneous sensing streams of the multiple unmanned devices based on the link characteristics and synchronization signaling mechanism of the multiple unmanned devices, so as to obtain the original data array of the multiple unmanned devices. The situation reconstruction module is used to map the spatiotemporal features of the original data array to preset environmental prior knowledge, and to perform situational structure reconstruction with the feature dimensions and spatial topology relationship of the environmental prior knowledge to obtain the situation map of the multi-unmanned equipment. The strategy optimization module is used to perform joint strategy optimization on the multi-agent behavioral intentions of the situation map based on the credit allocation and strategy evaluation mechanism of the multi-agent decision network in the multi-unmanned equipment, so as to obtain the collaborative behavioral intention trajectory of the multi-unmanned equipment. The instruction encoding module is used to encode the collaborative behavior intention trajectory at the instruction granularity according to the dynamic constraints and task capability description of the multi-unmanned equipment, so as to obtain the atomic operation instruction sequence of the multi-unmanned equipment; The synchronous execution module is used to distribute the atomic operation instruction sequence to the multiple unmanned devices, establish a global execution barrier for the multiple unmanned devices, and obtain the synchronous execution event of the multiple unmanned devices; The network update module is used to perform gradient updates on the policy generation parameters of the multi-agent decision network based on the multi-dimensional performance indicators and environmental disturbance signals generated by the synchronous execution events, so as to obtain the collaborative decision optimization network of the multi-unmanned equipment.

[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention combines the link characteristics and synchronization signaling mechanism of multiple unmanned devices to standardize and encapsulate heterogeneous sensing streams, enabling the sensing and status information of different devices to form a unified and complete original data array, ensuring data consistency and availability. Furthermore, by accurately mapping the spatiotemporal characteristics of the original data array with pre-defined environmental prior knowledge, and relying on feature dimensions and spatial topological relationships, the situational awareness structure is reconstructed, allowing the situational map to comprehensively capture the topological connections, orientation constraints, and relative motion events between devices, realistically presenting the device operating environment and interrelationships, and significantly improving the accuracy and comprehensiveness of situational awareness.

[0016] 2. This invention utilizes a multi-agent credit allocation and strategy evaluation mechanism to perform targeted coupling and joint strategy optimization of multi-agent behavioral intentions, making the collaborative behavioral intention trajectory more reasonable and adaptable. Furthermore, it combines equipment dynamics constraints and task capability descriptions for instruction granularity encoding, ensuring the executability of atomic operation instruction sequences. Global execution barriers enable synchronized equipment execution. Simultaneously, the decision network is dynamically updated based on multi-dimensional performance indicators and environmental disturbance signals, continuously optimizing strategy generation parameters. This significantly improves the scientific nature of collaborative decision-making among multiple unmanned devices, the synchronization of execution, and environmental adaptability, enhancing the overall efficiency and reliability of collaborative operations. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating a collaborative behavior decision-making method for multiple unmanned devices according to an embodiment of the present invention. Figure 2 A functional block diagram of a multi-unmanned equipment collaborative behavior decision-making system provided in an embodiment of the present invention; The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0018] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0019] This application provides a method for collaborative behavior decision-making among multiple unmanned devices. The executing entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the method for collaborative behavior decision-making among multiple unmanned devices can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0020] Reference Figure 1 The diagram shown is a flowchart illustrating a multi-unmanned equipment collaborative behavior decision-making method according to an embodiment of the present invention. In this embodiment, the multi-unmanned equipment collaborative behavior decision-making method includes: S1. Based on the link characteristics and synchronization signaling mechanism of multiple unmanned devices, the heterogeneous sensing streams of the multiple unmanned devices are normalized and encapsulated to obtain the original data array of the multiple unmanned devices. In this embodiment of the invention, the normalization and encapsulation of the heterogeneous sensing streams of the multiple unmanned devices based on the link characteristics and synchronization signaling mechanism of the multiple unmanned devices to obtain the original data array of the multiple unmanned devices includes: Based on the link characteristics of multiple unmanned devices, the continuous sensing flow and state flow of the multiple unmanned devices are captured to obtain the heterogeneous sensing flow of the multiple unmanned devices. The communication protocol is parsed to obtain the sensing data stream of the multiple unmanned devices; Based on the global clock source of the synchronization signaling mechanism in the multi-unmanned equipment, a synchronization timestamp is applied to the sensing data stream to obtain the original data array of the multi-unmanned equipment.

[0021] Based on the inherent characteristics of multi-unmanned equipment links, including transmission bandwidth, signal stability, and data transmission latency, a data acquisition channel precisely adapted to these characteristics is built. The channel adjusts the data reception rate according to the actual transmission capacity of the link to avoid data loss or congestion. It receives environmental perception information acquired by each unmanned equipment through its own sensors in real time. This information covers spatial features and environmental parameters around the equipment. At the same time, it collects the operating status information of each equipment, including battery level, current location, and working mode. The continuous perception information and status information of all equipment are summarized and integrated. Due to the differences in sensor types and data formats of different equipment, the summarized information exhibits heterogeneous characteristics, ultimately resulting in a heterogeneous perception stream of multi-unmanned equipment.

[0022] The heterogeneous sensing streams are analyzed using communication protocols. First, the communication protocol type used by the unmanned device corresponding to each data point is identified, and the fixed data structure and field meanings of the protocol are clarified. The data is then disassembled one by one according to the parsing rules specified in the protocol, and information unrelated to core sensing, such as control commands, device identifiers, and check codes, contained in the protocol header is removed. The core sensing data and status data carried by the protocol are then accurately extracted. All the extracted data are then arranged in a standardized manner according to a unified field definition and data format to ensure that the data is complete in content and uniform and recognizable in format, thereby obtaining a sensing data stream from multiple unmanned devices.

[0023] Relying on the global clock source preset by the synchronization signaling mechanism in multiple unmanned devices, this clock source provides a unified and accurate time reference for all unmanned devices. The operation and data acquisition of all devices follow the time calibration standard of this clock source. On each data record in the sensing data stream, the specific time information corresponding to the global clock source when the data was collected by the device is accurately marked, realizing strict unified alignment of different devices and different types of data in the time dimension. Then, all sensing data with synchronization timestamps are classified according to data type and arranged in an orderly manner according to the order of collection time, forming a unified structure, time synchronization, and original data array of multiple unmanned devices that can be directly used for subsequent situation reconstruction.

[0024] The beneficial effects are that by leveraging the link characteristics of multiple unmanned devices, continuous sensing and status streams can be accurately captured, ensuring that heterogeneous sensing streams can comprehensively cover device sensing information and operational status information. Then, by parsing and removing irrelevant and redundant content through targeted communication protocols, the sensing data stream has a unified format and complete core information. Combined with the global clock source of the synchronization signaling mechanism to apply a synchronization timestamp, the precise alignment of different devices and different types of data in the time dimension is achieved. The resulting raw data array has the characteristics of format uniformity, time synchronization, and content integrity, providing high-quality data support for subsequent spatiotemporal feature mapping, situational structure reconstruction, and other links, effectively improving the data reliability and processing efficiency of the entire process of collaborative decision-making of multiple unmanned devices.

[0025] S2. Map the spatiotemporal features of the original data array to preset environmental prior knowledge, and perform situational structure reconstruction with the feature dimensions and spatial topology relationship of the environmental prior knowledge to obtain the situational map of the multi-unmanned equipment. In this embodiment of the invention, the step of mapping the spatiotemporal features of the original data array to preset environmental prior knowledge, and performing situational structure reconstruction with the feature dimensions and spatial topological relationships of the environmental prior knowledge to obtain the situational map of the multi-unmanned equipment includes: Data decoupling is performed on the original data array to obtain the time series components and spatial distribution components of the multiple unmanned devices; Based on preset environmental prior knowledge, semantic infusion is performed on the time series component and the spatial distribution component to obtain the instantiated feature vector of the multi-unmanned equipment. Based on the spatial topological relationships in the prior environmental knowledge, spatial association reasoning is performed on the entities of the instantiated feature vector to obtain the topological connections, orientation constraints and relative motion events of the multiple unmanned devices; The topological connections, orientation constraints, and relative motion events are woven and integrated to construct the spatial relationship event network of the multiple unmanned devices, thereby obtaining the situation map of the multiple unmanned devices.

[0026] The original data array is decoupled. First, the time-related features in the data are identified, including the logic of data acquisition sequence and the evolution of state over time. Then, the spatial-related features are extracted, covering the spatial coverage of the device's sensing range and the spatial distribution pattern of data from different devices. By separating these two types of features and organizing them separately, the time series components and spatial distribution components of multiple unmanned devices are obtained.

[0027] Based on pre-defined environmental prior knowledge, which includes fixed information such as the inherent attributes of the equipment's operating environment, spatial layout rules, and common scene association logic, the data of different time periods in the time series component are matched with the scene features of the corresponding time periods in the environmental prior knowledge, giving the time series component a clear temporal semantics. At the same time, the data of different locations in the spatial distribution component are associated with the attribute information of the corresponding regions in the environmental prior knowledge, giving the spatial distribution component a clear spatial semantics. Through this semantic assignment process, the instantiated feature vectors of multiple unmanned devices are obtained.

[0028] Based on the spatial topological relationships in the prior knowledge of the environment, this relationship clarifies the spatial association rules, positional constraints, and common connection patterns of different elements in the environment. The spatial attributes of each entity in the instantiated feature vector are analyzed one by one, including the spatial coordinates and coverage of the entity. Based on the spatial topological relationships, it is determined whether there are direct or indirect associations between entities, and the orientational relationships between entities are determined. Combined with the changing trends reflected by the time series components, the dynamic changes in the relative positions between entities are captured, thereby obtaining the topological connections, orientational constraints, and relative motion events of multiple unmanned devices.

[0029] The obtained topological connections, orientation constraints, and relative motion events are woven and integrated. According to the time sequence and spatial logical relationship, the entity relationship reflected by the topological connections, the spatial position restriction of the orientation constraints, and the dynamic changes reflected by the relative motion events are systematically organized so that various types of information echo each other and form an organic whole. This constructs a spatial relationship event network that can comprehensively reflect the spatial distribution, interrelationship, and dynamic changes of multiple unmanned devices, and finally obtains the situation map of multiple unmanned devices.

[0030] The beneficial effects are as follows: by decoupling the original data array, the time series component and the spatial distribution component are accurately separated, so that the spatiotemporal characteristics of the data are clearly presented. Then, based on the pre-set environmental prior knowledge, semantic infusion is performed, so that the time and spatial dimension components have clear semantic orientations, forming instantiated feature vectors with practical meaning. Based on the spatial topological relationships in the environmental prior knowledge, spatial correlation reasoning is carried out to comprehensively capture the topological connections, orientation constraints and relative motion events of multiple unmanned devices. Finally, by weaving and integrating, a spatial relationship event network is constructed. The resulting situation map can completely and accurately reflect the spatial distribution, interrelationships and dynamic changes of multiple unmanned devices, providing high-quality and high-reliability situational basis for the joint strategy optimization of the behavior intentions of multiple intelligent agents, and effectively improving the accuracy and rationality of collaborative decision-making of multiple unmanned devices.

[0031] S3. Based on the credit allocation and strategy evaluation mechanism of the multi-agent decision-making network in the multi-unmanned equipment, the joint strategy optimization solution is performed on the multi-agent behavioral intentions of the situation map to obtain the collaborative behavioral intention trajectory of the multi-unmanned equipment. In this embodiment of the invention, the credit allocation and policy evaluation mechanism based on the multi-agent decision-making network of the multi-unmanned equipment, which performs joint policy optimization to solve the multi-agent behavioral intentions of the situation map, to obtain the collaborative behavioral intention trajectory of the multi-unmanned equipment, includes: By performing a graph structure traversal on the situation map, the independent state representations and interaction relationship representations of the multiple unmanned devices are obtained. Based on the credit allocation mechanism of the multi-agent decision-making network in the multi-unmanned equipment, the overall effectiveness of the group collaborative task in the multi-unmanned equipment is traced back to the independent state representation to obtain the differentiated contribution of the multi-unmanned equipment. Based on the strategy evaluation mechanism of the multi-agent decision network, the candidate joint strategies of the situation map are prospectively deduced to obtain the strategy superiority order of the multi-unmanned equipment. Based on the differential contribution degree and the superiority-inferiority relationship, the behavioral intentions of the multiple unmanned devices are targetedly coupled to obtain the optimized joint strategy of the multiple unmanned devices; The optimized joint strategy is smoothly connected to obtain the collaborative behavioral intention trajectory of the multiple unmanned devices.

[0032] The situation map is traversed structurally. Starting from each entity node in the situation map, the unique information of the unmanned equipment corresponding to each node, such as its own attributes, operating status, and perception data, is checked one by one. At the same time, the information interaction paths and functional cooperation logic between different devices are sorted out along the links between nodes. Through comprehensive investigation and sorting, the independent characteristics of each device and the interaction characteristics between devices are completely extracted, and the independent state representation and interaction relationship representation of multiple unmanned devices are obtained.

[0033] Based on the credit allocation mechanism of the multi-agent decision-making network in multi-unmanned equipment, we first clarify the overall performance of the group's collaborative task. This performance is determined based on core dimensions such as the quality and efficiency of task completion. Then, we map the overall performance to the equipment behavior corresponding to each independent state representation according to the behavioral association logic of the equipment. We analyze the specific role of each equipment behavior in improving or ensuring the overall performance, distinguish the differences in the size of the contributions of different equipment, and obtain the differentiated contribution degree of multi-unmanned equipment.

[0034] Based on the strategy evaluation mechanism of multi-agent decision-making network, all possible candidate joint strategies are listed for the task scenario corresponding to the situation map. The execution of each candidate joint strategy is simulated and deduced. During the deduction process, various constraints and interaction scenarios in the actual task environment are reproduced. The key performance of each candidate strategy in the execution process, such as the ability to achieve the task goal, resource consumption, and smoothness of behavior coordination, is observed. Based on these performances, the candidate strategies are comprehensively evaluated and ranked according to the degree of superiority and inferiority of the evaluation results to obtain the strategy superiority and inferiority relationship of multiple unmanned devices.

[0035] Based on the relationship between differentiated contribution and superiority / inferiority, the role weight of each device in the collaborative task is first determined according to the differentiated contribution. Devices with higher contribution play a more core role in the strategy. Then, the candidate strategy with the best ranking is selected as the basic framework in combination with the strategy superiority / inferiority relationship. The behavioral intentions of each device are then integrated in a targeted manner according to the role weight and the requirements of the basic framework to ensure that the behavioral intentions of each device are consistent with the overall collaborative goal, and that the behavioral intentions of the devices cooperate with each other without conflict or redundancy, thus obtaining an optimized joint strategy for multiple unmanned devices.

[0036] To ensure a smooth transition in the optimized joint strategy, we first analyze the connection nodes between different device behaviors and different execution stages in the optimized joint strategy, clarify the behavior switching logic at each connection node, and then supplement transitional behavior guidelines based on node characteristics. We adjust the execution sequence and intensity of relevant behaviors to eliminate execution gaps or conflicts between different behaviors, so that the entire joint strategy execution process is coherent and smooth, forming a complete, continuous, and implementable collaborative behavior intention trajectory of multiple unmanned devices.

[0037] The beneficial effects are as follows: by traversing the situation map structure, the independent state representations and interaction relationship representations of multiple unmanned devices are comprehensively obtained, providing complete and accurate basic information for joint strategy optimization. By leveraging the credit allocation mechanism of the multi-agent decision network, the overall effectiveness of the group collaborative task is traced back to the independent state representations, clarifying the differentiated contribution of each device and ensuring that strategy optimization can take into account both individual value and group goals. Based on the strategy evaluation mechanism, the candidate joint strategies are prospectively extrapolated to obtain the superiority order relationship, providing a reliable basis for selecting the optimal strategy. The behavioral intentions are targeted and coupled by combining the differentiated contribution and superiority order relationship, making the optimized joint strategy more targeted and adaptable. Then, through smooth connection processing, the collaborative behavioral intention trajectory is made coherent and smooth with strong executability, providing accurate and efficient behavioral guidance for the collaborative operation of multiple unmanned devices, significantly improving the scientific nature of collaborative decision-making and the overall collaborative operation effect.

[0038] S4. Based on the dynamic constraints and task capability description of the multi-unmanned equipment, the collaborative behavior intention trajectory is encoded at the instruction granularity to obtain the atomic operation instruction sequence of the multi-unmanned equipment; In this embodiment of the invention, the step of encoding the cooperative behavior intent trajectory at the instruction granularity according to the dynamic constraints and task capability description of the multiple unmanned devices to obtain the atomic operation instruction sequence of the multiple unmanned devices includes: Based on the temporal ordering and state reachability constraints of the multi-unmanned equipment, key state nodes are extracted from the intention trajectory to obtain the target state sequence and expected behavior pattern of the multi-unmanned equipment. Based on the dynamic constraints of the multi-unmanned equipment, the expected behavior pattern is adjusted for boundary compliance to obtain the executable behavior skeleton of the multi-unmanned equipment. Based on the task capability description of the multi-unmanned equipment, the executable behavior skeleton is mapped to the functional operation set of the multi-unmanned equipment to obtain the specific function call instructions of the multi-unmanned equipment. The scheduling gaps of the specific function call instructions are filled to obtain the atomic operation instruction sequence of the multiple unmanned devices.

[0039] Based on the temporal ordering constraints and state reachability constraints of multiple unmanned devices, the temporal ordering constraints clarify the logical sequence of device behavior execution, and the state reachability constraints define the feasibility boundary for the device to smoothly transition from the current state to the target state. By traversing the collaborative behavior intention trajectory, key state points that play a decisive role in task completion and cannot be omitted are identified. These nodes include the task start baseline state, core action switching nodes, and phased goal achievement states. These key state points are organized in chronological order to form the target state sequence of multiple unmanned devices. Then, based on the behavioral correlation logic between key state points, the action flow and behavioral rules that the device needs to execute in a coherent manner are summarized to obtain the expected behavior pattern of multiple unmanned devices.

[0040] Based on the dynamic constraints of multiple unmanned devices, these constraints cover the inherent attributes of the devices, such as the motion limit range, power output upper limit, operation response speed threshold, and physical structure limitations. Each action link in the expected behavior pattern is checked one by one to determine whether the execution parameters of the action exceed the boundaries of the device's dynamic constraints. For example, whether the execution rate of the action exceeds the maximum response capability of the device, whether the motion amplitude exceeds the physical structure limitations of the device, and whether the force intensity exceeds the power output upper limit. Actions that exceed the constraint boundaries are specifically corrected by adjusting the execution rhythm, amplitude, or method of the action to ensure that each action link conforms to the dynamic characteristics of the device, forming a well-structured, conflict-free, and implementable executable behavior skeleton for multiple unmanned devices.

[0041] Based on the task capability description of multiple unmanned devices, this description clarifies the specific functional types, operational coverage, suitable task scenarios, and functional implementation conditions that the device can perform. The function operation set of the multiple unmanned devices contains all the preset function instructions that can be directly triggered and executed. Each behavioral link in the executable behavior skeleton is precisely matched with the instructions in the function operation set. For example, the behavioral link of "environmental information collection" in the behavior skeleton corresponds to the "environmental data collection instruction" in the function operation set, and the behavioral link of "position adjustment" corresponds to the "precise positioning and movement instruction". This ensures that each behavioral link can find a unique corresponding function instruction that conforms to the task capability of the device, thus obtaining the specific function call instruction of the multiple unmanned devices.

[0042] To fill the scheduling gaps between specific function call instructions for multiple unmanned devices, we first analyze the execution time, device response latency, and other characteristics of each specific function call instruction to clarify the required connection time interval between adjacent instructions, avoiding instruction execution overlap or execution gaps. Then, based on the device's operation coordination logic, we supplement connecting operation guidelines or time calibration instructions to fill the scheduling gaps between instructions, so that all specific function call instructions are arranged in an orderly manner according to the task execution logic, forming an atomic operation instruction sequence for multiple unmanned devices where each instruction is indivisible, the execution process is coherent and smooth, and there are no connection breaks.

[0043] The beneficial effects are as follows: by transmitting atomic operation instruction sequences to multiple unmanned devices and receiving instruction reception confirmation signals, the instructions are ensured to be accurately delivered to each device. Based on the completeness requirement, the confirmation signals are verified to ensure the completeness of instruction distribution, guaranteeing that all devices successfully receive the instructions without omission. Relying on the judgment rules and heartbeat detection protocol in the multiple unmanned devices, distributed heartbeat detection is performed on the collaborative suspension status flag. The constructed global execution barrier can unify the device execution benchmark. Then, by injecting synchronous trigger signals, synchronous execution events of multiple unmanned devices are formed, which effectively avoids execution deviations or timing disorders between devices, ensures the consistency and coordination of execution actions, provides solid support for the smooth progress of multi-unmanned device collaborative operations, and significantly improves the reliability of collaborative execution and overall operational efficiency.

[0044] S5. Distribute the atomic operation instruction sequence to the multiple unmanned devices and establish a global execution barrier for the multiple unmanned devices to obtain the synchronous execution events of the multiple unmanned devices; In this embodiment of the invention, the step of distributing the atomic operation instruction sequence to the multiple unmanned devices and establishing a global execution barrier for the multiple unmanned devices to obtain the synchronous execution events of the multiple unmanned devices includes: The atomic operation instruction sequence is transmitted to the multi-unmanned equipment, and the instruction reception confirmation signal of the multi-unmanned equipment is received. Based on the completeness requirements of the multi-unmanned equipment, the instruction reception confirmation signal is verified to obtain the instruction distribution completeness of the multi-unmanned equipment; Based on the judgment rules and heartbeat detection protocol in the multi-unmanned equipment, distributed heartbeat detection is performed on the collaborative suspension status identifier of the instruction distribution integrity to obtain the global execution barrier of the multi-unmanned equipment. Based on the global execution barrier, a synchronization trigger signal is injected into the multiple unmanned devices to obtain the synchronization execution events of the multiple unmanned devices.

[0045] The atomic operation command sequence is transmitted to multiple unmanned devices. First, the unique identifier of each device is accurately identified using its built-in hardware encoding, network address, or dedicated identification chip, ensuring that each device's identifier is unique and can be quickly matched. Based on the identified unique identifier, a dedicated encrypted data transmission link is assigned to each device. This link first detects the device's communication interface type, distinguishing between wired and wireless interfaces and the differences in transmission protocols between different interfaces. Then, it dynamically adjusts the transmission mode based on the link's transmission bandwidth, signal stability, latency, and other characteristics. For example, high-bandwidth links use batch data packet transmission, low-latency links use continuous small packet transmission, and links prone to signal interference use redundant check transmission. Data packet loss and errors are monitored throughout the transmission process, and real-time retransmissions are made to ensure that every instruction in the atomic operation command sequence is sent to the target device completely and without errors. After the instruction sequence has been completely transmitted, an instruction reception confirmation signal is received from each device. This signal contains verification information corresponding to the device's unique identifier, a clear instruction reception completion marker, and a data integrity verification result generated based on the instruction data checksum and data length comparison. By checking whether this information is consistent with the instruction data recorded by the sending end, it is accurately confirmed that the instruction has been successfully received by the corresponding device and that there is no missing or tampered data.

[0046] Based on the completeness requirement of multiple unmanned devices, this requirement stipulates that all devices participating in this collaborative task, regardless of their functional role or deployment location, must fully receive the atomic operation command sequence, with no omissions or incomplete reception allowed. According to the preset target device list, each received command confirmation signal is checked one by one. First, it is checked whether the device identifier in the signal completely corresponds to the device in the list, ensuring that no device fails to return a signal and that no extra signals from non-target devices are mixed in. Then, for each confirmation signal, the data integrity check result is verified to be valid. Specifically, the checksum is checked to see if it matches the result calculated by the sender, whether the data length matches the original length of the command sequence, and whether there are any missing or misaligned data segments. Simultaneously, the signal format is strictly verified to conform to the preset standards, including the order of signal fields, the data type and length of each field, and the correct position of the check bits, ensuring that no key information is missing, and that there are no format errors or invalid characters in the signal. Only when the confirmation signals from all target devices meet the three requirements of full device identifier coverage, valid data integrity, and compliant signal format is the comprehensive verification of command distribution completed, and the completeness of command distribution for multiple unmanned devices is obtained.

[0047] Based on the decision-making rules and heartbeat detection protocol in multi-unmanned equipment, the decision-making rules specify the conditions for a device to enter the collaborative suspension state: in addition to successfully receiving the atomic operation instruction sequence and passing data integrity verification, the device must also complete its own hardware self-test, local storage and indexing of instruction data, parameter initialization before execution, and other preparatory work. Only after all preparatory steps meet the preset standards will the device generate a collaborative suspension state identifier. The heartbeat detection protocol specifies the sending frequency, signal format, and transmission path of the device status signal, clarifying that the device must send status signals at fixed time intervals to continuously feedback its own operating status. When performing distributed heartbeat detection on the collaborative suspension state identifier corresponding to the integrity of instruction distribution, each device will broadcast a status signal containing its own unique identifier, collaborative suspension state identifier, and confirmation information of the completion of preparatory work to all other devices participating in the collaborative task through the global communication network, while simultaneously listening to and receiving status signals sent by all other devices in real time. The device will establish a status checklist to record the status information of each device one by one. It will continuously compare whether the status identifiers of all devices are in the collaborative suspension state. If a device has not reached the collaborative suspension state, it will maintain heartbeat detection until the device completes the preparation work and sends the corresponding status signal. Only when the status identifiers of all devices are consistent and stable for multiple consecutive times without any status fluctuations can the global execution barrier of the multi-unmanned device be obtained.

[0048] Based on a global execution barrier, this barrier has confirmed through distributed heartbeat detection that all devices have completed instruction reception and preparation, and are in a unified collaborative suspension state, possessing the basic conditions for synchronous execution. Through a pre-established low-latency, highly synchronous global communication link, a synchronization trigger signal is simultaneously injected into all participating unmanned devices. This signal is a preset, uniquely identifiable digital instruction code or a specific frequency electrical signal, containing key information such as execution start timestamp and instruction sequence execution priority. Multi-path transmission ensures that all devices receive the signal simultaneously. Upon receiving the synchronization trigger signal, each device's built-in signal receiving module immediately activates the execution controller, calling the locally stored atomic operation instruction sequence and sending precise operation commands to the device's execution components one by one according to the logical order of the instructions. During execution, all devices rely on a global clock source for real-time timing calibration, ensuring that the start time, execution duration, and interval of each instruction are strictly consistent, avoiding time differences or sequence errors in execution actions. Ultimately, the execution actions of all devices remain highly synchronized in time, forming a coordinated and unified execution process, resulting in a synchronous execution event of multiple unmanned devices.

[0049] The beneficial effects are as follows: by transmitting atomic operation instruction sequences to multiple unmanned devices and receiving instruction reception confirmation signals, the instructions are ensured to be accurately delivered to each target device. Based on the integrity requirements of multiple unmanned devices, the confirmation signals are verified to obtain the integrity of instruction distribution, ensuring that all devices participating in the collaborative task successfully receive the instructions without omission. Relying on the judgment rules and heartbeat detection protocol in the multiple unmanned devices, distributed heartbeat detection is performed on the collaborative suspension status flag. The constructed global execution barrier can unify the execution benchmark of all devices, ensuring that all devices are in a ready state. Then, based on the global execution barrier, synchronous trigger signals are injected into multiple unmanned devices to form synchronous execution events, effectively avoiding the timing disorder or deviation of execution actions between devices, ensuring the consistency and coordination of multi-unmanned device collaborative operation, providing a solid guarantee for the smooth progress of the overall collaborative task, and significantly improving the reliability and efficiency of collaborative execution.

[0050] S6. Based on the multi-dimensional performance indicators and environmental disturbance signals generated by the synchronous execution events, the policy generation parameters of the multi-agent decision-making network are updated in gradient to obtain the collaborative decision optimization network of the multi-unmanned equipment.

[0051] In this embodiment of the invention, the step of updating the policy generation parameters of the multi-agent decision-making network based on the multi-dimensional performance indicators and environmental disturbance signals generated by the synchronous execution event to obtain the collaborative decision optimization network of the multi-unmanned equipment includes: Data separation is performed on the synchronous execution events to obtain multi-dimensional performance indicators and environmental disturbance signals of the multiple unmanned devices; A dual-path attribution analysis was performed on the multi-dimensional performance indicators and the environmental disturbance signals to obtain the incremental adjustment direction and adjustment range of the multiple unmanned devices. Based on the incremental adjustment direction and the adjustment magnitude, the strategy generation parameters are progressively adjusted to obtain the network parameter configuration of the multiple unmanned devices. The network parameter configuration is mapped to the multi-agent decision network to obtain the collaborative decision optimization network of the multi-unmanned equipment.

[0052] The dual-path attribution analysis of the multi-dimensional performance indicators and the environmental disturbance signals yields the incremental adjustment direction and magnitude of the multiple unmanned devices, including: The multi-dimensional performance indicators are structured and analyzed to obtain the overall performance measurement of the group and the individual performance measurement of the multi-unmanned equipment. In the first path of the dual-path attribution analysis, based on the correlation analysis framework of the multiple unmanned devices, the change in the overall effectiveness metric of the group is attributed to the individual effectiveness metric to obtain the strategy contribution credit value of the multiple unmanned devices. In the second path of the dual-path attribution analysis, the environmental disturbance signal is abstracted to obtain the disturbance type and disturbance intensity of the environmental disturbance signal; The disturbance type and the disturbance intensity are mapped to the influence weight factors of the multi-agent decision network; Based on the policy gradient theory of the multi-unmanned equipment, the update intensity of the multi-unmanned equipment is obtained by combining the policy contribution credit value and the influence weight factor. Based on the historical update trends of the multiple unmanned devices, the update intensity is pruned in a standardized manner to obtain the incremental adjustment direction and adjustment range of the multiple unmanned devices.

[0053] The formula for calculating the update intensity is as follows: ; In the formula, The first of the multiple unmanned devices The update intensity of each agent The first of the multiple unmanned devices Each agent's strategy contributes a credit value. The aforementioned influencing weighting factor, The total number of intelligent agents in the multi-unmanned equipment. The arithmetic mean of the credit values ​​contributed by all agents to the strategies of the multi-unmanned equipment. The preset convergence adjustment coefficient is used. For a preset minimum positive number, This is for natural exponent calculations. This is for absolute value operations.

[0054] Data separation is performed on synchronous execution events, which include various types of data generated during the execution of atomic operation command sequences by multiple unmanned devices. Data related to task completion quality, execution efficiency, and resource consumption are classified into one category, while data related to external environmental changes and interference factors are classified into another. By identifying the core characteristics of the two types of data, such as performance-related data directly related to task objectives and disturbance-related data unrelated to the device's own execution, the two types of data are extracted and normalized to obtain multi-dimensional performance indicators and environmental disturbance signals of the multiple unmanned devices.

[0055] A dual-path attribution analysis was conducted on multi-dimensional performance indicators and environmental disturbance signals. First, the multi-dimensional performance indicators were structured and analyzed to distinguish between the overall group performance metric reflecting the overall effect of collaborative operation of all equipment and the individual performance metric reflecting the performance of individual equipment. Based on the correlation analysis framework between equipment, the changes in the overall group performance metric were correlated with the performance metrics of each individual equipment, clarifying the degree of influence of each individual on the overall performance and obtaining the strategy contribution credit value of the multi-unmanned equipment. Simultaneously, the environmental disturbance signals were abstracted into context, identifying the specific type and intensity of the disturbance. According to the influence rules of the disturbance type and intensity on the decision network, they were transformed into corresponding influence weight factors. Then, combined with historical update experience and the operating logic of the decision network, a comprehensive analysis of the strategy contribution credit value and influence weight factors was conducted to determine the specific direction of adjustment required by the decision network. Finally, the adjustment range was determined based on the degree of influence of the two types of data, obtaining the incremental adjustment direction and adjustment magnitude of the multi-unmanned equipment.

[0056] Based on incremental adjustment direction and magnitude, the strategy generation parameters are progressively controlled. First, the decision function module corresponding to each strategy generation parameter is identified. Then, the adjustment tendency of each parameter to be enhanced, weakened, or maintained is determined according to the incremental adjustment direction. The adjustment step size of the parameter is set according to the adjustment magnitude. Starting from the core functional parameters of the decision network, the adjustment is gradually advanced to the auxiliary functional parameters. After each adjustment, the compatibility of the parameter with other related parameters is verified to avoid network operation conflicts caused by the adjustment of a single parameter. This ensures that the parameter adjustment process is smooth and conforms to the network operation logic. After gradual adjustment, a network parameter configuration for multiple unmanned devices that adapts to the current task requirements and environmental conditions is formed.

[0057] The network parameter configuration is mapped to a multi-agent decision network, which contains multiple functional modules, each corresponding to a specific policy generation task. Each parameter in the network parameter configuration is precisely matched with the corresponding module of the decision network according to its functional attributes. For example, parameters related to credit allocation correspond to the credit allocation module of the network, and parameters related to policy evaluation correspond to the policy evaluation module of the network. The new parameter configuration replaces the original parameters in the network one by one. After the replacement is completed, the entire decision network is functionally verified to ensure that all modules can operate collaboratively under the new parameter configuration and can generate better collaborative decision-making strategies based on the new parameters. Finally, a collaborative decision optimization network for multiple unmanned devices is obtained.

[0058] The multi-dimensional performance indicators are structured and organized, covering various information related to the effectiveness of collaborative operations, such as task completion quality, execution efficiency, resource consumption, and smoothness of coordination. The indicators are classified and organized according to the object attributes reflected by the data. The indicators that can reflect the comprehensive effect of collaborative operations of all unmanned equipment are integrated and categorized to form the overall performance measurement of multiple unmanned equipment groups. At the same time, the indicators that only reflect the performance of a single unmanned equipment are extracted and standardized separately to obtain the individual performance measurement of multiple unmanned equipment groups.

[0059] In the first path of the dual-path attribution analysis, relying on the correlation analysis framework of multiple unmanned devices, this framework clarifies the correlation logic and mapping rules between individual device behavior and group synergy. First, it determines the changes in the overall group effectiveness measure compared to the expected standard, including specific manifestations of improvement or decline. Then, it analyzes the correspondence between the changes in the effectiveness measure of each individual device and the changes in the overall group effectiveness measure. By tracing the source of the changes in group effectiveness, it clarifies the driving role of the behavior of each individual device in the changes in the overall group effectiveness, and then quantifies the degree of contribution of each individual to obtain the strategy contribution credit value of multiple unmanned devices.

[0060] In the second path of the dual-path attribution analysis, environmental disturbance signals are abstracted into context. These signals are external interference information that affects the collaborative operation of equipment. By identifying the core features, generation scenarios, and modes of action of the signals, different disturbance types are distinguished. At the same time, the disturbance intensity is determined based on the severity and scope of the impact of the signals on the equipment's execution actions and decision-making logic. Then, according to the preset mapping rules, different disturbance types and their corresponding disturbance intensities are transformed into quantitative values ​​that reflect their impact on the multi-agent decision-making network, thus obtaining the influence weight factor of the multi-agent decision-making network.

[0061] Based on the policy gradient theory of multi-unmanned equipment, this theory focuses on the correlation between policy adjustment and efficiency improvement. First, it clarifies the value of an individual policy as reflected by the policy contribution credit value. Then, it combines the degree of environmental interference reflected by the influencing weight factors and integrates the two. If the individual policy contribution credit value is high and the environmental disturbance has a large impact, the corresponding policy adjustment needs to be more urgent and the adjustment intensity needs to be increased accordingly. Conversely, the adjustment intensity should be appropriately reduced. Through this fusion judgment, the specific intensity of policy adjustment required for each agent is determined, thus obtaining the update intensity of multi-unmanned equipment.

[0062] Based on the historical update trends of multiple unmanned devices, which record the direction, magnitude, and actual effects of past decision network adjustments, the current update intensity is standardized and pruned to eliminate extreme update intensities that exceed the historical effective adjustment range and may lead to network instability. Referring to effective adjustment patterns in similar historical scenarios, the remaining update intensities are refined and corrected to clarify the adjustment direction of decision network parameters corresponding to each update intensity. At the same time, the reasonable range of adjustment is defined, resulting in the incremental adjustment direction and magnitude of multiple unmanned devices.

[0063] The strategy contribution credit value comes from the first path of the dual-path attribution analysis. Based on the correlation analysis framework of multiple unmanned devices, the change in the overall group effectiveness measurement is attributed to the individual effectiveness measurement. It is obtained after sorting out and judging the correlation between individual and group effectiveness.

[0064] The influencing weight factor comes from the second path of the dual-path attribution analysis. The environmental disturbance signal is abstracted into the disturbance type and disturbance intensity. Then, according to the preset mapping rules, the disturbance type and disturbance intensity are transformed into values ​​that can reflect their influence on the multi-agent decision network.

[0065] The total number of intelligent agents is the actual number obtained by statistically analyzing all intelligent agents participating in multi-unmanned equipment collaborative tasks.

[0066] The arithmetic mean of the credit values ​​of all agents' policy contributions is obtained by summing the credit values ​​of all agents' policy contributions, and then dividing the sum by the total number of agents.

[0067] The convergence adjustment coefficient is a preset fixed value, determined based on historical collaborative decision-making experience and the stability requirements of multi-agent decision-making networks, after analyzing and summarizing past network adjustment data.

[0068] The smallest positive number is a preset fixed value. To avoid the abnormal situation of the denominator being zero during the calculation process and to ensure the smooth progress of the calculation process, it is directly set based on the integrity requirements of the calculation logic.

[0069] The significance of this formula is to accurately determine the update intensity of each agent in a multi-agent unmanned device. It combines the policy contribution credit value of each agent to reflect the degree of individual contribution to the collaborative task, incorporates the influence weight factor to reflect the effect of environmental disturbances on the decision network, and calibrates the relative level of individual contribution by referring to the arithmetic mean of the policy contribution credit values ​​of all agents. It uses a convergence adjustment coefficient to control the change range of update intensity to avoid over-adjustment or under-adjustment, and adds a very small positive number to ensure that there are no abnormalities in the calculation process. The final update intensity can provide an accurate basis for adjusting the policy generation parameters of the multi-agent decision network, ensuring that the parameter adjustment intensity matches the individual contribution and environmental impact, and promoting the collaborative decision optimization network to better fit the current collaborative task requirements and environmental conditions.

[0070] The beneficial effects are as follows: by separating data from synchronously executed events, multi-dimensional performance indicators and environmental disturbance signals of multiple unmanned devices can be accurately obtained, providing comprehensive and core basic data support for decision network optimization. The multi-dimensional performance indicators are structured to obtain overall group performance metrics and individual performance metrics. Based on a correlation analysis framework, changes in overall group performance metrics are attributed to individual performance metrics to obtain policy contribution credit values. Simultaneously, environmental disturbance signals are abstracted into disturbance types and intensities, which are mapped to influence weight factors. This dual-path attribution analysis achieves a comprehensive consideration of individual contributions and environmental impacts, based on policy gradient theory. The update intensity of the contribution credit value and influence weight factor of the fusion strategy is obtained. By combining the historical update trend, the update intensity is standardized and pruned to obtain a scientific and reasonable incremental adjustment direction and adjustment range. Based on the adjustment direction and adjustment range, the strategy generation parameters are gradually adjusted to ensure that the parameter adjustment is stable and in line with actual needs, forming an adapted network parameter configuration. This network parameter configuration is mapped to a multi-agent decision network to obtain a collaborative decision optimization network. This enables the decision network to dynamically adapt to environmental changes and task execution feedback, significantly improving the accuracy, adaptability and reliability of collaborative decision-making of multiple unmanned devices, and ensuring the continuous and efficient advancement of collaborative operations.

[0071] like Figure 2 The diagram shown is a functional block diagram of a multi-unmanned equipment collaborative behavior decision-making system provided in an embodiment of the present invention.

[0072] The multi-unmanned equipment collaborative behavior decision-making system 100 described in this invention can be installed in an electronic device. Depending on the functions implemented, the multi-unmanned equipment collaborative behavior decision-making system 100 may include a data encapsulation module 101, a situation reconstruction module 102, a strategy optimization module 103, an instruction encoding module 104, a synchronization execution module 105, and a network update module 106. The modules described in this invention can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and are stored in the memory of the electronic device.

[0073] In this embodiment, the functions of each module / unit are as follows: The data encapsulation module 101 is used to normalize and encapsulate the heterogeneous sensing streams of the multiple unmanned devices based on the link characteristics and synchronization signaling mechanism of the multiple unmanned devices, so as to obtain the original data array of the multiple unmanned devices. The situation reconstruction module 102 is used to map the spatiotemporal features of the original data array to preset environmental prior knowledge, and to perform situational structure reconstruction with the feature dimensions and spatial topology relationship of the environmental prior knowledge to obtain the situation map of the multi-unmanned equipment. The strategy optimization module 103 is used to perform joint strategy optimization on the multi-agent behavioral intentions of the situation map based on the credit allocation and strategy evaluation mechanism of the multi-agent decision network in the multi-unmanned equipment, so as to obtain the collaborative behavioral intention trajectory of the multi-unmanned equipment. The instruction encoding module 104 is used to encode the collaborative behavior intention trajectory at the instruction granularity according to the dynamic constraints and task capability description of the multi-unmanned equipment, so as to obtain the atomic operation instruction sequence of the multi-unmanned equipment. The synchronous execution module 105 is used to distribute the atomic operation instruction sequence to the multiple unmanned devices and establish a global execution barrier for the multiple unmanned devices to obtain the synchronous execution event of the multiple unmanned devices; The network update module 106 is used to perform gradient updates on the policy generation parameters of the multi-agent decision network based on the multi-dimensional performance indicators and environmental disturbance signals generated by the synchronous execution event, so as to obtain the collaborative decision optimization network of the multi-unmanned equipment.

[0074] In the several embodiments provided by this invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.

[0075] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0076] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0077] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0078] This application embodiment can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for multi-unmanned device coordinated behavior decision-making, characterized in that, The method comprises: S1, based on the link characteristics of multiple unmanned devices and the synchronization signaling mechanism, the heterogeneous perception stream of the multiple unmanned devices is normalized and packaged to obtain the original data array of the multiple unmanned devices; S2, the spatio-temporal characteristics of the original data array are mapped to the preset environmental prior knowledge, and the feature dimension and spatial topological relationship of the environmental prior knowledge are situationally structured and reconstructed to obtain the situation spectrum of the multiple unmanned devices; S3, based on the credit distribution and policy evaluation mechanism of the multi-agent decision network in the multiple unmanned devices, the joint strategy optimization solution of the multi-agent behavior intention of the situation spectrum is obtained to obtain the cooperative behavior intention trajectory of the multiple unmanned devices; S4, according to the kinetic constraint and task capability description of the multiple unmanned devices, the instruction granularity coding of the cooperative behavior intention trajectory is carried out to obtain the atomic operation instruction sequence of the multiple unmanned devices; S5, the atomic operation instruction sequence is distributed to the multiple unmanned devices, and the global execution barrier of the multiple unmanned devices is established to obtain the synchronization execution event of the multiple unmanned devices; S6, based on the multi-dimensional performance index and environmental disturbance signal generated by the synchronization execution event, the gradient update of the strategy generation parameter of the multi-agent decision network is carried out to obtain the cooperative decision optimization network of the multiple unmanned devices.

2. The method of claim 1, wherein, Based on the link characteristics of multiple unmanned devices and the synchronization signaling mechanism, the heterogeneous perception stream of the multiple unmanned devices is normalized and packaged to obtain the original data array of the multiple unmanned devices, comprising: Based on the link characteristics of multiple unmanned devices, the continuous perception stream and state stream of the multiple unmanned devices are captured to obtain the heterogeneous perception stream of the multiple unmanned devices; The communication protocol analysis is carried out on the heterogeneous perception stream to obtain the perception data stream of the multiple unmanned devices; Based on the global clock source of the synchronization signaling mechanism in the multiple unmanned devices, the synchronization time stamp is applied to the perception data stream to obtain the original data array of the multiple unmanned devices.

3. The method of claim 1, wherein, The spatio-temporal characteristics of the original data array are mapped to the preset environmental prior knowledge, and the feature dimension and spatial topological relationship of the environmental prior knowledge are situationally structured and reconstructed to obtain the situation spectrum of the multiple unmanned devices, comprising: The data decoupling is carried out on the original data array to obtain the time series component and the spatial distribution component of the multiple unmanned devices; Based on the preset environmental prior knowledge, the semantic infusion is carried out on the time series component and the spatial distribution component to obtain the instantiated feature vector of the multiple unmanned devices; Based on the spatial topological relationship in the environmental prior knowledge, the spatial correlation reasoning of the entity of the instantiated feature vector is carried out to obtain the topological connection, the orientation constraint and the relative motion event of the multiple unmanned devices; The topological connection, the orientation constraint and the relative motion event are woven and integrated to construct the spatial relationship event network of the multiple unmanned devices to obtain the situation spectrum of the multiple unmanned devices.

4. The method of claim 1, wherein, The credit distribution and strategy evaluation mechanism based on the multi-agent decision network in the multiple unmanned devices optimizes and solves the joint strategy of the multi-agent behavior intention of the situation graph, and obtains the cooperative behavior intention trajectory of the multiple unmanned devices, including: The graph structure of the situation graph is traversed to obtain the independent state representation and interaction relationship representation of the multiple unmanned devices; According to the credit distribution mechanism of the multi-agent decision network in the multiple unmanned devices, the overall efficiency of the group cooperative task of the multiple unmanned devices is attributed to the independent state representation, and the differentiated contribution degree of the multiple unmanned devices is obtained; Based on the strategy evaluation mechanism of the multi-agent decision network, the candidate joint strategy of the situation graph is forward-looking deduced to obtain the strategy advantage and disadvantage sequence relationship of the multiple unmanned devices; Based on the differentiated contribution degree and the advantage and disadvantage sequence relationship, the behavior intention of the multiple unmanned devices is directionally coupled to obtain the optimized joint strategy of the multiple unmanned devices; The optimized joint strategy is smoothly connected to obtain the cooperative behavior intention trajectory of the multiple unmanned devices.

5. The method of claim 1, wherein, According to the kinetic constraints and task capability description of the multiple unmanned devices, the instruction granularity coding of the cooperative behavior intention trajectory is performed to obtain the atomic operation instruction sequence of the multiple unmanned devices, including: According to the time order and state reachability constraints of the multiple unmanned devices, the key state node of the intention trajectory is extracted to obtain the target state sequence and expected behavior mode of the multiple unmanned devices; Based on the kinetic constraints of the multiple unmanned devices, the boundary compliance of the expected behavior mode is adjusted to obtain the executable behavior skeleton of the multiple unmanned devices; Based on the task capability description of the multiple unmanned devices, the executable behavior skeleton is mapped to the function operation set of the multiple unmanned devices to obtain the specific function call instruction of the multiple unmanned devices; The specific function call instruction is scheduled and gap filled to obtain the atomic operation instruction sequence of the multiple unmanned devices.

6. The method of claim 1, wherein, The atomic operation instruction sequence is distributed to the multiple unmanned devices, and a global execution barrier is established for the multiple unmanned devices to obtain a synchronous execution event of the multiple unmanned devices, including: The atomic operation instruction sequence is directionally transmitted to the multiple unmanned devices, and an instruction receiving confirmation signal of the multiple unmanned devices is received; Based on the completeness requirement of the multiple unmanned devices, the instruction receiving confirmation signal is verified to obtain the instruction distribution integrity of the multiple unmanned devices; Based on the decision rule and heartbeat detection protocol in the multiple unmanned devices, the cooperative suspension state identification of the instruction distribution integrity is distributed heartbeat detected to obtain the global execution barrier of the multiple unmanned devices; Based on the global execution barrier, a synchronous trigger signal is injected into the multiple unmanned devices to obtain a synchronous execution event of the multiple unmanned devices.

7. The method of claim 1, wherein, Based on the multi-dimensional performance indicators and environmental disturbance signals generated by the synchronous execution event, the strategy generation parameters of the multi-agent decision network are gradient updated to obtain a cooperative decision optimization network of the multiple unmanned devices, including: Data is separated from the synchronous execution event, obtaining the multi-dimensional performance indicators and environmental disturbance signals of the multiple unmanned devices; Double-path attribution analysis is performed on the multi-dimensional performance indicators and the environmental disturbance signals, obtaining the incremental adjustment direction and adjustment amplitude of the multiple unmanned devices; Based on the incremental adjustment direction and the adjustment amplitude, the strategy generation parameters are gradually regulated, obtaining the network parameter configuration of the multiple unmanned devices; The network parameter configuration is mapped to the multi-agent decision network, obtaining the collaborative decision optimization network of the multiple unmanned devices.

8. The method of claim 7, wherein, The double-path attribution analysis on the multi-dimensional performance indicators and the environmental disturbance signals, obtaining the incremental adjustment direction and adjustment amplitude of the multiple unmanned devices, includes: The multi-dimensional performance indicators are structured and combed, obtaining the group overall performance measurement and individual performance measurement of the multiple unmanned devices; In the first path of the double-path attribution analysis, based on the correlation analysis framework of the multiple unmanned devices, the change of the group overall performance measurement is attributed to the individual performance measurement, obtaining the strategy contribution credit value of the multiple unmanned devices; In the second path of the double-path attribution analysis, the environmental disturbance signals are abstracted in the context, obtaining the disturbance type and disturbance intensity of the environmental disturbance signals; The disturbance type and the disturbance intensity are mapped as the influence weight factor of the multi-agent decision network; Based on the strategy gradient theory of the multiple unmanned devices, the strategy contribution credit value and the influence weight factor are obtained, obtaining the update strength of the multiple unmanned devices; Based on the historical update trend of the multiple unmanned devices, the update strength is pruned, obtaining the incremental adjustment direction and adjustment amplitude of the multiple unmanned devices.

9. The method of claim 8, wherein, The calculation formula of the update strength is as follows: ; In the formula, The first of the multiple unmanned devices The update intensity of each agent The first of the multiple unmanned devices Each agent's strategy contributes a credit value. The aforementioned influencing weighting factor, The total number of intelligent agents in the multi-unmanned equipment. The arithmetic mean of the credit values ​​contributed by all agents to the strategies of the multi-unmanned equipment. The preset convergence adjustment coefficient is used. For a preset minimum positive number, This is for natural index calculations. This is for absolute value operations.

10. A multi-unmanned device coordinated behavior decision system, characterized in that, A multi-unmanned device collaborative behavior decision method according to claim 1, the system comprises: A data encapsulation module for normalizing and encapsulating heterogeneous perception streams of multiple unmanned devices based on link characteristics and synchronization signaling mechanisms of the multiple unmanned devices, obtaining original data arrays of the multiple unmanned devices; A situation reconstruction module for mapping the time and space characteristics of the original data arrays to preset environmental prior knowledge, and structurally reconstructing the features and spatial topological relationships of the environmental prior knowledge, obtaining a situation atlas of the multiple unmanned devices; A strategy optimization module for jointly optimizing and solving multi-agent behavior intentions of the situation atlas based on credit allocation and strategy evaluation mechanisms of a multi-agent decision network in the multiple unmanned devices, obtaining collaborative behavior intention trajectories of the multiple unmanned devices; An instruction encoding module for encoding the collaborative behavior intention trajectories in instruction granularity according to the dynamics constraints and task capability descriptions of the multiple unmanned devices, obtaining atomic operation instruction sequences of the multiple unmanned devices; A synchronous execution module for distributing the atomic operation instruction sequences to the multiple unmanned devices and establishing a global execution barrier, obtaining synchronous execution events of the multiple unmanned devices; A network updating module is configured to perform gradient updating on policy generation parameters of the multi-agent decision network based on the multi-dimensional performance index generated by the synchronization execution event and the environmental disturbance signal, and obtain a collaborative decision optimization network of the multiple unmanned devices.