Cross-platform autonomous cooperative decision-making method and system
Patent Information
- Application Number
- CN202610843513.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-11
- Publication Date
- 2026-09-11
AI Technical Summary
但是无人机、无人车、巡检机器人等无人设备的运动能力存在差异,无人机的飞行速度较快,巡检机器人的行进速度较慢,各设备完成同一项作业任务中各自动作的耗时并不一致
1、本发明获取多个异构设备对采样目标采样得到的采样数据流,对采样数据流进行预处理得到当前时刻的采样阵列,基于采样阵列和预设的环境先验知识构建态势图谱,根据预训练的多智能体决策网络对态势图谱进行求解,得到各异构设备对应的协同行为数据,基于预设的时间步对协同行为数据进行拆解得到决策指令序列,下发各决策指令序列中同一时间步的决策指令至相应异构设备,并根据异构设备的响应信号同步下发下一时间步的决策指令。本发明将各异构设备的协同行为数据拆解为决策指令序列,并在收到异构设备的响应信号后再下发下一时间步的决策指令,可以在多异构设备协同过程中实现执行动作的同步管控,避免异构设备之间的相互等待或碰撞。
Smart Images

Figure CN122732115A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to data processing technology, and more particularly to a cross-machine type autonomous collaborative decision-making method and system. Background Technology
[0002] With the rapid development of unmanned systems technology, heterogeneous unmanned equipment such as drones, unmanned vehicles, and inspection robots are widely used in emergency rescue, industrial inspection, and security patrol. Cross-type collaborative operation technology, by integrating the perception and operation capabilities of different types of equipment, enables the efficient completion of complex tasks and has become one of the core technologies in the field of unmanned systems.
[0003] Currently, collaborative operations involving multiple heterogeneous devices primarily rely on a central controller to uniformly allocate tasks, after which each device executes independently, or on distributed negotiation among devices to determine their actions. However, unmanned devices such as drones, autonomous vehicles, and inspection robots exhibit varying mobility; drones fly at high speeds, while inspection robots move at slower speeds, resulting in inconsistent timeframes for each device to complete the same task. In existing collaborative operation methods, each heterogeneous device executes its task independently at its own pace after receiving it. Faster devices may begin subsequent actions before slower devices have completed their current ones, leading to staggered progress in the same task. Faster devices may have to wait for slower devices, or multiple devices may experience positional conflicts within the work area due to inconsistent progress.
[0004] Therefore, how to achieve synchronous control of actions executed during the collaboration of multiple heterogeneous devices, and avoid mutual waiting or collisions between heterogeneous devices, has become a key issue that urgently needs to be addressed. Summary of the Invention
[0005] This invention provides a cross-type autonomous collaborative decision-making method and system, which can realize synchronous control of the execution actions during the collaboration of multiple heterogeneous devices, and avoid mutual waiting or collision between heterogeneous devices.
[0006] A first aspect of the present invention provides a cross-aircraft autonomous collaborative decision-making method, comprising: S1, acquire sampling data streams obtained by multiple heterogeneous devices sampling the sampling target, preprocess the sampling data streams to obtain the sampling array at the current moment, and construct a situation map based on the sampling array and preset environmental prior knowledge; S2, Solve the situation map according to the pre-trained multi-agent decision network to obtain the cooperative behavior data corresponding to each heterogeneous device; S3, based on a preset time step, the collaborative behavior data is decomposed to obtain the decision instruction sequence corresponding to the collaborative behavior data; S4, issue decision instructions at the same time step in each decision instruction sequence to the corresponding heterogeneous devices, and synchronously issue decision instructions for the next time step according to the response signals of the heterogeneous devices.
[0007] Optionally, in one possible implementation of the first aspect, the preprocessing of the sampled data stream to obtain the sampling array at the current moment includes: Based on a preset reference frequency, time alignment processing is performed on each sampled data stream to obtain the aligned sampled data of each heterogeneous device at the current moment; The alignment sampling data is converted into image data of a preset size to obtain the size-adjusted alignment sampling data. The aligned sampled data after size adjustment are spliced together to obtain the sampled array at the current moment.
[0008] Optionally, in one possible implementation of the first aspect, the construction of the situation map based on the sampling array and preset environmental prior knowledge includes: Feature extraction is performed on the sampling array to obtain the situational features corresponding to each heterogeneous device and sampling target; Based on prior environmental knowledge, environmental constraint information corresponding to the situation features is retrieved, and corresponding environmental constraint information is added to the situation features to obtain the object features corresponding to each heterogeneous device and sampling target. Determine the spatial and interactive relationships between the heterogeneous devices, as well as between the heterogeneous devices and the sampling target; Using heterogeneous devices and sampling targets as nodes, nodes with spatial or interactive relationships are connected, and the object features are used as the node features of the corresponding nodes to obtain a situation map.
[0009] Optionally, in one possible implementation of the first aspect, step S2 includes: Obtain the multi-agent decision network trained based on the multi-agent proximal policy optimization algorithm, and input the situation map into the multi-agent decision network; The multi-agent decision network is used to solve the input situation map and output the collaborative behavior data of each heterogeneous device.
[0010] Optionally, in one possible implementation of the first aspect, step S3 includes: Retrieve a predefined decision instruction library, which includes basic action instructions corresponding to various types of heterogeneous devices; The collaborative behavior data is split along the time axis based on a preset time step to obtain the split behavior data corresponding to each time step. The split behavior data is matched with the decision instruction library to obtain the decision instructions corresponding to each split behavior data, and the decision instructions are concatenated to generate a decision instruction sequence.
[0011] Optionally, in one possible implementation of the first aspect, step S4 includes: S41, obtain the decision instruction of the first time step in each decision instruction sequence as the execution instruction; S42, the execution instruction is sent to the corresponding heterogeneous device, and a response signal from the heterogeneous device indicating that the execution instruction has been completed is received; S43, for the issued execution command, the response signal received within the preset reception time period will be used as the currently received response signal; S44, based on the currently received response signal, obtain the decision instruction for the next time step in the corresponding decision instruction sequence, and use it as the current execution instruction; S45, repeat steps S42-S44 until all decision instructions for all time steps have been issued.
[0012] Optionally, in one possible implementation of the first aspect, the step of using the response signal received within a preset reception duration as the currently received response signal includes: Obtain the sending time of the execution instruction, and based on the sending time and the preset receiving time, obtain the deadline time corresponding to the sending time; If a response signal is received before the deadline corresponding to the sending time, the corresponding response signal will be used as the currently received response signal.
[0013] A second aspect of the present invention provides a cross-aircraft autonomous collaborative decision-making system, comprising: The construction module is used to acquire the sampling data stream obtained by multiple heterogeneous devices sampling the sampling target, preprocess the sampling data stream to obtain the sampling array at the current time, and construct a situation map based on the sampling array and preset environmental prior knowledge; The computation module is used to solve the situation map based on the pre-trained multi-agent decision network to obtain the collaborative behavior data corresponding to each heterogeneous device; The decomposition module is used to decompose the collaborative behavior data based on a preset time step to obtain the decision instruction sequence corresponding to the collaborative behavior data; The decision module is used to issue decision instructions at the same time step in each decision instruction sequence to the corresponding heterogeneous devices, and to synchronously issue decision instructions for the next time step based on the response signals of the heterogeneous devices.
[0014] A third aspect of the present invention provides an electronic device comprising: a memory, a processor, and a computer program, the computer program being stored in the memory, and the processor executing the computer program to perform the methods described in the first aspect of the present invention and various possible methods related to the first aspect.
[0015] A fourth aspect of the present invention provides a storage medium storing a computer program, which, when executed by a processor, is used to implement the first aspect of the present invention and various methods possibly involved in the first aspect.
[0016] The beneficial effects of this invention are as follows: 1. This invention acquires sampling data streams from multiple heterogeneous devices sampling a target, preprocesses the sampling data streams to obtain a sampling array for the current time, constructs a situation map based on the sampling array and preset environmental prior knowledge, solves the situation map using a pre-trained multi-agent decision network to obtain the collaborative behavior data corresponding to each heterogeneous device, decomposes the collaborative behavior data into a decision instruction sequence based on a preset time step, issues decision instructions at the same time step in each decision instruction sequence to the corresponding heterogeneous device, and synchronously issues decision instructions for the next time step based on the response signals of the heterogeneous devices. This invention decomposes the collaborative behavior data of each heterogeneous device into a decision instruction sequence and issues the decision instruction for the next time step only after receiving the response signals from the heterogeneous devices. This enables synchronous control of actions executed during multi-heterogeneous device collaboration, avoiding mutual waiting or collisions between heterogeneous devices.
[0017] 2. This invention performs time alignment processing on the sampled data stream, extracts the aligned sampled data at the current moment, converts the aligned sampled data into image data of a preset size, and then stitches them together to obtain a sampled array, thereby achieving the unification of the sampling frequency and data format of the sampled data.
[0018] 3. This invention splits the collaborative behavior data and matches it with a decision instruction library to obtain a decision instruction sequence, thus converting the collaborative behavior of various heterogeneous devices into decision instructions at a unified time step. After issuing decision instructions at the same time step, this invention issues decision instructions for the next time step based on the response signals received within a preset reception period, avoiding the overall workflow from stalling when a heterogeneous device fails to provide a response signal for an extended period. Attached Figure Description
[0019] Figure 1 A flowchart illustrating a cross-aircraft autonomous collaborative decision-making method provided by this invention; Figure 2 A flowchart illustrating the generation of a decision instruction sequence for this invention; Figure 3 This is a schematic diagram of the structure of a cross-type autonomous collaborative decision-making system provided by the present invention; Figure 4 This is a schematic diagram of the hardware structure of an electronic device provided by the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein.
[0022] It should be understood that in the various embodiments of the present invention, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0023] It should be understood that in this invention, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or device.
[0024] It should be understood that in this invention, "multiple" refers to two or more. "And / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "Contains A, B, and C", "Contains A, B, and C" means that all three A, B, and C are contained; "Contains A, B, or C" means that one of A, B, and C is contained; "Contains A, B, and / or C" means that any one, two, or three of A, B, and C are contained.
[0025] It should be understood that in this invention, "B corresponding to A", "B corresponding to A", "A and B correspond", or "B and A correspond" means that B is associated with A, and B can be determined based on A. Determining B based on A does not mean determining B solely based on A; B can also be determined based on A and / or other information. Matching A and B is defined as a similarity between A and B that is greater than or equal to a preset threshold.
[0026] Depending on the context, "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection."
[0027] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0028] This invention provides a cross-aircraft autonomous collaborative decision-making method, such as... Figure 1 As shown, it includes: S1, acquire the sampling data stream obtained by multiple heterogeneous devices sampling the sampling target, preprocess the sampling data stream to obtain the sampling array at the current time, and construct a situation map based on the sampling array and preset environmental prior knowledge.
[0029] Heterogeneous equipment refers to work equipment with different operational capabilities or different types of sensors. In this invention, heterogeneous equipment includes drones equipped with RGB color cameras, unmanned vehicles equipped with lidar, and inspection robots equipped with thermal imaging cameras.
[0030] It should be noted that drones are equipped with image sensors, unmanned vehicles with point cloud sensors, and inspection robots with thermal imaging sensors. The sampling data collected by these heterogeneous devices differ in sampling frequency and data format. Therefore, this step performs unified preprocessing on the sampling data streams to standardize the format of different types of sampling data.
[0031] Understandably, the process begins by acquiring sampling data streams from drones, unmanned vehicles, and inspection robots sampling targets. The sampling targets are the task objects that perform collaborative tasks across these heterogeneous devices, and the sampling data streams are the data sequences continuously collected by the sensors of these devices. Since the sampling frequencies of these heterogeneous devices differ, it may be impossible to directly obtain sampling data from each device at the same time. Therefore, the sampling data streams are preprocessed to align them. After alignment, the sampling data for the current moment is extracted, and its format and size are standardized. Finally, the sampling data from each heterogeneous device is stitched together to obtain a multi-channel array, i.e., the sampling array, for the current moment. Next, pre-stored environmental prior knowledge is retrieved. This prior knowledge consists of pre-stored environmental information about the area where the sampling targets are located, including static maps and prohibited areas. Based on the sampling array and the environmental prior knowledge, a situational graph is constructed. This situational graph is a graph structure with heterogeneous devices and sampling targets as nodes and the relationships between nodes as edges.
[0032] In some embodiments, step S1 (preprocessing the sampled data stream to obtain the current sampling array) includes S11-S13: S11, perform time alignment processing on each sampled data stream based on a preset reference frequency to obtain the aligned sampled data of each heterogeneous device at the current moment.
[0033] The reference frequency refers to a pre-set common frequency used to align the sampled data streams of heterogeneous devices; the aligned sampled data refers to the sampled data of each heterogeneous device at the current moment after time alignment processing.
[0034] Understandably, a reference frequency is pre-defined. Since the original sampling frequencies of the data streams from different heterogeneous devices are not identical, interpolation or downsampling is performed on the sampling frequencies of different devices to unify the data streams to the reference frequency, ensuring that the data streams correspond to each other in time. After unifying the data streams to the reference frequency, a frame of data corresponding to the current moment is taken from each data stream to obtain the aligned sampling data for the drone, unmanned vehicle, and inspection robot at the current moment.
[0035] S12, the alignment sampling data is converted into image data of a preset size to obtain the size-adjusted alignment sampling data.
[0036] It should be noted that the image data from drones, the point cloud data from unmanned vehicles, and the thermal imaging data from inspection robots differ in data format and size. The resolutions of the image data and thermal imaging data are inconsistent, and the point cloud data is not in image format. Therefore, it is necessary to unify the data format of each data source. This step converts the aligned sampling data from these heterogeneous devices into image data of a preset size.
[0037] Understandably, a preset size is established, which refers to a pre-defined image size. The aligned sampling data is then converted into image data of the preset size, resulting in resized aligned sampling data. For example, if the preset size is 640×640 pixels, the resolution of image data from a drone is scaled down to 640×640 pixels; for point cloud data from an unmanned vehicle, the point cloud data is first voxelized, and then projected into a 640×640 pixel pseudo-depth image; for thermal images of inspection robots, their resolution is adjusted to 640×640 pixels. After these processes, the aligned sampling data from drones, unmanned vehicles, and inspection robots are all converted into 640×640 pixel image data, resulting in resized aligned sampling data.
[0038] S13, stitch together the aligned sampling data after size adjustment to obtain the sampling array at the current moment.
[0039] Understandably, the aligned sampling data of all heterogeneous devices at the same time, after size adjustments, are stitched together along the channel dimension, with the number of channels being the sum of the number of channels corresponding to each heterogeneous device. For example, a drone's RGB image has 3 channels, an unmanned vehicle's depth map has 1 channel, and an inspection robot's thermal image has 1 channel. After stitching, a sampling array with 5 channels is obtained, with a size of 640×640×5.
[0040] Among them, the sampling array refers to a multi-channel array obtained by stitching together the sampling data of various heterogeneous devices at the current moment to unify them into the same size.
[0041] In some embodiments, step S1 (constructing a situation map based on the sampling array and preset environmental prior knowledge) includes S14-S17: S14, perform feature extraction on the sampling array to obtain the situation features corresponding to each heterogeneous device and sampling target.
[0042] It is understandable that feature extraction is performed on the sampling array to identify heterogeneous devices and sampling targets from the sampling array, and the positions and motion states of the heterogeneous devices and sampling targets are extracted respectively. The motion states include the direction of motion and the speed of motion. The extracted positions and motion states are used as the situational features of the heterogeneous devices or sampling targets.
[0043] Among them, situational characteristics refer to the features that represent the position and motion state of heterogeneous devices and sampling targets.
[0044] S15, based on prior environmental knowledge, retrieve environmental constraint information corresponding to the situation features, add corresponding environmental constraint information to the situation features, and obtain object features corresponding to each heterogeneous device and sampling target.
[0045] Among them, environmental constraint information refers to the restriction information of the corresponding location determined from prior environmental knowledge based on the location of each heterogeneous device or sampling target.
[0046] Understandably, based on the location information of each heterogeneous device and sampling target, the environmental constraint information corresponding to that location is queried from prior environmental knowledge, such as whether it is a dangerous area or whether there are obstacles, and this information is added to the situation features to obtain object features containing environmental constraint information.
[0047] S16, determine the spatial and interactive relationships between the heterogeneous devices, as well as between the heterogeneous devices and the sampling target.
[0048] It should be noted that collaborative decision-making requires not only determining the state of each heterogeneous device and the sampling target, but also the relationships between the heterogeneous devices and between each heterogeneous device and the sampling target. For example, the distance between two heterogeneous devices, or whether there is a cooperative working relationship between the heterogeneous devices and the sampling target. Therefore, this step determines the spatial and interactive relationships between the heterogeneous devices and between each heterogeneous device and the sampling target.
[0049] Understandably, based on the positions of each heterogeneous device and the sampling target in the object characteristics, the spatial relationships between the heterogeneous devices and between each heterogeneous device and the sampling target are calculated. Spatial relationships refer to positional relationships, including relative distance and occlusion status. The occlusion status is determined by whether there is an obstruction between two objects, and the occlusion status can be represented by Boolean values to indicate whether an obstruction exists. Based on the collaborative tasks of each heterogeneous device, the interaction relationships between the heterogeneous devices and between each heterogeneous device and the sampling target are determined. Interaction relationships are the cooperative relationships of each heterogeneous device with other heterogeneous devices or the sampling target in the collaborative task, such as the cooperative relationship of a drone guiding an unmanned vehicle to the sampling target.
[0050] S17. Using each heterogeneous device and sampling target as nodes, connect nodes that have spatial or interactive relationships, and use the object features as the node features of the corresponding nodes to obtain a situation map.
[0051] Understandably, each heterogeneous device and sampling target is treated as a node in the situation map. For two nodes with spatial or interactive relationships, a connection is established between them to form edges in the situation map. For two nodes without spatial or interactive relationships, no connection is established. The object features of each heterogeneous device and sampling target obtained in step S15 are used as the node features of the corresponding nodes, thus obtaining the situation map. In the situation map, nodes represent the respective states of each heterogeneous device and sampling target, and edges represent the spatial and interactive relationships between the nodes.
[0052] S2, Solve the situation map based on the pre-trained multi-agent decision network to obtain the collaborative behavior data corresponding to each heterogeneous device.
[0053] It should be noted that when multiple heterogeneous devices perform the same collaborative task, they need to cooperate with each other, and a single heterogeneous device cannot determine its own actions in the collaborative task. This step solves the situational awareness graph based on a pre-trained multi-agent decision network to obtain the collaborative behavior data of the heterogeneous devices.
[0054] Among them, the multi-agent decision network is a network pre-trained based on the MAPPO algorithm, which can output the collaborative behavior data of each heterogeneous device according to the input situation map.
[0055] In some embodiments, step S2 includes S21-S22: S21, Obtain the multi-agent decision network trained based on the multi-agent proximal policy optimization algorithm, and input the situation map into the multi-agent decision network.
[0056] Understandably, the multi-agent decision network (MAD) uses the MAPPO algorithm as its foundation and introduces an intent-sharing mechanism. The MAD consists of an action network and an evaluation network. During training, the action network outputs collaborative behavior data from each heterogeneous device, and the evaluation network evaluates this data and adjusts its parameters accordingly. During usage, only the action network outputs the collaborative behavior of each heterogeneous device. The input to the MAD is the local observations and situational awareness map for each heterogeneous device. Local observations include the target location sampled by the device and the device's own pose. The action network outputs different collaborative behavior data for different types of heterogeneous devices: horizontal speed, vertical speed, yaw rate, and gimbal pitch for UAVs; linear speed and steering angle for unmanned vehicles; and track speed and robotic arm joint angles for inspection robots.
[0057] It should be noted that the training of the multi-agent decision network is conducted in a simulation environment built on Gazebo and CREW-Wildfire. The simulation environment includes drones, unmanned vehicles, inspection robots, and sampling targets. These heterogeneous devices interact multiple times within the simulation environment, generating training data consisting of situational maps and corresponding behaviors. The multi-agent decision network is then trained based on this training data. It is worth mentioning that data training based on multi-agent decision networks is existing technology and will not be elaborated upon here.
[0058] The present invention can also adjust the parameters of the action network and the evaluation network through a preset reward function, the reward function being as follows: in, As a reward value, For task completion rate, As a penalty for collision, For communication energy consumption, For synergistic gain, , , , These are the weights corresponding to task completion, collision penalty, communication energy consumption, and cooperation gain. Task completion refers to the progress of task completion, such as the number of target points covered; collision penalty is the penalty given when heterogeneous devices collide; communication energy consumption is the energy consumed by communication between heterogeneous devices; cooperation gain is the reward given when heterogeneous devices cooperate with each other, such as the reward given when a drone provides vision for an unmanned vehicle.
[0059] S22, the input situation map is solved based on the multi-agent decision network, and the collaborative behavior data of each heterogeneous device is output.
[0060] It is understandable that the current situation map is input into a pre-trained multi-agent decision network. Based on the node characteristics of each node in the situation map and the edges between nodes, the multi-agent decision network outputs the collaborative behavior data of each heterogeneous device in the future. The collaborative behavior data refers to the behavior of each heterogeneous device in the future, such as a drone reconnaissance of a certain area, an unmanned vehicle going to a certain location, or an inspection robot entering a certain pipeline.
[0061] S3, based on a preset time step, the collaborative behavior data is decomposed to obtain the decision instruction sequence corresponding to the collaborative behavior data.
[0062] It should be noted that the collaborative behavior data refers to the behavior of various heterogeneous devices over a period of time in the future. This step decomposes the collaborative behavior data based on a preset time step and converts the collaborative behavior data into a sequence of decision instructions arranged by time step that can be executed step by step by various heterogeneous devices.
[0063] Here, a time step refers to a pre-set time interval.
[0064] In some embodiments, step S3 includes S31-S33: S31, retrieve a predefined decision instruction library, which includes basic action instructions corresponding to various types of heterogeneous devices.
[0065] Among them, the decision instruction library refers to a collection of predefined and stored basic action instructions that can be executed by various types of heterogeneous devices; basic action instructions refer to action instructions that heterogeneous devices can directly execute.
[0066] It should be noted that drones, unmanned vehicles, and inspection robots can perform different actions. Drones can perform actions such as takeoff and hovering, unmanned vehicles can perform actions such as driving and loading, and inspection robots can perform actions such as rotating and grasping. The collaborative behavior data needs to be broken down to match executable actions for different types of heterogeneous devices.
[0067] Therefore, this step retrieves a predefined decision instruction library to provide executable basic action instructions for different types of heterogeneous devices.
[0068] S32, the collaborative behavior data is split on the time axis based on a preset time step to obtain the split behavior data corresponding to each time step.
[0069] Understandably, collaborative behavior data is broken down into segments based on preset time steps, resulting in segmented behavior data for each time step. For example, if a drone's collaborative behavior data is to fly to a certain coordinate within 5 seconds, and the time step is 1 second, this collaborative behavior data is broken down into 5 consecutive time steps, resulting in segmented behavior data for each time step. Each time step corresponds to the behavior that the drone should complete within that time step.
[0070] S33, match the split behavior data with the decision instruction library to obtain the decision instructions corresponding to each split behavior data, and concatenate the decision instructions to generate a decision instruction sequence.
[0071] Among them, decision instructions refer to the basic action instructions that are matched with the split behavior data and executed by heterogeneous devices at the corresponding time step.
[0072] It is understandable that, such as Figure 2 As shown, the split behavior data at each time step is matched with the basic action commands of the corresponding type of heterogeneous device in the decision command library to obtain the decision command corresponding to each split behavior data. For example, if the split behavior data of a drone at a certain time step is to move to a certain coordinate, this split behavior data is matched with the basic action command corresponding to the drone to obtain the decision command to move to the specified coordinate. The decision commands obtained from the matching of the same heterogeneous device at each time step are concatenated in chronological order to generate the decision command sequence of the heterogeneous device.
[0073] S4, issue decision instructions at the same time step in each decision instruction sequence to the corresponding heterogeneous devices, and synchronously issue decision instructions for the next time step according to the response signals of the heterogeneous devices.
[0074] It should be noted that the time required for heterogeneous devices such as drones, unmanned vehicles, and inspection robots to complete decision-making instructions at the same time step varies. If the decision-making instruction sequences of each heterogeneous device are sent out all at once and executed independently by each device, the execution progress of each device may be misaligned, leading to mutual waiting or collisions. This step sends out the decision-making instructions for the same time step from all heterogeneous devices together, and only sends out the decision-making instructions for the next time step after receiving response signals from all heterogeneous devices for the decision-making instructions of the current time step, so that all heterogeneous devices remain synchronized at the same time step.
[0075] Among them, the response signal refers to the signal fed back by the heterogeneous device after completing the decision instruction of the current time step.
[0076] In some embodiments, step S4 includes S41-S45: S41, obtain the decision instruction of the first time step in each decision instruction sequence as the execution instruction.
[0077] It is understandable that the decision instruction of the first time step in the decision instruction sequence of each heterogeneous device is obtained, and the obtained decision instruction of the first time step is used as the execution instruction.
[0078] S42, the execution instruction is sent to the corresponding heterogeneous device, and a response signal from the heterogeneous device indicating that the execution instruction has been completed is received.
[0079] It is understandable that each heterogeneous device executes the received execution instructions and sends back a response signal after completing the execution instructions, and receives the response signals from each heterogeneous device.
[0080] S43, for the issued execution command, the response signal received within the preset reception time period will be used as the currently received response signal.
[0081] It should be noted that after the execution command is issued, some heterogeneous devices may not respond for an extended period due to malfunctions or communication interruptions. Waiting indefinitely for responses from all heterogeneous devices could cause the entire process to stall. This step sets a preset reception duration, which is the pre-defined timeframe for receiving response signals. The response signals received within this preset duration are used as the currently received response signals, and subsequent commands are issued based on these signals.
[0082] Understandably, a preset reception duration is set, for example, a preset reception duration of 200 milliseconds. The response signal received within 200 milliseconds from the moment the execution command is issued is taken as the currently received response signal.
[0083] In some embodiments, step S43 (taking the response signal received within a preset reception duration as the currently received response signal) includes S431-S432: S431, obtain the sending time of the execution instruction, and obtain the cutoff time corresponding to the sending time based on the sending time and the preset receiving time.
[0084] It is understandable that the time when the execution command is sent is taken as the sending time, and the sending time is added to the preset receiving time to obtain the deadline time corresponding to the sending time.
[0085] S432, if a response signal is received before the deadline corresponding to the sending time, the corresponding response signal shall be used as the currently received response signal.
[0086] It is understandable that, for response signals fed back by various heterogeneous devices, if it is determined that the response signal was received before the deadline, the corresponding response signal will be taken as the currently received response signal.
[0087] S44: Based on the currently received response signal, obtain the decision instruction for the next time step in the corresponding decision instruction sequence, and use it as the current execution instruction.
[0088] It is understandable that, for the heterogeneous device corresponding to the currently received response signal, the decision instruction of the next time step in the decision instruction sequence of the corresponding heterogeneous device is obtained, and the obtained decision instruction of the next time step is used as the current execution instruction.
[0089] S45, repeat steps S42-S44 until all decision instructions for all time steps have been issued.
[0090] Understandably, steps S42 to S44 are repeated, with execution instructions sent and response signals received at each time step, until all decision instructions for all time steps have been issued.
[0091] It is worth mentioning that this invention uses heterogeneous devices that do not respond within a preset reception time as degraded devices, and re-plans the decision instruction sequence of the degraded devices to obtain a re-planned instruction sequence. Specifically, it acquires the device status of the degraded devices and the corresponding collaborative behavior data. The device status refers to the location of the degraded devices and the execution progress of the collaborative behavior. Based on the device status, the collaborative behavior data is adjusted to obtain adjusted collaborative behavior data. The adjusted collaborative behavior data is then decomposed to obtain the re-planning instruction sequence of the degraded devices.
[0092] In some embodiments, S51-S52 are also included: S51, obtain the actual execution data of the execution decision instruction sequence of each heterogeneous device.
[0093] It should be noted that the actual execution of decision command sequences by heterogeneous devices may differ from the collaborative behavior data. This step updates the parameters of the multi-agent decision network based on the differences between the actual execution data and the corresponding collaborative behavior data.
[0094] The actual execution data refers to the data actually generated after each heterogeneous device executes the decision instruction sequence, including the actual location reached by each heterogeneous device.
[0095] It is understandable that after each heterogeneous device executes the decision instruction sequence, the actual execution data of each heterogeneous device is obtained, and the actual position reached by each heterogeneous device at each time step forms the actual trajectory of each heterogeneous device.
[0096] S52, compare the actual execution data with the corresponding collaborative behavior data, update the parameters of the multi-agent decision network, and obtain the optimized multi-agent decision network.
[0097] Understandably, multiple time steps are considered as a single testing period, such as 10 time steps. The actual execution data and collaborative behavior data of each heterogeneous device within the testing period are compared to obtain the deviation between the actual trajectory corresponding to the actual execution data and the expected trajectory corresponding to the collaborative behavior data. Based on this deviation, the loss value is calculated using the loss function of the multi-agent decision network. The parameters of the evaluation network and action network in the multi-agent decision network are then updated according to the loss value, resulting in an optimized multi-agent decision network.
[0098] See Figure 3 This is a schematic diagram of a cross-aircraft autonomous collaborative decision-making system provided in an embodiment of the present invention. The system includes: The construction module is used to acquire the sampling data stream obtained by multiple heterogeneous devices sampling the sampling target, preprocess the sampling data stream to obtain the sampling array at the current time, and construct a situation map based on the sampling array and preset environmental prior knowledge; The computation module is used to solve the situation map based on the pre-trained multi-agent decision network to obtain the collaborative behavior data corresponding to each heterogeneous device; The decomposition module is used to decompose the collaborative behavior data based on a preset time step to obtain the decision instruction sequence corresponding to the collaborative behavior data; The decision module is used to issue decision instructions at the same time step in each decision instruction sequence to the corresponding heterogeneous devices, and to synchronously issue decision instructions for the next time step based on the response signals of the heterogeneous devices.
[0099] See Figure 4This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. The electronic device 40 includes: a processor 41, a memory 42, and a computer program; wherein... The memory 42 is used to store the computer program, and the memory may also be flash memory. The computer program is, for example, an application program or functional module that implements the above method.
[0100] The processor 41 is configured to execute the computer program stored in the memory to implement the various steps performed by the device in the above method. For details, please refer to the relevant descriptions in the preceding method embodiments.
[0101] Alternatively, the memory 42 can be either standalone or integrated with the processor 41.
[0102] When the memory 42 is a device independent of the processor 41, the device may further include: Bus 43 is used to connect the memory 42 and the processor 41.
[0103] The present invention also provides a readable storage medium storing a computer program, which, when executed by a processor, is used to implement the methods provided in the various embodiments described above.
[0104] The readable storage medium can be a computer storage medium or a communication medium. A communication medium includes any medium that facilitates the transfer of computer programs from one location to another. A computer storage medium can be any available medium accessible to a general-purpose or special-purpose computer. For example, a readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application-Specific Integrated Circuit (ASIC). Alternatively, the ASIC can be located in a user equipment. Of course, the processor and the readable storage medium can also exist as discrete components in a communication device. The readable storage medium can be a read-only memory (ROM), random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0105] The present invention also provides a program product including executable instructions stored in a readable storage medium. At least one processor of the device can read the executable instructions from the readable storage medium, and the at least one processor executes the executable instructions to cause the device to implement the methods provided in the various embodiments described above.
[0106] In the embodiments of the above-described device, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly manifested as execution by a hardware processor, or execution by a combination of hardware and software modules within the processor.
[0107] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A cross-aircraft autonomous collaborative decision-making method, characterized in that, include: S1, acquire sampling data streams obtained by multiple heterogeneous devices sampling the sampling target, preprocess the sampling data streams to obtain the sampling array at the current moment, and construct a situation map based on the sampling array and preset environmental prior knowledge; S2, Solve the situation map according to the pre-trained multi-agent decision network to obtain the cooperative behavior data corresponding to each heterogeneous device; S3, based on a preset time step, the collaborative behavior data is decomposed to obtain the decision instruction sequence corresponding to the collaborative behavior data; S4, issue decision instructions at the same time step in each decision instruction sequence to the corresponding heterogeneous devices, and synchronously issue decision instructions for the next time step according to the response signals of the heterogeneous devices.
2. The method according to claim 1, characterized in that, The preprocessing of the sampled data stream to obtain the current sampling array includes: Based on a preset reference frequency, time alignment processing is performed on each sampled data stream to obtain the aligned sampled data of each heterogeneous device at the current moment; The alignment sampling data is converted into image data of a preset size to obtain the size-adjusted alignment sampling data. The aligned sampled data after size adjustment are spliced together to obtain the sampled array at the current moment.
3. The method according to claim 1, characterized in that, The construction of the situation map based on the sampling array and preset environmental prior knowledge includes: Feature extraction is performed on the sampling array to obtain the situational features corresponding to each heterogeneous device and sampling target; Based on prior environmental knowledge, environmental constraint information corresponding to the situation features is retrieved, and corresponding environmental constraint information is added to the situation features to obtain the object features corresponding to each heterogeneous device and sampling target. Determine the spatial and interactive relationships between the heterogeneous devices, as well as between the heterogeneous devices and the sampling target; Using heterogeneous devices and sampling targets as nodes, nodes with spatial or interactive relationships are connected, and the object features are used as the node features of the corresponding nodes to obtain a situation map.
4. The method according to claim 1, characterized in that, Step S2 includes: Obtain the multi-agent decision network trained based on the multi-agent proximal policy optimization algorithm, and input the situation map into the multi-agent decision network; The multi-agent decision network is used to solve the input situation map and output the collaborative behavior data of each heterogeneous device.
5. The method according to claim 1, characterized in that, Step S3 includes: Retrieve a predefined decision instruction library, which includes basic action instructions corresponding to various types of heterogeneous devices; The collaborative behavior data is split along the time axis based on a preset time step to obtain the split behavior data corresponding to each time step. The split behavior data is matched with the decision instruction library to obtain the decision instructions corresponding to each split behavior data, and the decision instructions are concatenated to generate a decision instruction sequence.
6. The method according to claim 1, characterized in that, Step S4 includes: S41, obtain the decision instruction of the first time step in each decision instruction sequence as the execution instruction; S42, the execution instruction is sent to the corresponding heterogeneous device, and a response signal from the heterogeneous device indicating that the execution instruction has been completed is received; S43, for the issued execution command, the response signal received within the preset reception time period will be used as the currently received response signal; S44, based on the currently received response signal, obtain the decision instruction for the next time step in the corresponding decision instruction sequence, and use it as the current execution instruction; S45, repeat steps S42-S44 until all decision instructions for all time steps have been issued.
7. The method according to claim 6, characterized in that, The step of using the response signal received within a preset reception time as the currently received response signal includes: Obtain the sending time of the execution instruction, and based on the sending time and the preset receiving time, obtain the deadline time corresponding to the sending time; If a response signal is received before the deadline corresponding to the sending time, the corresponding response signal will be used as the currently received response signal.
8. A cross-aircraft autonomous collaborative decision-making system, characterized in that, include: The construction module is used to acquire the sampling data stream obtained by multiple heterogeneous devices sampling the sampling target, preprocess the sampling data stream to obtain the sampling array at the current time, and construct a situation map based on the sampling array and preset environmental prior knowledge; The computation module is used to solve the situation map based on the pre-trained multi-agent decision network to obtain the collaborative behavior data corresponding to each heterogeneous device; The decomposition module is used to decompose the collaborative behavior data based on a preset time step to obtain the decision instruction sequence corresponding to the collaborative behavior data; The decision module is used to issue decision instructions at the same time step in each decision instruction sequence to the corresponding heterogeneous devices, and to synchronously issue decision instructions for the next time step based on the response signals of the heterogeneous devices.
9. An electronic device, characterized in that, include: The method comprises a memory, a processor, and a computer program, wherein the computer program is stored in the memory and the processor executes the computer program to perform the method according to any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, is used to implement the method described in any one of claims 1 to 7.