A Method and System for Collaborative Search and Rescue by Quadruped Robot Dog Cluster Based on Multi-Agent Systems
Patent Information
- Application Number
- CN202610960735.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-11
AI Technical Summary
[0004]但基于集中式路径规划的方法依赖全局通信和较高计算资源,在通信受阻的废墟环境中鲁棒性较差
[0007]本发明公开的基于多智能体系统的四足机器狗集群协同搜救方法,根据实时获取的搜救区域内每一四足机器狗的单体感知数据,通过构建以机器狗为节点的动态搜救拓扑关系图进行处理,以得到表征集群内部空间关系与通信状态的搜救拓扑关系数据,实现了对集群动态协同状态的精细化建模;根据搜救区域的多源感知数据、单体感知数据及搜救拓扑关系数据,通过多层次特征融合进行处理,以得到融合全局环境、局部感知及拓扑协同信息的复合搜救状态数据,为后续协同决策提供了信息维度完备的数据基础;根据复合搜救状态数据,通过预设的强化学习主控智能体进行处理,以得到包含全局协同引导信息的多协同搜救约束矢量,为各机器狗的分布式决策提供了统一的全局协同约束;根据多协同搜救约束矢量及复合搜救状态数据,通过每一四足机器狗对应的强化学习执行智能体进行处理,以得到各机器狗的分布式步态参数与路径规划指令,实现了从全局协同约束到分布式动作指令的完整决策链路;根据分布式步态参数与路径规划指令实时获取搜救状态数据,通过结合搜救任务数据进行处理,以控制集群对搜救区域进行协同搜救,使得集群在未知复杂环境中能够实现全局协同一致、局部自适应灵活的高效搜救。
Smart Images

Figure CN122732907A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent robot and swarm control technology, specifically to a method and system for collaborative search and rescue of quadruped robot dogs based on a multi-agent system. Background Technology
[0002] Quadruped robot dogs, with their excellent obstacle-crossing ability and terrain adaptability, have become important search and rescue equipment in high-risk scenarios such as disaster relief and rubble search. However, due to limitations in sensor field of view, communication range, and battery life, a single quadruped robot dog cannot complete large-area, highly complex search and rescue tasks in a short period of time. Therefore, forming a cluster of multiple quadruped robot dogs to perform search and rescue tasks in a collaborative manner has become an inevitable technical path to improve search and rescue efficiency.
[0003] Currently, control methods for multi-robot collaborative search and rescue mainly fall into the following categories: First, methods based on centralized path planning, where a central controller plans the paths of all robots and issues them for execution. Second, methods based on pre-programmed formation control, such as master-slave following and virtual structure methods, where robots move according to preset formations and rules.
[0004] However, centralized path planning methods rely on global communication and high computing resources, resulting in poor robustness in rubble environments where communication is disrupted. Methods based on pre-programmed formation control struggle to cope with sudden environmental changes and cannot dynamically adjust formations and task allocations according to real-time search progress, easily leading to overlapping or missed search areas and thus reducing search and rescue efficiency. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention discloses a method and system for collaborative search and rescue using a cluster of quadruped robot dogs based on a multi-agent system, which aims to improve search and rescue efficiency.
[0006] To achieve the above objectives, this invention discloses a method for collaborative search and rescue using a cluster of quadruped robot dogs based on a multi-agent system, comprising: Real-time acquisition of individual perception data of each quadruped robot dog within the search and rescue area, so as to determine the search and rescue topology data of the quadruped robot dog cluster within the search and rescue area based on the individual perception data; Acquire multi-source sensing data of the search and rescue area, and determine the composite search and rescue status data of the quadruped robot dog cluster based on the multi-source sensing data, the individual sensing data, and the search and rescue topology data; Based on the composite search and rescue status data, the multi-cooperative search and rescue constraint vector of the quadruped robot dog cluster is obtained through a preset reinforcement learning master control agent. Based on the multi-cooperative search and rescue constraint vector and the composite search and rescue status data, the distributed gait parameters and path planning instructions of each quadruped robot dog are obtained through the reinforcement learning execution agent corresponding to each quadruped robot dog. The search and rescue status data of each quadruped robot dog is obtained in real time based on the distributed gait parameters and the path planning instructions, so as to control the quadruped robot dog cluster to carry out search and rescue in the search and rescue area according to the search and rescue status data and the search and rescue task data of the search and rescue area.
[0007] This invention discloses a quadruped robot dog swarm collaborative search and rescue method based on a multi-agent system. It utilizes real-time acquired individual perception data of each quadruped robot dog within the search and rescue area. This data is processed by constructing a dynamic search and rescue topology graph with the robot dogs as nodes to obtain search and rescue topology data characterizing the spatial relationships and communication status within the swarm, thus achieving refined modeling of the swarm's dynamic collaborative state. Furthermore, based on multi-source perception data, individual perception data, and search and rescue topology data of the search and rescue area, multi-level feature fusion is used to obtain composite search and rescue state data integrating global environment, local perception, and topology collaborative information, providing a complete data foundation for subsequent collaborative decision-making. Finally, based on the composite search and rescue state data, a preset reinforcement learning mechanism is used to... The system processes data through a control agent to obtain a multi-cooperative search and rescue constraint vector containing global collaborative guidance information, providing a unified global collaborative constraint for the distributed decision-making of each robot dog. Based on the multi-cooperative search and rescue constraint vector and composite search and rescue status data, the system processes the data through a reinforcement learning execution agent corresponding to each quadruped robot dog to obtain distributed gait parameters and path planning instructions for each robot dog, realizing a complete decision-making link from global collaborative constraints to distributed action instructions. Based on the distributed gait parameters and path planning instructions, the system acquires search and rescue status data in real time, and processes it in conjunction with search and rescue task data to control the cluster to conduct collaborative search and rescue in the search and rescue area. This enables the cluster to achieve efficient search and rescue with global coordination and local adaptation and flexibility in unknown and complex environments.
[0008] On the other hand, the present invention discloses a quadruped robot dog cluster collaborative search and rescue system based on a multi-agent system, including a single-agent perception module, a composite perception module, a global constraint module, a path planning module, and a search and rescue control module. The individual perception module is used to acquire the individual perception data of each quadruped robot dog in the search and rescue area in real time, so as to determine the search and rescue topology data of the quadruped robot dog cluster in the search and rescue area based on the individual perception data. The composite sensing module is used to acquire multi-source sensing data of the search and rescue area, so as to determine the composite search and rescue status data of the quadruped robot dog cluster based on the multi-source sensing data, the individual sensing data and the search and rescue topology data. The global constraint module is used to obtain the multi-cooperative search and rescue constraint vector of the quadruped robot dog cluster through a preset reinforcement learning master control agent based on the composite search and rescue status data. The path planning module is used to obtain the distributed gait parameters and path planning instructions of each quadruped robot dog through the reinforcement learning execution agent corresponding to each quadruped robot dog, based on the multi-cooperative search and rescue constraint vector and the composite search and rescue status data. The search and rescue control module is used to acquire the search and rescue status data of each quadruped robot dog in real time according to the distributed gait parameters and the path planning instructions, so as to control the quadruped robot dog cluster to carry out search and rescue in the search and rescue area according to the search and rescue status data and the search and rescue task data of the search and rescue area.
[0009] This invention discloses a quadruped robot dog swarm collaborative search and rescue system based on a multi-agent system. It utilizes real-time acquired individual perception data of each quadruped robot dog within the search and rescue area. This data is processed by constructing a dynamic search and rescue topology graph with the robot dogs as nodes to obtain search and rescue topology data representing the spatial relationships and communication status within the swarm, thus achieving refined modeling of the swarm's dynamic collaborative state. Furthermore, based on multi-source perception data, individual perception data, and search and rescue topology data of the search and rescue area, multi-level feature fusion is used to obtain composite search and rescue state data integrating global environment, local perception, and topology collaborative information, providing a complete data foundation for subsequent collaborative decision-making. Finally, based on the composite search and rescue state data, a preset reinforcement learning mechanism is used to... The system processes data through a control agent to obtain a multi-cooperative search and rescue constraint vector containing global collaborative guidance information, providing a unified global collaborative constraint for the distributed decision-making of each robot dog. Based on the multi-cooperative search and rescue constraint vector and composite search and rescue status data, the system processes the data through a reinforcement learning execution agent corresponding to each quadruped robot dog to obtain distributed gait parameters and path planning instructions for each robot dog, realizing a complete decision-making link from global collaborative constraints to distributed action instructions. Based on the distributed gait parameters and path planning instructions, the system acquires search and rescue status data in real time, and processes it in conjunction with search and rescue task data to control the cluster to conduct collaborative search and rescue in the search and rescue area. This enables the cluster to achieve efficient search and rescue with global coordination and local adaptation and flexibility in unknown and complex environments. Attached Figure Description
[0010] Figure 1 This is a flowchart illustrating the collaborative search and rescue method for quadruped robot dogs based on a multi-agent system disclosed in an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of the quadruped robot dog cluster collaborative search and rescue system based on a multi-agent system disclosed in an embodiment of the present invention. Detailed Implementation
[0011] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0012] See Figure 1 To improve search and rescue efficiency, one embodiment of the present invention provides a method for collaborative search and rescue using a cluster of quadruped robot dogs based on a multi-agent system, comprising: Step 101: Acquire the individual perception data of each quadruped robot dog in the search and rescue area in real time, so as to determine the search and rescue topology data of the quadruped robot dog cluster in the search and rescue area based on the individual perception data.
[0013] Step 102: Acquire multi-source perception data of the search and rescue area to determine the composite search and rescue status data of the quadruped robot dog cluster based on multi-source perception data, individual perception data and search and rescue topology data.
[0014] Step 103: Based on the composite search and rescue status data, obtain the multi-cooperative search and rescue constraint vector of the quadruped robot dog cluster through the preset reinforcement learning master control agent.
[0015] Step 104: Based on the multi-cooperative search and rescue constraint vector and composite search and rescue status data, obtain the distributed gait parameters and path planning instructions of each quadruped robot dog through the reinforcement learning execution agent corresponding to each quadruped robot dog.
[0016] Step 105: Obtain the search and rescue status data of each quadruped robot dog in real time based on the distributed gait parameters and path planning instructions, so as to control the quadruped robot dog cluster to carry out search and rescue in the search and rescue area based on the search and rescue status data and the search and rescue task data of the search and rescue area.
[0017] In this embodiment, individual perception data of each quadruped robot dog within the search and rescue area is acquired in real time, and the search and rescue topology data of the quadruped robot dog cluster is determined based on this individual perception data. Specifically, each quadruped robot dog can be equipped with a basic distance sensor and a communication module. When the quadruped robot dog moves within the search and rescue area, it periodically sends detection signals to the surrounding environment and receives signals from other robot dogs. By analyzing the strength or arrival time of the received signals, each robot dog can determine its relative distance to other robot dogs in the vicinity; after this relative distance information is collected, it can be used to construct a connection graph, in which robot dogs with a distance within a preset threshold are considered to be connected, thereby forming search and rescue topology data.
[0018] Secondly, multi-source perception data of the search and rescue area is acquired to determine the composite search and rescue status data of the quadruped robot dog cluster based on this multi-source perception data, the individual perception data, and the search and rescue topology data. Specifically, the multi-source perception data of the search and rescue area can be provided by external devices, such as setting up a drone to fly over the search and rescue area, taking pictures and transmitting aerial images of the area. These aerial images can be processed to identify large obstacle areas or explored areas. At the same time, individual perception data of each quadruped robot dog is collected. Then, the global image information provided by the drone is fused with the local information reported by each robot dog and the search and rescue topology data determined in the above manner to obtain composite search and rescue status data. For example, the robot dog's position can be marked on the aerial image, and different colors can be used to distinguish explored and unexplored areas, thereby forming a preliminary composite search and rescue status data.
[0019] Furthermore, based on the composite search and rescue status data, a pre-defined reinforcement learning master agent is used to obtain multi-cooperative search and rescue constraint vectors for the quadruped robot dog swarm. Specifically, the reinforcement learning master agent can be pre-trained, and its input is the aforementioned composite search and rescue status data. This master agent may contain a neural network model that learns from historical data and outputs cooperative constraints. For example, when the composite search and rescue status data indicates that the robot dog density in a certain area is too high, the master agent can output dispersed constraint vectors; when a certain area has not been explored for a long time, it can output concentrated exploration constraint vectors. These constraint vectors can be numerical values or flags, used to indicate the overall behavioral direction of the swarm.
[0020] Based on this, according to the multi-cooperative search and rescue constraint vector and the composite search and rescue state data, the distributed gait parameters and path planning instructions for each quadruped robot dog are obtained through the reinforcement learning execution agent corresponding to each quadruped robot dog. Specifically, each quadruped robot dog is equipped with a reinforcement learning execution agent. This execution agent receives the multi-cooperative search and rescue constraint vector from the master agent and plans its path and gait by combining it with its own composite search and rescue state data. The execution agent contains a neural network, which generates its current gait parameters (e.g., walking speed, turning angle) and path planning instructions (e.g., coordinates of the next target point) according to a preset strategy based on the received constraints and its own state.
[0021] Finally, the search and rescue status data of each quadruped robot dog is acquired in real time based on the distributed gait parameters and the path planning instructions. This data, along with the search and rescue task data for the search area, is then used to control the quadruped robot dog cluster to conduct search and rescue operations within that area. Specifically, each quadruped robot dog moves within the search area according to its received distributed gait parameters and path planning instructions. During movement, the robot dog continuously acquires its real-time position, posture, and whether any abnormalities (e.g., heat sources, sounds) through its sensors. This information constitutes the search and rescue status data for each robot dog. Simultaneously, the search and rescue task data for the search area (e.g., finding all survivors or covering the entire area) is input to the central control unit. Based on the search and rescue status data and task data of each robot dog, the central control unit controls the overall search and rescue behavior of the quadruped robot dog cluster through preset rules or manual intervention. For example, if a robot dog reports detecting a suspected target, the central control unit can instruct other robot dogs to go to the area to provide support or confirmation.
[0022] In this embodiment, step 101 includes: Step 1011: For any quadruped robot dog, after detecting the movement of the quadruped robot dog, acquire the individual perception data of the quadruped robot dog; wherein, the individual perception data includes at least its own posture data, local environmental obstacle distribution data, and target cue detection data; Step 1012: Obtain the real-time position coordinates and movement speed of the quadruped robot dog based on its own posture data, and fuse the real-time position coordinates, movement speed, and local environmental obstacle distribution data to obtain the node attribute feature vector of the quadruped robot dog; Step 1013: Match the quadruped robot dog based on its real-time position coordinates. The neighboring robot dogs are identified, and the communication status between the quadruped robot dog and the neighboring robot dogs is obtained. Based on the communication status, a connection edge is established between the quadruped robot dog and the neighboring robot dogs. Step 1014: For any connection edge, the relative search and rescue status data between the two quadruped robot dogs corresponding to the connection edge is obtained based on the node attribute feature vector, and the relative search and rescue status data is used as the edge attribute feature vector of the connection edge. Step 1015: Using the quadruped robot dogs as nodes, a dynamic search and rescue topology graph of the quadruped robot dog cluster is constructed based on the connection edge, edge attribute feature vector, and node attribute feature vector, as the search and rescue topology data of the quadruped robot dog cluster.
[0023] In this embodiment, taking a quadruped robot dog, including three quadruped robots, conducting search and rescue operations in a post-disaster rubble search and rescue scenario as an example, the three quadruped robots enter the search and rescue area in a preset formation and begin to move. Taking the first quadruped robot dog as an example, when its built-in inertial motion sensor detects that the body has accelerated and its position coordinates have changed, it is determined that the first quadruped robot dog has entered a moving state. At this time, the first quadruped robot dog synchronously collects individual perception data through its onboard depth camera or LiDAR and other data acquisition units.
[0024] Next, local environmental obstacle distribution data is extracted from depth images acquired by the depth camera. A semantic segmentation network is used to identify obstacle categories such as "rubble piles," "broken walls," and "tilted floors," along with their bounding boxes. The distance and orientation of each obstacle relative to the first robot dog are calculated to construct local environmental obstacle distribution data or to obtain detailed 3D terrain information from the LiDAR point cloud. The robot dog's own attitude data, including its three-axis acceleration, angular velocity, and attitude quaternions, is acquired using its inertial motion sensors. Simultaneously, target cue detection data is collected using the onboard infrared thermal imager and microphone array to detect the presence of human heat sources or distress signals. The real-time position coordinates and movement speed of the first robot dog are calculated from its own attitude data. The real-time position coordinates, movement speed, and local environmental obstacle distribution data (obstacle type, distance, and orientation) are concatenated into a multi-dimensional vector, serving as the node attribute feature vector of the first robot dog.
[0025] Based on the real-time location coordinates of the first robot dog, and using the communication radius as the neighborhood range, it searches whether other robot dogs are within this range. Heartbeat packets are sent to other robot dogs via the MESH networking module. The communication status is determined based on the received signal strength (RSSI value) and round-trip time, and connection edges are constructed between the first robot dog and its neighboring robot dogs based on the communication status.
[0026] Taking the connecting edge between the first robot dog and the second robot dog as an example, the relative search and rescue status data between the two robots is calculated based on the node attribute feature vectors of the first robot dog and the second robot dog respectively. For example, the Euclidean distance is calculated based on the coordinate difference to serve as the relative distance, relative speed, and relative azimuth angle (the azimuth of the second robot dog relative to the first robot dog is 45° east of north). These data are used as the edge attribute feature vector of the connecting edge.
[0027] Using the first, second, and third robotic dogs as nodes, a dynamic search and rescue topology graph is constructed based on the established connecting edges (first robotic dog-second robotic dog and first robotic dog-third robotic dog), the edge attribute feature vectors of each edge, and the node attribute feature vectors of each node. This topology graph is stored in the form of an adjacency matrix and a feature matrix, and is updated in real time as each robotic dog moves, serving as search and rescue topology data for subsequent modules to access.
[0028] The above implementation method uses real-time posture data, local environmental obstacle distribution data, and target cue detection data collected after each quadruped robot dog moves. By fusing real-time position coordinates, movement speed, and local environmental obstacle distribution data, it obtains node attribute feature vectors representing the individual state of each robot dog node, effectively integrating multi-dimensional perception information of the robot dog nodes. It matches neighboring robot dogs based on real-time position coordinates and obtains their communication status. By establishing connecting edges between robot dogs, it obtains topological connections representing the communication reachability between them, achieving explicit modeling of the cluster communication topology. Based on the node attribute feature vectors, it obtains relative search and rescue status data between the two robot dogs corresponding to the connecting edges. By processing the relative search and rescue status data as edge attribute feature vectors, it obtains interactive information representing the spatial and motion relationships between robot dogs, achieving a refined description of the collaborative relationship between them. Based on the connecting edges, edge attribute feature vectors, and node attribute feature vectors, it constructs a dynamic search and rescue topology graph to obtain search and rescue topology data representing the real-time collaborative state of the cluster, providing dynamic and structured topological support for the subsequent construction of composite search and rescue status data.
[0029] In this embodiment, step 102 includes: Step 1021: Rasterizing and semantically segmenting the multi-source environmental perception data of the search and rescue area to obtain a global environmental situation map of the search and rescue area; wherein, the global environmental situation map includes known obstacle distribution areas, searched areas, and unsearched areas; Step 1022: Extracting features from the individual perception data of each quadruped robot dog to generate a local search and rescue feature vector for each quadruped robot dog; Step 1023: Aggregating the neighborhood search and rescue information of each quadruped robot dog through a preset graph neural network based on the local search and rescue feature vector and the dynamic search and rescue topology diagram to obtain a neighborhood search and rescue feature vector for each quadruped robot dog; Step 1024: Aligning and stitching the local search and rescue feature vector and neighborhood search and rescue feature vector of each quadruped robot dog with the global environmental situation map to obtain composite search and rescue status data of each quadruped robot dog in the search and rescue area.
[0030] In this embodiment, a drone equipped with a high-definition aerial camera and LiDAR can perform high-altitude scanning of the search and rescue area to acquire high-precision optical orthophotos and 3D point cloud data, which serve as multi-source environmental perception data for the search and rescue area. Next, the multi-source environmental perception data is rasterized and semantically segmented to obtain a global environmental situation map of the search and rescue area. In this embodiment, the search and rescue area is divided into a multi-row, multi-column raster map, with each grid having the same area. Then, a pre-trained semantic segmentation model (based on the U-Net architecture) processes the optical images, classifying each grid into three types: "known obstacles," "searched area," or "unsearched area." The searched area refers to an area where at least one robot dog has traversed and collected data. Finally, a global environmental situation map is generated, presented as a color-coded image, which can be queried and downloaded by each robot dog.
[0031] Furthermore, feature extraction is performed on the individual perception data of each quadruped robot dog. Specifically, taking the first robot dog as an example, its individual perception data (its own posture, obstacle distribution, and target cues) collected within a recent period (e.g., 1 second) is input into a pre-trained convolutional neural network (CNN) for encoding, outputting a high-dimensional local search and rescue feature vector. This vector contains information such as the terrain accessibility assessment, obstacle density, and target presence probability within a preset range at the first robot dog's current location.
[0032] Based on the local search and rescue feature vector and the dynamic search and rescue topology graph, a pre-defined graph neural network aggregates neighborhood search and rescue information. This graph neural network uses a two-layer structure, starting with the local search and rescue feature vector of the first robot dog as the initial node feature. It performs message passing (two-hop iteration) on the dynamic search and rescue topology graph, aggregating the local search and rescue feature information of the second and third robot dogs to obtain the neighborhood search and rescue feature vector of the first robot dog. This vector contains the recent search progress and terrain assessment information of the second and third robot dogs. The local search and rescue feature vector of the first robot dog, the neighborhood search and rescue feature vector, and the encoded vector of the global environmental situation map are concatenated, and feature alignment and dimensionality reduction are performed through a fully connected layer to output the composite search and rescue status data of the first robot dog within the search and rescue area. The same operation is performed on the other robot dogs in the cluster to obtain their respective corresponding composite search and rescue status data.
[0033] The above implementation method processes multi-source environmental perception data of the search and rescue area through rasterization and semantic segmentation to obtain a global environmental situation map containing known obstacle distribution areas, searched areas, and unsearched areas, thus achieving semantic modeling of the overall environment of the search and rescue area. Based on the individual perception data of each quadruped robot dog, feature extraction is performed to obtain local search and rescue feature vectors representing the local perception information of each robot dog, achieving standardized encoding of individual perception data. Based on the local search and rescue feature vectors and the dynamic search and rescue topology graph, neighborhood search and rescue information is aggregated through a message passing mechanism in a pre-set graph neural network to obtain the neighborhood search and rescue feature vectors of each quadruped robot dog, achieving effective extraction of cluster collaborative structure information. Based on the local search and rescue feature vectors, neighborhood search and rescue feature vectors, and the global environmental situation map, feature alignment and stitching are performed to obtain composite search and rescue status data integrating individual perception, cluster collaboration, and global environmental three-dimensional information, providing complete input data for the subsequent accurate decision-making of the reinforcement learning agent.
[0034] In this embodiment, step 103 includes: Step 1031: For any quadruped robot dog, obtain the global interaction weight between each neighboring robot dog and the quadruped robot dog based on the neighborhood search and rescue feature vector and the multi-head self-attention mechanism preset in the reinforcement learning master control agent; Step 1032: Based on the local search and rescue feature vector, the dynamic search and rescue topology graph, and the global environmental situation map, perform multi-task processing through the multi-task decoupling network preset in the reinforcement learning master control agent to obtain the region segmentation boundary weight parameter, formation maintenance constraint parameter, and information sharing strength coefficient between each neighboring robot dog and the quadruped robot dog; Step 1033: Based on the global interaction weight, region segmentation boundary weight parameter, formation maintenance constraint parameter, and information sharing strength coefficient corresponding to each quadruped robot dog in the quadruped robot dog cluster, output the multi-cooperative search and rescue constraint vector of the quadruped robot dog cluster through the reinforcement learning master control agent.
[0035] In this embodiment, the reinforcement learning master agent is a decision-making entity designed based on the reinforcement learning paradigm. Its core lies in learning the optimal policy through interaction with the environment to maximize long-term cumulative rewards. In this application, it is responsible for coordinating the overall behavior of the quadruped robot dog swarm at a macro level and generating global collaborative constraints. This agent can be implemented based on a deep Q-network (DQN) or an actor-critic architecture, where the state input is composite search and rescue state data and the action output is a multi-collaborative search and rescue constraint vector. It is trained through interaction with a simulation environment or a real environment to optimize the swarm's search and rescue efficiency. Alternatively, this agent can also employ a policy gradient-based method, such as proximal policy optimization (PPO) or soft actor-critic (SAC), using the MAPPO algorithm (centralized value function) to output continuous constraint vectors, thereby better adapting to the demands of complex search and rescue tasks.
[0036] In this embodiment, a multi-head self-attention mechanism is used to evaluate the importance of interactions between an individual quadruped robot dog and its neighboring robot dogs. This mechanism can consist of multiple independent self-attention modules. Each module generates attention weights through query, key, and value vector calculations. The outputs of multiple heads are then concatenated and linearly transformed to obtain the final global interaction weights. Alternatively, the encoder portion of the Transformer architecture can be used, utilizing its multi-head self-attention layer to process neighborhood search and rescue feature vectors. By learning the influence of different neighboring robot dogs on the current robot dog's decision-making, the global interaction strength between them can be quantified. The global interaction weight represents the degree of mutual influence or importance between any quadruped robot dog and its neighboring robot dogs in a collaborative search and rescue task within the quadruped robot dog cluster. This weight ensures that important neighboring robot dogs receive higher priority in information sharing and task allocation, thereby optimizing the overall collaborative efficiency of the cluster.
[0037] In this embodiment, a multi-task decoupling network is used to separate constraint parameters of different dimensions, such as region segmentation, formation preservation, and information sharing, from a complex search and rescue situation. The multi-task decoupling network employs a hard parameter sharing approach, where all tasks share a single encoder, and each task has an independent decoder or output head to predict the region segmentation boundary weight parameters, formation preservation constraint parameters, and information sharing strength coefficients, respectively. Alternatively, a soft parameter sharing approach can be used, where each task has its own network, but techniques such as regularization or distillation are used to encourage similar network parameters, or a gating mechanism is used to dynamically select shared or independent features to achieve effective task decoupling.
[0038] The region segmentation boundary weight parameter defines the boundary strength or priority when a quadruped robot dog swarm assigns tasks and divides regions within a search and rescue area. This parameter guides the robot dogs on how to effectively divide search areas, avoid duplicate searches or omissions, and dynamically adjusts their sensitivity to region boundaries. Preferably, when a target clue is discovered, the boundary weight of the target region can be increased to encourage neighboring robot dogs to move closer to that region, or the weight can be reduced at the edge of an already searched region to prevent repeated entry. The formation maintenance constraint parameter regulates the strength of the quadruped robot dog swarm in maintaining a specific formation or relative positional relationship during movement and search and rescue. The information sharing intensity coefficient adjusts the frequency, scope, or importance of information exchange among quadruped robot dog swarm members. It determines the extent to which a robot dog needs to broadcast its own sensory data or receive neighboring information and dynamically balances the efficiency of information sharing with the communication load. For example, when critical information (such as the discovery of a suspected target) appears, the information sharing intensity can be increased to ensure rapid information dissemination; while during routine patrols, the intensity can be reduced to save communication resources. The multi-cooperative search and rescue constraint vector is a comprehensive vector that encompasses various cooperative rules and constraints that the quadruped robot dog swarm must follow when performing search and rescue missions. This vector is dynamically generated and can be adjusted according to the real-time environment and mission status. It serves as the basis for generating subsequent distributed gait parameters and path planning instructions, providing global guidance and local constraints for each quadruped robot dog, ensuring that the swarm can efficiently and robustly complete search and rescue missions in complex and ever-changing environments.
[0039] In this embodiment, during the cluster-based collaborative search and rescue process, the master control agent communicates with each robot dog via a network. Taking the first robot dog as an example, the reinforcement learning master control agent extracts its neighborhood search and rescue feature vector from the received composite search and rescue status data. Simultaneously, the master control agent also receives composite search and rescue status data from other robot dogs. Next, the neighborhood search and rescue feature vector of the first robot dog (as the query vector) and the neighborhood search and rescue feature vectors of the other robot dogs (as the key vector and value vector) are input into a multi-head self-attention module to calculate the dot product similarity between the query vector and the key vector. After normalization, the global interaction weight between the first robot dog and each of the other robot dogs is obtained (e.g., the weight between the first and second robot dogs is 0.65, and the weight between the first and third robot dogs is 0.35). This indicates that, in the current state, the search decision of the first robot dog should be more closely coordinated with that of the second robot dog.
[0040] Simultaneously, the encoded vectors of the first robot dog's local search and rescue feature vector, dynamic search and rescue topology graph (adjacency matrix), and global environmental situation map are concatenated and input into the multi-task decoupling network. This network contains three parallel sub-task heads: a region segmentation head, which outputs the region segmentation boundary weight parameters of the first robot dog relative to the second and third robot dogs; a formation preservation head, which outputs the formation preservation constraint parameters; and an information sharing head, which outputs the information sharing strength coefficient (for example, the first robot dog should share 80% of its search progress data with the second robot dog and 60% with the third robot dog).
[0041] Then, the multi-task decoupling network calculates the coupling coefficient between any two of the three sets of parameters mentioned above. For example, the coupling coefficient between the region segmentation boundary weight parameter and the formation preservation constraint parameter is 0.4, indicating that when adjusting the formation spacing, the region segmentation boundary also needs to be fine-tuned accordingly to avoid conflicts. Finally, each parameter is compensated and corrected according to the coupling coefficient (for example, the region segmentation boundary weight between the first and second robot dogs is corrected from 0.8 to 0.75). The corrected parameters are combined with the global interaction weight to form the multi-cooperative search and rescue constraint vector corresponding to the first robot dog. The same operation is performed on the other robot dogs in the cluster, and the master control agent outputs the cluster's multi-cooperative search and rescue constraint vector.
[0042] The above implementation method processes the neighborhood search and rescue feature vector of each quadruped robot dog through a multi-head self-attention mechanism preset in the reinforcement learning master agent to obtain the global interaction weight between each neighboring robot dog and the quadruped robot dog, thus achieving a comprehensive modeling of the complex interaction relationships within the cluster. Based on the local search and rescue feature vector, the dynamic search and rescue topology graph, and the global environmental situation map, multi-task processing is performed through a multi-task decoupling network preset in the reinforcement learning master agent to obtain the region segmentation boundary weight parameters, formation maintenance constraint parameters, and information sharing strength coefficient, thereby decomposing the complex global collaborative decision into multiple independently optimizable sub-tasks. Based on the global interaction weight, region segmentation boundary weight parameters, formation maintenance constraint parameters, and information sharing strength coefficient corresponding to each robot dog in the quadruped robot dog cluster, the reinforcement learning master agent performs combined processing to output a coordinated multi-cooperative search and rescue constraint vector, enabling effective coordination among multiple collaborative sub-objectives and providing accurate global guidance signals for the distributed decision-making of each executing agent, thereby improving the accuracy and efficiency of search and rescue.
[0043] In this embodiment, step 104 includes: Step 1041: For any quadruped robot dog, the local search and rescue feature vector of the quadruped robot dog is weighted and adjusted according to the region segmentation boundary weight parameter and information sharing intensity coefficient of the quadruped robot dog to obtain the quadruped robot dog's own search and rescue feature; Step 1042: Using the self-search and rescue feature as the query vector and the neighborhood search and rescue feature vector as the key vector and value vector, the collaborative attention weight between the quadruped robot dog and each neighboring robot dog is calculated through a pre-set collaborative attention mechanism in the reinforcement learning execution agent; Step 1043: The neighborhood search and rescue feature of the quadruped robot dog is calculated according to the collaborative attention weight. The vectors are weighted and aggregated to obtain the neighborhood search and rescue features of the quadruped robot dog; Step 1044: The quadruped robot dog's own search and rescue features and neighborhood search and rescue features are gated and fused according to the reinforcement learning executive agent corresponding to the quadruped robot dog to obtain the comprehensive search and rescue state representation vector of the quadruped robot dog; Step 1045: The comprehensive search and rescue state representation vector is input into the parameterized action network of the reinforcement learning executive agent to output the distributed gait parameters and path planning instructions of the quadruped robot dog; wherein, during the training process of the reinforcement learning executive agent, the network parameters of the parameterized action network are determined according to the composite reward function including search and rescue coverage, response time, energy consumption and task balance indicators.
[0044] In this embodiment, the various dimensions of the local search and rescue feature vector are linearly weighted and combined with the region segmentation boundary weight parameter and the information sharing intensity coefficient. Matrix multiplication or element-wise multiplication is used to highlight or suppress certain features. Alternatively, a neural network is used to take the local search and rescue feature vector, the region segmentation boundary weight parameter, and the information sharing intensity coefficient as input, and learns to obtain more representative self-search and rescue features.
[0045] Next, a self-attention mechanism based on the Transformer architecture is adopted, using the self-rescue feature as the query vector and the neighborhood rescue feature vector as the key vector and value vector. The attention weight is obtained by calculating the dot product of the query and the key and performing softmax normalization. Alternatively, the attention mechanism in the Graph Attention Network (GAT) can be adopted, which calculates the attention coefficient between the node and its neighbors through a shared linear transformation and the LeakyReLU activation function, thereby aggregating the neighbor features.
[0046] Furthermore, by calculating the similarity (e.g., dot product) between the self-rescue features (query vector) and the neighborhood search and rescue feature vectors (key vectors) of each neighboring robot dog, and then normalizing them using the softmax function, the attention weights of each neighboring robot dog can be obtained. Alternatively, a learnable weight matrix can be introduced, and the self-rescue features and neighborhood search and rescue feature vectors can be concatenated and transformed using this matrix, then passed through an activation function and a fully connected layer to obtain the attention weights. The neighborhood search and rescue feature vectors of the quadruped robot dog are weighted and aggregated according to the collaborative attention weights to obtain the neighborhood search and rescue features of the quadruped robot dog. Then, the neighborhood search and rescue feature vector of each neighboring robot dog is multiplied element-wise with its corresponding collaborative attention weight, and all weighted neighborhood feature vectors are summed to obtain the final neighborhood search and rescue features.
[0047] Furthermore, a gating mechanism similar to that in Long Short-Term Memory (LSTM) networks or Gated Recurrent Units (GRUs) is employed. The learned gating weights determine the fusion ratio of the self-rescue features and the neighborhood search and rescue features. Preferably, the self-rescue features and the neighborhood search and rescue features are used as inputs, and a fusion gating vector is generated through a sigmoid activation function. This vector is then used to weight and sum the two features. The comprehensive search and rescue state representation vector is a comprehensive, high-dimensional abstract representation of the quadruped robot dog's current state, local environment, neighborhood information, and global collaborative constraints. It integrates individual and collective information, providing a foundation for subsequent decision-making. It can be a fixed-dimensional floating-point vector, compressed and encoded by the encoding layer of a neural network to obtain all relevant information. The parameterized action network is a policy network in reinforcement learning. It takes the agent's state representation as input and outputs the action (or the probability distribution of the action) that the agent should take in the current state.
[0048] It's important to note that parameterization means the network's behavior is determined by its internal parameters (such as weights and biases), which are optimized through training. For example, it could be a multilayer perceptron (MLP) whose input is a comprehensive search and rescue state representation vector, and whose output is distributed gait parameters (such as joint angles and torques) and path planning instructions (such as target points and velocities); or it could be a deep neural network containing convolutional or recurrent layers to handle more complex temporal or spatial information and output more refined action instructions. The composite reward function is a metric used during reinforcement learning training to evaluate the agent's behavior. It comprehensively considers multiple sub-objectives related to the task objective to guide the agent to learn the optimal policy. Specifically, search and rescue coverage measures the extent to which the quadruped robot dog swarm explores the search and rescue area, encouraging the robots to explore unsearched areas; response time measures the robot dog's reaction speed to the search and rescue target or unexpected events, encouraging rapid response; energy consumption measures the energy consumed by the robot dog during task execution, encouraging energy-saving behavior; and task balance measures the evenness of task distribution among the robots in the swarm, avoiding overload or idleness for some robots. For example, the above indicators can be linearly weighted and summed to obtain the final composite reward value; or, a non-linear combination or hierarchical reward mechanism can be adopted, such as optimizing response time and energy consumption after a certain coverage rate is achieved.
[0049] In this embodiment, taking the first robot dog as an example, a dedicated reinforcement learning execution agent is deployed on its onboard computing unit. First, the execution agent extracts the region segmentation boundary weight parameters and information sharing intensity coefficients corresponding to the first robot dog from the multi-cooperative search and rescue constraint vectors issued by the master agent. It then extracts the local search and rescue feature vectors and neighborhood search and rescue feature vectors of the first robot dog from the composite search and rescue state data. The local search and rescue feature vectors are then weighted according to the region segmentation boundary weight parameters and the information sharing intensity coefficients: the mean of the region segmentation boundary weight parameters is used as an adjustment factor and multiplied element-wise with the local search and rescue feature vectors to obtain the task-guided self-search and rescue features, which have been modulated by the global cooperative target.
[0050] Using its own search and rescue features as the query vector Q, and the neighboring search and rescue feature vectors of other robot dogs as the key vector K and value vector V, the inputs are fed into the collaborative attention module. The similarity between Q and each K is calculated, and after softmax normalization, the collaborative attention weights are obtained (e.g., the attention weight for the second robot dog is 0.72, and the attention weight for the third robot dog is 0.28). V is weighted and aggregated according to the attention weights to obtain the neighboring search and rescue features of the first robot dog. Gated fusion is performed on the self-search and rescue features and the neighboring search and rescue features: the gating coefficient g (ranging from 0 to 1) is calculated through a fully connected layer and a sigmoid activation function; the self-feature is multiplied by g, and the neighboring features are multiplied by (1-g), then concatenated to obtain the comprehensive search and rescue state representation vector.
[0051] The comprehensive search and rescue state representation vector is input into a parameterized action network. This network consists of two branches: a continuous action branch (3 fully connected layers, with Tanh activation in the output layer): outputs distributed gait parameters, which are linearly mapped to obtain actual values—step frequency (corresponding to medium-speed walking in rubble terrain), stride length (adapting to rubble ground), and body height (lowering the center of gravity to improve stability); and a discrete action branch (3 fully connected layers, with Softmax activation in the output layer): outputs the probability distribution of path planning instructions, which are sampled by argmax to obtain specific actions—target heading angle, search speed level, and search mode.
[0052] It is important to note that the network parameters of the parameterized action network are determined during the training phase by optimizing the composite reward function. The composite reward function includes four metrics: search and rescue coverage reward (+1.0 for every 1% increase in coverage), response time penalty (-0.5 for every minute exceeding the preset time), energy consumption penalty (-0.2 for every 1% of battery consumed), and task balance reward (higher reward for more balanced area distribution, capped at +2.0). After training, the network parameters internalize the optimization capabilities for these four metrics, enabling the robot dog to autonomously make optimal decisions that balance efficiency, timeliness, energy saving, and balance in practical applications.
[0053] The above implementation method processes the local search and rescue feature vectors of each quadruped robot dog by weighting and adjusting the region segmentation boundary weight parameters and information sharing intensity coefficients to obtain self-search and rescue features that integrate global task guidance information. This allows the self-feature encoding of each robot dog to selectively integrate into the global collaborative target. Using the self-search and rescue features as the query vector and the neighborhood search and rescue feature vectors as the key and value vectors, a pre-set collaborative attention mechanism in the reinforcement learning agent is used to calculate the collaborative attention weights between the quadruped robot dog and its neighboring robot dogs, achieving selective attention to collaborative information within the neighborhood. The neighborhood search and rescue feature vectors are then weighted and aggregated according to the collaborative attention weights to obtain the neighborhood search and rescue features of the quadruped robot dog, enabling the robot dog to... The system dynamically focuses on neighbor information most valuable to its search and rescue decisions. Based on its own search and rescue characteristics and those of its neighboring areas, it uses gating fusion to obtain a comprehensive search and rescue state representation vector, enabling the robot dog to adaptively adjust its dependence on its own and neighboring states. This comprehensive search and rescue state representation vector is then processed through a parameterized action network of the executive agent using reinforcement learning to output distributed gait parameters and path planning instructions, achieving a complete decision output from high-level collaborative constraints to distributed action instructions. The network parameters of the parameterized action network are optimized during training using a composite reward function that includes search and rescue coverage, response time, energy consumption, and task balance indicators, ensuring that the executive agent can output optimal action instructions that balance multi-dimensional optimization goals in practical applications.
[0054] In this embodiment, step 105 includes: Step 1051: For each quadruped robot dog, when controlling the quadruped robot dog to move within the search and rescue area according to distributed gait parameters and path planning instructions, the search and rescue status data of the quadruped robot dog is obtained based on the individual perception data of the quadruped robot dog; Step 1052: When it is determined from the search and rescue status data that the quadruped robot dog has detected a suspected search and rescue target, the visual feature data and location data of the suspected search and rescue target are obtained, so as to obtain the individual confidence of the suspected search and rescue target based on the visual feature data; Step 1053: The location data is sent to several neighboring robot dogs corresponding to the quadruped robot dog to obtain the neighborhood confidence of each neighboring robot dog for the suspected search and rescue target; Step 1054: The individual confidence of the quadruped robot dog and all neighborhood confidences are weighted and summed according to the global interaction weight between the quadruped robot dog and each neighboring robot dog to obtain the search and rescue confirmation result of the suspected search and rescue target; Step 1055: Obtain the search and rescue mission for the search and rescue area, and control the quadruped robot dog cluster to conduct search and rescue in the search and rescue area based on the search and rescue mission and the search and rescue confirmation results of each quadruped robot dog.
[0055] In this embodiment, when each quadruped robot dog is controlled to move within the search and rescue area according to distributed gait parameters and path planning instructions, its search and rescue status data is obtained based on its individual perception data. Specifically, as the quadruped robot dog moves according to the distributed gait parameters and path planning instructions, its onboard sensors continuously collect environmental information and its own status information, i.e., individual perception data. This individual perception data, such as its own posture, local environmental obstacles, and target cues, is used to update the quadruped robot dog's search and rescue status data in real time, providing the latest basic information for subsequent target detection and collaborative decision-making.
[0056] When the quadruped robot dog detects a suspected search and rescue target based on search and rescue status data, it acquires the visual feature data and location data of the suspected target to obtain a single-target confidence score based on the visual feature data. Specifically, when the quadruped robot dog's search and rescue status data includes target detection information (e.g., human silhouettes, signs of life, etc., identified through image recognition, thermal imaging, etc.), the robot dog's computing unit extracts the visual feature data (such as color, texture, shape, thermal radiation pattern) and its location data within the search and rescue area of the suspected target. Subsequently, using a pre-trained classifier or deep learning model, the likelihood of the suspected target being a real search and rescue target is evaluated based on these visual feature data, thereby obtaining the quadruped robot dog's single-target confidence score for that target. It is important to note that the quadruped robot dog can be equipped with a visible light camera and an infrared thermal imager. When its visual processing module identifies an area in the image that is similar to the features of a preset search and rescue target (such as a human body or animal) using algorithms such as convolutional neural networks (CNN), it determines that a suspected search and rescue target has been detected.
[0057] The location data is sent to several neighboring robot dogs corresponding to the quadruped robot dog to obtain the neighborhood confidence score of each neighboring robot dog for the suspected search and rescue target. Specifically, when a quadruped robot dog detects a suspected search and rescue target and obtains its location data, it sends the location data to other quadruped robot dogs in its neighborhood through the cluster's internal communication network. After receiving the location data, these neighboring robot dogs will focus on and independently perceive the location area based on their current search and rescue status data and perception capabilities, and assess whether a search and rescue target exists at that location, thereby generating their respective neighborhood confidence scores for the suspected target. The communication protocol between the robot dogs can adopt point-to-point or broadcast methods to send the location data of the suspected search and rescue target to all neighboring robot dogs within the communication range. After receiving the information, the neighboring robot dog will adjust its sensors to face the location and run its own detection algorithm to calculate a confidence score value based on its perception results. For example, if the infrared sensor of a neighboring robot dog detects a heat source at the location, its neighborhood confidence score will increase accordingly; if it does not detect one, the confidence score will be lower.
[0058] The search and rescue confirmation result for a suspected target is obtained by weighting and summing the individual confidence score of the quadrupedal robot dog and the confidence scores of all neighboring robots according to their respective global interaction weights. Specifically, by weighting and summing the individual confidence score of the detected robot dog and the neighborhood confidence scores of all neighboring robots according to their respective global interaction weights, a comprehensive search and rescue confirmation result can be obtained. This result can effectively reduce the risk of misjudgment or missed detection by a single robot dog. The weighted summation can be performed linearly, i.e., Search and rescue confirmation result = Σ (Confidence score_i) / (Integer confidence score_i) Global interaction weight (_i), where i iterates through the detection robot dog itself and all neighboring robot dogs.
[0059] The system acquires search and rescue tasks for the search and rescue area, and controls the quadruped robot dog cluster to conduct search and rescue operations within that area based on these tasks and the search and rescue confirmation results of each quadruped robot dog. Specifically, search and rescue tasks can be preset or dynamically issued, such as "find all survivors" or "confirm whether there is a target in a specific area." Depending on the type of search and rescue task (e.g., single-target confirmation, multi-target confirmation, targetless search) and the search and rescue confirmation results of suspected targets (e.g., confirmation results exceeding a certain threshold), the cluster will adopt different control strategies, such as stopping the search, proceeding to the target point for detailed reconnaissance, adjusting the search area, or continuing a wide-area search, thereby optimizing search and rescue efficiency and resource allocation. Search and rescue tasks can be input into the master control agent through preset configuration files or operator instructions. The master control agent generates corresponding cluster control instructions based on the task type (e.g., "stop when one survivor is found," "find all survivors," "cover all areas") and the search and rescue confirmation results. For example, if the task is "stop once a survivor is found," and the search and rescue confirmation result for a certain target reaches a high confidence level, the master agent can issue an instruction to all robot dogs to stop their current search and converge on or report to the target location. Alternatively, the search and rescue task can be a dynamically adjusted strategy. For instance, when the search and rescue confirmation result indicates the presence of a high-value target in a certain area, the master agent can dynamically adjust the path planning instructions for some robot dogs in the cluster based on task priority, prioritizing their movement to that area for detailed reconnaissance, while other robot dogs continue to perform their original search tasks, thus achieving dynamic task scheduling and optimized resource allocation.
[0060] The above implementation method, based on the distributed gait parameters and path planning instructions of each quadruped robot dog, controls its movement within the search and rescue area and processes real-time individual perception data to obtain the search and rescue status data of the quadruped robot dog, thus achieving real-time monitoring of the robot dog's movement status and perception information. When a suspected search and rescue target is detected based on the search and rescue status data, visual feature data and location data are acquired and processed to obtain the individual confidence level, achieving a preliminary credibility assessment of the suspected target. The location data is sent to neighboring robot dogs, and each neighboring robot dog calculates and returns the results to the neighboring area. The confidence level is processed to obtain target evaluation information for multi-machine collaboration, achieving an effective combination of single-machine detection and multi-machine verification. Based on the global interaction weights between the quadruped robot dog and its neighboring robot dogs, the confidence level of the individual robot dog and the confidence levels of all neighboring robot dogs are weighted and summed to obtain the search and rescue confirmation result of the suspected target. This ensures that the confirmation result integrates multi-source information and takes into account the reliability differences of each information source. Based on the search and rescue confirmation result and search and rescue mission data, the quadruped robot dog cluster is controlled to carry out subsequent search and rescue processing, realizing a complete closed loop from target detection, collaborative confirmation to mission execution.
[0061] In this embodiment, step 1055 includes the following steps: Step S11: Obtain search and rescue mission data of the search and rescue area, and perform semantic parsing on the search and rescue mission data to determine the type of search and rescue mission; Step S12: When the search and rescue mission is a single-target confirmation mission and the confirmation result of any quadruped robot dog is higher than the preset confirmation threshold, a search and rescue termination command is generated, and the location data is synchronized to the other quadruped robot dogs in the quadruped robot dog cluster through the search and rescue topology data to control the quadruped robot dog cluster to complete the search and rescue mission; Step S13: When the search and rescue mission is a multi-target confirmation mission, for any quadruped robot dog, when the confirmation result of the quadruped robot dog is higher than the preset confirmation threshold, the location data of the quadruped robot dog's neighboring robot dog set and the suspected search and rescue target are obtained in real time, so as to obtain the updated distributed gait parameters and path planning instructions for each neighboring robot dog based on the location data; Step S14: The number of search and rescue targets in the quadruped robot dog cluster is obtained in real time. When the number of search and rescue targets is equal to or greater than the preset number of mission targets, a search and rescue termination command is generated, and the location data is synchronized to the other quadruped robot dogs in the quadruped robot dog cluster through the search and rescue topology data. The search and rescue process is synchronized with the other quadruped robot dogs in the quadruped robot dog cluster to control the quadruped robot dog cluster to complete the search and rescue mission; Step S15: When the search and rescue mission is a targetless search and rescue mission, for any quadruped robot dog, when the search and rescue confirmation result of the quadruped robot dog is higher than the preset confirmation threshold, the location data of the quadruped robot dog's neighboring robot dog set and the location data of the suspected search and rescue target are obtained in real time, so as to obtain the updated distributed gait parameters and path planning instructions of each neighboring robot dog based on the location data; Step S16: According to the path planning instructions of the quadruped robot dog and the updated path planning instructions of each neighboring robot dog, the search and rescue area coverage of the quadruped robot dog cluster is obtained; Step S17: When the search and rescue area coverage is greater than the preset coverage threshold, a search and rescue termination instruction is generated and the location data is synchronized with the other quadruped robot dogs in the quadruped robot dog cluster through the search and rescue topology relationship data to control the quadruped robot dog cluster to complete the search and rescue mission.
[0062] In this embodiment, search and rescue mission data for the search and rescue area is acquired, and semantic analysis is performed on the data to determine the type of search and rescue mission. Semantic analysis refers to extracting key information from unstructured or semi-structured mission descriptions using natural language processing (NLP) techniques or rule-based pattern matching. For example, the mission may involve finding a specific number of targets, confirming the existence of targets, or simply providing comprehensive coverage of the area. For instance, the mission type can be determined through keyword recognition (e.g., "finding one survivor" corresponds to single-target confirmation, "finding all survivors" corresponds to multi-target confirmation, and "comprehensive search area" corresponds to targetless search).
[0063] When the search and rescue mission is a single-target confirmation mission and the confirmation result of any quadruped robot dog exceeds a preset confirmation threshold, a search and rescue termination command is generated, and the location data is synchronized to the remaining quadruped robot dogs in the cluster via search and rescue topology data to control the quadruped robot dog cluster to complete the search and rescue mission. Specifically, when any quadruped robot dog in the cluster confirms the discovery of a suspected search and rescue target through its perception and collaborative confirmation mechanism (as described above), and the confirmation result of that target (e.g., confidence score) exceeds a preset confirmation threshold, it indicates that the mission objective has been achieved. At this time, the central control unit immediately generates a search and rescue termination command, along with the location data of the suspected search and rescue target. This command and location data are synchronized to all other quadruped robot dogs in the cluster via the search and rescue topology data of the quadruped robot dog cluster (e.g., broadcast or multi-hop transmission via a wireless communication network). The quadruped robot dog that receives the termination command will stop its current search and rescue behavior and can proceed to the target location for further assistance or return to standby based on the location data.
[0064] When the search and rescue mission is a multi-target confirmation mission, for any quadruped robot dog, when the confirmation result of that quadruped robot dog is higher than a preset confirmation threshold, the location data of its neighboring robot dogs and the location of the suspected search target are acquired in real time. Based on this location data, updated distributed gait parameters and path planning instructions are obtained for each neighboring robot dog. Specifically, when a quadruped robot dog discovers and confirms a suspected search target (the confirmation result is higher than the threshold), this does not mean the mission ends immediately. Instead, the cluster needs to reallocate resources to find other targets. To this end, the quadruped robot dog acquires its neighboring robot dog set (i.e., robot dogs with communication connections or geographical proximity) and the location data of the confirmed suspected search target in real time. Based on this location data, updated distributed gait parameters and path planning instructions are generated for each neighboring robot dog. These updated instructions aim to guide the neighboring robot dogs to avoid already searched areas, optimize search paths, or proceed to unsearched areas to continue searching for other targets, thereby avoiding repeated searches and improving overall search and rescue efficiency.
[0065] Based on this, the number of search and rescue targets for the quadruped robot dog swarm is acquired in real time. When the number of search and rescue targets equals or exceeds the preset number of mission targets, a search and rescue termination command is generated, and the location data is synchronized to the remaining quadruped robot dogs in the swarm via search and rescue topology data to control the swarm to complete the search and rescue mission. Specifically, during multi-target search and rescue, the central control unit continuously counts the total number of search and rescue targets confirmed by the swarm. When this real-time count of search and rescue targets reaches or exceeds the preset number of mission targets (for example, the mission requires finding at least 3 survivors, and the swarm has confirmed 3 or more), it indicates that the multi-target search and rescue mission is essentially complete. At this point, a search and rescue termination command is generated, and the location data of all confirmed targets is synchronized to all quadruped robot dogs in the swarm via search and rescue topology data, instructing them to stop search and rescue activities and complete the mission.
[0066] When the search and rescue mission is a targetless search and rescue mission, for any quadruped robot dog, when the search and rescue confirmation result of the quadruped robot dog is higher than a preset confirmation threshold, the system acquires the location data of the quadruped robot dog's neighboring robot dog set and the suspected search and rescue target in real time. Based on the location data, the system obtains updated distributed gait parameters and path planning instructions for each neighboring robot dog. Specifically, the main purpose of a targetless search and rescue mission is to fully cover a designated area, even if no target is found. However, if a quadruped robot dog unexpectedly discovers a suspected search and rescue target (the search and rescue confirmation result is higher than the threshold) in such a mission, the system still needs to respond. At this time, the quadruped robot dog acquires the location data of its neighboring robot dog set and the suspected search and rescue target. Based on this location data, the system generates updated distributed gait parameters and path planning instructions for the neighboring robot dogs. These instructions may guide the neighboring robot dogs to the target area for collaborative confirmation, or adjust their search path to ensure more detailed coverage of the area surrounding the target, while avoiding repeated searches of already confirmed target areas, thus optimizing coverage efficiency when a target is found.
[0067] Simultaneously, based on the path planning instructions of the quadruped robot dog and the updated path planning instructions of each neighboring robot dog, the search and rescue area coverage rate of the quadruped robot dog cluster is obtained. Specifically, the search and rescue area coverage rate is a key indicator for measuring the completion rate of a targetless search and rescue mission. The system comprehensively considers the current path planning instructions of the quadruped robot dog itself and the updated path planning instructions of its neighboring robot dogs to calculate the coverage of the entire quadruped robot dog cluster in real time. The coverage rate can be calculated based on the grid map method, that is, counting the proportion of the number of grids that the cluster has visited or sensed out of the total number of grids; or it can be based on the probabilistic coverage model, estimating the probability of each area being searched through sensor models and robot trajectories.
[0068] When the search area coverage exceeds a preset coverage threshold, a search termination command is generated, and the location data is synchronized to the remaining quadruped robot dogs in the cluster via search topology data to control the cluster to complete the search and rescue mission. Preferably, when the real-time calculated search area coverage reaches or exceeds a preset coverage threshold (e.g., 90% or 95%), it indicates that the search area has been sufficiently searched. At this point, the system generates a search termination command and synchronizes the location data of any discovered suspected targets (if any) to all quadruped robot dogs in the cluster via search topology data, instructing them to stop search activities and complete the mission. This ensures that the cluster can efficiently and accurately complete the area search task even without a clear target.
[0069] The above implementation method, based on the search and rescue task data in the search and rescue area, determines the type of search and rescue task through semantic parsing and processes it, realizing automatic identification and differentiation of single-target confirmation, multi-target confirmation, and targetless search and rescue tasks. When the search and rescue task is a single-target confirmation task and the search and rescue confirmation result is higher than the preset confirmation threshold, a search and rescue termination command is generated and synchronized to other robot dogs in the cluster through search and rescue topology data for processing, realizing rapid cluster response and task termination after single-target confirmation. When the search and rescue task is a multi-target confirmation task, it is processed by tracking the cumulative number of search and rescue targets, updating the neighborhood path planning command, and judging whether the preset task target number has been reached, realizing dynamic counting of confirmed targets, collaborative updating of neighborhood paths, and unified cluster termination after full target confirmation in multi-target scenarios, avoiding the problem of stopping the entire cluster upon discovering a single target, which affects search and rescue efficiency. When the search and rescue task is a targetless search and rescue task, it is processed by updating the neighborhood path planning command and obtaining the search and rescue area coverage rate, realizing cluster continuous search and adaptive termination driven by area coverage rate, enabling the cluster to complete area coverage search and rescue tasks efficiently and orderly even without a clear number of targets.
[0070] In this first embodiment, when a quadruped robot dog detects a search and rescue target and stops in the area where the target is located, in order to ensure the complete success of the search and rescue mission, it is necessary to control other quadruped robot dogs to search the areas not searched by the quadruped robot dog. The process of controlling other robot dogs to conduct the search and rescue includes: Step S21: Obtain the set of neighboring robot dogs of the quadruped robot dog according to the search and rescue topology data; Step S22: Associate the location data of suspected search and rescue targets whose search and rescue confirmation results are higher than a preset confirmation threshold with the real-time location coordinates of the quadruped robot dog to obtain the absolute location coordinates of the suspected search and rescue targets within the search and rescue area. Step S23: Obtain the dwell state of the quadruped robot dog, and determine the search area to be assigned to the quadruped robot dog based on the dwell state and the path planning instructions of the quadruped robot dog; Step S24: Update the neighborhood search and rescue feature vector and self-search and rescue features of each neighboring robot dog according to the search and rescue topology data; Step S25: Take the updated neighborhood search and rescue feature vector, updated self-search and rescue features, relative spatial relationship, boundary information of the search area to be assigned, and location data of the suspected search and rescue target of the neighboring robot dog as input, and regenerate the distributed gait parameters and path planning instructions of the neighboring robot dog through the reinforcement learning execution agent corresponding to the neighboring robot dog.
[0071] In this embodiment, the location data of the suspected search and rescue target is associated with the real-time location coordinates of the quadruped robot dog to obtain the absolute location coordinates of the suspected search and rescue target within the search and rescue area. If the location data of the suspected search and rescue target is relative to the local coordinate system of the quadruped robot dog, then the coordinate system can be transformed using the real-time location coordinates and attitude data of the quadruped robot dog to convert it into the absolute location coordinates in the global coordinate system of the search and rescue area.
[0072] Next, the stationary state of the quadruped robot dog is obtained. When determining the search area to be assigned to the quadruped robot dog based on its stationary state and path planning instructions, the stationary state can be determined based on the robot dog's movement speed. For example, if its movement speed is continuously below a certain preset threshold for a period of time, it is considered to be in a stationary state. Alternatively, the stationary state can also be obtained through the quadruped robot dog's internal state machine (such as "observing" or "analyzing"). Then, based on the quadruped robot dog's current path planning instructions and its stationary position, an unsearched area centered on that position or along its path direction is designated as the search area to be assigned. Alternatively, based on a raster map of the search area, adjacent raster areas not currently covered by the quadruped robot dog and not assigned by other robots can be designated as the search area to be assigned.
[0073] When updating the neighborhood search and rescue feature vector and its own search and rescue features of each neighboring robot dog based on search and rescue topology data, the update operation may include receiving information such as target location and confidence level from quadruped robot dogs that have discovered suspected search and rescue targets, and integrating this information into the neighboring robot dog's own search and rescue features. For example, this new information can be integrated into the neighborhood search and rescue feature vector through weighted averaging, feature concatenation, or by using graph neural networks to aggregate neighborhood information. Updating the own search and rescue features may involve adjusting the robot dog's internal priorities for search and rescue tasks and search strategies.
[0074] When the updated neighborhood search and rescue feature vector, updated self-search and rescue features, relative spatial relationships, boundary information of the search area to be assigned, and location data of the suspected search and rescue target are used as inputs, the reinforcement learning execution agent corresponding to the neighborhood robot dog regenerates the distributed gait parameters and path planning instructions of the neighborhood robot dog. The reinforcement learning execution agent receives this comprehensive information as a representation of its current state. The parameterized action network inside the agent (such as the parameterized action network of the aforementioned reinforcement learning execution agent) will output distributed gait parameters (e.g., adjusting walking speed, turning angle, and gait pattern to adapt to terrain or quickly approach the target area) and path planning instructions (e.g., new target point sequences or obstacle avoidance paths) adapted to the current new environment and task requirements, based on its trained policy. The relative spatial relationship can refer to the relative position and direction between the neighborhood robot dog and the suspected search and rescue target or the robot dog that has discovered the target.
[0075] In this first embodiment, taking the confirmation of the search and rescue target at the first coordinate (search and rescue confirmation result 0.80, higher than the preset confirmation threshold of 0.75) as an example, firstly, based on the search and rescue topology data, using the first robot dog as the reference node, the set of neighboring robot dogs with connecting edges to it is retrieved. The adjacency matrix of the dynamic search and rescue topology graph is queried, revealing that the first robot dog A and the second robot dog B (connecting edge AB, good communication quality) have a connecting edge; therefore, the set of neighboring robot dogs is {second robot dog B}.
[0076] The confirmed search target's position data relative to the first robot dog A is transformed with the first robot dog A's real-time position coordinates (120, 85) to obtain the target's absolute position coordinates within the search area. The first robot dog A determines that its search confirmation result is higher than the confirmation threshold, marks its status as "target under processing," and stops executing the original path planning instructions. Simultaneously, the first robot dog A divides its incomplete original search area into a pending search area and synchronizes the boundary information (coordinates of the four vertices) of the pending search area to neighboring robot dog B via the MESH network.
[0077] Specifically, taking the neighboring robot dog B as an example, robot dog A obtains the search area information corresponding to its original path planning instruction. In this embodiment, the original planned search area of robot dog A is a fan-shaped area (corresponding to a part of its initial search area) starting from the current position, along the navigation direction, with a preset radius and a preset search angle. Since robot dog A has stopped, the part of this fan-shaped area that has not yet been searched is divided into a search area to be assigned. Robot dog A extracts the boundary information of the search area to be assigned, including the coordinates of four vertices: (120, 85), (135, 105), (128, 112), (115, 95), and synchronizes the above boundary information and the absolute position coordinates of the search target to the neighboring robot dog B through search and rescue topology data.
[0078] After robot dog A stops, the node attributes of each robot dog in the cluster change: the node attribute feature vector of robot dog A changes, and this change is synchronized to the main control terminal and other robot dogs in the cluster through the real-time update mechanism of search and rescue topology data. Based on the updated search and rescue topology data, the neighborhood search and rescue feature vector and its own search and rescue feature of neighboring robot dog B are updated respectively: Updating its own search and rescue feature: robot dog B extracts the local search and rescue feature vector from its latest composite search and rescue status data, and adjusts it according to the weight parameters of the region segmentation boundary corresponding to robot dog B and the information sharing intensity coefficient in the latest multi-cooperative search and rescue constraint vector issued by the main control agent, to obtain the updated self-search and rescue feature of robot dog B (256 dimensions).
[0079] Update the neighborhood search and rescue feature vector: Based on the updated dynamic search and rescue topology graph, a pre-trained graph neural network is used to aggregate neighborhood information for robot dog B. Since the node state of robot dog A has changed, the graph neural network incorporates the new state features of robot dog A into the aggregation during message passing, resulting in the updated neighborhood search and rescue feature vector for robot dog B. In this vector, the contribution feature of robot dog A to robot dog B has changed from cooperative search partner to semantics indicating that the target has been confirmed and the area is waiting to be taken over.
[0080] For the neighboring robot dog B, the relative spatial relationship between robot dog B and the search and rescue target is calculated. The following data is input to the reinforcement learning agent of robot dog B: the updated neighboring search and rescue feature vector, the updated self-search and rescue features, the relative spatial relationship (relative distance and relative azimuth -15°), the boundary information of the search area to be assigned, and the absolute position coordinates of the suspected search and rescue target. The agent of robot dog B, through its parameterized action network, directly outputs the updated distributed gait parameters and path planning instructions. The newly generated path planning instructions integrate the search area to be assigned from robot dog A into robot dog B's search task, while automatically guiding robot dog B's search path to bypass the safe area (radius 10m) where the search and rescue target is located.
[0081] Through the above steps, after robot dog A confirms the target and stops, the change in cluster topology directly triggers path replanning for neighboring robot dog B. The search area not completed by robot dog A is seamlessly assigned to neighboring robot dogs, and at the same time, the search paths of each neighboring robot dog automatically avoid the location of the confirmed target. The entire process does not require checking spatial conflicts one by one, realizing efficient collaboration and task succession of the cluster in dynamic changes.
[0082] The above implementation method obtains a set of neighboring robot dogs based on search and rescue topology data. By associating the location data of suspected search targets whose search and rescue confirmation results are higher than a preset confirmation threshold with the real-time location coordinates of the current robot dog, the absolute location coordinates of the suspected search targets within the search and rescue area are obtained, providing a unified spatial reference benchmark for the path adjustment of the neighboring robot dogs. The search area to be assigned is determined based on the current dwell status of the robot dog and the path planning instructions. The neighboring search and rescue feature vectors and self-search and rescue features of each neighboring robot dog are updated through search and rescue topology data, realizing the dynamic division of search tasks for dwelling robot dogs and the real-time synchronization of collaborative information. Based on the updated neighboring search and rescue feature vectors, self-search and rescue features, boundary information of the search area to be assigned, and location data of suspected search and rescue targets, the distributed gait parameters and path planning instructions are regenerated through the reinforcement learning execution agent corresponding to each neighboring robot dog. This allows the neighboring robot dogs to integrate additional search needs generated by the dwelling of companions on the basis of the original search task, avoiding the search vacuum caused by the dwelling of a single robot dog, and realizing seamless handover and continuous and efficient advancement of cluster search tasks in dynamic changes.
[0083] On the other hand, refer to Figure 2 This embodiment also discloses a quadruped robot dog cluster collaborative search and rescue system based on a multi-agent system, which mainly includes a single-agent perception module 201, a composite perception module 202, a global constraint module 203, a path planning module 204, and a search and rescue control module 205.
[0084] The individual perception module 201 is used to acquire the individual perception data of each quadruped robot dog in the search and rescue area in real time, so as to determine the search and rescue topology data of the quadruped robot dog cluster in the search and rescue area based on the individual perception data.
[0085] The composite sensing module 202 is used to acquire multi-source sensing data of the search and rescue area, so as to determine the composite search and rescue status data of the quadruped robot dog cluster based on the multi-source sensing data, individual sensing data and search and rescue topology data.
[0086] The global constraint module 203 is used to obtain the multi-cooperative search and rescue constraint vector of the quadruped robot dog cluster through a preset reinforcement learning master control agent based on the composite search and rescue status data.
[0087] The path planning module 204 is used to obtain the distributed gait parameters and path planning instructions of each quadruped robot dog by means of the reinforcement learning execution agent corresponding to each quadruped robot dog, based on the multi-cooperative search and rescue constraint vector and composite search and rescue status data.
[0088] The search and rescue control module 205 is used to obtain the search and rescue status data of each quadruped robot dog in real time based on the distributed gait parameters and path planning instructions, so as to control the quadruped robot dog cluster to carry out search and rescue in the search and rescue area based on the search and rescue status data and the search and rescue task data of the search and rescue area.
[0089] In this embodiment, the single-unit perception module 201 includes a data acquisition unit, a node attribute unit, an edge unit, and a topology unit.
[0090] The data acquisition unit is used to acquire individual perception data of any quadruped robot dog after detecting its movement; wherein, the individual perception data includes at least its own posture data, local environmental obstacle distribution data, and target cue detection data; The node attribute unit is used to obtain the real-time position coordinates and movement speed of the quadruped robot dog based on its own posture data, so as to fuse the real-time position coordinates, movement speed and local environmental obstacle distribution data to obtain the node attribute feature vector of the quadruped robot dog. The edge unit is used to match the quadruped robot dog with the neighboring robot dogs based on the real-time location coordinates, and to obtain the communication status between the quadruped robot dog and the neighboring robot dogs, so as to establish the connection edge between the quadruped robot dog and the neighboring robot dogs according to the communication status; for any connection edge, the relative search and rescue status data between the two quadruped robot dogs corresponding to the connection edge is obtained according to the node attribute feature vector, and the relative search and rescue status data is used as the edge attribute feature vector of the connection edge; The topology unit is used to construct a dynamic search and rescue topology graph of the quadruped robot dog cluster based on the connecting edges, edge attribute feature vectors, and node attribute feature vectors, so as to serve as the search and rescue topology data of the quadruped robot dog cluster.
[0091] This embodiment discloses a method and system for collaborative search and rescue of quadruped robot dogs based on a multi-agent system. It utilizes real-time acquired individual perception data of each quadruped robot dog within the search and rescue area. This data is processed by constructing a dynamic search and rescue topology graph with the robot dogs as nodes to obtain search and rescue topology data characterizing the spatial relationships and communication status within the cluster, thus achieving refined modeling of the cluster's dynamic collaborative state. Furthermore, based on multi-source perception data, individual perception data, and search and rescue topology data of the search and rescue area, multi-level feature fusion is used to obtain composite search and rescue state data integrating global environment, local perception, and topology collaborative information, providing a complete data foundation for subsequent collaborative decision-making. Finally, based on the composite search and rescue state data, a preset reinforcement algorithm is used to... The master control agent processes the data to obtain a multi-cooperative search and rescue constraint vector containing global collaborative guidance information, providing a unified global collaborative constraint for the distributed decision-making of each robot dog. Based on the multi-cooperative search and rescue constraint vector and composite search and rescue status data, the reinforcement learning execution agent corresponding to each quadruped robot dog processes the data to obtain the distributed gait parameters and path planning instructions for each robot dog, realizing a complete decision-making link from global collaborative constraints to distributed action instructions. Based on the distributed gait parameters and path planning instructions, search and rescue status data is acquired in real time and processed in conjunction with search and rescue task data to control the cluster to conduct collaborative search and rescue in the search and rescue area. This enables the cluster to achieve efficient search and rescue with global coordination and local adaptive flexibility in unknown and complex environments.
[0092] The above specific embodiments further illustrate the purpose, technical solution, and beneficial effects of the present invention in detail. It should be understood that the above are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A method for collaborative search and rescue using quadruped robot dogs based on a multi-agent system, characterized in that, include: Real-time acquisition of individual perception data of each quadruped robot dog within the search and rescue area, so as to determine the search and rescue topology data of the quadruped robot dog cluster within the search and rescue area based on the individual perception data; Acquire multi-source sensing data of the search and rescue area, and determine the composite search and rescue status data of the quadruped robot dog cluster based on the multi-source sensing data, the individual sensing data, and the search and rescue topology data; Based on the composite search and rescue status data, the multi-cooperative search and rescue constraint vector of the quadruped robot dog cluster is obtained through a preset reinforcement learning master control agent. Based on the multi-cooperative search and rescue constraint vector and the composite search and rescue status data, the distributed gait parameters and path planning instructions of each quadruped robot dog are obtained through the reinforcement learning execution agent corresponding to each quadruped robot dog. The search and rescue status data of each quadruped robot dog is obtained in real time based on the distributed gait parameters and the path planning instructions, so as to control the quadruped robot dog cluster to carry out search and rescue in the search and rescue area according to the search and rescue status data and the search and rescue task data of the search and rescue area.
2. The method for collaborative search and rescue of quadruped robot dogs based on a multi-agent system according to claim 1, characterized in that, The real-time acquisition of individual perception data for each quadruped robot dog within the search and rescue area, and the determination of search and rescue topology data for the quadruped robot dog cluster within the search and rescue area based on the individual perception data, includes: For any of the quadruped robot dogs, when the movement of the quadruped robot dog is detected, the individual perception data of the quadruped robot dog is acquired; wherein, the individual perception data includes at least its own posture data, local environmental obstacle distribution data, and target cue detection data; The real-time position coordinates and movement speed of the quadruped robot dog are obtained from its own posture data. The real-time position coordinates, the movement speed and the local environmental obstacle distribution data are fused to obtain the node attribute feature vector of the quadruped robot dog. Based on the real-time location coordinates, the quadruped robot dog is matched with neighboring robot dogs, and the communication status between the quadruped robot dog and the neighboring robot dogs is obtained, so as to establish a connection edge between the quadruped robot dog and the neighboring robot dogs according to the communication status. For any connecting edge, the relative search and rescue status data between the two quadruped robot dogs corresponding to the connecting edge is obtained according to the node attribute feature vector, and the relative search and rescue status data is used as the edge attribute feature vector of the connecting edge. Using the quadruped robot dog as a node, a dynamic search and rescue topology graph of the quadruped robot dog cluster is constructed based on the connecting edges, the edge attribute feature vectors, and the node attribute feature vectors, as the search and rescue topology data of the quadruped robot dog cluster.
3. The method for collaborative search and rescue of quadruped robot dogs based on a multi-agent system according to claim 2, characterized in that, The step of acquiring multi-source sensing data of the search and rescue area, and determining composite search and rescue status data of the quadruped robot dog cluster based on the multi-source sensing data, the individual sensing data, and the search and rescue topology data, includes: The multi-source environmental perception data of the search and rescue area is rasterized and semantically segmented to obtain a global environmental situation map of the search and rescue area; wherein, the global environmental situation map includes known obstacle distribution areas, searched areas and unsearched areas; Feature extraction is performed on the individual perception data of each quadruped robot dog to generate a local search and rescue feature vector for each quadruped robot dog. Based on the local search and rescue feature vector and the dynamic search and rescue topology graph, the neighborhood search and rescue information of each quadruped robot dog is aggregated through a preset graph neural network to obtain the neighborhood search and rescue feature vector of each quadruped robot dog. The local search and rescue feature vector and the neighborhood search and rescue feature vector of each quadruped robot dog are aligned and stitched together with the global environmental situation map to obtain the composite search and rescue status data of each quadruped robot dog in the search and rescue area.
4. The method for collaborative search and rescue of quadruped robot dogs based on a multi-agent system according to claim 3, characterized in that, The step of obtaining the multi-cooperative search and rescue constraint vector of the quadruped robot dog cluster through a preset reinforcement learning master control agent based on the composite search and rescue status data includes: For any quadruped robot dog, the global interaction weight between each neighboring robot dog and the quadruped robot dog is obtained according to the neighborhood search and rescue feature vector and the multi-head self-attention mechanism preset in the reinforcement learning master control agent. Based on the local search and rescue feature vector, the dynamic search and rescue topology graph and the global environmental situation map, multi-task processing is performed through a multi-task decoupling network preset in the reinforcement learning master control agent to obtain the region segmentation boundary weight parameters, formation preservation constraint parameters and information sharing strength coefficient between each neighboring robot dog and the quadruped robot dog. Based on the global interaction weight, the region segmentation boundary weight parameter, the formation preservation constraint parameter, and the information sharing strength coefficient corresponding to each quadruped robot dog in the quadruped robot dog cluster, the reinforcement learning master control agent outputs the multi-cooperative search and rescue constraint vector of the quadruped robot dog cluster.
5. The method for collaborative search and rescue of quadruped robot dogs based on a multi-agent system according to claim 4, characterized in that, The step of obtaining distributed gait parameters and path planning instructions for each quadruped robot dog through a reinforcement learning agent corresponding to each quadruped robot dog, based on the multi-cooperative search and rescue constraint vector and the composite search and rescue state data, includes: For any quadruped robot dog, the local search and rescue feature vector of the quadruped robot dog is weighted and adjusted according to the region segmentation boundary weight parameter and the information sharing intensity coefficient to obtain the self-search and rescue feature of the quadruped robot dog. Using its own search and rescue features as the query vector and the neighborhood search and rescue feature vector as the key vector and value vector, the collaborative attention weights between the quadruped robot dog and each neighboring robot dog are calculated through a collaborative attention mechanism preset in the reinforcement learning execution agent. The neighborhood search and rescue feature vector of the quadruped robot dog is weighted and aggregated according to the collaborative attention weight to obtain the neighborhood search and rescue feature of the quadruped robot dog. The quadruped robot dog performs gating fusion of its own search and rescue features and the neighborhood search and rescue features according to the reinforcement learning execution agent corresponding to the quadruped robot dog to obtain the comprehensive search and rescue state representation vector of the quadruped robot dog. The comprehensive search and rescue state representation vector is input into the parameterized action network of the reinforcement learning executive agent to output the distributed gait parameters and path planning instructions of the quadruped robot dog; wherein, during the training process of the reinforcement learning executive agent, the network parameters of the parameterized action network are determined according to a composite reward function including search and rescue coverage, response time, energy consumption and task balance indicators.
6. The method for collaborative search and rescue of quadruped robot dogs based on a multi-agent system according to claim 1, characterized in that, The step of acquiring search and rescue status data of each quadruped robot dog in real time based on the distributed gait parameters and the path planning instructions, and controlling the quadruped robot dog cluster to conduct search and rescue in the search and rescue area based on the search and rescue status data and the search and rescue task data of the search and rescue area, includes: For each of the quadruped robot dogs, when the quadruped robot dog is controlled to move within the search and rescue area according to the distributed gait parameters and the path planning instructions, the search and rescue status data of the quadruped robot dog is obtained according to the individual perception data of the quadruped robot dog. When the search and rescue status data determines that the quadruped robot dog has detected a suspected search and rescue target, the visual feature data and location data of the suspected search and rescue target are obtained, so as to obtain the individual confidence level of the suspected search and rescue target based on the visual feature data. The location data is sent to several neighboring robot dogs corresponding to the quadruped robot dog to obtain the neighborhood confidence level of each neighboring robot dog for the suspected search and rescue target. The search and rescue confirmation result of the suspected target is obtained by weighting and summing the individual confidence score of the quadruped robot dog and all the neighborhood confidence scores based on the global interaction weight between the quadruped robot dog and each of the neighboring robot dogs. The search and rescue mission for the search and rescue area is obtained, and the quadruped robot dog cluster is controlled to conduct a search and rescue operation in the search and rescue area based on the search and rescue mission and the search and rescue confirmation result of each quadruped robot dog.
7. The method for collaborative search and rescue of quadruped robot dogs based on a multi-agent system according to claim 6, characterized in that, The step of acquiring the search and rescue task for the search and rescue area, and controlling the quadruped robot dog cluster to conduct a search and rescue operation in the search and rescue area based on the search and rescue task and the search and rescue confirmation result of each quadruped robot dog, includes: Acquire search and rescue mission data for the search and rescue area, and perform semantic parsing on the search and rescue mission data to determine the type of search and rescue mission; When the search and rescue mission is a single-target confirmation mission and the search and rescue confirmation result of any of the quadruped robot dogs is higher than the preset confirmation threshold, a search and rescue termination command is generated and the location data is synchronized to the other quadruped robot dogs in the quadruped robot dog cluster through the search and rescue topology data, so as to control the quadruped robot dog cluster to complete the search and rescue mission. When the search and rescue mission is a multi-target confirmation mission, for any quadruped robot dog, when the search and rescue confirmation result of the quadruped robot dog is higher than the preset confirmation threshold, the location data of the neighboring robot dog set and the suspected search and rescue target of the quadruped robot dog are obtained in real time, so as to obtain the updated distributed gait parameters and path planning instructions of each neighboring robot dog based on the location data. The number of search and rescue targets of the quadruped robot dog cluster is obtained in real time. When the number of search and rescue targets is equal to or greater than the preset number of task targets, a search and rescue termination command is generated and the location data is synchronized to the other quadruped robot dogs in the quadruped robot dog cluster through search and rescue topology data, so as to control the quadruped robot dog cluster to complete the search and rescue task. When the type of the search and rescue mission is a targetless search and rescue mission, for any quadruped robot dog, when the search and rescue confirmation result of the quadruped robot dog is higher than the preset confirmation threshold, the location data of the neighboring robot dog set and the suspected search and rescue target of the quadruped robot dog are obtained in real time, so as to obtain the updated distributed gait parameters and path planning instructions of each neighboring robot dog based on the location data. The search and rescue area coverage of the quadruped robot dog cluster is obtained based on the path planning instructions of the quadruped robot dog and the updated path planning instructions of each neighboring robot dog. When the coverage of the search and rescue area is greater than a preset coverage threshold, a search and rescue termination command is generated and the location data is synchronized to the other quadruped robot dogs in the quadruped robot dog cluster through search and rescue topology data, so as to control the quadruped robot dog cluster to complete the search and rescue task.
8. The method for collaborative search and rescue of quadruped robot dogs based on a multi-agent system according to claim 7, characterized in that, For any quadruped robot dog, when the search and rescue confirmation result of the quadruped robot dog is higher than a preset confirmation threshold, the location data of the neighboring robot dog set and the suspected search and rescue target are acquired in real time. Based on the location data, updated distributed gait parameters and path planning instructions for each neighboring robot dog are obtained, including: Based on the search and rescue topology data, obtain the set of neighboring robot dogs of the quadruped robot dog; The location data of suspected search targets whose search and rescue confirmation results are higher than the preset confirmation threshold are associated with the real-time location coordinates of the quadruped robot dog to obtain the absolute location coordinates of the suspected search and rescue targets within the search and rescue area. The stationary state of the quadruped robot dog is obtained, and the search area to be assigned to the quadruped robot dog is determined according to the stationary state and the path planning instructions of the quadruped robot dog. Update the neighborhood search and rescue feature vector and its own search and rescue features for each of the neighboring robot dogs based on the search and rescue topology data; The updated neighborhood search and rescue feature vector of the neighborhood robot dog, the updated self-search and rescue features, the relative spatial relationship, the boundary information of the search area to be assigned, and the location data of the suspected search and rescue target are used as inputs. The distributed gait parameters and path planning instructions of the neighborhood robot dog are regenerated by the reinforcement learning execution agent corresponding to the neighborhood robot dog.
9. A quadruped robot dog swarm collaborative search and rescue system based on a multi-agent system, characterized in that, It includes a single-unit perception module, a composite perception module, a global constraint module, a path planning module, and a search and rescue control module; The individual perception module is used to acquire the individual perception data of each quadruped robot dog in the search and rescue area in real time, so as to determine the search and rescue topology data of the quadruped robot dog cluster in the search and rescue area based on the individual perception data. The composite sensing module is used to acquire multi-source sensing data of the search and rescue area, so as to determine the composite search and rescue status data of the quadruped robot dog cluster based on the multi-source sensing data, the individual sensing data and the search and rescue topology data. The global constraint module is used to obtain the multi-cooperative search and rescue constraint vector of the quadruped robot dog cluster through a preset reinforcement learning master control agent based on the composite search and rescue status data. The path planning module is used to obtain the distributed gait parameters and path planning instructions of each quadruped robot dog through the reinforcement learning execution agent corresponding to each quadruped robot dog, based on the multi-cooperative search and rescue constraint vector and the composite search and rescue status data. The search and rescue control module is used to acquire the search and rescue status data of each quadruped robot dog in real time according to the distributed gait parameters and the path planning instructions, so as to control the quadruped robot dog cluster to carry out search and rescue in the search and rescue area according to the search and rescue status data and the search and rescue task data of the search and rescue area.
10. The quadruped robot dog swarm collaborative search and rescue system based on a multi-agent system according to claim 9, characterized in that, The single-unit perception module includes a data acquisition unit, a node attribute unit, an edge unit, and a topology unit; The data acquisition unit is used to acquire individual perception data of any quadruped robot dog after detecting its movement; wherein, the individual perception data includes at least its own posture data, local environmental obstacle distribution data, and target cue detection data. The node attribute unit is used to obtain the real-time position coordinates and movement speed of the quadruped robot dog based on its own posture data, and to fuse the real-time position coordinates, the movement speed and the local environmental obstacle distribution data to obtain the node attribute feature vector of the quadruped robot dog. The edge unit is used to match the quadruped robot dog's neighboring robot dogs based on the real-time location coordinates, and to obtain the communication status between the quadruped robot dog and the neighboring robot dogs, so as to establish a connection edge between the quadruped robot dog and the neighboring robot dogs according to the communication status; for any connection edge, the relative search and rescue status data between the two quadruped robot dogs corresponding to the connection edge is obtained according to the node attribute feature vector, and the relative search and rescue status data is used as the edge attribute feature vector of the connection edge; The topology unit is used to construct a dynamic search and rescue topology graph of the quadruped robot dog cluster based on the connecting edges, the edge attribute feature vectors, and the node attribute feature vectors, using the quadruped robot dog as a node, so as to serve as the search and rescue topology data of the quadruped robot dog cluster.