A multi-animal robot collaborative human-computer interactive navigation system
Through a centralized control architecture and multi-agent reinforcement learning algorithm, the problems of perception range and coordination difficulties of single-species animal robots in complex environments were solved, and efficient collaborative navigation and improved mission success rate of multi-animal robot systems were achieved.
Patent Information
- Application Number
- CN202510980411.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-16
AI Technical Summary
In complex and ever-changing real-world environments, animal robots of a single species have limited local perception capabilities and find it difficult to fully understand the surrounding environment, which leads to decision-making bias and affects the task execution effect of multi-agent systems. In addition, traditional control methods have problems such as high communication overhead and difficult coordination in large-scale intelligent agent systems.
A centralized control architecture is adopted to collect information from each animal robot through a central control system, a collaborative perception module is used to generate a semantic map from a bird's-eye view, and control instructions are generated through a collaborative control module. Combined with multi-agent reinforcement learning algorithms and knowledge distillation technology, a globally optimal navigation solution is achieved.
It improves the environmental perception range and navigation task success rate of the multi-animal robot system, optimizes the collaborative operation capability between different types of animal robots, and enhances the stability and overall collaborative capability of the system.
Smart Images

Figure CN120510386B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of navigation technology, and in particular to a multi-animal robot collaborative human-computer interaction navigation system. Background Art
[0002] In modern rescue and search missions, animal robots, with their flexible movements, long endurance, and excellent concealment, have become a focus of research. For example, in confined and complex environments like earthquake ruins and underground pipelines, rat robots, with their compact size, excellent climbing abilities, and adaptability to dim lighting, can accomplish tasks that are difficult for humans and mechanical robots. However, single-species animal robots also face significant limitations and challenges when performing these tasks.
[0003] In complex and ever-changing real-world environments, individual agents, limited by their sensors and field of view, often only acquire partial environmental information. This localized perception makes it difficult for agents to fully understand their surroundings, potentially leading to biased decision-making and compromising the performance of the entire multi-agent system. For example, in a search and rescue mission, a single agent might be unable to detect a target hidden behind an obstacle, resulting in inefficient rescue efforts. Therefore, expanding the agent's environmental perception range and improving its accuracy are key challenges facing multi-agent systems.
[0004] Coordinated control of multi-agent systems is key to achieving efficient collaboration. Different agents may have varying capabilities and goals. Coordinating the actions of these agents in a complex environment so they can collaborate and achieve the overall task is a challenging problem. While traditional distributed control methods offer some flexibility, they are prone to high communication overhead and coordination difficulties when handling complex tasks and large-scale agent systems. Centralized control methods, while enabling unified scheduling and decision-making, remain a pressing challenge in effectively aggregating and processing information from each agent and generating appropriate control instructions.
[0005] Therefore, a multi-animal robot collaborative human-computer interaction navigation system has become an urgent problem to be solved. Summary of the Invention
[0006] The purpose of this invention is to propose a multi-animal robot collaborative human-computer interaction navigation system to enhance the perception and navigation capabilities of the multi-animal robot navigation system in complex environments, and to solve the problems of limited perception range and motion capabilities as well as low control and navigation efficiency in the existing technology.
[0007] To achieve the above objectives, the present invention provides a technical solution: a multi-animal robot collaborative human-machine interactive navigation system, comprising a central control system and a navigation system, wherein the central control system collects and processes information provided by each animal robot in the navigation system and implements unified control decisions;
[0008] The central control system includes a collaborative perception module and a collaborative control module. The collaborative perception module generates a semantic map from a bird's-eye view. The collaborative control module generates control instructions for each animal robot and integrates a multi-agent reinforcement learning algorithm with personalized training and distillation execution.
[0009] The navigation system includes several animal robots and a central control system; the animal robots carry micro cameras for shooting videos of the navigation environment and transmitting them to the computer of the central control system in real time. The central control body views these videos in real time on the computer and executes the navigation control method, sending control instructions to each animal robot through the wireless communication module and the remote control interface to control it to complete the predetermined navigation task.
[0010] Furthermore, the specific content of the control decision is as follows:
[0011] All animal robots are in a shared environment, and the state space is , each agent has an observation space and action space , all animal robots work together to complete a given task ;
[0012] Fuse the local observation information obtained by each robot into a global observation;
[0013] Use multi-agent reinforcement learning algorithms to generate globally optimal navigation solutions;
[0014] The decision results are handed over to the operator for control and adjustment.
[0015] Furthermore, the collaborative perception module includes an image super-resolution unit, a BEV feature extraction unit and a semantic segmentation unit. The image super-resolution unit uses a lightweight SwinTransformer model trained on low-resolution images to perform image super-resolution. The BEV feature extraction unit uses an attention mechanism in time and space to learn unified bird's-eye view features (BEV features). The semantic segmentation unit uses MaskDecoder to generate a semantic map from a bird's-eye view based on the extracted BEV features.
[0016] Furthermore, the specific working steps of the collaborative perception module are as follows:
[0017] S21. Use low-resolution images to train a lightweight Transformer image super-resolution model, create low-resolution-high-resolution pseudo training pairs through multi-scale enhancement, and use a lightweight SwinTransformer model based on the window multi-head self-attention mechanism to super-resolve the low-resolution images, using the mean square error loss. To train:
[0018] (x);
[0019] ;
[0020] ;
[0021] in, represents the original input image, represents the number of channels of the original input image, represents the width of the original input image, represents the height of the original input image, represents the nearest neighbor downsampling, is the result of downsampling, where , is the downsampling ratio, represents bicubic upsampling, is the result of upsampling, where , is the upsampling ratio, is the image obtained after super resolution;
[0022] S22, the observation image of the animal robot i at time t The feature map is obtained by sequentially passing the image super-resolution model and the pre-trained ResNet50 model , create a BEV query First, use the BEV features and BEV query of the previous moment to do temporal self-attention, and then do spatial cross-attention with the feature map of the current time to generate the BEV features of the current moment ;
[0023] S23, input the obtained BEV features into the pre-trained MaskDecoder, output the semantic map, and use the pixel-level cross entropy loss function Perform model training:
[0024] ;
[0025] ;
[0026] in, represents the predicted semantic map, y represents the true label of the pixel, represents the probability that the model predicts the pixel to be the true label, Represents a set of pixel categories.
[0027] Furthermore, the multi-agent reinforcement learning algorithm uses massive prior data to pre-train the model and optimizes the control instruction generation strategy through global environmental information;
[0028] During the training phase, the global information is customized into personalized global information for each agent through the global information personalization module;
[0029] During the execution phase, knowledge distillation is used to simulate the effects of personalized global information using only the agent’s local information.
[0030] Furthermore, the specific working steps of the collaborative control module are as follows:
[0031] S31. Use recurrent neural network (RNN) to process the local information trajectory of each agent, the local observation of the agent and the action at the previous moment Encoded as a hidden state :
[0032] ;
[0033] S32, use the global information personalization module to extract and generate information tailored to the specific needs of each agent from the global state, and use the semantic map predicted by the collaborative perception module As global information :
[0034] ;
[0035] ;
[0036] in, Represents a personalized parameter network based on the local information of each animal robot Output a set of parameters and , Represents a personalized network, using the parameters output by the personalized parameter network to process Generate animal robots Customized global information ;
[0037] S33. Use knowledge distillation to achieve decentralized execution, using the agent's local information to distill personalized global information:
[0038] ;
[0039] ;
[0040] ;
[0041] in, It is a teacher network. Use global information and local information The generated personalized global information, It's a student network. Use local information Simulate the personalized global information generated by the teacher network and minimize the mean square error loss during training. accomplish, Indicates the an intelligent agent, is the number of agents;
[0042] S34. Calculate the action value function of each agent, expressed as:
[0043] ;
[0044] Select the optimal action based on the calculated value of the action value function, represents the action-value function, Represents the observation and action of the i-th agent at time t, and uses the QMIX algorithm for strategy training;
[0045] S35, the training process is divided into two stages. The first stage provides each agent with personalized global information to calculate the action value function and individual strategy. The second stage uses offline knowledge distillation to distill personalized global information using local information. The execution stage directly uses local information to replace personalized global information. For each set of data, a semantic map representing global information is generated through the collaborative perception module. express.
[0046] Furthermore, the process of using the QMIX algorithm for strategy training is as follows:
[0047] (1) Initialize hybrid network parameters , the number of agents ;
[0048] (2) Initialize the parameters of the single agent Q network and experience buffer ;
[0049] (4) Initialize the discount factor , Experience buffer capacity , learning rate and batch size ;
[0050] (5) At each time step t, all agents perform actions according to their respective Q networks and obtain observations and rewards at the next moment, converting the tuple Store in experience buffer ;
[0051] (6) When the number of tuples in the experience buffer is greater than , randomly sample a batch of data from the experience buffer;
[0052] (7) For each observation in the data set , a semantic map representing global information is generated through the collaborative perception module, and express;
[0053] (8) For each agent, use a recurrent network to generate its hidden state , use the global information personalization module to generate customized global information ;
[0054] (9) Calculate the target Q value of each agent and the global target Q value , the Q-value loss function of each agent is ;
[0055] (10) Calculate the global Q value using the hybrid network , the loss function of the hybrid network is:
[0056] ;
[0057] (11) Update the Q network parameters of each agent using the current batch of experience data And the parameters of the hybrid network ;
[0058] (12) Repeat steps (5)-(11) until the strategy converges or the preset maximum number of training steps is reached.
[0059] The advantages of the present invention compared with the prior art are:
[0060] The present invention establishes a hybrid decision-making mechanism of machine intelligence optimization and human cognition guidance, integrating machine intelligence and human judgment in the navigation system. Machine intelligence generates machine decisions based on observations, and humans correct machine decisions based on professional knowledge and experience to generate effective control instructions.
[0061] This invention utilizes a centralized control architecture. A central control system aggregates and processes information from each robot, providing the results to operators for decision-making and generating a globally optimal navigation solution. This system enables functional complementarity between different types of animal robots, optimizes their collaborative operation, and increases the success rate of navigation tasks.
[0062] The collaborative perception module improves image quality through the SwinTransformer model and generates a global semantic map through BEV feature extraction and semantic segmentation, effectively solving the problem of low-resolution images and expanding the perception range, thereby ensuring the accuracy of the global navigation solution and a high success rate of navigation tasks;
[0063] The collaborative control module uses a multi-agent reinforcement learning algorithm with personalized training and distillation execution, and pre-trains the model based on a large amount of prior data. It can quickly adapt to the needs of different types of agents and generate customized control instructions, thereby improving the reliability of the instructions and the overall collaborative capabilities of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 The present invention provides a workflow diagram of a multi-animal robot collaborative human-machine interactive navigation system.
[0065] Figure 2 This is a schematic diagram of the image super-resolution algorithm.
[0066] Figure 3 This is a schematic diagram of the global information personalization module.
[0067] Figure 4 This is a schematic diagram of the knowledge distillation module. DETAILED DESCRIPTION
[0068] Various exemplary embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present invention.
[0069] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the invention, its application, or uses.
[0070] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.
[0071] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.
[0072] The following is a further detailed description of the multi-animal robot collaborative human-computer interaction navigation system of the present invention in conjunction with the accompanying drawings.
[0073] Combined with attachment Figure 1-4 The specific implementation process of the multi-animal robot collaborative human-machine interactive navigation system of the present invention is as follows:
[0074] As attached Figure 1 As shown in the figure, a multi-animal robot collaborative human-machine interactive navigation system adopts a centralized control architecture. The central control system collects and processes information from each robot. The central control system includes a collaborative perception module and a collaborative control module:
[0075] The centralized control system aggregates and processes information provided by each animal robot through a central control system to implement unified control decisions. This structure first integrates the local observation information obtained by each animal robot into a global observation. It then uses a multi-agent reinforcement learning algorithm to generate a globally optimal navigation solution, and ultimately passes the decision results to the operator for control and adjustment. This centralized architecture effectively promotes efficient collaboration between the animal robots, optimizes the task execution process, and significantly improves the stability and overall performance of the system.
[0076] The collaborative perception module aims to generate a bird's-eye view semantic map. It consists of three core components: an image super-resolution unit, a BEV feature extraction unit, and a semantic segmentation unit. The image super-resolution unit uses a lightweight SwinTransformer model trained only on low-resolution images for image super-resolution. The BEV feature extraction unit uses an attention mechanism across time and space to learn unified BEV features. The semantic segmentation unit uses a MaskDecoder to generate a bird's-eye view semantic map based on the extracted BEV features. This approach effectively expands the system's environmental perception range and significantly improves perception accuracy.
[0077] The collaborative control module is responsible for generating control instructions for each animal robot and adopts a multi-agent reinforcement learning algorithm with personalized training and distillation execution. The algorithm uses massive prior data to pre-train the model and optimizes the control instruction generation strategy through global environmental information, thereby improving the reliability of the instructions. During the training phase, the global information is customized into personalized global information for each agent through the global information personalization module to enhance the performance of each agent and learn the optimal control strategy. Subsequently, in the execution phase, the knowledge distillation method is used to simulate the effects of these personalized global information, ensuring the effect during execution.
[0078] The navigation system consists of multiple animal robots and a central control system. The animal robots include an electrostimulation backpack, a miniature camera, and the animal itself. The electrostimulation backpack includes a wireless communication module and an electrostimulation module. The central control system includes a central control body, a wireless communication module, a computer, and a remote control interface. The miniature cameras carried by the animal robots capture the navigation environment, and these videos are transmitted in real time to the computer in the central control system, where the central control body can view these videos in real time. The central control body executes the navigation control method and sends control instructions to each animal robot via the wireless communication module and the remote control interface, controlling it to complete the predetermined navigation task.
[0079] As attached Figure 2 As shown, the image super-resolution part trains a lightweight Transformer image super-resolution model, creates low-resolution-high-resolution pseudo training pairs through multi-scale enhancement, and uses a lightweight SwinTransformer model based on the window multi-head self-attention mechanism to super-resolve low-resolution images, using the mean square error loss To train:
[0080] (x);
[0081] ;
[0082] ;
[0083] in, represents the original input image, represents the number of channels of the original input image, represents the width of the original input image, represents the height of the original input image, represents the nearest neighbor downsampling, is the result of downsampling, where , is the downsampling ratio, represents bicubic upsampling, is the result of upsampling, where , is the upsampling ratio, is the image obtained after super resolution;
[0084] The BEV feature extraction and semantic map process is as follows: the observation image of animal robot i at time t The feature map is obtained by sequentially passing the image super-resolution model and the pre-trained ResNet50 model , create a BEV query First, use the BEV features and BEV query of the previous moment to do temporal self-attention, and then do spatial cross-attention with the feature map of the current time to generate the BEV features of the current moment ; Input the obtained BEV features into the pre-trained MaskDecoder, output the semantic map, and use the pixel-level cross entropy loss function Perform model training:
[0085] ;
[0086] ;
[0087] in, represents the predicted semantic map, y represents the true label of the pixel, represents the probability that the model predicts the pixel to be the true label, Represents a set of pixel categories.
[0088] As attached Figure 3 As shown in the figure, the global information personalization module extracts and generates information tailored to the specific needs of each agent from the global state, and integrates the semantic map predicted by the collaborative perception module into the As global information :
[0089] ;
[0090] ;
[0091] in, Represents a personalized parameter network based on the local information of each animal robot Output a set of parameters and , Represents a personalized network, using the parameters output by the personalized parameter network to process Generate animal robots Customized global information ;
[0092] As attached Figure 4 As shown, the knowledge distillation method uses the local information of the agent to distill personalized global information:
[0093] ;
[0094] ;
[0095] ;
[0096] in, It is a teacher network. Use global information and local information The generated personalized global information, It's a student network. Use local information Simulate the personalized global information generated by the teacher network and minimize the mean square error loss during training. accomplish, Indicates the an intelligent agent, is the number of agents.
[0097] The process of using the QMIX algorithm for strategy training is as follows:
[0098] (1) Initialize hybrid network parameters .
[0099] (2) Initialize the parameters of the single agent Q network and experience buffer .
[0100] (4) Initialize the discount factor , Experience buffer capacity , learning rate and batch size .
[0101] (5) At each time step t, all agents perform actions according to their respective Q networks and obtain observations and rewards at the next moment, converting the tuple Store in experience buffer .
[0102] (6) When the number of tuples in the experience buffer is greater than , randomly sample a batch of data from the experience buffer.
[0103] (7) For each observation in the data set , a semantic map representing global information is generated through the collaborative perception module, and express.
[0104] (8) For each agent, use a recurrent network to generate its hidden state , use the global information personalization module to generate customized global information .
[0105] (9) Calculate the target Q value of each agent and the global target Q value , the Q-value loss function of each agent is .
[0106] (10) Calculate the global Q value using the hybrid network , the loss function of the hybrid network is:
[0107] .
[0108] (11) Update the Q network parameters of each agent using the current batch of experience data And the parameters of the hybrid network .
[0109] (12) Repeat steps (5) to (11) until the strategy converges or the preset maximum number of training steps is reached.
[0110] The present invention and its embodiments are described above. This description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this and, without departing from the purpose of the present invention, designs structures and embodiments similar to this technical solution without inventiveness, they shall fall within the scope of protection of the present invention.
Claims
1. A multi-animal robot collaborative human-machine interactive navigation system, characterized by: It includes a central control system and a navigation system. The central control system collects and processes information provided by each animal robot in the navigation system and implements unified control decisions. The central control system includes a collaborative perception module and a collaborative control module. The collaborative perception module generates a semantic map from a bird's-eye view. The collaborative control module generates control instructions for each animal robot and integrates a multi-agent reinforcement learning algorithm with personalized training and distillation execution. The navigation system includes several animal robots and a central control system; the animal robots carry miniature cameras for capturing videos of the navigation environment and transmitting them in real time to a computer in the central control system. The central control entity views the videos in real time on the computer and executes a navigation control method, sending control instructions to each animal robot via a wireless communication module and a remote control interface to control it to complete a predetermined navigation task; The specific content of the control decision is as follows: All animal robots are in a shared environment. Each agent has an observation space and action space. All animal robots collaborate to complete a given task. Fuse the local observation information obtained by each robot into a global observation; Use multi-agent reinforcement learning algorithms to generate globally optimal navigation solutions; The decision results are handed over to the operator for control and adjustment; The collaborative perception module includes an image super-resolution unit, a BEV feature extraction unit, and a semantic segmentation unit. The image super-resolution unit uses a lightweight SwinTransformer model trained on low-resolution images for image super-resolution. The BEV feature extraction unit uses an attention mechanism in time and space to learn unified BEV features. The semantic segmentation unit uses MaskDecoder to generate a semantic map from a bird's-eye view based on the extracted BEV features. The specific working steps of the collaborative perception module are as follows: S21. Use low-resolution images to train a lightweight Transformer image super-resolution model, create low-resolution-high-resolution pseudo training pairs through multi-scale enhancement, and use a lightweight SwinTransformer model based on the window multi-head self-attention mechanism to super-resolve the low-resolution images, using the mean square error loss. To train: (x); ; ; in, represents the original input image, represents the nearest neighbor downsampling, is the result of downsampling, where , represents bicubic upsampling, is the result of upsampling, where , is the image obtained after super resolution; S22, the observation image of the animal robot i at time t The feature map is obtained by sequentially passing the image super-resolution model and the pre-trained ResNet50 , create a BEV query First, use the BEV features and BEV query of the previous moment to do temporal self-attention, and then do spatial cross-attention with the feature map of the current time to generate the BEV features of the current moment ; S23, input the obtained BEV features into the pre-trained MaskDecoder, output the semantic map, and use the pixel-level cross entropy loss function Perform model training: ; ; in, represents the predicted semantic map, y represents the true label of the pixel, Represents the probability that the model predicts the pixel to be the true label.
2. The multi-animal robot collaborative human-machine interactive navigation system according to claim 1, characterized in that: The multi-agent reinforcement learning algorithm uses massive prior data to pre-train the model and optimizes the control instruction generation strategy through global environmental information; During the training phase, the global information is customized into personalized global information for each agent through the global information personalization module; In the execution phase, knowledge distillation is used to simulate the effect of personalized global information using only the local information of the agent.
3. The multi-animal robot collaborative human-machine interactive navigation system according to claim 2, characterized in that: The specific working steps of the collaborative control module are as follows: S31, use RNN to process the local information trajectory of each agent, the local observation of the agent and the action at the previous moment Encoded as a hidden state : ; S32, use the global information personalization module to extract and generate information tailored to the specific needs of each agent from the global state, and use the semantic map predicted by the collaborative perception module As global information : ; ; in, Represents a personalized parameter network based on the local information of each animal robot Output a set of parameters and , Represents a personalized network, using the parameters output by the personalized parameter network to process Generate animal robots Customized global information ; S33. Use knowledge distillation to achieve decentralized execution, using the agent's local information to distill personalized global information: ; ; ; in, It is a teacher network. Use global information and local information The generated personalized global information, It's a student network. Use local information Simulate the personalized global information generated by the teacher network and minimize the mean square error loss during training. accomplish; S34. Calculate the action value function of each agent, expressed as: ; Select the optimal action based on the calculated value of the action value function and use the QMIX algorithm for strategy training; S35, the training process is divided into two stages. The first stage provides each agent with personalized global information to calculate the action value function and individual strategy. The second stage uses offline knowledge distillation to distill personalized global information using local information. The execution stage directly uses local information to replace personalized global information. For each set of data, a semantic map representing global information is generated through the collaborative perception module. express.
4. The multi-animal robot collaborative human-machine interactive navigation system according to claim 3, characterized in that: The process of using the QMIX algorithm for strategy training is as follows: (1) Initialize hybrid network parameters ; (2) Initialize the parameters of the single agent Q network and experience buffer ; (4) Initialize the discount factor , Experience buffer capacity , learning rate and batch size ; (5) At each time step t, all agents perform actions according to their respective Q networks and obtain observations and rewards at the next moment, converting the tuple Store in experience buffer ; (6) When the number of tuples in the experience buffer is greater than , randomly sample a batch of data from the experience buffer; (7) For each observation in the data set , a semantic map representing global information is generated through the collaborative perception module, and express; (8) For each agent, use a recurrent network to generate its hidden state , use the global information personalization module to generate customized global information ; (9) Calculate the target Q value of each agent and the global target Q value , the Q-value loss function of each agent is ; (10) Calculate the global Q value using the hybrid network , the loss function of the hybrid network is: ; (11) Update the Q network parameters of each agent using the current batch of experience data And the parameters of the hybrid network ; (12) Repeat steps (5)-(11) until the strategy converges or the preset maximum number of training steps is reached.
Citation Information
Patent Citations
Multi-agent vehicle and road cloud integrated cooperative decision control architecture system and method based on federal reinforcement learning
CN118379878A
Multi-agent cooperation decision-making and training method
US20200125957A1