A Multi-Robot Vision Sharing and Recognition Collaboration Method and System
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-18
- Publication Date
- 2026-08-14
AI Technical Summary
其主要目的在于解决当前多机器人视觉识别单点感知局限、信息冗余与通信瓶颈、多源异构数据融合困难、协同效率低下等问题
Smart Images

Figure CN122560001A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of intelligent control technology, and in particular to a method and system for multi-robot visual sharing, recognition and collaboration. Background Technology
[0002] With the development of robotics technology, single robots have been widely used in various fields such as industrial manufacturing, logistics warehousing, and security inspection. In terms of visual perception, traditional solutions typically rely on the visual sensors carried by a single robot for environmental perception and target recognition. However, these solutions are limited by the field of view and perception capabilities of a single robot, and in complex scenarios, they are prone to blind spots, missed detections, or false detections due to occlusion, limited perspective, or environmental changes, making it difficult to meet the requirements for high robustness and full coverage perception.
[0003] To address these issues, solutions have emerged that utilize collaborative visual perception across multiple robots. One common approach involves a central node that centrally processes and fuses the visual information collected by each robot to achieve a more comprehensive understanding of the environment. While this method improves the consistency of perception, it places high demands on the bandwidth and stability of the communication network. Furthermore, the failure of the central node can lead to system-wide paralysis, limiting reliability and real-time performance. In addition, dynamically allocating tasks such as recognition and tracking based on real-time perception information during multi-robot collaborative operations to avoid task overlap or resource idleness is a key issue affecting system efficiency. Existing task allocation mechanisms are mostly based on fixed rules or simple scheduling strategies, making it difficult to adapt to dynamically changing environments and task requirements, and prone to problems such as response delays and uneven resource utilization.
[0004] Therefore, there is an urgent need for a multi-robot visual sharing and recognition collaboration method to solve the problems that the current field of multi-robot visual recognition still faces, such as the limitations of single-point perception, information redundancy and communication bottlenecks, difficulties in fusion of multi-source heterogeneous data, and low collaboration efficiency. Summary of the Invention
[0005] This disclosure provides a method and system for multi-robot visual sharing and collaborative recognition. Its main purpose is to solve the problems of single-point perception limitations, information redundancy and communication bottlenecks, difficulties in fusion of multi-source heterogeneous data, and low collaborative efficiency in current multi-robot visual recognition.
[0006] According to a first aspect of this disclosure, a multi-robot visual shared recognition collaborative method is provided, comprising:
[0007] Multiple robots are used to collect multi-source visual data of the environment, and the multi-source visual data is preprocessed to obtain local compressed visual feature data of each robot. The local compressed visual feature data of each robot is shared through a wireless communication network to generate a shared visual feature dataset. When a robot loses its target due to occlusion or malfunction, the robot closest to it automatically takes over its recognition task. Based on the shared visual feature dataset, a multimodal feature fusion and spatiotemporal alignment algorithm is used to process the data and generate globally consistent environmental perception results. Based on the globally consistent environmental perception results, a task allocation and scheduling strategy is adopted to assign identification, tracking or localization tasks to each robot. When the speed or acceleration of the tracked target exceeds a preset threshold or some robots fail, the robot autonomously adjusts the cooperative topology and task division among the robots.
[0008] Preferably, the step of using multiple robots to collect multi-source visual data of the environment and preprocessing the multi-source visual data to obtain compressed visual feature data for each robot includes: Multiple robots are used to collect at least one of the following environmental images: monocular images, binocular images, depth images, or infrared images. The multi-source visual data is processed through feature extraction, keyframe filtering, and semantic compression to obtain compressed visual feature data for each robot. The semantic compression is used to remove redundant information from the visual data and retain feature information related to the preset recognition task.
[0009] Preferably, sharing the compressed visual feature data of each robot via a wireless communication network includes: Each robot exchanges data based on a preset distributed communication protocol, which includes efficient encoding of the local compressed visual feature data and state synchronization within the robot group.
[0010] Preferably, the step of having the nearest robot automatically take over the identification task when a robot loses the target due to occlusion or malfunction includes: If the robot's confidence level in identifying the target is detected to be lower than the first threshold or communication is interrupted, other robots with the same perception coverage area will continue to identify and maintain the target's status based on shared information.
[0011] Preferably, the step of processing the shared visual feature dataset using a multimodal feature fusion and spatiotemporal alignment algorithm to generate a globally consistent environment perception result includes: The local compressed visual feature data from different robots are time-stamped and transformed to unify them under the same spatiotemporal reference system, generating globally consistent environmental perception results.
[0012] Preferably, the task allocation and scheduling strategy includes: constructing a graph model of the robot, task, environment, and relationships based on the globally consistent environmental perception results, and dynamically allocating identification, tracking, or localization tasks to each robot based on the optimized solution or strategy search of the graph model.
[0013] Preferably, the autonomous adjustment of the collaborative topology and task division among robots includes: When the speed or acceleration of the tracked target exceeds a preset threshold or some robots fail, the perception capabilities and task load of each robot are reassessed, and the data sharing links and task dependencies between robots are re-established.
[0014] Preferably, it further includes: assessing and arbitrating the confidence levels of recognition results from different robots targeting the same target or region, and fusing them to form a final recognition result with higher confidence.
[0015] According to a second aspect of this disclosure, a multi-robot vision-sharing and collaborative recognition system is provided, comprising: The visual acquisition module is deployed in each robot to collect multi-source visual data of the environment and perform local feature extraction, keyframe filtering and semantic compression. The communication and sharing module is used to establish a distributed data sharing mechanism among robots to transmit and share compressed visual feature data and support task takeover. The data fusion and alignment module is used to perform multimodal feature fusion and spatiotemporal alignment on visual feature data from different robots to generate globally consistent environmental perception results. The collaboration identification and task allocation module is used to perform task allocation and dynamic scheduling based on the globally consistent environmental perception results, and to autonomously adjust the collaborative topology and task division among robots when the environment changes suddenly or some robots fail.
[0016] Preferably, it further includes: a flexible collaboration management module; the flexible collaboration management module is communicatively connected to the communication and sharing module and the collaboration identification and task allocation module, and is used to generate collaboration strategy adjustment instructions and send them to the collaboration identification and task allocation module based on the monitored robot failure events or environmental dynamic change events.
[0017] In the embodiments provided in this disclosure, the limitations of a single robot's field of vision and the problem of occlusion are effectively overcome by sharing and complementing visual information from multiple robots. Even when some robots fail or their viewpoints are obstructed, the system can still maintain continuous and stable environmental perception and target recognition capabilities through visual information provided by other robots. By performing feature extraction and semantic compression on visual data locally, only high-value information is shared, and combined with an efficient encoding and transmission mechanism, the bandwidth load required for communication between multiple robots is significantly reduced, alleviating network bottlenecks. Through multimodal feature fusion and spatiotemporal alignment algorithms, the heterogeneity problem between visual data from different sensors, different locations, and different times is effectively solved, generating a globally unified and spatiotemporally consistent environmental perception model, providing a reliable basis for collaborative decision-making. Based on a globally perceptive intelligent task allocation and dynamic scheduling strategy, the system can autonomously and flexibly adjust the collaborative relationship and task division among robots according to real-time environmental changes and system status, thereby optimizing overall task execution efficiency and resource utilization.
[0018] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0019] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 A flowchart illustrating a multi-robot visual shared recognition and collaboration method provided for related technologies; Figure 2 This is a schematic diagram of the structure of a multi-robot vision sharing and recognition collaborative system provided in an embodiment of this disclosure. Detailed Implementation
[0020] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0021] A multi-robot visual sharing recognition and collaboration method and system according to an embodiment of the present disclosure is described below with reference to the accompanying drawings.
[0022] Figure 1 This is a flowchart illustrating a multi-robot visual shared recognition and collaboration method provided in an embodiment of this disclosure. Figure 1 As shown, the method includes the following steps: Step 101: Use multiple robots to collect multi-source visual data of the environment, and preprocess the multi-source visual data to obtain local compressed visual feature data of each robot. Step 102: Share the local compressed visual feature data of each robot through a wireless communication network to generate a shared visual feature dataset. When a robot loses its target due to occlusion or malfunction, the robot with the closest location will automatically take over its recognition task. Step 103: Based on the shared visual feature dataset, a multimodal feature fusion and spatiotemporal alignment algorithm is used to process the data and generate a globally consistent environment perception result. Step 104: Based on the globally consistent environmental perception results, a task allocation and scheduling strategy is adopted to assign identification, tracking or localization tasks to each robot, and when the speed or acceleration of the tracked target exceeds a preset threshold or some robots fail, the robot autonomously adjusts the cooperative topology and task division among the robots.
[0023] Based on the above embodiments, this embodiment will provide a detailed description of step 101: In one embodiment, multiple robots are used to collect at least one of monocular images, binocular images, depth images, or infrared images of the environment. The multi-source visual data is processed by feature extraction, keyframe filtering, and semantic compression to obtain compressed visual feature data of each robot. The semantic compression is used to remove redundant information in the visual data and retain feature information related to the preset recognition task.
[0024] Specifically, multiple robots are first dispersed in the environment to be perceived, according to task requirements or a pre-set deployment strategy. The types of visual sensors carried by each robot can be independent or partially overlapping, with typical configurations including, but not limited to, monocular cameras, stereo cameras, depth cameras, infrared thermal imagers, and multispectral or hyperspectral imaging devices, thus forming a multi-source visual data acquisition capability of the environment. After each robot starts, its visual sensors begin acquiring raw visual data streams at a preset sampling frequency. During this process, different robots may acquire raw images or image sequences with different viewpoints, illuminations, resolutions, and modalities due to differences in their motion states or environmental conditions. Subsequently, each robot independently performs a preprocessing process locally, which mainly includes the following steps: Initial quality assessment and screening: The robot first performs a preliminary quality assessment on the acquired raw visual data, such as detecting whether the image is blurred due to motion, overexposed or underexposed due to drastic changes in lighting, or invalid frames generated due to momentary sensor malfunction. Only frames that pass the quality assessment are sent to the subsequent processing pipeline, while unqualified data are discarded in real time or marked as low-confidence data.
[0025] Feature Extraction and Semantic Understanding: For the visual data that has passed quality screening, the robot runs a locally deployed feature extraction algorithm. This algorithm can be configured according to the task objective. For example, for object recognition tasks, the algorithm can be a convolutional neural network to extract the bounding boxes, category probability distributions, and key point features of objects in the image; for scene understanding tasks, the algorithm can be a semantic segmentation network to generate pixel-level semantic label maps of the image (such as ground, obstacles, target objects, etc.). This step transforms the raw pixel-level data into a structured, semantic high-level feature representation.
[0026] Keyframe Selection and Data Compression: To avoid the significant redundancy caused by transmitting continuous video streams, the robot employs a keyframe selection strategy. This strategy is based on the degree of information change between consecutive frames. For example, by calculating the differences in feature representations between adjacent frames (such as the cosine distance of feature vectors or the change ratio of the scene semantic map), when the change exceeds a preset threshold, the current frame is determined to be a keyframe. For the selected keyframes, the robot compresses and encodes their corresponding structured features (such as target feature vectors and semantic maps). The encoding method can employ lossless or lossy compression algorithms for feature data (such as quantization, entropy coding, etc.), aiming to retain the most core recognition and representation information with minimal data volume.
[0027] Metadata Association and Encapsulation: While generating compressed visual feature data, the robot generates and associates necessary metadata. This metadata includes at least the timestamp when the data was acquired, the robot's spatial pose information, the robot's own identifier, and the calibration parameters of the sensors used. Finally, the compressed visual feature data and metadata are encapsulated into a standardized data packet format.
[0028] Through the above preprocessing steps, each robot ultimately generates its own unique local compressed visual feature data package. Compared to the original visual data, this data package is significantly smaller in size and rich in semantic information relevant to subsequent tasks, laying a solid foundation for subsequent distributed sharing and efficient fusion.
[0029] Based on the above embodiments, this embodiment will provide a detailed description of step 102: In one embodiment, each robot exchanges data based on a preset distributed communication protocol, which includes efficient encoding of the local compressed visual feature data and state synchronization within the robot group.
[0030] If the robot's confidence level in identifying the target is detected to be lower than the first threshold or communication is interrupted, other robots with the same perception coverage area will continue to identify and maintain the target's status based on shared information.
[0031] Specifically, after each robot generates its local compressed visual feature data, it initiates a distributed data sharing mechanism via a wireless communication network. The sharing process follows a predefined communication protocol that specifies the rules for broadcasting, multicasting, or selective unicasting of data packets. Robots can periodically send their generated data packets to other robots or designated neighbor nodes in the network, while simultaneously listening for and receiving similar data packets from other nodes. All received local compressed visual feature data packets (including their metadata) from different robots are logically or physically aggregated, indexed, and managed, collectively forming a dynamically updated shared visual feature dataset.
[0032] Based on the above embodiments, this embodiment will provide a detailed description of step 103: In one embodiment, local compressed visual feature data from different robots are time-stamped and transformed to unify them under the same spatiotemporal reference system, generating globally consistent environmental perception results.
[0033] Specifically, a set of visual feature data to be processed is retrieved from a shared visual feature dataset. This set of data may come from two or more different robot nodes, and their corresponding original sensor types differ. Each data packet mainly contains: compressed visual feature vectors (or feature maps), timestamps, the robot pose (position and orientation) that acquired the data, and metadata such as sensor type identifiers.
[0034] The received data is then spatiotemporally aligned, an operation that includes two key sub-processes: Time alignment: Due to slight drift in the clocks of each robot or delays in data packet transmission, the timestamps of the data they report are not strictly synchronized. The system performs time interpolation or resampling on the feature data based on the timestamp sequence, projecting it onto a unified processing time window to achieve time axis synchronization.
[0035] Spatial Alignment: Since each robot has a different pose when collecting data, their observations are performed in different local coordinate systems. The system utilizes the pose information reported by the robots and preset sensor intrinsic and extrinsic parameters to uniformly transform the spatial location information contained in each piece of visual feature data to the same global coordinate system (e.g., the world coordinate system with the system's starting point as the origin) using a coordinate transformation matrix. This makes data that were originally scattered across different viewpoints and locations spatially comparable and superimposed.
[0036] After spatiotemporal alignment, data from different robots and sensors are initially aligned in time and space. The next step is multimodal feature fusion, including: Homogeneous data fusion: For visual features describing the same spatial region extracted by similar sensors (e.g., both RGB cameras), feature-level weighted fusion, selective fusion (e.g., selecting the highest confidence level), or redundancy elimination can be used to obtain a more robust and accurate feature representation. Heterogeneous data fusion: For features extracted by different types of sensors (e.g., visible light cameras and infrared thermal imagers), due to their different feature distributions, the system employs a pre-defined multimodal fusion strategy. For example, feature vectors from different modalities can be projected into a common shared subspace before fusion, or a multimodal neural network model can be used, with heterogeneous features as different input branches, performing feature interaction and fusion in the intermediate layers of the network. The purpose of fusion is to integrate the complementary advantages of different modal information. For example, visible light images provide rich texture and color information, while infrared images provide thermal radiation information; fusion can provide a more comprehensive understanding of the target.
[0037] After multimodal fusion and spatiotemporal alignment, the system integrates all input data to construct and output a globally consistent environmental perception result.
[0038] Based on the above embodiments, this embodiment will provide a detailed description of step 104: In one embodiment, a graph model of the robot, task, environment, and relationships is constructed based on the globally consistent environmental perception results, and the identification, tracking, or localization tasks of each robot are dynamically allocated based on the optimization solution or policy search of the graph model.
[0039] When the speed or acceleration of the tracked target exceeds a preset threshold or some robots fail, the perception capabilities and task load of each robot are reassessed, and the data sharing links and task dependencies between robots are re-established.
[0040] Specifically, after obtaining globally consistent environmental perception results, the task allocation process is initiated. These results include complete target state information, environmental structure data, and the real-time status of each robot. Based on this, the task allocation module invokes a preset task allocation and scheduling strategy. First, it decomposes the perception results into specific executable identification tasks, tracking tasks, and localization tasks. Then, based on each robot's sensor configuration, current position, remaining energy, and computational load, it performs capability matching and efficiency evaluation. An initial task allocation scheme is generated through optimization algorithms, assigning each task to the most suitable robot and issuing execution instructions.
[0041] During task execution, the system continuously monitors dynamic events and analyzes the motion state of all tracked targets in real time. When the instantaneous velocity or acceleration value of any target exceeds a preset safety threshold, the system immediately marks the event as a dynamic anomaly. Simultaneously, it monitors the communication status and task report data of each robot in the network. When communication timeouts, consecutive task failures, or proactive fault reporting occur, the system determines that a robot failure event has occurred. These two types of events will serve as the criteria for triggering the system's adjustment mechanism.
[0042] Once a dynamic anomaly or failure event is confirmed, the system immediately initiates an autonomous adjustment procedure. The system first assesses the scope of the event's impact, determining the set of tasks requiring reallocation and the currently available robot resources. Based on the latest global perception results, the task allocation module re-runs the optimization strategy, generating an adjusted task allocation scheme: for high-speed targets, robots with greater mobility or those ahead of the movement path will be prioritized to take over tracking; tasks undertaken by failed robots will be reallocated to lighter-loaded replacement units with corresponding perception capabilities. Simultaneously, the system updates the data flow and collaborative relationships between robots, establishing a new collaborative topology. After the adjustment scheme is issued, each robot completes task handover and switches collaborative modes, ensuring the continuity of overall task execution and system stability.
[0043] This embodiment also includes: performing confidence assessment and arbitration on the recognition results from different robots for the same target or area, and fusing them to form a final recognition result with higher confidence.
[0044] Specifically, during operation, visual perception data from multiple robots are collected. When the system detects that the recognition results reported by different robots point to the same target or area in the environment, it initiates a confidence arbitration process. First, all recognition result records for that target or area are extracted from the shared dataset. Each record contains the robot's recognition conclusion, the corresponding visual feature data, and the original recognition confidence score and sensor type information.
[0045] The system quantifies and evaluates each recognition result based on a pre-defined confidence assessment model. Evaluation metrics include the inherent confidence of the recognition algorithm, sensor data quality score, the quality of the robot's observation perspective, and the weight of the robot's historical recognition accuracy. The system calculates a comprehensive confidence score for each recognition result based on these multi-dimensional metrics, reflecting the reliability of the result under the current observation conditions.
[0046] After completing the comprehensive confidence assessment, the system initiates an arbitration mechanism. All recognition results are sorted according to their comprehensive confidence, and the results with the highest confidence are selected as candidates. If there are category conflicts among the results with the highest confidence, the system calls a pre-trained consistency verification model to adjudicate. This model determines the most reasonable recognition conclusion by analyzing the consistency of underlying visual features. If the results with the highest confidence are consistent across categories, they are directly adopted as the benchmark.
[0047] Finally, the system performs a feature-level fusion operation, weighting and fusing the visual feature data corresponding to the adopted recognition results. The weights are proportional to the overall confidence level. This fusion generates a richer and more robust description at the feature level, and based on this, outputs a final recognition result with higher confidence. This result will be used to update the global environmental perception state and serve as an authoritative basis for subsequent task allocation and collaborative decision-making.
[0048] This embodiment provides a multi-robot visual sharing and recognition collaboration method. By sharing and complementing visual information from multiple robots, it effectively overcomes the limitations of a single robot's field of vision and occlusion problems. Even when some robots fail or their viewpoints are obstructed, the system can still maintain continuous and stable environmental perception and target recognition capabilities through visual information provided by other robots. By performing feature extraction and semantic compression on visual data locally, only high-value information is shared, and combined with an efficient encoding and transmission mechanism, the bandwidth load required for communication between multiple robots is significantly reduced, alleviating network bottlenecks. Through multimodal feature fusion and spatiotemporal alignment algorithms, the heterogeneity problem between visual data from different sensors, locations, and times is effectively solved, generating a globally unified and spatiotemporally consistent environmental perception model, providing a reliable basis for collaborative decision-making. Based on a globally perceptive intelligent task allocation and dynamic scheduling strategy, the system can autonomously and flexibly adjust the collaborative relationships and task division among robots according to real-time environmental changes and system status, thereby optimizing overall task execution efficiency and resource utilization.
[0049] Based on the above embodiments, this embodiment describes a multi-robot vision sharing and recognition collaborative system, such as... Figure 2 As shown, the details are as follows: The visual acquisition module is deployed in each robot to collect multi-source visual data of the environment and perform local feature extraction, keyframe filtering and semantic compression. The system in this embodiment is a multi-robot collaborative perception and decision-making platform. Its core functional modules can be distributed across a robot cluster or partially centralized on an edge server. Assume the system consists of N (N≥2) mobile robot nodes. Each robot node is equipped with: Hardware layer: at least one visual sensor (such as a camera), computing unit (such as an embedded AI chip), wireless communication module, and navigation and positioning unit.
[0050] Software / System Layer: This layer runs instances of the functional modules described in the system claims of this invention. These modules can be independent processes or services at the software level, connected via internal communication interfaces (such as ROS topics, services, or custom message buses).
[0051] After the system is powered on, the initialization process is executed: Network establishment: The communication and sharing modules of each robot are activated, and they discover and establish connections with each other through wireless communication networks (such as Ad-hoc networks or connected to the same access point), forming a distributed communication network.
[0052] Parameter synchronization: Each module loads unified configuration parameters, such as: visual feature extraction network model version, keyframe selection threshold, data compression algorithm identifier, task allocation strategy model, etc.
[0053] Spatiotemporal reference alignment: The data fusion and alignment module (which may exist on one or more designated nodes) coordinates the clock synchronization of each robot and agrees on a unified global coordinate system. Each robot reports its pose to this module through its localization unit for initial registration.
[0054] The operation process of each module in the system: The visual acquisition module is deployed in each robot to collect multi-source visual data of the environment and perform local feature extraction, keyframe filtering and semantic compression. Specifically, the trigger is as follows: after the system starts, the vision acquisition modules of each robot remain active, driving the local vision sensors to scan the environment.
[0055] Processing: For each frame of raw visual data captured, this module immediately executes a local preprocessing pipeline: Feature extraction: Invoking a pre-trained neural network model to extract high-dimensional feature maps from images or generate structured semantic information for specific tasks (such as detection and segmentation).
[0056] Keyframe filtering: Based on the feature similarity or information entropy change between consecutive frames, determine whether the current frame is a keyframe carrying new information.
[0057] Semantic compression: Compress and encode the feature information or semantic results corresponding to keyframes to generate "local compressed visual feature data" with a significantly reduced data volume.
[0058] Output: This module encapsulates the generated compressed data, along with metadata such as timestamps, local pose, and sensor IDs, into a standard message and publishes it to the local system's internal message bus for consumption by other modules.
[0059] The communication and sharing module is used to establish a distributed data sharing mechanism among robots to transmit and share compressed visual feature data and support task takeover. Specifically, subscription and sending: The communication and sharing module subscribes to compressed data messages from the vision acquisition module. Once received, the module serializes the data and encapsulates it into network data packets according to a preset communication protocol (such as an efficient multicast protocol based on UDP), and sends it to other robots in the cluster via a wireless network.
[0060] Reception and Distribution: Simultaneously, this module continuously monitors the network, receiving similar data packets from other robots. Upon receiving the data, it performs deserialization and integrity verification, and then publishes the valid "other-machine compressed visual feature data" to the local internal message bus.
[0061] Status Maintenance: This module maintains a neighbor robot status table, recording communication link quality, most recent reception time, etc., providing a basis for communication-level status judgment for task takeover.
[0062] Generate a shared dataset: At the system level, all compressed visual feature data exchanged through this module together constitute a dynamic shared visual feature dataset. This dataset can be stored in a distributed manner in the memory caches of each robot, or it can be centrally indexed by a single node.
[0063] The data fusion and alignment module is used to perform multimodal feature fusion and spatiotemporal alignment on visual feature data from different robots to generate globally consistent environmental perception results. Specifically, data aggregation: The data fusion and alignment module simultaneously subscribes to compressed visual feature data streams from the local visual acquisition module and the communication and sharing module (forwarding data from other machines).
[0064] Spatiotemporal alignment: For the aggregated data, this module first uses the timestamps and pose metadata in the data packets to invoke a spatiotemporal alignment algorithm. This algorithm transforms all data to a unified global coordinate system and the same time reference. For data with asynchronous time, interpolation or prediction may be used for alignment.
[0065] Feature fusion: Under a unified spatiotemporal reference frame, this module performs multimodal feature fusion on feature data from different robots (possibly different sensor modalities) describing the same environmental region or the same target. The fusion algorithm can be feature-level concatenation, weighted averaging, or attention-based fusion, etc.
[0066] Generate global awareness results: After alignment and fusion, this module generates globally consistent environment awareness results. This result can be a global semantic map labeled with various targets, obstacles, and regions of interest, or a list of target states (position, velocity, category, etc.) with global coordinates. This result is then encapsulated as a message publication.
[0067] The collaboration identification and task allocation module is used to perform task allocation and dynamic scheduling based on the globally consistent environmental perception results, and to autonomously adjust the collaborative topology and task division among robots when the environment changes suddenly or some robots fail.
[0068] Specifically, the perception input is the globally consistent environmental perception results published by the data fusion and alignment module, which are subscribed to by the collaborative recognition and task allocation module.
[0069] Task Decision: Based on the global perception results, this module executes a task allocation and scheduling strategy. This strategy may be an optimization algorithm or a learned strategy model. It comprehensively considers various target attributes (such as priority), robot states (such as position, battery level, sensor capabilities, and current task load), and environmental constraints to calculate an optimized task allocation scheme. For example, robot A may be assigned to track target T1, and robot B may be assigned to scan region Z2.
[0070] Command Issuance: This module issues task commands (including task type, target description, behavioral parameters, etc.) to the corresponding robots. The commands are transmitted to the actuators or controllers of each robot via internal communication.
[0071] Dynamic scheduling: During task execution, this module continuously monitors the global perception results and the status feedback from each robot. When a new high-priority target appears, or when an abnormality in robot task execution is detected (such as the target position not being updated for a long time), dynamic scheduling is triggered to recalculate and reassign tasks.
[0072] The flexible collaboration management module is communicatively connected to the communication and sharing module and the collaboration identification and task allocation module. It is used to generate collaboration strategy adjustment instructions based on the monitored robot failure events or dynamic environmental change events and send them to the collaboration identification and task allocation module.
[0073] Specifically, system monitoring: The flexible collaboration management module serves as the monitoring and coordination hub of the system, continuously monitoring network status reports (such as robot offline) from the communication and sharing module, local health status reports from each robot, and abnormal task execution events from the collaboration identification and task allocation module.
[0074] Event Judgment: This module defines and identifies “environmental mutation events” (such as the sudden appearance of a large number of unmodeled dynamic obstacles in the global perception results) and “robot failure events” (such as communication interruption accompanied by task failure).
[0075] Triggering Reconfiguration: Once the above-mentioned event is detected, this module sends a collaboration strategy adjustment instruction to the collaboration identification and task allocation module. This instruction may include the event type, information about the affected robot / area, and suggested adjustment directions (such as reassigning tasks or changing the collaboration formation).
[0076] Coordinated execution: After receiving the instruction, the collaboration recognition and task allocation module uses it as a higher-order input and incorporates it into the next round of task allocation decisions, thereby realizing the autonomous adjustment of the collaborative topology and task division among robots.
[0077] This embodiment provides a multi-robot visual sharing and recognition collaborative system. Through real-time sharing and intelligent fusion of visual information from multiple robots, the system constructs a global environmental perception model that surpasses the capabilities of any single robot. This system effectively overcomes the inherent blind spots, occlusion limitations, and sensor limitations of single-point perception. Even if individual robots malfunction or encounter severe interference, their perception tasks can be seamlessly taken over by neighboring nodes based on shared data. This maintains continuous, stable, and full-coverage perception capabilities in dynamic and complex environments, greatly improving the overall system's task reliability and environmental adaptability. Local preprocessing and semantic compression significantly reduce the amount of raw data, and combined with an efficient distributed communication mechanism, significantly reduce network bandwidth pressure and communication latency. The intelligent task allocation strategy dynamically assigns tasks such as recognition, tracking, and localization based on the global situation and the real-time status of each robot, avoiding task overlap or resource idleness. This achieves efficient coordination and optimized utilization of computing resources, communication resources, and robot maneuvering resources, thereby improving overall task execution efficiency.
[0078] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0079] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0080] The target detection method provided in this disclosure has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this disclosure. The descriptions of the embodiments above are only for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this disclosure without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this disclosure.
Claims
1. A multi-robot visual shared recognition collaborative method, characterized in that, include: Multiple robots are used to collect multi-source visual data of the environment, and the multi-source visual data is preprocessed to obtain local compressed visual feature data of each robot. The local compressed visual feature data of each robot is shared through a wireless communication network to generate a shared visual feature dataset. When a robot loses its target due to occlusion or malfunction, the robot closest to it automatically takes over its recognition task. Based on the shared visual feature dataset, a multimodal feature fusion and spatiotemporal alignment algorithm is used to process the data and generate globally consistent environmental perception results. Based on the globally consistent environmental perception results, a task allocation and scheduling strategy is adopted to assign identification, tracking or localization tasks to each robot. When the speed or acceleration of the tracked target exceeds a preset threshold or some robots fail, the robot autonomously adjusts the cooperative topology and task division among the robots.
2. The multi-robot vision sharing and recognition collaboration method according to claim 1, characterized in that, The process of using multiple robots to collect multi-source visual data of the environment and preprocessing the multi-source visual data to obtain compressed visual feature data for each robot includes: Multiple robots are used to collect at least one of the following environmental images: monocular images, binocular images, depth images, or infrared images. The multi-source visual data is processed through feature extraction, keyframe filtering, and semantic compression to obtain compressed visual feature data for each robot. The semantic compression is used to remove redundant information from the visual data and retain feature information related to the preset recognition task.
3. The multi-robot vision sharing and recognition collaboration method according to claim 1, characterized in that, The sharing of compressed visual feature data of each robot via a wireless communication network includes: Each robot exchanges data based on a preset distributed communication protocol, which includes efficient encoding of the local compressed visual feature data and state synchronization within the robot group.
4. The multi-robot vision sharing and recognition collaboration method according to claim 1, characterized in that, When a robot loses its target due to occlusion or malfunction, the robot closest to it automatically takes over the identification task, including: If the robot's confidence level in identifying the target is detected to be lower than the first threshold or communication is interrupted, other robots with the same perception coverage area will continue to identify and maintain the target's status based on shared information.
5. The multi-robot vision sharing and recognition collaboration method according to claim 1, characterized in that, The process of generating globally consistent environment perception results based on the shared visual feature dataset, using multimodal feature fusion and spatiotemporal alignment algorithms, includes: The local compressed visual feature data from different robots are time-stamped and transformed to unify them under the same spatiotemporal reference system, generating globally consistent environmental perception results.
6. The multi-robot vision sharing and recognition collaboration method according to claim 1, characterized in that, The task allocation and scheduling strategy includes: constructing a graph model of robots, tasks, environment and relationships based on the globally consistent environmental perception results, and dynamically allocating identification, tracking or localization tasks to each robot based on the optimization solution or strategy search of the graph model.
7. The multi-robot vision sharing and recognition collaboration method according to claim 6, characterized in that, The autonomous adjustment of the collaborative topology and task division among robots includes: When the speed or acceleration of the tracked target exceeds a preset threshold or some robots fail, the perception capabilities and task load of each robot are reassessed, and the data sharing links and task dependencies between robots are re-established.
8. The multi-robot vision sharing and recognition collaboration method according to claim 1, characterized in that, Also includes: The confidence level of recognition results from different robots targeting the same target or area is assessed and arbitrated, and then merged to form a final recognition result with higher confidence.
9. A multi-robot vision sharing and recognition collaborative system, characterized in that, include: The visual acquisition module is deployed in each robot to collect multi-source visual data of the environment and perform local feature extraction, keyframe filtering and semantic compression. The communication and sharing module is used to establish a distributed data sharing mechanism among robots to transmit and share compressed visual feature data and support task takeover. The data fusion and alignment module is used to perform multimodal feature fusion and spatiotemporal alignment on visual feature data from different robots to generate globally consistent environmental perception results. The collaboration identification and task allocation module is used to perform task allocation and dynamic scheduling based on the globally consistent environmental perception results, and to autonomously adjust the collaborative topology and task division among robots when the environment changes suddenly or some robots fail.
10. The multi-robot vision sharing and recognition collaborative system according to claim 9, characterized in that, Also includes: Flexible collaboration management module; The flexible collaboration management module is communicatively connected to the communication and sharing module and the collaboration identification and task allocation module. It is used to generate collaboration strategy adjustment instructions based on the monitored robot failure events or dynamic environmental change events and send them to the collaboration identification and task allocation module.