Cluster robot cross-view collaborative sensing method and system for aircraft manufacturing
By constructing a point cloud collaborative three-dimensional detection model with cross-view angle consistent spatial mapping, cluster robots can achieve high-precision three-dimensional object detection during aircraft manufacturing, solving the problems of limited vision and inconsistency of view of a single robot, and improving perceptual consistency and robustness.
Patent Information
- Application Number
- CN202510990804.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-07-18
AI Technical Summary
During the aircraft manufacturing process, the perceived field of view and detection range of a single robot is limited, and it is impossible to perceive the entire manufacturing and assembly scene in full scene. There are spatial deviations and feature inconsistencies in the perceived data at different perspectives, which affects the information fusion and perceived consistency.
A point cloud collaborative three-dimensional detection model based on cross-view angle consistent spatial mapping is constructed. Through the cluster robot collaborative perception platform, three-dimensional point cloud data is obtained using lidar, and a motion capture system is used to obtain the robot position pose. Multi-agent feature extraction, cross-view angle feature fusion and feature reconstruction modules are constructed. The model parameters are optimized using joint loss function to realize multi-robot perception collaboration.
It improves the adaptability and expression integrity of cluster robots to complex structural environments, realizes efficient and reliable collaborative perception of multiple agents, and can conduct high-precision three-dimensional detection of large components, staff and operation robots in the aircraft manufacturing process, supporting target detection of large-scale manufacturing scenarios.
Smart Images

Figure CN120495643A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of collaborative three-dimensional perception technology of cluster robots in large-scale intelligent manufacturing scenarios, and in particular relates to a cross-perspective collaborative perception method and system of cluster robots for aircraft manufacturing. Background Art
[0002] The manufacturing and assembly of aviation equipment is a core component of economic development, especially during the development of domestically produced large aircraft. The quality of manufacturing and assembly directly determines the overall performance and flight safety of the aircraft. During the assembly of typical large, thin-walled structures such as aircraft skins, wings, and fuselages, systematic inspections of key areas are required. This includes comprehensive inspections of the initial state of large components like skins, including stress distribution, structural deformation, attitude inclination, and ground support status. This allows for the early identification of potential structural anomalies or safety risks, providing reliable data support and assurance for subsequent high-precision assembly operations.
[0003] Currently, the assembly and inspection process of complex components on large-scale aircraft still relies mainly on manual operation to identify key structural features, assess status, and troubleshoot anomalies. This method is not only inefficient and has limited accuracy, but also has problems such as large human subjective errors and heavy operational burdens, making it difficult to meet the needs of high-reliability and high-consistency inspections. With the continuous advancement of robotics and artificial intelligence, promoting the unmanned and intelligent upgrade of aircraft manufacturing inspections has become an important part of intelligent manufacturing and a key measure to ensure aircraft assembly quality and flight safety. Therefore, conducting research on inspection robot perception technology for aircraft manufacturing scenarios is a trend in the high-quality development of major equipment.
[0004] When robots conduct inspections of aircraft manufacturing processes, achieving high-precision three-dimensional target monitoring of large-scale manufacturing scenarios is key. Three-dimensional target detection can provide position and size information for multiple targets in aircraft manufacturing scenarios, including large aircraft components, workers, and operating robots. Only with high-precision three-dimensional target detection can robots complete intelligent and accurate safety inspections during the aircraft manufacturing process. In recent years, deep learning-based three-dimensional target detection methods have gradually gained attention, but they still cannot meet the engineering requirements for high-precision three-dimensional target detection by robots in aircraft manufacturing scenarios. The following difficulties are faced during the specific implementation process: 1. Aircraft manufacturing scenarios cover a wide area. The perception field and detection range of a single robot are limited, making it impossible to fully perceive the entire aircraft manufacturing and assembly scene. Therefore, collaborative perception through swarm robots is required to cover a larger area and a wider field of view, thereby more comprehensively detecting targets in the environment.
[0005] 2. Due to the diverse distribution of robots, the acquired LiDAR point clouds exhibit significant perspective differences. This leads to spatial deviations and feature inconsistencies in the perception data of the same target from different perspectives, compromising information fusion and perception consistency. To achieve effective collaboration in swarm perception systems, it is imperative to address the alignment and fusion of multi-perspective point cloud data to improve the consistency and robustness of perception results. Summary of the Invention
[0006] In response to the above technical problems, the present invention provides a cross-perspective collaborative perception method and system for cluster robots in aircraft manufacturing. Its purpose is to optimize the technical problems faced by the robot's three-dimensional target detection process in large-scale aircraft manufacturing scenes, such as limited perception field of view and detection range, and inconsistency in cross-perspective perception.
[0007] The technical solution adopted by the present invention to solve the technical problem is: A cross-view collaborative perception method for swarm robots in aircraft manufacturing, comprising the following steps: S100: Build a collaborative 3D perception platform for swarm robots. Each robot uses LiDAR to scan and acquire 3D point cloud data from a wide range of aircraft manufacturing scenes. The motion capture system then acquires each robot's real-time position and posture. S200: Label and segment the 3D point cloud data scanned by the swarm robot, and complete the production of collaborative point cloud detection datasets for large-scale scenarios in aircraft manufacturing. S300: Builds a point cloud collaborative 3D detection model based on cross-view consistency spatial mapping, including a point cloud feature extraction module, a feature fusion module, a feature reconstruction module, and a detection head; S400: pre-processing the point cloud data in the training set and inputting it into the point cloud feature extraction module for feature encoding to obtain multi-agent features; S500: Input the multi-agent features into the feature fusion module for cross-view consistency feature mapping and cross-view difference feature mapping to obtain consistency features and difference features respectively, and perform feature fusion on the consistency features and difference features to obtain multi-agent fusion features; S600: Inputting the multi-agent fusion features into the feature reconstruction module to perform feature reconstruction to obtain reconstructed features; S700: Input the multi-agent fusion features into the detection head to obtain the prediction results. Based on the consistency features, difference features, reconstructed features and prediction results, the parameters of the point cloud collaborative 3D detection model are updated through the set loss function until the point cloud collaborative 3D detection model converges; the validation set and test set are respectively input into the trained point cloud collaborative 3D detection model for calculation to obtain the 3D detection results, verify the effectiveness and generalization of the model, and then deploy it on the platform.
[0008] Preferably, S100 includes: S110: Assemble and debug the mechanical and electrical structures of the cluster robots. The robots are numbered as follows: 、 、 A Livox non-repeating scanning LiDAR is installed in the slot directly in front of the swarm robot, with a reserved LiDAR installation location. The Livox LiDAR is connected to the Nvidia Orin edge computing device inside the swarm robot to perform laser scanning, data transmission, and model deployment. The Livox LiDAR follows the swarm robot's movements. S120: Using swarm robots 、 、 The Livox laser radar carried out point cloud data collection for a wide range of aircraft manufacturing scenes, and obtained point cloud data respectively. 、 、 ,The real-time pose of each robot is obtained through the motion capture system.
[0009] Preferably, S200 includes: S210: Use the point cloud annotation software LabelCloud to label the point cloud data: 、 、 Perform coordinate system transformation, and the real-time position of the robot obtained by the motion capture system will be 、 、 Mapped to the unified world coordinate system, forming a 、 、 Global collaborative point cloud data of point cloud information , for point cloud data Perform three-dimensional detection frame and category annotation, and the labeled data set is recorded as ; S220: Labeling the dataset Divide and label the dataset Randomly take out 65% of the data for training, as the training set, denoted as , the labeled dataset Randomly take out the equivalent of 15% of the data is used for verification, as the verification set, denoted as , the labeled dataset The rest of is used as the test set, which is , used to test the generalization of the model.
[0010] Preferably, S400 includes: S410: Point cloud data acquired by each robot 、 、 With the laser radar as the center of the coordinate axis, the robot's forward direction is the x-axis direction, the robot's left direction is the y-axis direction, and the robot's upward direction is the z-axis direction. The point cloud is mathematically represented as follows: A vector representation of size, the content is ,in are the three-dimensional coordinates of the point cloud data, is the reflection intensity of point cloud data, is the number of point clouds; S420: Select point cloud data 、 、 The coordinate range of the point cloud on the coordinate axis is specified as , within this range, the point cloud is divided into columns, and the specifications of each point cloud column are , take the center of each point cloud column as the center point coordinate, and then calculate the relative coordinates of each point and the center of the point cloud column , and use it as a supplementary representation of the point cloud, specifically A vector representation of size, the content is ; S430: Use three PointPillars point cloud feature extraction networks with the same structure to extract 、 、 Point cloud data is feature encoded to obtain multi-agent features 、 、 .
[0011] Preferably, S500 includes: S510: Yes 、 、 Perform cross-view consistency feature mapping, specifically: ; ; ; in is a cross-view consistency spatial mapping layer, is the fully connected layer, is the activation function, is layer normalization, 、 、 Share a cross-view consistency spatial mapping layer, through Will 、 、 Mapped to the consistent feature space, the output is ; S520: Yes 、 、 Perform cross-view difference feature mapping, specifically: ; ; ; in 、 、 It is a spatial mapping layer that maps differences across perspectives. is the fully connected layer, is the activation function, is layer normalization, 、 、 Using different cross-view difference spatial mapping layers, 、 、 Will 、 、 Map to different feature spaces to obtain differential features 、 、 ; S530: consistency features and differential characteristics 、 、 Perform feature fusion, specifically: ; ; ; in Is a splicing operation, use Consistency characteristics of cluster robots and differential characteristics 、 、 Perform splicing at the feature level to obtain splicing features , and They are maximum pooling and average pooling operations, which are used for feature selection and dimensionality reduction. The pooled features are then concatenated to obtain features. , is the feature fusion layer, is a three-dimensional convolutional layer, is the activation function, through Perform feature fusion to obtain multi-agent fusion features .
[0012] Preferably, S600 includes: For consistency features and differential characteristics 、 、 Perform feature reconstruction, specifically: ; ; in Is a consistent feature and differential characteristics of and, It is the feature reconstruction layer, through Reconstruct the consistent features and the different features, and 、 、 The multi-agent features of swarm robots correspond one to one.
[0013] Preferably, in S700, the multi-agent fusion features are input to the detection head to obtain the prediction results, including: The detection head completes the classification and regression tasks, Predictions are made through different network layers, specifically: ; ; in, Used to predict classification scores, it consists of a 1×1 two-dimensional convolution layer. The regression of the three-dimensional detection box consists of a 1×1 two-dimensional convolutional layer.
[0014] Preferably, the loss function in S700 is: ; ; ; ; ; in, , , , , is the prediction box classification loss, is the 3D detection box regression loss, is the cross-view consistency-difference feature reconstruction loss, is the cross-view consistency distance loss, For the total loss.
[0015] Preferably, in S700, the validation set and the test set are respectively input into the trained point cloud collaborative 3D detection model to perform calculations to obtain 3D detection results, verify the validity and generalization of the model, and perform platform deployment, including: After completing the training batches, the trained 3D detection model is evaluated. The validation set is used to validate the model and select the best cluster robot 3D detection model. , the test set Input 3D detection model , get the test results ; 3D detection model It is deployed on the NVIDIA Orin edge computing platform of cluster robots, first converting it into ONNX intermediate data, and then converting the model into TensorRT for model acceleration, enabling cluster robots to collaboratively detect targets in a large range of scenes.
[0016] A swarm robot cross-view collaborative perception system for aircraft manufacturing, including a 3D point cloud data acquisition module, a data set creation module, a 3D detection model building module, a feature extraction module, a feature fusion module, a feature reconstruction module, and a training and deployment module; The 3D point cloud data acquisition module is used to build a collaborative 3D perception platform for swarm robots. Each robot uses a lidar to scan and acquire 3D point cloud data in a large-scale aircraft manufacturing scene, and the real-time position and posture of each robot is obtained through a motion capture system. The dataset creation module is used to label and segment the 3D point cloud data scanned by the swarm robot, completing the creation of collaborative point cloud detection datasets for large-scale scenarios in aircraft manufacturing. 3D detection model building module, used to build a point cloud collaborative 3D detection model based on cross-view consistency spatial mapping, including point cloud feature extraction module, feature fusion module, feature reconstruction module and detection head; The feature extraction module is used to pre-process the point cloud data in the training set and perform feature encoding to obtain multi-agent features; The feature fusion module is used to map the cross-view consistency feature and the cross-view difference feature to obtain the consistency feature and difference feature respectively, and fuse the consistency feature and the difference feature to obtain the multi-agent fusion feature; The feature reconstruction module is used to reconstruct the multi-agent fusion features and obtain the reconstructed features; The training and deployment module is used to input the multi-agent fusion features into the detection head to obtain the prediction results. Based on the consistency features, difference features, reconstructed features and prediction results, the parameters of the point cloud collaborative 3D detection model are updated through the set loss function until the point cloud collaborative 3D detection model converges; the validation set and test set are respectively input into the trained point cloud collaborative 3D detection model for calculation to obtain the 3D detection results, verify the effectiveness and generalization of the model, and then deploy it on the platform.
[0017] The above-mentioned cross-perspective collaborative perception method and system of swarm robots for aircraft manufacturing can perform high-precision three-dimensional detection of large aircraft components, workers, operating robots, etc. in the aircraft manufacturing process, and provide swarm robots with the position and spatial information of targets in large-scale manufacturing scenarios; in view of the problem of perspective perception differences faced by swarm robots in large-scale aircraft manufacturing scenarios, a point cloud three-dimensional detection model based on cross-perspective consistent spatial mapping is proposed. This collaborative fusion mechanism not only ensures the spatial consistency of multi-agent perception information, but also retains the differences in local details under different perspectives, thereby improving the system's adaptability and expression integrity to complex structural environments, making multi-agent collaborative perception more efficient and reliable; and it has the potential value of being applied to three-dimensional detection in any large-scale industrial manufacturing scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a flow chart of a cross-view collaborative perception method of swarm robots for aircraft manufacturing in one embodiment of the present invention; Figure 2 Schematic diagram of a collaborative three-dimensional perception platform of swarm robots in one embodiment of the present invention; Figure 3 Schematic diagram of a point cloud collaborative 3D detection model based on cross-view consistency spatial mapping in one embodiment of the present invention. DETAILED DESCRIPTION
[0019] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is further described in detail below with reference to the accompanying drawings.
[0020] In one embodiment, Figure 1 As shown, a cross-view collaborative perception method of swarm robots for aircraft manufacturing includes the following steps: S100: Build a collaborative 3D perception platform for swarm robots. Each robot uses LiDAR to scan and acquire 3D point cloud data from a wide range of aircraft manufacturing scenes. The motion capture system then acquires each robot's real-time position and posture. S200: Label and segment the 3D point cloud data scanned by the swarm robot, and complete the production of collaborative point cloud detection datasets for large-scale scenarios in aircraft manufacturing. S300: Builds a point cloud collaborative 3D detection model based on cross-view consistency spatial mapping, including a point cloud feature extraction module, a feature fusion module, a feature reconstruction module, and a detection head; S400: pre-processing the point cloud data in the training set and inputting it into the point cloud feature extraction module for feature encoding to obtain multi-agent features; S500: Input the multi-agent features into the feature fusion module for cross-view consistency feature mapping and cross-view difference feature mapping to obtain consistency features and difference features respectively, and perform feature fusion on the consistency features and difference features to obtain multi-agent fusion features; S600: Inputting the multi-agent fusion features into the feature reconstruction module to perform feature reconstruction to obtain reconstructed features; S700: Input the multi-agent fusion features into the detection head to obtain the prediction results. Based on the consistency features, difference features, reconstructed features and prediction results, the parameters of the point cloud collaborative 3D detection model are updated through the set loss function until the point cloud collaborative 3D detection model converges; the validation set and test set are respectively input into the trained point cloud collaborative 3D detection model for calculation to obtain the 3D detection results, verify the effectiveness and generalization of the model, and then deploy it on the platform.
[0021] The method proposed in the present invention can realize collaborative three-dimensional inspection of cluster robots in large-scale aircraft manufacturing scenarios, and can quickly provide cluster robots with position and spatial information of large aircraft components, staff, and operating robots, so that cluster robots can conduct systematic inspections of key areas in aircraft manufacturing scenarios, and help cluster robots avoid obstacles, plan and control in large-scale manufacturing scenarios, so that they can better complete operations and assembly tasks, improve the assembly efficiency of cluster robots, increase the safety and stability of cluster robots' operations, and promote the intelligent, high-quality and high-speed development of domestic large aircraft manufacturing.
[0022] In one embodiment, Figure 2 As shown, S100 includes: S110: Assemble and debug the mechanical and electrical structures of the cluster robots. The robots are numbered as follows: 、 、 A Livox non-repeating scanning LiDAR is installed in the slot directly in front of the swarm robot, with a reserved LiDAR installation location. The Livox LiDAR is connected to the Nvidia Orin edge computing device inside the swarm robot to perform laser scanning, data transmission, and model deployment. The Livox LiDAR follows the swarm robot's movements. S120: Using swarm robots 、 、 The Livox laser radar carried out point cloud data collection for a wide range of aircraft manufacturing scenes, and obtained point cloud data respectively. 、 、 ,The real-time pose of each robot is obtained through the motion capture system.
[0023] In one embodiment, S200 includes: S210: Use the point cloud annotation software LabelCloud to label the point cloud data: 、 、 Perform coordinate system transformation, and the real-time position of the robot obtained by the motion capture system will be 、 、 Mapped to the unified world coordinate system, forming a 、 、 Global collaborative point cloud data of point cloud information , for point cloud data Perform three-dimensional detection frame and category annotation, and the labeled data set is recorded as ; S220: Labeling the dataset Divide and label the dataset Randomly take out 65% of the data for training, as the training set, denoted as , the labeled dataset Randomly take out the equivalent of 15% of the data is used for verification, as the verification set, denoted as , the labeled dataset The rest of is used as the test set, which is , used to test the generalization of the model.
[0024] In one embodiment, Figure 3 As shown, S400 includes: S410: Point cloud data acquired by each robot 、 、 With the laser radar as the center of the coordinate axis, the robot's forward direction is the x-axis direction, the robot's left direction is the y-axis direction, and the robot's upward direction is the z-axis direction. The point cloud is mathematically represented as follows: A vector representation of size, the content is ,in are the three-dimensional coordinates of the point cloud data, is the reflection intensity of point cloud data, is the number of point clouds; S420: Select point cloud data 、 、 The coordinate range of the point cloud on the coordinate axis is specified as , within this range, the point cloud is divided into columns, and the specifications of each point cloud column are , take the center of each point cloud column as the center point coordinate, and then calculate the relative coordinates of each point and the center of the point cloud column , and use it as a supplementary representation of the point cloud, specifically A vector representation of size, the content is ; S430: Use three PointPillars point cloud feature extraction networks with the same structure to extract 、 、 Point cloud data is feature encoded to obtain multi-agent features 、 、 .
[0025] In one embodiment, Figure 3 As shown, S500 includes: S510: Yes 、 、 Perform cross-view consistency feature mapping, specifically: ; ; ; in is a cross-view consistency spatial mapping layer, is the fully connected layer, is the activation function, is layer normalization, 、 、 Share a cross-view consistency spatial mapping layer, through Will 、 、 Mapped to the consistent feature space, the output is ; S520: Yes 、 、 Perform cross-view difference feature mapping, specifically: ; ; ; in 、 、 It is a spatial mapping layer that maps differences across perspectives. is the fully connected layer, is the activation function, is layer normalization, 、 、 Using different cross-view difference spatial mapping layers, 、 、 Will 、 、 Map to different feature spaces to obtain differential features 、 、 ; S530: consistency features and differential characteristics 、 、 Perform feature fusion, specifically: ; ; ; in Is a splicing operation, use Consistency characteristics of cluster robots and differential characteristics 、 、 Perform splicing at the feature level to obtain splicing features , and They are maximum pooling and average pooling operations, which are used for feature selection and dimensionality reduction. The pooled features are then concatenated to obtain features. , is the feature fusion layer, is a three-dimensional convolutional layer, is the activation function, through Perform feature fusion to obtain multi-agent fusion features .
[0026] In one embodiment, Figure 3 As shown, S600 includes: For consistency features and differential characteristics 、 、 Perform feature reconstruction, specifically: ; ; in Is a consistent feature and differential characteristics of and, It is the feature reconstruction layer, through Reconstruct the consistent features and the different features, and 、 、 The multi-agent features of swarm robots correspond one to one.
[0027] In one embodiment, in S700, the multi-agent fusion features are input to the detection head to obtain a prediction result, including: The detection head completes the classification and regression tasks, Predictions are made through different network layers, specifically: ; ; in, Used to predict classification scores, it consists of a 1×1 two-dimensional convolution layer. The regression of the three-dimensional detection box consists of a 1×1 two-dimensional convolutional layer.
[0028] In one embodiment, the loss function in S700 is specifically: ; ; ; ; ; in, , , , , is the prediction box classification loss, is the 3D detection box regression loss, is the cross-view consistency-difference feature reconstruction loss, is the cross-view consistency distance loss, For the total loss.
[0029] In one embodiment, in S700, the validation set and the test set are respectively input into the trained point cloud collaborative 3D detection model to perform calculations to obtain 3D detection results, verify the validity and generalization of the model, and perform platform deployment, including: After completing the training batches, the trained 3D detection model is evaluated. The validation set is used to validate the model and select the best cluster robot 3D detection model. , the test set Input 3D detection model , get the test results ; 3D detection model It is deployed on the NVIDIA Orin edge computing platform of cluster robots, first converting it into ONNX intermediate data, and then converting the model into TensorRT for model acceleration, enabling cluster robots to collaboratively detect targets in a large range of scenes.
[0030] The aforementioned cross-view collaborative perception method for aircraft manufacturing swarm robots uses the NVIDIA Orin edge computing platform to acquire large-scale scene point cloud information using Livox non-repetitive scanning lidar. The swarm robots then construct a collaborative 3D point cloud detection model based on cross-view consistent spatial mapping. This model comprises a multi-agent point cloud feature extraction module, a multi-agent feature fusion module based on cross-view consistent spatial mapping, a cross-view consistent-differential feature reconstruction module, and a detection head module. The multi-agent feature fusion module, based on cross-view consistent spatial mapping, maps the information acquired by multiple agents from different viewpoints into a consistent feature space and combines the differential representations between these viewpoints, achieving deep fusion of multi-source perception information. In this module, multiple agents share a consistent mapping structure, ensuring feature consistency while preserving key structural information. Simultaneously, each agent performs differential mapping to highlight local feature differences from its own viewpoint. This collaborative modeling approach of consistency and difference not only enhances the robustness and completeness of feature representation but also improves the system's adaptability to complex spatial structures. The fusion process employs feature concatenation, pooling compression, and convolutional fusion to further enhance the hierarchical expression and semantic abstraction capabilities of features. The resulting fused features more comprehensively reflect the multi-perspective, multi-agent perception of environmental information. In aircraft manufacturing, a typical large-scale and complex industrial scenario, this module effectively supports the collaborative perception and intelligent decision-making of swarm robots for large-scale components, long-distance structures, and multi-process areas. This improves global perception and operational precision in tasks such as inspection, docking, and assembly, ensuring operational efficiency and safety.
[0031] In one embodiment, a cross-view collaborative perception system of swarm robots for aircraft manufacturing is also provided, including a 3D point cloud data acquisition module, a data set production module, a 3D detection model building module, a feature extraction module, a feature fusion module, a feature reconstruction module, and a training and deployment module; The 3D point cloud data acquisition module is used to build a collaborative 3D perception platform for swarm robots. Each robot uses a lidar to scan and acquire 3D point cloud data in a large-scale aircraft manufacturing scene, and the real-time position and posture of each robot is obtained through a motion capture system. The dataset creation module is used to label and segment the 3D point cloud data scanned by the swarm robot, completing the creation of collaborative point cloud detection datasets for large-scale scenarios in aircraft manufacturing. 3D detection model building module, used to build a point cloud collaborative 3D detection model based on cross-view consistency spatial mapping, including point cloud feature extraction module, feature fusion module, feature reconstruction module and detection head; The feature extraction module is used to pre-process the point cloud data in the training set and then perform feature encoding to obtain multi-agent features; The feature fusion module is used to map the cross-view consistency feature and the cross-view difference feature to obtain the consistency feature and difference feature respectively, and fuse the consistency feature and the difference feature to obtain the multi-agent fusion feature; The feature reconstruction module is used to reconstruct the multi-agent fusion features and obtain the reconstructed features; The training and deployment module is used to input the multi-agent fusion features into the detection head to obtain the prediction results. Based on the consistency features, difference features, reconstructed features and prediction results, the parameters of the point cloud collaborative 3D detection model are updated through the set loss function until the point cloud collaborative 3D detection model converges; the validation set and test set are respectively input into the trained point cloud collaborative 3D detection model for calculation to obtain the 3D detection results, verify the effectiveness and generalization of the model, and then deploy it on the platform.
[0032] For the specific definition of the cross-perspective collaborative perception system of cluster robots for aircraft manufacturing, please refer to the definition of the cross-perspective collaborative perception method of cluster robots for aircraft manufacturing above, which will not be repeated here. Each module in the above-mentioned cross-perspective collaborative perception system of cluster robots for aircraft manufacturing can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0033] In typical large-scale industrial scenarios, such as aircraft manufacturing, the manufacturing and assembly processes place extremely high demands on efficiency and precision. As manufacturing systems rapidly evolve towards intelligent and unmanned operations, swarm robots are becoming a key enabler for the coordinated execution of complex tasks. Their 3D object detection capabilities directly impact the system's performance in precision manufacturing and high-precision assembly. To adapt to operating environments characterized by large aircraft structures and widely distributed components, a swarm robot 3D detection model with efficient response and accurate recognition capabilities is needed. By deploying sensing equipment such as lidar across multiple mobile robots, the system can achieve multi-perspective coverage of large-scale spatial environments and collect high-density point clouds, providing comprehensive spatial structure and position information. However, current mainstream point cloud 3D detection methods often suffer from limited detection range and insufficient recognition capabilities for complex or uniquely shaped objects. This is particularly true for critical components, such as flat, narrow, and fuzzy aircraft skins, which often experience reduced recognition accuracy and amplified positioning errors. This capability bottleneck is particularly pronounced in large-scale operations, hindering the swarm robot's comprehensive perception and precise control of the manufacturing site. Therefore, it is necessary to develop a multi-agent 3D detection mechanism for large-scale aircraft manufacturing environments, improve the system's adaptability and stability in the process of multi-scale and multi-structure target recognition, and thus support swarm robots to achieve efficient collaborative operations and precise process execution in complex industrial environments. Therefore, the innovations of this invention are as follows: (1) This invention proposes a cross-viewpoint consistent collaborative perception method and system for swarm robots in aircraft manufacturing. This method can perform high-precision three-dimensional detection of large aircraft components, workers, and operating robots during the aircraft manufacturing process, and provide swarm robots with position and spatial information of targets in a wide range of manufacturing scenarios. (2) Swarm robots face the problem of perspective perception differences in large-scale aircraft manufacturing scenarios. This paper proposes a point cloud 3D detection model based on cross-perspective consistency spatial mapping. The multi-agent feature fusion module based on cross-perspective consistency spatial mapping maps the perception information obtained by different agents from their respective perspectives into a unified feature space, and combines their respective differential feature expressions to construct a comprehensive representation of fusion consistency and diversity. This collaborative fusion mechanism not only ensures the spatial consistency of multi-agent perception information, but also retains the differences in local details from different perspectives, thereby improving the system's adaptability to complex structural environments and expression integrity, making multi-agent collaborative perception more efficient and reliable.
[0034] (3) The present invention is not only applicable to three-dimensional detection in aircraft manufacturing scenarios, but also has the potential value of being applied to three-dimensional detection in any large-scale industrial manufacturing scenarios. Its three-dimensional detection performance is characterized by high efficiency, high precision, and cluster robot collaboration, promoting the high-quality development of intelligent manufacturing.
[0035] The above is a detailed introduction to the cross-perspective collaborative perception method and system of cluster robots for aircraft manufacturing provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the core idea of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, the present invention can also be improved and modified, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A cross-view collaborative perception method for swarm robots in aircraft manufacturing, characterized by: The method comprises the following steps: S100: Build a collaborative 3D perception platform for swarm robots. Each robot uses LiDAR to scan and acquire 3D point cloud data from a wide range of aircraft manufacturing scenes. The motion capture system then acquires each robot's real-time position and posture. S200: Label and segment the 3D point cloud data scanned by the swarm robot, and complete the production of collaborative point cloud detection datasets for large-scale scenarios in aircraft manufacturing. S300: Builds a point cloud collaborative 3D detection model based on cross-view consistency spatial mapping, including a point cloud feature extraction module, a feature fusion module, a feature reconstruction module, and a detection head; S400: pre-processing the point cloud data in the training set and inputting it into the point cloud feature extraction module for feature encoding to obtain multi-agent features; S500: Input the multi-agent features into the feature fusion module for cross-view consistency feature mapping and cross-view difference feature mapping to obtain consistency features and difference features respectively, and perform feature fusion on the consistency features and difference features to obtain multi-agent fusion features; S600: Inputting the multi-agent fusion features into the feature reconstruction module to perform feature reconstruction to obtain reconstructed features; S700: Input the multi-agent fusion features into the detection head to obtain the prediction results. Based on the consistency features, difference features, reconstructed features and prediction results, the parameters of the point cloud collaborative 3D detection model are updated through the set loss function until the point cloud collaborative 3D detection model converges; the validation set and test set are respectively input into the trained point cloud collaborative 3D detection model for calculation to obtain the 3D detection results, verify the effectiveness and generalization of the model, and then deploy it on the platform.
2. The method according to claim 1, characterized in that S100 includes: S110: Assemble and debug the mechanical and electrical structures of the cluster robots. The robots are numbered as follows: 、 、 A Livox non-repeating scanning LiDAR is installed in the slot directly in front of the swarm robot, with a reserved LiDAR installation location. The Livox LiDAR is connected to the Nvidia Orin edge computing device inside the swarm robot to perform laser scanning, data transmission, and model deployment. The Livox LiDAR follows the swarm robot's movements. S120: Using swarm robots 、 、 The Livox laser radar carried out point cloud data collection for a wide range of aircraft manufacturing scenes, and obtained point cloud data respectively. 、 、 ,The real-time pose of each robot is obtained through the motion capture system.
3. The method according to claim 2, characterized in that S200 includes: S210: Use the point cloud annotation software LabelCloud to label the point cloud data: 、 、 Perform coordinate system transformation, and the real-time position of the robot obtained by the motion capture system will be 、 、 Mapped to the unified world coordinate system, forming a 、 、 Global collaborative point cloud data of point cloud information , for point cloud data Perform three-dimensional detection frame and category annotation, and the labeled data set is recorded as ; S220: Labeling the dataset Divide and label the dataset Randomly take out 65% of the data for training, as the training set, denoted as , the labeled dataset Randomly take out the equivalent of 15% of the data is used for verification, as the verification set, denoted as , the labeled dataset The rest of is used as the test set, which is , used to test the generalization of the model.
4. The method according to claim 3, characterized in that S400 includes: S410: Point cloud data acquired by each robot 、 、 With the laser radar as the center of the coordinate axis, the robot's forward direction is the x-axis direction, the robot's left direction is the y-axis direction, and the robot's upward direction is the z-axis direction. The point cloud is mathematically represented as follows: A vector representation of size, the content is ,in are the three-dimensional coordinates of the point cloud data, is the reflection intensity of point cloud data, is the number of point clouds; S420: Select point cloud data 、 、 The coordinate range of the point cloud on the coordinate axis is specified as , within this range, the point cloud is divided into columns, and the specifications of each point cloud column are , take the center of each point cloud column as the center point coordinate, and then calculate the relative coordinates of each point and the center of the point cloud column , and use it as a supplementary representation of the point cloud, specifically A vector representation of size, the content is ; S430: Use three PointPillars point cloud feature extraction networks with the same structure to extract 、 、 Point cloud data is feature encoded to obtain multi-agent features 、 、 .
5. The method according to claim 4, characterized in that S500 includes: S510: Yes 、 、 Perform cross-view consistency feature mapping, specifically: ; ; ; in is a cross-view consistency spatial mapping layer, is the fully connected layer, is the activation function, is layer normalization, 、 、 Share a cross-view consistency spatial mapping layer, through Will 、 、 Mapped to the consistent feature space, the output is ; S520: Yes 、 、 Perform cross-view difference feature mapping, specifically: ; ; ; in 、 、 It is a spatial mapping layer that maps differences across perspectives. is the fully connected layer, is the activation function, is layer normalization, 、 、 Using different cross-view difference spatial mapping layers, 、 、 Will 、 、 Map to different feature spaces to obtain differential features 、 、 ; S530: consistency features and differential characteristics 、 、 Perform feature fusion, specifically: ; ; ; in Is a splicing operation, use Consistency characteristics of cluster robots and differential characteristics 、 、 Perform splicing at the feature level to obtain splicing features , and They are maximum pooling and average pooling operations, which are used for feature selection and dimensionality reduction. The pooled features are then concatenated to obtain features. , is the feature fusion layer, is a three-dimensional convolutional layer, is the activation function, through Perform feature fusion to obtain multi-agent fusion features .
6. The method according to claim 1, wherein: S600 includes: For consistency features and differential characteristics 、 、 Perform feature reconstruction, specifically: ; ; in Is a consistent feature and differential characteristics of and, It is the feature reconstruction layer, through Reconstruct the consistent features and the different features, and 、 、 The multi-agent features of swarm robots correspond one to one.
7. The method according to claim 6, characterized in that In S700, the multi-agent fusion features are input to the detection head to obtain the prediction results, including: The detection head completes the classification and regression tasks, Predictions are made through different network layers, specifically: ; ; in, Used to predict classification scores, it consists of a 1×1 two-dimensional convolution layer. The regression of the three-dimensional detection box consists of a 1×1 two-dimensional convolutional layer.
8. The method according to claim 7, characterized in that The S700 loss function is specifically: ; ; ; ; ; in, , , , , is the prediction box classification loss, is the 3D detection box regression loss, is the cross-view consistency-difference feature reconstruction loss, is the cross-view consistency distance loss, For the total loss.
9. The method according to claim 8, characterized in that In S700, the validation and test sets are fed into the trained point cloud collaborative 3D detection model to calculate and obtain 3D detection results. This verifies the model's effectiveness and generalization, and then deploys the platform, including: After completing the training batches, the trained 3D detection model is evaluated. The validation set is used to validate the model and select the best cluster robot 3D detection model. , the test set Input 3D detection model , get the test results ; 3D detection model It is deployed on the NVIDIA Orin edge computing platform of cluster robots, first converting it into ONNX intermediate data, and then converting the model into TensorRT for model acceleration, enabling cluster robots to collaboratively detect targets in a large range of scenes.
10. A cross-perspective collaborative perception system of swarm robots for aircraft manufacturing, characterized by: It includes 3D point cloud data acquisition module, data set production module, 3D detection model building module, feature extraction module, feature fusion module, feature reconstruction module, and training and deployment module; The 3D point cloud data acquisition module is used to build a collaborative 3D perception platform for swarm robots. Each robot uses a lidar to scan and acquire 3D point cloud data in a large-scale aircraft manufacturing scene, and the real-time position and posture of each robot is obtained through a motion capture system. The dataset creation module is used to label and segment the 3D point cloud data scanned by the swarm robot, completing the creation of collaborative point cloud detection datasets for large-scale scenarios in aircraft manufacturing. 3D detection model building module, used to build a point cloud collaborative 3D detection model based on cross-view consistency spatial mapping, including point cloud feature extraction module, feature fusion module, feature reconstruction module and detection head; The feature extraction module is used to pre-process the point cloud data in the training set and then perform feature encoding to obtain multi-agent features; The feature fusion module is used to map the cross-view consistency feature and the cross-view difference feature to obtain the consistency feature and difference feature respectively, and fuse the consistency feature and the difference feature to obtain the multi-agent fusion feature; The feature reconstruction module is used to reconstruct the multi-agent fusion features and obtain the reconstructed features; The training and deployment module is used to input the multi-agent fusion features into the detection head to obtain the prediction results. Based on the consistency features, difference features, reconstructed features and prediction results, the parameters of the point cloud collaborative 3D detection model are updated through the set loss function until the point cloud collaborative 3D detection model converges; the validation set and test set are respectively input into the trained point cloud collaborative 3D detection model for calculation to obtain the 3D detection results, verify the effectiveness and generalization of the model, and then deploy it on the platform.
Citation Information
Patent Citations
Multi-robot co-localization and fusion mapping method under multi-view in open space
CN109579843A
Three-dimensional target detection method and system based on single-line laser radar and monocular camera
CN118068356A
Unified three-dimensional target detection method for large-scale assembly process full scene
CN118864827A
Multi-robot collaborative three-dimensional target detection method based on visual state space model
CN119169606A
Multi-sensor online three-dimensional detection method and device for large assembly scene
CN119291714A
Cited By
Double-robot collaborative three-dimensional target identification method and system based on mutual information estimation feature unwrapping
CN121267936A