Cluster robot cross-view cooperative perception method and system for aircraft manufacturing
Through the cross-view collaborative perception method of cluster robots, a point cloud detection model based on cross-view consistency spatial mapping was constructed, which solved the problems of limited field of view and detection range in aircraft manufacturing, achieved high-precision three-dimensional target detection and collaborative perception, and improved the efficiency and safety of the manufacturing process.
Patent Information
- Application Number
- CN202510990804.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-07-18
AI Technical Summary
In the aircraft manufacturing process, the perception field of view and detection range of a single robot are limited, and the differences in perspective between different robots lead to inconsistent perception data, making it difficult to achieve high-precision three-dimensional target detection in all scenarios.
By adopting the cross-view collaborative perception method of swarm robots and building a collaborative 3D perception platform, using Livox lidar and NVIDIA Orin edge computing devices, a point cloud collaborative 3D detection model based on cross-view consistent spatial mapping is constructed. Feature extraction, fusion and reconstruction are performed to achieve unified expression of multi-agent features and differentiated preservation.
It improves the adaptability and detection consistency of swarm robots in complex structural environments, provides high-precision three-dimensional target detection results, and supports efficient collaborative operations and precise control in a wide range of manufacturing scenarios.
Smart Images

Figure CN120495643B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of large-scale intelligent manufacturing scene cluster robot collaborative three-dimensional perception, in particular to a cluster robot cross-view collaborative perception method and system for aircraft manufacturing. BACKGROUND
[0002] The manufacturing and assembly of aviation equipment is the core link of economic development, especially in the development process of domestic large aircraft, the manufacturing and assembly quality directly determines the comprehensive performance and flight safety of the aircraft. In the assembly process of typical large-size thin-walled structures such as aircraft skin, wing, fuselage, etc., systematic inspection needs to be carried out on the key areas. Specifically, it includes comprehensive inspection of the initial state of large components such as skin, including stress distribution, structural deformation, attitude inclination and ground support state, etc., to identify potential structural abnormalities or safety risks in advance, and provide reliable data support and protection for subsequent high-precision assembly operations.
[0003] At present, in the assembly inspection process of large-size aircraft complex components, it still mainly relies on manual operation to identify key structural features, state evaluation and abnormality checking. This way not only has low efficiency and limited precision, but also has problems such as large subjective error and heavy operation burden, which is difficult to meet the needs of high reliability and high consistency inspection. With the continuous progress of robot technology and artificial intelligence, it has become an important part of intelligent manufacturing to promote the unmanned and intelligent upgrading of aircraft manufacturing inspection, and it is also a key measure to ensure the quality of aircraft assembly and flight safety. Therefore, the research on inspection robot perception technology for aircraft manufacturing scene is the trend of high-quality development of major equipment.
[0004] When the robot inspects the aircraft manufacturing process, it is crucial to achieve high-precision three-dimensional target monitoring of the robot in a large-scale manufacturing scene. Three-dimensional target detection can provide position information and size information of multiple targets such as large aircraft components, workers, and operating robots in the aircraft manufacturing scene. On the basis of high-precision three-dimensional target detection, the robot can complete intelligent and accurate safety inspection in the aircraft manufacturing process. In recent years, three-dimensional target detection methods based on deep learning have gradually attracted attention, but they still cannot meet the engineering requirements of high-precision three-dimensional target detection for robots in the aircraft manufacturing scene. The following difficulties are faced in the implementation process:
[0005] 1. The aircraft manufacturing scene has the characteristics of wide coverage, and the single robot's perception field and detection range are limited, which cannot cover the entire aircraft manufacturing and assembly scene for full-scene perception. Therefore, cluster robots need to be used for collaborative perception to cover a larger area and wider field of view, so as to detect targets in the environment more comprehensively.
[0006] 2. Due to different robot distribution positions, the obtained laser radar point cloud has significant viewing angle difference, resulting in spatial deviation and feature inconsistency of the same target under different viewing angles, affecting information fusion and perception consistency. In order to realize effective cooperation of the cluster perception system, it is urgent to solve the difference alignment and fusion problem between multi-view point cloud data, and improve the consistency and robustness of the perception result. SUMMARY
[0007] In view of the above technical problems, the application provides a cluster robot cross-view cooperative perception method and system for aircraft manufacturing, which aims to optimize the technical problems such as limited perception field and detection range, cross-view perception inconsistency in the process of three-dimensional target detection of robots in large-scale scenes of aircraft manufacturing.
[0008] The technical solution adopted by the application to solve its technical problems is:
[0009] The cluster robot cross-view cooperative perception method for aircraft manufacturing comprises the following steps:
[0010] S100: Build a cluster robot cooperative three-dimensional perception platform, each robot uses a laser radar to scan and obtain three-dimensional point cloud data in a large-scale scene of aircraft manufacturing, and obtains the real-time pose of each robot through a motion capture system;
[0011] S200: Data labeling and division are performed on the three-dimensional point cloud data scanned by the cluster robots, and a cooperative point cloud detection data set for the large-scale scene of aircraft manufacturing is completed;
[0012] S300: Construct a point cloud cooperative three-dimensional detection model based on cross-view consistency space mapping, including a point cloud feature extraction module, a feature fusion module, a feature reconstruction module and a detection head;
[0013] S400: The point cloud data in the training set is preprocessed and input into the point cloud feature extraction module for feature coding to obtain multi-agent features;
[0014] S500: The multi-agent features are input into the feature fusion module for cross-view consistency feature mapping and cross-view difference feature mapping, respectively, to obtain consistency features and difference features, and the consistency features and the difference features are fused to obtain multi-agent fusion features;
[0015] S600: The multi-agent fusion features are input into the feature reconstruction module for feature reconstruction to obtain reconstructed features;
[0016] S700: Input the multi-agent fusion features into the detection head to obtain the prediction results. Based on the consistency features, difference features, reconstructed features and prediction results, the parameters of the point cloud collaborative 3D detection model are updated through the set loss function until the point cloud collaborative 3D detection model converges; the validation set and test set are respectively input into the trained point cloud collaborative 3D detection model for calculation to obtain the 3D detection results, verify the effectiveness and generalization of the model, and then deploy it on the platform.
[0017] Preferably, S100 includes:
[0018] S110: Assemble and debug the mechanical and electrical structures of the cluster robots. The robots are numbered as follows: 、 、 A Livox non-repeating scanning LiDAR is installed in the slot directly in front of the swarm robot, with a reserved LiDAR installation location. The Livox LiDAR is connected to the Nvidia Orin edge computing device inside the swarm robot to perform laser scanning, data transmission, and model deployment. The Livox LiDAR follows the swarm robot's movements.
[0019] S120: Using swarm robots 、 、 The Livox laser radar carried out point cloud data collection for a wide range of aircraft manufacturing scenes, and obtained point cloud data respectively. 、 、 ,The real-time pose of each robot is obtained through the motion capture system.
[0020] Preferably, S200 includes:
[0021] S210: Use the point cloud annotation software LabelCloud to label the point cloud data: 、 、 Perform coordinate system transformation, and the real-time position of the robot obtained by the motion capture system will be 、 、 Mapped to the unified world coordinate system, forming a 、 、 Global collaborative point cloud data of point cloud information , for point cloud data Perform three-dimensional detection frame and category annotation, and the labeled data set is recorded as ;
[0022] S220: Labeling the dataset Divide and label the dataset Randomly take out 65% of the data for training, as the training set, denoted as , the labeled dataset Randomly take out the equivalent of 15% of the data is used for verification, as the verification set, denoted as , the labeled dataset The rest of is used as the test set, which is , used to test the generalization of the model.
[0023] Preferably, S400 includes:
[0024] S410: Point cloud data acquired by each robot 、 、 With the laser radar as the center of the coordinate axis, the robot's forward direction is the x-axis direction, the robot's left direction is the y-axis direction, and the robot's upward direction is the z-axis direction. The point cloud is mathematically represented as follows: The vector representation of size is ,in are the three-dimensional coordinates of the point cloud data, is the reflection intensity of point cloud data, is the number of point clouds;
[0025] S420: Select point cloud data 、 、 The coordinate range of the point cloud on the coordinate axis is specified as , within this range, the point cloud is divided into columns, and the specifications of each point cloud column are , take the center of each point cloud column as the center point coordinate, and then calculate the relative coordinates of each point and the center of the point cloud column , and use it as a supplementary representation of the point cloud, specifically The vector representation of size is ;
[0026] S430: Use three PointPillars point cloud feature extraction networks with the same structure to extract 、 、 Point cloud data is feature encoded to obtain multi-agent features 、 、 .
[0027] Preferably, S500 includes:
[0028] S510: mapping cross-view consistency features, specifically: , , ;
[0029] ;
[0030] ;
[0031] ;
[0032] wherein is a cross-view consistency spatial mapping layer, is a fully connected layer, is an activation function, is layer normalization, , , share a cross-view consistency spatial mapping layer, and maps , , to a consistent feature space, and the output is ;
[0033] S520: mapping cross-view difference features, specifically: , , ;
[0034] ;
[0035] ;
[0036] ;
[0037] wherein , , are cross-view difference spatial mapping layers, is a fully connected layer, is an activation function, is layer normalization, , , use different cross-view difference spatial mapping layers, and , , map , , to different feature spaces to obtain difference features , , ;
[0038] S530: consistency features and differential characteristics 、 、 Perform feature fusion, specifically:
[0039] ;
[0040] ;
[0041] ;
[0042] in Is a splicing operation, use Consistency characteristics of cluster robots and differential characteristics 、 、 Perform splicing at the feature level to obtain splicing features , and They are maximum pooling and average pooling operations, which are used for feature selection and dimensionality reduction. The pooled features are then concatenated to obtain features. , is the feature fusion layer, is a three-dimensional convolutional layer, is the activation function, through Perform feature fusion to obtain multi-agent fusion features .
[0043] Preferably, S600 includes:
[0044] Consistency Features and differential characteristics 、 、 Perform feature reconstruction, specifically:
[0045] ;
[0046] ;
[0047] in Is a consistent feature and differential characteristics of and, is the feature reconstruction layer, through Reconstruct the consistent features and the different features, and 、 、 The multi-agent features of swarm robots correspond one to one.
[0048] Preferably, in S700, the multi-agent fusion features are input to the detection head to obtain the prediction results, including:
[0049] The detection head completes the classification and regression tasks, Predictions are made through different network layers, specifically:
[0050] ;
[0051] ;
[0052] in, Used to predict classification scores, it consists of a 1×1 two-dimensional convolution layer. The regression of the three-dimensional detection box consists of a 1×1 two-dimensional convolutional layer.
[0053] Preferably, the loss function in S700 is specifically:
[0054] ;
[0055] ;
[0056] ;
[0057] ;
[0058] ;
[0059] in, , , , , is the prediction box classification loss, is the 3D detection box regression loss, is the cross-view consistency-difference feature reconstruction loss, is the cross-view consistency distance loss, For the total loss.
[0060] Preferably, in S700, the validation set and the test set are respectively input into the trained point cloud collaborative 3D detection model to perform calculations to obtain 3D detection results, verify the validity and generalization of the model, and perform platform deployment, including:
[0061] After completing the training batches, the trained 3D detection model is evaluated. The validation set is used to validate the model and select the best cluster robot 3D detection model. , the test set Input 3D detection model , get the test results ;
[0062] 3D detection model Deployed on the NVIDIA Orin edge computing platform of cluster robots, it is first converted into ONNX intermediate data, and then the model is converted into TensorRT for model acceleration, enabling cluster robots to collaboratively detect targets in a large range of scenes in three dimensions.
[0063] A swarm robot cross-view collaborative perception system for aircraft manufacturing, including a 3D point cloud data acquisition module, a data set creation module, a 3D detection model building module, a feature extraction module, a feature fusion module, a feature reconstruction module, and a training and deployment module;
[0064] The 3D point cloud data acquisition module is used to build a collaborative 3D perception platform for swarm robots. Each robot uses a lidar to scan and acquire 3D point cloud data in a large-scale aircraft manufacturing scene, and the motion capture system obtains the real-time position and posture of each robot.
[0065] The dataset creation module is used to label and segment the 3D point cloud data scanned by the swarm robot, completing the creation of collaborative point cloud detection datasets for large-scale scenarios in aircraft manufacturing.
[0066] 3D detection model building module, used to build a point cloud collaborative 3D detection model based on cross-view consistency spatial mapping, including point cloud feature extraction module, feature fusion module, feature reconstruction module and detection head;
[0067] The feature extraction module is used to pre-process the point cloud data in the training set and perform feature encoding to obtain multi-agent features;
[0068] The feature fusion module is used to map the cross-view consistency feature and the cross-view difference feature to obtain the consistency feature and difference feature respectively, and then fuse the consistency feature and the difference feature to obtain the multi-agent fusion feature;
[0069] The feature reconstruction module is used to reconstruct the multi-agent fusion features and obtain the reconstructed features;
[0070] The training and deployment module is used to input the multi-agent fusion features into the detection head to obtain the prediction results. Based on the consistency features, difference features, reconstructed features and prediction results, the parameters of the point cloud collaborative 3D detection model are updated through the set loss function until the point cloud collaborative 3D detection model converges; the validation set and test set are respectively input into the trained point cloud collaborative 3D detection model for calculation to obtain the 3D detection results, verify the effectiveness and generalization of the model, and then deploy it on the platform.
[0071] The above-mentioned cross-perspective collaborative perception method and system of swarm robots for aircraft manufacturing can perform high-precision three-dimensional detection of large aircraft components, workers, operating robots, etc. in the aircraft manufacturing process, and provide swarm robots with the position and spatial information of targets in large-scale manufacturing scenarios; in view of the problem of perspective perception differences faced by swarm robots in large-scale aircraft manufacturing scenarios, a point cloud three-dimensional detection model based on cross-perspective consistent spatial mapping is proposed. This collaborative fusion mechanism not only ensures the spatial consistency of multi-agent perception information, but also retains the differences in local details under different perspectives, thereby improving the system's adaptability and expression integrity to complex structural environments, making multi-agent collaborative perception more efficient and reliable; and it has the potential value of being applied to three-dimensional detection in any large-scale industrial manufacturing scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 This is a flow chart of a cross-view collaborative perception method of swarm robots for aircraft manufacturing in one embodiment of the present invention;
[0073] Figure 2 Schematic diagram of a collaborative three-dimensional perception platform of swarm robots in one embodiment of the present invention;
[0074] Figure 3 Schematic diagram of a point cloud collaborative 3D detection model based on cross-view consistency spatial mapping in one embodiment of the present invention. DETAILED DESCRIPTION
[0075] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is further described in detail below with reference to the accompanying drawings.
[0076] In one embodiment, Figure 1 As shown, a cross-view collaborative perception method of swarm robots for aircraft manufacturing includes the following steps:
[0077] S100: Build a collaborative 3D perception platform for swarm robots. Each robot uses LiDAR to scan and acquire 3D point cloud data from a wide range of aircraft manufacturing scenes. The motion capture system then acquires each robot's real-time position and posture.
[0078] S200: Label and segment the 3D point cloud data scanned by the swarm robot, and complete the production of collaborative point cloud detection datasets for large-scale scenarios in aircraft manufacturing.
[0079] S300: Builds a point cloud collaborative 3D detection model based on cross-view consistency spatial mapping, including a point cloud feature extraction module, a feature fusion module, a feature reconstruction module, and a detection head;
[0080] S400: input the point cloud data in the training set after preprocessing into the point cloud feature extraction module for feature coding to obtain multi-agent features;
[0081] S500: input the multi-agent features into the feature fusion module for cross-view consistency feature mapping and cross-view difference feature mapping to respectively obtain consistency features and difference features, and perform feature fusion on the consistency features and the difference features to obtain multi-agent fusion features;
[0082] S600: input the multi-agent fusion features into the feature reconstruction module for feature reconstruction to obtain reconstructed features;
[0083] S700: input the multi-agent fusion features into the detection head to obtain a prediction result, and based on the consistency features, the difference features, the reconstructed features and the prediction result, update parameters of the point cloud collaborative three-dimensional detection model through a set loss function until the point cloud collaborative three-dimensional detection model converges; input the verification set and the test set into the trained point cloud collaborative three-dimensional detection model for operation to obtain a three-dimensional detection result, verify the effectiveness and generalization of the model, and deploy the platform.
[0084] The method can realize cluster robot collaborative three-dimensional detection in a wide range of aircraft manufacturing scenes, can quickly provide position and spatial information of aircraft large components, workers and working robots for the cluster robot, can enable the cluster robot to systematically patrol key areas of the aircraft manufacturing scene, can help the cluster robot to avoid obstacles, plan and control in a wide range of manufacturing scenes, can make the cluster robot better complete work and assembly tasks, can improve the assembly efficiency of the cluster robot, can increase the work safety and stability of the cluster robot, and can promote the intelligent high-quality high-speed development of domestic large aircraft manufacturing.
[0085] In one embodiment, as shown in Figure 2 S100 includes:
[0086] S110: assemble and debug the mechanical and electrical structure of the cluster robot, and the robot numbers are , , , reserve a laser radar installation position in front of the cluster robot, install a Livox non-repeating scanning laser radar in the reserved position in front of the cluster robot, connect the Livox laser radar with the Nvidia Orin edge computing device inside the cluster robot, so as to complete laser scanning, data transmission and model deployment, wherein the Livox laser radar moves with the cluster robot;
[0087] S120: use the cluster robot , , The Livox laser radar carried out point cloud data collection for a wide range of aircraft manufacturing scenes, and obtained point cloud data respectively. 、 、 ,The real-time pose of each robot is obtained through the motion capture system.
[0088] In one embodiment, S200 includes:
[0089] S210: Use the point cloud annotation software LabelCloud to label the point cloud data: 、 、 Perform coordinate system transformation, and the real-time position of the robot obtained by the motion capture system will be 、 、 Mapped to the unified world coordinate system, forming a 、 、 Global collaborative point cloud data of point cloud information , for point cloud data Perform three-dimensional detection frame and category annotation, and the labeled data set is recorded as ;
[0090] S220: Labeling the dataset Divide and label the dataset Randomly take out 65% of the data for training, as the training set, denoted as , the labeled dataset Randomly take out the equivalent of 15% of the data is used for verification, as the verification set, denoted as , the labeled dataset The rest of is used as the test set, which is , used to test the generalization of the model.
[0091] In one embodiment, Figure 3 As shown, S400 includes:
[0092] S410: Point cloud data acquired by each robot 、 、 With the laser radar as the center of the coordinate axis, the robot's forward direction is the x-axis direction, the robot's left direction is the y-axis direction, and the robot's upward direction is the z-axis direction. The point cloud is mathematically represented as follows: The vector representation of size is ,in are the three-dimensional coordinates of the point cloud data, For point cloud data reflection intensity, For point cloud quantity;
[0093] S420: Select point cloud data , , Coordinate range, define the range of point cloud on coordinate axis as , the point cloud is divided into columns in this range, and the specification of each point cloud column is , and the center of each point cloud column is taken as the center point coordinate, and then the relative coordinates of each point and the center of the point cloud column are calculated , and it is taken as a supplementary representation of the point cloud, specifically Vector representation with size ;
[0094] S430: Feature encoding is performed on , , Point cloud data using three structure-similar PointPillars point cloud feature extraction networks respectively, and multi-agent features , , are obtained respectively.
[0095] In an embodiment, as shown in Figure 3 , S500 includes:
[0096] S510: Cross-view consistency feature mapping is performed on , , , specifically:
[0097] ;
[0098] ;
[0099] ;
[0100] Wherein is a cross-view consistency spatial mapping layer, is a fully connected layer, is an activation function, is layer normalization, , , Share a cross-view consistency spatial mapping layer, and maps , , to a consistent feature space, and the output is ;
[0101] S520: mapping the cross-view difference feature, specifically: , , ,
[0102] ;
[0103] ;
[0104] ;
[0105] wherein , , is a cross-view difference spatial mapping layer, is a fully connected layer, is an activation function, is layer normalization, , , uses different cross-view difference spatial mapping layers to map , , to different feature spaces to obtain difference features , , ; , , ;
[0106] S530: feature fusion of the consistency feature and the difference feature , , , specifically:
[0107] ;
[0108] ;
[0109] ;
[0110] wherein is a concatenation operation, using to concatenate the consistency feature and the difference feature , , of the swarm robot on the feature level to obtain the concatenated feature , and are max-pooling and average-pooling operations, respectively, for feature selection and dimension reduction, and the pooled features are again concatenated to obtain the feature , is a feature fusion layer, is a three-dimensional convolution layer, is an activation function, and feature fusion is performed to obtain multi-agent fusion features .
[0111] In one embodiment, as shown in Figure 3 S600 includes:
[0112] consistency features and difference features , , feature reconstruction is performed, specifically:
[0113] ;
[0114] ;
[0115] wherein is the sum of the consistency features and the difference features , is a feature reconstruction layer that reconstructs the consistency features and the difference features by , and , , corresponds to the multi-agent features of the swarm robot.
[0116] In one embodiment, the multi-agent fusion features are input to a detection head in S700 to obtain a prediction result, including:
[0117] The detection head completes classification and regression tasks, and is predicted through different network layers, specifically:
[0118] ;
[0119] ;
[0120] wherein is used to predict a classification score and is composed of a 1x1 two-dimensional convolution layer, is used to regress a three-dimensional detection box and is composed of a 1x1 two-dimensional convolution layer.
[0121] In one embodiment, the loss function in S700 is specifically:
[0122] ;
[0123] ;
[0124] ;
[0125] ;
[0126] ;
[0127] wherein, , , , , is a prediction box classification loss, is a three-dimensional detection box regression loss, is a cross-view consistency-difference feature reconstruction loss, is a cross-view consistency distance loss, is a total loss.
[0128] In one embodiment, the verification set and the test set in S700 are respectively input into the trained point cloud collaborative three-dimensional detection model to obtain a three-dimensional detection result, the effectiveness and the generalization of the model are verified, and platform deployment is performed, including:
[0129] After the number of training batches is completed, the trained three-dimensional detection model is evaluated, and the verification set is used for model verification, and the three-dimensional detection model of the cluster robot with the best effect is selected , the test set is input into the three-dimensional detection model , and a test result is obtained;
[0130] The three-dimensional detection model is deployed on the NVIDIA Orin edge computing platform of the cluster robot, is first converted into an ONNX intermediate quantity, and then the model is converted into TensorRT for model acceleration, so as to realize the collaborative three-dimensional detection of the cluster robot on targets in a large range of scenes.
[0131] The cluster robot for aircraft manufacturing based on the above-mentioned cross-view collaborative perception method of the cluster robot obtains the point cloud information of the large-scale scene of aircraft manufacturing through the Livox non-repeating scanning laser radar based on the NVIDIA Orin edge computing platform, and constructs a point cloud collaborative three-dimensional detection model based on cross-view consistency spatial mapping. The model includes a multi-agent point cloud feature extraction module, a multi-agent feature fusion module based on cross-view consistency spatial mapping, a cross-view consistency-difference feature reconstruction module, and a detection head module. The multi-agent feature fusion module based on cross-view consistency spatial mapping maps the information obtained by the multi-agent in different views to a consistent feature space, and combines the difference between the views to realize the deep fusion of multi-source perception information. In the module, the multi-agent shares a consistent mapping structure, so that the key structure information is retained while the consistency of the features is maintained; at the same time, each agent also performs difference mapping to highlight the local feature difference under the individual view. Through the collaborative modeling mode of consistency and difference, the robustness and completeness of the feature representation are enhanced, and the adaptability of the system to complex spatial structures is also improved. In the fusion process, the feature splicing, pooling compression and convolution fusion are adopted to further strengthen the hierarchical expression and semantic abstraction ability of the features, and the finally obtained fusion features can more comprehensively reflect the environment information perceived by the multi-view and multi-agent. In this typical large-scale complex industrial scene of aircraft manufacturing, the module can effectively support the collaborative perception and intelligent decision of the cluster robot on large-size components, long-distance structures and multi-process areas, improve the global perception ability and operation precision in tasks such as inspection, docking and assembly, and ensure the efficiency and safety of the operation.
[0132] In one embodiment, a cluster robot for aircraft manufacturing cross-view collaborative perception system is also provided, including a three-dimensional point cloud data acquisition module, a data set making module, a three-dimensional detection model building module, a feature extraction module, a feature fusion module, a feature reconstruction module, and a training and deployment module;
[0133] The three-dimensional point cloud data acquisition module is used to build a cluster robot collaborative three-dimensional perception platform, and each robot uses a laser radar to scan and obtain three-dimensional point cloud data in a large-scale scene of aircraft manufacturing, and a motion capture system is used to obtain the real-time pose of each robot;
[0134] The data set making module is used to make data annotations and divisions on the three-dimensional point cloud data scanned by the cluster robot, and complete the making of a collaborative point cloud detection data set for a large-scale scene of aircraft manufacturing;
[0135] The three-dimensional detection model building module is used to construct a point cloud collaborative three-dimensional detection model based on cross-view consistency spatial mapping, including a point cloud feature extraction module, a feature fusion module, a feature reconstruction module, and a detection head;
[0136] a feature extraction module, configured to perform feature coding on the point cloud data in the training set after preprocessing, to obtain multi-agent features;
[0137] a feature fusion module, configured to respectively obtain consistency features and difference features by performing feature mapping on the cross-view consistency features and the cross-view difference features, and perform feature fusion on the consistency features and the difference features to obtain multi-agent fusion features;
[0138] a feature reconstruction module, configured to perform feature reconstruction on the multi-agent fusion features to obtain reconstructed features;
[0139] a training and deployment module, configured to input the multi-agent fusion features into a detection head to obtain a prediction result, update parameters of the point cloud collaborative three-dimensional detection model by using a loss function set based on the consistency features, the difference features, the reconstructed features and the prediction result, until the point cloud collaborative three-dimensional detection model converges, input the verification set and the test set into the trained point cloud collaborative three-dimensional detection model to obtain three-dimensional detection results, verify effectiveness and generalization of the model, and deploy the platform.
[0140] Specific limitations of the cluster robot cross-view collaborative perception system for aircraft manufacturing can be seen in the limitations of the cluster robot cross-view collaborative perception method for aircraft manufacturing in the foregoing, which will not be repeated here. Each module in the cluster robot cross-view collaborative perception system for aircraft manufacturing can be realized by software, hardware and combinations thereof, in whole or in part. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to each module.
[0141] In typical large-scale industrial scenarios such as aircraft manufacturing, the manufacturing and assembly links put forward high requirements for the efficiency and accuracy of the work. With the acceleration of the manufacturing system towards intelligent and unmanned direction, cluster robots become the key force to support the collaborative execution of complex tasks, and their three-dimensional target detection capability directly affects the performance of the system in fine manufacturing and high-precision assembly. In order to adapt to the work environment with huge body structure and widely distributed components, it is necessary to build a cluster robot three-dimensional detection model with high response capability and accurate identification capability. By deploying laser radar and other sensing devices in multiple mobile robots, the system can realize multi-view coverage and high-density point cloud collection of large-scale space environment, thereby providing complete spatial structure and location information support. However, the current mainstream point cloud three-dimensional detection method has the problems of limited detection distance, insufficient recognition ability for complex or special structure targets, especially when facing key components such as aircraft skin, which is flat and narrow with fuzzy edges, often resulting in decreased recognition accuracy and enlarged positioning error. This capability bottleneck is particularly evident in large-scale operations, affecting the overall perception and precise control of cluster robots in manufacturing sites. Therefore, it is necessary to develop a multi-agent three-dimensional detection mechanism for large-scale aircraft manufacturing environment to improve the adaptability and stability of the system in multi-scale and multi-structure target recognition process, thereby supporting cluster robots to achieve efficient collaborative work and precise process execution in complex industrial environments. Therefore, the innovation points of the present application are as follows:
[0142] (1) The present application proposes a cluster robot cross-view consistency collaborative perception method and system for aircraft manufacturing, which can perform high-precision three-dimensional detection on aircraft large components, workers, and work robots in the aircraft manufacturing process, and provide position and spatial information of targets in large-scale manufacturing scenarios for cluster robots;
[0143] (2) Cluster robots face the problem of visual perception difference in large-scale aircraft manufacturing scenarios. The present application proposes a point cloud three-dimensional detection model based on cross-view consistency spatial mapping, wherein the multi-agent feature fusion module based on cross-view consistency spatial mapping maps the perception information obtained by different agents in their respective views to a unified feature space, combines their different feature expressions, and constructs a comprehensive representation that integrates consistency and diversity. This collaborative fusion mechanism not only ensures the spatial consistency of multi-agent perception information, but also retains the local detail differences under different views, thereby improving the adaptability and completeness of the system to complex structure environment, making the multi-agent collaborative perception more efficient and reliable.
[0144] (3) The present application is not only suitable for three-dimensional detection in aircraft manufacturing scenarios, but also has potential value for three-dimensional detection in any large-scale industrial manufacturing scenario. Its three-dimensional detection performance has the characteristics of high efficiency, high precision, and cluster robot collaboration, promoting the high-quality development of intelligent manufacturing.
[0145] The above is a detailed introduction to the cross-perspective collaborative perception method and system of cluster robots for aircraft manufacturing provided by the present invention. This article uses specific examples to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the core idea of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, the present invention can also be improved and modified, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. A cross-view collaborative perception method for swarm robots in aircraft manufacturing, characterized by: The method comprises the following steps: S100: Build a collaborative 3D perception platform for swarm robots. Each robot uses LiDAR to scan and acquire 3D point cloud data from a wide range of aircraft manufacturing scenes. The motion capture system then acquires each robot's real-time position and posture. S200: Label and segment the 3D point cloud data scanned by the swarm robot, and complete the production of collaborative point cloud detection datasets for large-scale scenarios in aircraft manufacturing. S300: Builds a point cloud collaborative 3D detection model based on cross-view consistency spatial mapping, including a point cloud feature extraction module, a feature fusion module, a feature reconstruction module, and a detection head; S400: Preprocessing the point cloud data in the training set and inputting it into the point cloud feature extraction module for feature encoding to obtain a plurality of intelligent agent features corresponding to the number of swarm robots, each intelligent agent feature corresponding to the three-dimensional point cloud data scanned by one robot; S500: Input multiple agent features into the feature fusion module, and map them to the consistent feature space through a shared cross-view consistency space mapping layer, thereby obtaining consistent features with the same number as the agent features; each agent feature is then mapped to a different feature space through an independent different cross-view difference space mapping layer, thereby obtaining difference features with the same number as the agent features; and the consistent features and difference features of all robots are fused to obtain multi-agent fusion features; S600: Inputting the consistency feature and difference feature of each robot into the feature reconstruction module for feature reconstruction to obtain the reconstructed features corresponding to each robot; S700: Input the multi-agent fusion features into the detection head to obtain the prediction results. Based on the consistency features, difference features, the features after reconstruction of each robot, the agent features of each robot and the prediction results, the parameters of the point cloud collaborative 3D detection model are updated through the set loss function until the point cloud collaborative 3D detection model converges; the validation set and test set are respectively input into the trained point cloud collaborative 3D detection model for calculation to obtain the 3D detection results, verify the effectiveness and generalization of the model, and then deploy it on the platform.
2. The method according to claim 1, characterized in that S100 includes: S110: Assemble and debug the mechanical and electrical structures of the cluster robots. The robots are numbered as follows: 、 、 A Livox non-repeating scanning LiDAR is installed in the slot directly in front of the swarm robot, with a reserved LiDAR installation location. The Livox LiDAR is connected to the Nvidia Orin edge computing device inside the swarm robot to perform laser scanning, data transmission, and model deployment. The Livox LiDAR follows the swarm robot's movements. S120: Using swarm robots 、 、 The Livox laser radar on board collects point cloud data of a large range of aircraft manufacturing scenes, and obtains point cloud data respectively 、 、 ,The real-time pose of each robot is obtained through the motion capture system.
3. The method according to claim 2, characterized in that S200 includes: S210: Use the point cloud annotation software LabelCloud to label the point cloud data: 、 、 Perform coordinate system transformation, and the real-time position of the robot obtained by the motion capture system will be 、 、 Mapped to the unified world coordinate system, forming a 、 、 Global collaborative point cloud data of point cloud information , for point cloud data Perform three-dimensional detection frame and category annotation, and the labeled data set is recorded as ; S220: Labeling the dataset Divide and label the dataset Randomly take out 65% of the data for training, as the training set, denoted as , the labeled dataset Randomly take out the equivalent of 15% of the data is used for verification, as the verification set, denoted as , the labeled dataset The rest of is used as the test set, which is , used to test the generalization of the model.
4. The method according to claim 3, characterized in that S400 includes: S410: Point cloud data acquired by each robot 、 、 With the laser radar as the center of the coordinate axis, the robot's forward direction is the x-axis direction, the robot's left direction is the y-axis direction, and the robot's upward direction is the z-axis direction. The point cloud is mathematically represented as follows: The vector representation of size is ,in are the three-dimensional coordinates of the point cloud data, is the reflection intensity of point cloud data, is the number of point clouds; S420: Select point cloud data 、 、 The coordinate range of the point cloud on the coordinate axis is specified as , within this range, the point cloud is divided into columns, and the specifications of each point cloud column are , take the center of each point cloud column as the center point coordinate, and then calculate the relative coordinates of each point and the center of the point cloud column , and use it as a supplementary representation of the point cloud, specifically The vector representation of size is ; S430: Use three PointPillars point cloud feature extraction networks with the same structure to extract 、 、 The point cloud data is feature encoded to obtain the intelligent body features 、 、 .
5. The method according to claim 4, characterized in that S500 includes: S510: Yes 、 、 Perform cross-view consistency feature mapping, specifically: ; ; ; in is a cross-view consistency spatial mapping layer, is the fully connected layer, is the activation function, is layer normalization, 、 、 Share a cross-view consistency spatial mapping layer, through Will 、 、 Mapped to the consistent feature space, the output is ; S520: Yes 、 、 Perform cross-view difference feature mapping, specifically: ; ; ; in 、 、 It is a spatial mapping layer that maps differences across perspectives. is the fully connected layer, is the activation function, is layer normalization, 、 、 Using different cross-view difference spatial mapping layers, 、 、 Will 、 、 Map to different feature spaces to obtain differential features 、 、 ; S530: consistency features and differential characteristics 、 、 Perform feature fusion, specifically: ; ; ; in Is a splicing operation, use Consistency characteristics of cluster robots and differential characteristics 、 、 Perform splicing at the feature level to obtain splicing features , and They are maximum pooling and average pooling operations, which are used for feature selection and dimensionality reduction. The pooled features are then concatenated to obtain features. , is the feature fusion layer, is a three-dimensional convolutional layer, is the activation function, through Perform feature fusion to obtain multi-agent fusion features .
6. The method according to claim 5, characterized in that S600 includes: For consistency features and differential characteristics 、 、 Perform feature reconstruction, specifically: ; ; in Is a consistent feature and differential characteristics of and, It is the feature reconstruction layer, through Reconstruct the consistent features and the different features, and 、 、 The intelligent agent features of swarm robots correspond one to one.
7. The method according to claim 6, characterized in that In S700, the multi-agent fusion features are input to the detection head to obtain the prediction results, including: The detection head completes the classification and regression tasks, Predictions are made through different network layers, specifically: ; ; in, Used to predict classification scores, it consists of a 1×1 two-dimensional convolution layer. The regression of the three-dimensional detection box consists of a 1×1 two-dimensional convolutional layer.
8. The method according to claim 7, characterized in that The S700 loss function is specifically: ; ; ; ; ; in, , , , , is the prediction box classification loss, is the 3D detection box regression loss, is the cross-view consistency-difference feature reconstruction loss, is the cross-view consistency distance loss, For the total loss.
9. The method according to claim 8, characterized in that In S700, the validation and test sets are fed into the trained point cloud collaborative 3D detection model to calculate and obtain 3D detection results. This verifies the model's effectiveness and generalization, and then deploys the platform, including: After completing the training batches, the trained 3D detection model is evaluated. The validation set is used to validate the model and select the best cluster robot 3D detection model. , the test set Input 3D detection model , get the test results ; 3D detection model It is deployed on the NVIDIA Orin edge computing platform of cluster robots, first converting it into ONNX intermediate data, and then converting the model into TensorRT for model acceleration, enabling cluster robots to collaboratively detect targets in a large range of scenes.
10. A cross-perspective collaborative perception system of swarm robots for aircraft manufacturing, characterized by: It includes 3D point cloud data acquisition module, data set production module, 3D detection model building module, feature extraction module, feature fusion module, feature reconstruction module, and training and deployment module; The 3D point cloud data acquisition module is used to build a collaborative 3D perception platform for swarm robots. Each robot uses a lidar to scan and acquire 3D point cloud data in a large-scale aircraft manufacturing scene, and the real-time position and posture of each robot is obtained through a motion capture system. The dataset creation module is used to label and segment the 3D point cloud data scanned by the swarm robot, completing the creation of collaborative point cloud detection datasets for large-scale scenarios in aircraft manufacturing. 3D detection model building module, used to build a point cloud collaborative 3D detection model based on cross-view consistency spatial mapping, including point cloud feature extraction module, feature fusion module, feature reconstruction module and detection head; The feature extraction module is used to pre-process the point cloud data in the training set and then perform feature encoding to obtain multiple agent features corresponding to the number of cluster robots. Each agent feature corresponds to the 3D point cloud data scanned by one robot. Feature fusion module: Multiple agent features are mapped to the consistent feature space through a shared cross-view consistency space mapping layer, and the number of consistent features is the same as that of the agent features. The features of each agent are then mapped to different feature spaces through independent different cross-view difference space mapping layers, obtaining the same number of difference features as the agent features. The consistency features and difference features of all robots are then fused to obtain multi-agent fusion features. The feature reconstruction module is used to reconstruct the consistent features and differential features of each robot to obtain the reconstructed features corresponding to each robot; The training and deployment module is used to input the multi-agent fusion features into the detection head to obtain the prediction results. Based on the consistency features, difference features, the features after reconstruction of each robot, the agent features of each robot and the prediction results, the parameters of the point cloud collaborative 3D detection model are updated through the set loss function until the point cloud collaborative 3D detection model converges; the validation set and test set are respectively input into the trained point cloud collaborative 3D detection model for calculation to obtain the 3D detection results, verify the effectiveness and generalization of the model, and then deploy it on the platform.
Citation Information
Patent Citations
Three-dimensional target detection method and system based on single-line laser radar and monocular camera
CN118068356A
Multi-Task Multi-Sensor Fusion for Three-Dimensional Object Detection
US20200160559A1