Battery replacement robot target point cloud completion method based on dynamic graph convolution

The target fastener point cloud is reconstructed through dynamic graph convolution network and pyramid point fractal generator, which solves the problem of point cloud data loss caused by viewing angle occlusion and uneven lighting, improves the geometric structure integrity and environmental perception capabilities of point clouds, and reduces the operation error of the battery swap robot.

CN120451468APending Publication Date: 2025-08-08SOUTHEAST UNIV

Patent Information

Application Number
CN202510536640.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the new energy vehicle battery swap system, the target fastener point cloud data is missing due to factors such as viewing angle occlusion and light unevenness, which affects the expressive ability of the point cloud geometric structure and the robustness and accuracy of subsequent pose estimation.

Method used

The target point cloud completion method of battery swap robot based on dynamic graph convolution is adopted to construct the target fastener point cloud data set through three-dimensional reconstruction and data enhancement, and the DGCNN-Trans feature extractor and pyramid point fractal generator are used to optimize point cloud generation with multi-objective joint loss function to reconstruct the missing geometric structure.

Benefits of technology

It significantly improves the integrity and structural fidelity of point clouds, improves the three-dimensional environment perception ability of the battery swap robot to target fasteners, and reduces subsequent operation errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451468A_ABST
    Figure CN120451468A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic graph convolution-based target point cloud completion method for a battery replacement robot, and the method comprises the steps: 1, reconstructing a complete target fastener model from a multi-view image through SFM and MVS technologies, and obtaining complete point cloud data; 2, constructing an incomplete-complete point cloud pair as training data by using a geometric constraint-based adaptive cutting strategy; 3, multi-resolution point cloud processing is adopted, and feature extraction and fusion are carried out according to three-level resolution; 4, on the basis of DGCNN dynamic graph convolution, in combination with multi-stage Edge Conv dynamic edge convolution and a Transform coding module, local geometric feature capture and global relation perception are realized; 5, point cloud generation adopts a pyramid step-by-step refining method, geometric details are added to each layer based on a previous layer result, and the number of points is gradually expanded; and 6, the training process is guided through a multi-target joint loss function, and the point cloud quality is improved while the precision is ensured. The geometric structure integrity of the point cloud of the target fastener is improved, and the operation error of follow-up operation of the battery replacement robot is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent battery-swapping robots, and specifically discloses a target point cloud completion method for a battery-swapping robot based on dynamic graph convolution. Background Art

[0002] In the battery swapping system of new energy vehicles, the target fasteners are important components connecting the battery pack and the vehicle body. Their structural integrity directly affects the accuracy and safety of automated disassembly and assembly operations. However, since the depth camera carried by the battery swapping robot is often interfered with by factors such as perspective occlusion, uneven lighting, and material surface reflection during the shooting process, the point cloud data of the target fasteners actually collected often have varying degrees of structural loss. This incompleteness of the point cloud not only impairs the ability to express the geometric structure of the target fasteners, but also significantly affects the robustness and accuracy of subsequent key links such as pose estimation and three-dimensional registration. Traditional point cloud completion methods mostly rely on intermediate representations (such as voxel grids, depth maps, etc.), but such methods often introduce quantization errors during the conversion process, resulting in information loss, and are limited in effectiveness when processing industrial point clouds with fine geometric structures and complex missing patterns. In recent years, end-to-end cloud completion methods based on deep learning have gradually emerged, especially those that directly use the original point cloud as input, which can improve flexibility and expression capabilities without sacrificing accuracy. However, existing methods often suffer from insufficient local feature extraction and limited ability to reason about missing structures, making them unable to cope with the non-random occlusion, irregular deformation, and high-density detail required in the point cloud of battery target fasteners. Therefore, effectively improving the ability to recover the semantic structure of missing regions in target fasteners is a key issue in enhancing 3D perception and high-precision operation capabilities in intelligent battery swapping systems.

[0003] Compared with the prior art, the differences are as follows:

[0004] Technical comparison with patent CN117635699A "Power battery pose estimation method based on machine vision point cloud segmentation";

[0005] Patent CN117635699A proposes a power battery pose estimation method based on machine vision point cloud segmentation, aiming to accurately locate the battery pack locking mechanism through point cloud segmentation technology, improve the positioning accuracy of the battery swap robot, and is applicable to a variety of electric vehicle models and battery pack locking mechanism forms. It can achieve high-precision pose estimation, effectively reduce battery swap operating costs, and improve battery swap efficiency. This patent proposes a target point cloud completion method for a battery swap robot based on dynamic graph convolution, focusing on solving the problem of missing target fastener point cloud data caused by factors such as limited viewing angle and uneven lighting. Through the optimized dynamic graph convolution network and pyramid point fractal generator architecture, it effectively reconstructs the missing geometric structure, significantly improving the point cloud integrity and structural fidelity. It is mainly used in point cloud data completion scenarios in electric vehicle battery swap systems, providing high-quality geometric data support for subsequent point cloud alignment and pose estimation. There are essential differences in the application scenarios of the two.

[0006] Patent CN117635699A captures RGBD images of the battery pack, uses a convolutional neural network to segment the locking mechanism foreground, and converts it into a point cloud. Point cloud preprocessing is achieved through point cloud filtering, outlier removal, and surface smoothing. A region growing method combined with the RANSAC algorithm is used to segment and cluster the point cloud, assigning semantic labels. The FPFH feature descriptor is calculated and fused with the point cloud segmentation results. Finally, the ICP fine registration algorithm is used to estimate the pose and complete the battery pack positioning. This patent uses 3D reconstruction and data enhancement methods to construct a complete target fastener point cloud dataset, and an adaptive cropping strategy based on geometric constraints is used to generate incomplete-complete point cloud pairs. The point cloud space is divided into inner, annular, and outer ring regions to simulate the loss caused by limited viewing angles. The FPS algorithm is used to construct a three-level resolution point cloud input. The DGCNN-Trans feature extractor is used to fuse local geometric features with global features. The point cloud skeleton is gradually refined using a pyramid point fractal generator. A multi-objective joint loss function is used to guide model training, including point cloud chamfer distance loss and hierarchical supervision loss. The two technical solutions differ fundamentally.

[0007] Patent CN117635699A uses a 3D camera to acquire RGBD images, then uses a convolutional neural network to perform image segmentation to extract the foreground of the locking mechanism. The two-dimensional image is projected into a three-dimensional point cloud, and point cloud filtering, outlier removal, and surface smoothing are performed. The region growing method and the RANSAC method are used to segment and cluster the point cloud for different parts, and semantic labels are assigned to the point cloud through spatial geometric features. The FPFH 3D descriptor is calculated, and feature fusion is performed based on the point cloud segmentation results. Finally, the ICP precise registration algorithm is used to achieve pose estimation. This patent uses an adaptive cropping strategy based on geometric constraints to construct training data, dividing the point cloud space into three regions: inner ring, annular ring, and outer ring. It also uses preset angle intervals to simulate structural loss caused by camera shooting angles. It uses the FPS algorithm to construct a three-level resolution point cloud input DGCNN-Trans feature extractor, which fuses four layers of Edge Conv dynamic edge convolution and a Transformer encoder to capture local geometric features and global features. It uses a pyramid point fractal generator to gradually refine the point cloud, using multiple fully connected networks to reduce the dimensionality of global features to generate a reference point cloud skeleton. It then gradually generates more refined point clouds based on features of different dimensions. During training, a multi-objective joint loss function consisting of a fine point cloud chamfer distance loss, a key point level supervision loss, a rejection loss, a uniformity loss, and a normal vector consistency loss is used to guide model optimization. The two technologies differ fundamentally in their technical approaches.

[0008] Patent CN117635699A enables precise positioning of locking mechanisms for multiple battery pack types, with an average registration error within ±2mm and a positioning error less than 3mm, while accommodating up to 30% occlusion. This accelerates the battery swapping process for electric vehicles, improves swapping accuracy, reduces operating costs, and enables the system to adapt to different vehicle chassis topologies and diverse battery pack models. This patent utilizes an improved dynamic graph convolutional network to achieve high-precision completion of missing parts in point clouds, essentially restoring the structure of circular regions and areas with limited viewing angles. Multi-resolution feature extraction and a pyramid point fractal generation strategy significantly enhance detail recovery, resulting in significantly improved geometric consistency and surface uniformity in the generated point cloud. The combination of dynamic graph convolution and a Transformer encoder enhances the network's ability to capture global structural relationships, enabling the completed point cloud to maintain the geometric characteristics of the original fastener. Through a multi-objective joint loss function optimization, the generated point cloud maintains geometric accuracy while significantly improving point distribution quality and local structure preservation, providing more accurate 3D environmental perception support for battery swapping robots and significantly reducing subsequent operational errors. The two technologies differ significantly in their technical effectiveness.

[0009] Technical comparison with patent CN119251305A "A method for estimating the pose of a battery-swap robot based on point cloud component segmentation and registration";

[0010] Patent CN119251305A proposes a pose estimation method for a battery-swapping robot based on point cloud component segmentation and registration. This method aims to capture images and point cloud data of the battery pack unlocking device using a 3D vision sensor, segment the lock head point cloud using a PointNet network, and perform point cloud registration in combination with principal component analysis and FPFH feature descriptions. This method achieves accurate pose estimation of the battery pack unlocking device, improving the battery-swapping robot's compatibility with battery packs of different vehicle models and structural sizes. This patent proposes a target point cloud completion method for a battery-swapping robot based on dynamic graph convolution. This method focuses on addressing the issue of missing point cloud data for target fasteners caused by factors such as limited viewing angles and uneven lighting. Through an optimized dynamic graph convolutional network and a pyramid point fractal generator architecture, it effectively reconstructs missing geometric structures, significantly improving point cloud integrity and structural fidelity. This method is primarily used in point cloud data completion scenarios within electric vehicle battery-swapping systems, providing high-quality geometric data support for subsequent point cloud registration and pose estimation. The application scenarios of the two methods differ fundamentally.

[0011] Patent CN119251305A uses a 3D vision sensor to collect color and depth images of the battery pack on the vehicle chassis in a smart battery swap station; applies image instance segmentation technology to the collected color image to output the instance segmentation mask for locking and unlocking; uses point cloud statistical filtering and voxel filtering to preprocess the point cloud; trains the PointNet point cloud segmentation network to extract the locking and unlocking lock point cloud; uses principal component analysis to calculate the important parts of the lock point cloud and calculates the FPFH features; uses the RANSAC method for coarse alignment and the ICP algorithm for fine alignment to finally obtain the estimated position of the locking and unlocking. This patent uses 3D reconstruction and data augmentation methods to construct a complete target fastener point cloud dataset. It employs a geometrically constrained adaptive cropping strategy to generate incomplete-complete point cloud pairs. The point cloud space is divided into inner, annular, and outer ring regions to simulate the loss caused by a restricted view angle. The FPS algorithm is used to construct a three-level resolution point cloud input. The DGCNN-Trans feature extractor is used to fuse local geometric features with global features. A pyramid point fractal generator is used to progressively refine the point cloud skeleton. Furthermore, a multi-objective joint loss function, including a point cloud chamfer distance loss and a hierarchical supervision loss, is applied to guide model training. The two technical solutions differ fundamentally.

[0012] Patent CN119251305A first uses a convolutional neural network to segment the locking and unlocking instances in the image and obtain instance masks; uses the alignment relationship and internal parameters of the color camera and depth camera to project the depth values into the original point cloud of the locking and unlocking; preprocesses the point cloud through point cloud statistical filtering and voxel filtering; uses the PointNet network to segment the point cloud and identify the lock point cloud; uses principal component analysis to extract the important parts of the lock point cloud and calculate the FPFH local features; finally, uses RANSAC for coarse alignment and ICP for fine alignment to obtain the final pose estimate of the locking and unlocking. This patent uses an adaptive cropping strategy based on geometric constraints to construct training data, dividing the point cloud space into three regions: inner ring, annular ring, and outer ring. It also uses preset angle intervals to simulate structural loss caused by camera shooting angles. It uses the FPS algorithm to construct a three-level resolution point cloud input DGCNN-Trans feature extractor, which fuses four layers of Edge Conv dynamic edge convolution and a Transformer encoder to capture local geometric features and global features. It uses a pyramid point fractal generator to gradually refine the point cloud, using multiple fully connected networks to reduce the dimensionality of global features to generate a reference point cloud skeleton. It then gradually generates more refined point clouds based on features of different dimensions. During training, a multi-objective joint loss function consisting of a fine point cloud chamfer distance loss, a key point level supervision loss, a rejection loss, a uniformity loss, and a normal vector consistency loss is used to guide model optimization. The two technologies differ fundamentally in their technical approaches.

[0013] Patent CN119251305A obtains the locking and unlocking lock point cloud through image instance segmentation and PointNet point cloud component segmentation, which can effectively adapt to the locking and unlocking devices of battery packs of different vehicle models and different structural sizes; the principal component analysis method is used to extract the important parts of the lock point cloud, which can effectively avoid the influence of different angles and distances of the visual sensor on the locking and unlocking pose estimation results; through FPFH feature description and RANSAC alignment method, it can achieve accurate alignment of the scene lock point cloud and the template lock point cloud; the pose estimation method uses the powerful nonlinear data feature learning ability of the neural network to improve the compatibility and accuracy of the battery pack locking and unlocking positioning. This patent achieves high-precision completion of missing parts in point clouds through an improved dynamic graph convolutional network, which can basically restore the structure of annular areas and areas with limited viewing angles; multi-resolution feature extraction and pyramid point fractal generation strategies significantly improve detail recovery capabilities, and the geometric consistency and surface uniformity of the generated point cloud are significantly improved; the combination of dynamic graph convolution and Transformer encoder enhances the network's ability to capture global structural relationships, allowing the completed point cloud to maintain the geometric characteristics of the original fasteners; through multi-objective joint loss function optimization, the generated point cloud maintains geometric accuracy while significantly improving point distribution quality and local structure retention, providing more accurate three-dimensional environmental perception support for battery-swap robots and significantly reducing subsequent operational errors. There is an essential difference between the two in terms of technical effects.

[0014] Technical comparison with patent CN115272655A "Visual positioning method and system device for multi-type battery packs for battery swapping robots";

[0015] Patent CN115272655A proposes a multi-type battery pack visual positioning method and system device for battery swapping robots. It aims to obtain the three-dimensional point cloud of the battery pack through visual sensors, combine color features and geometric features to identify unlocking holes, calculate the battery pack pose, and solve the problem of lack of compatibility of battery packs of multiple models and brands in battery swapping stations. It can accurately locate battery packs with different locking methods (snap-on, bolt-on, spinning, etc.), reduce dependence on vehicle parking accuracy, and reduce battery swapping costs and time. This patent proposes a target point cloud completion method for battery swapping robots based on dynamic graph convolution, focusing on solving the problem of missing target fastener point cloud data caused by factors such as limited viewing angle and uneven lighting. Through the optimized dynamic graph convolution network and pyramid point fractal generator architecture, it effectively reconstructs the missing geometric structure, significantly improving the point cloud integrity and structural fidelity. It is mainly used in point cloud data completion scenarios in electric vehicle battery swapping systems, providing high-quality geometric data support for subsequent point cloud alignment and pose estimation. There are essential differences in the application scenarios of the two.

[0016] Patent CN115272655A uses a visual sensor to obtain the image of the battery swap scene, and back-projects it to obtain a three-dimensional point cloud of the battery pack; uses a voxel filter to process the point cloud and adopts Euclidean clustering to distinguish the battery pack point cloud from the environmental point cloud; uses PCA principal component analysis and RANSAC algorithm to fit the vehicle chassis battery pack plane; uses HSV color space transformation to identify red feature marks and Hough circle transformation to detect unlocking holes; uses the least squares method to fit the normal and center position of the unlocking hole; and converts the fitting results into the workspace coordinate system of the battery swap station to guide the operation of the battery swap robot. This patent uses 3D reconstruction and data augmentation methods to construct a complete target fastener point cloud dataset. It employs a geometrically constrained adaptive cropping strategy to generate incomplete-complete point cloud pairs. The point cloud space is divided into inner, annular, and outer ring regions to simulate the loss caused by a restricted view angle. The FPS algorithm is used to construct a three-level resolution point cloud input. The DGCNN-Trans feature extractor is used to fuse local geometric features with global features. A pyramid point fractal generator is used to progressively refine the point cloud skeleton. Furthermore, a multi-objective joint loss function, including a point cloud chamfer distance loss and a hierarchical supervision loss, is applied to guide model training. The two technical solutions differ fundamentally.

[0017] Patent CN115272655A fixes the visual sensor 0.5 meters from the bottom of the battery swap platform to collect the three-dimensional point cloud of the battery pack; removes noise through voxel filtering and uses Euclidean clustering to separate the battery pack point cloud from the environmental point cloud; applies PCA principal component analysis to obtain the initial posture of the battery pack, and combines the RANSAC algorithm to accurately fit the battery pack plane; uses HSV color space transformation to identify the red logo, and uses Hough circle transform to detect the geometric shape of the unlocking hole, and fuses the two detection results to confirm the position of the unlocking hole; applies the least squares method to fit the normal direction and center coordinates of each unlocking hole point cloud; finally, converts all information into the global coordinate system of the battery swap station to provide operational guidance for the battery swap robot. This patent uses an adaptive cropping strategy based on geometric constraints to construct training data, dividing the point cloud space into three regions: inner ring, annular ring, and outer ring. It also uses preset angle intervals to simulate structural loss caused by camera shooting angles. It uses the FPS algorithm to construct a three-level resolution point cloud input DGCNN-Trans feature extractor, which fuses four layers of Edge Conv dynamic edge convolution and a Transformer encoder to capture local geometric features and global features. It uses a pyramid point fractal generator to gradually refine the point cloud, using multiple fully connected networks to reduce the dimensionality of global features to generate a reference point cloud skeleton. It then gradually generates more refined point clouds based on features of different dimensions. During training, a multi-objective joint loss function consisting of a fine point cloud chamfer distance loss, a key point level supervision loss, a rejection loss, a uniformity loss, and a normal vector consistency loss is used to guide model optimization. The two technologies differ fundamentally in their technical approaches.

[0018] Patent CN115272655A uses a multimodal recognition method based on color and geometric features to simultaneously be compatible with battery packs with different locking methods such as snap-on, bolt-on and spin-on types, thereby improving the compatibility of the battery swap robot with battery packs of multiple brands and models; precise three-dimensional plane fitting and unlocking hole positioning enable the battery swap robot to accurately dock with the vehicle battery pack, reducing the requirements for vehicle parking accuracy and eliminating the need for additional positioning and adjustment devices; the system adopts an "eye-to-hand" visual sensor configuration method to perceive the battery pack position in real time and can dynamically adapt to changes in the battery pack position under different vehicle chassis morphologies; the positioning system realizes information sharing and collaboration with the battery swap station and robot controller, improving the efficiency and success rate of the entire battery swap process. This patent achieves high-precision completion of missing parts in point clouds through an improved dynamic graph convolutional network, which can basically restore the structure of annular areas and areas with limited viewing angles; multi-resolution feature extraction and pyramid point fractal generation strategies significantly improve detail recovery capabilities, and the geometric consistency and surface uniformity of the generated point cloud are significantly improved; the combination of dynamic graph convolution and Transformer encoder enhances the network's ability to capture global structural relationships, allowing the completed point cloud to maintain the geometric characteristics of the original fasteners; through multi-objective joint loss function optimization, the generated point cloud maintains geometric accuracy while significantly improving point distribution quality and local structure retention, providing more accurate three-dimensional environmental perception support for battery-swap robots and significantly reducing subsequent operational errors. There is an essential difference between the two in terms of technical effects.

[0019] Technical comparison with patent CN119594979A "A method for measuring the navigation posture of a battery-swap robot based on three-dimensional visual features";

[0020] Patent CN119594979A proposes a navigation pose measurement method for a battery swapping robot based on three-dimensional visual features. This method uses a 3D visual sensor to capture color and depth images of the battery pack unlocking area, then uses a convolutional neural network to locate and segment the target. Combined with the FPS sampling method, it extracts key points and guide vectors, establishes a correspondence between the template point cloud and the scene point cloud, and achieves high-precision pose estimation for battery pack unlocking. This improves the battery swapping robot's positioning compatibility for battery packs of different vehicle models and sizes, accelerating the battery swapping efficiency of new energy vehicles. This patent proposes a target point cloud completion method for battery swapping robots based on dynamic graph convolution. This method focuses on addressing the issue of missing point cloud data for target fasteners caused by factors such as limited viewing angle and uneven lighting. By utilizing an optimized dynamic graph convolutional network and a pyramid point fractal generator architecture, it effectively reconstructs missing geometric structures, significantly improving point cloud integrity and structural fidelity. This method is primarily used for point cloud data completion in electric vehicle battery swapping systems, providing high-quality geometric data support for subsequent point cloud registration and pose estimation. The two application scenarios differ significantly.

[0021] Patent CN119594979A installs a 3D vision sensor on a battery-swapping robot to collect color and depth images of the battery pack; uses a convolutional neural network to segment the encryption and unlocking instances in the image to obtain the encryption and unlocking mask; projects the depth value into an initial point cloud based on the alignment relationship between the color camera and the depth camera; performs statistical filtering and voxel filtering preprocessing on the point cloud; uses the FPS sampling method to collect 5 key points from the template point cloud, calculates the point cloud object coordinate system and obtains the guide vector; uses the guide vector to select points in the scene point cloud that correspond to the key points of the template point cloud; finally, based on the key point correspondence, the least squares method is used to fit the initial pose, and the ICP algorithm is used for precise alignment optimization. This patent uses 3D reconstruction and data augmentation methods to construct a complete target fastener point cloud dataset. It employs a geometrically constrained adaptive cropping strategy to generate incomplete-complete point cloud pairs. The point cloud space is divided into inner, annular, and outer ring regions to simulate the loss caused by a restricted view angle. The FPS algorithm is used to construct a three-level resolution point cloud input. The DGCNN-Trans feature extractor is used to fuse local geometric features with global features. A pyramid point fractal generator is used to progressively refine the point cloud skeleton. Furthermore, a multi-objective joint loss function, including a point cloud chamfer distance loss and a hierarchical supervision loss, is applied to guide model training. The two technical solutions differ fundamentally.

[0022] Patent CN119594979A uses 3D vision sensors to collect battery swap scene data, and uses a deep learning network to intelligently segment the battery pack unlocking and unlocking areas, combining the segmentation results with depth information to generate an initial point cloud; statistical filtering is used to remove outliers, and voxel filtering is used to reduce the point cloud density and improve processing efficiency; the FPS sampling method is innovatively introduced to extract representative key points from the template point cloud and calculate the guide vector relative to the point cloud object coordinate system; the corresponding key points are searched in the scene point cloud through the angle and length constraints of the guide vector to achieve feature matching under different shooting perspectives; finally, the least squares method and SVD decomposition technology are used to calculate the transformation matrix, combined with ICP precise alignment to achieve high-precision pose estimation. This patent uses an adaptive cropping strategy based on geometric constraints to construct training data, dividing the point cloud space into three regions: inner ring, annular ring, and outer ring. It also uses preset angle intervals to simulate structural loss caused by camera shooting angles. It uses the FPS algorithm to construct a three-level resolution point cloud input DGCNN-Trans feature extractor, which fuses four layers of Edge Conv dynamic edge convolution and a Transformer encoder to capture local geometric features and global features. It uses a pyramid point fractal generator to gradually refine the point cloud, using multiple fully connected networks to reduce the dimensionality of global features to generate a reference point cloud skeleton. It then gradually generates more refined point clouds based on features of different dimensions. During training, a multi-objective joint loss function consisting of a fine point cloud chamfer distance loss, a key point level supervision loss, a rejection loss, a uniformity loss, and a normal vector consistency loss is used to guide model optimization. The two technologies differ fundamentally in their technical approaches.

[0023] Patent CN119594979A greatly enhances the system's adaptability to point cloud deformation and missing points by matching feature point correspondences rather than overall shape matching, and can maintain high-precision pose estimation even when the point cloud is incomplete; the use of a key point matching strategy guided by a guide vector effectively overcomes the point cloud hole problem that is prone to occur when the object is facing away from the camera, and improves the robustness of feature expression; the technical route that combines deep learning with traditional geometric calculation methods enables the system to intelligently identify different types of battery pack unlocking structures, and improves the compatibility of the battery swap robot with diverse battery packs; instance segmentation technology combined with 3D point cloud processing allows the system to accurately distinguish between unlocking areas and other parts, reducing environmental interference and improving algorithm stability. This patent achieves high-precision completion of missing parts in point clouds through an improved dynamic graph convolutional network, which can basically restore the structure of annular areas and areas with limited viewing angles; multi-resolution feature extraction and pyramid point fractal generation strategies significantly improve detail recovery capabilities, and the geometric consistency and surface uniformity of the generated point cloud are significantly improved; the combination of dynamic graph convolution and Transformer encoder enhances the network's ability to capture global structural relationships, allowing the completed point cloud to maintain the geometric characteristics of the original fasteners; through multi-objective joint loss function optimization, the generated point cloud maintains geometric accuracy while significantly improving point distribution quality and local structure retention, providing more accurate three-dimensional environmental perception support for battery-swap robots and significantly reducing subsequent operational errors. There is an essential difference between the two in terms of technical effects. Summary of the Invention

[0024] In order to solve the above technical problems, the present invention proposes a target point cloud completion method for a battery-swapping robot based on dynamic graph convolution, which improves the geometric structure integrity of the target fastener point cloud and reduces the operational errors of subsequent operations of the battery-swapping robot.

[0025] To achieve the above object, the technical solution adopted by the present invention is:

[0026] A target point cloud completion method for a battery-swapping robot based on dynamic graph convolution includes the following steps:

[0027] (1) Reconstruct the 3D model to obtain the 3D point cloud data of the target fastener with complete structure; comprehensively utilize the rigid transformation method, non-rigid transformation method, and data expansion method to enhance the complete point cloud of the target fastener, and construct the target fastener point cloud dataset FastComp3D that can simulate the multi-pose, multi-view, and multi-deformation missing forms under real working conditions, denoted as point cloud P0;

[0028] (2) For each original point cloud P0 data, an adaptive cropping strategy based on geometric constraints is used to construct an incomplete-complete point cloud pair for deep learning model training and evaluation. The specific implementation idea is to divide the point cloud space P0 into three areas: inner ring, annular ring, and outer ring to simulate the loss of the annular area inside the hole, and combine it with the preset angle interval to simulate the loss of the back-to-camera surface structure caused by the camera shooting angle. Then, the cropping strategy is dynamically adjusted according to the batch index to achieve structure-aware cropping of the point cloud data. Based on the complete point cloud, 1536 points are cropped to obtain an incomplete point cloud P1 containing only 2560 points.

[0029] (3) The FPS algorithm is used to downsample the incomplete point cloud P1 to a medium-resolution point cloud P2 with only 1024 points and a coarse-resolution point cloud P3 with only 512 points. A three-level resolution point cloud list [P1, P2, P3] is constructed and input into the DGCNN-Tr ans feature extractor to extract geometric features F1, F2, and F3. Subsequently, the multi-scale information is integrated through feature splicing and an MLP layer to finally generate a 2048-dimensional global feature vector F containing geometric details and structural semantics. global ;

[0030] (4) The DGCNN-Trans feature extractor is based on the DGCNN dynamic graph convolution framework. First, the input point cloud data P i Extract geometric features F of different scales through four layers of Edge Conv operation i1 、F i2 、F i3 、F i4 , then the geometric features of different scales are spliced and the dimension order is adjusted to obtain F cat , F cat Input Transformer Encoder module to obtain the global feature F between points trans , use the 5th layer Edge Conv to perform global feature F trans Perform another side convolution operation to obtain the high-dimensional feature F i5 ,Then, the high-dimensional feature F5 is aggregated through global maximum pooling and average pooling to ,aggregate spatial dimension information and generate compact global feature vectors F1, F2, ,F3;

[0031] (5) Using the pyramid point fractal generator: the 2048-dimensional global feature F global Gradually reduce the dimension to F through a three-layer fully connected network 1024 、F 512 、F 256 , construct multi-level point features. Then, use the lowest dimensional feature F 256 Generate 64 reference points as the coarse-grained point cloud skeleton P coarse ; Then based on the middle-level feature F512 , calculate the geometric offset for each reference point and generate 2 local points to form a medium-grained point cloud P containing 128 points medium ; Finally, using the highest dimensional feature F 1024 Through a complex feature transformation network, each medium-precision point is expanded into multiple fine points, and finally a high-precision missing point cloud P containing 1536 points is generated. final ;

[0032] (6) During the model training process, the fine point cloud chamfer distance loss L is comprehensively utilized CD , key point level supervision loss α1L key1 +α2L key2 , rejection loss λ1L rep , uniformity loss λ2L uni , normal vector consistency loss λ3L norm The multi-objective joint loss function is used to guide the training process of the point cloud completion network to optimize the point cloud generation direction.

[0033] As a further improvement of the present invention, the method for preparing the structurally complete target fastener point cloud FastComp3D dataset in step (1) is as follows:

[0034] (1-1) Use a depth camera to shoot a video around the target fastener and convert it into a continuous frame image;

[0035] (1-2) In the WORKFLOW panel of the software, click "Folder" to import the continuous frame image folder. Click "Start" to start the automatic reconstruction process, including steps such as point cloud generation, map generation, and model reconstruction. In the "TOOLS" panel, use the "Lasso / Rect / Box" selection tool to select and delete unnecessary background.

[0036] (1-3) Convert the generated obj format model file into a point cloud file in ply or txt format;

[0037] (1-4) The rigid transformation method in data enhancement sets the scaling factor range to 0.8-1.2 times, the translation range to ±10% of the size range, and the rotation angle to ±180° for rigid transformation; the non-rigid transformation method randomly generates 5 control points in the point cloud space, sets the deformation degree to 0.1, and then assigns a random displacement to each control point. Then, the radial basis function RBF is used for TPS interpolation to calculate the displacement of all points. Finally, the calculated displacement is applied to the original point cloud to obtain the deformation result; the data expansion method collects 3D models of similar bolted joints from public resources and converts them into point cloud data format; In addition, Gaussian noise is used in the enhancement process to simulate the error in the actual scanning process;

[0038] (1-5) Repeat the above operations multiple times using the target fastener after initial 3D reconstruction and the target fastener after data enhancement.

[0039] As a further improvement of the present invention, the clipping strategy in step (2) includes the following steps:

[0040] (2-1) Divide the point cloud space into inner ring area R according to spherical coordinates inner , middle annular area R ring and the outer ring area R outer , cropping priority from R ring Area deletion points to simulate the missing structure at the edge of the target fastener hole;

[0041] (2-2) Combined with the spherical coordinate azimuth angle θ∈[θ1,θ2], the area of the point cloud facing away from the camera is culled within a specific angle range to simulate the visibility loss caused by different shooting angles;

[0042] (2-3) According to the current sample number i∈[1,N] and batch index b i , cropping masks are dynamically generated to achieve cyclic combination of multi-modal cropping strategies.

[0043] As a further improvement of the present invention, the specific implementation of FPS downsampling, feature splicing and fusion in step (3) is as follows:

[0044] (3-1) For the generated incomplete point cloud P1, randomly select an initial point p0∈P1 and add it to the sampling point set S={p0}, and calculate each point p i ∈P1 is the minimum distance d from the current sampling point set S i , select d i The largest point p next As the next sampling point, and p next Add the sampling point set S and iterate t=1, 2, ..., K-1 times. When the size of the sampling point set S reaches the number of sampling points K, the algorithm terminates.

[0045] (3-2) For each resolution of the point cloud, input it into a DGCNN-Trans feature extractor, and concatenate its output features F1, F2, and F3 using the Concat operation in PyTorch to form a multi-scale feature F;

[0046] (3-3) Feature F is compressed by one-dimensional convolution Conv1D, batch normalization BatchNormalization, BN and ReLU activation function to obtain the global feature vector F global .

[0047] As a further improvement of the present invention, the feature extractor in step (4) is specifically implemented as follows:

[0048] (4-1)EdgeConv: For the input point cloud P i Or upper-level feature F i=1,2,3 , use KNN algorithm for each point p i Calculate its K nearest neighbor points Use the nearest neighbor point set to build a local area graph, and then use the directed edge v ij Calculate each center point p i The K nearest neighbor edge feature e ij ; For each center point, its K edge features e ij Aggregate through the maximum pooling operation to obtain the fusion feature F K :

[0049] (4-2) Continuously perform four layers of Edge Conv edge convolution to obtain geometric features F of different scales i1 、F i2 、F i3 、F i4 , where each layer of Edge Conv will recalculate the K nearest neighbor point set based on the new local relationship, so that the network can dynamically adjust the connection relationship between points in the feature space; then the geometric features of different scales are spliced and the dimension order is adjusted to obtain F cat ;

[0050] (4-3) Transformer encoder: The fused multi-scale geometric features F cat As input, it is linearly transformed using multi-head self-attention to obtain the query Q, key K and value V; then the attention weight score A is calculated by scaling the dot product, and the attention weight A is weighted summed with the value vector V to obtain the feature F att ; Then the original F cat With F att After achieving residual connection, normalization is performed to obtain feature F norm ; Finally, it is input into the two-layer fully connected network to obtain the globally enhanced feature representation F fin .

[0051] As a further improvement of the present invention, the pyramid point fractal generator in step (5) is specifically implemented as follows:

[0052] (5-1) The 2048-dimensional global feature F global Gradually reduce the dimension to F through a three-layer fully connected network 1024 、F 512 、F 256 , construct multi-level point features;

[0053] (5-2) Then, using the lowest dimensional feature F 256 Generate 64 reference points as the coarse-grained point cloud skeleton P coarse ; Then based on the middle-level feature F 512 , calculate the geometric offset for each reference point and generate 2 local points to form a medium-grained point cloud P containing 128 points medium ; Finally, using the highest dimensional feature F 1024 Through a complex feature transformation network, each medium-precision point is expanded into multiple fine points, and finally a high-precision missing point cloud P containing 1536 points is generated. final .

[0054] As a further improvement of the present invention, the method for constructing the multi-objective joint loss function in step (6) is:

[0055] (6-1) The fine point cloud chamfer distance loss L CD , key point level supervision loss L key1 , L key2 , rejection loss L rep , uniformity loss L uni , normal vector consistency loss L norm Add the weights of α1, α2, λ1, λ2, and λ3 to get the final overall loss L for training total .

[0056] Beneficial effects:

[0057] The present invention discloses a target point cloud completion method for a battery-swapping robot based on dynamic graph convolution. This method constructs a target fastener point cloud dataset FastComp3D that can simulate multi-posture, multi-view, and multi-deformation missing forms under real working conditions based on three-dimensional model reconstruction and data enhancement, providing data support for model training. Before the start of model training, an adaptive cropping algorithm based on geometric constraints is used to construct an incomplete-complete point cloud pair that can simulate the real camera shooting to ensure that the trained model can well complete the real incomplete point cloud. In addition, the DGCNN feature extractor improved based on Transformer is used, combined with multi-resolution point cloud input and pyramid point fractal generator, to improve the model's ability to extract geometric details and perceive global structure. Finally, a multi-objective joint loss function is used to optimize the point cloud generation direction to ensure the geometric structure integrity and point distribution uniformity of the completion result. This scheme can effectively reconstruct the missing parts of the target fasteners, provide accurate three-dimensional environmental perception capabilities for the battery-swapping robot, and help improve the accuracy of target fastener point cloud registration and pose estimation. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 It is a flow chart of the method disclosed in the present invention. DETAILED DESCRIPTION

[0059] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:

[0060] The present invention provides a target point cloud completion method for a battery-swapping robot based on dynamic graph convolution, which can reconstruct the incomplete geometric structure parts in the target fastener point cloud and improve the accuracy of subsequent operations of the battery-swapping robot.

[0061] As an embodiment of the present invention, the present invention provides a target point cloud completion method for a battery-swapping robot based on dynamic graph convolution, wherein the target fastener point cloud completion flow chart is as follows: Figure 1 As shown, this is achieved through the following steps:

[0062] (1) Reality Capture 1.4 software was used to reconstruct the 3D model and obtain the 3D point cloud data of the target fastener with complete structure. The rigid transformation method, non-rigid transformation method, and data augmentation method were used to enhance the complete point cloud of the target fastener. The FastComp3D point cloud dataset of the target fastener was constructed, which can simulate the multi-pose, multi-view, and multi-deformation missing forms under real working conditions. It is denoted as point cloud P0.

[0063] (2) For each original point cloud P0 data, an adaptive cropping strategy based on geometric constraints is adopted to construct an incomplete-complete point cloud pair for deep learning model training and evaluation. The specific implementation idea is to divide the point cloud space P0 into three areas: inner ring, annular ring and outer ring to simulate the loss of the annular area inside the hole, and combine it with the preset angle interval to simulate the loss of the back-to-camera surface structure caused by the camera shooting angle. Then, the cropping strategy is dynamically adjusted according to the batch index to achieve structure-aware cropping of the point cloud data. Based on the complete point cloud, 1536 points are cropped to obtain an incomplete point cloud P1 containing only 2560 points.

[0064] (3) The FPS algorithm is used to downsample the incomplete point cloud P1 to a medium-resolution point cloud P2 containing only 1024 points and a coarse-resolution point cloud P3 containing only 512 points. A three-level resolution point cloud list [P1, P2, P3] is constructed and input into the DGCNN-Trans feature extractor to extract geometric features F1, F2, and F3. Subsequently, the multi-scale information is integrated through feature concatenation and an MLP layer to finally generate a 2048-dimensional global feature vector F containing geometric details and structural semantics. global ;

[0065] (4) The DGCNN-Trans feature extractor is based on the DGCNN dynamic graph convolution framework. First, the input point cloud data P i Extract geometric features F of different scales through four layers of Edge Conv operation i1 、Fi2 、F i3 、F i4 , then the geometric features of different scales are spliced and the dimension order is adjusted to obtain F cat , F cat Input Transformer Encoder module to obtain the global feature F between points trans ; Use the 5th layer Edge Conv to perform global feature F trans Perform another side convolution operation to obtain the high-dimensional feature F i5 ,Then, the high-dimensional feature F5 is aggregated through global maximum pooling and average pooling to ,aggregate spatial dimension information and generate compact global feature vectors F1, F2, ,F3;

[0066] (5) The 2048-dimensional global feature F global Gradually reduce the dimension to F through a three-layer fully connected network 1024 、F 512 、F 256 , construct multi-level point features. Then, use the lowest dimensional feature F 256 Generate 64 reference points as the coarse-grained point cloud skeleton P coarse ; Then based on the middle-level feature F 512 , calculate the geometric offset for each reference point and generate 2 local points to form a medium-grained point cloud P containing 128 points medium ; Finally, using the highest dimensional feature F 1024 Through a complex feature transformation network, each medium-precision point is expanded into multiple fine points, and finally a high-precision missing point cloud P containing 1536 points is generated. final ;

[0067] (6) During the model training process, the fine point cloud chamfer distance loss L is comprehensively utilized CD , key point level supervision loss α1L key1 +α2L key2 , rejection loss λ1L rep , uniformity loss λ2L uni , normal vector consistency loss λ3L norm The multi-objective joint loss function is used to guide the training process of the point cloud completion network to optimize the point cloud generation direction.

[0068] The method for producing the structurally complete target fastener point cloud FastComp3D dataset in step (1) is as follows:

[0069] (1-1) Use a depth camera to shoot a video around the target fastener and convert it into a continuous frame image;

[0070] (1-2) In the Workflow panel of Reality Capture 1.4, click "Folder" to import the folder of consecutive frame images. Click "Start" to start the automatic reconstruction process, which includes steps such as point cloud generation, map generation, and model reconstruction. In the Tools panel, use the "Lasso / Rect / Box" selection tool to select and delete the unnecessary background.

[0071] (1-3) Use Cloud Compare software to convert the generated obj format model file into a point cloud file in ply or txt format;

[0072] (1-4) The rigid transformation method in data enhancement sets the scaling factor range to 0.8-1.2 times, the translation range to ±10% of the size range, and the rotation angle to ±180° for rigid transformation. The non-rigid transformation method randomly generates 5 control points in the point cloud space, sets the deformation degree to 0.1, and then assigns a random displacement to each control point. Then, the radial basis function (RBF) is used for TPS interpolation to calculate the displacement of all points. Finally, the calculated displacement is applied to the original point cloud to obtain the deformation result. The data augmentation method collects 3D models of similar bolted joints from public resources and converts them into point cloud data format. In addition, Gaussian noise is used in the enhancement process to simulate the error in the actual scanning process:

[0073]

[0074] The geometric constraint-based adaptive clipping strategy in step (2) includes the following steps:

[0075] (2-1) Calculate the center of the point cloud:

[0076]

[0077] (2-2) Calculate the distance from each point to the center and the angle from each point to the center:

[0078]

[0079] (2-3) Determine the adaptive annular area parameters:

[0080] inner_ring=percentile(D,35%-42%)

[0081] outer_ring=percentile(D,70%-75%)

[0082] (2-4) If the batch index b i If it is an even number, the starting point of the angle is randomly selected Construct the angle interval and angle mask to generate the cropping area mask combination:

[0083]

[0084] M θ =θ∈A r

[0085] M=M θ &(Μ ring |Μ inner )

[0086] (2-5) If the batch index b i If it is an odd number, calculate the number of points N in the annular area ring , if N ring <0.8K, then expand the outer ring radius. If N ring >1.5K, the adaptive angle range is calculated to generate a local cropping area mask;

[0087] (2-6) Extracting the cropped point cloud P crop , set the clipping coordinates to 0 and update the residual point cloud P remain .

[0088] The specific implementation of FPS downsampling, feature splicing and fusion in step (3) is as follows:

[0089] (3-1) For the generated incomplete point cloud P1, randomly select an initial point p0∈P1 and add it to the sampling point set S={p0}, and calculate each point p i ∈P1 is the minimum distance d from the current sampling point set S i :

[0090]

[0091] Here, ||·||2 represents the Euclidean distance.

[0092] (3-2) Select d i The largest point p next As the next sampling point, and p next Add the sampling point set S and iterate t=1, 2, ..., K-1 times. When the size of the sampling point set S reaches the number of sampling points K, the algorithm terminates.

[0093] (3-3) For each resolution of the point cloud, it is input into a DGCNN-Trans feature extractor, and its output features F1, F2, and F3 are spliced together using the Concat operation in PyTorch to form a multi-scale feature F:

[0094] F=concat(F (1) ,F(2) ,F (3) )

[0095] Among them, F (m) Represents the feature vector obtained after the point cloud of the mth layer resolution passes through the feature extractor;

[0096] (3-4) Feature F is compressed by one-dimensional convolution Conv1D, batch normalization BatchNormalization and ReLU activation function to obtain the global feature vector F global .

[0097] The DGCNN-Trans feature extractor in step (4) includes the following steps:

[0098] (4-1)EdgeConv: For the input point cloud P i Or upper-level feature F i=1,2,3 , use KNN algorithm for each point p i Calculate the k nearest points:

[0099]

[0100] in Represents the kth nearest neighbor point, the nearest neighbor point set satisfy:

[0101]

[0102] (4-2) Using the nearest neighbor point set Construct a local area map, each layer of the map is composed of the input point cloud P i Or upper-level feature F i=1,2,3 Each center point p i and directed edge v ij Composition. ij It is p i Relative to its neighboring point p ik Directed edges of :

[0103]

[0104] Among them, v ij =p ij -p i |j=1,2,...,k, k is the number of selected neighbor points, n is the number of point clouds. ij Each center point p can be calculated i The K nearest neighbor edge feature e ij , the calculation formula is:

[0105] e ij =h Θ (p i,v ij )

[0106] Where: h Θ It is a feature transformation function, which is implemented by expanding the center point feature P and concatenating it with the feature V in the neighborhood graph.

[0107] (4-3) For each center point, its K edge features are aggregated through the maximum pooling operation:

[0108]

[0109] Get the fusion feature F K .

[0110] (4-5) Continuously perform four layers of Edge Conv edge convolution to obtain geometric features F of different scales i1 、F i2 、F i3 、F i4 , where each layer of Edge Conv will recalculate the K nearest neighbor point set based on the new local relationship:

[0111]

[0112] Among them, l represents the number of network layers, X l is the feature map of this layer. This enables the network to dynamically adjust the connection relationship between points in the feature space; then the geometric features of different scales are spliced and the dimension order is adjusted to obtain F cat ;

[0113] (4-6) The fused multi-scale geometric features F cat As input, we use multi-head self-attention to perform a linear transformation to obtain the query Q, key K and value V:

[0114] Q=F cat W Q ,K=F cat W K ,V=F cat W V

[0115] Among them, W Q , W K , W V is a learnable parameter. The attention score is then calculated by scaling the dot product:

[0116]

[0117] Among them, d k is the dimension of the key vector. The output feature is the weighted sum of the attention weight and the value vector:

[0118]

[0119] (4-7) To avoid the gradient disappearance, the original F cat With F att After achieving residual connection, normalization is performed to obtain feature F norm :

[0120] F norm =LayerNorm(F cat ,F att )

[0121] (4-8) The feature F norm Enter a two-layer fully connected network:

[0122] F fin =Droupout(σ(W2Droupout(GeLu(W1F norm ))))

[0123] Get the globally enhanced feature representation F fin .

[0124] The specific implementation of the pyramid point fractal generator in step (5) is:

[0125] (5-1) The 2048-dimensional global feature F global Gradually reduce the dimension to F through a three-layer fully connected network 1024 、F 512 、F 256 , construct multi-level point features:

[0126] F 1024 =Linear(F global )

[0127] F 512 =Linear(F 1024 )

[0128] F 256 =Linear(F 512 )

[0129] (5-2) Using the lowest dimensional feature F 256 Generate 64 reference points as the coarse-grained point cloud skeleton P coarse :

[0130] P coarse =FC(Reshape(F 256 ))

[0131] (5-3) Based on the middle-level feature F 512 , calculate the geometric offset for each reference point and generate 2 local points to form a medium-grained point cloud P containing 128 pointsmedium :

[0132] P medium =Reshape(Conv1D(Linear(F 512 )))+P coarse

[0133] (5-4) Using the highest dimensional feature F 1024 Through the feature transformation network, each medium-precision point is expanded into multiple fine points, and finally a high-precision missing point cloud P containing 1536 points is generated. final :

[0134] P final =Reshape(Conv1D(Conv1D(Reshape(Linear(F 1024 )))))+P medium

[0135] The construction method of the multi-objective joint loss function in step (6) is:

[0136] (6-1) Fine point cloud chamfer distance loss L CD :

[0137]

[0138] Among them, p pre To generate point cloud, p gt is the real point cloud p gt ;

[0139] (6-2) Key point level supervision loss L key1 , L key2 : For the first-level key points (64 points) and second-level key points (128 points) generated by the decoder, calculate the chamfer distance loss between them and the corresponding level key points of the real point cloud:

[0140]

[0141] (6-3) Rejection loss L rep : Calculate when the distance between two points is less than the distance threshold h, generate a penalty inversely proportional to the distance, forcing the network to generate more dispersed points:

[0142]

[0143] Among them, the distance threshold h is set to 0.03.

[0144] (6-4) Uniformity loss L uni: Calculate the variance of the distances of the k nearest neighbor points near each point, and encourage the minimization of this variance, so that the point cloud is more evenly distributed in the local area:

[0145]

[0146] Among them, D i = ||p i - p j ||, j ∈ N i (K) is the set of distances from the point p i to its K nearest neighbors, and the number of nearest neighbor points k = 10.

[0147] (6 - 5) Normal vector consistency loss L norm : Calculate the consistency between the generated point cloud p pre and the real point cloud p gt in the direction of the local surface normal vector, and its calculation formula is:

[0148]

[0149] Among them, n i is the normal vector of the point in the input point cloud that is closest to the point p i in the generated points, c i is the nearest neighbor point of p i in the input point cloud, and the number of nearest neighbor points k = 10.

[0150] (6 - 6) Add the above loss functions according to the weights α1, α2, λ1, λ2, λ3 to obtain the overall loss L total used for training finally:

[0151] L total = L CD + α1L key1 + α2L key2 + λ1L rep + λ2L uni + λ3L norm

[0152] Among them, the weight coefficients α1, α2 are dynamically adjusted according to the number of training epochs. At the beginning of training when Epoch ≤ 30, set α1 = 0.01, α2 = 0.02 to focus on the overall shape reconstruction first; as the training progresses, when 30 < Epoch ≤ 80, increase to α1 = 0.05, α2 = 0.1 to strengthen the detail recovery; when 80 < Epoch ≤ 200, finally reach α1 = 0.1, α2 = 0.2. For the regularization terms, set λ1 = 0.1, λ2 = 0.05, λ3 = 0.1 to improve the quality of the point cloud while ensuring geometric accuracy.

[0153] The above description is merely a preferred embodiment of the present invention and does not constitute any other form of limitation to the present invention. Any modification or equivalent variation based on the technical essence of the present invention shall still fall within the scope of protection claimed by the present invention.

Claims

1. A target point cloud completion method for a battery-swapping robot based on dynamic graph convolution, characterized in that: The steps include: (1) Reconstruct the 3D model to obtain the 3D point cloud data of the target fastener with complete structure; comprehensively utilize the rigid transformation method, non-rigid transformation method, and data expansion method to enhance the complete point cloud of the target fastener, and construct the target fastener point cloud dataset FastComp3D that can simulate the multi-pose, multi-view, and multi-deformation missing forms under real working conditions, denoted as point cloud P0; (2) For each original point cloud P0 data, an adaptive cropping strategy based on geometric constraints is used to construct an incomplete-complete point cloud pair for deep learning model training and evaluation. The specific implementation idea is to divide the point cloud space P0 into three areas: inner ring, annular ring, and outer ring to simulate the loss of the annular area inside the hole, and combine it with the preset angle interval to simulate the loss of the back-to-camera surface structure caused by the camera shooting angle. Then, the cropping strategy is dynamically adjusted according to the batch index to achieve structure-aware cropping of the point cloud data. Based on the complete point cloud, 1536 points are cropped to obtain an incomplete point cloud P1 containing only 2560 points. (3) The FPS algorithm is used to downsample the incomplete point cloud P1 to a medium-resolution point cloud P2 with only 1024 points and a coarse-resolution point cloud P3 with only 512 points. A three-level resolution point cloud list [P1, P2, P3] is constructed and input into the DGCNN-Trans feature extractor to extract geometric features F1, F2, and F3. Subsequently, the multi-scale information is integrated through feature splicing and an MLP layer to finally generate a 2048-dimensional global feature vector F containing geometric details and structural semantics. global ; (4) The DGCNN-Trans feature extractor is based on the DGCNN dynamic graph convolution framework. First, the input point cloud data P i Extract geometric features F of different scales through four layers of Edge Conv operation i1 、F i2 、F i3 、F i4 , then the geometric features of different scales are spliced and the dimension order is adjusted to obtain F cat , F cat Input Transformer Encoder module to obtain the global feature F between points trans , use the 5th layer Edge Conv to perform global feature F trans Perform another side convolution operation to obtain the high-dimensional feature F i5 ,Then, the high-dimensional feature F5 is aggregated through global maximum pooling and average pooling to ,aggregate spatial dimension information and generate compact global feature vectors F1, F2, ,F3; (5) Using the pyramid point fractal generator: the 2048-dimensional global feature F global Gradually reduce the dimension to F through a three-layer fully connected network 1024 、F 512 、F 256 , construct multi-level point features. Then, use the lowest dimensional feature F 256 Generate 64 reference points as the coarse-grained point cloud skeleton P coarse ; Then based on the middle-level feature F 512 , calculate the geometric offset for each reference point and generate 2 local points to form a medium-grained point cloud P containing 128 points medium ; Finally, using the highest dimensional feature F 1024 Through a complex feature transformation network, each medium-precision point is expanded into multiple fine points, and finally a high-precision missing point cloud P containing 1536 points is generated. final ; (6) During the model training process, the fine point cloud chamfer distance loss L is comprehensively utilized CD , key point level supervision loss α1L key1 +α2L key2 , rejection loss λ1L rep , uniformity loss λ2L uni , normal vector consistency loss λ3L norm The multi-objective joint loss function is used to guide the training process of the point cloud completion network to optimize the point cloud generation direction.

2. The target point cloud completion method for a battery-swapping robot based on dynamic graph convolution according to claim 1 is characterized in that: The method for preparing the structurally complete target fastener point cloud FastComp3D dataset described in step (1) is as follows: (1-1) Use a depth camera to shoot a video around the target fastener and convert it into a continuous frame image; (1-2) In the WORKFLOW panel of the software, click "Folder" to import the continuous frame image folder. Click "Start" to start the automatic reconstruction process, including steps such as point cloud generation, map generation, and model reconstruction. In the "TOOLS" panel, use the "Lasso / Rect / Box" selection tool to select and delete unnecessary background. (1-3) Convert the generated obj format model file into a point cloud file in ply or txt format; (1-4) The rigid transformation method in data enhancement sets the scaling factor range to 0.8-1.2 times, the translation range to ±10% of the size range, and the rotation angle to ±180° for rigid transformation; the non-rigid transformation method randomly generates 5 control points in the point cloud space, sets the deformation degree to 0.1, and then assigns a random displacement to each control point. Then, the radial basis function RBF is used for TPS interpolation to calculate the displacement of all points. Finally, the calculated displacement is applied to the original point cloud to obtain the deformation result; the data expansion method collects 3D models of similar bolted joints from public resources and converts them into point cloud data format; In addition, Gaussian noise is used in the enhancement process to simulate the errors in the actual scanning process; (1-5) Repeat the above operations multiple times using the target fastener after initial 3D reconstruction and the target fastener after data enhancement.

3. The target point cloud completion method for a battery-swapping robot based on dynamic graph convolution according to claim 1 is characterized in that: The clipping strategy in step (2) includes the following steps: (2-1) Divide the point cloud space into inner ring area R according to spherical coordinates inner , middle annular area R ring and the outer ring area R outer , cropping priority from R ring Area deletion points to simulate the missing structure at the edge of the target fastener hole; (2-2) Combined with the spherical coordinate azimuth angle θ∈[θ1,θ2], the area of the point cloud facing away from the camera is culled within a specific angle range to simulate the visibility loss caused by different shooting angles; (2-3) According to the current sample number i∈[1,N] and batch index b i , cropping masks are dynamically generated to achieve cyclic combination of multi-modal cropping strategies.

4. The target point cloud completion method for a battery-swapping robot based on dynamic graph convolution according to claim 1 is characterized in that: The specific implementation of FPS downsampling, feature splicing and fusion in step (3) is as follows: (3-1) For the generated incomplete point cloud P1, randomly select an initial point p0∈P1 and add it to the sampling point set S={p0}, and calculate each point p i ∈P1 is the minimum distance d from the current sampling point set S i , select d i The largest point p next As the next sampling point, and p next Add the sampling point set S and iterate t=1, 2, ..., K-1 times. When the size of the sampling point set S reaches the number of sampling points K, the algorithm terminates. (3-2) For each resolution of the point cloud, input it into a DGCNN-Trans feature extractor, and concatenate its output features F1, F2, and F3 using the Concat operation in PyTorch to form a multi-scale feature F; (3-3) Feature F is compressed by one-dimensional convolution Conv1D, batch normalization Batch Normalization and ReLU activation function to obtain the global feature vector F global .

5. The target point cloud completion method for a battery-swapping robot based on dynamic graph convolution according to claim 1 is characterized in that: The specific implementation of the feature extractor in step (4) is: (4-1)Edge Conv: For the input point cloud P i Or upper-level feature F i=1,2,3 , use KNN algorithm for each point p i Calculate its K nearest neighbor points Use the nearest neighbor point set to build a local area graph, and then use the directed edge v ij Calculate each center point p i The K nearest neighbor edge feature e ij ; For each center point, its K edge features e ij Aggregate through the maximum pooling operation to obtain the fusion feature F K : (4-2) Continuously perform four layers of Edge Conv edge convolution to obtain geometric features F of different scales i1 、F i2 、F i3 、F i4 , where each layer of Edge Conv will recalculate the K nearest neighbor point set based on the new local relationship, so that the network can dynamically adjust the connection relationship between points in the feature space; then the geometric features of different scales are spliced and the dimension order is adjusted to obtain F cat ; (4-3) Transformer encoder: The fused multi-scale geometric features F cat As input, use multi-head self-attention to perform linear transformation on it to obtain query Q, key K and value V; Then the attention weight score A is calculated by scaling the dot product, and the attention weight A is weighted summed with the value vector V to obtain the feature F att ; Then the original F cat With F att After achieving residual connection, normalization is performed to obtain feature F norm ; Finally, it is input into the two-layer fully connected network to obtain the globally enhanced feature representation F fin .

6. The target point cloud completion method for a battery-swapping robot based on dynamic graph convolution according to claim 1 is characterized in that: The specific implementation of the pyramid point fractal generator in step (5) is: (5-1) The 2048-dimensional global feature F global Gradually reduce the dimension to F through a three-layer fully connected network 1024 、F 512 、F 256 , construct multi-level point features; (5-2) Then, using the lowest dimensional feature F 256 Generate 64 reference points as the coarse-grained point cloud skeleton P coarse ; Then based on the middle-level feature F 512 , calculate the geometric offset for each reference point and generate 2 local points to form a medium-grained point cloud P containing 128 points medium ; Finally, using the highest dimensional feature F 1024 Through a complex feature transformation network, each medium-precision point is expanded into multiple fine points, and finally a high-precision missing point cloud P containing 1536 points is generated. final .

7. The target point cloud completion method for a battery-swapping robot based on dynamic graph convolution according to claim 1 is characterized in that: The method for constructing the multi-objective joint loss function in step (6) is: (6-1) The fine point cloud chamfer distance loss L CD , key point level supervision loss L key1 、L key2 , rejection loss L rep , uniformity loss L uni , normal vector consistency loss L norm Add the weights of α1, α2, λ1, λ2, and λ3 to get the final overall loss L for training total .

Citation Information

Patent Citations

  • Multi-type battery pack visual positioning method and system device for battery replacement robot

    CN115272655A

  • Power battery pose estimation method based on machine vision point cloud segmentation

    CN117635699A

  • Position and posture estimation method for battery replacement robot based on point cloud component segmentation and registration

    CN119251305A

  • Battery replacement robot navigation pose measurement method based on three-dimensional visual features

    CN119594979A

Cited By

  • High-fidelity point cloud completion method and system based on double-path attention and fractal structure

    CN120655840A

  • Steel bar binding method, device and equipment and storage medium

    CN121861391A

  • A reinforcing bar binding method, device, equipment and storage medium

    CN121861391B