Dynamic plant area unmanned logistics vehicle rapid relocation method based on multi-scale map

Through a deep learning network based on multi-scale map processing point clouds, filtering out dynamic elements, combining GOD map descriptors and Kalman filtering, the relocation accuracy and real-time problems of unmanned logistics vehicles in dynamic factories are solved, and efficient and fast relocation is achieved.

CN120274743APending Publication Date: 2025-07-08SUZHOU DACHENGYUNHE INTELLIGENT TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510298375.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the dynamic factory area, the relocation technology of the existing unmanned logistics vehicle has problems such as poor real-time performance and low map retrieval accuracy, especially due to GNSS signal occlusion and dynamic element interference, which affects the positioning effect and accuracy.

Method used

Using a multi-scale map-based method, the point cloud is processed through a deep learning network, the front view and bird's eye view are generated, movable elements are filtered out, and the GOD map descriptor is combined to match the global and local maps, and the pose is adjusted using Kalman filtering to achieve rapid repositioning.

Benefits of technology

The repositioning accuracy and real-time performance of the unmanned logistics vehicle are improved, dynamic target interference is eliminated, the positioning accuracy is improved by 50%, and the training efficiency is high.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120274743A_ABST
    Figure CN120274743A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale map-based rapid relocation method for an unmanned logistics vehicle in a dynamic factory, and the method comprises the following steps: S1, filtering point cloud dynamic elements based on binary semantic segmentation of a deep convolutional neural network, S2, processing static point cloud, and S3, carrying out the rapid relocation of the unmanned logistics vehicle based on global and local map search. According to the method, movable elements in the cloud are filtered based on the deep convolutional neural network, and interference of a dynamic target on subsequent vehicle positioning is eliminated. The method is also combined with a GOD map descriptor to process the static point cloud, so that the subsequent positioning step is clear and rapid. According to the method, local map matching and global map matching are combined, and rapid relocation of the unmanned vehicle is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of driverless, and particularly to a method for quickly repositioning a dynamic factory area unmanned logistics vehicle based on a multi-scale map. Background Art

[0002] When a dynamic factory area unmanned logistics vehicle performs vehicle positioning, due to the occlusion of GNSS signals, the positioning information may be lost. Therefore, in this scenario, the unmanned logistics vehicle often uses a LiDAR sensor to collect environmental information.

[0003] However, most of the current repositioning technologies have the problem of poor real-time performance; in addition, due to the interference of dynamic elements (such as moving vehicles) in the factory area on point cloud matching, the map retrieval accuracy is poor. These problems affect the effect and accuracy of repositioning and have become urgent problems to be solved in the industry. Summary of the Invention

[0004] In order to solve the problems mentioned in the background art, the present invention provides a method for quickly repositioning a dynamic factory area unmanned logistics vehicle with good real-time performance and relatively high repositioning accuracy.

[0005] For this purpose, the present invention adopts the following technical solutions:

[0006] A method for quickly repositioning a dynamic factory area unmanned logistics vehicle based on a multi-scale map, comprising the following steps:

[0007] S1, converting the original factory area 3D point cloud into the expression forms of a front view and a bird's-eye view, and generating movable element markers therein; then constructing a deep learning network architecture, training it through a weighted cross-entropy loss function, and then filtering the movable element markers in the 3D point cloud to obtain a static point cloud; the original factory area 3D point cloud is obtained by the unmanned logistics vehicle.

[0008] S2, processing the static point cloud, including:

[0009] S2.1, obtaining a binary descriptor, including the following steps:

[0010] S2.1.1, downsampling the static point cloud and converting it into a bird's-eye view representation, and then performing circular division and sector division in the horizontal space with the point cloud origin of the static point cloud as the center to obtain the divided point cloud.

[0011] S2.1.2, converting the divided point cloud in S2.1.1 into a descriptor f = {f r,s}; where r is the number of rings after the circular division of the point cloud, s is the number of sectors obtained by the sector division of the point cloud, and the position of the sub-descriptor f r,s is determined by the number of rings and the number of sectors.

[0012] S2.1.3, Assign a binary value to each descriptor f according to the occupancy distribution of the factory area point cloud to obtain a binary descriptor, f r,s = 0 indicates that the position (r, s) is unoccupied, f r,s = 1 indicates that the position (r, s) is occupied;

[0013] S2.2, Construct the GOD map descriptor T of the point cloud map d ;

[0014] S3, Fast relocalization of the unmanned logistics vehicle based on global and local map search: First, determine the initial position of the unmanned logistics vehicle through global and local map search, and then filter the positioning error by using the Kalman filter combined with the GOD map descriptor T d and the adaptive pose adjustment to obtain the final pose of the unmanned logistics vehicle.

[0015] Step S1 includes the following sub-steps:

[0016] S1.1, Generate the front view

[0017] Arrange the original 3D point cloud of the factory area in the form of a 2D array, and convert the point cloud from the Euclidean coordinate system to the spherical coordinate system to obtain the front view where H represents the height of the spherical image, W represents the width of the spherical image, and C is the number of channels of the spherical image;

[0018] S1.2, Generate the bird's-eye view

[0019] In the top-down state, generate a two-dimensional grid with a resolution of S meters in the area of X×Y around the body of the unmanned logistics vehicle, and convert the coordinates of each point in the original 3D point cloud of the factory area into positions on the grid, and then obtain the bird's-eye view I after data processing and normalization BE , where H ′ is the length of the bird's-eye view, W' is the width of the bird's-eye view, and C ′ is the feature of the bird's-eye view;

[0020] S1.3, Generate movable element markers: Generate movable element markers by using the 3D guiding bounding boxes in the open-source factory area scene dataset. Among them, mark the movable elements in the front view as Y FR , and mark the movable elements in the bird's-eye view as Y BE , which is used to evaluate the movable elements in the real factory area scene;

[0021] S1.4, Establish a network architecture for movable element segmentation: Establish a contraction-expansion network architecture, where the contraction part in the contraction-expansion network architecture is used to reduce the spatial resolution of the front view and the bird's-eye view and increase the depth of features, and the expansion part in the contraction-expansion network architecture restores the spatial resolution of the front view and the bird's-eye view through upsampling and convolution operations, while reducing the depth of features; the contraction-expansion network architecture includes a front view network architecture and a bird's-eye view network architecture;

[0022] S1.5, Take the front view I FR , the bird's-eye view I BE and the movable elements marked as Y FR and Y Be obtained in S1.3 as inputs, and perform training iterations on the network architecture of S1.4 until the final loss function converges to obtain a movable element segmenter. The movable element segmenter is used to filter out movable objects in the front view and the bird's-eye view generated by the current point cloud, and then map the front view and the bird's-eye view segmented by the movable element segmenter back to the representational form of the 3D point cloud to obtain the point cloud data after filtering out movable objects; the final loss function is determined by a weighted cross-entropy loss function;

[0023] S1.6, Cluster the point cloud data obtained in S1.5, filter out the errors caused by movable objects, and obtain a static point cloud.

[0024] Step S2.2 includes the following sub-steps:

[0025] S2.2.1, Create candidate points:

[0026] Divide the original plant area 3D point cloud in S1 into grids with equal side lengths in the horizontal space. The center points of the grids are used as initial candidate points, and the side length of the grid is set to r c ; then save all the initial candidate points in the form of a kd-tree, and the value of r c is determined according to the actual situation; finally, the initial candidate points where the unmanned logistics vehicle is likely to pass are used as candidate points;

[0027] S2.2.2, Create the GOD map descriptor T d :

[0028] Stack the sub-descriptors f i,j at all the candidate points to form a matrix. (i, j) represents the i-th row and j-th column of the grid divided by the map. f i,j = 0 indicates that the grid (i, j) is unoccupied, and f i,j= 1 indicates that the grid (i,j) is occupied, and then save this matrix in the form of a kd - tree to obtain the GOD map descriptor T d .

[0029] S3 includes the following sub - steps:

[0030] S3.1, Global map matching:

[0031] Set the transformation matrix between the real - time point cloud data and the original plant area 3D point cloud in S1 as T. Map the binary descriptor obtained in S2.1.3 into the point cloud according to T, and then calculate the loss function between the mapped point cloud descriptor and the GOD map descriptor T d in S2.2.2. Through iterative transformation, obtain the optimal map descriptor f that matches when the loss function is minimized * and the optimal transformation T * ;

[0032] S3.2, Local map matching:

[0033] First, match the optimal map descriptor at the previous moment and the optimal map descriptor at the current moment to reduce candidate points and obtain the optimal candidate point set f best ; where t represents the moment of obtaining the optimal map descriptor f * ;

[0034] Then, obtain the final vehicle pose by weighted averaging the optimal candidate points in the optimal candidate point set f best , and the weight of the weighting is inversely proportional to the loss function d(·).

[0035] Preferably, in S1.2, X×Y = 50 m×60 m and S = 0.1 m.

[0036] S1.4 includes the following sub - steps:

[0037] S1.4.1, Establish a front - view network architecture, which includes an initial convolutional layer, a contraction part, an expansion part, and skip connections, where:

[0038] The initial convolutional layer uses a stride ratio of 1:2 in its first convolutional layer to quickly reduce the size of the input front - view;

[0039] The contraction part further reduces the resolution of the front - view by performing two contraction operations to enable the view network to capture extensive context information;

[0040] The expansion part restores the resolution of the front - view through a de - convolutional layer;

[0041] The skip connection is to add a skip connection between the contraction part and the expansion part, which is used to realize the combination of high-level semantic information and low-level detailed information;

[0042] S1.4.2. Establish the bird's-eye view network architecture:

[0043] Add Fire Modules to the front view network architecture established in S1.4.1 to obtain the bird's-eye view network architecture.

[0044] In step S1.5, the weighted cross-entropy loss function is specifically shown as follows:

[0045]

[0046] where I n represents the front view or the bird's-eye view, Y n is the corresponding movable element label in the front view or the bird's-eye view, ω is the class imbalance weighting function, which is determined according to the actual situation, Id is the indexing function, and the Id is used to select whether it corresponds to a movable object or an immovable object. represents the feature extraction function;

[0047] The final loss function of the front view network and the bird's-eye view network is specifically shown as follows:

[0048]

[0049] where λ m is the regularization weight of each resolution loss, 0 ≤ λ m ≤ 1, and m is the resolution level for calculating the loss function.

[0050] Preferably, when calculating the front view, m = 3; when calculating the bird's-eye view, m = 5; where the regularization weight λ m of the bird's-eye view is set to twice that of the front view.

[0051] In step S2.1.1, the voxel grid is used to downsample the static point cloud obtained in S1, and then the downsampled result is converted into a bird's-eye view representation. In the polar coordinate system, the horizontal space is circularly divided and sectorially divided with the point cloud origin as the center to obtain the divided point cloud.

[0052] In S3.1, global map matching is performed through the following formula:

[0053]

[0054] where Denote the point cloud descriptor obtained by mapping the binary descriptor obtained in S2.1.3 to the 3D point cloud through T; d(·) is the loss function for global map matching, and denote the real-time GOD map descriptor, denote the descriptor in the said T d ; f * is the map descriptor with the lowest loss;

[0055] In step S3.2, local map matching is still calculated through the loss function d(·), specifically:

[0056]

[0057] wherein, f best is the optimal candidate point set in the point cloud, is the optimal map descriptor at the previous moment, is the optimal map descriptor at the current moment, and the subscript t represents the moment of obtaining the optimal map descriptor f * ; The Kalman filter is used to improve the robustness of online positioning.

[0058] Compared with the prior art, the present invention has the following beneficial effects:

[0059] 1. The method of the present invention filters out the movable elements in the point cloud based on a deep convolutional neural network, eliminating the interference of dynamic targets on subsequent vehicle positioning; a relatively simple network architecture is also used in step S1, improving the efficiency of training and detection; since the present invention only needs to perform binary semantic segmentation of moving or static classes, the loss function design is relatively simple and the efficiency of semantic segmentation is high.

[0060] 2. The method of the present invention combines the GOD map descriptor to process the static point cloud, enabling clear and rapid global map matching and local map matching in subsequent steps.

[0061] 3. The method of the present invention combines local map matching and global map matching on the basis of removing the interference of dynamic targets, realizing rapid repositioning of driverless vehicles. The positioning accuracy after removing the interference of dynamic targets is improved by 50% compared with the positioning accuracy without removal. Brief Description of the Drawings

[0062] Figure 1 is the flow chart of the present invention;

[0063] Figure 2 is the segmentation effect diagram of moving and static targets obtained by using the method of step S1 of the present invention;

[0064] Figure 3 is the positioning trajectory diagram of Network 1 in the comparative example;

[0065] Figure 4 It is the positioning trajectory diagram of Network 2 in the comparative example;

[0066] Figure 5 It is the comparison diagram of the positioning errors of Network 1 and Network 2 in the comparative example. Specific implementation manners

[0067] The following further elaborates in detail on the method for rapid repositioning of an unmanned logistics vehicle in a dynamic factory area based on a multi-scale map of the present invention in conjunction with the accompanying drawings and embodiments.

[0068] As Figure 1 shown, the method for rapid repositioning of an unmanned logistics vehicle in a dynamic factory area of the present invention includes the following steps:

[0069] S1. Filter the dynamic elements of the point cloud based on binary semantic segmentation of a deep convolutional neural network. Specifically, convert the original 3D point cloud of the factory area into the expression forms of a front view and a bird's-eye view, and generate movable element markers therein; then construct a deep learning network architecture, train it through a weighted cross-entropy loss function, and then filter the movable element markers in the 3D point cloud to obtain a static point cloud; the original 3D point cloud of the factory area is obtained by an unmanned logistics vehicle, and specifically includes the following sub-steps:

[0070] S1.1. Generate a front view

[0071] Arrange the original 3D point cloud of the factory area obtained by the lidar mounted on the body of the unmanned logistics vehicle in the form of a 2D array, and convert the point cloud from the Euclidean coordinate system to the spherical coordinate system to obtain a front view where H represents the height of the spherical image, W represents the width of the spherical image, and these two parameters are obtained by discretizing the point cloud by formulating the azimuth step based on the parameters of the lidar; C is the number of image channels of the spherical image, and one image channel is used to save the range value of each point in the point cloud, and the other image channel is used to save the reflectivity of each point in the point cloud.

[0072] S1.2. Generate a bird's-eye view

[0073] Generate a two-dimensional grid with a resolution of S meters in the area of X×Y around the vehicle body, and convert the coordinates of each point in the point cloud into grid positions. In an embodiment of the present invention, X×Y = 50 m × 60 m, S = 0.1 m. After data processing and normalization, a bird's-eye view is obtained where H ′ is the length of the bird's-eye view, W′ is the width of the bird's-eye view, C ′Features of the bird's-eye view, including: binary occupancy terms with zero values, absolute occupancy terms, representing the total number of points in the cell, the average reflectivity value of each point, the average height value of the projection elements, the minimum height value of the projection elements, and the maximum height value of the projection elements.

[0074] S1.3, Generate movable element markers: The movable element markers are generated by using the 3D oriented bounding boxes of the open-source plant area scene dataset. Among them, the movable elements in the front view are marked as Y FR , and the movable elements in the bird's-eye view are marked as Y BE , which are used to evaluate the movable elements in the real plant area scene;

[0075] S1.4, Establish a network architecture: For the front view and bird's-eye view proposed in the above steps, the present invention proposes a contraction-expansion architecture. The contraction part is used to reduce the spatial resolution of the feature map and increase the depth of the features, while the expansion part restores the spatial resolution of the feature map through upsampling and convolution operations, and at the same time reduces the depth of the features. This way of contracting and then expanding effectively realizes the extraction of point cloud features. In addition, by combining skip connections, the network can retain low-level detail information while learning high-level features. Generally speaking, the contraction-expansion architecture constructs a powerful feature learning framework by combining the contraction part, the expansion part, and skip connections. It specifically includes the following sub-steps:

[0076] S1.4.1, Establish the front view network architecture, including an initial convolutional layer, a contraction part, an expansion part, and skip connections, where:

[0077] Initial convolutional layer: In the first convolutional layer of the front view network architecture, a stride ratio of 1:2 is used to quickly reduce the size of the input feature map, thereby reducing the computational complexity and extracting more abstract feature representations.

[0078] Contraction part: By performing two contraction operations to further reduce the feature map resolution, enabling the network to capture a wider range of context information.

[0079] Expansion part: The resolution of the feature map is restored through a deconvolution layer to facilitate subsequent fine segmentation.

[0080] Skip connections: By adding skip connections between the contraction and expansion parts, the combination of high-level semantic information and low-level detail information is achieved, thereby retaining more edge information.

[0081] S1.4.2, Establish the bird's-eye view network architecture:

[0082] Since the moving objects in the bird's-eye view occupy a relatively small proportion of each frame's grid, for the bird's-eye view network architecture, a Fire Module is added on the basis of the front view network architecture established in S1.4.1 to keep the number of network parameters small while still providing high accuracy. Specifically, the bird's-eye view network architecture first uses a convolutional layer to reduce the number of feature maps, then two new convolutional layers with different filter sizes are applied in parallel, and finally the different results are connected to obtain robust features with local and rich context-aware information.

[0083] The Fire Module can be found in the reference B. Wu, X. Zhou, S. Zhao, X. Yue, and K. Keutzer, “Squeezesegv2: Improved model structure and unsupervised domain adaptation for road object segmentation from a lidar point cloud,” in IEEE International Conference on Robotics and Automation (ICRA), 2019.

[0084] S1.5, Implement model training based on the weighted cross-entropy loss function: Use the weighted cross-entropy loss function Train the front view and the bird's-eye view to address the class imbalance problem, defined as:

[0085]

[0086] where, I n represents the front view or the bird's-eye view, Y n is the label of the corresponding movable element in the front view or the bird's-eye view. At the same time, calculate the class imbalance weighting function ω as the inverse ratio between the vehicle class and the background class in the training set samples. ω is determined according to the actual situation, and Id is an index function used to select the prediction probability related to the expected true class, that is, whether it corresponds to a movable object or an immovable object. represents the feature extraction function. Through the weighted cross-entropy loss function The model can better balance the prediction performance between movable objects and immovable objects, especially when the class distribution of the background and movable objects is severely unbalanced, improving the detection ability of movable objects. To guide the network to obtain the correct results faster, the present invention additionally implements a multi-scale solution to the segmentation problem by introducing intermediate predictions and losses of different resolutions, that is, valuable gradients are inserted at the intermediate levels, and the network can be gradually optimized at each layer, avoiding getting feedback only at the last layer, thereby improving the training efficiency and the results. Therefore, this step independently calculates the final loss functions of the front view network and the bird's-eye view network.

[0087]

[0088] where λ m is the regularization weight of each resolution loss, 0 ≤ λ m ≤ 1, m is the resolution level for calculating the loss function. In an embodiment of the present invention, the front view network calculates the losses of 3 resolution levels, that is, when calculating the front view, m = 3; the bird's-eye view network calculates the losses of 5 resolution levels, that is, when calculating the bird's-eye view, m = 5; the regularization weight of the bird's-eye view is set to twice that of the front view.

[0089] In the model training stage, the present invention uses the KITTI tracking dataset to train the network architecture, from which the ground truth of movable objects can be obtained. In order to retain the geometric attributes of the driving scene, the dataset is augmented by horizontal flipping in the front view and vertical flipping in the bird's-eye view with a probability of 50%. During the training process, Adam optimization with parameters β1 = 0.9 and β2 = 0.999 is used. Using the front view I FR and the bird's-eye view I BE and the Y FR and Y BE obtained from S1.5 are iterated in the final loss function until the final loss function converges to obtain a movable element segmenter, which is used to filter out the movable objects in the front view and bird's-eye view generated by the current point cloud, and then map the front view and bird's-eye view segmented by the movable element segmenter back to the 3D point cloud representation form to obtain the point cloud data with movable objects filtered out.

[0090] In an embodiment of the present invention, each network is independently trained on an Nvidia 3080Ti GPU, and the batch size for 400,000 iterations is set to 10. At the same time, from 10 -3Starting from the learning rate, after the first 150,000 iterations, it is halved every 50,000 iterations. Considering the ratio relationship between classes on each domain, the present invention sets the regularizer ω to 25 in the front view and 1000 in the bird's-eye view.

[0091] S1.6. Cluster the point cloud data obtained in S1.6, filter the projection errors caused by movable objects, and obtain static point clouds, specifically:

[0092] Based on the point clouds of the front view and the bird's-eye view after filtering out movable objects obtained in S1.5, cluster the predicted points in the two types respectively, and discard the interference clusters according to the predicted average probability. Among them, when performing the clustering operation, the contribution weight of the bird's-eye view samples is twice that of the front view samples to reduce the influence of noise therein. After obtaining the final predicted clustering, filter the original input LiDAR data by eliminating all points within a radius of 10 cm from the predicted points, thereby limiting the influence of possible projection errors.

[0093] S2. Process the static point cloud, including the following steps:

[0094] S2.1. Obtain binary descriptors:

[0095] S2.1.1. Downsample the static point cloud and convert it into a bird's-eye view representation, and then perform circular division and sector division centered on the origin of the point cloud in the horizontal space to obtain the divided point cloud;

[0096] Specifically, the downsampling is to downsample the static point cloud obtained in S1 using a voxel grid, that is, select one or more points in each voxel to represent the area, thereby reducing the number of points, effectively maintaining the spatial structure of the point cloud, and reducing the computational burden at the same time.

[0097] Then convert the result of the downsampling into a bird's-eye view representation. For the convenience of subsequent feature processing, circular division and sector division are performed centered on the origin of the point cloud in the polar coordinate system, making subsequent processing and feature extraction more efficient, and obtaining the divided point cloud.

[0098] S2.1.2. Convert the divided point cloud into a descriptor f = {f r,s}:

[0099] Among them, r and s are the ring number and sector number where the sub-descriptor f r,s is located.

[0100] Let N(r, s) be the set of points in the overlapping area of the rth ring and the sth sector in the descriptor f. Since the descriptor f proposed by the present invention focuses on the occupancy of the vertical structure, the set of points does not include points on the ground and the ceiling (factory roof).

[0101] S2.1.3, assign a binary value to each descriptor using the following encoding function:

[0102]

[0103] where N(·) is a function that calculates the number of points within a descriptor. N min and N max are determined according to the actual situation, 0 represents that this area is an occupied area, and 1 represents that this area is an unoccupied area.

[0104] Obtain a binary descriptor for judging the distribution of occupied areas within the static point cloud in S1.

[0105] S2.2, Data processing of the point cloud map:

[0106] In this step, the process of generating the offline map data is described, that is, the map candidate points and map descriptors in the converted 2D grid map are described. In this process, by constructing an offline kd-tree, the search for map candidate points and map descriptors is accelerated. Specifically:

[0107] S2.2.1, Create candidate points:

[0108] Divide the original plant 3D point cloud obtained in S1 into grids with equal side lengths in the horizontal space. The center point of the grid is used as the initial candidate point, and the side length of the grid is set to r c , save all the initial candidate points in the form of a kd-tree, and the value of the r c is determined according to the actual situation; finally, the initial candidate points that the unmanned logistics vehicle has a probability of passing through are used as candidate points;

[0109] S2.2.2, Create a Grid Occupied Description (GOD):

[0110] Stack the sub-descriptors f i,j at all candidate points to form a matrix. (i, j) represents the i-th row and j-th column of the grid divided by the map. f i,j =0 indicates that the grid (i, j) is unoccupied, f i,j =1 indicates that the grid (i, j) is occupied, and then construct this matrix in the form of a kd-tree to obtain the GOD map descriptor T d ; the GOD map descriptor T d can be used for fast map indexing to achieve vehicle positioning;

[0111] S3, Fast repositioning of the unmanned logistics vehicle based on global and local map search. First, determine the initial position of the vehicle through global and local map search, and then by using the combined GOD map descriptor Td Kalman filtering and adaptive pose adjustment are used to filter out errors and ensure high-precision positioning. The specific steps are as follows:

[0112] S3.1, Global map matching: Global map matching uses the binary descriptor obtained in S2.1.3 to find the best matching descriptor corresponding to it from the GOD map descriptor T d and calculate the pose of the vehicle. Therefore, global map matching can be formalized as:

[0113]

[0114] where T is the transformation matrix between the real-time point cloud data and the original factory area 3D point cloud; d(·) is the loss function of global map matching, and represents the point cloud descriptor obtained by mapping the binary descriptor obtained in S2.1.3 to the 3D point cloud through T; represents the local map descriptor of the GOD map descriptor T d at the grid i, j; f * is the optimal map descriptor obtained by matching when the loss function is the smallest.

[0115] Global map matching will find the best match of the online descriptor in the map descriptor every time. The map data used for global map matching is created offline, while the search process is online.

[0116] S3.2, Local map matching:

[0117] Since the optimal map descriptor f * obtained by global map matching still contains several candidate points, local map matching is needed to reduce the candidate points and obtain the optimal candidate point set f best , improving the robustness of online positioning. Match the optimal map descriptor at the previous moment and the optimal map descriptor at the current moment. Local map matching is still calculated through the loss function d(·), specifically:

[0118]

[0119] where, is the optimal map descriptor at the previous moment, is the optimal map descriptor at the current moment, and the subscript t here represents the moment when the optimal map descriptor f * is obtained.

[0120] For the optimal candidate point set f best after local map matching, and then through the optimal candidate point set fbest The optimal candidate points in it are weighted and averaged to obtain the final vehicle pose. The weight of the weighting is inversely proportional to the loss function d(·), and the online positioning robustness is improved through Kalman filtering.

[0121] Embodiment 1

[0122] Figure 2 It is the dynamic and static target segmentation effect diagram obtained by using the method of step S1 of the present invention. Among them, A is the class of point clouds of movable targets, and other point clouds are the class of point clouds of static targets. This frame of point cloud is collected when the vehicle enters the factory entrance.

[0123] After removing the point clouds of the movable class, the point clouds of the static class can achieve a more accurate point-to-point matching relationship with the prior static map, so as to obtain a more accurate positioning result. Compared with the fast repositioning accuracy of the unmanned logistics vehicle based on global and local map search without removing dynamic targets, which is about 0.26m, the fast repositioning accuracy of the unmanned logistics vehicle based on global and local map search after removing dynamic targets is about 0.13m, an increase of 50%.

[0124] Comparative Example 1

[0125] For the comparison of the positioning effects, the method of the present invention is taken as Network 1, and the method of the present invention without the step of removing dynamic targets is taken as Network 2. As Figure 3 and Figure 4 shown, the positioning trajectory of Network 1 has reduced a large number of miscellaneous points compared with the training result of Network 2, reduced the interference of dynamic targets, and improved the positioning robustness.

[0126] Regarding the comparison of the positioning errors after training of Network 1 and Network 2, as Figure 5 shown, Network 2 has a large error due to the interference of dynamic targets at some moments. However, since Network 1 removes dynamic targets, the positioning accuracy has been significantly improved.

[0127] Comparative Example 2

[0128] For the comparison of the training efficiency, the existing three-dimensional target detection method for autonomous driving vehicles (Zhao Yunche. Research on Three-Dimensional Target Detection of Autonomous Driving Vehicles Based on Image and Point Cloud Fusion [D]. Shijiazhuang Tiedao University, 2023. DOI: 10.27334 / d.cnki.gstdy.2023.000114) is taken as Network 3, and then Network 1 and Network 3 are independently trained simultaneously on an Nvidia 3080Ti GPU; for each network, the batch size at 400,000 iterations is set to 10, and at the same time, from 10 -3Starting from the learning rate, after the first 150,000 iterations, it is halved every 50,000 iterations. The decay rate is 50,000 and the decay coefficient is 0.5. In Network Three, the decay rate is 30,000 and the decay coefficient is 0.8, and the training efficiency is much lower than that of Network One.

Claims

1. A method for rapid relocalization of an unmanned logistics vehicle in a dynamic factory area based on a multi-scale map, characterized in that It includes the following steps: S1. Convert the original 3D point cloud of the factory area into the expression forms of the front view and the bird's-eye view, and generate movable element markers therein; then construct a deep learning network architecture, train it through a weighted cross-entropy loss function, and then filter the movable element markers in the 3D point cloud to obtain a static point cloud; The original 3D point cloud of the factory area is obtained by an unmanned logistics vehicle; S2. Process the static point cloud, including: S2.

1. Obtain a binary descriptor, including the following steps: S2.1.

1. Downsample the static point cloud and convert it into a bird's-eye view representation, and then perform circular division and sector division centered on the point cloud origin in the horizontal space to obtain the divided point cloud; S2.1.

2. Convert the point cloud divided in S2.1.1 into a descriptor \(f = \{f r,s \}\); where \(r\) is the number of rings after the point cloud is divided by the ring division, \(s\) is the number of sectors obtained by the sector division of the point cloud, and the position of the sub-descriptor \(f r,s \) is determined by the number of rings and the number of sectors; S2.1.3, assign a binary value to each descriptor f according to the occupancy distribution of the plant area point cloud to obtain a binary descriptor, f r,s = 0 indicates that the position (r, s) is unoccupied, f r,s = 1 indicates that the position (r, s) is occupied; S2.2, construct the GOD map descriptor T of the point cloud map d ; S3, Fast relocalization of the unmanned logistics vehicle based on global and local map search: First, determine the initial position of the unmanned logistics vehicle through global and local map search, and then filter the positioning error and obtain the final pose of the unmanned logistics vehicle by using the Kalman filter combined with the GOD map descriptor T d and adaptive pose adjustment.

2. The dynamic plant unmanned logistics vehicle rapid relocalization method based on a multi-scale map according to claim 1, wherein Step S1 includes the following sub-steps: S1.1, generate a front view Arrange the original 3D point cloud of the factory area in the form of a 2D array, and convert the point cloud from the Euclidean coordinate system to the spherical coordinate system to obtain the front view where H represents the height of the spherical image, W represents the width of the spherical image, and C is the number of channels of the spherical image; S1.2, Generate a bird's-eye view In the top-down view, a two-dimensional grid with a resolution of S meters is generated in the X×Y area around the body of the unmanned logistics vehicle, and the coordinates of each point in the original 3D point cloud of the factory area are converted into positions on the grid. After data processing and normalization, an aerial view I is obtained. BE , where H ′ is the length of the aerial view, W′ is the width of the aerial view, and C ′ is the feature of the aerial view. S1.3, Generate movable element markers: Generate movable element markers by using the 3D oriented bounding boxes in the open-source plant area scene dataset. Among them, the movable elements in the front view are marked as Y FR , and the movable elements in the bird's-eye view are marked as Y BE , which are used to evaluate the movable elements in the real plant area scene; S1.

4. Establish a network architecture for movable element segmentation: Establish a contraction-expansion network architecture, where the contraction part in the contraction-expansion network architecture is used to reduce the spatial resolution of the front view and the bird's-eye view and increase the depth of features, and the expansion part in the contraction-expansion network architecture restores the spatial resolution of the front view and the bird's-eye view through upsampling and convolution operations while reducing the depth of features; the contraction-expansion network architecture includes a front view network architecture and a bird's-eye view network architecture; S1.5, project the front view I FR , the bird's-eye view I BE and the movable elements marked as Y FR and Y BE obtained in S1.3 as inputs, and perform training iterations on the network architecture of S1.4 until the final loss function converges to obtain a movable element segmenter, which is used to filter out movable objects in the front view and bird's-eye view generated from the current point cloud, and then map the front view and bird's-eye view segmented by the movable element segmenter back to the representational form of the 3D point cloud to obtain the point cloud data after filtering out movable objects; the final loss function is determined by a weighted cross-entropy loss function; S1.

6. Cluster the point cloud data obtained in S1.5, filter the errors caused by movable objects, and obtain a static point cloud.

3. The dynamic plant unmanned logistics vehicle rapid repositioning method based on a multi-scale map according to claim 1, characterized in that Step S2.2 includes the following sub-steps: S2.2.

1. Create candidate points: Divide the original 3D point cloud of the factory area in S1 into grids with equal side lengths in the horizontal space. The center point of the grid is used as the initial candidate point, and the side length of the grid is set to r c ; then save all the initial candidate points in the form of a kd-tree, and the value of r c is determined according to the actual situation; finally, the initial candidate points that the unmanned logistics vehicle is likely to pass through are used as candidate points; S2.2.2, Create the GOD map descriptor T d : Stack the sub - descriptors f at all the candidate points i,j to form a matrix. Let (i, j) represent the i - th row and j - th column of the grid divided by the map. When f i,j = 0, it means that the grid (i, j) is unoccupied. When f i,j = 1, it means that the grid (i, j) is occupied. Then save this matrix in the form of a kd - tree to obtain the GOD map descriptor T d .

4. The method for rapid repositioning of an unmanned logistics vehicle in a dynamic factory area based on a multi-scale map according to claim 1, characterized in that, S3 includes the following sub-steps: S3.

1. Global map matching: Set the transformation matrix between the real-time point cloud data and the original plant 3D point cloud in S1 as T. Map the binary descriptors obtained in S2.1.3 to the point cloud according to T, and then calculate the loss function between the mapped point cloud descriptors and the GOD map descriptors T in S2.2.2 d Through iterative transformation, obtain the optimal map descriptor f that matches when the loss function is minimized * and the optimal transformation T * ; S3.

2. Local map matching: First, match the optimal map descriptor at the previous moment and the optimal map descriptor at the current moment to reduce candidate points and obtain the optimal candidate point set f best ; where t represents the moment when the optimal map descriptor f * is obtained; Then, the final vehicle pose is obtained by weighted averaging the optimal candidate points in the optimal candidate point set f best , and the weights for weighting are inversely proportional to the loss function d(·).

5. The method for fast relocating an unmanned logistics vehicle in a dynamic factory area based on a multi-scale map according to claim 2, wherein: In S1.2, X×Y = 50 m×60 m, S = 0.1 m.

6. The method for rapid relocalization of an unmanned logistics vehicle in a dynamic factory area based on a multi-scale map according to claim 2, characterized in that, S1.4 includes the following sub-steps: S1.4.

1. Establish a front view network architecture, and the front view network architecture includes an initial convolutional layer, a contraction part, an expansion part, and a skip connection, where: The initial convolutional layer uses a stride ratio of 1:2 in its first convolutional layer to quickly reduce the size of the input front view; The contraction part further reduces the resolution of the front view by performing two contraction operations to enable the view network to capture extensive context information; The expansion part restores the resolution of the front view through a deconvolutional layer; The skip connection adds a skip connection between the contraction part and the expansion part to realize the combination of high-level semantic information and low-level detail information; S1.4.

2. Establish a bird's-eye view network architecture: Add Fire Modules to the front view network architecture established in S1.4.1 to obtain a bird's-eye view network architecture.

7. According to the method for fast repositioning of an unmanned logistics vehicle in a dynamic factory area based on a multi-scale map as claimed in claim 2, wherein: In step S1.5, the weighted cross-entropy loss function is specifically shown as follows: Among them, I n represents a front view or a bird's-eye view, Y n is the corresponding movable element label in the front view or the bird's-eye view, ω is a class imbalance weighting function, determined according to the actual situation, Id is an index function, and the Id is used to select whether the corresponding object is a movable object or an immovable object, represents a feature extraction function; The final loss functions of the front view network and the bird's-eye view network Specifically, it is shown as the following formula: Among them, λ m is the regularization weight of each resolution loss, 0 ≤ λ m ≤ 1, and m is the resolution level at which the loss function is calculated.

8. The method for rapid relocalization of an unmanned logistics vehicle in a dynamic factory area based on a multi-scale map according to claim 7, wherein: When calculating the front view, m = 3; when calculating the bird's-eye view, m = 5; where the regularization weight λ of the bird's-eye view m is set to be twice that of the front view.

9. According to the method for fast repositioning of an unmanned logistics vehicle in a dynamic factory area based on a multi-scale map as claimed in claim 3, wherein: In step S2.1.1, a voxel grid is used to downsample the static point cloud obtained in S1, and then the downsampled result is converted into a bird's-eye view representation. In the polar coordinate system, the horizontal space is divided into rings and sectors centered on the origin of the point cloud to obtain the divided point cloud.

10. The dynamic plant unmanned logistics vehicle rapid relocalization method based on a multi-scale map according to claim 4, characterized in that: In S3.1, global map matching is performed by the following formula: Among them, denotes mapping the binary descriptor obtained from S2.1.3 by T to the point cloud descriptor of the 3D point cloud; d(·) is the loss function of global map matching, and denotes the real-time GOD map descriptor, denotes the descriptor in the said T d ; f * is the map descriptor with the lowest loss. In step S3.2, local map matching is still calculated through the loss function d(·), specifically: Among them, f best is the optimal candidate point set in the point cloud, is the optimal map descriptor at the previous moment, is the optimal map descriptor at the current moment, and the subscript t represents the moment when the optimal map descriptor f * is obtained; the Kalman filter is used to improve the robustness of online positioning.

Citation Information

Cited By

  • Dynamic multi-access-point smooth switching method and system for multi-source fusion positioning of unmanned logistics vehicle

    CN121140812A

  • Dynamic multi-access point smooth switching method and system for unmanned streamer vehicle multi-source fusion positioning

    CN121140812B

  • Inspection robot SLAM positioning mapping method based on deep learning multi-sensor fusion

    CN121612285A