Multi-sensor fusion AGV dynamic target elimination method and device based on semantic segmentation

Through multi-sensor fusion technology, combined with the advantages of lidar and vision sensors, high-precision positioning and map construction of AGV in dynamic environments is achieved, solving the problem of difficult to guarantee positioning accuracy and stability in the existing technology.

CN119152211BActive Publication Date: 2025-05-13ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411614755.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-13
Publication Date
2025-05-13
Estimated Expiration
2044-11-13

AI Technical Summary

Technical Problem

The positioning accuracy of AGV in a dynamic environment is affected by dynamic target separation methods, especially the visual sensor is susceptible to light and has high computational complexity, which makes it difficult to guarantee positioning accuracy and stability.

Method used

The multi-sensor fusion method is adopted, combining the data of two-dimensional lidar and monocular cameras, and the removal of dynamic targets is achieved through target detection and semantic segmentation technology. Lidar is used to initially detect dynamic targets, and vision sensors are used to assist in confirmation and semantic segmentation to reduce the consumption of computing resources.

Benefits of technology

It improves the positioning accuracy and map construction quality of AGV in a dynamic environment, reduces the problems of misjudgment and excessive removal of feature points, and enhances the stability and real-time nature of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119152211B_ABST
    Figure CN119152211B_ABST
Patent Text Reader

Abstract

Multi-sensor fusion AGV dynamic target rejection method and device based on object detection and semantic segmentation technology, the method comprising: jointly calibrating a monocular camera and a two-dimensional lidar to obtain the internal parameter matrix of the camera and the transformation matrix T from the lidar coordinate system to the monocular camera coordinate system LC ; the laser SLAM enters the laser tracking thread to detect the dynamic threshold of the minimum adjacent point pair distance of the front and rear groups of dynamic point clouds; match the timestamps of the images collected by the monocular camera and the laser point clouds, perform interpolation calculation to obtain visual key frames; project the laser dynamic points onto the image coordinate system of the visual key frame Frame L at the tc moment, and input the lidar dynamic points into the object detection thread; perform object detection on the visual key frame at the tc moment, extract the feature points outside the object detection box for updating the sub-map; perform semantic segmentation in the object detection box area, filter the feature points, screen out the static feature points, and update the grid probability in the map construction stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision technology, and in particular relates to a multi-sensor fusion AGV dynamic target elimination method and device based on semantic segmentation technology. Background Art

[0002] AGV (Automated Guided Vehicle) is an advanced automatic transport tool, a wheeled mobile robot commonly used in the industrial and logistics fields, assisting factories and workshops in unmanned material transportation and improving production efficiency. Currently, AGV has implemented a back-end optimization method from state estimation to map optimization, as well as a SLAM (Simultaneous Localization And Mapping) architecture based on visual technology and precise positioning composed of multi-sensor fusion. The above technologies have achieved accurate positioning and mapping in static environments, but there are still some shortcomings when applied to dynamic environments.

[0003] Traditional visual SLAM has high accuracy and excellent loopback effect, but in dynamic environments, the existence of dynamic feature points will make the constructed map disordered, affecting the positioning accuracy. Therefore, it is particularly important to separate dynamic objects when building maps. The separation methods of dynamic objects are usually divided into three methods: target detection, semantic segmentation, and instance segmentation. The target detection method detects the image captured by the visual sensor, selects the target with the minimum circumscribed rectangular box and removes it, but the static feature points in the box are also removed at the same time, affecting the positioning accuracy. The semantic segmentation method detects and assigns semantics to the image captured by the visual sensor, can segment along the contour of the detected object, remove the feature points of the detected object, and use the remaining feature points for map construction. The instance segmentation method detects and assigns semantics to the image captured by the visual sensor, can segment and track specific objects, dynamically track moving objects, and optimize dynamic feature points and static feature points. The tracking and segmentation of dynamic targets requires the help of clustering methods, and there is a problem of regional mismatching. In addition, both semantic segmentation methods and instance segmentation methods are based on visual SLAM, which are easily affected by ambient lighting and have high computational complexity, and their stability and real-time performance are difficult to guarantee.

[0004] Visual sensors are generally divided into depth cameras, binocular cameras, fisheye cameras, monocular cameras and panoramic cameras. Depth cameras have limited working distance and loss of depth information accuracy beyond a certain range. Binocular cameras have high computational complexity and fixed baseline length. Fisheye cameras have data redundancy, severe distortion and low accuracy in edge areas. Monocular cameras have missing depth information, scale uncertainty and difficulty in initialization. Panoramic cameras have missing depth information and complex stitching.

[0005] Traditional two-dimensional laser SLAM is widely used because it can detect farther scales, is less affected by light, and has fast calculation speed. However, since the feature points are limited to the plane and are sparse, it will cause mismatching when encountering similar environmental structures.

[0006] It can be seen that the fusion of vision and laser sensors can make up for their respective shortcomings. The high-frequency characteristics of the laser radar can be used to detect dynamic targets that may exist in the scene, and then the visual target detection and semantic segmentation methods can be used to remove dynamic targets. A grid map method for phased feature projection is constructed. A visual thread is added based on the framework of laser SLAM. The visual thread is divided into two stages: confirming the dynamic target area and cutting the dynamic target. In the stage of confirming the dynamic target area, the static feature points are projected first, and then the feature points are supplemented in the stage of cutting the dynamic target, so as to improve the running speed and accuracy of SLAM in dynamic environments. In the traditional multi-sensor fusion dynamic object removal method, the features of vision and laser point clouds are extracted. At the same time, the dynamic objects need to be completely cut before the feature points are projected, which reduces the running speed of SLAM in dynamic environments.

[0007] There are at least the following problems in the prior art:

[0008] The SLAM optimization method based on semantic segmentation only uses visual sensors and is easily affected by lighting conditions, which affects the semantic segmentation objects. For the tracking of dynamic targets, it is necessary to first use a clustering method to determine the tracking area. The tracking feature points will be affected by the clustering method, and there is a possibility of mismatching. For the detection of visual targets, certain complex calculations must be performed first. Summary of the invention

[0009] The present invention aims to solve the problem that visual sensors need to perform preliminary complex calculations for target detection and the problem that two-dimensional laser radars detect dynamic points incorrectly, and proposes a multi-sensor fusion AGV dynamic target elimination method and device based on target detection and semantic segmentation technology.

[0010] The present invention fuses the sensor information of a monocular camera, a two-dimensional laser radar, an odometer, and an IMU (Inertial Measurement Unit), and constructs it based on the cartographer and YOLOv8 algorithms. The differential calculation and target detection algorithm of the two-dimensional laser radar are used to mark dynamic objects. The characteristics of the two-dimensional laser radar with small data volume and high-frequency scanning can be used to quickly determine whether there are dynamic targets. The target detection algorithm is used for auxiliary judgment to eliminate the situation where the two-dimensional laser radar mistakenly detects dynamic targets. Only semantic segmentation is performed on the target detection area to reduce the amount of semantic segmentation calculations. The quality of mapping in complex dynamic environments is optimized to improve the positioning accuracy of AGVs, and to solve the problem of misjudgment or excessive removal of feature points in complex environments with poor lighting, slow-moving objects, and dynamic changes.

[0011] To achieve the above objectives, the first aspect of the present invention provides a multi-sensor fusion AGV dynamic target elimination method based on target detection and semantic segmentation technology, and the specific steps are as follows:

[0012] (1) Install the 2D laser radar and monocular camera, calibrate the monocular camera and 2D laser radar together, and obtain the camera's intrinsic parameter matrix and the transformation matrix T from the laser radar coordinate system to the monocular camera coordinate system. LC ;

[0013] (2) Laser SLAM enters the laser tracking thread, pre-processes the laser point cloud, integrates the IMU and odometer posture information to scan and match the acquired two-dimensional laser point cloud, corrects the laser point cloud through scanning and matching, uses the laser radar data frame to perform inter-frame difference calculation, and performs dynamic threshold detection on the minimum neighbor point pair distance of the two sets of dynamic point clouds;

[0014] (3) Match the timestamps of the images captured by the monocular camera and the laser point cloud in the factory, based on the time t to which the dynamic point belongs. l Interpolate the visual data frame to obtain the visual data frame Frame corresponding to time tc L , taking it as the visual keyframe;

[0015] (4) Convert the laser dynamic point into coordinates and project it to the visual key frame Frame at time tc L The image coordinate system is used to input the laser radar dynamic points into the target detection thread;

[0016] (5) Perform target detection on the visual keyframe at time tc, track the target detection frame, use the target detection method to assist in confirming the dynamic object, and extract feature points outside the target detection frame for updating the sub-graph;

[0017] (6) Perform semantic segmentation in the target detection box area, filter the feature points according to the semantic segmentation results, screen out static feature points, and update the grid probability during the map construction stage.

[0018] Furthermore, the specific process of step (1) is as follows:

[0019] Step 1-1: Use the lidar and camera to collect data in the same environment, use calibration objects that can be recognized by both sensors to extract features, project visual features into the lidar coordinate system, and establish a connection between the lidar coordinate system and the camera coordinate system;

[0020] Step 1-2: Solve the optimal rotation matrix and translation vector through the optimization algorithm to obtain the external parameter matrix of the two-dimensional laser radar and monocular camera.

[0021] Step 1-3: In the process of solving the external parameter matrix, the laser radar and the monocular camera extract features of the same calibration object to obtain a data point group P in the laser radar coordinate system. l and the corresponding point group P in the camera coordinate system c , where R LC is the rotation matrix, t is the translation vector, so that:

[0022] (1)

[0023] P c and P l分 Respectively represent the position of the same physical point in the camera coordinate system and the lidar coordinate system.

[0024] Step 1-4: Solve R by minimizing the reprojection error using the Levenberg-Marquardt (LM) method LC Project the points in the camera coordinate system to the image coordinate system and obtain the camera's intrinsic parameter matrix K, such that:

[0025] (2)

[0026] Where Pi is the point coordinate in the image coordinate system.

[0027] Furthermore, the specific process of step (2) is as follows:

[0028] Step 2-1, laser radar scans to obtain point cloud, and performs dynamic point cloud detection by laser radar inter-frame difference method. Scan at regular intervals to obtain a set of point cloud data sets P(t) and P(t+Δt), where Δt is the time interval between two scans. The two-dimensional laser point clouds obtained by the two scans are iteratively registered using the ICP (Iterative Closest Point algorithm).

[0029] Step 2-2: Introduce the current two-dimensional lidar data into the KD tree to improve the retrieval speed of the nearest neighbor point pairs. In order to accelerate the detection of the nearest neighbor point pairs matched by ICP, select the X-axis and Y-axis as the segmentation axes, and use the maximum variance within the axis as the segmentation basis to recursively divide the data.

[0030] Step 2-3: In the nearest neighbor search process, starting from the root node, recursively point the query point downward to the left subtree or right subtree according to the split axis and the maximum variance until the leaf node is reached. Update the nearest neighbor pair. At the leaf node, calculate the distance between the query point and the data point in the leaf node, and update the current nearest neighbor and the shortest distance.

[0031] (3)

[0032] Step 2-4, where d(Q,P) is the Euclidean distance between the query point Q and the data point P, and are the coordinates of point Q and point P in the i-th dimension respectively. Backtrack the parent node to observe whether the subtree on the other side contains a point whose distance to the split plane is less than the current shortest distance. Recursively search the subtree on the other side downward and update the nearest neighbor. After the search is completed, the current nearest neighbor is the nearest matching point of the query point.

[0033] Step 2-5: During the iteration of the ICP algorithm, the nearest neighbor points calculated by the KD tree are used to calculate the rotation and translation matrix R of the current lidar frame and the reference lidar frame to achieve the optimal alignment of the two sets of point clouds. Suppose a set of corresponding points is ,in is the point of the current point cloud, is the corresponding nearest neighbor of the reference point cloud. The transformation consists of the rotation matrix R and the translation vector t, which constitutes the minimization error function:

[0034] (4)

[0035] Step 2-6: retrieve the minimum neighboring point pair with the reference frame to obtain the optimal rotation matrix R and translation vector t, apply the rotation and translation matrix to the lidar coordinate system of the current frame, obtain the lidar pose in the world coordinate system, convert the three-dimensional lidar pose to the two-dimensional plane, assign the pose to the current lidar data frame, and if the two sets of point clouds are matched successfully, record the lidar pose ξ in the current three-dimensional space.

[0036] Step 2-7: If a dynamic object enters or leaves any continuous frame, the position of the laser point cloud changes, and the difference can be simplified to ΔP(x,y) = P(t+Δt)(x,y)-P(t)(x,y). For the laser points of the previous and next frames, determine whether there is a displacement change based on the difference result, set the threshold δ, and if ||ΔP(x,y)|| > δ, it is considered that the point has a dynamic change.

[0037] Furthermore, the specific process of step (3) is as follows:

[0038] Step 3-1: The sampling rates of monocular cameras and lidar sensors are different, so they need to be matched through known data points. For visual sensors, their sampling rates are much lower than those of radar sensors, so interpolation matching methods need to be used to estimate the visual data frames. L Time is the time when the dynamic point of the laser radar is generated.

[0039] Step 3-2, t c1 and t c2 Frame is the time when two adjacent data frames of the monocular camera are located.L is the estimated data frame, Frame C1 and Frame C2 The content of two adjacent data frames is expressed by the following formula: L Perform calculations.

[0040] (5)

[0041] Furthermore, the specific process of step (4) is as follows:

[0042] The laser dynamic point is transformed into coordinates. The arbitrary dynamic laser point P obtained in step 2 is transformed by the rotation matrix R LC and the translation vector t, obtained in step 3 as Frame L Perform coordinate transformation in the visual data frame to obtain Frame L The laser point set P_c in the camera coordinate system. The laser point P_c in the camera coordinate system is projected through the camera's pinhole projection model, which is described by the intrinsic parameter matrix K. After rotation and translation, any laser point P_c (x c ,y c ,z c ) is converted to the image coordinate system to obtain the image laser point P_t(x t ,y t ,1), get the image laser point set P_t.

[0043] Furthermore, the specific process of step (5) is as follows:

[0044] Step 5-1: Use the target detection method to assist in identifying dynamic objects. The dynamic detection box in YOLOv8 is (x c ,y c ,w,h), the coordinates of the upper left corner of the detection box can be expressed as (x1,x2)=(x c -w / 2,y c -h / 2), the coordinates of the lower right corner of the detection box can be expressed as (x2,y2)=(x c +w / 2,y c +h / 2), for any laser point P_t(x t ,y t ,1), only need to satisfy , , then the target detection area can be obtained;

[0045] Step 5-2: Dynamically track the target detection frame, obtain the real detection frame in the previous frame, calculate it with the current frame, and use the center point coordinates of the detection frame (c x , c y) to calculate the displacement of the center point, and obtain the displacement distance d, and then based on the displacement distance d and the time interval between two consecutive frames Calculate speed v and set speed threshold v d , if the velocity v is greater than v d , it is considered that there are dynamic objects in the current area.

[0046] Step 5-3: Perform ORB (Oriented FAST and Rotated BRIEF) feature extraction on the original image, and mark the feature points of the image area outside the target detection frame area obtained in step 5-1 as static feature points.

[0047] Step 5-4: Construct a static feature point set P j , the laser radar pose ξ obtained in step 2 is transformed by the rotation and translation matrix R LC The depth value Z is obtained by converting the current frame’s visual sensor pose and the previous frame’s visual sensor pose into a depth solution. The feature points and their depth values ​​are projected onto the AGV coordinate system of the current frame.

[0048] (6)

[0049] f is the focal length of the camera, B is the baseline, and They are the horizontal coordinates of the feature points of the current frame and the previous frame respectively.

[0050] Step 5-4: Project the three-dimensional feature points on the robot coordinate system onto a two-dimensional plane and overlay them with the lidar point cloud to convert them into a raster map.

[0051] Furthermore, the specific process of step (6) is as follows:

[0052] Step 6-1: perform semantic segmentation on the target detection frame area obtained in step 5-1 to obtain a semantic segmentation mask, mark the feature points within the mask as dynamic semantic feature points, and mark the feature points outside the mask as static semantic feature points.

[0053] Step 6-2: filter the feature points after semantic segmentation, retain the static semantic feature points, convert the coordinates of the static semantic feature points to the world coordinate system, convert the coordinates of the semantic feature points in the world coordinate system to the grid index, and calculate the logarithmic probability of the grid where the semantic feature points are located. , set a log-odds increment , get the logarithmic probability of the semantic feature point after projection .

[0054] (7)

[0055] The log odds Convert it into grid probability P, get the new grid probability of the grid coordinate, and realize the update of grid probability.

[0056] The second aspect of the present invention relates to a multi-sensor fusion AGV dynamic target elimination device based on semantic segmentation, comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the multi-sensor fusion AGV dynamic target elimination method based on semantic segmentation of the present invention.

[0057] The third aspect of the present invention relates to a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the multi-sensor fusion AGV dynamic target elimination method based on semantic segmentation of the present invention.

[0058] A fourth aspect of the present invention relates to a computer program product, comprising a computer program, which, when executed by a processor, implements the multi-sensor fusion AGV dynamic target elimination method based on semantic segmentation of the present invention.

[0059] The present invention also relates to a multi-sensor fusion AGV dynamic target removal system based on semantic segmentation, including a mobile chassis, a feature removal module, and a motion control module; the feature removal module includes a monocular camera data processing module, a two-dimensional laser radar data processing module, a target detection module, and a semantic segmentation module; the feature points and depth map of the environment are extracted by the camera data processing module, and the laser radar data processing module performs pose inference, sub-graph matching, and graph optimization. A two-dimensional laser radar is used to perform dynamic rough detection of the surrounding environment to achieve preliminary dynamic detection, and the rough detected dynamic laser point group is projected into the image coordinate system. The dynamic laser point group is projected into the image coordinate system, and the target detection frame range detected by the target detection algorithm in the image coordinate system is retrieved. If a dynamic laser point falls into the target detection frame, the target object is considered to be a dynamic object and is initially removed, otherwise the feature points in the area are converted into a grid map. The image of the target detection frame area is used as the input of the semantic segmentation thread, and a semantic label is assigned. The segmented semantic object is used as a dynamic semantic object, and the remaining area in the target detection frame is used as a static semantic object. The feature points are filtered according to the semantic segmentation result to screen out the static feature points. In the map construction stage, the grid probability is updated to remove the dynamic feature points and increase the representation of the remaining static feature points in the grid map. The present invention enriches the map features of two-dimensional laser SLAM and constructs the map in a dynamic environment.

[0060] The beneficial effects of the present invention are as follows: the dynamic feature points of the laser radar are used as the input of the target detection thread to assist in confirming the existence of the moving object and reduce the segmentation area of ​​the semantic segmentation. The real-time dynamic feedback in the environment is provided by the laser radar dynamic points, and the target detection method is used to confirm the existence of the dynamic target and optimize the dynamic detection effect. Semantic segmentation is performed only when it is confirmed to be a real dynamic object, thereby reducing the consumption of computing resources. The acquired feature points are processed in multiple stages to improve the accuracy of the map. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 is a flow chart of the method of the present invention.

[0062] Figure 2 It is a flow chart of the laser SLAM based on the method of removing dynamic objects of the present invention.

[0063] Figure 3 It is a specific flow chart of removing dynamic objects in the present invention.

[0064] Figure 4 It is a schematic diagram of removing dynamic objects according to the present invention.

[0065] Figure 5 Schematic diagram of the device of the present invention. DETAILED DESCRIPTION

[0066] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0067] Example 1

[0068] like Figure 1 , Figure 2 As shown, the present invention provides a laser SLAM method for removing dynamic semantic objects, comprising:

[0069] Step 1: Install the 2D laser radar and monocular camera, calibrate the monocular camera and 2D laser radar together, and obtain the camera's intrinsic parameter matrix and the transformation matrix T from the laser radar coordinate system to the monocular camera coordinate system. LC ;

[0070] Step 2: Laser SLAM enters the laser tracking thread, pre-processes the laser point cloud, integrates the posture information of IMU and odometer to scan and match the acquired two-dimensional laser point cloud, corrects the laser point cloud through scanning and matching, uses the laser radar data frame to perform inter-frame difference calculation, and performs dynamic threshold detection on the minimum neighbor point pair distance of the two groups of dynamic point clouds;

[0071] Step 3: Match the timestamps of the images captured by the monocular camera and the laser point cloud in the factory based on the time t of the dynamic point. l Interpolate the visual data frame to obtain the visual data frame Frame corresponding to time tc L , taking it as the visual keyframe;

[0072] Step 4: Convert the coordinates of the laser dynamic point and project it to the visual key frame Frame at time tc L The image coordinate system is used to input the lidar dynamic points into the target detection thread.

[0073] Step 5: Detect the target on the visual key frame at time tc, track the target detection frame, use the target detection method to assist in confirming the dynamic object, and extract the feature points outside the target detection frame for updating the sub-image;

[0074] Step 6: Perform semantic segmentation in the target detection box area, filter the feature points according to the semantic segmentation results, screen out static feature points, and update the grid probability during the map construction phase.

[0075] Furthermore, if Figure 3 As shown in the figure, in the same environment, laser radar and camera are used to collect data at the same time, and calibration objects that can be recognized by both sensors are used to extract features, and visual features are projected into the laser radar coordinate system. The relationship between the laser radar coordinate system and the camera coordinate system is established, and the optimal rotation matrix and translation vector are solved by the optimization algorithm to obtain the extrinsic parameter matrix of the two-dimensional laser radar and monocular camera. Step 1 In the process of solving the extrinsic parameter matrix, the laser radar and monocular camera extract features from the same calibration object to obtain the data point group P in the laser radar coordinate system. l and the corresponding point group P in the camera coordinate system c , where R LC is the rotation matrix, t is the translation vector, so that:

[0076] (1)

[0077] P c and P l分 The relationship can be used to establish an optimization problem and solve R by minimizing the reprojection error using the Levenberg-Marquardt (LM) method. LC and t.

[0078] Project the points in the camera coordinate system into the image coordinate system and obtain the camera's intrinsic parameter matrix K, such that:

[0079] (2)

[0080] Where Pi is the point coordinate in the image coordinate system.

[0081] Furthermore, in step 2, the laser radar scans to obtain the point cloud, and the dynamic point cloud detection is performed by the laser radar frame difference method. A set of point cloud data sets P(t) and P(t+Δt) are obtained by scanning at a certain time interval, where Δt is the time interval between the two scans. The two-dimensional laser point clouds obtained by the two scans are iteratively registered by the ICP algorithm.

[0082] The current two-dimensional lidar data is introduced into the KD tree to improve the retrieval speed of the nearest neighbor point pairs. In order to accelerate the detection of the nearest neighbor point pairs of ICP matching, the X-axis and Y-axis are selected as the segmentation axes, and the maximum variance within the axis is used as the segmentation basis to recursively divide the data.

[0083] In the nearest neighbor search process, starting from the root node, the query point is recursively directed downward to the left subtree or right subtree according to the split axis and the maximum variance until a leaf node is reached. The nearest neighbor pair is updated. At the leaf node, the distance between the query point and the data point in the leaf node is calculated, and the current nearest neighbor point and the shortest distance are updated.

[0084] (3)

[0085] Where d(Q,P) is the Euclidean distance between the query point Q and the data point P, and are the coordinates of point Q and point P in the i-th dimension respectively. Backtrack the parent node to observe whether the subtree on the other side contains a point whose distance to the split plane is less than the current shortest distance. Recursively search the subtree on the other side downward and update the nearest neighbor. After the search is completed, the current nearest neighbor is the nearest matching point of the query point.

[0086] In the iterative process of the ICP algorithm, the nearest neighbor points calculated by the KD tree are used to calculate the rotation and translation matrix R of the current lidar frame and the reference lidar frame to achieve the optimal alignment of the two sets of point clouds. Suppose a set of corresponding points is ,in is the point of the current point cloud, is the corresponding nearest neighbor of the reference point cloud. The transformation consists of the rotation matrix R and the translation vector t, which constitutes the minimization error function:

[0087] (4)

[0088] The minimum neighboring point pair is retrieved with the reference frame to obtain the optimal rotation matrix R and translation vector t. The rotation and translation matrix is ​​applied to the lidar coordinate system of the current frame to obtain the lidar pose in the world coordinate system. The three-dimensional lidar pose is converted to a two-dimensional plane and the pose is assigned to the current lidar data frame. If the two sets of point clouds are matched successfully, the lidar pose ξ in the current three-dimensional space is recorded.

[0089] If a dynamic object enters or leaves any continuous frame, the position of the laser point cloud changes, and the difference can be simplified to ΔP(x,y) = P(t+Δt)(x,y)-P(t)(x,y). For the laser points of the previous and next frames, determine whether there is a displacement change based on the difference result, set the threshold δ, and if ||ΔP(x,y)|| > δ, then the point is considered to have a dynamic change.

[0090] Furthermore, in step 3, the sampling rates of the monocular camera and the lidar sensor are different, and they need to be matched through known data points. For the visual sensor, its sampling rate is much lower than that of the radar sensor, and the interpolation matching method needs to be used to estimate the visual data frame. L Time is the time when the dynamic point of the laser radar is generated, t c1 and t c2 Frame is the time when two adjacent data frames of the monocular camera are located. L is the estimated data frame, Frame C1 and Frame C2 The content of two adjacent data frames is expressed by the following formula: L Perform calculations.

[0091] (5)

[0092] Furthermore, in step 4, the laser dynamic point is transformed into coordinates. The arbitrary dynamic laser point P obtained in step 2 is transformed into coordinates by the rotation matrix R LC and the translation vector t, obtained in step 3 as Frame L Coordinate transformation is performed in the visual data frame to obtain Frame L The laser point set P_c in the camera coordinate system. The laser point P_c in the camera coordinate system is projected through the camera's pinhole projection model, which is described by the intrinsic parameter matrix K. After rotation and translation, any laser point P_c (x c ,y c ,z c ) is converted to the image coordinate system to obtain the image laser point P_t(x t ,y t ,1), get the image laser point set P_t.

[0093] Furthermore, if Figure 4 As shown in the figure, step 5 uses the target detection method to assist in confirming dynamic objects. The dynamic detection box in YOLOv8 is (x c ,y c ,w,h), the coordinates of the upper left corner of the detection box can be expressed as (x1,x2)=(x c -w / 2,y c -h / 2), the coordinates of the lower right corner of the detection box can be expressed as (x2,y2)=(x c +w / 2,y c +h / 2), for any laser point P_t(x t ,y t ,1), only need to satisfy , , we can get that the target detection area is a dynamic area, and we can use ORB features on the original image, and regard the feature points outside the target detection area as static feature points.

[0094] Construct a static feature point set P j , the laser radar pose ξ obtained in step 2 is transformed by the rotation and translation matrix R LC The depth value Z is obtained by converting the current frame’s visual sensor pose and the previous frame’s visual sensor pose into a depth solution. The feature points and their depth values ​​are projected onto the AGV coordinate system of the current frame.

[0095] (6)

[0096] f is the focal length of the camera, B is the baseline, and are the horizontal coordinates of the feature points of the current frame and the previous frame respectively. Project the three-dimensional feature points in the robot coordinate system onto the two-dimensional plane and update the grid probability in the map.

[0097] Furthermore, for step six, semantic segmentation is performed on the target detection frame area obtained in step five to obtain a semantic segmentation mask, and the feature points within the mask are marked as dynamic semantic feature points, and the feature points outside the mask are marked as static semantic feature points.

[0098] Filter the feature points after semantic segmentation, retain the static semantic feature points, convert the coordinates of the static semantic feature points to the world coordinate system, convert the coordinates of the semantic feature points in the world coordinate system to the grid index, and calculate the logarithmic probability of the grid where the semantic feature points are located. , set a log-odds increment , get the logarithmic probability of the semantic feature point after projection .

[0099] (7)

[0100] The log odds Convert it into grid probability P, get the new grid probability of the grid coordinate, and realize the update of grid probability.

[0101] Example 2

[0102] like Figure 5 As shown, this embodiment relates to a multi-sensor fusion AGV dynamic target elimination device based on semantic segmentation, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the multi-sensor fusion AGV dynamic target elimination method based on semantic segmentation of Example 1.

[0103] At the hardware level, the device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 Of course, in addition to the software implementation, the present invention does not exclude other implementations, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0104] For the improvement of a technology, it can be clearly distinguished whether it is a hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or a software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages ​​and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.

[0105] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that, in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, such a controller may be considered as a hardware component, and the devices for implementing various functions included therein may also be considered as structures within the hardware component. Or even, the devices for implementing various functions may be considered as both software modules for implementing the method and structures within the hardware component.

[0106] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0107] For the convenience of description, the above device is described as being divided into various units according to their functions. Of course, when implementing the present invention, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0108] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0109] Example 3

[0110] This embodiment relates to a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the multi-sensor fusion AGV dynamic target elimination method based on semantic segmentation in Example 1 is implemented.

[0111] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0112] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0113] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0114] The present invention may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0115] Example 4

[0116] This embodiment relates to a computer program product, including a computer program, which, when executed by a processor, implements the multi-sensor fusion AGV dynamic target elimination method based on semantic segmentation in Example 1.

[0117] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0118] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0119] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0120] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0121] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0122] Each embodiment of the present invention is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0123] The above description is only an embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention should be included in the scope of the claims of the present invention.

Claims

1. A multi-sensor fusion AGV dynamic target elimination method based on semantic segmentation, characterized in that: The steps include: Step (1) install the 2D laser radar and the monocular camera, calibrate the monocular camera and the 2D laser radar together, and obtain the camera's intrinsic parameter matrix and the transformation matrix T from the laser radar coordinate system to the monocular camera coordinate system. LC ; Step (2), laser SLAM enters the laser tracking thread, pre-processes the laser point cloud, integrates the posture information of IMU and odometer to scan and match the acquired two-dimensional laser point cloud, corrects the laser point cloud through scanning and matching, uses the laser radar data frame to perform inter-frame difference calculation, and performs dynamic threshold detection on the minimum neighbor point pair distance of the two groups of dynamic point clouds; Step (3) is to match the timestamps of the images captured by the monocular camera and the laser point cloud in the factory, based on the time t to which the dynamic point belongs. l Interpolate the visual data frame to obtain the visual data frame Frame corresponding to time tc L , taking it as the visual keyframe; Step (4): transform the coordinates of the laser dynamic point and project it to the visual key frame Frame at time tc L The image coordinate system is used to input the laser radar dynamic points into the target detection thread; Step (5), perform target detection on the visual key frame at time tc, track the target detection frame, use the target detection method to assist in confirming the dynamic object, and extract feature points outside the target detection frame for updating the sub-graph; The method of using target detection to assist in identifying dynamic objects includes: Step 5-1, using the target detection method to assist in identifying dynamic objects, the dynamic detection box in YOLOv8 is (x c ,y c ,w,h), the coordinates of the upper left corner of the detection box are expressed as (x1,y1)=(x c -w / 2,y c -h / 2), the coordinates of the lower right corner of the detection box are expressed as (x2,y2)=(x c +w / 2,y c +h / 2), for the image laser point P_t(x t ,y t ,1), only need to satisfy , , then the target detection area is obtained; Step 5-2: Dynamically track the target detection frame, obtain the real detection frame in the previous frame, calculate it with the current frame, and use the center point coordinates of the detection frame (c x , c y ) to calculate the displacement of the center point, and obtain the displacement distance d, and then based on the displacement distance d and the time interval between two consecutive frames Calculate speed v and set speed threshold v d , if the velocity v is greater than v d , it is considered that there are dynamic objects in the current area; Step (6) performs semantic segmentation in the target detection box area, filters the feature points according to the semantic segmentation results, selects static feature points, and updates the grid probability during the map construction phase.

2. The multi-sensor fusion AGV dynamic target elimination method based on semantic segmentation as claimed in claim 1, characterized in that: The specific process of step (1) is as follows: Step 1-1: Use the lidar and camera to collect data in the same environment, use calibration objects that can be recognized by both sensors to extract features, project visual features into the lidar coordinate system, and establish a connection between the lidar coordinate system and the camera coordinate system; Step 1-2: Solve the optimal rotation matrix and translation vector through the optimization algorithm to obtain the external parameter matrix of the two-dimensional laser radar and monocular camera; Step 1-3: In the process of solving the external parameter matrix, the laser radar and the monocular camera extract features of the same calibration object to obtain a data point group P in the laser radar coordinate system. l and the corresponding point group P in the camera coordinate system c , where R LC is the rotation matrix, t is the translation vector, so that: , P c and P l Respectively represent the positions of the same physical point in the camera coordinate system and the lidar coordinate system; Step 1-4: Solve R by minimizing the reprojection error using the LM method LC and t; project the points in the camera coordinate system into the image coordinate system and obtain the camera's intrinsic parameter matrix K, such that: , where Pi is the point coordinate in the image coordinate system.

3. The multi-sensor fusion AGV dynamic target elimination method based on semantic segmentation as claimed in claim 1, characterized in that: The specific process of step (2) is as follows: Step 2-1, laser radar scans to obtain point cloud, and performs dynamic point cloud detection by laser radar inter-frame difference method. Scan at regular time intervals to obtain a set of point cloud data sets P(t) and P(t+Δt), where Δt is the time interval between two scans; the two-dimensional laser point clouds obtained by the two scans are iteratively registered by ICP algorithm; Step 2-2, introduce the current two-dimensional laser radar data into the KD tree to improve the retrieval speed of the nearest neighbor point pairs. In order to accelerate the detection of the nearest neighbor point pairs matched by ICP, select the X-axis and Y-axis as the segmentation axes, take the maximum variance within the axis as the segmentation basis, and recursively divide the data; Step 2-3: In the nearest neighbor search process, starting from the root node, the query point is recursively directed downward to the left subtree or the right subtree according to the split axis and the maximum variance until a leaf node is reached; Update the nearest neighbor point pair. At the leaf node, calculate the distance between the query point and the data point in the leaf node, and update the current nearest neighbor point and the shortest distance. Step 2-4, where d(Q,P) is the Euclidean distance between the query point Q and the data point P, and are the coordinates of point Q and point P in the i-th dimension respectively; backtrack the parent node to observe whether the subtree on the other side contains a point whose distance to the split plane is less than the current shortest distance, recursively search downward on the subtree on the other side and update the nearest neighbor point; After the search is completed, the current nearest neighbor point is the closest matching point of the query point; Step 2-5: During the iteration of the ICP algorithm, the nearest neighbor points calculated by the KD tree are used to calculate the rotation and translation matrix R of the current lidar frame and the reference lidar frame to achieve the optimal alignment of the two sets of point clouds. Suppose a set of corresponding points is {( , )},in is the point of the current point cloud, is the corresponding nearest neighbor of the reference point cloud. The transformation consists of the rotation and translation matrix R and the translation vector t, which constitutes the minimization error function: Step 2-6, search for the minimum neighboring point pair with the reference frame to obtain the optimal rotation matrix R and translation vector t, apply the rotation and translation matrix to the laser radar coordinate system of the current frame, obtain the laser radar pose in the world coordinate system, convert the three-dimensional laser radar pose to the two-dimensional plane, assign the pose to the current laser radar data frame, and if the two sets of point clouds are matched successfully, then record the laser radar pose ξ in the current three-dimensional space; Step 2-7: If a dynamic object enters or leaves any continuous frame, the position of the laser point cloud changes, and the difference is simplified to ΔP(x,y) = P(t+Δt)(x,y)-P(t)(x,y); for the laser points of the previous and next frames, determine whether there is a displacement change based on the difference result, set a threshold δ, and if ||ΔP(x,y)|| > δ, then the point is considered to have a dynamic change.

4. The multi-sensor fusion AGV dynamic target elimination method based on semantic segmentation as claimed in claim 1, characterized in that: The specific process of step (3) is as follows: Step 3-1: The sampling rates of the monocular camera and the lidar sensor are different, so they need to be matched through known data points. For the visual sensor, its sampling rate is much lower than that of the radar sensor, so the interpolation matching method needs to be used to estimate the visual data frame; L Time is the time when the dynamic point of the laser radar is generated. Step 3-2, t c1 and t c2 Frame is the time when two adjacent data frames of the monocular camera are located. L is the estimated data frame, Frame C1 and Frame C2 The content of two adjacent data frames is expressed by the following formula: L Perform the calculation: (5)。 5. The multi-sensor fusion AGV dynamic target elimination method based on semantic segmentation as claimed in claim 1, characterized in that: The specific process of step (4) is as follows: The coordinates of the laser dynamic point are transformed. The arbitrary dynamic laser point P obtained in step (2) is transformed by the rotation matrix R LC and the translation vector t, obtained in step (3) L Coordinate transformation is performed in the visual data frame to obtain Frame L The laser point set P_c in the camera coordinate system; the laser point P_c in the camera coordinate system is projected through the pinhole projection model of the camera, which is described by the intrinsic parameter matrix K. After rotation and translation, any laser point P_c (x c ,y c ,z c ) is converted to the image coordinate system to obtain the image laser point P_t(x t ,y t ,1), get the image laser point set P_t.

6. The multi-sensor fusion AGV dynamic target elimination method based on semantic segmentation as claimed in claim 1, characterized in that: The step (5) of extracting feature points outside the target detection frame for updating the sub-image includes: Step 5-3, perform ORB feature extraction on the original image, and mark the feature points of the image area outside the target detection frame area obtained in step 5-1 as static feature points; Step 5-4: Construct a static feature point set P j , the laser radar pose ξ obtained in step (2) is transformed by the rotation and translation matrix R LC and t are converted into visual sensor poses, the visual sensor poses of the current frame and the visual sensor poses of the reference frame are used for solution, the static feature points of the current frame and the static feature points of the previous frame are used for depth solution to obtain the depth value Z, and the feature points and their depth values ​​are projected onto the AGV coordinate system of the current frame; , f is the focal length of the camera, B is the baseline, and are the horizontal coordinates of the feature points of the current frame and the previous frame respectively; Step 5-4: Project the three-dimensional feature points on the robot coordinate system onto a two-dimensional plane and overlay them with the lidar point cloud to convert them into a raster map.

7. The multi-sensor fusion AGV dynamic target elimination method based on semantic segmentation as claimed in claim 1, characterized in that: The specific process of step (6) is as follows: Step 6-1, perform semantic segmentation on the target detection frame area obtained in step 5-1, obtain a semantic segmentation mask, mark the feature points within the mask as dynamic semantic feature points, and mark the feature points outside the mask as static semantic feature points; Step 6-2: filter the feature points after semantic segmentation, retain the static semantic feature points, convert the coordinates of the static semantic feature points to the world coordinate system, convert the coordinates of the semantic feature points in the world coordinate system to the grid index, and calculate the logarithmic probability of the grid where the semantic feature points are located. , set a log-odds increment , get the logarithmic probability of the semantic feature point after projection ; , the log-odds Convert it into grid probability P, get the new grid probability of the grid coordinate, and realize the update of grid probability.

8. A multi-sensor fusion AGV dynamic target elimination device based on semantic segmentation, characterized in that: It includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the multi-sensor fusion AGV dynamic target elimination method based on semantic segmentation as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the multi-sensor fusion AGV dynamic target elimination method based on semantic segmentation described in any one of claims 1 to 7 is implemented.

10. A computer program product, characterized in that It includes a computer program, which, when executed by a processor, implements the multi-sensor fusion AGV dynamic target elimination method based on semantic segmentation as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Dynamic environment-oriented camera and solid-state laser radar fusion repositioning method

    CN115718303A

  • Monocular vision guided AGV obstacle avoidance distance measurement method and device based on instance segmentation technology

    CN117036464A