Tracking navigation method for positioning crop center point based on laser radar and image fusion

By combining the lidar and image fusion positioning method with the U-Net model and imaging parameter calibration, the problem of navigation technology being affected by weather and environment in unmanned agricultural operations was solved, high-precision crop center point navigation was achieved, and the operating accuracy and safety of agricultural robots were improved.

CN120762047APending Publication Date: 2025-10-10HEBEI AGRICULTURAL UNIV.
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202511136322.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing navigation technology is greatly affected by weather, lighting and environment in unmanned agricultural operations, resulting in reduced navigation accuracy and efficiency, especially in weeding operations.

Method used

A lidar and image fusion positioning method is adopted. The lidar generates point cloud data and the camera takes images. The U-Net model is combined for semantic segmentation and imaging parameters are calibrated to achieve the classification of crops, weeds and ground, and the crop center point is calculated for navigation.

Benefits of technology

It achieves centimeter-level precise navigation, improves the accuracy and safety of agricultural robot operations, and reduces dependence on weather and environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120762047A_ABST
    Figure CN120762047A_ABST
Patent Text Reader

Abstract

The invention relates to a tracking navigation method for positioning a crop center point based on laser radar and image fusion, and belongs to the technical field of navigation, and the method comprises the steps: carrying out the marking of a farmland image after distortion correction, generating a training sample, inputting the training sample into a U-net model, carrying out the training, obtaining a crop segmentation model, and carrying out the classification of crops and weeds; performing feature fusion on the filtered laser radar point cloud data and the farmland image after distortion correction to obtain crop point cloud with semantic information; calculating a crop center point according to the crop point cloud with the semantic information; and providing navigation for the agricultural robot by taking the crop center point as a preview point. The spatial structure information of the target crop is obtained through the laser radar and the camera, the semantic point cloud is generated through feature fusion, the crop center point is dynamically extracted from the point cloud to serve as a real-time preview target, a fixed preview distance mechanism in a tracking algorithm is replaced, centimeter-level precise navigation is achieved, and the precision of the target crop is improved. And the operation precision and the operation safety of the agricultural robot are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of navigation technology, in particular to a tracking navigation method for locating a crop center point based on laser radar and image fusion. BACKGROUND

[0002] In agricultural unmanned operation, the navigation system is a very important link. The current mainstream navigation schemes mainly include satellite positioning navigation (GNSS), computer vision positioning navigation, laser radar navigation and other methods.

[0003] Satellite positioning mainly obtains position information through satellite data and local base stations to control the autonomous driving of robots. The machine vision system usually processes and analyzes the two-dimensional image information of the farmland through various methods to automatically obtain various instruction data to control the autonomous driving of robots. Laser radar mainly realizes positioning and mapping through laser reflection and accurately obtains the spatial position information of the crops on the farmland surface.

[0004] In the field, the key technologies of machine vision navigation and laser radar navigation are crop row extraction and seedling strip recognition. In the vision direction, the main method of crop row extraction is feature point extraction and navigation reference line fitting. For example, the color components of the RGB color image are normalized, and then the green component is enhanced to strengthen the seedling strip features, and then the corn seedling strip is segmented. Or after point cloud transformation and filtering, dynamic region calibration is performed to find the crop center point, and then crop row recognition and navigation line extraction are performed.

[0005] The existing satellite positioning navigation technical scheme can walk along the prescribed path outdoors, is suitable for rotary tillage and seeding, but for weeding operation, the positioning information may be affected by the weather to cause poor signal, and the vehicle may be affected by the terrain to cause vehicle yaw, which will affect the weeding effect.

[0006] Laser radar navigation is not affected by light, and can work stably at night, in strong light and in shadow area, but is sensitive to some environments, such as point cloud attenuation in rainy and foggy weather, and noise points are easily generated when encountering a mirror, which will affect the navigation effect and reduce the weeding rate. Visual navigation can accurately identify crop rows and weeds through semantic segmentation, and adapt to crop growth changes, but the image is overexposed under strong light, the recognition rate is greatly reduced, the features are lost in the shadow area, and dust and mud will pollute the lens and need to be cleaned frequently, so that the navigation system cannot normally navigate. SUMMARY

[0007] To solve the above problems, the purpose of the embodiment of the present application is to provide a tracking navigation method for locating a crop center point based on laser radar and image fusion.

[0008] The tracking navigation method for locating a crop center point based on laser radar and image fusion comprises the following steps:

[0009] Step 1: Use a laser radar to scan the target farmland to generate laser radar point cloud data, and perform through-filtering and outlier filtering on the laser radar point cloud data to obtain filtered laser radar point cloud data;

[0010] Step 2: Use a camera to capture a target farmland image, and perform distortion correction on the captured target farmland image to obtain a distortion-corrected farmland image;

[0011] Step 3: Annotate the distortion-corrected farmland image to generate training samples, which are then fed into a U-net model for training to obtain a crop segmentation model. Dynamic convolution is used in place of standard convolution in the convolutional layers of the U-net model, and cross-layer residual connections are introduced in each decoder layer.

[0012] Step 4: Use the crop segmentation model to perform semantic segmentation on the farmland image to classify crops, weeds, and ground;

[0013] Step 5: Calibrate the imaging parameters of the camera and lidar to obtain the calibrated imaging parameters;

[0014] Step 6: Use the calibrated imaging parameters to perform feature fusion on the filtered lidar point cloud data and the distortion-corrected farmland image to obtain a crop point cloud with semantic information;

[0015] Step 7: Compute the crop center point based on the crop point cloud with semantic information;

[0016] Step 8: Use the crop center point as a pre-aiming point to provide navigation for the agricultural robot.

[0017] Preferably, the step 5 of calibrating the imaging parameters of the camera and the lidar to obtain calibrated imaging parameters includes:

[0018] Step 5.1: Use a camera to capture the black and white grid calibration plate to form a calibration image, and use a radar laser to scan the black and white grid calibration plate to obtain a calibration point cloud;

[0019] Step 5.2: Calibrate the camera intrinsic parameters using the calibration image and obtain the camera intrinsic parameter matrix;

[0020] Step 5.3: Solve the camera distortion coefficient using the camera intrinsic parameter matrix;

[0021] Step 5.4: Calibrate the camera's extrinsic parameters based on the calibration point cloud to obtain the camera's extrinsic parameters;

[0022] Step 5.5: Jointly optimize the camera intrinsic parameter matrix, camera distortion coefficient, and camera extrinsic parameters to obtain the calibrated camera intrinsic parameters, camera distortion coefficient, and camera extrinsic parameters;

[0023] Step 5.6: Verify the calibration accuracy of the camera intrinsic parameters, camera distortion coefficients, and camera extrinsic parameters. If the calibration accuracy does not meet the preset conditions, recalibrate.

[0024] Preferably, in step 5.5, the objective function is used:

[0025]

[0026] The camera intrinsic parameter matrix, camera distortion coefficient and camera extrinsic parameter are jointly optimized to obtain the calibrated camera intrinsic parameter, camera distortion coefficient and camera extrinsic parameter; among them, represents the optimization target, M represents the number of image frames, N represents the number of feature points in each frame, and p ij represents the actual observed pixel coordinates of the jth feature point in the i-th frame image, π() represents the camera projection function, K represents the camera intrinsic parameter matrix, dist represents the distortion parameter set, R i , t i Represents the camera external parameters of the i-th frame image, P j represents the j-th 3D point coordinate in the world coordinate system, Q j represents the observation coordinates of the jth point in the laser radar coordinate system, T lc represents the homogeneous transformation matrix from the camera coordinate system to the lidar coordinate system, φ(R i , t i , P j ) represents the function that transforms the world coordinate system point to the camera coordinate system through the extrinsic parameters of the i-th frame.

[0027] Preferably, step 6: using the calibrated imaging parameters to perform feature fusion on the filtered lidar point cloud data and the distortion-corrected farmland image to obtain a crop point cloud with semantic information, includes:

[0028] Step 6.1: Extract the timestamp of the lidar point cloud and the midpoint timestamp of the camera exposure, and use the interpolation compensation formula to synchronize the time;

[0029] Step 6.2: Perform motion distortion correction on the time-synchronized lidar point cloud data to obtain corrected lidar point cloud data;

[0030] Step 6.3: Use the calibrated imaging parameters to transform the LiDAR coordinate system corresponding to the corrected LiDAR point cloud data into the camera coordinate system, completing the association between the point cloud data and the pixel data.

[0031] Step 6.4: Perform 3D semantic reconstruction on the associated data to obtain a reconstructed 3D point cloud with semantic information;

[0032] Step 6.5: removing invalid point clouds other than target crops from the three-dimensional point cloud with semantic information to obtain a crop point cloud with semantic information.

[0033] Preferably, in the step 6.3, the formula is used:

[0034]

[0035] The association of point cloud data and pixel data is completed; wherein, P l represents the lth3D point in the point cloud, Proj(P l ) represents the function of projecting the point P l to the image plane, [u, v] T represents the candidate pixel coordinates, Reflectance(P l ) represents the reflectivity of the point P l , Grayscale(u, v) represents the grayscale value of the image at the pixel (u, v), λ represents the weight coefficient, (u * , v * ) represents the optimal matching pixel coordinates, P represents the candidate 3D point, represents the gradient of the image at (u, v), represents the depth gradient at the point P, μ represents the weight coefficient, P * represents the optimal matching 3D point.

[0036] Preferably, in the step 6.4, the formula is used:

[0037]

[0038] The associated data is three-dimensionally reconstructed to obtain a reconstructed three-dimensional point cloud with semantic information; wherein, TSDF(v) is the truncated signed distance function value at the voxel v, v is the voxel position in the three-dimensional space, i is the observation data index, z i is the depth value of the ithobservation, d i is the distance from the voxel v to the ithobservation viewpoint, o i is the camera optical center position of the ithobservation, w i is the fusion weight of the ithobservation, C l is the confidence of the lidar data, C v is the confidence of the visual data.

[0039] Preferably, the step 7: calculating the crop center point according to the crop point cloud with semantic information, comprises:

[0040] Step 7.1: removing invalid point clouds from the crop point cloud with semantic information, and screening out point clouds at the target crop stem;

[0041] Step 7.2: Grid the selected point cloud to obtain a gridded density map;

[0042] Step 7.3: Find the density peak area in the gridded density map, determine the edge points of the area from the density peak area, and select three edge points that meet the preset conditions from the edge points;

[0043] Step 7.4: Fit a circle through the three edge points and use the coordinates of the circle center as the crop center point.

[0044] Preferably, in step 7.1, the following steps are included:

[0045] Remove invalid point clouds in the ground plane through point cloud height filtering; the point cloud height filtering formula is:

[0046]

[0047] in, is the original point cloud set, p i is a point in the point cloud, z i Point p i The z coordinate, μ z is the mean height of the point cloud, σ z is the standard deviation of the point cloud height, Represents the valid point cloud set retained after height filtering;

[0048] Through semantic label filtering, other invalid point clouds are removed; the semantic label filtering formula is:

[0049]

[0050] in, Represents the effective point cloud set after height filtering, p i express A point in i The semantic probability distribution vector represented by max(s i ) represents the point p i The maximum value in the semantic probability distribution, s corn represents the semantic category corresponding to the maximum value, τ p represents the probability threshold, Represents the final set of point clouds that belong to the target crop category.

[0051] Preferably, the step 7.3: finding a density peak region from the gridded density map, determining edge points of the region from the density peak region, and selecting three edge points that meet preset conditions from the edge points, includes:

[0052] Step 7.4.1: Find the density peak area in the gridded density map. The density peak area is defined as:

[0053]

[0054] in, represents the density peak area, ρ(x, y) represents the density value of the grid unit (x, y), ρ max The maximum density value of the entire density map, ρ(x, y) represents the density value of the grid cell (x, y), cell(x, y) represents the unit area in the two-dimensional grid, p i Indicates the point belonging to this grid cell, z i Indicates the height value of the point, μ z represents the mean height of the points within the grid cell, σ z represents the standard deviation of the height of the points within the grid cell, and exp represents the Gaussian weight function;

[0055] Step 7.4.2: Calculate the edge points of the density peak area using the polar coordinate edge detection method; the edge point calculation formula is:

[0056]

[0057] Among them, θ k is the angle of the kth search direction, K is the total number of directions, k is the direction index, and p k is the direction θ k The most edge point on x i ,y i For point p i The plane coordinate, x p ,y p is the density peak coordinate, cosθ k , sinθ k is the direction vector component;

[0058] Step 7.4.3: From the edge points of the density peak area, select the points that meet the selection strategy:

[0059] Three edge points; A is the first edge point, B is the second edge point, C is the third edge point, θ k is the polar angle of the point.

[0060] Preferably, the step 8: using the crop center point as a preview point to provide navigation for the agricultural robot includes:

[0061] The nearest crop center point is used as a preview point for navigation. If the preview point is no longer visible, the next nearest crop center point is searched for as a preview point to continue navigation. The strategy for selecting the nearest crop center point is as follows:

[0062]

[0063] Among them, P nearest The optimal crop center point represents the target crop center position that the navigation system will aim at, P i The center point of the candidate crop represents the center coordinate of the i-th target crop detected by the sensor, V is the real-time coordinate of the center point of the agricultural robot, θ v is the heading of the agricultural robot, θ row Indicates the angle between the crop and the current direction of the agricultural robot.

[0064] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0065] The present invention relates to a tracking and navigation method for locating the center point of a crop based on laser radar and image fusion. Compared with the existing technology, the present invention obtains spatial structural information of the target crop through laser radar and camera, generates a semantic point cloud through feature fusion, and dynamically extracts the crop center point from the point cloud as a real-time preview target, replacing the fixed preview distance mechanism in the tracking algorithm, achieving centimeter-level precise navigation and greatly improving the operating accuracy and safety of agricultural robots.

[0066] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0068] Figure 1 This is a flow chart of a tracking and navigation method for locating the center point of a crop based on laser radar and image fusion in an embodiment of the present invention;

[0069] Figure 2 The point cloud retained after the straight-through filtering provided by the present invention;

[0070] Figure 3 The black and white grid calibration plate provided by the present invention;

[0071] Figure 4The pure tracking algorithm schematic diagram provided by the present application is shown in the following figure:

[0072] Figure 5 The improved schematic diagram of the pure tracking algorithm provided by the present application is shown in the following figure. DETAILED DESCRIPTION

[0073] In the description of the present application, it should be understood that the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise" and the like indicate the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application.

[0074] In addition, the terms "first" and "second" are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more of the features. In the description of the present application, the meaning of "multiple" is two or more, unless otherwise specifically limited.

[0075] In the present application, unless otherwise specifically defined and limited, the terms "mounting", "connection", "connection", "fixing" and the like should be understood broadly, for example, it can be fixedly connected, or it can be detachably connected, or it can be integrally connected; it can be mechanically connected, or it can be electrically connected; it can be directly connected, or it can be indirectly connected through an intermediate medium; it can be the communication inside two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0076] Please refer to Figure 1 The tracking navigation method for locating the center point of crops based on laser radar and image fusion includes:

[0077] Step 1: using laser radar to scan the target farmland to generate laser radar point cloud data, and performing straight-through filtering and outlier filtering processing on the laser radar point cloud data to obtain filtered laser radar point cloud data;

[0078] In this embodiment of the present invention, the target crop is corn seedlings. To accurately obtain the spatial position of corn seedlings in a farmland, a lidar sensor can be used. A lidar sensor uses a laser to emit laser pulses. These pulses propagate in a straight line through a propagation medium and are reflected back to a receiver upon encountering an object. By measuring the time between the emission of the pulse and the reception of the reflected pulse and multiplying it by the speed of light, the distance between the object and the lidar can be accurately calculated.

[0079] 1) Point cloud direct filtering

[0080] like Figure 2 As shown in the figure, in order to effectively focus on the forward perception area of ​​the environment and reduce the computational complexity of subsequent processing, it is necessary to perform straight-through filtering on the original point cloud and spatially crop the original LiDAR point cloud data. The 120-degree field of view centered on the sensor's own coordinate system is defined as the key perception area. Only valid observation points within the specified forward fan-shaped area are retained, reducing the storage and computing power consumption of irrelevant data. This improves the collaborative work effect of LiDAR data and visual data in fusion. The mathematical principle of straight-through filtering is expressed as follows:

[0081]

[0082] Among them, PointClound Valid represents valid point cloud data, that is, points retained by filtering, PointClound Invalid represents invalid point cloud data, that is, points removed by filtering, Point Position refers to the azimuth of each point, and Range specifies the angle range.

[0083] 2) Point cloud outlier filtering

[0084] When operating in complex environments, lidar point cloud data is susceptible to interference from outliers caused by atmospheric particles such as dust, smoke, and raindrop reflections. These noise points typically exhibit sparse spatial distribution, significantly lower density than points on the surface of real objects, and exhibit unusual jumps in ranging values. To effectively suppress this noise and improve the quality of point cloud data, outlier filtering is employed to process the point cloud. The core principle of this algorithm is based on statistical analysis of local neighborhood point clouds. For each query point in the point cloud, the algorithm calculates the average distance to its k nearest neighbors and calculates the mean (μ) and standard deviation (σ) of these average distances across the entire point cloud. Outliers are then removed. This algorithm effectively identifies and removes sparse, isolated noise points generated by discrete suspended particles in the air, while effectively preserving the dense structural information representing the surface of real objects. This step significantly reduces the negative impact of noise on key tasks such as feature extraction, target recognition, and 3D reconstruction, enhancing the reliability of the perception system under non-ideal atmospheric conditions. The mathematical principles of this algorithm are as follows.

[0085] This function will traverse and calculate the distance from point P to these points:

[0086]

[0087] Among them, distance(p i ,p) represents point p i The distance between point p and point p i is the i-th point, point p is the point currently considered, (x i ,y i ,z i ) is point p i The coordinates of point p are (x, y, z).

[0088] Find the average distance from point P to these points:

[0089]

[0090] Mean distance represents the average distance from point p to K points;

[0091] Compute the global average distance:

[0092]

[0093] For all points, calculate the global average distance standard deviation:

[0094]

[0095] σ represents the standard deviation; determine whether a point P is an outlier:

[0096] If meandistance p >Globalmeandistance+std*σ

[0097] Then the point P is an outlier, and std is the set threshold.

[0098] Step 2: Use a camera to capture a target farmland image, and perform distortion correction on the captured target farmland image to obtain a distortion-corrected farmland image;

[0099] When shooting a pattern, the pattern may be distorted by the lens, affecting the quality of the pattern. For example, radial distortion is caused by the curvature of the lens. The mathematical model is:

[0100] x corrected =x d (1+k1r 2 +k2r 4 +k3r 6 )

[0101] The tangential distortion is mainly caused by the non-parallelism between the lens and the sensor, and the mathematical model is as follows:

[0102]

[0103] In order to obtain more accurate images, it is necessary to eliminate the distortion, and the mathematical model for correcting the radial distortion is as follows, which can effectively eliminate the bending of the edge of the wide-angle lens, such as the deformation of the corn stalk:

[0104] x corrected =x d (1+k1r 2 +k2r 4 +k3r 6 )

[0105] y corrected =y d (1+k1r 2 +k2r 4 +k3r 6 )

[0106]

[0107] The mathematical model for correcting the tangential distortion is as follows, which mainly corrects the image shear caused by the non-parallelism between the lens and the sensor, and reduces the influence on the straightness of the ridge line:

[0108]

[0109] Wherein, x d , y d are the coordinates of the distortion point (i.e. the coordinates before correction), r is the distance from the distortion point to the center point of the image, k1, k2, k3 are the radial distortion coefficients, p1, p2 are the tangential distortion coefficients, x corrected , y corrected are the corrected coordinates.

[0110] Step 3: labeling the farmland image after distortion correction to generate training samples, inputting the training samples into the U-net model for training to obtain a crop segmentation model; wherein, in the convolution layer of the U-net model, dynamic convolution is used instead of standard convolution, and cross-layer residual connection is introduced in each layer of the decoder;

[0111] In order to realize the semantic segmentation of visual data, the U-Net model can be used, and the U-Net model is modified, and a model capable of identifying corn seedlings is trained.

[0112] In this scheme, the U-Net model will be modified, and the modification scheme of the U-Net model is as follows: in the convolution layer, dynamic convolution is used instead of standard convolution. Cross-layer residual connection is introduced in each layer of the decoder.

[0113] In the original model, the standard convolution kernel had the following limitations: the fixed weights could not adapt to the changing shape of cornfields, were sensitive to leaf occlusion and lighting changes, and had a significant drop in recognition rate under strong light. The limitations of standard convolution are as follows:

[0114] Y=X*W static +b

[0115] Among them, X is the input feature map, W static is the static convolution kernel weight, b is the bias term, and Y is the feature map of the convolution output.

[0116] Therefore, dynamic convolution is used this time. The mathematical principle is as follows:

[0117] π k =Softmax(GAP(X)·V k )

[0118]

[0119] Y=X*W dyn +b

[0120] Among them, GAP(X) is the global average pooling of the input feature map, V k is a learnable weight vector used to calculate the attention score of the kth convolution kernel. The Softmax function can convert the score into a probability distribution. K is the number of convolution kernels used in dynamic convolution, and W k is the weight of the kth static convolution kernel, π k is the attention weight of the kth convolution kernel, W dyn The convolution kernel weights are dynamically generated.

[0121] At the same time, for the farmland environment, the weight of the farmland crop feature channel, such as corn stalks, is strengthened:

[0122]

[0123] Among them, X c is the cth channel of the input feature map, ChannelAttn is the channel attention function, α c is the learnable weight parameter of channel c, V k is the vector used to calculate the attention score of the k-th convolution kernel.

[0124] At the same time, the introduction of cross-layer residual connections can improve noise suppression, reduce soil false detection rate, and improve edge fidelity. The mathematical expression of the improved residual connection is as follows:

[0125] R i =DynamicConv(E i )

[0126]

[0127] Among them, E i represents the feature map of the encoder at layer i, R i Represents the result of dynamic convolution on the encoder feature map, U i Denotes the upsampled feature map, D i is the output feature map of the decoder layer i, and DynamicConv represents the dynamic convolution of the input of the decoder layer i.

[0128] After completing the model modification, data collection is required. A large number of pictures containing corn seedlings are taken. These pictures are then classified into training and test sets.

[0129] The training set images are then labeled and preprocessed. First, the images are resized to a uniform size, the corn seedlings are labeled, and their boundaries are marked. The data is then normalized to a standard normal distribution with a mean of 0 and a standard deviation of 1. Images can also be rotated to increase sample diversity. These images are then fed into the U-net model for training. During training, the model is continuously iterated and optimized to improve recognition accuracy.

[0130] Step 4: Use the crop segmentation model to perform semantic segmentation on the farmland image to classify crops, weeds, and ground, and identify information such as crops, weeds, and ground to facilitate subsequent fusion processing;

[0131] Step 5: Calibrate the imaging parameters of the camera and lidar to obtain the calibrated imaging parameters;

[0132] Furthermore, step 5 includes:

[0133] Step 5.1: Use a camera to capture the black and white grid calibration plate to form a calibration image, and use a radar laser to scan the black and white grid calibration plate to obtain a calibration point cloud;

[0134] Step 5.2: Calibrate the camera intrinsic parameters using the calibration image and obtain the camera intrinsic parameter matrix;

[0135] Step 5.3: Solve the camera distortion coefficient using the camera intrinsic parameter matrix;

[0136] Step 5.4: Calibrate the camera's extrinsic parameters based on the calibration point cloud to obtain the camera's extrinsic parameters;

[0137] Step 5.5: Jointly optimize the camera intrinsic parameter matrix, camera distortion coefficient, and camera extrinsic parameters to obtain the calibrated camera intrinsic parameters, camera distortion coefficient, and camera extrinsic parameters;

[0138] Step 5.6: Verify the calibration accuracy of the camera intrinsic parameters, camera distortion coefficients, and camera extrinsic parameters. If the calibration accuracy does not meet the preset conditions, recalibrate.

[0139] In practical applications, step 5 is specifically as follows:

[0140] like Figure 3 As shown in the figure, a black and white grid calibration plate is used about 50 cm in front of the lidar and camera to save point cloud data and visual data.

[0141] First, perform camera intrinsic calibration based on the recorded data, and first calculate the homography matrix:

[0142]

[0143] H=K[r1 r2 t]

[0144] Where λ is the scale factor used to scale the homogeneous coordinates, u,v are the pixel coordinates on the image plane, and X,Y are the point coordinates in the world coordinate system. H is the homography matrix that maps the world coordinate point (X,Y,1) to the image pixel point (u,v,1). K is the camera intrinsic parameter matrix, r1,r2 are the first two columns of the rotation matrix R, and t is the translation vector.

[0145] Then perform internal parameter solution:

[0146] The closed-form solution is:

[0147]

[0148] h ij is the i-th row and j-th column element of the homography matrix H.

[0149] The constraint equation is:

[0150]

[0151] Among them, h1,h2 are the column vectors of the homography matrix H, K -T The inverse transpose of the internal parameter matrix K, K -1 The inverse matrix of the intrinsic parameter matrix K.

[0152] The internal parameter matrix is:

[0153]

[0154] Among them, f x ,f y The focal length of the camera in the x-axis and y-axis directions, c x ,c y The pixel coordinates of the camera's principal point in the image.

[0155] After solving the internal parameters, the distortion coefficient needs to be solved. The distortion mathematical model is:

[0156]

[0157] Among them, u, v are the normalized image coordinates without distortion, u d ,v d The normalized image coordinates after distortion, r is the radial distance from the image center to the point (u, v), k1, k2, k3 are the radial distortion coefficients, and p1, p2 are the tangential distortion coefficients.

[0158] Minimize the reprojection error:

[0159]

[0160] in, To optimize the target, the error is minimized by adjusting the distortion coefficient, P ij is the actual observed pixel coordinate of the jth point in the i-th image, the proj projection function projects the 3D point to the image plane, K is the camera intrinsic parameter matrix, dist is the distortion parameter, and R,t is the camera extrinsic parameter.

[0161] After the camera intrinsic calibration, the lidar-camera extrinsic calibration is also required:

[0162] First, perform point cloud corner extraction and plane fitting:

[0163]

[0164] Where a, b, c, d are the coefficients of the plane equation, and x, y, z are the 3D point coordinates in the point cloud.

[0165] Then solve the external parameters and first obtain the centralized coordinates:

[0166]

[0167] in, is the center of mass of the point in the camera coordinate system, C is the coordinate of the i-th point in the camera coordinate system, is the center of mass of the point in the laser radar coordinate system, L j is the coordinate of the jth point in the lidar coordinate system, and N is the number of points.

[0168] Then we get the covariance matrix:

[0169]

[0170] Where H is the covariance matrix, is the center of mass of the point in the camera coordinate system, C i is the coordinate of the i-th point in the camera coordinate system, is the center of mass of the point in the laser radar coordinate system, L i is the coordinate of the i-th point in the lidar coordinate system.

[0171] And perform SVD decomposition:

[0172]

[0173] Where SVD(H) represents the singular value decomposition of the matrix H, [U, S, V] represents the result matrix of the SVD decomposition, and R represents the rotation matrix from the lidar coordinate system to the camera coordinate system.

[0174] Finally, translate the vector:

[0175]

[0176] t represents the translation vector from the lidar coordinate system to the camera coordinate system.

[0177] After completing the solution, you need to import the motion transformation equation and external parameter matrix:

[0178]

[0179]

[0180] Where I9 represents the 9×9 identity matrix, represents the Kronecker product of matrices A and B, vec(X) represents the operation of vectorizing matrix X, X represents the homogeneous transformation matrix from the lidar coordinate system to the camera coordinate system, R lc represents the rotation matrix from the lidar coordinate system to the camera coordinate system, t lc Represents the translation vector from the lidar coordinate system to the camera coordinate system.

[0181] Finally, we need to jointly optimize the data of both parties. The objective function is as follows:

[0182]

[0183] in, represents the optimization goal (minimizing camera intrinsic parameters, distortion parameters, and camera extrinsic parameters), M represents the number of image frames, N represents the number of feature points in each frame, and p ij represents the actual observed pixel coordinates of the jth feature point in the i-th frame image, π() is the machine projection function (projecting the 3D point to the image plane), K represents the camera intrinsic parameter matrix, dist represents the distortion parameter set, R i ,t i Represents the camera external parameters of the i-th frame image, P j represents the j-th 3D point coordinate in the world coordinate system, Qj represents the observed coordinates of the jth point in the lidar coordinate system, T lc represents the homographic transformation matrix from the camera coordinate system to the lidar coordinate system (i.e., the inverse of X), φ(R i ,t i , P j ) represents the function of transforming a world coordinate system point through the ith frame extrinsic to the camera coordinate system.

[0184] The Jacobian matrix needs to be introduced to calculate the camera projection derivative and the point cloud distance derivative, and the mathematical principles are as follows:

[0185]

[0186] where J cam is the Jacobian matrix of the camera projection, (u, v) represents the pixel coordinates on the plane, (X c , Y c , Z c ) represents the 3D coordinates in the camera coordinate system, (R, t) is the extrinsic of the camera, J lidar represents the Jacobian matrix of the point cloud distance, Q is the point coordinates in the lidar coordinate system, T lc is the homographic transformation matrix from the camera coordinate system to the lidar coordinate system, P c is the coordinate of the point in the camera coordinate system.

[0187] After completing the calibration, the calibration accuracy needs to be verified:

[0188] First, verify the reprojection error:

[0189]

[0190] where ∈ cam is the average reprojection error, N is the total number of feature points, p ij is the observed pixel coordinates of the jth feature point in the ith image, π() is the camera projection function, K is the camera intrinsic matrix, dist is the distortion parameter set, R i , t i is the camera extrinsic of the ith image, P j is the 3D point coordinates in the world coordinate system.

[0191] Then verify the radar point cloud alignment error:

[0192]

[0193] where ∈ lidar represents the average point cloud alignment error, Q j is the observed coordinates of the jth point in the lidar coordinate system, T lcis the transformation matrix from camera to lidar, is the coordinate of the j-th point in the camera coordinate system.

[0194] Finally, we need to verify the time and space synchronization errors between the two:

[0195] Δt=|t cam -t lidar |<0.01s

[0196] δx=v·Δt

[0197] Where Δt is the time synchronization error, t cam The timestamp of the camera collecting data, t lidar The timestamp of the lidar data acquisition;

[0198] Step 6: Use the calibrated imaging parameters to perform feature fusion on the filtered lidar point cloud data and the distortion-corrected farmland image to obtain a crop point cloud with semantic information;

[0199] Furthermore, step 6 includes:

[0200] Step 6.1: Extract the timestamp of the lidar point cloud and the midpoint timestamp of the camera exposure, and use the interpolation compensation formula to synchronize the time;

[0201] Step 6.2: Perform motion distortion correction on the time-synchronized lidar point cloud data to obtain corrected lidar point cloud data;

[0202] Step 6.3: Use the calibrated imaging parameters to transform the LiDAR coordinate system corresponding to the corrected LiDAR point cloud data into the camera coordinate system, completing the association between the point cloud data and the pixel data.

[0203] Step 6.4: Perform 3D semantic reconstruction on the associated data to obtain a reconstructed 3D point cloud with semantic information;

[0204] Step 6.5: Remove invalid point clouds other than the target crop from the 3D point cloud with semantic information to obtain the crop point cloud with semantic information.

[0205] In practical applications, step 6 is specifically as follows:

[0206] After calibration is completed, the data of the two will be fused. After the host obtains the data of the two, time synchronization will be performed first. The host will extract the timestamp of the lidar point cloud and the midpoint time of the camera exposure, and use the interpolation compensation formula for time synchronization:

[0207]

[0208] Among them, P syncrepresents the point cloud coordinates after time synchronization, P t1 ,P t2 Represents the point cloud coordinates of adjacent timestamps and t1, t2, t cam It represents the camera exposure midpoint timestamp, and t1 and t2 represent the lidar point cloud timestamps.

[0209] Then the radar point cloud motion distortion correction is performed, mainly for the radar's own motion and dynamic object compensation:

[0210]

[0211] Among them, x, y, z are the original point cloud coordinates (uncorrected), x′, y′, z′ are the corrected point cloud coordinates, and R imu is the rotation matrix measured by IMU, t imu is the translation vector measured by IMU, R,t is the rotation matrix and translation vector to be optimized, P i is the point coordinate in the target point cloud, Q j The coordinates of the point in the source point cloud.

[0212] The coordinate data of the two need to be normalized. The main method is to transform the radar coordinate system to the camera coordinate system by using the rotation matrix:

[0213]

[0214] Among them, x l ,y l ,z l Indicates the point coordinates in the laser radar coordinate system, x c ,y c ,y z Represents the point coordinates in the camera coordinate system, R lc represents the rotation matrix from the laser radar to the camera, t lc represents the translation vector from the laser radar to the camera, T lc Represents a homogeneous transformation matrix.

[0215] At this point, the data of the two are synchronized to the camera coordinate system. Next, the data in the camera coordinate system is synchronized to the image pixels:

[0216]

[0217] Among them, u, v represent the image pixel coordinates, s represents the scaling factor, and f x ,f y Indicates the focal length of the camera, c x ,c y The coordinates of the camera principal point are shown, x c ,y c ,z cdenotes the point coordinate in the camera coordinate system.

[0218] Now the point cloud to pixel projection can be performed, the point cloud data is associated with the pixel data by the following method, the accurate point-by-point correspondence is established between the two modal data, and the problem of "which point cloud point corresponds to which pixel" is solved.

[0219]

[0220] where P l denotes the lth3D point in the point cloud, Proj(P l ) denotes the function of projecting the point P l to the image plane, [u, v] T denotes the candidate pixel coordinate, Reflectance(P l ) denotes the reflectance of the point P l , Grayscale(u, v) denotes the grayscale value of the image at the pixel (u, v), λ denotes the weight coefficient, (u * , v * ) denotes the optimal matching pixel coordinate, P denotes the candidate 3D point, denotes the gradient of the image at (u, v), denotes the depth gradient at the point P, μ denotes the weight coefficient, P * denotes the optimal matching 3D point.

[0221] Next, the associated data is subjected to three-dimensional semantic reconstruction, and a reconstructed three-dimensional point cloud with semantic information is obtained:

[0222]

[0223] where TSDF(v) is the truncated signed distance function value at voxel v, v is the voxel position in three-dimensional space, i is the observation data index (all observations are traversed), z i is the depth value of the ith observation, d i is the distance from the voxel v to the ith observation viewpoint, o i is the camera optical center position of the ith observation, w i is the fusion weight of the ith observation, C l is the confidence of the lidar data, C v is the confidence of the visual data.

[0224] Thus, a three-dimensional point cloud with semantic information is obtained.

[0225] Step 7: calculating the crop center point according to the crop point cloud with semantic information;

[0226] Further, step 7 includes:

[0227] Step 7.1: Remove invalid point clouds from the crop point cloud with semantic information and select the point cloud at the target crop stem;

[0228] Step 7.2: Grid the selected point cloud to obtain a gridded density map;

[0229] Step 7.3: Find the density peak area in the gridded density map, determine the edge points of the area from the density peak area, and select three edge points that meet the preset conditions from the edge points;

[0230] Step 7.4: Fit a circle through the three edge points and use the coordinates of the circle center as the crop center point.

[0231] After obtaining the three-dimensional point cloud with semantic information, invalid point clouds other than corn seedlings, such as ground point clouds and weed point clouds, will be removed based on the three-dimensional point cloud with semantic information.

[0232] First, remove the ground plane point cloud by filtering the point cloud height:

[0233]

[0234] in, is the original point cloud set, p i is a point in the point cloud, z i Click p i The z coordinate, μ z The mean value of the point cloud height (z coordinate), σ z The standard deviation of the point cloud height (z coordinate), The valid point cloud set retained after height filtering.

[0235] Next, we filter the semantic labels to remove invalid point clouds:

[0236]

[0237] in, Represents the effective point cloud set after height filtering (input), p i express A point in i The semantic probability distribution vector represented by max(s i ) represents the point p i The maximum value in the semantic probability distribution, s corn Indicates the semantic category corresponding to the maximum value, τ p Represents the probability threshold, requiring the maximum probability to be no less than this value, Represents the final set of point clouds that belong to the corn category.

[0238] At this point, we have obtained a valid corn seedling point cloud with semantic information. Next, we will calculate the center point of the crop based on the semantic information of the corn seedling point cloud to provide a preview point for subsequent navigation.

[0239] Based on the fused point cloud with semantic information, crop row center detection is carried out to determine the spatial coordinates of the collected corn seedlings, providing data support for subsequent navigation operations.

[0240] When calculating the center point of the crop, first find the point cloud at the corn stalk:

[0241] P base ={p i ∈P corn |z i ∈[z min +0.05,z min +0.3]}

[0242] in, represents the final retained point cloud set belonging to the corn category, p i is a point in the point cloud, z i For point p i The z coordinate, z min is the lowest height of the entire point cloud, P base It is a point cloud collection of the filtered corn stalk parts.

[0243] The point cloud is then gridded to obtain a gridded density map:

[0244]

[0245] Among them, ρ(x, y) represents the density value of the grid cell (x, y), cell(x, y) represents the unit area in the two-dimensional grid, and p i Indicates the point belonging to this grid cell, z i Indicates the height value of the point, μ z represents the mean height of the points within the grid cell, σ z represents the standard deviation of the height of the points within the grid cell, and exp represents the Gaussian weight function.

[0246] Find the peak density area, the peak area is defined as

[0247]

[0248] in, represents the density peak area, ρ(x,y) represents the density value of the grid unit (x,y), ρ max The maximum density value of the entire density map.

[0249] Then perform density peak area detection

[0250]

[0251] Among them, (x p ,y p ) is the grid coordinate with the largest density, and ρ(x, y) is the density value of the grid unit (x, y).

[0252] After finding the density peak area, we can find the edge points of the area by using the polar coordinate edge detection method. The principle is as follows:

[0253]

[0254] Among them, θ k is the angle (radian) of the kth search direction, K is the total number of directions (such as K = 8 means 8 directions), k is the direction index, p k is the direction θ k The most edge point on x i ,y i For point p i The plane coordinate, x p ,y p is the density peak coordinate, cosθ k , sinθ k is the direction vector component.

[0255] Then select the three most marginal points. The three-point selection strategy is:

[0256]

[0257] Among them, θ k is the polar coordinate angle of the point, p argmax Indicates that from all edge points, the polar coordinate angle θ is selected k The largest point, the angle θ k The largest point is defined as point A, and point A is selected as the first edge point. Then, after the relative angle is offset by 120°, point B with the largest value is selected as the second edge point. Then, after the relative angle is offset by 240°, point C with the largest value is selected as the third edge point.

[0258] Through these three points, a circle can be fitted, and the equation of the circle is:

[0259] (x-x0) 2 +(y-y0) 2 =r 2

[0260] Based on this circle, the center can be solved:

[0261]

[0262] The coordinates of points A, B, and C are:

[0263] A=(x A ,y A )

[0264] B=(x B ,y B )

[0265] C=(x C ,y C )

[0266] And finally the center of the circle is:

[0267] C corn =(x0, y0, z min )

[0268] This point is the center of the corn seedling.

[0269] Step 8: Use the crop center point as a pre-aiming point to provide navigation for the agricultural robot.

[0270] This solution uses an improved pure tracking algorithm for navigation. By selecting the center of the crop as the preview point for navigation, we first introduce the pure tracking algorithm. The algorithm principle is as follows: Figure 4 shown.

[0271] First, establish the following kinematic equations

[0272] Vehicle position change rate:

[0273]

[0274] Vehicle heading change rate:

[0275] θ·=vtan(δ) / L

[0276] The front wheel angle is then derived using a pure tracking algorithm:

[0277]

[0278] Among them, (x r ,y r ) represents the coordinates of a preview point on the path; l d It is the connection between the origin of the body coordinate system and the preview coordinate (x r ,y r ) is the chord length of the arc segment, also called the preview distance; R is the radius of the arc segment, that is, the turning radius; L is the wheelbase of the vehicle body; δ is the front wheel turning angle; α is the angle between the vehicle body and the preview point of the path; φ is the yaw angle of the vehicle body; (x, y) are the coordinates of the center of the tractor's rear wheel.

[0279] Pure tracking algorithms use a fixed preview distance to find a preview point, then calculate the front wheel steering angle based on that preview point. However, this scheme ignores heading angle deviation. When the initial heading error is large, relying solely on position correction can introduce lateral overshoot, causing snaking oscillations in the vehicle trajectory and significantly slowing convergence. Furthermore, its performance is highly dependent on the choice of preview distance. If a fixed preview distance is used, a trade-off between response speed and stability is necessary on paths with variable curvature. Preview points with too short a distance can lead to increased oscillations, while too long a distance can cause lag in corners. Dynamically adjusting the preview distance requires considering both vehicle speed and path curvature, significantly increasing the complexity of parameter tuning. This makes it difficult to implement on a large scale.

[0280] Therefore, to obtain a more accurate navigation preview point, this solution uses the crop center as the preview point instead of using a fixed preview distance for navigation. The positioning of the crop center point is completed by LiDAR and visual cameras.

[0281] After obtaining the crop center point, the nearest crop center point is used as the preview point, and the device will continue to navigate with this preview point until the target point is no longer visible. Figure 5 .

[0282] The pure tracking algorithm uses a fixed preview distance to find the preview point. To find the preview point, a preset trajectory must be obtained first. However, in farmland scenarios, trajectory acquisition is more difficult, and due to interference from weeds and seedlings, the acquired trajectory may also be inaccurate, which can lead to navigation errors. Furthermore, because the pure tracking algorithm uses a fixed preview distance to find the preview point, after the vehicle has traveled a certain distance, it will select subsequent path points as preview points based on the preview distance (the dotted line in the figure indicates the vehicle's future travel position). This method may cause the device to deviate from the preset trajectory due to heading angle deviation.

[0283] Therefore, the present invention proposes an improved solution. By selecting the center point of the crop as the preview point and not using a fixed preview distance for navigation, this method is closer to the complex environment of the field and has stronger adaptability. At the same time, there is no need to preset the path in advance, which reduces the workload. At the same time, the preview point selection strategy is optimized to avoid the occurrence of crop damage caused by fixed-distance preview.

[0284] When the system is running, it will continue to follow this preview point until it is no longer visible in the system. At this time, the system will continue to look for the next nearest crop center point as the preview point for navigation, and this cycle will continue. This solution can effectively improve the navigation accuracy of the equipment in the field and reduce damage to seedlings.

[0285] The selection strategy of the nearest crop center point in this scheme is:

[0286]

[0287] Among them, P nearest The optimal crop center point represents the center position of the corn plant that the navigation system will target, P i The center point of the candidate crop represents the center coordinate of the i-th corn detected by the sensor, V is the current position of the vehicle, that is, the real-time coordinate of the center point of the agricultural machinery, θ v is the vehicle heading, θ row represents the angle between the crop and the current direction of the vehicle, and λ represents the weight coefficient, which is used to balance the weight of distance and direction deviation.

[0288] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0289] 1. By fusing the spatial information of lidar with the semantic information of vision, a three-dimensional point cloud with semantic labels is generated, which significantly improves the recognition accuracy and positioning precision of corn seedlings under complex lighting (strong light, shadows) and partial occlusion conditions.

[0290] 2. By directly using the dynamically detected crop center point as the preview point, the system overcomes the limitations of traditional pure tracking algorithms that rely on preset paths and fixed preview distances. This significantly reduces trajectory oscillation and lag caused by heading angle deviation, improves navigation smoothness and real-time dynamic adaptability, and effectively reduces the risk of crushing seedlings due to navigation errors.

[0291] 3. The improved U-Net model (dynamic convolution and cross-layer residuals) enhances the model's recognition accuracy for variable farmland morphology (such as leaf occlusion, varying lighting conditions, and soil contamination). This improves semantic segmentation accuracy, particularly for key features like corn stalks. This provides a foundation for precise crop center location.

[0292] 4. Point cloud preprocessing (through filtering and outlier filtering) effectively removes interference from environmental noise (such as rain, fog, dust, etc.), improves the quality and credibility of point cloud data, and ensures the accuracy of subsequent fusion and center point calculation.

[0293] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technical solution that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A tracking and navigation method for locating the center point of a crop based on laser radar and image fusion, characterized in that: include: Step 1: Use a laser radar to scan the target farmland to generate laser radar point cloud data, and perform through-filtering and outlier filtering on the laser radar point cloud data to obtain filtered laser radar point cloud data; Step 2: Use a camera to capture a target farmland image, and perform distortion correction on the captured target farmland image to obtain a distortion-corrected farmland image; Step 3: Annotate the distortion-corrected farmland image to generate training samples, which are then fed into a U-net model for training to obtain a crop segmentation model. Dynamic convolution is used in place of standard convolution in the convolutional layers of the U-net model, and cross-layer residual connections are introduced in each decoder layer. Step 4: Use the crop segmentation model to perform semantic segmentation on the farmland image to classify crops, weeds, and ground; Step 5: Calibrate the imaging parameters of the camera and lidar to obtain the calibrated imaging parameters; Step 6: Use the calibrated imaging parameters to perform feature fusion on the filtered lidar point cloud data and the distortion-corrected farmland image to obtain a crop point cloud with semantic information; Step 7: Compute the crop center point based on the crop point cloud with semantic information; Step 8: Use the crop center point as a pre-aiming point to provide navigation for the agricultural robot.

2. The tracking and navigation method for locating the center point of a crop based on laser radar and image fusion according to claim 1, characterized in that: Step 5: calibrating the imaging parameters of the camera and the lidar to obtain calibrated imaging parameters, including: Step 5.1: Use a camera to capture the black and white grid calibration plate to form a calibration image, and use a radar laser to scan the black and white grid calibration plate to obtain a calibration point cloud; Step 5.2: Calibrate the camera intrinsic parameters using the calibration image and obtain the camera intrinsic parameter matrix; Step 5.3: Solve the camera distortion coefficient using the camera intrinsic parameter matrix; Step 5.4: Calibrate the camera's extrinsic parameters based on the calibration point cloud to obtain the camera's extrinsic parameters; Step 5.5: Jointly optimize the camera intrinsic parameter matrix, camera distortion coefficient, and camera extrinsic parameters to obtain the calibrated camera intrinsic parameters, camera distortion coefficient, and camera extrinsic parameters; Step 5.6: Verify the calibration accuracy of the camera intrinsic parameters, camera distortion coefficients, and camera extrinsic parameters. If the calibration accuracy does not meet the preset conditions, recalibrate.

3. The tracking and navigation method for locating the center point of a crop based on laser radar and image fusion according to claim 2, characterized in that: In step 5.5, use the objective function: The camera intrinsic parameter matrix, camera distortion coefficient and camera extrinsic parameter are jointly optimized to obtain the calibrated camera intrinsic parameter, camera distortion coefficient and camera extrinsic parameter; among them, represents the optimization target, M represents the number of image frames, N represents the number of feature points in each frame, and p ij represents the actual observed pixel coordinates of the jth feature point in the i-th frame image, π( ) represents the camera projection function, K represents the camera intrinsic parameter matrix, dist represents the distortion parameter set, R i , t i Represents the camera external parameters of the i-th frame image, P j represents the j-th 3D point coordinate in the world coordinate system, Q j represents the observation coordinates of the jth point in the laser radar coordinate system, T lc represents the homogeneous transformation matrix from the camera coordinate system to the lidar coordinate system, φ(R i , t i , P j ) represents the function that transforms the world coordinate system point to the camera coordinate system through the extrinsic parameters of the i-th frame.

4. The tracking and navigation method for locating the center point of a crop based on laser radar and image fusion according to claim 1, characterized in that: Step 6: Using the calibrated imaging parameters to perform feature fusion on the filtered lidar point cloud data and the distortion-corrected farmland image to obtain a crop point cloud with semantic information, including: Step 6.1: Extract the timestamp of the lidar point cloud and the midpoint timestamp of the camera exposure, and use the interpolation compensation formula to synchronize the time; Step 6.2: Perform motion distortion correction on the time-synchronized lidar point cloud data to obtain corrected lidar point cloud data; Step 6.3: Use the calibrated imaging parameters to transform the LiDAR coordinate system corresponding to the corrected LiDAR point cloud data into the camera coordinate system, completing the association between the point cloud data and the pixel data. Step 6.4: Perform 3D semantic reconstruction on the associated data to obtain a reconstructed 3D point cloud with semantic information; Step 6.5: Remove invalid point clouds other than the target crop from the 3D point cloud with semantic information to obtain the crop point cloud with semantic information.

5. The tracking and navigation method for locating the center point of a crop based on laser radar and image fusion according to claim 4, characterized in that: In step 6.3, the formula is used: ||Reflectance(P l )-Grayscale(u,v)||) Complete the association between point cloud data and pixel data; where P l Represents the lth 3D point in the point cloud, Prok(P l ) means point P l Function for projection onto the image plane, [u, v] T Represents the candidate pixel coordinates, Reflectance(P l ) represents point P l Grayscale(u, v) represents the grayscale value of the image at pixel (u, v), λ represents the weight coefficient, (u * , v * ) represents the optimal matching pixel coordinates, P represents the candidate 3D point, Represents the gradient of the image at (u, v), represents the depth gradient at point P, μ represents the weight coefficient, P * Represents the best matching 3D point.

6. The tracking and navigation method for locating the center point of a crop based on laser radar and image fusion according to claim 5, characterized in that: In step 6.4, use the formula: The associated data is reconstructed into 3D semantics to obtain a reconstructed 3D point cloud with semantic information; where TSDF(v) is the truncated signed distance function value at voxel v, v is the voxel position in 3D space, i is the observation data index, z i is the depth value of the i-th observation, d i is the distance from voxel v to the i-th observation viewpoint, o i is the optical center position of the camera at the i-th observation, w i is the fusion weight of the i-th observation, C l is the confidence of the lidar data, C v is the confidence level of the visual data.

7. The tracking and navigation method for locating the center point of a crop based on laser radar and image fusion according to claim 6, characterized in that: The step 7: computing the crop center point based on the crop point cloud with semantic information includes: Step 7.1: Remove invalid point clouds from the crop point cloud with semantic information and select the point cloud at the target crop stem; Step 7.2: Grid the selected point cloud to obtain a gridded density map; Step 7.3: Find the density peak area in the gridded density map, determine the edge points of the area from the density peak area, and select three edge points that meet the preset conditions from the edge points; Step 7.4: Fit a circle through the three edge points and use the coordinates of the circle center as the crop center point.

8. The tracking and navigation method for locating the center point of a crop based on laser radar and image fusion according to claim 7, characterized in that: In step 7.1, it includes: Remove invalid point clouds in the ground plane through point cloud height filtering; the point cloud height filtering formula is: in, is the original point cloud set, p i is a point in the point cloud, z i Point p i The z coordinate, μ z is the mean height of the point cloud, σ z is the standard deviation of the point cloud height, is the valid point cloud set retained after height filtering; Through semantic label filtering, other invalid point clouds are removed; the semantic label filtering formula is: in, Represents the effective point cloud set after height filtering, p i express A point in i The semantic probability distribution vector represented by max(s i ) represents the point p i The maximum value in the semantic probability distribution, s corn represents the semantic category corresponding to the maximum value, τ p represents the probability threshold, Represents the final set of point clouds that belong to the target crop category.

9. The tracking and navigation method for locating the center point of a crop based on laser radar and image fusion according to claim 8, characterized in that: Step 7.3: Finding a density peak region from the gridded density map, determining edge points of the region from the density peak region, and selecting three edge points that meet preset conditions from the edge points, including: Step 7.4.1: Find the density peak area in the gridded density map. The density peak area is defined as: in, represents the density peak area, ρ max Represents the maximum density value of the entire density map, ρ(x, y) represents the density value of the grid cell (x, y), cell(x, y) represents the unit area in the two-dimensional grid, and p i Indicates the point belonging to this grid cell, z i Indicates the height value of the point, μ z represents the mean height of the points within the grid cell, σ z represents the standard deviation of the height of the points within the grid cell, and exp represents the Gaussian weight function; Step 7.4.2: Calculate the edge points of the density peak area using the polar coordinate edge detection method; the edge point calculation formula is: Among them, θ k is the angle of the kth search direction, K is the total number of directions, k is the direction index, and p k is the direction θ k The most edge point on x i ,y i For point p i The plane coordinate, x p ,y p is the density peak coordinate, cosθ k ,sinθ k is the direction vector component; Step 7.4.3: From the edge points of the density peak area, select the points that meet the selection strategy: Three edge points; where pargmax represents the selection of polar coordinate angle θ from all edge points k The largest point, the angle θ k The largest point is defined as point A, and point A is selected as the first edge point. After the relative angle is offset by 120°, point B with the largest value is selected as the second edge point. After the relative angle is offset by 240°, point C with the largest value is selected as the third edge point. k is the polar angle of the point.

10. The tracking and navigation method for locating the center point of a crop based on laser radar and image fusion according to claim 9, characterized in that: Step 8: Using the crop center point as a pre-aiming point to provide navigation for the agricultural robot includes: The nearest crop center point is used as a preview point for navigation. If the preview point is no longer visible, the next nearest crop center point is searched for as a preview point to continue navigation. The strategy for selecting the nearest crop center point is as follows: Among them, P nearest The optimal crop center point represents the target crop center position that the navigation system will aim at, P i The center point of the candidate crop represents the center coordinate of the i-th target crop detected by the sensor, V is the real-time coordinate of the center point of the agricultural robot, θ v is the heading of the agricultural robot, θ row represents the angle between the crop and the current direction of the agricultural robot, and λ represents the weight coefficient, which is used to balance the weight of distance and direction deviation.

Citation Information

Cited By

  • AGV environment sensing method and system based on laser beams

    CN120927009A

  • Blind guiding robot path planning method and system fusing laser radar and camera semantics

    CN121540170A

  • Unmanned aerial vehicle DSM-based paddy field positioning model construction method, positioning method and system

    CN121810806A

  • Garden-oriented graph-free autonomous navigation method based on laser radar, plant protection method and robot

    CN121957027A