A citrus recognition and positioning method, device, equipment and storage medium

通过YOLOV4网络和激光雷达相结合的方法,解决了柑橘识别定位在复杂场景下的精度和实时性问题,实现了高效的柑橘目标检测和定位。

CN114332689BActive Publication Date: 2025-07-11HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111527626.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-14
Publication Date
2025-07-11
Estimated Expiration
2041-12-14

AI Technical Summary

Technical Problem

The existing citrus identification and positioning methods have poor recognition accuracy when occlusion is severe in complex scenarios, the positioning process calculation complexity is high, and the positioning accuracy and real-timeness are difficult to guarantee.

Method used

The YOLOV4 network is used to obtain the center position information of citrus, combine camera internal parameters and lidar external parameters calibration, and project point clouds onto the image through coordinate transformation matrix, fuse the image and point cloud data to achieve the positioning of citrus.

Benefits of technology

It improves the accuracy and real-time nature of citrus recognition, reduces the calculation amount of the positioning process, adapts to citrus detection in complex scenarios, and reduces dependence on light.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332689B_ABST
    Figure CN114332689B_ABST
Patent Text Reader

Abstract

The present invention discloses a citrus recognition and positioning method, device, equipment and storage medium. The method includes: inputting the collected image into the YOLOV4 network, and using the YOLOV4 network to obtain the position information of the citrus center in pixel coordinates; calibrating the internal parameters of the camera; calibrating the external parameters of the camera and the lidar; fusing the point cloud and the image by combining the obtained internal and external parameters, and projecting the point cloud onto the image by using the coordinate transformation matrix; finding the point cloud corresponding to the target citrus to obtain its depth value information, and completing the positioning of the citrus. The advantages of the present invention are: relatively high recognition accuracy for citrus, small calculation amount in the positioning process, and ensuring positioning accuracy and real-time performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision / multi-sensor data fusion, and more particularly to a citrus recognition and positioning method, device, equipment and storage medium. Background Art

[0002] The recognition and positioning of citrus are important parts for a picking robot to achieve automatic picking, which are mainly divided into two parts: target detection and target positioning. With the development and application of deep learning, target detection networks based on deep learning have emerged. In terms of target detection, traditional citrus recognition relies on the conversion of color spaces and image segmentation clustering to distinguish citrus fruits from the background. Such methods have poor detection accuracy for citrus with severe occlusion in complex scenarios. Using a convolutional neural network to automatically extract feature information of the target area can adapt to complex natural environments and has stronger generalization ability. However, a convolutional neural network usually runs slowly in target detection and it is difficult to balance detection speed and detection accuracy. In terms of target positioning, the method of using a binocular camera to calculate the distance using parallax is mostly adopted to obtain the target position information. However, a binocular camera is too sensitive to environmental light, is not applicable to monotonous scenes lacking texture, and has high computational complexity, making it difficult to ensure accuracy and real-time performance.

[0003] For example, Chinese Patent Publication No. CN109711317A discloses a method for segmenting and recognizing mature citrus fruits and branches based on regional features. First, a feature vector is generated based on the color features of a color image, and a feature mapping table is used to reduce the dimension of the color features to reduce the dimension of the feature vector. Then, the ROI size of the target object is determined based on the working space of the picking robot, the field of view size of the binocular camera, and the size of the citrus fruit. The proportion of the number of pixel points in the target range in the R and B channels is used as the basis for selecting the ROI. Finally, the ROIs with a large degree of overlap among the obtained multiple preliminary ROIs are sorted by score, and the ROI with the highest score is selected as the best segmentation and recognition area. The test results of this patent application show that under the condition of light change, the comprehensive recognition accuracy of this method for citrus fruits, background and branches reaches 94%, and the single-image segmentation time reaches 0.2 s, meeting the real-time requirement. However, this patent application relies on the conversion of color spaces and image segmentation clustering to distinguish citrus fruits from the background, has poor detection accuracy for citrus with severe occlusion in complex scenarios, uses a binocular camera to calculate the distance using parallax to obtain the target position information, but a binocular camera is too sensitive to environmental light, is not applicable to monotonous scenes lacking texture, and has high computational complexity, making it difficult to ensure accuracy and real-time performance. Summary of the Invention

[0004] The technical problem to be solved by the present invention is that the existing citrus recognition and positioning method has poor recognition accuracy for citrus with serious occlusion in complex scenes, high computational complexity in the positioning process, and it is difficult to ensure positioning accuracy and real-time performance.

[0005] The present invention solves the above technical problems through the following technical means: A citrus recognition and positioning method, the method comprising:

[0006] Step 1: Input the collected image into the YOLOV4 network, and use the YOLOV4 network to obtain the position information of the citrus center in pixel coordinates;

[0007] Step 2: Calibrate the internal parameters of the camera;

[0008] Step 3: Calibrate the external parameters of the camera and the lidar;

[0009] Step 4: Combine the obtained internal and external parameters to fuse the point cloud and the image, and project the point cloud onto the image using the coordinate transformation matrix;

[0010] Step 5: Find the point cloud corresponding to the target citrus to obtain its depth value information, and complete the positioning of the citrus.

[0011] The present invention inputs the collected image into the YOLOV4 network to obtain the position information of the citrus center in pixel coordinates. The YOLOV4 network has a faster recognition speed and higher recognition accuracy compared to other networks. It applies the lidar widely used in the autonomous driving scenario to citrus positioning. The position information of the point cloud output by scanning the target is more accurate and has higher real-time performance compared to the binocular camera. It fuses the output data of the radar and the camera, projects the point cloud onto the image after completing the joint calibration of the lidar and the camera, establishes a corresponding relationship between the pixels and the point cloud of the target, and processes the position information of the pixel and point cloud data to achieve target positioning. The computational amount in the positioning process is small, further improving the positioning accuracy and real-time performance.

[0012] Further, the step 1 includes:

[0013] The YOLOV4 network divides an image into S×S grids, multiplies the class information predicted by each grid and the confidence truth value of the prediction box containing the object, and the result is the coincidence degree between the prediction box and the truth value and the probability that the object belongs to a certain class; in the final output of the YOLOV4 network, each prediction box contains the position information of the object, that is, the center point coordinates and side length parameters of the prediction box. Thus, the detection of the citrus and the acquisition of the position information of the citrus center in pixel coordinates are completed using the YOLOV4 network.

[0014] Even further, the step 2 includes:

[0015] Define oxy as the image coordinate system, O cis the optical center of the camera, O c X c Y c is the world coordinate system where the camera is located, oO c The distance to is f, then through the formula

[0016]

[0017]

[0018] Solve the transformation relationship between the world coordinate system and the image coordinate system;

[0019] Convert the image coordinate system to the pixel coordinate system. Assume that the pixel coordinate system is scaled by α times on the x-axis and β times on the y-axis, and the origin is translated by [c x , c y T , then the point [u, v] on the pixel coordinate system T is expressed as:

[0020]

[0021] Substitute Equation (1) into Equation (3) and combine αf into f x , and combine βf into f y , we get:

[0022]

[0023] Convert Equation (3) into matrix form:

[0024]

[0025] The middle matrix of Equation (5) is the internal parameter matrix of the required camera.

[0026] Furthermore, the second step further includes:

[0027] Considering the non-linear distortion of the camera, assume any point p on the normalized plane, with coordinates [x, y] T , [x distored , y distored T is the normalized coordinate of the distorted point, r is the distance between point p and the coordinate origin, then

[0028] x distored = x(1 + k1r 2 + k2r 4 + k3r 6 ) (6)

[0029] y distored = y(1 + k1r 2 + k2r​​4 +k3r 6 ) (7)

[0030] In addition, tangential distortion is corrected using two other parameters:

[0031] x distored = x + 2p1xy + p2(r 2 + 2x 2 ) (8)

[0032] y distored = y + p1(r 2 + 2y 2 ) + 2p2xy (9)

[0033] where k1, k2, k3, p1, p2 are the five distortion parameters of the camera;

[0034] For the internal parameter calibration, the checkerboard calibration method is adopted. The findChessboardCorners function in OpenCV is used to extract the inner corner points of the calibration board. On the premise of knowing the size of the calibration board, the corresponding relationship between the three-dimensional space points and the pixel points is established to solve each parameter in the above formula. The solution of the internal parameter matrix and the distortion parameters completes the internal parameter calibration of the camera.

[0035] Furthermore, the third step includes:

[0036] The image data captured by the camera is represented by (u, v), and the point cloud position information captured by the lidar is represented by (X, Y, Z). The conversion relationship between the two is expressed as

[0037]

[0038] where f x , f y , c x , c y are the internal parameter matrix parameters of the camera, and R, t are the rotation and translation matrices of the relative pose between the camera and the lidar. The process of external parameter calibration is the process of solving the parameters R, t;

[0039] Solve the center point coordinates, plane normal vectors, and four corner point coordinates of the calibration board in the lidar coordinate system and the camera coordinate system. After collecting multiple groups of data at different positions, construct an objective function to optimize and solve the external parameters, and obtain the parameters R, t.

[0040] Furthermore, the fourth step includes: projecting the point cloud position information captured by the lidar onto the image data captured by the camera, and fusing the image and point cloud information. The fused model not only retains the original RGB image information but also contains the position and depth value information of the lidar point cloud.

[0041] Further, the fifth step includes:

[0042] Assume that a point P(Xc, Yc, Zc) in the camera coordinate system is the three-dimensional spatial coordinate of the center point of the citrus fruit, and its corresponding coordinate in the pixel coordinate system is (u, v). After fusing the image and the point cloud, depth value information is assigned to the inside of the prediction box by the lidar. Therefore, the depth value Zc of point P is measured by the lidar, and the pixel coordinates (u, v) corresponding to point P are the center point coordinates of the prediction box of the citrus fruit detected and output by the YOLOV4 network. Based on the above information, the expressions for Xc and Yc are solved as follows:

[0043]

[0044] Obtaining the coordinate values of point P(Xc, Yc, Zc) completes the positioning of the citrus fruit.

[0045] The present invention also provides a citrus fruit recognition and positioning device, and the device includes:

[0046] A pixel coordinate recognition module, configured to input the collected image into the YOLOV4 network, and use the YOLOV4 network to obtain the position information of the center of the citrus fruit in the pixel coordinates;

[0047] An internal parameter calibration module, configured to calibrate the internal parameters of the camera;

[0048] An external parameter calibration module, configured to calibrate the external parameters of the camera and the lidar;

[0049] A projection module, configured to fuse the point cloud and the image by combining the obtained internal and external parameters, and project the point cloud onto the image using a coordinate transformation matrix;

[0050] A positioning module, configured to find the point cloud corresponding to the target citrus fruit to obtain its depth value information, and complete the positioning of the citrus fruit.

[0051] Further, the pixel coordinate recognition module is further configured to:

[0052] The YOLOV4 network divides an image into S×S grids, and the class information predicted by each grid is multiplied by the confidence true value of the prediction box containing the object. The result is the overlap degree between the prediction box and the true value and the probability that the object belongs to a certain class; in the final output of the YOLOV4 network, the position information of the object contained in each prediction box, that is, the center point coordinates and side length parameters of the prediction box, thus completing the detection of the citrus fruit using the YOLOV4 network and obtaining the position information of the center of the citrus fruit in the pixel coordinates.

[0053] Furthermore, the internal parameter calibration module is further configured to:

[0054] Define oxy as the image coordinate system, and O c as the optical center of the camera, and Oc X c Y c is the world coordinate system where the camera is located, oO c The distance is f, then through the formula

[0055]

[0056]

[0057] Solve the transformation relationship between the world coordinate system and the image coordinate system;

[0058] Convert the image coordinate system to the pixel coordinate system. Assume that the pixel coordinate system is scaled by α times on the x-axis and β times on the y-axis, and the origin is translated by [c x , c y T , then the point [u, v] on the pixel coordinate system T is expressed as:

[0059]

[0060] Substitute Equation (1) into Equation (3) and combine αf into f x , and combine βf into f y , we get:

[0061]

[0062] Convert Equation (3) into matrix form:

[0063]

[0064] The middle matrix of Equation (5) is the internal parameter matrix of the required camera.

[0065] Furthermore, the internal parameter calibration module is also used for:

[0066] Considering the non-linear distortion of the camera, assume any point p on the normalized plane, with coordinates [x, y] T , [x distored , y distored T is the normalized coordinate of the distorted point, r is the distance between point p and the coordinate origin, then

[0067] x distored = x(1 + k1r 2 + k2r 4 + k3r 6 ) (6)

[0068] y distored = y(1 + k1r 2 + k2r​​4 +k3r 6 ) (7)

[0069] In addition, tangential distortion is corrected using two other parameters:

[0070] x distored = x + 2p1xy + p2(r 2 + 2x 2 ) (8)

[0071] y distored = y + p1(r 2 + 2y 2 ) + 2p2xy (9)

[0072] where k1, k2, k3, p1, and p2 are the five distortion parameters of the camera;

[0073] For the internal parameter calibration, the checkerboard calibration method is adopted. The internal corner points of the calibration board are extracted using the findchessboardCorners function in OpenCV. On the premise of knowing the size of the calibration board, the corresponding relationship between the three-dimensional space points and the pixel points is established to solve each parameter in the above formula. The solution of the internal parameter matrix and the distortion parameters completes the internal parameter calibration of the camera.

[0074] Furthermore, the external parameter calibration module is also used for:

[0075] The image data captured by the camera is represented by (u, v), and the point cloud position information captured by the lidar is represented by (X, Y, Z). The conversion relationship between the two is expressed as

[0076]

[0077] where f x , f y , c x , c y are the internal parameter matrix parameters of the camera, and R and t are the rotation and translation matrices of the relative pose between the camera and the lidar. The process of external parameter calibration is the process of solving the parameters R and t;

[0078] Solve the center point coordinates, plane normal vector, and four corner point coordinates of the calibration board in the lidar coordinate system and the camera coordinate system. After collecting multiple groups of data at different positions, construct an objective function to optimize and solve the external parameters, and obtain the parameters R and t.

[0079] Furthermore, the projection module is also used for: projecting the point cloud position information captured by the lidar onto the image data captured by the camera, and fusing the image and point cloud information. The fused model not only retains the original RGB image information but also contains the position and depth value information of the lidar point cloud.

[0080] Furthermore, the positioning module is also used for:

[0081] Assume that a point P(Xc, Yc, Zc) in the camera coordinate system is the three-dimensional space coordinate of the center point of the citrus fruit, and its corresponding coordinate in the pixel coordinate system is (u, v). After fusing the image and the point cloud, depth value information is given to the inside of the prediction box by the lidar. Therefore, the depth value Zc of point P is measured by the lidar, and the pixel coordinates (u, v) corresponding to point P are the center point of the prediction box of the citrus fruit detected and output by the YOLOV4 network. Based on the above information, the expressions for Xc and Yc are solved as follows:

[0082]

[0083] Obtaining the coordinate values of point P(Xc, Yc, Zc) completes the positioning of the citrus fruit.

[0084] The present invention also provides an electronic device, including a processor and a memory. The memory stores computer program instructions that can be executed by the processor. When the processor executes the computer program instructions, the method steps described above are implemented.

[0085] The present invention also provides a computer-readable storage medium, storing computer program instructions. When the computer program instructions are called and executed by a processor, the method steps described above are implemented.

[0086] The advantages of the present invention are as follows:

[0087] (1) In the present invention, the collected image is input into the YOLOV4 network to obtain the position information of the center of the citrus fruit in the pixel coordinates. The YOLOV4 network has a faster recognition speed and higher recognition accuracy compared to other networks. The lidar widely used in the autonomous driving scenario is transferred to citrus positioning. The position information of the point cloud output by scanning the target is more accurate and has higher real-time performance compared to the binocular camera. The output data of the lidar and the camera are fused. After completing the joint calibration of the lidar and the camera, the point cloud is projected onto the image, and a corresponding relationship is established between the pixels and the point cloud of the target. The position information of the pixel and the point cloud data is processed to achieve target positioning. The calculation amount in the positioning process is small, further improving the positioning accuracy and real-time performance.

[0088] (2) In order to make the camera calibration result more accurate, the non-linear distortion of the camera should be considered during camera calibration to correct the ideal projection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] Figure 1 It is a schematic diagram of the YOLOV4 network structure in a citrus recognition and positioning method disclosed in an embodiment of the present invention;

[0090] Figure 2The flow chart of citrus positioning in a citrus recognition and positioning method disclosed in an embodiment of the present invention;

[0091] Figure 3 The schematic diagram of a pinhole camera model in a citrus recognition and positioning method disclosed in an embodiment of the present invention;

[0092] Figure 4 The schematic diagram of the basic principle of the joint calibration of a camera and a lidar in a citrus recognition and positioning method disclosed in an embodiment of the present invention. Specific embodiments

[0093] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0094] Embodiment 1

[0095] A citrus recognition and positioning method, the method comprising:

[0096] Step 1: Input the collected image into the YOLOV4 network, and use the YOLOV4 network to obtain the position information of the center of the citrus in the pixel coordinates;

[0097] Step 2: Intrinsic calibration of the camera;

[0098] Step 3: Extrinsic calibration of the camera and the lidar;

[0099] Step 4: Combine the obtained intrinsic and extrinsic parameters to fuse the point cloud and the image, and project the point cloud onto the image using the coordinate transformation matrix;

[0100] Step 5: Find the point cloud corresponding to the target citrus to obtain its depth value information, and complete the positioning of the citrus. The specific processes of each step are introduced in detail in the following sections.

[0101] 1. Citrus detection

[0102] The core idea of the YOLO algorithm is to use the entire image as the input of the network, divide an image into S×S grids, and if the center of a target to be detected is located in this grid, then this grid is responsible for detecting the target. The product of the class information predicted by each grid and the confidence truth value of the bounding box containing the object is the coincidence degree between the predicted box and the truth value and the probability that the object belongs to a certain class. The formula is as follows:

[0103]

[0104] where Pr(Class i |Object) represents the probability that the target belongs to a certain class, represents the confidence ground truth that the bounding box (prediction box) contains the object, represents the intersection of the prediction value and the ground truth. After obtaining the confidence score of each bounding box, set a threshold to remove the part with a lower score, and perform NMS processing on the remaining bounding boxes to obtain the final detection result. The detection result includes three parts: the type information of the target, the coordinate information of the target, and the target class probability. Its network structure is as Figure 1 shown.

[0105] In the final output of the YOLOV4 network, each bounding box contains the position information of the object, that is, the center point coordinates and side length parameters of the bounding box. So far, the detection of citrus and the acquisition of the position information of the center of the citrus in the pixel coordinates have been completed using the YOLOV4 network, but its depth value has not been obtained yet. The next step is to obtain the depth value information of the citrus.

[0106] 2. Citrus localization

[0107] To obtain the depth value information of the citrus, it is roughly divided into the following steps: 1. Intrinsic calibration of the camera, 2. Extrinsic calibration of the camera and lidar, 3. Fuse the point cloud and the image by combining the obtained intrinsic and extrinsic parameters, project the point cloud onto the image using the coordinate transformation matrix, 4. Find the point cloud corresponding to the target citrus to obtain its depth value information, that is, complete the localization of the citrus. The flow chart is as Figure 2 shown.

[0108] Two models are mainly used for the intrinsic calibration of the camera: the pinhole model and the distortion model.

[0109] Figure 3 is the pinhole camera model, where the oxy coordinates are the image coordinate system, and O c is the optical center of the camera. To make the model more in line with the actual situation, the imaging plane oxy can be equivalently placed symmetrically in front of the camera, and together with the three-dimensional world space point P on the same side of the camera coordinate system. Given that △ABO c is similar to △OCO c , and △PBO c is similar to △pCO c , it can be deduced that:

[0110]

[0111]

[0112]

[0113] The above formula completes the solution of the transformation relationship between the world coordinate system and the image coordinate system, and then converts the image coordinate system into the pixel coordinate system. There is a difference in scaling and origin translation between the pixel coordinate system and the imaging plane. Suppose the pixel coordinate system is scaled by α times on the x-axis and β times on the y-axis, and the origin is translated by [c x , c y T , then the point [u, v] on the pixel coordinate system T can be expressed as:

[0114]

[0115] Substitute Equation (1) and combine αf into f x , and combine βf into f y , to get:

[0116]

[0117] Convert Equation (3) into matrix form:

[0118]

[0119] The middle matrix of Equation (5) is the internal parameter matrix of the required camera.

[0120] The ideal camera model is the pinhole model, but the actual lens does not conform to this assumption. In order to make the camera calibration result more accurate, the non-linear distortion of the camera should be taken into account when performing camera calibration to correct the ideal projection model. Assume an arbitrary point p on the normalized plane, and its coordinates are [x, y] T , [x distored , y distored T is the normalized coordinate of the distorted point. It is usually assumed that these distortions are polynomial relationships, and r is the distance between point p and the coordinate origin.

[0121] x distored = x(1 + k1r 2 + k2r 4 + k3r 6 ) (6)

[0122] y distored = y(1 + k1r 2 + k2r 4 + k3r 6 ) (7)

[0123] In addition, the tangential distortion can be corrected with two other parameters:

[0124] x distored ​​= x + 2p1xy + p2(r 2 + 2x 2 ) (8)

[0125] y distored = y + p1(r 2 + 2y 2 ) + 2p2xy (9)

[0126] In summary, the distortion of the camera can be represented by five parameters (k1, k2, k3, p1, p2). The internal parameter calibration adopts the checkerboard calibration method. The findChessboardCorners function in OpenCV is used to extract the internal corner points of the calibration board. On the premise of knowing the size of the calibration board, the corresponding relationship between the three-dimensional space points and the pixel points is established to solve the parameters in the above formula. The solution of the internal parameter matrix and the distortion parameters completes the internal parameter calibration of the camera.

[0127] The solved internal parameter matrix of the camera is:

[0128]

[0129] The distortion coefficients (k1, k2, k3, p1, p2) are: -0.063009, 0.163677, -0.000323, 0.001588, 0.000000.

[0130] The basic principle model of the joint calibration of the camera and the lidar is as Figure 4 shown. The image data captured by the camera is represented by (u, v), and the point cloud position information captured by the lidar is represented by (X, Y, Z). The conversion relationship can be expressed as:

[0131]

[0132] where f x , f y , c x , c y are the internal parameter matrix parameters of the camera, and R, t are the rotation and translation matrices of the relative pose between the camera and the lidar. The process of joint calibration is the process of solving the parameters R, t. To remove the uninteresting regions in the lidar point cloud data, the present invention uses rqt_reconfigure to dynamically adjust the size of each coordinate limit value in the lidar coordinate system, obtain the ROI of the point cloud data, reduce the possibility of false detection, and facilitate the fitting of the calibration board plane. Even without an empty calibration site, the joint calibration work can be completed more accurately.

[0133] The fitting of lidar point cloud adopts the Random Sample Consensus algorithm (RANSAC). It generates candidate solutions by estimating the minimum number of observations (data points) required for the basic model parameters to fit the calibration board point cloud. However, in fact, the segmented and fitted point cloud is not on an exact plane. The segmented point cloud is projected onto the fitted plane through the ProjectInliers function, and the normal vector of the calibration board is obtained according to the fitting result. The starting and ending points of each line of the calibration board point cloud are obtained, and the four sides of the calibration board point cloud are obtained by using the Random Sample Consensus algorithm. The four corner points of the calibration board point cloud plane are obtained through the lineWithLineIntersection function, and then the center point coordinates are calculated. Thus, the coordinates of the four corner points, the center point coordinates, and the plane normal vector of the calibration board in the lidar coordinate system are obtained.

[0134] For the extraction of camera features, first convert the RGB image into a grayscale image, and use the findchessboardCorners function to extract the sub-pixel accuracy internal corner point data of the calibration board to find the center coordinates of the calibration board. Given the size information of the checkerboard, the pixel coordinates and the coordinates in the camera coordinate system of each edge corner point can be obtained, and the pose of the calibration board in the camera coordinate system is solved using the PnP algorithm to obtain the plane normal vector of the calibration board.

[0135] The above content respectively obtains the center point coordinates, plane normal vector, and four corner point coordinates of the calibration board in the lidar coordinate system and the camera coordinate system. After collecting multiple groups of data at different positions, a target function is constructed to optimize and solve the external parameters of the sensor. The rotation and translation matrices of the relative pose between the camera and the lidar are solved, where the R rotation matrix is represented in the form of Euler angles, namely roll roll angle, pitch pitch angle, yaw yaw angle, and t contains the translation amounts in the x, y, and z directions. The final results are as follows:

[0136] R = [-1.52033, 0.0242735, -1.50977] T ,

[0137] t = [1.93773, -0.741232, -0.144967] T

[0138] After obtaining the internal and external parameters, the point cloud can be projected onto the image to fuse the image and point cloud information. The fused model not only retains the original RGB image information but also contains the position and depth value information of the lidar point cloud.

[0139] It can be obtained from Equation (5). Assume that a point P(Xc, Yc, Zc) in the camera coordinate system is the three-dimensional spatial coordinate of the center point of the citrus. Its corresponding coordinate in the pixel coordinate system is (u, v). After fusing the image and the point cloud, depth value information is given to the inside of the bounding box by the lidar. Therefore, the depth value Zc of point P can be measured by the lidar, and the pixel coordinates (u, v) corresponding to point P are the center points of the bounding box of the citrus detected by the YOLOV4 network. Based on the above information, the expressions for Xc and Yc are solved as follows:

[0140]

[0141] The parameter f in Equation (11) x , f y , c x , c y can all be obtained by the aforementioned camera internal parameter calibration. After obtaining the coordinate values of point P(Xc, Yc, Zc), the positioning of the citrus is completed.

[0142] Through the above technical solutions, the present invention uses the YOLOV4 network for citrus target detection, adjusts the network parameters, enables it to improve the speed of citrus detection on the premise of ensuring the recognition accuracy, and is more suitable for citrus target detection in real scenarios. On this basis, the data information of the camera and the lidar is fused, the point cloud is projected onto the image, and depth information is given to the image, so as to realize the solution of the three-dimensional spatial position of the citrus target. Compared with the method of obtaining depth information using two cameras, the depth information of the point cloud data is more accurate and the processing calculation amount is smaller, not affected by light, and has certain advantages in positioning time consumption and accuracy, which better meets the technical requirements of citrus picking robots.

[0143] Embodiment 2

[0144] Based on Embodiment 1, the present invention also provides a citrus recognition and positioning device, which includes:

[0145] A pixel coordinate recognition module, configured to input the collected image into the YOLOV4 network and use the YOLOV4 network to obtain the position information of the citrus center in the pixel coordinates;

[0146] An internal parameter calibration module, configured to calibrate the internal parameters of the camera;

[0147] An external parameter calibration module, configured to calibrate the external parameters of the camera and the lidar;

[0148] A projection module, configured to fuse the point cloud and the image by combining the obtained internal and external parameters, and project the point cloud onto the image using the coordinate transformation matrix;

[0149] A positioning module, configured to find the point cloud corresponding to the target citrus to obtain its depth value information and complete the positioning of the citrus.

[0150] Specifically, the pixel coordinate recognition module is also used for:

[0151] The YOLOV4 network divides an image into SxS grids, and multiplies the predicted category information of each grid by the confidence value of the object contained in the prediction box. The result is the overlap between the prediction box and the true value and the probability that the object belongs to a certain class. In the final output of the YOLOV4 network, each prediction box contains the location information of the object, that is, the center point coordinates and side length parameters of the prediction box. At this point, the YOLOV4 network is used to complete the detection of citrus and obtain the location information of the center of the citrus in pixel coordinates.

[0152] More specifically, the internal reference calibration module is also used for:

[0153] Define oxy as the image coordinate system, O c is the optical center of the camera, O c X c Y c is the world coordinate system where the camera is located, oO c The distance is f, then the formula

[0154]

[0155]

[0156] Solve the transformation relationship between the world coordinate system and the image coordinate system;

[0157] Convert the image coordinate system to the pixel coordinate system. Suppose the pixel coordinate system is scaled by α times on the x-axis and β times on the y-axis. At the same time, the origin is translated by [c x , c y ] T , then the point [u, v] on the pixel coordinate system T It is expressed as:

[0158]

[0159] Substitute equation (1) into equation (3) and combine αf into f x , merge βf into f y ,have to:

[0160]

[0161] Convert equation (3) into matrix form:

[0162]

[0163] The middle matrix of formula (5) is the intrinsic parameter matrix of the required camera.

[0164] More specifically, the internal parameter calibration module is further configured to:

[0165] Considering the non - linear distortion of the camera, assume an arbitrary point p on the normalized plane with coordinates [x, y] T , [x distored , y distored T being the normalized coordinates of the distorted point, r being the distance between point p and the coordinate origin, then

[0166] x distored = x(1 + k1r 2 + k2r 4 + k3r 6 ) (6)

[0167] y distored = y(1 + k1r 2 + k2r 4 + k3r 6 ) (7)

[0168] In addition, for tangential distortion, it is corrected with two other parameters:

[0169] x distored = x + 2p1xy + p2(r 2 + 2x 2 ) (8)

[0170] y distored = y + p1(r 2 + 2y 2 ) + 2p2xy (9)

[0171] where k1, k2, k3, p1, p2 are the five distortion parameters of the camera;

[0172] The internal parameter calibration adopts the checkerboard calibration method. The findChessboardCorners function in OpenCV is used to extract the inner corner points of the calibration board. On the premise of knowing the size of the calibration board, the corresponding relationship between the three - dimensional space points and pixel points is established to solve each parameter in the above formula. The solution of the internal parameter matrix and distortion parameters completes the internal parameter calibration of the camera.

[0173] More specifically, the external parameter calibration module is further configured to:

[0174] The image data captured by the camera is represented by (u, v), and the point cloud position information captured by the lidar is represented by (X, Y, Z). Their conversion relationship is expressed as

[0175]

[0176] where f x , f​y c x c y are the internal parameter matrix parameters of the camera, R and t are the rotation and translation matrices of the relative pose between the camera and the lidar. The process of extrinsic parameter calibration is the process of solving the parameters R and t.

[0177] Solve the central point coordinates, plane normal vectors, and four corner point coordinates of the calibration board in the lidar coordinate system and the camera coordinate system. After collecting multiple groups of data at different positions, construct an objective function to optimize and solve the extrinsic parameters, and obtain the parameters R and t.

[0178] More specifically, the projection module is further configured to project the point cloud position information captured by the lidar onto the image data captured by the camera, fuse the image and point cloud information. The fused model not only retains the original RGB image information but also contains the position and depth value information of the lidar point cloud.

[0179] More specifically, the positioning module is further configured to:

[0180] Assume that a point P(Xc, Yc, Zc) in the camera coordinate system is the three-dimensional space coordinate of the center point of the citrus. Its corresponding coordinate in the pixel coordinate system is (u, v). After fusing the image and the point cloud, the depth value information is given to the inside of the prediction box by the lidar. Therefore, the depth value Zc of point P is measured by the lidar, and the pixel coordinate (u, v) corresponding to point P is the center point of the prediction box of the citrus detected by the YOLOV4 network. Based on the above information, the expressions of Xc and Yc are solved as follows:

[0181]

[0182] Obtaining the coordinate values of point P(Xc, Yc, Zc) completes the positioning of the citrus.

[0183] Embodiment 3

[0184] The present invention also provides an electronic device, including a processor and a memory. The memory stores computer program instructions that can be executed by the processor. When the processor executes the computer program instructions, the method steps described in Embodiment 1 are implemented.

[0185] Embodiment 4

[0186] The present invention also provides a computer-readable storage medium, storing computer program instructions. The computer program instructions implement the method steps described in Embodiment 4 when being called and executed by a processor.

[0187] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A citrus recognition and positioning method, characterized in that The method comprises: Step 1: Input the collected image into the YOLOV4 network, and use the YOLOV4 network to obtain the position information of the center of the citrus in pixel coordinates; Step 2: Camera internal calibration; Step 3: Calibrate the external parameters of the camera and lidar; the image data captured by the camera is represented by (u,v), and the point cloud position information captured by the lidar is represented by (X,Y,Z). The conversion relationship between the two is expressed as where f x , f y , c x , c y are the internal parameter matrix parameters of the camera, R and t are the rotation and translation matrices of the relative pose between the camera and the lidar, and the process of external parameter calibration is the process of solving the parameters R and t; Solve the coordinates of the center point, plane normal vector, and four corner points of the calibration plate in the laser radar coordinate system and the camera coordinate system. After collecting multiple sets of data at different positions, construct an objective function to optimize and solve the external parameters, and solve the parameters R, t; Step 4: Combine the required internal and external parameters to fuse the point cloud and the image, and use the coordinate transformation matrix to project the point cloud onto the image; project the point cloud position information captured by the lidar onto the image data captured by the camera, and fuse the image and point cloud information. The fused model not only retains the original RGB image information, but also contains the position and depth value information of the lidar point cloud; Step 5: Find the point cloud corresponding to the target citrus to obtain its depth value information and complete the positioning of the citrus; Assume that a point P (Xc, Yc, Zc) in the camera coordinate system is the three-dimensional spatial coordinate of the center point of the citrus, and its corresponding coordinates in the pixel coordinate system are (u, v). After fusing the image with the point cloud, the depth value information is given to the inside of the prediction box by the laser radar, so the depth value Zc of point P is measured by the laser radar, and the pixel coordinates (u, v) corresponding to point P are the center point of the prediction box output by the YOLOV4 network detection. Combining the above information, the expression of Xc, Yc is solved as follows: The location of the orange is completed by obtaining the coordinate value of point P (Xc, Yc, Zc).

2. The citrus recognition and positioning method according to claim 1, characterized in that The step one comprises: The YOLOV4 network divides an image into SxS grids, and multiplies the predicted category information of each grid by the confidence value of the object contained in the prediction box. The result is the overlap between the prediction box and the true value and the probability that the object belongs to a certain class. In the final output of the YOLOV4 network, each prediction box contains the location information of the object, that is, the center point coordinates and side length parameters of the prediction box. At this point, the YOLOV4 network is used to complete the detection of citrus and obtain the location information of the center of the citrus in pixel coordinates.

3. The citrus recognition and positioning method according to claim 2, characterized in that The second step comprises: Define oxy as the image coordinate system, with O c being the optical center of the camera, and O c X c Y c being the world coordinate system in which the camera is located. If the distance from o to O c is f, then through the formula Solve the transformation relationship between the world coordinate system and the image coordinate system; Convert the image coordinate system to the pixel coordinate system. Assume that the pixel coordinate system is scaled by α times on the x-axis and β times on the y-axis, and the origin is translated by [c x , c y T . Then the point [u, v] in the pixel coordinate system T is expressed as:​ Substitute Equation (1) into Equation (3) and combine αf into f x , combine βf into f y , we get: Convert equation (3) into matrix form: The middle matrix of formula (5) is the intrinsic parameter matrix of the required camera.

4. The citrus recognition and positioning method according to claim 3, wherein The step 2 also includes: Considering the non - linear distortion of the camera, assume an arbitrary point p on the normalized plane with coordinates [x, y] T , [x distored , y distored T is the normalized coordinate of the distorted point, r is the distance between point p and the coordinate origin, then​ x distored = x(1 + k1r 2 + k2r 4 + k3r 6 ) (6) y distored = y(1 + k1r 2 + k2r 4 + k3r 6 ) (7) In addition, two other parameters are used to correct the tangential distortion: x distored = x + 2p1xy + p2(r 2 + 2x 2 ) (8) y distored = y + p1(r 2 + 2y 2 ) + 2p2xy (9) Among them, k1, k2, k3, p1, p2 are the five distortion parameters of the camera; The intrinsic calibration adopts the chessboard calibration method, and uses the findchessboardCorners function in OpenCV to extract the inner corner points in the calibration plate. Under the premise of knowing the size of the calibration plate, the correspondence between the three-dimensional space points and the pixel points is established to complete the solution of the parameters in the above formula. The solution of the intrinsic parameter matrix and the distortion parameter completes the intrinsic calibration of the camera.

5. A citrus recognition and positioning device, characterized in that, The device comprises: A pixel coordinate recognition module, which is used to input the collected image into the YOLOV4 network and utilize the YOLOV4 network to obtain the position information of the center of the citrus in pixel coordinates; An internal parameter calibration module, which is used for calibrating the internal parameters of the camera; An external parameter calibration module, which is used for calibrating the external parameters between the camera and the lidar; the image data captured by the camera is represented by (u, v), and the position information of the point cloud captured by the lidar is represented by (X, Y, Z), and the conversion relationship between the two is expressed as where f x , f y , c x , c y are the internal parameter matrix parameters of the camera, and R and t are the rotation and translation matrices of the relative pose between the camera and the lidar. The process of external parameter calibration is the process of solving the parameters R and t; Solve the center point coordinates, plane normal vector, and four corner point coordinates of the calibration board in the lidar coordinate system and the camera coordinate system. After collecting multiple groups of data at different positions, construct an objective function to optimize and solve the external parameters, and obtain the parameters R and t; A projection module, which is used to fuse the point cloud and the image by combining the obtained internal and external parameters, and project the point cloud onto the image using the coordinate transformation matrix; the position information of the point cloud captured by the lidar is projected onto the image data captured by the camera, fusing the image and point cloud information. The fused model not only retains the original RGB image information but also contains the position and depth value information of the lidar point cloud; A positioning module, which is used to find the point cloud corresponding to the target citrus to obtain its depth value information and complete the positioning of the citrus; assume that a point P(Xc, Yc, Zc) in the camera coordinate system is the three-dimensional space coordinate of the center point of the citrus, and its corresponding coordinate in the pixel coordinate system is (u, v). After fusing the image and the point cloud, the depth value information is given to the inside of the prediction box by the lidar. Therefore, the depth value Zc of point P is measured by the lidar, and the pixel coordinates (u, v) corresponding to point P are the center point of the prediction box of the citrus detected and output by the YOLOV4 network. Based on the above information, the expressions for Xc and Yc are solved as follows: Obtaining the coordinate values of point P(Xc, Yc, Zc) completes the positioning of the citrus.

6. An electronic device, characterized in that, It includes a processor and a memory. The memory stores computer program instructions that can be executed by the processor. When the processor executes the computer program instructions, the method steps described in any one of claims 1-4 are implemented.

7. A computer-readable storage medium, characterized in that, Stores computer program instructions, and when the computer program instructions are called and executed by the processor, the method steps described in any one of claims 1-4 are implemented.

Citation Information

Patent Citations

  • A region feature-based segmentation recognition method for mature citrus fruits and branches and leaves

    CN109711317A

  • Space non-cooperative target three-dimensional reconstruction method based on projection matrix

    CN107680159A

  • A method for object recognition and registration based on event triggered camera and three-dimensional laser radar fusion system

    CN109146929A

  • Maritime floating wharf detection method and system based on vision and laser radar fusion

    CN113436258A