Outdoor Apple Recognition and Localization Method and System Based on Multi-Sensor Fusion

Through multi-sensor fusion technology, combined with camera and lidar data, edge feature matching and fitting algorithms are used to solve the accuracy of Apple's three-dimensional recognition in strong light environments, and improve the robustness and accuracy of the picking task.

CN119360366BActive Publication Date: 2025-05-30SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411918158.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-30
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

The existing three-dimensional recognition methods are difficult to accurately identify apple's three-dimensional information in strong light environments, which affects the accuracy and stability of the picking task.

Method used

Using a multi-sensor fusion method, combining camera and lidar data, the edge feature matching algorithm and RANSAC fitting algorithm are used to accurately locate Apple's three-dimensional information under complex lighting conditions.

Benefits of technology

Improves Apple's robustness in strong light interference environments, achieves more accurate Apple position and shape estimation, and enhances the accuracy and stability of the picking task.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360366B_ABST
    Figure CN119360366B_ABST
Patent Text Reader

Abstract

The present disclosure provides an outdoor apple recognition and positioning method and system based on multi-sensor fusion, which relates to the technical field of target detection. It includes obtaining image data and lidar point cloud data of outdoor apples, inputting the image data into a target detection model to obtain an RGB image; obtaining an ROI depth information area through calibrating the external reference matrix; using an edge detection algorithm to extract image edge features from the RGB image to obtain an ROI edge information map; denoising and detecting discontinuous points in the ROI depth information area to obtain depth edge data; using a double-edge matching algorithm to perform edge feature matching on the depth edge data and the ROI edge information map, and using a RANSAC fitting algorithm to generate a smooth edge curve of the apple; constructing a sphere model according to the fitting result, and solving the center coordinates and radius by minimizing the sum of squared residuals to obtain the three-dimensional positioning information of the apple.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of object detection, and particularly to an outdoor apple recognition and positioning method and system based on multi-sensor fusion. Background Art

[0002] The statements in this part merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.

[0003] In the automated picking of fruits and vegetables, the three-dimensional information recognition technology of the target is crucial for improving the picking efficiency and accuracy. The three-dimensional recognition of apples is particularly challenging because apples are usually blocked by dense leaves and branches on the tree, and their complex surface reflectivity, variable-shaped contours, and visibility differences under different lighting conditions all pose problems for high-precision three-dimensional positioning and recognition. This makes it a key requirement for intelligent picking systems to obtain reliable three-dimensional information in various lighting environments.

[0004] Most of the current mainstream three-dimensional recognition methods are based on structured light or binocular cameras, and achieve precise positioning of the target by obtaining images on the camera plane and combining depth information. In these methods, structured light projects special patterns to estimate the scene depth, but in strong lighting environments, the projected light is often interfered by sunlight, resulting in a significant decline in the depth estimation effect of structured light or even complete failure. At the same time, binocular cameras rely on two sets of images for stereo matching to calculate depth. However, in strong light, light interference and shadow changes often lead to recognition and positioning errors of matching points, resulting in inaccurate depth estimation, which is usually unacceptable in picking tasks with high precision requirements. Moreover, the failure of three-dimensional information recognition under complex lighting conditions directly affects the accuracy and stability of apple recognition, and further brings unexpected challenges to the picking task. Summary of the Invention

[0005] In order to solve the above problems, the present disclosure proposes an outdoor apple recognition and positioning method and system based on multi-sensor fusion, which performs multi-modal fusion on the data of two sensors, uses an edge feature matching algorithm to accurately locate the three-dimensional information of apples under complex lighting conditions. At the same time, through the matching method of a planar lidar and multiple robotic arms, a more stable detection effect can be achieved with limited recognition hardware costs.

[0006] According to some embodiments, the present disclosure adopts the following technical solutions:

[0007] An outdoor apple recognition and positioning method based on multi-sensor fusion, comprising:

[0008] Obtaining image data and lidar point cloud data of outdoor apples and preprocessing them;

[0009] Input the image data into the target detection model to obtain the RGB image of the apple's two-dimensional bounding box; project the lidar point cloud into the camera coordinate system through the camera intrinsic parameters and the calibrated extrinsic parameter matrix of the multi-sensor to obtain the ROI depth information area that matches the two-dimensional bounding box.

[0010] Use the edge detection algorithm to extract the image edge features from the RGB image, and use the inverse distance weighted transformation to assign edge degrees to non-edge pixels, extending the neighborhood edge information to non-edge pixels to obtain the complete ROI edge information map; perform denoising and discontinuous point detection on the ROI depth information area to obtain the depth edge data under the denoised ROI point cloud.

[0011] Use the double-edge matching algorithm to perform edge feature matching on the depth edge data and the ROI edge information map, identify the true edge contour of the apple, and use the RANSAC fitting algorithm to generate the smooth edge curve of the apple.

[0012] Construct a sphere model according to the fitting result, solve the center coordinates and radius by minimizing the sum of squared residuals to obtain the three-dimensional positioning information of the apple.

[0013] According to some embodiments, the present disclosure adopts the following technical solutions:

[0014] An outdoor apple recognition and positioning system based on multi-sensor fusion, including:

[0015] A data acquisition module for acquiring the image data and lidar point cloud data of outdoor apples and performing preprocessing.

[0016] A target recognition module for inputting the image data into the target detection model to obtain the RGB image of the apple's two-dimensional bounding box; projecting the lidar point cloud into the camera coordinate system through the camera intrinsic parameters and the calibrated extrinsic parameter matrix of the multi-sensor to obtain the ROI depth information area that matches the two-dimensional bounding box.

[0017] An edge feature extraction module for using the edge detection algorithm to extract the image edge features from the RGB image, and using the inverse distance weighted transformation to assign edge degrees to non-edge pixels, extending the neighborhood edge information to non-edge pixels to obtain the complete ROI edge information map; performing denoising and discontinuous point detection on the ROI depth information area to obtain the depth edge data under the denoised ROI point cloud.

[0018] An edge matching module for using the double-edge matching algorithm to perform edge feature matching on the depth edge data and the ROI edge information map, identifying the true edge contour of the apple, and using the RANSAC fitting algorithm to generate the smooth edge curve of the apple.

[0019] A positioning calculation module, configured to construct a sphere model according to the fitting result, solve the center coordinates and radius by minimizing the sum of squared residuals, and obtain the three-dimensional positioning information of the apple.

[0020] According to some embodiments, the present disclosure adopts the following technical solutions:

[0021] A computer program product, including a computer program, which when executed by a processor implements the outdoor apple recognition and positioning method based on multi-sensor fusion described above.

[0022] According to some embodiments, the present disclosure adopts the following technical solutions:

[0023] A non-transitory computer-readable storage medium, which is used to store computer instructions, and when the computer instructions are executed by a processor, the outdoor apple recognition and positioning method based on multi-sensor fusion described above is implemented.

[0024] According to some embodiments, the present disclosure adopts the following technical solutions:

[0025] An electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory, so that the electronic device executes and implements the outdoor apple recognition and positioning method based on multi-sensor fusion described above.

[0026] Compared with the prior art, the beneficial effects of the present disclosure are:

[0027] The outdoor apple recognition and positioning method based on multi-sensor fusion of the present disclosure, an apple three-dimensional information recognition method based on a camera and a lidar, is used to improve the robustness of recognition in a strong light interference environment. The lidar has the characteristic of being unaffected by ambient light, can still maintain accurate depth detection under high-intensity light, and combines the image information of the camera to achieve more accurate apple position and shape estimation. By performing multi-modal fusion on the data of the two sensors of the camera and the lidar, the limitations of a single method under strong light can be effectively overcome, laying a technical foundation for the efficient recognition and picking of outdoor apple orchards. At the same time, through the matching method of the area array lidar and multiple robotic arms, a more stable detection effect can be achieved under the condition of limited recognition hardware cost. This innovative solution not only improves the adaptability of automated picking equipment in complex environments, but also provides new ideas for the application of high-robustness target recognition technology in the agricultural field.

[0028] The outdoor apple recognition and positioning method based on multi-sensor fusion of the present disclosure uses a double-edge matching algorithm to perform edge feature matching on depth edge data and ROI edge information maps, identify the true edge contour of the apple, and uses the RANSAC fitting algorithm to generate a smooth edge curve of the apple. By using the double-edge matching method of image edge and depth edge, the true contour information of the target fruit can be obtained more stably under complex outdoor lighting conditions, effectively reducing the influence of overlapping occlusion of leaves, branches and other fruits.

[0029] The outdoor apple recognition and positioning method based on multi-sensor fusion of the present disclosure constructs a sphere model according to the fitting result, solves the center coordinates and radius by minimizing the sum of squared residuals, and obtains the three-dimensional positioning information of the apple. Constructing the sphere model can better establish the three-dimensional pose information of the apple, including the center of the apple object and the three-dimensional contour information of the apple, so as to better perform the operation of grasping the target by the robotic arm subsequently. Brief Description of the Drawings

[0030] The specification drawings forming a part of the present disclosure are used to provide a further understanding of the present disclosure. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure.

[0031] Figure 1 It is a flowchart of the outdoor apple recognition and positioning method based on multi-sensor fusion according to an embodiment of the present disclosure;

[0032] Figure 2 It is a schematic diagram of the specific implementation process of the outdoor apple recognition and positioning method based on multi-sensor fusion according to an embodiment of the present disclosure. Detailed Description of the Invention

[0033] The present disclosure will be further described below in conjunction with the drawings and embodiments.

[0034] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.

[0035] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0036] Embodiment 1

[0037] In one embodiment of the present disclosure, an outdoor apple recognition and localization method based on multi-sensor fusion is provided, including:

[0038] Step 1: Obtain the image data and lidar point cloud data of outdoor apples and preprocess them;

[0039] Step 2: Input the image data into the target detection model to obtain the RGB image of the apple's two-dimensional bounding box; project the lidar point cloud into the camera coordinate system through the camera internal parameters and the calibrated external parameter matrix of the multi-sensor to obtain the ROI depth information area matching the two-dimensional bounding box;

[0040] Step 3: Use the edge detection algorithm to extract the image edge features from the RGB image, and use the inverse distance weighted transformation to assign edge degrees to non-edge pixels, and expand the neighborhood edge information to non-edge pixels to obtain the complete ROI edge information map; perform denoising and discontinuous point detection on the ROI depth information area to obtain the depth edge data under the denoised ROI point cloud;

[0041] Step 4: Use the double-edge matching algorithm to match the edge features of the depth edge data and the ROI edge information map, identify the true edge contour of the apple, and use the RANSAC fitting algorithm to generate the smooth edge curve of the apple;

[0042] Step 5: Construct a sphere model according to the fitting result, solve the center coordinates and radius by minimizing the sum of squared residuals, and obtain the three-dimensional localization information of the apple.

[0043] As an embodiment, the specific implementation process of an outdoor apple recognition and localization method based on multi-sensor fusion of the present disclosure is as follows:

[0044] Step 1: Obtain the image data and lidar point cloud data of outdoor apples and preprocess them;

[0045] Specifically, perform distortion correction through the calibration results (camera internal and external parameters) to eliminate lens distortion. The point cloud data is processed by using statistical filtering (Statistical Outlier Removal, SOR) to remove outliers, reduce the data volume, and improve the processing efficiency.

[0046] Step 2: Input the image data into the target detection model to obtain the RGB image of the apple's two-dimensional bounding box; project the lidar point cloud into the camera coordinate system through the camera internal parameters and the calibrated external parameter matrix of the multi-sensor to obtain the ROI depth information area matching the two-dimensional bounding box;

[0047] Specifically, for the image collected by the camera, use deep learning target detection models such as Yolo to obtain the two-dimensional bounding box (Bbox) of the apple and record its pixel coordinate information, and the recognized Bbox pixel coordinates are 。

[0048] Furthermore, through the intrinsic matrix of the camera and the extrinsic matrix of the calibration of multiple sensors, the lidar point cloud is projected into the camera coordinate system. Let the extrinsic matrix be , and the intrinsic matrix of the camera K can map the lidar point cloud to the camera plane, and then intercept to form a ROI depth information area matching the Bbox.

[0049] Specifically, in order to map the lidar point cloud to the camera plane, the geometric calibration and projection process between multiple sensors must be completed. The intrinsic matrix of the camera K is obtained through a camera calibration tool (such as the OpenCV or MATLAB calibration method), and the distortion of the image is corrected.

[0050]

[0051] Among them, obtaining the camera calibration from the projection matrix means extracting the intrinsic parameters of the camera from the projection matrix of the camera, including the focal length and the calibration parameters of the optical center .

[0052] Then, the extrinsic matrix between the area array lidar and a certain camera is obtained by using the PNP method. The PNP (Perspective-n-Point) method is a technology for calculating the pose of the camera (i.e., the rotation matrix R and the translation vector t) by known three-dimensional points and their two-dimensional projections on the image plane. Specifically, the input data 3D point set (lidar coordinate system): the known three-dimensional point set P = {(X, Y, Z)}, usually obtained through a calibration object (such as a checkerboard), and (X, Y, Z) is the three-dimensional coordinate position of the point cloud in the coordinate system.

[0053] The corresponding 2D pixel points: the projection points p = {(u, v)} of these three-dimensional points on the image plane are usually obtained through image processing. The intrinsic matrix of the camera K has been described above. Then, calculations are performed according to the projection formula.

[0054]

[0055] Among them, R ∣ t is the pose of the camera relative to the lidar coordinate system. S is the scaling factor, which is related to the depth Z . Kis the camera intrinsic matrix. By minimizing the error (such as the reprojection error), the rotation matrix and translation vector of the camera are iteratively solved. (X, Y, Z) are the three-dimensional coordinates of the point cloud in the coordinate system, and the specific positions of X, Y, and Z on the X, Y, and Z axes of the coordinate system respectively. (u, v) are the coordinates of the projection points of the three-dimensional points on the image plane.

[0056] Finally, the extrinsic matrix is obtained .

[0057] Among them, R is the rotation matrix and t is the translation vector.

[0058] Step 3: Use the edge detection algorithm to extract the image edge features of the RGB image, and use the inverse distance weighted transformation to assign the edge degree to the non-edge pixels, expand the neighborhood edge information to the non-edge pixels, and obtain the complete ROI edge information map; perform denoising and discontinuous point detection on the ROI depth information area to obtain the depth edge data under the denoised ROI point cloud;

[0059] Specifically, 1) The process of using the edge detection algorithm to extract the image edge features of the RGB image and using the inverse distance weighted transformation to assign the edge degree to the non-edge pixels and expand the neighborhood edge information to the non-edge pixels to obtain the complete ROI edge information map is as follows:

[0060] Due to different fields of view, the lidar point cloud will have depth interference noise caused by occlusion of leaves, branches or fruits, which cannot be eliminated during conversion and will be brought into the ROI area. In order to exclude the interference of non-apple objects (such as leaves and branches), first grayscale the RGB image, and then use edge detection algorithms such as Canny or Sobel to extract the image edges to obtain the binary edge map T.

[0061] At each edge pixel point ( i , j ), its edge degree is calculated by measuring the maximum absolute value of the surrounding 8-neighborhood pixels M ( i , j ). Specifically in the calculation, first obtain the absolute value of the difference in grayscale values between the target pixel and its 8-neighborhood pixels (the 8-neighborhood is defined as the pixels adjacent to the pixel point ( i , j ) in the horizontal, vertical and diagonal directions), and then take the maximum value of these differences as the edge degree of the pixel . The larger the edge degree, the more significant the grayscale difference between the pixel and its surrounding pixels, and the more likely it is to be located in the edge area of the image.

[0062] Furthermore, the inverse distance weighted transformation is used to assign the edge degree to the non-edge pixels , and the formula is:

[0063]

[0064] Among them, is the edge degree of the pixel, T is the binary edge map, where each pixel value represents the edge degree value at that place, β is the distance weighting coefficient between each non-edge point and points at other positions, the closer the greater, and γ is set to 0.5, representing the weighted sum of the maximum radiation amounts from other positions to this position for each non-edge pixel. is the coordinate of the target pixel, that is, the pixel point whose edge degree is currently being calculated, represents the coordinates of other pixel points, which are other points participating in the calculation in the entire image. represents the coordinate position The edge degree value of the pixel at the place. The radiation amount is related to the edge degree and distance at other positions. This assignment method extends the neighborhood edge information to non-edge pixels, obtains a complete ROI edge information map, and at the same time sets a threshold to remove points that do not conform to the edge, providing a more robust edge description for subsequent feature matching.

[0065] 2) Denoise the ROI depth information area and detect discontinuous points to obtain the depth edge data under the denoised ROI point cloud, including:

[0066] In the point cloud data, in order to remove the interference of leaves and branches, find the points with discontinuous depth between adjacent points in the ROI depth information area. Let the i point in the point cloud be p, Calculate its distance difference from adjacent points and calculate its depth discontinuity value ,

[0067]

[0068] Among them, d represents the depth of this point, that is, the z value after the point cloud obtained by the lidar is transformed into the camera coordinate system (x, y, z). The physical meaning is the distance of this point from the origin of the camera coordinate system in the z-axis direction. γ is a weight parameter used to control the influence of the distance difference on the calculation of depth discontinuity, which is set to 0.5 here.

[0069] Filter all points with depth discontinuity less than the threshold (that is, continuous points) to obtain the depth edge data under the denoised ROI point cloud for further analysis.

[0070] Step 4: Use the double-edge matching algorithm to perform edge feature matching on the depth edge data and the ROI edge information map, identify the true edge contour of the apple, and use the RANSAC fitting algorithm to generate a smooth edge curve of the apple;

[0071] Specifically, the depth edge data under the filtered and denoised ROI point cloud is matched with the ROI edge information map to identify the true edge of the apple. This includes:

[0072] Obtain the set of image edge points: Obtain the set of edge contours of the two-dimensional image of the ROI edge information map ; Set of depth discontinuity points: Set of depth discontinuity points of the lidar point cloud , which contains M discontinuity points, representing potential edges regarded as interference in the depth map.

[0073] Calculate the distance between edges through the Chamfer matching method, and find the edge region with the highest similarity to filter out interference items.

[0074]

[0075] Among them, represents an image edge point , and the Euclidean distance to the nearest point in the set of depth discontinuity points . For each d point, find its nearest point in , sum up all the minimum distances and take the average. Through this distance metric, if the Chamfer distance is small, it indicates that the edge features match well, and the found region can be considered as the set of edge contours of the target apple and .

[0076] Furthermore, after obtaining the filtered edge data, use the RANSAC fitting algorithm to generate a smooth edge curve of the apple. The RANSAC method fits the edge curve by random sampling and minimizing the residuals.

[0077] Let the fitted edge curve be C, then the sum of the distance residuals from the set of edge points and to the fitted curve is defined as

[0078]

[0079] Among them, and represent the projection points of the curve C at the positions of points d and x. Represents the Euclidean distance calculation. By minimizing Error, the best-fitting curve C can be found, thereby generating a smooth apple edge curve. This curve can be used as the basis for subsequent apple positioning and shape estimation.

[0080] Step 5: Construct a sphere model based on the fitting results, solve for the center coordinates and radius by minimizing the sum of squared residuals, and obtain the three-dimensional localization information of the apple.

[0081] Specifically, assuming that the apple is spherical or approximately spherical in structure, using the inner depth points of the fitting curve, select multiple uniformly distributed sample points and calculate the depth value of each point. By fitting the sample depths within the point cloud, the least squares sphere fitting method is used to calculate the center position and three-dimensional information of the apple. Assume that N uniformly distributed point cloud depth points are selected on the inner side of the fitting curve, denoted as (where i = 1, 2, …, N). Then the goal is to find a sphere model such that these points are as close as possible to the surface of the sphere.

[0082] The mathematical model of the sphere can be expressed as

[0083]

[0084] where represents the center point of the sphere.

[0085] Solve for the center coordinates and radius by minimizing the sum of squared residuals r , and the sum of squared residuals is defined as

[0086]

[0087] After expanding the above formula, a non-linear equation system about and r is obtained. To minimize Error, the least squares method or a numerical optimization algorithm (such as the Levenberg-Marquardt algorithm) can be used to solve it, and the solution result is the three-dimensional information of the apple accurately located under complex lighting conditions.

[0088] Through the above multi-sensor fusion and edge feature matching algorithm, the present disclosure can accurately locate the three-dimensional information of the apple under complex lighting conditions, improving the robustness and accuracy of the picking system.

[0089] Embodiment 2

[0090] In an embodiment of the present disclosure, a multi-sensor fusion-based outdoor apple recognition and localization system is provided, including:

[0091] A data acquisition module for acquiring image data and lidar point cloud data of outdoor apples and performing preprocessing;

[0092] A target recognition module, which is used to input image data into a target detection model to obtain an RGB image of the two-dimensional bounding box of the apple; by using the camera internal parameters and the calibrated external parameter matrix of multiple sensors, project the lidar point cloud into the camera coordinate system to obtain an ROI depth information area that matches the two-dimensional bounding box.

[0093] An edge feature extraction module, which is used to extract image edge features from the RGB image using an edge detection algorithm, assign an edge degree to non-edge pixels using inverse distance weighted transformation, expand the neighborhood edge information to non-edge pixels, and obtain a complete ROI edge information map; perform denoising and discontinuous point detection on the ROI depth information area to obtain depth edge data under the denoised ROI point cloud.

[0094] An edge matching module, which is used to perform edge feature matching on the depth edge data and the ROI edge information map using a double-edge matching algorithm, identify the true edge contour of the apple, and use the RANSAC fitting algorithm to generate a smooth edge curve of the apple.

[0095] A positioning calculation module, which is used to construct a sphere model according to the fitting result, solve the center coordinates and radius by minimizing the sum of squared residuals, and obtain the three-dimensional positioning information of the apple.

[0096] Embodiment 3

[0097] In an embodiment of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, it implements the outdoor apple recognition and positioning method based on multi-sensor fusion described above.

[0098] Embodiment 4

[0099] In an embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, and the non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, it implements the outdoor apple recognition and positioning method based on multi-sensor fusion described above.

[0100] Embodiment 5

[0101] In an embodiment of the present disclosure, an electronic device is provided, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes and implements the outdoor apple recognition and positioning method based on multi-sensor fusion described above.

[0102] This disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the specified functions in the process Figure 1 one process or multiple processes and / or blocks Figure 1 means for implementing the specified functions in one block or multiple blocks.

[0103] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in the process Figure 1 one process or multiple processes and / or blocks Figure 1 steps for implementing the specified functions in one block or multiple blocks.

[0104] Although the specific implementation manners of the present disclosure have been described above in conjunction with the accompanying drawings, they are not limitations on the protection scope of the present disclosure. Those skilled in the art should understand that, based on the technical solutions of the present disclosure, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present disclosure.

Claims

1. An outdoor apple recognition and positioning method based on multi-sensor fusion is characterized by: include: Obtain outdoor apple image data and lidar point cloud data, and pre-process them; The image data is input into the object detection model to obtain the RGB image of the apple's two-dimensional bounding box. The lidar point cloud is projected into the camera coordinate system through the camera's intrinsic parameters and the multi-sensor calibration extrinsic parameter matrix to obtain the ROI depth information area that matches the two-dimensional bounding box. Use edge detection algorithm to extract image edge features from RGB images, and use inverse distance weighted transformation to assign edge degree to non-edge pixels, extend neighborhood edge information to non-edge pixels, and obtain a complete ROI edge information map; perform denoising and discontinuous point detection on the ROI depth information area to obtain the depth edge data under the denoised ROI point cloud; The double edge matching algorithm is used to match the edge features of the deep edge data with the ROI edge information map, identify the real edge contour of the apple, and use the RANSAC fitting algorithm to generate the smooth edge curve of the apple; specifically, the deep edge data under the filtered and denoised ROI point cloud is matched with the ROI edge information map to identify the real edge of the apple, including: Get the image edge point set: Get the edge contour set of the two-dimensional image of the ROI edge information map ; Depth discontinuity point set: depth discontinuity point set of lidar point cloud , contains M discontinuities, representing potential edges that are considered as interference in the depth map; The distance between edges is calculated by Chamfer matching method, and the edge area with the highest similarity is found to filter out interference items. in, Represents the edge points of the image , and the set of depth discontinuities The Euclidean distance of the nearest point in d Click and find its The point with the closest distance in the middle is summed up and averaged; through this distance metric, if the Chamfer distance is small, the edge feature match is good, and the found area is considered to be the edge contour set of the target apple. and ; A spherical model is constructed based on the fitting results, and the center coordinates and radius are solved by minimizing the residual sum of squares to obtain the three-dimensional positioning information of the apple.

2. The outdoor apple recognition and positioning method based on multi-sensor fusion as claimed in claim 1, characterized in that: The YOLO target detection model is used to obtain the apple image captured by the camera, obtain the two-dimensional bounding box Bbox of the apple and record its pixel coordinate information; the lidar point cloud is projected to the camera coordinate system through the camera intrinsic parameter matrix and the calibrated extrinsic parameter matrix of the multi-sensor; the camera intrinsic parameter matrix maps the lidar point cloud to the camera plane, and then intercepts it to form a ROI depth information area that matches the two-dimensional bounding box Bbox.

3. The outdoor apple recognition and positioning method based on multi-sensor fusion as claimed in claim 1, characterized in that: The RGB image is grayed out, and the edge detection algorithm is used to extract the image edge to obtain a binary edge map. At each edge pixel, the edge degree is measured by the maximum absolute value of the surrounding 8 neighborhood pixels, and the inverse distance weighted transform is used to assign the edge degree to the non-edge pixels, which is: in, T is a binary edge map, where each pixel value represents the edge value at that location. β The distance weighting coefficient of each non-edge point to other position points is larger the closer the point is, and γ Set to 0.5, representing the weighted sum of the maximum radiation of each non-edge pixel at that position. i , j ) is the coordinate of the target pixel, that is, the pixel whose edge degree is currently being calculated, (x, y) represents the coordinate of other pixels, which are other points in the entire image involved in the calculation, and the subscripts x and y of T represent the position.

4. The outdoor apple recognition and positioning method based on multi-sensor fusion as claimed in claim 1, characterized in that: In the point cloud data, remove the interference of various noises, find the points with discontinuous depth between adjacent points in the ROI depth information area, and set the first i Point p , calculate the distance difference between it and its neighboring points, and calculate its depth discontinuity value , in, d is the distance from the point to the origin of the camera coordinate system in the z-axis direction, γ is the weight parameter; Filter all points whose depth discontinuity is less than the set threshold to obtain the depth edge data under the denoised ROI point cloud.

5. The outdoor apple recognition and positioning method based on multi-sensor fusion as claimed in claim 1, characterized in that: A set of pixel points of the two-dimensional image is extracted from the ROI edge information map, and a set of depth discontinuous points of the lidar point cloud is obtained from the depth edge data. The distance between edges is calculated by the Chamfer matching method. Through distance measurement, if the Chamfer distance is small, it means that the edge feature matches well, and the found area is considered to be the edge contour of the target apple. After obtaining the filtered edge contour data, the RANSAC fitting algorithm is used to generate a smooth edge curve of the apple.

6. The outdoor apple recognition and positioning method based on multi-sensor fusion as claimed in claim 1, characterized in that: Construct a spherical model of an apple with a spherical or spherical structure. Use the inner depth points of the fitting curve to select multiple evenly distributed sample points, calculate the depth value of each point, and calculate the center position of the apple and its three-dimensional information by fitting the sample depth in the point cloud using the least squares spherical fitting method. The mathematical model of the sphere is expressed as: Among them, multiple evenly distributed point cloud depth points are selected inside the fitting curve and expressed as , r represents the radius of the sphere, c The center point of the sphere.

7. The outdoor apple recognition and positioning system based on multi-sensor fusion is characterized by: include: Data acquisition module, used to acquire outdoor apple image data and LiDAR point cloud data, and pre-process them; The target recognition module is used to input the image data into the target detection model to obtain the RGB image of the apple's two-dimensional bounding box; through the camera intrinsic parameters and the calibration extrinsic parameter matrix of the multi-sensor, the lidar point cloud is projected to the camera coordinate system to obtain the ROI depth information area that matches the two-dimensional bounding box; The edge feature extraction module is used to extract the edge features of the RGB image using the edge detection algorithm, and assign the edge degree to the non-edge pixels using the inverse distance weighted transformation, and extend the neighborhood edge information to the non-edge pixels to obtain a complete ROI edge information map; denoise and detect discontinuous points in the ROI depth information area to obtain the depth edge data under the denoised ROI point cloud; The edge matching module is used to match the edge features of the deep edge data with the ROI edge information map using a dual edge matching algorithm, identify the real edge contour of the apple, and use the RANSAC fitting algorithm to generate a smooth edge curve of the apple; specifically, the deep edge data under the filtered and denoised ROI point cloud is matched with the ROI edge information map to identify the real edge of the apple, including: Get the image edge point set: Get the edge contour set of the two-dimensional image of the ROI edge information map ; Depth discontinuity point set: depth discontinuity point set of lidar point cloud , contains M discontinuities, representing potential edges that are considered as interference in the depth map; The distance between edges is calculated by Chamfer matching method, and the edge area with the highest similarity is found to filter out interference items. in, Represents the edge points of the image , and the set of depth discontinuities The Euclidean distance of the nearest point in d Click and find its The point with the closest distance in the middle is summed up and averaged; through this distance metric, if the Chamfer distance is small, the edge feature match is good, and the found area is considered to be the edge contour set of the target apple. and ; The positioning calculation module is used to build a spherical model based on the fitting results, and to solve the center coordinates and radius by minimizing the residual sum of squares to obtain the three-dimensional positioning information of the apple.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the outdoor apple recognition and positioning method based on multi-sensor fusion described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the outdoor apple identification and positioning method based on multi-sensor fusion as described in any one of claims 1-6 is implemented.

10. An electronic device, characterized in that: include: A processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the outdoor apple recognition and positioning method based on multi-sensor fusion as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Shielded vegetable and fruit harvesting method based on depth association perception algorithm

    CN110033487A

  • Method for fusing depth maps by using binocular camera and laser radar

    CN110942477A