Inspection robot real-time position recognition and tomato counting yield evaluation method
By using multi-sensor fusion and deep learning algorithms, the problems of error and time consumption in counting tomatoes by inspection robots in facility agriculture environments have been solved, achieving precise positioning and accurate counting of tomatoes.
Patent Information
- Application Number
- CN202310559502.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-17
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2043-05-17
AI Technical Summary
In existing technologies, inspection robots have problems with large errors and long time consumption when counting tomatoes in real time in facility agriculture environments, especially in complex unstructured growth environments, where it is difficult to accurately identify the location and quantity.
Employing multi-sensor fusion technology, including metal detectors, RGB sensors, RGB-D cameras, and odometry, combined with deep learning algorithms, the system achieves precise location and counting of tomatoes through feature matching and target tracking.
It enables precise location and accurate counting of tomatoes in a facility agriculture environment, reduces repeated counting errors, and improves counting efficiency and accuracy.
Smart Images

Figure CN116681964B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-sensor fusion, specifically to a method for real-time location recognition and tomato counting yield assessment of an inspection robot. Background Technology
[0002] The tomato cultivation in this facility agriculture environment utilizes a Venlo-style greenhouse, characterized by its large space and excellent ventilation. Tomatoes are grown year-round, with pruning, vine thinning, and leaf removal ensuring they ripen and are harvested within their optimal growing areas. Tomatoes grow in planting troughs along a standard bidirectional track. These troughs, stretching hundreds of meters longitudinally and dozens to hundreds laterally, are arranged parallel to the track. The tomato fruits grow at a 120° spatial angle on different plants, resulting in overlapping fruits, branches, leaves, and stems. The complexity of the background spatial distribution, with multiple rows of overlapping information between the fruit and the background, makes it extremely difficult to distinguish the red, fading red, and green hues of ripe fruits during the growing season—a typical unstructured growth environment.
[0003] Overlapping tomato planting presents challenges in yield assessment and counting: deep overlap, unclear layering, multiple images per frame, and multiple interpretations of a single image. An inspection robot, moving along its track, obtains real-time information on the tomatoes in the current row and counts the ripe tomatoes. This provides the grower with accurate maturity data for the current track, helping them manage harvesting, pesticide application, and marketing based on the ripeness count. Traditionally, workers rely on experience to judge the number of fruits and their ripening time, then allocate harvesting personnel and estimate harvesting times. Environmental data monitoring requires specialized instruments for collection, detection, and analysis. Manual analysis is prone to significant errors due to individual subjectivity, and the data collection and analysis process is tedious and unstable. The work is also labor-intensive and highly repetitive.
[0004] Existing technologies have several shortcomings: Inspection robots rely on a single sensor for alignment, leading to inaccurate alignment and long latency. The robot's position coordinates on the track are inaccurate and lack integration with tomato counting, making it difficult to obtain real-time tomato quantity information. Traditional deep learning-based object detection methods directly count crops by multiplying the number of detected fruits in an image by the frame number. This method results in many tomatoes being counted repeatedly in adjacent frames, leading to overestimation errors. Furthermore, traditional methods lack timely, location-based, and quantitative tomato yield estimation for plantation managers. Summary of the Invention
[0005] To address the shortcomings of existing technologies, the purpose of this invention is to provide a method for real-time location recognition and tomato counting yield assessment of an inspection robot.
[0006] A method for real-time position recognition and tomato counting yield assessment of an inspection robot according to the present invention includes the following steps:
[0007] Steps for obtaining row coordinate information: The inspection robot moves between multiple tracks to obtain row coordinate information;
[0008] Tomato inspection and counting steps: The inspection robot acquires tomato images in real time, processes them through deep learning algorithms to obtain tomato targets at different stages of maturity, and tracks and counts the number of tomatoes based on feature matching;
[0009] Location-quantity association steps: Establish the relationship between the current number of mature tomatoes on the ridge and the current track position.
[0010] Preferably, the step of obtaining the row coordinate information includes:
[0011] Track construction steps: Build a track map, label the track IDs, and obtain track coordinate information;
[0012] Track identification steps: The inspection robot automatically identifies the track based on a metal detector and an RGB sensor;
[0013] Track perception switching steps: The inspection robot senses its position based on sensors and switches to different tracks.
[0014] Preferably, the metal detector and RGB sensor are mounted on the inspection robot, wherein:
[0015] The metal detector acquires information about the head of the track, and the inspection robot moves along the direction perpendicular to the track to detect the position of the rail and achieve preliminary positioning.
[0016] The RGB sensor acquires the marking information at the head of the track, determines the track centerline, and achieves precise track positioning.
[0017] Preferably, the tomato inspection and counting steps include:
[0018] Dataset construction steps: Use an RGB-D camera to acquire multi-dimensional spatial information of tomatoes and construct a dataset for detecting tomatoes at different ripeness levels;
[0019] The steps for constructing the feature extraction dataset are as follows: data augmentation and preprocessing are performed on the tomato detection datasets of different maturity levels; the YOLO v7 object detection model is trained using the tomato detection datasets of different maturity levels; and feature extraction datasets of tomatoes of different maturity levels are constructed using YOLO v7 object detection.
[0020] Preprocessing steps: Perform data augmentation and preprocessing on the dataset of tomatoes with different maturity levels for feature extraction;
[0021] Feature extraction training steps: Train the EfficientNet model to extract features using a feature extraction dataset of tomatoes at different maturity levels;
[0022] Clustering mapping steps: Perform surface clustering segmentation on the target row tomato point cloud to obtain the target row tomato point cloud, and map the target row tomato point cloud to a color image to obtain the target row tomato color image;
[0023] Object detection steps: Use the YOLO v7 object detection model to perform object detection, and use the DeepSORT algorithm to perform object tracking, completing the task of counting tomatoes at different maturity levels with multiple targets.
[0024] Preferably, the YOLO v7 object detection model is an improved YOLO v7 object detection model based on the CA attention mechanism.
[0025] Preferably, the clustering mapping step includes:
[0026] The tomato depth point cloud acquired by the vision system is segmented into inter-row segments based on a fixed threshold according to the spacing between planting rows to obtain the tomato point cloud of the target row.
[0027] Tomato point clouds in the target row are obtained by performing surface clustering segmentation on the tomato point cloud of the target ridge. Spatial position constraints are applied to the clustering plane based on the planting row spacing to reduce the search space of the clustering plane.
[0028] The target row of tomato color image is obtained by mapping the tomato point cloud to the color image.
[0029] Preferably, the location quantity association step includes:
[0030] Mileage calculation steps: Use an odometer to record the inspection robot's position information on the track;
[0031] Movement distance estimation steps: Based on the odometry and the inspection robot's speed and position perception, estimate the distance the inspection robot has moved relative to its initial position;
[0032] Inspection robot location acquisition steps: The inspection robot obtains its positioning coordinates in real time and then inputs them into the industrial control computer;
[0033] Steps for obtaining the number of tomatoes: The inspection robot acquires tomato images in real time using an RGB-Deep depth camera, processes them using a deep learning algorithm to obtain tomato targets, and tracks and counts the number of tomatoes based on feature matching;
[0034] Fusion Step: The image information after counting tomatoes in each frame is fused with the current coordinate position of the inspection robot to obtain the number of tomatoes at each coordinate of the inspection robot at each time.
[0035] Preferably, a laser ranging sensor is provided at the end of the track. When the inspection robot moves to a set distance from the end of the track, the laser ranging sensor detects the inspection robot, and the inspection robot returns to its original position.
[0036] Preferably, in the dataset construction step, the multi-dimensional spatial information includes color maps, depth maps, and video streams.
[0037] Preferably, the data augmentation method includes color gamut transformation, translation, rotation, and mosaic methods, and the preprocessing includes normalization and transformation into a fixed-size image.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] 1. This invention uses multiple sensors to achieve precise orbit determination, with accurate position coordinates and short processing time.
[0040] 2. This invention calculates the number of tomatoes using a deep learning algorithm and achieves target tracking through clustering mapping, thereby completing the task of counting tomatoes of different ripeness levels for multiple targets and avoiding duplicate counting of tomatoes.
[0041] 3. This invention integrates the track position and the number of tomatoes to obtain the number of tomatoes at each coordinate of the inspection robot at each time.
[0042] 4. Replace the DeepSORT algorithm's feature extraction network with the EfficientNet model to obtain more accurate features.
[0043] 5. The size of the bounding box for tomato target detection is adaptively adjusted, and some neighboring tomato images are obtained. Further features such as the positional relationship of neighboring tomatoes are extracted, which is beneficial to the accuracy of tomato ID matching between consecutive frames and prevents frequent ID changes.
[0044] 6. Based on the characteristics of the planting ridge spacing, the depth map is segmented by point cloud surface clustering to obtain the tomato point cloud of the target row. Then, the color image is mapped to obtain the tomato image of the target row, eliminating the interference of different row background tomatoes. Attached Figure Description
[0045] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0046] Figure 1 This is a schematic diagram of the orbital coordinate ID information.
[0047] Figure 2 This is a schematic diagram of the automatic identification guide rail structure for inspection robots.
[0048] Figure 3 This is a schematic diagram of an inspection robot automatically moving up and down the track.
[0049] Figure 4 Flowchart for counting tomatoes at different stages of maturity.
[0050] Figure 5A schematic diagram illustrating the acquisition of tomato image information by an RGB-D camera during inspection.
[0051] Figure 6 This is a flowchart of color image mapping based on depth map point cloud segmentation.
[0052] Figure 7 A schematic diagram of the adaptive detection results for the detection box.
[0053] Figure 8 This is a schematic diagram of a tomato counting robot for inspection purposes.
[0054] Figure 9 A schematic diagram showing the fusion of the tomato count and the current position coordinates of the inspection robot. Detailed Implementation
[0055] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.
[0056] like Figures 1 to 9 As shown, this invention provides a method for real-time position recognition and tomato counting yield assessment of an inspection robot. It employs a multi-sensor fusion approach, including a visual RGB camera, an RGB-Deep depth camera, a metal sensor, an odometer, and a laser rangefinder. This allows the inspection robot to automatically acquire track position information and tomato growth information, which are then highly integrated. This enables real-time position recognition and tomato counting yield assessment for the inspection robot. Specifically, the metal sensor detects information at the head of the metal track for initial track positioning. The RGB camera at the front of the vehicle captures the red marker on the head of the track, determining the centerline of the guide rail for precise guide rail positioning. This yields the positioning coordinates of the inspection robot in one direction (denoted as the X-coordinate). Once the inspection robot enters the track, the odometer is activated to measure the distance the robot travels along the track in real time, thus obtaining the positioning coordinates of the robot in the longitudinal direction of the track (denoted as the Y-coordinate). The inspection robot inputs these positioning coordinates into the industrial control computer in real time. As the inspection robot moves along the longitudinal direction of the track, a laser sensor based on a non-contact measurement method is installed at the end of the guide rail. When the sensor detects the inspection robot, the robot automatically stops and begins its return journey. The inspection robot is equipped with 1-3 RGB-Deep depth cameras at the upper part of the tomato fruiting area to acquire tomato point cloud images in real time. These images are then processed using deep learning algorithms to determine the tomato count. Each frame of the image after tomato counting is labeled with the robot's current coordinates, thus obtaining the number of tomatoes at each coordinate point and time. This lays the foundation for subsequent yield assessment.
[0057] More specifically, the process begins by setting up the inspection robot to automatically move up and down the track, thus obtaining the coordinate information of the rows. For example... Figure 1 As shown, first, a guide rail map is constructed, guide rail ID numbers are labeled, and guide rail coordinate information is obtained. The track numbers are then labeled and stored in an orderly manner, as follows: Figure 1 Three parallel tracks are shown. The inspection robot is equipped with a metal detector and RGB sensors, enabling it to automatically identify the tracks. Specifically, firstly, using the metal detector at the front of the robot to detect the metal track head, the robot translates along the X-axis perpendicular to the track, achieving initial positioning. Secondly, precise track positioning is achieved based on the visual RGB sensors. By using the RGB sensors at the front of the robot to obtain the red mark on the track head, the centerline of the track is determined, achieving precise track positioning.
[0058] The inspection robot, equipped with laser sensors, can automatically stop, return, and derail. First, a drive motor propels the robot onto the track, where it moves longitudinally. A laser rangefinder is installed at the end of the track for non-contact remote sensing of the robot's body. When the robot is 30-50cm away from the sensor, the laser rangefinder detects the robot, causing it to automatically stop and begin its return journey.
[0059] The inspection robot uses sensors to detect in real time when it reaches the end of the guide rail. Upon reaching the end position, the robot retreats back to the starting point on the track, derails, and moves to another track. The robot's derailment process is detailed below. Figure 3 Specifically, it includes: laser rangefinder sensor sensing the roving robot, stopping, returning to the starting position 1 – lower rail 2 – translation 3 – guide rail recognition 4 – translation adjustment 5 – upper rail 6 – forward movement 7.
[0060] The inspection robot is equipped with 1-3 RGB-Deep depth cameras at the upper part of the tomato fruiting area to acquire tomato images in real time. These images are then processed using deep learning algorithms to identify the target tomatoes, and feature matching is used to track and count the number of tomatoes. See [link / details]. Figure 4As shown. To address the issues of overlapping rows of tomatoes in yield assessment and counting, problems exist such as depth overlap, unclear layering, multiple images per frame, and multiple interpretations of a single image. This invention proposes a real-time tomato counting method based on RGB-D information mapping. First, multiple rows of background tomatoes are removed using a depth map, and then the effective depth information is mapped onto the color image, thus creating a masking effect for the multiple rows of background tomatoes and eliminating interference from multiple rows of background tomatoes. This method includes registering the point clouds of multiple rows of tomatoes and the single row of tomatoes in the current row, thereby constructing global and local growth information of tomatoes and improving the accuracy of single-tomato counting. The tomato point cloud is segmented by surface clustering to obtain the target row tomato point cloud (i.e., depth map), and the depth map is mapped onto the color image to eliminate interference from background tomatoes (i.e., tomatoes not in the target row). The specific process is as follows: First, the tomato depth point cloud collected by the vision system is segmented into inter-row segments with a fixed threshold based on the planting row spacing to obtain the tomato point cloud of the target row. Then, the tomato point cloud of the target row is segmented into a surface clustering to obtain the tomato point cloud of the target row. In order to improve the surface clustering speed, the clustering plane is constrained in terms of spatial position based on the planting row spacing to reduce the search space of the clustering plane. Finally, the tomato point cloud of the target row is mapped to a color image to obtain the color image of the tomato of the target row.
[0061] To further explain, this method utilizes an RGB-D camera to acquire multi-dimensional spatial information about tomatoes, including color images, depth maps, and video streams, to construct detection datasets for tomatoes at different maturity levels. Data augmentation and preprocessing are then performed on these datasets. An improved YOLO v7 object detection model with a CA attention mechanism is trained using these datasets. A feature extraction dataset for tomatoes at different maturity levels is then constructed using the improved YOLO v7 object detection model. Data augmentation and preprocessing are performed on this feature extraction dataset. An improved DeepSORT algorithm is obtained by replacing the DeepSORT algorithm's feature extraction network with the EfficientNet model, and the EfficientNet model is trained using the feature extraction datasets for tomatoes at different maturity levels to extract features. Object detection is achieved using the improved YOLO v7 model, and object tracking is achieved using the improved DeepSORT algorithm, thus completing the multi-target tomato counting task at different maturity levels.
[0062] More detailed explanation: The tomato detection dataset for different maturity levels includes three categories: immature, semi-ripe, and ripe, and contains corresponding spatial physical location information (multiple rows of tomatoes, single bunch of tomatoes in the current row), with bounding box positions labeled. The verification videos for the tomato counting algorithm at different maturity levels include depth videos and color videos.
[0063] The data augmentation methods include color gamut transformation, translation, rotation, and mosaic methods. Preprocessing includes normalization and transformation into a fixed-size image.
[0064] Improving the YOLO v7 algorithm using the CA attention mechanism is beneficial for enhancing model detection accuracy. The YOLO v7 model consists of three parts: the input part, the backbone network, and the detection head. The backbone network, based on YOLO v5, introduces the ELAN and MP structures. The detection head consists of the SPPC module, the PAN module with the introduced ELAN structure, and the RepConv module.
[0065] To address the requirements for identifying information on different ripeness levels and spatial locations of tomatoes (multiple rows of tomatoes, single bunch of tomatoes in the current row), multi-dimensional camera spatial location relationship cloud registration is performed. The Iterative Closest Point (ICP) algorithm is used to improve the accuracy of point cloud registration from different perspectives.
[0066] The original DeepSORT algorithm's feature extraction network is relatively simple, lacking sufficient feature extraction capabilities. Furthermore, it only extracts features from tomatoes within the bounding boxes obtained from object detection, resulting in tomatoes with similar features that are difficult to distinguish. Replacing the DeepSORT feature extraction network with the EfficientNet model yields more accurate features. It also adaptively expands the bounding box by a factor of two, acquiring images of nearby tomatoes and further extracting features such as the positional relationships between neighboring tomatoes. This improves the accuracy of tomato ID matching between consecutive frames and prevents frequent ID changes.
[0067] Due to background tomato interference, the depth map is first segmented by point cloud surface clustering based on the characteristics of planting row spacing to obtain the tomato point cloud of the target row, eliminating background tomato interference from different rows. The effective depth information is then mapped onto the color map to obtain a color map with effective depth information, which is used for detection by the YOLO v7 model and counting by the DeepSORT model, improving the model detection speed and counting accuracy.
[0068] The input color image has undergone depth mapping to eliminate interference from multiple rows of background tomatoes. The improved YOLO v7 model, as a detector, achieves higher detection accuracy. In the improved DeepSORT algorithm, each new target can be initialized with a new trajectory. The Hungarian algorithm enables the tracking of new targets and the generation of new target IDs. The cosine distance between features extracted by the EfficientNet model is used to filter the targets to be matched. The EfficientNet model, as a feature extractor, provides more accurate ID matching.
[0069] The implementation method of the tomato inspection and counting method based on RGB-D information mapping of the present invention is as follows: Figure 4 As shown, it specifically includes:
[0070] Step 1: A multi-view Realsense camera is used to acquire depth and color video with a resolution of 1080*720 and a depth detection range of 0.2m-10m. The camera device is equipped with two infrared cameras (left and right) and an infrared dot projector. The infrared cameras are used for ranging, and the infrared dot projector is used to improve the accuracy of depth calculation.
[0071] Step 2: Construct a dataset for detecting tomatoes at different maturity levels, namely a training set, a test set, and a validation set.
[0072] Half of the real-time captured color video was used as the detection dataset, saved in JPG format every 5 frames. Labelme software was used for labeling, classifying tomatoes into three categories: unripe, semi-ripe, and ripe. The dataset was then divided into training and testing sets in a 4:1 ratio. Data normalization and data augmentation were performed, including color gamut transformation, translation, rotation, and mosaic techniques. Color gamut transformation converted RGB to HSV color gamut. Figure 1 As shown.
[0073] Step 4: Improve YOLO v7 using the CA attention mechanism, training and parameter tuning using the training set. The CA attention mechanism can improve the detection accuracy of YOLO v7. The backbone feature extraction network, based on YOLO v5, introduces the ELAN and MP structures. The detection head consists of the SPPC module, the PAN module with the introduced ELAN structure, and the RepConv module.
[0074] Step 5: Improve YOLO v7's detection on the training set by adaptively cropping the detected bounding boxes. Adaptation refers to expanding the width and height of the bounding boxes to half of the original size. This allows for the acquisition of the positional relationships of neighboring tomatoes, the extraction of more unique features, better differentiation of tomatoes of the same class, reduced ID changes, and more accurate counting.
[0075] Step 6: Construct a feature extraction dataset for tomatoes at different maturity levels. Normalize and augment the dataset using this dataset. Data augmentation techniques include gamut transformation, translation, rotation, and mosaic methods. Gamut transformation converts RGB to HSV color gamut.
[0076] Step 7: Replace the DeepSORT algorithm feature extraction network with the EfficientNet model to obtain an improved DeepSORT algorithm, and train and optimize the Efficient model using a tomato feature extraction dataset with different maturity levels.
[0077] Step 8: To identify information on different ripeness levels and spatial locations of tomatoes (multiple rows of tomatoes, single bunch of tomatoes in the current row), perform multi-dimensional camera spatial location relationship cloud registration, and use the Iterative Closest Point (ICP) algorithm to improve the accuracy of point cloud registration from different perspectives.
[0078] Step 9: Perform surface clustering segmentation on the depth map point cloud based on the characteristics of the planting ridge spacing to obtain the tomato point cloud of the target row, and map the effective depth information onto the color map.
[0079] Step 10: Use the improved YOLO v7 to detect the color image and obtain the detection boxes and categories, such as... Figure 4 As shown. The yellow dashed lines represent ripe tomatoes, and the red solid lines represent unripe tomatoes.
[0080] Step 11: Input the image within the adaptive detection box into the Efficient model for feature extraction. Obtain the feature vector.
[0081] Step 12: Perform concatenated matching on the feature vectors. If successful, input a list of successfully matched targets and trajectories; otherwise, perform IoU matching. If IoU matching is successful, input a list of successfully matched targets and trajectories; if the number of failures reaches a certain threshold, abandon the process. Then, perform Kalman filtering prediction on the list of successfully matched targets and trajectories.
[0082] Step 13: Acquire another frame and repeat steps 8-12 to achieve the target count.
[0083] like Figure 9 As shown, the relationship between the current number of ripe tomatoes on the ridge and the track position is established. An odometer-equipped inspection robot is mounted on the track wheels. Data from motion sensors obtained by the wheeled odometer is used to estimate the robot's position change over time. Based on the robot's speed and position perception using the odometer, the distance the robot has moved relative to its initial position is estimated. Using this method, the robot's positioning coordinates in one direction (denoted as the X coordinate) and its positioning coordinates in the track's depth direction (denoted as the Y coordinate) are obtained. The robot receives these positioning coordinates in real time and inputs them into the industrial control computer. The inspection robot is equipped with 1-3 RGB-Deep depth cameras at the upper part of the tomato-bearing area to acquire tomato images in real time. These images are processed using a deep learning algorithm to obtain the tomato targets, and then target tracking and tomato counting are achieved based on feature matching. The image information after each tomato count is fused with the robot's current coordinate position. This yields the number of tomatoes at each coordinate of the inspection robot at each time point.
[0084] In the description of this application, it should be understood that the terms "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.
[0085] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.
Claims
1. A method for real-time position recognition and tomato counting yield assessment of an inspection robot, characterized in that, Includes the following steps: Steps for obtaining row coordinate information: The inspection robot moves between multiple tracks to obtain row coordinate information; Tomato inspection and counting steps: The inspection robot acquires tomato images in real time, processes them through deep learning algorithms to obtain tomato targets at different stages of maturity, and tracks and counts the number of tomatoes based on feature matching; Location-quantity association steps: Establish the relationship between the current number of ripe tomatoes on the ridge and the current track position; The steps for counting tomatoes during inspection include: Dataset construction steps: Use an RGB-D camera to acquire multi-dimensional spatial information of tomatoes and construct a dataset for detecting tomatoes at different ripeness levels; The steps for constructing the feature extraction dataset are as follows: data augmentation and preprocessing are performed on the tomato detection datasets of different maturity levels; the YOLO v7 object detection model is trained using the tomato detection datasets of different maturity levels; and feature extraction datasets of tomatoes of different maturity levels are constructed using YOLO v7 object detection. Preprocessing steps: Perform data augmentation and preprocessing on the dataset of tomatoes with different maturity levels for feature extraction; Feature extraction training steps: Train the EfficientNet model to extract features using a feature extraction dataset of tomatoes at different maturity levels; Clustering mapping steps: Perform surface clustering segmentation on the target row tomato point cloud to obtain the target row tomato point cloud, and map the target row tomato point cloud to a color image to obtain the target row tomato color image; Target detection steps: Use the YOLO v7 target detection model to implement target detection, and use the DeepSORT algorithm to implement target tracking, to complete the task of counting tomatoes with different maturity levels and multiple targets; The clustering mapping step includes: The tomato depth point cloud acquired by the vision system is segmented into inter-row segments based on a fixed threshold according to the spacing between planting rows to obtain the tomato point cloud of the target row. Tomato point clouds in the target row are obtained by performing surface clustering segmentation on the tomato point cloud of the target ridge. Spatial position constraints are applied to the clustering plane based on the planting row spacing to reduce the search space of the clustering plane. The target row of tomato color image is obtained by mapping the tomato point cloud to the color image.
2. The method for real-time position recognition and tomato counting yield assessment of the inspection robot according to claim 1, characterized in that, The steps for obtaining the row coordinate information include: Track construction steps: Build a track map, label the track IDs, and obtain track coordinate information; Track identification steps: The inspection robot automatically identifies the track based on a metal detector and an RGB sensor; Track perception switching steps: The inspection robot senses its position based on sensors and switches to different tracks.
3. The method for real-time position recognition and tomato counting yield assessment of the inspection robot according to claim 2, characterized in that, The metal detector and RGB sensor are mounted on the inspection robot, wherein: The metal detector acquires information about the head of the track, and the inspection robot moves along the direction perpendicular to the track to detect the position of the rail and achieve preliminary positioning. The RGB sensor acquires the marking information at the head of the track, determines the track centerline, and achieves precise track positioning.
4. The method for real-time position recognition and tomato counting yield assessment of the inspection robot according to claim 1, characterized in that, The YOLO v7 object detection model is an improved version of the CA attention mechanism.
5. The method for real-time position recognition and tomato counting yield assessment of the inspection robot according to claim 1, characterized in that, The location quantity association step includes: Mileage calculation steps: Use an odometer to record the inspection robot's position information on the track; Movement distance estimation steps: Based on the odometry and the inspection robot's speed and position perception, estimate the distance the inspection robot has moved relative to its initial position; Inspection robot location acquisition steps: The inspection robot obtains its positioning coordinates in real time and then inputs them into the industrial control computer; Steps for obtaining the number of tomatoes: The inspection robot acquires tomato images in real time using an RGB-Deep depth camera, processes them using a deep learning algorithm to obtain tomato targets, and tracks and counts the number of tomatoes based on feature matching; Fusion Step: The image information after counting tomatoes in each frame is fused with the current coordinate position of the inspection robot to obtain the number of tomatoes at each coordinate of the inspection robot at each time.
6. The method for real-time position recognition and tomato counting yield assessment of the inspection robot according to claim 1, characterized in that, A laser rangefinder is installed at the end of the track. When the inspection robot moves to a set distance from the end of the track, the laser rangefinder detects the inspection robot, and the inspection robot returns to its original position.
7. The method for real-time position recognition and tomato counting yield assessment of the inspection robot according to claim 1, characterized in that, In the dataset construction step, the multi-dimensional spatial information includes color maps, depth maps, and video streams.
8. The method for real-time position recognition and tomato counting yield assessment of the inspection robot according to claim 1, characterized in that, The data augmentation methods include color gamut transformation, translation, rotation, and mosaic methods. Preprocessing includes normalization and transformation into a fixed-size image.
Citation Information
Patent Citations
Method for extracting interesting buildings from three-dimensional laser point cloud data
CN101726255A
Robot automatic switch system used for agriculture greenhouse
CN106171651A
3D target detection algorithm based on camera and laser radar data fusion
CN113985445A
KR20200080450A