Soil block removing method for potato harvester based on E-NET
Through E-NET algorithm and machine vision identification of soil blocks, combined with the strike device, efficient and automated removal of soil blocks during potato harvesting is achieved, solving the problems of high labor intensity and high cost in the existing technology, and improving harvest efficiency and automation.
Patent Information
- Application Number
- CN202510811903.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-18
AI Technical Summary
The prior art is difficult to efficiently and automatically remove soil blocks during potato harvesting, resulting in high labor intensity and low efficiency, and the existing equipment or system cost, complex structure or difficulty in popularizing it.
Deep learning and machine vision methods based on E-NET algorithm are used to identify soil blocks in real time and determine target hit points. Large soil blocks are broken into small blocks through the strike device, and removed by the gap between the conveying chain rollers, which is suitable for continuous operations.
It realizes the precise removal of soil blocks during potato harvesting, improves the degree of automation, reduces hardware costs and structural complexity, and is suitable for continuous operations.
Smart Images

Figure CN120339295A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of potato harvesting, and in particular, to a method for removing soil clods for a potato harvester based on E-NET. Background Art
[0002] Potatoes are the most important food crops in the world after wheat, rice, and corn, and China is currently the country with the largest total potato output in the world. However, large-scale planting has made potato harvesting a difficult problem. At present, it is still a difficult problem to remove soil clods in the potato flow during the potato harvesting process, and this problem still needs to be solved manually. Manual operation not only has low work efficiency but also high labor intensity.
[0003] The impact separation device designed by Feller R et al. uses the difference in the rebound trajectories of steel drums to achieve separation. Although the separation success rate is relatively high, it is necessary to strictly control the falling height, resulting in limited processing speed, and the mechanical structure is complicated, increasing the maintenance cost.
[0004] The manipulator picking system based on visual recognition by Liu Wendong et al. identifies soil clods through external features. Although it provides ideas for automatic cleaning, it relies on high-precision sensors and computer processing, with high hardware costs, and the dynamic response speed is difficult to meet the requirements of continuous operation.
[0005] The machine vision cleaning control system developed by Wang Zedong of Heilongjiang Bayi Agricultural University has a high system integration level and a long debugging cycle, making it difficult to popularize in small and medium-sized agricultural machinery. Summary of the Invention
[0006] In order to solve the above technical problems, the present application provides a method for removing soil clods for a potato harvester based on E-NET. The method includes the following steps: S1, collecting images in real time during the impurity removal process; S2, preprocessing the images to reduce image noise; S3, using the E-NET algorithm to perform instance segmentation on the preprocessed images for soil clods; S4, outputting the masked layer image after instance segmentation; wherein, instances and the background are marked with different colors; S5, determining the target hitting points according to the masked layer image; S6, determining the coordinates of the target hitting points; S7, performing a hitting action according to the coordinates of the target hitting points.
[0007] In some embodiments of the present application, in step S4, determining the target hitting points according to the masked layer image includes the following steps: S41, determining the hittable trajectories according to the movement paths of the hitting execution components and the conveyor belt; S42. Determine the intersection points of the hittable trajectory and the soil block mask based on the hittable trajectory and the masked layer image after instance segmentation. S43. Determine the target hitting points based on the intersection points of the hittable trajectory and the soil block mask.
[0008] In some embodiments of the present application, in step S43, configure the center point of the two intersection points as the target hitting point.
[0009] In some embodiments of the present application, in step S6, execute the hitting action using a hitting device; the hitting device includes a conveyor belt, a hitting execution component, and a camera; potatoes and soil blocks to be identified are placed on the conveyor belt; the camera and the hitting execution component are respectively arranged above the conveyor belt along the conveying direction of the potatoes.
[0010] In some embodiments of the present application, the hitting execution component includes two sets of cylinder assemblies arranged side by side, and each set of cylinder assemblies includes a plurality of cylinders arranged side by side.
[0011] In some embodiments of the present application, in step S6, assume that the coordinate system for the camera to detect the object position is , and the coordinate system for the hitting execution component to execute the action is , then the relationship between the coordinate system for the camera to detect the object position and the coordinate system for the hitting execution component to execute the action is as shown in the following formula:
[0012] Where: X represents a 3×3 homogeneous transformation matrix, and its specific structure is as follows;
[0013] and are the rotation components in the X-axis direction; is the translation amount in the X-axis direction; and are the rotation components in the Y-axis direction; is the translation amount in the Y-axis direction; The calculation process of X is as follows: Configure a standard target with 9 target points in the hitting area, and configure the coordinates of the target in the coordinate system as ( , ), and configure the coordinates of the target in the coordinate system as ( , ) Then the coordinate relationship of each target point is as follows: .
[0014] In some embodiments of the present application, in step S2, preprocessing the image includes cropping and filtering; The cropping is used to adjust the size and shape of the image; The filtering uses bilateral filtering.
[0015] In some embodiments of the present application, the E-NET algorithm includes a linearly deformable convolutional layer.
[0016] In some embodiments of the present application, the E-NET algorithm sets three detection heads, namely the P3 / 8 detection head, the P4 / 16 detection head, and the P5 / 32 detection head; The size of the detection feature map corresponding to the P3 / 8 detection head is 80*80, and it is used to detect targets larger than 8*8; The size of the detection feature map corresponding to the P4 / 16 detection head is 40*40, and it is used to detect targets larger than 16*16; The size of the detection feature map corresponding to the P5 / 32 detection head is 20*20, and it is used to detect targets larger than 32*32.
[0017] In some embodiments of the present application, the detection head includes a bounding box prediction branch and a classification prediction branch; The bounding box prediction branch includes two convolutional layers, a two-dimensional convolution, and a bounding box loss module; The classification prediction branch includes two depthwise separable convolutions, two convolutional layers, a two-dimensional convolution, and a classification loss module.
[0018] Compared with the prior art, the present invention has the following advantages and beneficial effects: The method for removing soil clods from a potato harvester based on E-NET of the present application proposes an E-NET model with a simple structure and good real-time performance based on the E-NET algorithm. This model has good real-time performance, small computational complexity, and small memory occupancy, and can meet the requirements of the potato and impurity separation scenario; It uses deep learning and machine vision algorithms to realize the recognition and positioning of soil clods, and then determines the target hitting points of the soil clods, improving the accuracy of target hitting, breaking large soil clods into small soil clods, and the small soil clods fall to the ground through the gaps between the conveying chain rollers, realizing the removal of soil clods. This method can accurately remove soil clods from potatoes and is suitable for continuous operation.
[0019] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings that form a part of this specification are used to provide a further understanding of this specification. The schematic embodiments and descriptions thereof in this specification are used to explain this specification and do not constitute an improper limitation of this specification. In the drawings: Figure 1 It is a principle block diagram of a soil clod removal method for a potato harvester based on E-NET provided by an exemplary embodiment of the present application; Figure 2 It is a perspective view of a striking device provided by an exemplary embodiment of the present application; Figure 3 It is a side view of a potato harvester provided by an exemplary embodiment of the present application; Figure 4 It is a comparison diagram before and after filtering processing provided by an exemplary embodiment of the present application; Figure 5 It is a process diagram of determining a target striking point provided by an exemplary embodiment of the present application; Figure 6 It is an E-NET algorithm model diagram provided by an exemplary embodiment of the present application; Figure 7 It is a structure model diagram of a linearly deformable convolution layer in the E-NET algorithm provided by an exemplary embodiment of the present application; Figure 8 It is a structure model diagram of a convolution layer in the E-NET algorithm provided by an exemplary embodiment of the present application; Figure 9 It is a structure model diagram of a residual block in the E-NET algorithm provided by an exemplary embodiment of the present application; Figure 10 It is a structure model diagram of a detection head in the E-NET algorithm provided by an exemplary embodiment of the present application; Figure 11 It is an experimental effect diagram of the Yolo v11n algorithm provided by an exemplary embodiment of the present application; Figure 12 It is an experimental effect diagram of the E-NET algorithm provided by an exemplary embodiment of the present application.
[0021] In the figure: 1. Camera; 2. Light source; 3. Conveyor belt; 4. Striking execution component. Detailed implementation manners
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other arbitrarily.
[0023] Potatoes are the world's most important food crop after wheat, rice, and corn, and China is currently the country with the largest total potato output in the world. However, large-scale planting has made potato harvesting a difficult problem. At present, the removal of soil clods from the potato stream during the harvesting process remains a difficult problem, which still needs to be solved manually. Manual operation not only has low work efficiency but also high labor intensity.
[0024] The impact separation device designed by Feller R et al. uses the difference in the rebound trajectories of steel drums to achieve separation. Although the separation success rate is relatively high, it is necessary to strictly control the falling height, resulting in limited processing speed, and the complication of the mechanical structure increases the maintenance cost.
[0025] The mechanical hand picking system based on visual recognition by Liu Wendong et al. identifies soil clods through external features. Although it provides ideas for automatic cleaning, it relies on high-precision sensors and computer processing, with high hardware costs, and the dynamic response speed is difficult to meet the requirements of continuous operation.
[0026] The machine vision cleaning control system developed by Wang Zedong of Heilongjiang Bayi Agricultural University has a high degree of system integration and a long debugging cycle, making it difficult to popularize in small and medium-sized agricultural machinery.
[0027] Based on this, this application exemplarily provides a method for removing soil clods for a potato harvester based on E-NET. Based on the E-NET algorithm, an E-NET model with a simple structure and good real-time performance is proposed. This model has good real-time performance, small computational complexity, and small memory occupancy, and can meet the requirements of the potato and impurity separation scenario; deep learning and machine vision algorithms are used to realize the recognition and positioning of soil clods, and then determine the target hitting points of the soil clods, improving the accuracy of target hitting, breaking large soil clods into small soil clods, and the small soil clods fall to the ground through the gaps between the conveying chain rollers, realizing the removal of soil clods. This method can accurately remove soil clods from potatoes and is suitable for continuous operation.
[0028] An exemplary embodiment of this application provides a method for removing soil clods for a potato harvester based on E-NET, as Figure 1 shown. This method includes the following steps: S1, collect images in real time during the impurity removal process; S2, preprocess the images to reduce image noise; S3, use the E-NET algorithm to perform instance segmentation on the preprocessed images for soil clods; S4, output the masked layer image after instance segmentation; where the instances and the background are marked with different colors; S5, determine the target hitting points according to the masked layer image; S6, determine the coordinates of the target hitting points; S7, perform a hitting action according to the coordinates of the target hitting points.
[0029] This application uses a striking device to implement and execute the striking action; as Figure 2 and 3 shown, after the harvested potatoes are preliminarily screened to remove small soil clods, the remaining potatoes and larger soil clods are transported to the striking device together. The striking device includes a conveyor belt 3, a striking execution component 4, a light source 2, a bracket, and a camera 1; the bracket is erected above the conveyor belt 3, and both the camera 1 and the light source 2 are arranged on the bracket. Preferably, the camera is a Hikvision industrial camera; the light source 2 is preferably a strip light source. The two strip light sources are symmetrically installed, which can not only provide sufficient brightness for the system but also prevent the problem of poor quality of the collected pictures caused by light source reflection; particularly, a light-shielding cover is equipped outside the bracket to avoid interference from external light sources. The conveyor belt 3 is placed with potatoes and soil clods to be identified; the striking execution component 4 is arranged above the conveyor belt 3, and the cylinder assembly is arranged downstream of the bracket along the conveying direction of the potatoes.
[0030] Preferably, the striking execution component 4 includes two groups of cylinder assemblies arranged side by side, each group including a plurality of cylinders arranged side by side; the cylinders of the two groups of cylinder assemblies are alternately arranged. Exemplarily, the first group of cylinder assemblies includes 7 cylinders, the second group of cylinder assemblies includes 6 cylinders, and the cylinders of the first group of cylinder assemblies and the second group of cylinder assemblies are staggered in the direction perpendicular to the conveying direction of the conveyor belt, so that the distance between the striking points is 3.5 cm, while the distance between adjacent conveying chain rollers of the potato harvester is 5 cm. In this way, the broken small soil clods can fall through the gaps of the next-level conveying chain to remove the soil clods.
[0031] In step S1, image acquisition is performed using a camera.
[0032] In step S2, preprocessing of the image includes cropping and filtering; among them, cropping is used to adjust the size and shape of the image; preferably, the size of the cropped picture is 640*640. In this step, bilateral filtering is used for filtering, which is a non-linear filtering method that combines the spatial proximity and pixel value similarity of the image, and takes into account both the spatial domain information and the gray similarity to achieve the purpose of edge-preserving denoising. This method is simple, non-iterative, and localized, and is particularly suitable for processing the edge information of images The formula for bilateral filtering can be expressed as:
[0033] Among them: is the pixel value after filtering; is the pixel value of the original image; is the normalized weight, which is used to normalize the filtering result; represents the neighborhood window of the filter; is a spatial domain kernel function used to calculate the spatial distance weight between pixels; is a pixel value range kernel function used to calculate the similarity weight between pixel values.
[0034] Since excessive filtering will lose the original texture, color and other features of the target, in actual operation, we generally only perform mild filtering. In this method, we set the value of d to 3; the value of sigmaColor to 20; the value of sigmaSpace to 50. If the value of sigmaSpace is large, it means that pixels that are far away but have similar colors will affect each other, so that sufficiently similar colors in a larger area will obtain the same color. The comparison of the images before and after filtering is as Figure 4 shown.
[0035] S3. Use the E-NET algorithm to perform instance segmentation on the preprocessed image for soil clods, as Figure 5 shown. In the segmented mask image, white represents soil clods and black represents the background. The E-NET algorithm is improved based on the Yolo v11n algorithm to form the E-NET algorithm. Its network model structure is as Figures 6 to 10 shown. The E-NET algorithm includes modules such as input, convolutional layer, linear deformable convolutional layer, residual block, and detection head. Among them, the convolutional layer includes two-dimensional convolution, batch normalization, and activation function; the residual block includes a convolutional layer, channel splitting, a residual bottleneck module, and feature fusion; the residual bottleneck module includes 2 convolutional layers.
[0036] The E-NET model is proposed based on yolo v11n. The E-NET proposes a simple and practical structure for the requirements of the soil clod hitting scenario. Since the shapes of soil clods are diverse, standard convolution samples in a local window with a fixed shape and size, which is difficult to dynamically adapt to the shapes of different objects. Although deformable convolution allows flexible sampling positions, the number of its parameters increases quadratically with the size of the convolution kernel, and the computational efficiency is low. Therefore, in this application, a linear deformable convolutional layer is introduced to adapt to soil clods with different shapes and sizes. The linear deformable convolutional layer includes two-dimensional convolution, pointwise convolution, and linear deformable convolution.
[0037] The linear deformable convolution proposes an algorithm to generate the initial sampling coordinates of a convolution kernel of any size and adjusts the sampling shape through an offset amount to enable it to adapt to the changes in the target shape. At the same time, the linear deformable convolution allows the convolution kernel to have any number of parameters, such as 1, 2, 3, 4, 5, 6, 7, etc., thus providing more flexibility for network design; at the same time, the linear deformable convolution changes the growth trend of the number of parameters from quadratic to linear, thereby reducing the requirements for the hardware environment.
[0038] For the input X, the implementation process of linearly deformable convolution is as follows: First, generate the initial sampling coordinates: According to the convolution kernel size, use an algorithm to generate the initial sampling coordinates, which can be of any shape, such as triangles, rectangles, rhombuses, etc. Second, calculate the offset: Learn the offset through convolution operations and add it to the initial sampling coordinates to generate new sampling coordinates, thereby adjusting the sampling shape. Third, extract features: Through interpolation and resampling of the feature map, obtain the features corresponding to the new sampling coordinates, and use the corresponding convolution operations to extract features.
[0039] In the E-NET algorithm, the input image first enters the convolutional layer to extract preliminary features. Subsequently, the extracted features pass through a combined module of linearly deformable convolutional layers and residual blocks in sequence. The linearly deformable convolutional layer further extracts flexible and diverse features, while the residual block helps the network better train and learn deep features. After being processed by 3 groups of combined modules of linearly deformable convolutional layers and residual blocks, they are input into the corresponding detection heads at the 4th layer, 6th layer, and 8th layer respectively. The detection heads make predictions related to object detection based on these features and output information such as the category and location of the objects.
[0040] The E-NET algorithm sets three detection heads, namely the P3 / 8 detection head, the P4 / 16 detection head, and the P5 / 32 detection head; the size of the detection feature map corresponding to the P3 / 8 detection head is 80*80, which is used to detect objects larger than 8*8; the size of the detection feature map corresponding to the P4 / 16 detection head is 40*40, which is used to detect objects larger than 16*16; the size of the detection feature map corresponding to the P5 / 32 detection head is 20*20, which is used to detect objects larger than 32*32.
[0041] As Figure 10 shown, the detection head includes a bounding box prediction branch and a classification prediction branch; the bounding box prediction branch includes two convolutional layers, a two-dimensional convolution, and a bounding box loss module; the features output by the detection head first pass through a convolutional layer to extract preliminary features, and then are further processed by subsequent convolutional layers and two-dimensional convolution to continuously refine the features, and finally are used to calculate the bounding box loss to optimize the target position prediction. The classification prediction branch includes two depthwise separable convolutions, two convolutional layers, a two-dimensional convolution, and a classification loss module; the detection head features enter the depthwise separable convolution, utilize its efficient feature extraction method, and then combine with convolutional layers to further process the features. After multiple operations, the results are output through a two-dimensional convolution to calculate the classification loss and optimize the target category prediction.
[0042] On the same device, with a python3.9 environment and a GTX 1650 (4G) graphics card, the same batch of test sets (200 images) were tested, and the comparison experimental results are as Figure 11 、 Figure 12 and Table 1 show.
[0043] Table 1 Comparison Test Data Table of Yolo v11n Algorithm and E-NET Algorithm
[0044] From Figure 11 and Figure 12 it can be seen that the E-NET model is slightly inferior to Yolo v11n in terms of segmentation accuracy, but the E-NET model is generally superior to Yolo v11n in terms of confidence, and the actual effect of the E-NET model can already meet the requirements of the soil block separation scenario.
[0045] As can be seen from Table 1, the E-NET model has an increase of 0.001 in terms of accuracy compared to Yolo v11n, is on par with Yolo v11n in terms of recall rate, and has a decrease of 0.01 in terms of mean average precision. Since the E-NET redesigned the network structure and introduced linearly variable linear convolution, the E-NET model has significant advantages in terms of model size and real-time performance. The E-NET has a reduction in model size from 6.0MB to 3.6MB compared to Yolo v11n, and the frame rate has also increased from 7 frames to 13 frames.
[0046] In step S4, determine the target hitting point according to the mask layer image, continue to refer to Figure 5 , including the following steps: S41, determine the hittable trajectory according to the hitting execution component and the movement path of the conveyor belt; Exemplarily, in this application, 13 cylinders are set, and the cylinders are arranged in a direction perpendicular to the conveying direction of the conveyor belt. Each cylinder performs a piston movement in the vertical direction to hit the soil block. When the conveyor belt moves, each cylinder will form a hittable trajectory within the hitting area, for a total of 13 hittable trajectories.
[0047] S42, determine the intersection points of the hittable trajectory and the soil block mask according to the hittable trajectory and the mask layer image after instance segmentation; Along the extension direction of the hittable trajectory, each hittable trajectory and the soil block mask have two intersection points.
[0048] S43, determine the target hitting point according to the intersection points of the hittable trajectory and the soil block mask.
[0049] Preferably, in order to accurately hit the soil block and break the soil block into relatively uniform small soil blocks, configure the center point of the two intersection points of each hittable trajectory and the soil block mask as the target hitting point.
[0050] S6, determine the coordinates of the target hitting point; In this application, the coordinates of the hitting point are obtained by converting the pixel coordinates to actual coordinates. Specifically, the pixel coordinates * the transformation matrix gives the coordinates of the hitting point. Assume that the pixel coordinate system is the coordinate system for the camera to detect the position of the object The coordinate system of the striking point is , then the relationship between the coordinate system of the camera detecting the object position and the coordinate system of the striking execution component performing the action is shown in the following formula:
[0051] Where: X represents a 3×3 homogeneous transformation matrix, and its specific structure is as follows;
[0052] and is the rotation component in the X-axis direction; is the translation in the X-axis direction; and is the rotation component in the Y-axis direction; is the translation in the Y-axis direction; The calculation process of X is as follows: Configure 9 standard targets at the striking area and place the targets in the coordinate system. The coordinate configuration in is ( , ), place the target in the coordinate system The coordinate configuration in is ( , ) The coordinate relationship of each target point is as follows:
[0053] By calculating the coordinate relationship of the 9 target points and solving the equations simultaneously, X can be solved and the solved X can be substituted into the formula The coordinates of the hitting point can be obtained.
[0054] S7, execute the striking action according to the coordinates of the target striking point; illustratively, after determining the coordinates of the target striking point, the control system sends the coordinates of the target striking point to the lower computer, and the lower computer receives the striking instruction. When the soil block moves with the conveyor belt to the bottom of the striking execution component, the striking execution component executes the striking action to break the large soil block into small soil blocks. Exemplarily, in the coordinates of the striking point, C represents the cylinder number, and its value range is 01-13, corresponding to the 13 cylinders of the striking execution component, and each cylinder is equipped with a solenoid valve; E represents the number of pulses of the photoelectric encoder, and its value range is 0-9999. For example, (02, 1000) means that when the microcontroller collects 1000 pulses of the photoelectric encoder, it outputs a high level to the solenoid valve of cylinder No. 2, and the solenoid valve controls the opening and closing of the cylinder air inlet and outlet to make the cylinder produce a striking action.
[0055] In this application, the terms "comprises", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that an article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such article or device. In the absence of more restrictions, the elements defined by the sentence "comprising..." do not exclude the presence of other identical elements in the article or device comprising the elements.
[0056] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0057] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the intention of the present application also includes these modifications and variations.
Claims
1. A method for removing soil clods for a potato harvester based on E-NET, characterized in that, The method comprises the following steps: S1, acquiring images in real time during the impurity removal process; S2, preprocessing the images to reduce image noise; S3, using the E-NET algorithm to perform instance segmentation on the preprocessed images for soil blocks; S4, outputting the masked layer image after instance segmentation; wherein, instances and the background are marked with different colors; S5, determining the target hitting points according to the masked layer image; S6, determining the coordinates of the target hitting points; S7, performing a hitting action according to the coordinates of the target hitting points.
2. The method for removing soil clods for a potato harvester based on E-NET according to claim 1, wherein In step S4, determining the target hitting points according to the masked layer image comprises the following steps: S41, determining the hittable trajectory according to the hitting execution component and the movement path of the conveyor belt; S42, determining the intersection points of the hittable trajectory and the soil block mask according to the hittable trajectory and the masked layer image after instance segmentation; S43, determining the target hitting points according to the intersection points of the hittable trajectory and the soil block mask.
3. The method for removing soil clods for a potato harvester based on E-NET according to claim 2, wherein In the said step S43, the center point of the two intersection points is configured as the target hitting point.
4. The method for removing soil clods for a potato harvester based on E-NET according to claim 1, wherein In the said step S6, a hitting device is used to perform the hitting action; the hitting device comprises a conveyor belt, a hitting execution component and a camera; potatoes and soil blocks to be identified are placed on the conveyor belt; the camera and the hitting execution component are respectively arranged above the conveyor belt along the conveying direction of the potatoes.
5. The method for removing soil clods for a potato harvester based on E-NET according to claim 4, characterized in that, The hitting execution component comprises two groups of cylinder assemblies arranged side by side, and each group of cylinder assemblies comprises a plurality of cylinders arranged side by side.
6. The method for removing soil clods for a potato harvester based on E-NET according to claim 1, wherein In step S6, assuming that the coordinate system for the camera to detect the object position is , and the coordinate system for the hitting execution component to perform the action is , then the relationship between the coordinate system for the camera to detect the object position and the coordinate system for the hitting execution component to perform the action is shown by the following formula: Wherein: X represents a 3×3 homogeneous transformation matrix, and its specific structure is as follows; and is the rotational component in the X-axis direction; is the translation amount in the X-axis direction; and is the rotational component in the Y-axis direction; is the translation amount in the Y-axis direction; The calculation process of X is as follows: Configure a standard target with 9 target points in the hitting area, and configure the coordinates of the target in the coordinate system as ( , ), and configure the coordinates of the target in the coordinate system as ( , ) Then the coordinate relationship of each target point is as follows: 。 7. The method for removing soil clods for a potato harvester based on E-NET according to claim 1, characterized in that, In the said step S2, preprocessing the images includes cropping and filtering; The cropping is used to adjust the size and shape of the images; The filtering uses bilateral filtering.
8. The method for removing soil clods for a potato harvester based on E-NET according to claim 1, characterized in that, The E-NET algorithm includes a linearly deformable convolutional layer.
9. The method for removing soil clods for a potato harvester based on E-NET according to claim 1, wherein, The E-NET algorithm sets three detection heads, namely the P3 / 8 detection head, the P4 / 16 detection head and the P5 / 32 detection head; The detection feature map corresponding to the P3 / 8 detection head has a size of 80*80 and is used to detect targets with a size of 8*8 or more; The detection feature map corresponding to the P4 / 16 detection head has a size of 40*40 and is used to detect targets with a size of 16*16 or more; The detection feature map corresponding to the P5 / 32 detection head has a size of 20*20 and is used to detect targets with a size of 32*32 or more.
10. The method for removing soil clods for a potato harvester based on E-NET according to claim 9, characterized in that, The detection head includes a bounding box prediction branch and a classification prediction branch; The bounding box prediction branch includes two convolutional layers, a two-dimensional convolution and a bounding box loss module; The classification prediction branch includes two depthwise separable convolutions, two convolutional layers, a two-dimensional convolution and a classification loss module.
Citation Information
Patent Citations
Control method of manipulator grabbing control system based on binocular stereoscopic vision
CN106041937A
Potato harvester loss reduction and separation device with soil crushing device
CN109348825A
Method and device for disorderly grabbing workpieces based on industrial robot and intelligent terminal
CN114347008A
Grabbing method and device and robot
CN115922703A
Stacked object identification method, device and equipment and computer storage medium
CN116171463A