A fruit picking robot vision bionics method, device, equipment and medium
The fruit-harvesting robot, equipped with a cascaded vision system, utilizes a combination of large and small field-of-view sensors to achieve dynamic tracking and precise image acquisition of target fruits, thereby improving the harvesting accuracy and efficiency of the robot.
Patent Information
- Application Number
- CN202510623713.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2045-05-15
AI Technical Summary
The existing visual perception function of harvesting robots has poor ability to perceive target fruits, resulting in low harvesting accuracy and poor results.
A cascaded vision system is adopted, which consists of a large field-of-view sensor and a small field-of-view sensor. By using a visual tracking network model and a multi-target segmentation network, the system can dynamically track the target fruit and acquire accurate image information, thereby controlling the harvesting mechanism to harvest precisely.
It improves the accuracy and efficiency of fruit harvesting, and solves the problem of low harvesting accuracy in existing technologies.
Smart Images

Figure CN120244989B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of fruit sensing and harvesting technology, and in particular, to a visual bionic method, device, equipment and medium for a fruit harvesting robot. Background Technology
[0002] With the continuous development and application of robot performance, robots are being used in various fields to replace traditional manual labor. In the field of fruit harvesting, harvesting robots are gradually being applied. However, the visual perception capabilities of existing harvesting robots still suffer from poor ability to perceive target fruits, leading to low harvesting accuracy and poor harvesting results. Therefore, improving the fruit perception and harvesting capabilities of harvesting robots has become a major challenge in this field. Summary of the Invention
[0003] This application provides a visual bionic method, device, equipment, and medium for fruit picking robots to solve one or more technical problems existing in the prior art, and at least provide a beneficial option or create conditions.
[0004] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.
[0005] According to one aspect of the embodiments of this application, a visual biomimetic method for a fruit-picking robot is provided, applied to a fruit-picking robot, the fruit-picking robot including a large field-of-view sensor, a small field-of-view sensor, and a picking mechanism, the method comprising:
[0006] Construct a visual tracking network model related to the large field-of-view sensor;
[0007] The first original image information of the target fruit uploaded by the large field of view sensor is input into the visual tracking network model to obtain the first image information, so as to control the small field of view sensor to move to the pre-picking point according to the first image information.
[0008] A multi-target segmentation network associated with the small field-of-view sensor is constructed. After the small field-of-view sensor moves to the pre-harvesting point, the second original image information of the target fruit uploaded by the small field-of-view sensor is input into the multi-target segmentation network to obtain the second image information. Based on the second image information, the position of the target fruit is determined and the harvesting mechanism is controlled to move to the position to harvest the target fruit.
[0009] In one embodiment of this application, based on the foregoing scheme, the step of inputting the first original image information of the target fruit uploaded by the large field-of-view sensor into the visual tracking network model to obtain the first image information includes:
[0010] The first original image information is input into a preset image adaptive control network to obtain the processed initial image information;
[0011] The target fruit in the initial image information is dynamically tracked based on the visual tracking network model to obtain the first image information.
[0012] In one embodiment of this application, based on the foregoing scheme, the image adaptive modulation network includes a first processing module for inverse mapping adjustment and a second processing module for learning image signal parameters; the step of inputting the first original image information into a preset image adaptive modulation network to obtain processed initial image information includes:
[0013] The first raw image captured by the large field-of-view sensor on the target fruit is input into the first processing module and the second processing module respectively to obtain the processed initial image information.
[0014] In one embodiment of this application, based on the foregoing scheme, the dynamic tracking of the target fruit in the initial image information based on the visual tracking network model includes:
[0015] The target fruit region in the initial image information is determined based on the fruit recognition classifier in the visual tracking network model.
[0016] The MASNet single-frame image recognition result corresponding to the target fruit region is determined based on the MASNet backbone network in the visual tracking network model.
[0017] The depth distance information between the target fruit and the large field-of-view sensor is obtained based on the visual tracking network model.
[0018] The target fruit is dynamically tracked based on the MASNet single-frame image recognition results and the depth and distance information.
[0019] In one embodiment of this application, based on the foregoing scheme, the dynamic tracking of the target fruit based on the MASNet single-frame image and the depth distance information includes:
[0020] Based on the depth distance information and the third coordinate system of the large field of view sensor, pseudo point cloud information is constructed, and a bird's-eye view coordinate system is established based on the pseudo point cloud information, so as to perceive the target fruit according to the bird's-eye view coordinate system.
[0021] Based on the MASNet single-frame image recognition results, the center pixel position of the target fruit region is determined, and the center pixel position is mapped to the bird's-eye view coordinate system to obtain the three-dimensional spatial coordinates of the target fruit in the bird's-eye view coordinate system, so as to dynamically track the target fruit according to the three-dimensional spatial coordinates.
[0022] In one embodiment of this application, based on the foregoing scheme, the step of inputting the second original image information of the target fruit uploaded by the small field-of-view sensor into the multi-target segmentation network to obtain the second image information includes:
[0023] The second original image information is segmented based on the target recognition network in the multi-target segmentation network to obtain a salient image;
[0024] The saliency image is input into the layer aggregation network in the multi-target segmentation network to obtain the target image;
[0025] The target image is subjected to multi-level extraction of fruit, branches and leaves to obtain a segmentation mask of the target fruit, and the segmentation mask is used as the second image information.
[0026] In one embodiment of this application, based on the foregoing scheme, after obtaining the segmentation mask of the target fruit, the method further includes:
[0027] Obtain the distance information between the target fruit and the small field-of-view sensor;
[0028] A three-dimensional coordinate system related to the small field-of-view sensor is constructed based on the distance information and the segmentation mask;
[0029] Obtain the fruit area and average depth of the fruit in the segmentation mask in the three-dimensional coordinate system, and obtain the branch and leaf area and average depth of the branches and leaves in the segmentation mask in the three-dimensional coordinate system;
[0030] The shading rate of the fruit due to shading by the branches and leaves is calculated based on the fruit area, the average depth of the fruit, the area of the branches and leaves, and the average depth of the branches and leaves.
[0031] According to one aspect of the embodiments of this application, a visual bionic device for a fruit-picking robot is provided, applied to a fruit-picking robot, the fruit-picking robot including a large field-of-view sensor, a small field-of-view sensor, and a picking mechanism, the device comprising:
[0032] The first network construction unit is used to construct a visual tracking network model related to the large field-of-view sensor;
[0033] An image processing unit is used to input the first original image information of the target fruit uploaded by the large field of view sensor into the visual tracking network model to obtain the first image information, so as to control the small field of view sensor to move to the pre-picking point according to the first image information.
[0034] The second network construction unit is used to construct a multi-target segmentation network related to the small field-of-view sensor. After the small field-of-view sensor moves to the pre-picking point, the second original image information of the target fruit uploaded by the small field-of-view sensor is input into the multi-target segmentation network to obtain the second image information. Based on the second image information, the position of the target fruit is determined and the picking mechanism is controlled to move to the position to pick the target fruit.
[0035] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided that stores a computer program thereon, the computer program including executable instructions that, when executed by a processor, implement the method described in the above embodiments.
[0036] According to one aspect of the embodiments of this application, an electronic device is provided, including: one or more processors; and a memory for storing executable instructions of the processors, which, when executed by the one or more processors, cause the one or more processors to perform the method as described in the above embodiments.
[0037] The beneficial effects of this application are as follows: The fruit-picking robot proposed in this application includes a large field-of-view sensor, a small field-of-view sensor, and a picking mechanism. The large field-of-view sensor and the small field-of-view sensor can form a cascaded vision system. That is, the large field-of-view sensor can obtain a first original image information of a large range for the target fruit, and input the first original image information into the visual tracking network model related to the large field-of-view sensor to obtain the first image information. This allows the large field-of-view sensor to dynamically track the target fruit. Then, based on the first image information obtained from the dynamic tracking, the small field-of-view sensor is controlled to move to the pre-picking point for the next operation.
[0038] After the small field-of-view sensor moves to the pre-harvesting point, the second original image information of the target fruit uploaded by the small field-of-view sensor is input into the multi-target segmentation network. This enables the small field-of-view sensor to capture more accurate images of the target fruit at the pre-harvesting point, obtaining second image information with higher pixel accuracy. Then, the position of the target fruit can be determined based on the second image information, and the harvesting mechanism can be controlled to move to the position to harvest the target fruit.
[0039] The visual biomimetic method for fruit-picking robots provided in this application enables cascaded visual perception. By using a large field-of-view sensor to search for and perceive fruits over a wider range, and after perceiving the target fruit, a small field-of-view sensor can be moved to a pre-picking point closer to the target fruit to obtain higher-precision second image information. This information is then used to control the picking mechanism to move to the location of the target fruit, thereby achieving precise perception and picking of the target fruit and improving the accuracy of fruit picking.
[0040] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0041] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:
[0042] Figure 1 This is a flowchart illustrating a visual biomimetic method for a fruit-picking robot according to an embodiment of this application;
[0043] Figure 2 This is a logical schematic diagram of a small-field-of-view sensor-related multi-target segmentation network according to an embodiment of this application;
[0044] Figure 3 This is a block diagram illustrating a vision bionic device for a fruit-picking robot according to an embodiment of this application;
[0045] Figure 4 This is a structural diagram of an electronic device according to this application. Detailed Implementation
[0046] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.
[0047] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.
[0048] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller node devices.
[0049] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0050] It should be noted that "multiple" in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0051] The implementation details of the technical solutions in the embodiments of this application are described in detail below:
[0052] According to one aspect of this application, a visual biomimetic method for a fruit-picking robot is provided. Figure 1 The flowchart below illustrates a visual biomimetic method for a fruit-harvesting robot according to an embodiment of this application. This method includes at least steps S1 to S3, which are described in detail below:
[0053] In step S1, a visual tracking network model related to the large field of view sensor is constructed.
[0054] In step S2, the first original image information of the target fruit uploaded by the large field of view sensor is input into the visual tracking network model to obtain the first image information, so as to control the small field of view sensor to move to the pre-picking point according to the first image information;
[0055] In step S3, a multi-target segmentation network related to the small field-of-view sensor is constructed. After the small field-of-view sensor moves to the pre-picking point, the second original image information of the target fruit uploaded by the small field-of-view sensor is input into the multi-target segmentation network to obtain the second image information. Based on the second image information, the position of the target fruit is determined and the picking mechanism is controlled to move to the position to pick the target fruit.
[0056] First, the principle of the cascaded vision system implemented in this application is as follows:
[0057] Define {B} as the base coordinate system of the fruit-picking robot body, {E} as the coordinate system of the picking mechanism, {S1} as the coordinate system of the large field-of-view sensor, {P} as the coordinate system of the pre-picking point, {S2} as the coordinate system of the small field-of-view sensor, and {H} as the coordinate system of the target fruit. Once the large field-of-view sensor detects the target fruit, and the distance between the pre-picking point and the target fruit is relatively short (typically 50 cm), the coordinates of the pre-picking point can still be obtained.
[0058] The first transformation relationship between the coordinate system of the harvesting mechanism and the base coordinate system of the robot body. The small field-of-view sensor can be moved to the pre-picking point, as expressed by the following formula:
[0059]
[0060] in This represents the transformation matrix from the large field-of-view sensor to the fruit-picking robot itself. The transformation matrix is the coordinate system from the pre-harvest point to the large field-of-view sensor. This is the inverse of the transformation matrix of the small-field-of-view sensor relative to the first coordinate system of the harvesting mechanism. After the small-field-of-view sensor detects the location of the target fruit at the pre-harvest point, it uses the second transformation relationship between the coordinate system of the harvesting mechanism and the base coordinate system of the fruit-harvesting robot. This allows the harvesting mechanism to move to the location of the target fruit to perform the harvesting task, as expressed by the following formula:
[0061]
[0062] in, This is the transformation matrix from the small field-of-view sensor to the coordinate system of the harvesting mechanism. This is the transformation matrix from the target fruit's location point to the coordinate system of the small-field-of-view sensor. This application constructs a cascaded control system of a large-field-of-view sensor, a small-field-of-view sensor, and a harvesting mechanism using the above method to achieve the perception of the target fruit.
[0063] In one embodiment of this application, the step of inputting the first raw image information of the target fruit uploaded by the large field-of-view sensor into the visual tracking network model to obtain the first image information includes:
[0064] The first original image information is input into a preset image adaptive control network to obtain the processed initial image information;
[0065] The target fruit in the initial image information is dynamically tracked based on the visual tracking network model to obtain the first image information.
[0066] The image adaptive modulation network includes a first processing module for inverse mapping adjustment and a second processing module for learning image signal parameters; the step of inputting the first original image information into the preset image adaptive modulation network to obtain processed initial image information includes:
[0067] The first raw image captured by the large field-of-view sensor on the target fruit is input into the first processing module and the second processing module respectively to obtain the processed initial image information.
[0068] Specifically, a large-field bionic visual coarse perception and planning model (composed of an image adaptive control network and a visual tracking network model) is constructed. In the actual perception of the target fruit, it will be affected by a variety of environmental factors, such as lighting and ambient wind. Lighting will affect the captured image, while ambient wind will cause the target fruit to shake. Therefore, the large-field bionic visual coarse perception and planning model is constructed to get rid of the interference caused by these environmental factors.
[0069] First, an image adaptive control network related to the large field-of-view sensor is constructed, suitable for eliminating interference from lighting factors. For the complex lighting conditions in the orchard environment, the constructed image adaptive control network includes two independent processing flows: a branch L (i.e., the first processing module) for inverse mapping adjustment and a branch G (the second processing module) for learning ISP (Image Signal Processor) parameters.
[0070] Branch L is proposed to consist of three pixel enhancement modules. These modules use 1×1 convolutions and efficient depthwise convolutions to aggregate contextual pixels across channels, ensuring that the global contextual relationships between pixels are implicitly modeled with low computational cost. Branch G is proposed to use an Attention module to obtain global information to generate the color matrix and gamma values. The network's loss function is considered to be a weighted sum of the Smooth L1 (first loss function) and Charbonnier (second loss function) loss functions to synthesize multi-scale feature information from different levels, initially defined as:
[0071]
[0072] Among them, R s,n T represents the output image of feature layer s and scale n in an image adaptive modulation network. n Let λ represent the expected image at scale n (Ground Truth), where λ is the weighting factor. Let be the total loss function of the image adaptive modulation network. For the first loss function, This is the second loss function.
[0073] In one embodiment of this application, the dynamic tracking of the target fruit in the initial image information based on the visual tracking network model includes:
[0074] The target fruit region in the initial image information is determined based on the fruit recognition classifier in the visual tracking network model.
[0075] The MASNet single-frame image recognition result corresponding to the target fruit region is determined based on the MASNet backbone network in the visual tracking network model.
[0076] The depth distance information between the target fruit and the large field-of-view sensor is obtained based on the visual tracking network model.
[0077] The target fruit is dynamically tracked based on the MASNet single-frame image recognition results and the depth and distance information.
[0078] The dynamic tracking of the target fruit based on the MASNet single-frame image and the depth distance information includes:
[0079] Based on the depth distance information and the coordinate system of the large field of view sensor, pseudo point cloud information is constructed, and a bird's-eye view coordinate system is established based on the pseudo point cloud information, so as to perceive the target fruit according to the bird's-eye view coordinate system.
[0080] Based on the MASNet single-frame image recognition results, the center pixel position of the target fruit region is determined, and the center pixel position is mapped to the bird's-eye view coordinate system to obtain the three-dimensional spatial coordinates of the target fruit in the bird's-eye view coordinate system, so as to dynamically track the target fruit according to the three-dimensional spatial coordinates.
[0081] Specifically, addressing dynamic factors such as wind in the orchard environment, the project first uses a fruit recognition classifier to accurately perceive and adaptively control target fruits in images. Then, based on this, a visual tracking network model using a Kalman filter is employed to achieve human-eye-like biomimetic perception and motion tracking of the target fruits. The constructed fruit recognition classifier includes a MASNet backbone network, referred to here as MASNet (Multi-level Auxiliary Information Supervision Citr-YOLO, MASNet). Specifically, an efficient layer aggregation network is inserted between the MASNet backbone network and the feature pyramid layers to combine multi-level auxiliary gradient information returned from different prediction heads. This ensures that the features of the main branch pyramid layers are no longer affected by background object information such as branches and leaves, mitigating information loss and corruption issues in deep supervision.
[0082] In natural environments, the position of fruits in images changes non-linearly over time due to disturbances caused by dynamic factors such as wind. Therefore, based on the single-frame image recognition results of MASNet, depth and distance information is integrated into the tracking process to mimic the target-following mechanism of the human eye, enabling motion tracking and trajectory prediction of disturbed fruit targets.
[0083] Specifically, firstly, using the large field-of-view sensor API, pseudo-point cloud information is constructed in the corresponding image coordinate system (i.e., the coordinate system of the large field-of-view sensor) by combining depth and distance information; then, a BEV (Bird's Eye View) coordinate system is established, and the center pixel position of the target fruit region is calculated based on the MASNet single-frame image recognition results, and mapped to the BEV coordinate system to represent its three-dimensional spatial position, thus obtaining the three-dimensional spatial coordinates of the target fruit; finally, the relative position (i.e., depth and distance information) of the target fruit to the large field-of-view sensor is calculated and matched, thereby realizing the tracking and trajectory prediction of the target fruit.
[0084] During trajectory prediction and update, the Kalman filter algorithm is used to simultaneously predict and update spatial location information. For detection boxes and trajectories that fail to match, their Euclidean distance in the BEV coordinate system is calculated, and the following cost matrix is constructed: (The purpose of constructing this cost matrix is for rematching and fine-tuning).
[0085]
[0086] The purpose of constructing the cost matrix is to perform rematching and fine-tuning, making the trajectory prediction of the target fruit more accurate.
[0087] In the cost matrix and formula above, a and b represent the detection box and trajectory of the failed match, respectively. x and a yLet represent the horizontal and vertical coordinates of the detection box relative to the origin of the large field-of-view sensor in the BEV coordinate system, respectively, and b x and b y This represents a matching trajectory. The Euclidean distance between two points is denoted as Dis. ab The cost matrix is denoted as CM. To achieve matching with minimum cost, the Lapjv algorithm is used, which filters out matches that do not meet the conditions by setting a threshold t.
[0088] In one embodiment of this application, the step of inputting the second original image information of the target fruit uploaded by the small field-of-view sensor into the multi-target segmentation network to obtain the second image information includes:
[0089] The second original image information is segmented based on the target recognition network in the multi-target segmentation network to obtain a salient image;
[0090] The saliency image is input into the layer aggregation network in the multi-target segmentation network to obtain the target image;
[0091] The target image is subjected to multi-level extraction of fruit, branches and leaves to obtain a segmentation mask of the target fruit, and the segmentation mask is used as the second image information.
[0092] After obtaining the segmentation mask of the target fruit, the method further includes:
[0093] Obtain the distance information between the target fruit and the small field-of-view sensor;
[0094] A three-dimensional coordinate system related to the small field-of-view sensor is constructed based on the distance information and the segmentation mask;
[0095] Obtain the fruit area and average depth of the fruit in the segmentation mask in the three-dimensional coordinate system, and obtain the branch and leaf area and average depth of the branches and leaves in the segmentation mask in the three-dimensional coordinate system;
[0096] The shading rate of the fruit due to shading by the branches and leaves is calculated based on the fruit area, the average depth of the fruit, the area of the branches and leaves, and the average depth of the branches and leaves.
[0097] Specifically, an active visual fine perception and picking decision-making model is constructed that links the target fruit with a small field-of-view sensor. This active visual fine perception and picking decision-making model is the multi-target segmentation network described in this application, and its principle is as follows: Figure 2As shown, taking the target fruit as an example, the image of the citrus fruit in the figure is the second original image information. After passing through the target recognition network, the layer aggregation network and the recognition model, the bounding box, class label and mask can be obtained through the recognition model, and finally the rightmost fruit branch and leaf segmentation mask is obtained.
[0098] In the constructed multi-object segmentation network, in order to avoid the fruit-picking robot vision system becoming bloated, a Masks branch is introduced on the basis of the MASNet backbone network to generate pixel-level masks (i.e. segmentation masks) of the target fruit, thereby achieving instance segmentation of multi-class targets.
[0099] The loss function is adjusted to measure the difference between the predicted and actual masks. The mask loss of the network is considered to be the improved cross-entropy loss, VFL Loss. Here, the mask loss function is defined as:
[0100]
[0101] Where q is the intersection-over-union ratio of the model's predicted values and the true values, p is the model's class prediction probability, α and γ are weighting factors, and VFL Loss consists of the loss functions corresponding to the bounding box and the mask.
[0102] For fruit-harvesting robots, target fruit occlusion is an existing problem that is unavoidable in reality. To address this problem, a multi-information fusion-based active fruit recognition method is proposed for calculating fruit size, occlusion rate, and growth posture from multiple perspectives. First, after obtaining the segmentation mask, image processing morphology methods are used to eliminate mask noise. Then, combined with a depth image (an image representing depth information, obtained using a small field of view), which is the distance information described in this application, the fruit, branch, and leaf regions are extracted separately. This is mapped to obtain the actual area A (i.e., any one or more of the fruit, branch, and leaf regions) and average depth D (i.e., any one of the average depth values of the fruit and branches / leaves) of each target (i.e., any one or more of the fruit, branch, and leaf regions) in three-dimensional space (the three-dimensional coordinate system related to the small field of view sensor described in this application). Next, background fruits with an area less than 1500 pixels are removed, and the mask of the foreground branches and leaves is extracted. Finally, the depth of the branch / leaf region and the depth of the fruit region are combined to determine whether the fruit is occluded, and the fruit occlusion rate is calculated based on the above information. The formula for calculating the fruit occlusion rate is defined as:
[0103]
[0104] Where D and C are the average depth values of the fruit and branches / leaves, respectively; E, F, and B represent the actual areas of the fruit, shaded area, and branches / leaves, respectively; X represents whether the fruit is shaded by branches / leaves; H is an empirical value used to account for errors caused by fruit clusters; and Y represents the degree of shading, i.e., the shading rate of the fruit.
[0105] In actual harvesting, to improve the success rate of robot harvesting, it is necessary to know the harvesting point and growth posture information of the target fruit. Since the harvesting mechanism has a certain tolerance range, the estimation of the growth posture of the target fruit does not need to be very precise, which means that an approximation can be provided to determine the rotation angle of the end effector (harvesting mechanism). Based on this, according to biological knowledge and previous experimental experience, the fruit stem tilt angle range [45, 135] is divided into three gradients. At this time, the fruit posture estimation problem can be regarded as a classification problem. In order to predict the growth tilt angle of the fruit, a fruit posture estimation method based on convolutional network is proposed. First, the fruit region image is extracted based on multi-target segmentation mask; then, the fruit region image is scaled to 50x50, and the aspect ratio w / h of the original image is input into the model; finally, the trained model is used to predict the fruit growth posture angle, realize the active recognition of the fruit target, and provide phenotypic information such as fruit occlusion rate, position and posture for the robot to calculate the optimal harvesting view and optimal harvesting point.
[0106] In summary, the solution provided in this application can achieve the perception of the target fruit and realize the biomimetic mechanism of the target fruit. By perceiving the target fruit through a cascaded vision system, it is possible to track and predict the trajectory of the target fruit during its movement. Based on this, a small field-of-view sensor is precisely moved to the pre-harvesting point corresponding to the target fruit. It should be noted that there is a linear relationship between the position of the target fruit and the pre-harvesting point; that is, the distance and height between the position of the target fruit and the pre-harvesting point are constant values. Due to the movement of the target fruit, the coordinates of the pre-harvesting point will also change accordingly.
[0107] Furthermore, after the small field-of-view sensor reaches the pre-harvest point, it can accurately perceive the target fruit, thereby enabling the harvesting mechanism to accurately harvest the target fruit. This solves the problems of low harvesting accuracy and poor harvesting effect in the existing technology and improves the harvesting accuracy of the target fruit.
[0108] According to one aspect of the embodiments of this application, a visual bionic device 300 for a fruit-picking robot is proposed, such as... Figure 3 As shown, Figure 3 A block diagram of a vision-inspired bionic device 300 for a fruit-harvesting robot, the device 300 comprising:
[0109] The first network construction unit 301 is used to construct a visual tracking network model related to the large field of view sensor;
[0110] Image processing unit 302 is used to input the first original image information of the target fruit uploaded by the large field of view sensor into the visual tracking network model to obtain the first image information, so as to control the small field of view sensor to move to the pre-picking point according to the first image information.
[0111] The second network construction unit 303 is used to construct a multi-target segmentation network related to the small field-of-view sensor. After the small field-of-view sensor moves to the pre-picking point, the second original image information of the target fruit uploaded by the small field-of-view sensor is input into the multi-target segmentation network to obtain the second image information. Based on the second image information, the position of the target fruit is determined and the picking mechanism is controlled to move to the position to pick the target fruit.
[0112] In another aspect, this application also provides a computer-readable storage medium storing a program product capable of implementing the methods provided above in this specification. In some possible implementations, various aspects of this application may also be implemented as a program product comprising program code that, when run on a terminal device, causes the terminal device to perform the steps described in the "Embodiment Methods" section of this specification according to various exemplary embodiments of this application.
[0113] According to the embodiments of this application, the program product used to implement the above-described method may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of this application is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0114] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0115] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0116] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0117] Program code for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0118] The following reference Figure 4 To describe an electronic device 400 according to this embodiment of the present application. Figure 4 The electronic device 400 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0119] like Figure 4 As shown, the electronic device 400 is manifested in the form of a general-purpose computing device. The components of the electronic device 400 may include, but are not limited to: at least one processing unit 410, at least one storage unit 420, and a bus 430 connecting different system components (including storage unit 420 and processing unit 410).
[0120] The storage unit stores program code that can be executed by the processing unit 410, causing the processing unit 410 to perform the steps described in the "Embodiment Methods" section above according to various exemplary embodiments of this application.
[0121] Storage unit 420 may include a readable medium in the form of a volatile storage unit, such as random access memory (RAM) 421 and / or cache memory 422, and may further include a read-only memory (ROM) 423.
[0122] Storage unit 420 may also include a program / utility 424 having a set (at least one) of program modules 425, such program modules 425 including but not limited to: an operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.
[0123] Bus 430 can represent one or more of several types of bus structures, including a memory cell bus or memory cell control node, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0124] Electronic device 400 can also communicate with one or more external devices 1200 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 400, and / or with any device that enables electronic device 400 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 450. Furthermore, electronic device 400 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 460. As shown, network adapter 460 communicates with other modules of electronic device 400 via bus 430. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 400, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0125] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this application.
[0126] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this application, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0127] It should be understood that this application is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A visual biomimetic method for a fruit-picking robot, characterized in that, An application to a fruit-harvesting robot, the fruit-harvesting robot comprising a large field-of-view sensor, a small field-of-view sensor, and a harvesting mechanism, the method comprising: Construct a visual tracking network model related to the large field-of-view sensor; The first original image information of the target fruit uploaded by the large field of view sensor is input into the visual tracking network model to obtain the first image information, so as to control the small field of view sensor to move to the pre-picking point according to the first image information. A multi-target segmentation network related to the small field-of-view sensor is constructed. After the small field-of-view sensor moves to the pre-picking point, the second original image information of the target fruit uploaded by the small field-of-view sensor is input into the multi-target segmentation network to obtain the second image information. The position of the target fruit is determined according to the second image information, and the picking mechanism is controlled to move to the position to pick the target fruit. The step of inputting the first raw image information of the target fruit uploaded by the large field-of-view sensor into the visual tracking network model to obtain the first image information includes: The first original image information is input into a preset image adaptive control network to obtain the processed initial image information; Based on the visual tracking network model, the target fruit in the initial image information is dynamically tracked to obtain the first image information; The image adaptive modulation network includes a first processing module for inverse mapping adjustment and a second processing module for learning image signal parameters; the step of inputting the first original image information into the preset image adaptive modulation network to obtain processed initial image information includes: The first raw image captured by the large field-of-view sensor on the target fruit is input into the first processing module and the second processing module respectively to obtain the processed initial image information; The dynamic tracking of the target fruit in the initial image information based on the visual tracking network model includes: The target fruit region in the initial image information is determined based on the fruit recognition classifier in the visual tracking network model. The MASNet single-frame image recognition result corresponding to the target fruit region is determined based on the MASNet backbone network in the visual tracking network model. The depth distance information between the target fruit and the large field-of-view sensor is obtained based on the visual tracking network model. The target fruit is dynamically tracked based on the MASNet single-frame image recognition results and the depth and distance information. The dynamic tracking of the target fruit based on the MASNet single-frame image and the depth distance information includes: Based on the depth distance information and the third coordinate system of the large field of view sensor, pseudo point cloud information is constructed, and a bird's-eye view coordinate system is established based on the pseudo point cloud information, so as to perceive the target fruit according to the bird's-eye view coordinate system. Based on the MASNet single-frame image recognition results, the center pixel position of the target fruit region is determined, and the center pixel position is mapped to the bird's-eye view coordinate system to obtain the three-dimensional spatial coordinates of the target fruit in the bird's-eye view coordinate system, so as to dynamically track the target fruit according to the three-dimensional spatial coordinates.
2. The visual bionic method for fruit picking robots according to claim 1, characterized in that, The step of inputting the second original image information of the target fruit uploaded by the small field-of-view sensor into the multi-target segmentation network to obtain the second image information includes: The second original image information is segmented based on the target recognition network in the multi-target segmentation network to obtain a salient image; The saliency image is input into the layer aggregation network in the multi-target segmentation network to obtain the target image; The target image is subjected to multi-level extraction of fruit, branches and leaves to obtain a segmentation mask of the target fruit, and the segmentation mask is used as the second image information.
3. The visual bionic method for fruit picking robots according to claim 2, characterized in that, After obtaining the segmentation mask of the target fruit, the method further includes: Obtain the distance information between the target fruit and the small field-of-view sensor; A three-dimensional coordinate system related to the small field-of-view sensor is constructed based on the distance information and the segmentation mask; Obtain the fruit area and average depth of the fruit in the segmentation mask in the three-dimensional coordinate system, and obtain the branch and leaf area and average depth of the branches and leaves in the segmentation mask in the three-dimensional coordinate system; The shading rate of the fruit due to shading by the branches and leaves is calculated based on the fruit area, the average depth of the fruit, the area of the branches and leaves, and the average depth of the branches and leaves.
4. A visual bionic device for a fruit-picking robot, characterized in that, An application to a fruit-harvesting robot, the fruit-harvesting robot comprising a large field-of-view sensor, a small field-of-view sensor, and a harvesting mechanism, the device comprising: The first network construction unit is used to construct a visual tracking network model related to the large field-of-view sensor; An image processing unit is used to input the first original image information of the target fruit uploaded by the large field of view sensor into the visual tracking network model to obtain the first image information, so as to control the small field of view sensor to move to the pre-picking point according to the first image information. The second network construction unit is used to construct a multi-target segmentation network related to the small field of view sensor. After the small field of view sensor moves to the pre-picking point, the second original image information of the target fruit uploaded by the small field of view sensor is input into the multi-target segmentation network to obtain the second image information. Based on the second image information, the position of the target fruit is determined and the picking mechanism is controlled to move to the position to pick the target fruit. The step of inputting the first raw image information of the target fruit uploaded by the large field-of-view sensor into the visual tracking network model to obtain the first image information includes: The first original image information is input into a preset image adaptive control network to obtain the processed initial image information; Based on the visual tracking network model, the target fruit in the initial image information is dynamically tracked to obtain the first image information; The image adaptive modulation network includes a first processing module for inverse mapping adjustment and a second processing module for learning image signal parameters; the step of inputting the first original image information into the preset image adaptive modulation network to obtain processed initial image information includes: The first raw image captured by the large field-of-view sensor on the target fruit is input into the first processing module and the second processing module respectively to obtain the processed initial image information; The dynamic tracking of the target fruit in the initial image information based on the visual tracking network model includes: The target fruit region in the initial image information is determined based on the fruit recognition classifier in the visual tracking network model. The MASNet single-frame image recognition result corresponding to the target fruit region is determined based on the MASNet backbone network in the visual tracking network model. The depth distance information between the target fruit and the large field-of-view sensor is obtained based on the visual tracking network model. The target fruit is dynamically tracked based on the MASNet single-frame image recognition results and the depth and distance information. The dynamic tracking of the target fruit based on the MASNet single-frame image and the depth distance information includes: Based on the depth distance information and the third coordinate system of the large field of view sensor, pseudo point cloud information is constructed, and a bird's-eye view coordinate system is established based on the pseudo point cloud information, so as to perceive the target fruit according to the bird's-eye view coordinate system. Based on the MASNet single-frame image recognition results, the center pixel position of the target fruit region is determined, and the center pixel position is mapped to the bird's-eye view coordinate system to obtain the three-dimensional spatial coordinates of the target fruit in the bird's-eye view coordinate system, so as to dynamically track the target fruit according to the three-dimensional spatial coordinates.
5. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the visual bionic method for fruit picking robots according to any one of claims 1 to 3.
6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the visual bionic method for fruit picking robots according to any one of claims 1 to 3.
Citation Information
Patent Citations
String type fruit distributed visual active sensing method and application thereof
CN111602517A