A Self-Supervised Method and Device for Estimating Kiwi Blossom Cluster Depth Based on Monocular Vision

By employing a self-supervised kiwifruit flower cluster depth estimation method based on monocular vision, and utilizing an improved Yolov8 network and Criss Cross AT module, the problems of low manual efficiency and resource waste in kiwifruit pollination were solved, and precision pollination was achieved.

CN120525941BActive Publication Date: 2025-10-31SHAANXI LOGISTICS GRP IND RES INST CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511013179.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-10-31
Estimated Expiration
2045-07-23

AI Technical Summary

Technical Problem

Existing kiwi pollination technologies suffer from low efficiency in manual pollination and waste of resources in mechanical pollination, especially the lack of precise positioning of flower clusters.

Method used

A self-supervised kiwifruit flower cluster depth estimation method based on monocular vision was adopted. An improved Yolov8 network was used for flower cluster detection and depth estimation. The depth perception capability was enhanced by combining the Criss Cross AT module. Precise pollination was achieved by adjusting the pose of the monocular camera and pollen injector.

Benefits of technology

Precise pollination of kiwifruit flower clusters was achieved with limited computing power, reducing equipment complexity and resource waste, and improving pollination efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120525941B_ABST
    Figure CN120525941B_ABST
Patent Text Reader

Abstract

This invention discloses a self-supervised method and device for estimating the depth of kiwifruit flower clusters based on monocular vision, solving the problems of low efficiency in manual pollination and waste of resources in mechanical pollination in existing assisted pollination techniques for kiwifruit. The method includes: acquiring images of kiwifruit flower clusters collected in the field by a monocular camera; cropping the images to obtain cropped images of the kiwifruit flower clusters; inputting the cropped images of the kiwifruit flower clusters into a pre-trained improved Yolov8 network to obtain flower cluster detection results and depth estimation results; adjusting the posture of the pollen injector based on the flower cluster detection results and depth estimation results to complete the kiwifruit flower cluster pollination; using a backbone network for feature extraction to obtain features at different layers; using a neck network for feature fusion based on features at different layers to obtain different fused features; and using a head network to obtain the flower cluster detection results; achieving accurate pollination with limited computing power.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a self-supervised method and apparatus for estimating the depth of kiwifruit flower clusters based on monocular vision. Background Technology

[0002] Shaanxi Province in my country is a major kiwifruit producing province. Its kiwifruit is rich in vitamin C, dietary fiber, vitamin E, potassium, and other nutrients, making it very popular with consumers. Kiwifruit not only offers numerous benefits such as promoting digestion, boosting immunity, regulating blood sugar, and improving cardiovascular health, but also holds an important position in domestic and international markets due to its unique flavor and rich nutritional value. However, in the kiwifruit production process, the application of pollination techniques directly affects the yield and quality of the fruit.

[0003] Because kiwifruit has poor natural pollination ability and a low fruit set rate, assisted pollination is a crucial step in kiwifruit production. Existing assisted pollination technologies for kiwifruit mainly include manual pollination and mechanical pollination. Manual pollination typically uses a spot-applied technique, where temporary workers are hired during the kiwifruit flowering period to manually apply pollen to the stigma of the female flower. However, this method is ill-suited to the short flowering period of about one week, and the varying skill levels of the workers make it difficult to guarantee stable fruit yield and quality. To reduce the cost of manual pollination, mechanized assisted pollination technologies are widely used. Currently, common mechanical pollination methods include spray pollination and dry powder pollination. Spray pollination uses an atomizing device to suspend pollen in a liquid and spray it onto the surface of the flower clusters. This method can cover a large area of ​​flowers, but there is a potential impact of the liquid medium on pollen activity, and the amount of pollen carried by the liquid is difficult to control, resulting in low pollination uniformity. Dry powder pollination uses mechanical equipment to evenly spread dried pollen among the kiwifruit flowers. Although this method avoids the influence of liquid media on pollen activity, the lack of precise positioning of flower clusters means that a large amount of pollen may be lost in the air or to non-target areas, resulting in pollen waste and reduced pollination efficiency.

[0004] In summary, current methods of assisted pollination for kiwifruit suffer from problems such as low efficiency of manual pollination and waste of resources in mechanical pollination. Summary of the Invention

[0005] This invention provides a self-supervised method and device for estimating the depth of kiwi flower clusters based on monocular vision, which solves the problems of low efficiency of manual pollination and waste of resources in mechanical pollination in the existing technology of kiwi assisted pollination, and realizes accurate pollination with low computing power.

[0006] In a first aspect, the present invention provides a self-supervised method for estimating the depth of kiwifruit flower clusters based on monocular vision, the method comprising:

[0007] Acquire images of kiwifruit flower clusters captured in the field using a monocular camera, and then crop the images of the kiwifruit flower clusters to obtain cropped images of the kiwifruit flower clusters.

[0008] The cropped image of the kiwifruit flower cluster is input into a pre-trained improved Yolov8 network to obtain the flower cluster detection results and depth estimation results. The posture of the pollen injector is adjusted according to the flower cluster detection results and the depth estimation results to complete the pollination of the kiwifruit flower cluster. The improved Yolov8 network includes a Backbone network, a Neck network and a Head network.

[0009] The Backbone network is used to extract hierarchical features from the cropped kiwi flower cluster image to obtain shallow features, medium features, and deep features.

[0010] The Neck network is used to fuse the shallow features, the middle features and the deep features to obtain the first depth estimation feature output by the depth estimation pre-module and the first fused feature corresponding to the instance detection pre-module.

[0011] The Head network is used to extract cross-attention features from the first depth estimation features to obtain a Criss Cross feature map; to perform depth estimation on each pixel in the Criss Cross feature map to obtain a depth estimation result; and to detect flower cluster targets using the first fused features to obtain a flower cluster detection result.

[0012] Secondly, the present invention provides a self-supervised kiwifruit flower cluster depth estimation device based on monocular vision, the device comprising:

[0013] An image acquisition and processing unit is used to acquire images of kiwifruit flower clusters captured by a monocular camera in the field, and to crop the images of kiwifruit flower clusters to obtain cropped images of kiwifruit flower clusters.

[0014] The image processing unit is used to perform flower cluster detection and depth estimation on the cropped kiwi flower cluster image based on the pre-trained improved Yolov8 network, and obtain flower cluster detection results and depth estimation results, so as to adjust the posture of the pollen injector according to the flower cluster detection results and the depth estimation results to complete the pollination of kiwi flower clusters.

[0015] One or more technical solutions provided in this invention have at least the following technical effects or advantages:

[0016] This invention reduces the number of sensors required by the device by using a monocular camera, making it more suitable for application in field environments with obstacles; it improves the Yolov8 network to reduce computational complexity, and the Criss Cross AT module uses cross-local correlation instead of window correlation in the convolutional layers of the prior art to enhance the depth perception capability of flower clusters; it solves the problems of low efficiency of manual pollination and waste of resources in mechanical pollination of kiwifruit in the prior art, and realizes accurate pollination with low computing power. Attached Figure Description

[0017] Figure 1 A flowchart illustrating the steps of a self-supervised kiwifruit flower cluster depth estimation method based on monocular vision provided in an embodiment of the present invention;

[0018] Figure 2 This is a schematic diagram of the structure of the improved Yolov8 network provided in an embodiment of the present invention;

[0019] Figure 3 This is a schematic diagram of the processing flow corresponding to the Criss Cross AT module provided in an embodiment of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0021] A self-supervised method for estimating the depth of kiwifruit flower clusters based on monocular vision, see [link to relevant documentation]. Figure 1 The method includes the following steps S101 to S102.

[0022] S101, acquire images of kiwifruit flower clusters collected in the field by a monocular camera, crop the kiwifruit flower cluster images to obtain cropped kiwifruit flower cluster images.

[0023] For example, a monocular camera acquires images of kiwifruit flower clusters in a field environment, filters out images containing kiwifruit flower clusters, and crops the acquired kiwifruit flower cluster images into 640 pixels. A cropped image of a cluster of kiwi blossoms, 640 pixels in size.

[0024] S102, the cropped kiwi flower cluster image is input into the pre-trained improved Yolov8 network to obtain the flower cluster detection results and depth estimation results, so as to adjust the posture of the pollen injector according to the flower cluster detection results and depth estimation results to complete the pollination of kiwi flower clusters.

[0025] See Figure 2 Improvements to the Yolov8 network include: Backbone network, Neck network, and Head network.

[0026] The Backbone network is used to extract hierarchical features from cropped images of kiwifruit flower clusters, obtaining shallow, medium, and deep features.

[0027] Specifically, the Backbone network comprises three sequentially connected Stage 1, Stage 2, and Stage 3 modules; it performs hierarchical feature extraction on the cropped kiwifruit flower cluster image to obtain shallow, mid-level, and deep features, including:

[0028] The Stage1 module performs shallow feature extraction on the cropped kiwi flower cluster image to obtain shallow features; wherein, the Stage1 module includes a first Conv layer, a first Backbone C2F layer, a second Conv layer, and a second Backbone C2F layer connected in sequence.

[0029] The Stage2 module extracts mid-level features from shallow features to obtain mid-level features; the Stage2 module includes a third Conv layer and a third Backbone C2f layer connected in sequence.

[0030] The Stage3 module extracts deep features from the mid-level features to obtain deep features; the Stage3 module includes the fourth Conv layer, the fourth Backbone C2f layer, the fifth Conv layer, and the SPFF layer connected in sequence.

[0031] The Neck network includes a feature enhancement module, a depth estimation pre-module, and an instance detection pre-module connected in sequence. The Neck network is used to fuse shallow features, mid-level features, and deep features to obtain the first depth estimation feature output by the depth estimation pre-module and the first fused feature corresponding to the instance detection pre-module.

[0032] Specifically, the Neck network fuses shallow, mid-level, and deep features using a feature enhancement module, a depth estimation pre-module, and an instance detection pre-module to obtain first depth estimation features and first fused features, including:

[0033] The Neck network fuses mid-level and deep-level features using a feature enhancement module to obtain the first enhanced feature; the feature enhancement module includes a first Upsample layer, a first Concat layer, and a first NeckC2f layer connected in sequence.

[0034] The Neck network fuses the first enhanced features and shallow features according to the depth estimation pre-module to obtain the first depth estimation features; wherein, the depth estimation pre-module includes a second Upsample layer, a second Concat layer and a second Neck C2f layer connected in sequence;

[0035] The Neck network performs feature fusion on the first enhanced feature, the first depth estimation feature, and the deep feature according to the instance detection pre-module to obtain the first fused feature; wherein, the instance detection pre-module includes the sixth Conv layer, the third Concat layer, the third Neck C2f layer, the seventh Conv layer, the fourth Concat layer, and the fourth Neck C2f layer connected in sequence.

[0036] For example, the Neck network consists of multiple alternating Upsample layers, Conv layers, Concat layers, and C2f layers.

[0037] In the feature enhancement module of the Backbone network, the input of the first Upsample layer is the deep features; the input of the first Concat layer is the output of the first Upsample layer and the mid-level features.

[0038] The input to the second Upsample layer of the depth estimation pre-module of the Backbone network is the output of the first Neck C2f layer; the input to the second Concat layer is the output of the second Upsample layer and shallow features.

[0039] The first Conv layer of the instance detection pre-module of the Backbone network takes the output of the second Neck C2f layer as input; the third Concat layer takes the output of the first Neck C2f layer and the output of the first Conv layer as input; and the fourth Concat layer takes the deep features and the output of the second Conv layer as input.

[0040] The Head network is used to extract the CrissCross feature map by performing cross-attention feature extraction on the first depth estimation feature map; depth estimation is performed on each pixel in the CrissCross feature map to obtain the depth estimation result; and flower cluster target detection is performed on the first fused feature map to obtain the flower cluster detection result.

[0041] Specifically, the Criss Cross AT module in the Head network includes: three sequentially connected... The system consists of convolutional layers, attention weight generation layers, softmax layers, aggregation layers, and fusion layers. It extracts Criss Cross feature maps from the first depthwise estimated features using cross-attention technology, including:

[0042] According to three The convolutional layers extract depth features from the first depth estimation features to obtain the query matrix, key matrix, and value matrix.

[0043] Based on the attention weight generation layer, the similarity of corresponding positions is calculated using the query matrix and the key matrix to obtain the attention weight vector matrix;

[0044] The attention map is obtained by converting the attention weight vector matrix into a probability distribution using the Softmax layer.

[0045] Based on the Aggregation layer, the attention map is used to aggregate the value matrix to obtain the position features corresponding to each position in the first depth estimation features;

[0046] The fusion layer performs an Add operation on the location features corresponding to each location and the depth features at the corresponding locations of the first depth estimation features to obtain the Criss Cross feature map.

[0047] For example, see Figure 3 The Criss Cross AT module allows YOLOv8 to incorporate flower cluster shapes, which helps obtain depth maps that conform to the flower cluster geometry.

[0048] First, the deep feature map obtained from the second Neck C2f layer of the aforementioned Neck network is sequentially processed through three... The convolutional layer obtains the query matrix. Key matrix Sum matrix , ;in, Indicates the height of the depth feature map; Indicates the width of the depth feature map; This represents the number of channels in the depth feature map. It represents the set of real numbers.

[0049] Then, by querying the matrix Key matrix Perform Affinity calculation to obtain the attention weight vector matrix;

[0050] Next, perform a softmax operation on the attention weight vector matrix to obtain the attention map. ;

[0051] Finally, the output is obtained after feature fusion with deep features through aggregation operations. That is, the location features corresponding to each position in the first depth estimation features; the specific steps are as follows:

[0052] The aforementioned deep features are processed through three independent convolutional layers to generate a query matrix. Key matrix Sum matrix ;

[0053] Next, for the first depth estimation feature, the position is... query vector Collection location is Key-value pairs in the row and column:

[0054] ;

[0055] ;

[0056] Calculate query vector Set of keys The similarity between elements in the dataset is used to generate attention weights, which is the Affinity operation.

[0057] Then, Softmax is performed to obtain the attention map. The location is Attention weights ;

[0058] Next, through attention weights The weighted set, used as an aggregation operation, yields the feature at position . Thus, the features of all positions are obtained as follows: ;

[0059] Finally, the location features corresponding to each location are added to the depth features at the corresponding locations of the first depth estimation features to obtain the Criss Cross feature map. .

[0060] In this invention, the improved Yolov8 network is trained using a self-supervised training strategy, and the training process includes:

[0061] (1) Obtain Images of kiwi blossom clusters at different times and Images of kiwi blossom clusters at a specific moment;

[0062] (2) Images of kiwi blossom clusters at specific times are input into an improved Yolov8 network to obtain... Depth estimation results at any given time; simultaneously, acquiring depth data from the monocular camera... Time to A sequence of camera pose changes at different times;

[0063] (3) According to Images of kiwi blossom clusters at different times Time to The sequence of camera pose changes at different times and The depth estimation results at time step 1 are then reprojected using a reprojection method to obtain... The predicted depth estimation results at time;

[0064] (4) Calculation Depth estimation results at time and The self-supervised error between the predicted depth estimation results at time step is used to update the parameters of the improved Yolov8 network, resulting in a pre-trained improved Yolov8 network.

[0065] For example, will Images of kiwi blossom clusters at specific times are input into an improved Yolov8 network to obtain... Depth estimation results at time step Read from the δMove module in the monocular camera Time to Camera pose change sequence at time points Camera pose change sequence Used to describe the relative rotation and translation of a camera between two consecutive frames;

[0066] Next, according to Images of kiwi blossom clusters at different times Time to The sequence of camera pose changes at different times The depth estimation results at time step are combined with the monocular camera's intrinsic parameters O and reprojected using the reprojection method. The reprojected data is then used to generate... Predicted depth estimation results at time .

[0067] Predicted depth estimation results at time and The kiwi blossom cluster images at time points are used to constrain errors using a self-supervised loss function. Training continues until the self-supervised loss function no longer decreases. The specific self-supervised loss function is expressed as follows:

[0068] ;

[0069] in, express Local windows in the depth estimation results at time step; express Local windows in the predicted depth estimation results at time points; Represents a local window The average value of the middle pixels; Represents a local window The average value of the middle pixels; Represents a local window The variance of the mid-pixel; Represents a local window The variance of the mid-pixel; Represents a local window Medium pixels and local windows Covariance of mid-pixel; It is the first constant; It is the second constant; express The depth estimation results at time or The total number of pixels in the predicted depth estimation result at time step; express The depth estimation results at time 1 1 pixel; express The first time in the predicted depth estimation results 1 pixel; This represents the improved self-supervised loss of the Yolov8 network.

[0070] The loss function of the improved Yolov8 network during training is expressed as follows:

[0071] ;

[0072] in, This represents the ratio of the intersection to the union of the ground truth bounding boxes and the predicted bounding boxes. Indicates the first weighting coefficient; This represents the Euclidean distance between the ground truth bounding box and the predicted bounding box before calculation. Indicates the position of the center point of the actual bounding box; Indicates the position of the center point of the prediction box; This represents the diagonal length of the bounding box between the ground truth bounding box and the predicted bounding box. This represents the second weighting coefficient; This represents the width difference between the ground truth bounding box and the predicted bounding box. This represents the height difference between the ground truth bounding box and the predicted bounding box; Indicates the third weighting coefficient; Indicates the predicted probability; This represents the fourth weighting coefficient; This represents the total loss of the improved Yolov8 network.

[0073] During the training of the improved Yolov8 network, images of kiwi flower clusters from the training set are input into the improved Yolov8 network to obtain a flower cluster detection result set and a depth estimation result set. The flower cluster detection result set is scanned row by row. Based on the coordinate position and depth estimation result of the corresponding flower cluster detection result, starting from the coordinate position at the top left corner, the camera is moved horizontally and vertically using the difference between the center point position of the training set and the coordinate position, until the flower cluster is located in the center of the image of the monocular camera. The camera pose is then adjusted using the depth estimation result corresponding to the coordinate position until the imaging area includes the entire flower cluster.

[0074] Specifically, in step S102, the orientation of the pollen injector is adjusted based on the flower cluster detection results and depth estimation results to complete the pollination of the kiwifruit flower clusters, including:

[0075] S1021, Based on the depth estimation results and flower cluster detection results, determine the relative displacement between the flower cluster and the pollen injector;

[0076] S1022, adjust the posture of the pollen injector according to the relative displacement to complete the pollination of the flower cluster.

[0077] For example, the relative displacement between the flower cluster and the pollen injector is determined using depth estimation results and flower cluster detection results;

[0078] A monocular camera is mounted on a pollen injector. The attitude of the monocular camera in the pollen injector is adaptively adjusted according to the relative displacement so that the flower cluster is located at the center of the camera image. At the same time, the image sequence received by the monocular camera sensor and the corresponding camera attitude change are recorded frame by frame during the attitude adjustment process.

[0079] Secondly, the present invention provides a self-supervised kiwifruit flower cluster depth estimation device based on monocular vision, the device comprising:

[0080] The image acquisition and processing unit is used to acquire images of kiwifruit flower clusters collected by a monocular camera in the field, and to crop the images of kiwifruit flower clusters to obtain cropped images of kiwifruit flower clusters.

[0081] The image processing unit is used to perform flower cluster detection and depth estimation on the cropped kiwi flower cluster image based on the pre-trained improved Yolov8 network, and obtain the flower cluster detection results and depth estimation results. Based on the flower cluster detection results and depth estimation results, the posture of the pollen injector is adjusted to complete the pollination of the kiwi flower cluster.

[0082] For example, when using the device provided by the present invention, firstly, images of kiwifruit flower clusters in the field are acquired, and the kiwifruit flower cluster images are cropped to obtain cropped kiwifruit flower cluster images; then, the cropped kiwifruit flower cluster images are input into a pre-trained improved Yolov8 network to obtain flower cluster detection results and depth estimation results.

[0083] Subsequently, the monocular camera mounted on the pollen injector is adaptively adjusted based on the depth estimation results, and then the flower clusters are pollinated based on the flower cluster detection results to complete the pollination task.

[0084] The various embodiments described in this specification are presented in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. All or part of this invention can be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, mobile communication terminals, multiprocessor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.

[0085] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the present invention.

Claims

1. A self-supervised method for estimating the depth of kiwifruit flower clusters based on monocular vision, characterized in that, include: Acquire images of kiwifruit flower clusters captured in the field using a monocular camera, and then crop the images of the kiwifruit flower clusters to obtain cropped images of the kiwifruit flower clusters. The cropped image of the kiwifruit flower cluster is input into a pre-trained improved Yolov8 network to obtain the flower cluster detection results and depth estimation results. The posture of the pollen injector is adjusted according to the flower cluster detection results and the depth estimation results to complete the pollination of the kiwifruit flower cluster. The improved Yolov8 network includes a Backbone network, a Neck network and a Head network. The Backbone network is used to extract hierarchical features from the cropped kiwi flower cluster image to obtain shallow features, medium features, and deep features. The Neck network is used to fuse the shallow features, the middle features, and the deep features to obtain a first depth estimation feature output by the depth estimation pre-module and a first fused feature corresponding to the instance detection pre-module; the Neck network includes: a feature enhancement module, a depth estimation pre-module, and an instance detection pre-module connected in sequence; The Neck network performs feature enhancement on the shallow features, the mid-layer features, and the deep features to obtain a first depth estimation feature output by the depth estimation pre-module and a first fusion feature corresponding to the instance detection pre-module, including: The feature enhancement module performs feature enhancement on the middle-layer features and the deep-layer features to obtain the first enhanced feature; The depth estimation pre-module performs depth estimation on the first enhanced feature and the shallow feature to obtain the first depth estimation feature; The instance detection pre-module performs instance detection on the first enhanced feature, the first depth estimation feature, and the deep feature to obtain the first fused feature. The Head network is used to extract cross-attention features from the first depth estimation features to obtain a Criss Cross feature map; to perform depth estimation on each pixel in the Criss Cross feature map to obtain a depth estimation result; and to detect flower cluster targets using the first fused features to obtain a flower cluster detection result. The Head network includes a Criss Cross AT module, which comprises three sequentially connected modules. Convolutional layers, attention weight generation layers, softmax layers, aggregation layers, and fusion layers; The step of extracting Criss Cross feature maps from the first depth estimation features includes: According to the three The convolutional layers extract depth features from the first depth estimation features to obtain the query matrix, key matrix, and value matrix; Based on the attention weight generation layer, the similarity of corresponding positions is calculated using the query matrix and the key matrix to obtain the attention weight vector matrix; The attention weight vector matrix is ​​converted into a probability distribution based on the Softmax layer to obtain the attention map; Based on the Aggregation layer, the attention map is used to aggregate the value matrix to obtain the position features corresponding to each position in the first depth estimation features; The fusion layer performs an Add operation on the positional features corresponding to each position and the depth features at the corresponding positions of the first depth estimation features to obtain a Criss Cross feature map.

2. The self-supervised kiwifruit flower cluster depth estimation method based on monocular vision according to claim 1, characterized in that, The Backbone network includes three Stage1 modules, Stage2 modules, and Stage3 modules connected in sequence; The Backbone network extracts features from the cropped kiwifruit flower cluster image, obtaining shallow, medium, and deep features, including: The Stage1 module is used to extract shallow features from the cropped kiwi flower cluster image to obtain shallow features. The Stage2 module is used to extract mid-level features from the shallow features to obtain mid-level features. The Stage3 module is used to extract deep features from the mid-level features to obtain deep features.

3. The self-supervised kiwifruit flower cluster depth estimation method based on monocular vision according to claim 1, characterized in that, The improved Yolov8 network is trained using a self-supervised training strategy. The training process includes: Get Images of kiwi blossom clusters at different times and Images of kiwi blossom clusters at a specific moment; The The image of the kiwi blossom cluster at a given time is input into the improved Yolov8 network to obtain... Depth estimation results at any given time; simultaneously, acquiring depth data from the monocular camera... Time to A sequence of camera pose changes at different times; According to the above Images of kiwi blossom clusters at different times Time to The sequence of camera pose changes at time points and the aforementioned The depth estimation results at time step 1 are then reprojected using a reprojection method to obtain... The predicted depth estimation results at time; Calculate the The depth estimation results at time and the above The self-supervised error between the predicted depth estimation results at time step is used to update the parameters of the improved Yolov8 network, resulting in a pre-trained improved Yolov8 network.

4. The self-supervised kiwifruit flower cluster depth estimation method based on monocular vision according to claim 3, characterized in that, The calculation The depth estimation results at time and the above The self-supervised error between the predicted depth estimation results at time points includes: calculating the self-supervised error based on the self-supervised loss function, wherein the self-supervised loss function is expressed as: ; in, express Local windows in the depth estimation results at time step; express Local windows in the predicted depth estimation results at time points; Represents a local window The average value of the middle pixels; Represents a local window The average value of the middle pixels; Represents a local window The variance of the mid-pixel; Represents a local window The variance of the mid-pixel; Represents a local window Medium pixels and local windows Covariance of mid-pixel; It is the first constant; It is the second constant; express The depth estimation results at time or The total number of pixels in the predicted depth estimation result at time step; express The depth estimation results at time 1 1 pixel; express The first time in the predicted depth estimation results 1 pixel; This represents the improved self-supervised loss of the Yolov8 network.

5. The self-supervised kiwifruit flower cluster depth estimation method based on monocular vision according to claim 1, characterized in that, The loss function of the improved Yolov8 network during training is expressed as follows: ; in, This represents the ratio of the intersection to the union of the ground truth bounding boxes and the predicted bounding boxes. Indicates the first weighting coefficient; This represents the Euclidean distance between the ground truth bounding box and the predicted bounding box before calculation. Indicates the position of the center point of the actual bounding box; Indicates the position of the center point of the prediction box; This represents the diagonal length of the bounding box between the ground truth bounding box and the predicted bounding box. This represents the second weighting coefficient; This represents the width difference between the ground truth bounding box and the predicted bounding box. This represents the height difference between the ground truth bounding box and the predicted bounding box; Indicates the third weighting coefficient; Indicates the predicted probability; This represents the fourth weighting coefficient; This represents the total loss of the improved Yolov8 network.

6. The self-supervised kiwifruit flower cluster depth estimation method based on monocular vision according to claim 1, characterized in that, The step of adjusting the pollen injector's orientation based on the flower cluster detection results and the depth estimation results to complete kiwi flower cluster pollination includes: Based on the depth estimation results and the flower cluster detection results, the relative displacement between the flower cluster and the pollen injector is determined; The orientation of the pollen injector is adjusted according to the relative displacement to complete the pollination of the flower clusters.

7. A self-supervised kiwifruit blossom cluster depth estimation device based on monocular vision, characterized in that, include: An image acquisition and processing unit is used to acquire images of kiwifruit flower clusters captured by a monocular camera in the field, and to crop the images of kiwifruit flower clusters to obtain cropped images of kiwifruit flower clusters. An image processing unit is used to input the cropped kiwi flower cluster image into a pre-trained improved Yolov8 network to obtain flower cluster detection results and depth estimation results, so as to adjust the posture of the pollen injector according to the flower cluster detection results and the depth estimation results to complete the pollination of kiwi flower clusters; wherein, the improved Yolov8 network includes: Backbone network, Neck network and Head network; The Backbone network is used to extract hierarchical features from the cropped kiwi flower cluster image to obtain shallow features, medium features, and deep features. The Neck network is used to fuse the shallow features, the middle features, and the deep features to obtain a first depth estimation feature output by the depth estimation pre-module and a first fused feature corresponding to the instance detection pre-module; the Neck network includes: a feature enhancement module, a depth estimation pre-module, and an instance detection pre-module connected in sequence; The Neck network performs feature enhancement on the shallow features, the mid-layer features, and the deep features to obtain a first depth estimation feature output by the depth estimation pre-module and a first fusion feature corresponding to the instance detection pre-module, including: The feature enhancement module performs feature enhancement on the middle-layer features and the deep-layer features to obtain the first enhanced feature; The depth estimation pre-module performs depth estimation on the first enhanced feature and the shallow feature to obtain the first depth estimation feature; The instance detection pre-module performs instance detection on the first enhanced feature, the first depth estimation feature, and the deep feature to obtain the first fused feature. The Head network is used to extract cross-attention features from the first depth estimation features to obtain a Criss Cross feature map; to perform depth estimation on each pixel in the Criss Cross feature map to obtain a depth estimation result; and to detect flower cluster targets using the first fused features to obtain a flower cluster detection result. The Head network includes a Criss Cross AT module, which comprises three sequentially connected modules. Convolutional layers, attention weight generation layers, softmax layers, aggregation layers, and fusion layers; The step of extracting Criss Cross feature maps from the first depth estimation features includes: According to the three The convolutional layers extract depth features from the first depth estimation features to obtain the query matrix, key matrix, and value matrix; Based on the attention weight generation layer, the similarity of corresponding positions is calculated using the query matrix and the key matrix to obtain the attention weight vector matrix; The attention weight vector matrix is ​​converted into a probability distribution based on the Softmax layer to obtain the attention map; Based on the Aggregation layer, the attention map is used to aggregate the value matrix to obtain the position features corresponding to each position in the first depth estimation features; The fusion layer performs an Add operation on the positional features corresponding to each position and the depth features at the corresponding positions of the first depth estimation features to obtain a Criss Cross feature map.

Citation Information

Patent Citations

  • Automatic pollination method and device for kiwi fruits, pollination equipment and storage medium

    CN113902675A

  • Automatic pollination method, device and equipment for strawberry flowers, medium and pollination equipment

    CN118140808A

  • Lightweight improved pepper flower small target detection method and system based on YOLOv8n

    CN120298670A