Fruit Detection Method, Device, Electronic Device and Storage Medium
By fusing two-dimensional images and point cloud image features, the problem of inaccurate spatial position determination in fruit detection is solved, and the precise identification and positioning of the fruit picking robot is achieved, which improves the accuracy of the detection.
Patent Information
- Application Number
- CN202210626390.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-02
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-06-02
AI Technical Summary
In the prior art, fruit detection cannot accurately determine the spatial position of the fruit, two-dimensional coordinate detection cannot characterize the spatial position, the depth image field angle is small and depends on the accuracy of the acquisition equipment, and the fruit boundary cannot be divided during occlusion.
By acquiring two-dimensional images and point cloud images, integrating two-dimensional image features and point cloud features, using image features to compensate for the edge and surface features of the fruit, enhancing the point cloud feature characterization ability, and combining depth information to determine the spatial position of the fruit.
It improves the accuracy of fruit detection, ensures that the fusion characteristics can fully reflect the spatial position of the fruit, and realizes the precise identification and positioning of the fruit picking robot.
Smart Images

Figure CN115018789B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a fruit detection method, device, electronic device, and storage medium. Background Art
[0002] With the rapid development of computer vision, the application scope of image recognition technology is becoming more and more extensive. In the field of agricultural production, by combining image recognition technology with robots, fruit picking robots can be developed to replace manual picking and achieve the intelligentization of agricultural production. For fruit picking robots, detecting the position of fruits is the key to performing tasks such as fruit and vegetable picking and pest and disease identification.
[0003] Currently, most perform two-dimensional object detection on fruit images to determine the coordinate position of fruits in the fruit images. However, the two-dimensional coordinate position cannot accurately represent the spatial position of fruits. Based on this, replacing the fruit image with a depth image and performing semantic segmentation on the depth image to obtain the fruit boundary. However, the depth image has the problems of a small field of view angle and the accuracy of the depth information of the depth image being too dependent on the acquisition device of the depth image, resulting in an inability to accurately obtain the spatial position of fruits.
[0004] In summary, how to accurately locate the spatial position of fruits is an urgent problem to be solved currently. Summary of the Invention
[0005] The present invention provides a fruit detection method, device, electronic device, and storage medium to solve the defect of low accuracy in fruit detection in the prior art.
[0006] The present invention provides a fruit detection method, including:
[0007] Obtain a two-dimensional image of the fruit region to be recognized and a point cloud image of the fruit region to be recognized;
[0008] Fuse the image features of the two-dimensional image and the point cloud features of the point cloud image to obtain fused features, where the image features are used to represent the position information of the fruits in the two-dimensional image;
[0009] Perform fruit detection based on the fused features to obtain the spatial position of the fruits in the fruit region to be recognized.
[0010] According to the fruit detection method provided by the present invention, the fusing the image features of the two-dimensional image and the point cloud features of the point cloud image to obtain fused features includes:
[0011] Obtain the image features of the two-dimensional image at at least two levels;
[0012] Fuse the point cloud features of the point cloud image at the current level with the image features at at least one level among the at least two levels to obtain cross features at the current level;
[0013] Extract features from the cross features at the current level to obtain point cloud features at the next level, and use the next level as the current level until the current level is the last level;
[0014] Determine the fusion features based on the point cloud features at the last level.
[0015] According to a fruit detection method provided by the present invention, the step of fusing the point cloud features of the point cloud image at the current level with the image features at at least one level among the at least two levels to obtain cross features at the current level includes:
[0016] Determine the fusion weights corresponding to the image features at at least one level among the at least two levels based on the image features at at least one level among the at least two levels;
[0017] Based on the fusion weights corresponding to the image features at at least one level, perform feature enhancement processing on the point cloud features at the current level respectively to obtain enhanced features corresponding to the image features at at least one level;
[0018] Fuse the enhanced features corresponding to the image features at at least one level with the point cloud features at the current level respectively to obtain cross-fusion features corresponding to the image features at at least one level;
[0019] Fuse the cross-fusion features corresponding to the image features at at least one level to obtain cross features at the current level.
[0020] According to a fruit detection method provided by the present invention, the step of fusing the image features of the two-dimensional image and the point cloud features of the point cloud image to obtain fusion features includes:
[0021] Obtain the image features of the two-dimensional image at one level;
[0022] Fuse the point cloud features of the point cloud image at the current level with the image features at the one level to obtain cross features at the current level;
[0023] Extract features from the cross features at the current level to obtain point cloud features at the next level, and use the next level as the current level until the current level is the last level;
[0024] Determine the fusion features based on the point cloud features at the last level.
[0025] According to a fruit detection method provided by the present invention, the point cloud features of the point cloud image at the first level are determined based on the following steps:
[0026] Based on the internal parameter matrix and external parameter matrix of the acquisition device of the point cloud image, the point cloud image is converted into a two-dimensional pseudo-image;
[0027] Based on the two-dimensional pseudo-image, the point cloud features of the point cloud image at the first level are determined.
[0028] According to a fruit detection method provided by the present invention, after obtaining the two-dimensional image of the fruit area to be recognized, the following steps are further included:
[0029] Based on the rot area detection model, the image features of the two-dimensional image are detected for rot areas to obtain the fruit rot areas;
[0030] The rot area detection model is trained based on the first sample two-dimensional image marked with the fruit area and the second sample two-dimensional image marked with the fruit area and the rot area.
[0031] According to a fruit detection method provided by the present invention, the step of detecting the rot area of the two-dimensional image based on the rot area detection model to obtain the fruit rot area includes:
[0032] Based on the multi-channel feature extraction layer in the rot area detection model, multi-channel feature extraction is performed on the image features to obtain channel features of at least two channels;
[0033] Based on the feature fusion layer in the rot area detection model, weighted fusion is performed on the channel features of the at least two channels to obtain fused image features;
[0034] Based on the semantic segmentation layer in the rot area detection model, rot area segmentation is performed on the fused image features to obtain the fruit rot area.
[0035] According to a fruit detection method provided by the present invention, after detecting the rot area of the two-dimensional image based on the rot area detection model to obtain the fruit rot area, the following steps are further included:
[0036] Obtain the first confidence corresponding to the spatial position and the second confidence corresponding to the fruit rot area;
[0037] When the first confidence is greater than the first threshold and the second confidence is greater than the second threshold, it is determined that the fruit is a rotten fruit;
[0038] When the first confidence level is less than or equal to the first threshold and it is determined that the second confidence level is greater than the second threshold, it is determined that a fruit exists;
[0039] When the first confidence level is less than or equal to the first threshold and it is determined that the second confidence level is less than or equal to the second threshold, it is determined that no fruit exists;
[0040] When the first confidence level is greater than the first threshold and it is determined that the second confidence level is less than or equal to the second threshold, it is determined that the fruit has no rotten area.
[0041] According to a fruit detection method provided by the present invention, the fruit detection based on the fusion feature to obtain the spatial position of the fruit in the to-be-identified fruit area includes:
[0042] Based on the fusion feature, perform fruit position detection to obtain a fruit area;
[0043] Based on the fruit area and the depth information of the point cloud image, determine the spatial position of the fruit in the to-be-identified fruit area.
[0044] The present invention also provides a fruit detection device, including:
[0045] An acquisition module, configured to acquire a two-dimensional image of a to-be-identified fruit area and a point cloud image of the to-be-identified fruit area;
[0046] A fusion module, configured to fuse the image feature of the two-dimensional image and the point cloud feature of the point cloud image to obtain a fusion feature, where the image feature is used to characterize the position information of the fruit in the two-dimensional image;
[0047] A detection module, configured to perform fruit detection based on the fusion feature to obtain the spatial position of the fruit in the to-be-identified fruit area.
[0048] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of any one of the above-mentioned fruit detection methods are implemented.
[0049] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned fruit detection methods are implemented.
[0050] The fruit detection method, device, electronic device and storage medium provided by the present invention. The image features of the two-dimensional image can be used to characterize the position information of the fruit in the two-dimensional image. Based on this, by fusing the image features of the two-dimensional image with the point cloud features of the point cloud image, the edge features and surface features of the fruit can be compensated by the image features of the two-dimensional image, so as to enhance the characterization ability of the point cloud features, enabling the fusion features for fruit detection to consider the position information of the fruit in the two-dimensional image, thus ensuring that the fusion features can more completely reflect the position information of the fruit, and further making the spatial position obtained by fruit detection based on the fusion features more accurate, that is, improving the accuracy of fruit detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0052] Figure 1 One of the schematic flowcharts of the fruit detection method provided by the present invention;
[0053] Figure 2 Another schematic flowchart of the fruit detection method provided by the present invention;
[0054] Figure 3 Schematic diagram of the cross-fusion method provided by the present invention;
[0055] Figure 4 One of the schematic flowcharts of the fruit detection method provided by the present invention;
[0056] Figure 5 Another schematic flowchart of the fruit detection method provided by the present invention;
[0057] Figure 6 Schematic diagram of the structure of the fruit detection device provided by the present invention;
[0058] Figure 7 Schematic diagram of the structure of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0060] In modern agricultural production, with the rapid development of computer vision, computer vision has been widely applied to tasks such as automatic fruit picking and pest and disease identification in fruits and vegetables. Traditional fruit picking requires farmers to pick fruits beside fruit trees. Manual picking is very labor-consuming, costly, and inefficient. Therefore, with the rapid development of computer vision, by combining image recognition technology with robots, fruit picking robots can be developed to replace manual picking and achieve the intelligence of agricultural production. For fruit picking robots, accurate identification and positioning of fruits are the key to distinguishing good fruits from rotten fruits and also the key to determining the size of fruits.
[0061] In the prior art, through target detection algorithms, target detection and instance segmentation of fruits are performed to determine the coordinate positions of fruits in fruit images. However, the determined coordinate positions are only coordinate values in two-dimensional images and cannot obtain three-dimensional positions, and thus cannot accurately represent the spatial positions of fruits.
[0062] In another prior art, semantic segmentation of the depth image of a fruit is performed to obtain the fruit boundary. However, the depth image has the problem of a small field of view angle and cannot cover a large area; and there is also the problem that the accuracy of the depth information of the depth image depends too much on the acquisition device of the depth image, resulting in the inability to accurately obtain the spatial position of the fruit; in addition, when the fruits on the fruit tree are blocked, semantic segmentation of the depth image cannot be performed to obtain the boundary of the fruit.
[0063] In view of the above problems, the present invention provides a fruit detection method. Figure 1 One of the flow schematic diagrams of the fruit detection method provided by the present invention is as Figure 1 shown. The method includes:
[0064] Step 110, obtaining a two-dimensional image of the area of the fruit to be recognized and a point cloud image of the area of the fruit to be recognized.
[0065] Here, the area of the fruit to be recognized is the area where fruit detection needs to be performed, which can be the area where the fruit tree is located or the area where the orchard is located, etc. The embodiments of the present invention do not make specific limitations on this. The area of the fruit to be recognized may include 0 fruits, 1 fruit, or multiple fruits.
[0066] Here, the two-dimensional image can be an RGB (Red Green Blue) image, an HSV (Hue Saturation Value) image, a grayscale image, etc. The embodiments of the present invention do not make specific limitations thereto. The two-dimensional image can be obtained by image acquisition using an image acquisition device; it can also be obtained by further processing the acquired image. For example, an HSV image can be obtained by further processing the RGB image obtained by an RGB image acquisition device.
[0067] In a specific embodiment, the method provided by the embodiments of the present invention is applied to a fruit picking robot. The two-dimensional image can be obtained by using an image acquisition device deployed on the fruit picking robot to acquire images of the area where the fruit picking robot advances; it can also be obtained by using an image acquisition device deployed on the orchard site to acquire images of the orchard site, and then the image acquisition device sends the two-dimensional image to the execution subject of the method provided by the embodiments of the present invention. The orchard site is the area where the fruit picking robot works.
[0068] It should be noted that the acquisition angle of the two-dimensional image can be set according to the actual situation. For example, for fruit trees of different heights, the image acquisition device for the two-dimensional image can be set at different heights or different angles.
[0069] Here, the point cloud image can be obtained by point cloud acquisition using a point cloud acquisition device; it can also be obtained by performing coordinate transformation on the acquired two-dimensional image and depth image. The embodiments of the present invention do not make specific limitations thereto.
[0070] In a specific embodiment, the method provided by the embodiments of the present invention is applied to a fruit picking robot. The point cloud image can be obtained by using a point cloud acquisition device deployed on the fruit picking robot to acquire point clouds of the area where the fruit picking robot advances; it can also be obtained by using a point cloud acquisition device deployed on the orchard site to acquire point clouds of the orchard site, and then the point cloud acquisition device sends the point cloud image to the execution subject of the method provided by the embodiments of the present invention. The orchard site is the area where the fruit picking robot works.
[0071] It should be noted that the acquisition angle of the point cloud image can be set according to the actual situation. For example, for fruit trees of different heights, the point cloud acquisition device for the point cloud image can be set at different heights or different angles.
[0072] It can be understood that the two-dimensional image and the point cloud image are images corresponding to the same area, so as to ensure that the fusion features of the two can be accurately obtained subsequently.
[0073] Step 120: Fuse the image features of the two-dimensional image and the point cloud features of the point cloud image to obtain fused features. The image features are used to represent the position information of the fruits in the two-dimensional image.
[0074] Here, the image features of the two-dimensional image are the features obtained after feature extraction from the two-dimensional image. There are various ways to extract features from a two-dimensional image. For example, features can be extracted from the two-dimensional image through CNN (Convolutional Neural Networks), Resnet50, etc.
[0075] Here, the point cloud features of the point cloud image can be the features obtained by further preprocessing the point cloud image; they can also be the features obtained after feature extraction from the point cloud image; they can also be the features obtained after feature extraction from the preprocessed point cloud image; or the features obtained by fusing the point cloud features of the previous level with the image features of the two-dimensional image and then performing further feature extraction on the fused features as the point cloud features of the next level. There are various ways to extract point cloud features. For example, point cloud features can be extracted through CNN (Convolutional Neural Networks), Resnet50, etc.
[0076] Here, the fused features are the features used for fruit detection. The image features of the two-dimensional image and the point cloud features of the point cloud image can be fused by means such as concatenation, addition, weighted fusion, etc.
[0077] It should be noted that when fusing the image features of the two-dimensional image and the point cloud features of the point cloud image, since the image features are used to represent the position information of the fruits in the two-dimensional image, the image features of the two-dimensional image can complement the edge features and surface features of the fruits, so as to enhance the representation ability of the point cloud features, enabling the fused features for subsequent fruit detection to consider the position information of the fruits in the two-dimensional image.
[0078] Furthermore, to better fuse the features of the two-dimensional image and the point cloud image, the two-dimensional image and the point cloud image are registered. Specifically, based on the internal parameter matrix and external parameter matrix of the acquisition device of the point cloud image, the point cloud image can be converted into a two-dimensional pseudo-image to fuse the image features of the two-dimensional image and the point cloud features of the two-dimensional pseudo-image to obtain fused features.
[0079] Step 130: Perform fruit detection based on the fused features to obtain the spatial positions of the fruits in the fruit region to be recognized.
[0080] It should be noted that the image features of the two-dimensional image are two-dimensional features. Therefore, it is necessary to convert the point cloud features of the point cloud image into two-dimensional features to better fuse the image features and the point cloud features. Based on this, the fused features are also two-dimensional features.
[0081] Specifically, based on the fused features and the depth information of the point cloud image, fruit detection is performed to obtain the spatial position of the fruit in the fruit region to be recognized.
[0082] In one embodiment, based on the fused features, two-dimensional fruit position detection is performed to obtain a two-dimensional fruit region; then, the two-dimensional fruit region is fused with the depth information of the point cloud image to obtain the spatial position of the fruit in the fruit region to be recognized.
[0083] In another embodiment, the fused features are fused with the depth information of the point cloud image to obtain three-dimensional features; based on the three-dimensional features, three-dimensional fruit position detection is performed to obtain the spatial position of the fruit in the fruit region to be recognized.
[0084] Here, the spatial position is a three-dimensional position, which can represent the accurate position of the fruit in space. It can be understood that the fruit region to be recognized may include multiple fruits, and correspondingly, each fruit corresponds to a spatial position.
[0085] In the method provided by the embodiment of the present invention, the image features of the two-dimensional image can be used to represent the position information of the fruit in the two-dimensional image. Based on this, by fusing the image features of the two-dimensional image and the point cloud features of the point cloud image, the edge features and surface features of the fruit can be compensated by the image features of the two-dimensional image, so as to enhance the representation ability of the point cloud features, so that the fused features for fruit detection can consider the position information of the fruit in the two-dimensional image, thereby ensuring that the fused features can more completely reflect the position information of the fruit, and further making the spatial position obtained by fruit detection based on the fused features more accurate, that is, improving the accuracy of fruit detection.
[0086] Based on the above embodiments, Figure 2 is the second schematic flow chart of the fruit detection method provided by the present invention. As Figure 2 shown, the above step 120 includes:
[0087] Step 121, obtaining the image features of the two-dimensional image at at least two levels.
[0088] Here, the two-dimensional image can be subjected to multiple feature extractions to obtain image features at multiple levels. The scale sizes of the image features at different levels are different.
[0089] In a specific embodiment, a two-dimensional image is input into an image feature extraction layer to obtain image features at at least two levels output by each extraction layer in the image feature extraction layer. Among them, each extraction layer in the image feature extraction layer is connected in sequence. More specifically, assuming that the image feature extraction layer includes four extraction layers, the two-dimensional image is input into the first extraction layer of the image feature extraction layer to obtain the image features at the first level output by the first extraction layer; the image features at the first level are input into the second extraction layer of the image feature extraction layer to obtain the image features at the second level output by the second extraction layer; the image features at the second level are input into the third extraction layer of the image feature extraction layer to obtain the image features at the third level output by the third extraction layer; the image features at the third level are input into the fourth extraction layer of the image feature extraction layer to obtain the image features at the fourth level output by the fourth extraction layer.
[0090] Among them, each extraction layer in the image feature extraction layer can be a residual network layer. The size of each residual network layer can be the same or different.
[0091] In one embodiment, the image feature extraction layer includes 4 residual network layers, and the number of residual blocks included in each residual network layer can be the same or different. For example, the first residual network layer of the image feature extraction layer includes 2 residual blocks, the second residual network layer of the image feature extraction layer includes 4 residual blocks, the third residual network layer of the image feature extraction layer includes 6 residual blocks, and the fourth residual network layer of the image feature extraction layer includes 2 residual blocks.
[0092] Step 122: Fuse the point cloud features of the point cloud image at the current level with the image features at at least one level among the at least two levels to obtain the cross features at the current level.
[0093] Here, the features corresponding to the point cloud image include point cloud features at at least one level.
[0094] If the point cloud features of the point cloud image at the current level are the point cloud features at the first level, then the point cloud features at the first level can be the features obtained by further preprocessing the point cloud image; they can also be the features obtained by extracting features from the point cloud image; or they can be the features obtained by extracting features from the preprocessed point cloud image.
[0095] If the point cloud features of the point cloud image at the current level are the point cloud features at other levels except the first level, then the point cloud features at the current level can be the features obtained by extracting features from the cross features at the previous level.
[0096] In one embodiment, the image features at at least one level may include the image features at one level. In this case, the point cloud features of the point cloud image at the current level are fused with the image features at one level to obtain the cross features at the current level.
[0097] In another embodiment, the image features at at least one level may include the image features at multiple levels. In this case, the point cloud features of the point cloud image at the current level are respectively fused with the image features at multiple levels to obtain multiple cross-fused features. Then, the multiple cross-fused features are fused to obtain the cross features at the current level.
[0098] Step 123: Extract features from the cross features at the current level to obtain the point cloud features at the next level, and use the next level as the current level until the current level is the last level.
[0099] In one embodiment, the features corresponding to the point cloud image include the point cloud features at multiple levels. In this case, after extracting the features of the cross features at the current level to obtain the point cloud features at the next level, the next level is used as the current level, and the above step 122 is returned until the current level is the last level.
[0100] In this case, the image features at at least two levels can be cross-fused with the point cloud features at at least two levels. That is to say, the image features at different levels can be cross-fused with the point cloud features at different levels, and the cross-fusion method can be set according to actual needs. For example, the cross-fusion method can be set based on the fusion of high-order features with low-order features or low-order features with high-order features. The embodiments of the present invention do not make specific limitations on this.
[0101] For ease of understanding, a specific embodiment is used here to illustrate one of the cross-fusion methods. As Figure 3 shown, the two-dimensional image is input into the feature extraction layer of the two-dimensional image, and the point cloud image is input into the feature extraction layer of the point cloud image; among them, the feature extraction layer of the two-dimensional image includes four extraction layers, and the feature extraction layer of the point cloud image includes four extraction layers; the image features of the two-dimensional image at the first level are respectively fused with the point cloud features of the point cloud image at the second level and the third level, the image features of the two-dimensional image at the second level are respectively fused with the point cloud features of the point cloud image at the second level and the third level, and the image features of the two-dimensional image at the third level are respectively fused with the point cloud features of the point cloud image at the first level and the second level; among them, the decay area detection layer for decay area detection and the fruit detection layer for fruit detection share the feature extraction layer of the two-dimensional image.
[0102] Here, the feature extraction methods for cross - features at different levels can be the same or different.
[0103] In a specific embodiment, the features corresponding to the point - cloud image include point - cloud features at 4 levels. The point - cloud features at the first level are the features obtained after pre - processing the point - cloud image. The feature extraction layer for obtaining the point - cloud features at the second level can be a residual network layer, and the residual network layer corresponding to the second level can include 2 residual blocks. The feature extraction layer for obtaining the point - cloud features at the third level can be a residual network layer, and the residual network layer corresponding to the third level can include 2 residual blocks. The feature extraction layer for obtaining the point - cloud features at the fourth level can be a residual network layer, and the residual network layer corresponding to the fourth level can include 2 residual blocks.
[0104] In another embodiment, the features corresponding to the point - cloud image include point - cloud features at one level. In this case, feature extraction is performed on the cross - features at the current level to obtain the point - cloud features at the last level.
[0105] In this case, the image features at at least two levels can be fused with the point - cloud features at one level, and feature extraction is performed on the fused features to obtain the point - cloud features at the last level. That is to say, the image features at different levels can be fused with the point - cloud features at a single level.
[0106] Step 124, determine the fused feature based on the point - cloud features at the last level.
[0107] Specifically, the point - cloud features at the last level can be directly used as the fused feature, or further feature extraction can be performed on the point - cloud features at the last level to obtain the fused feature.
[0108] The method provided in the embodiments of the present invention fuses the image features of a two - dimensional image at multiple levels with the point - cloud features of a point - cloud image at at least one level. The edge features and surface features of the fruit can be compensated for by the image features at multiple levels, thereby enhancing the representation ability of the point - cloud features at each level, so that the finally obtained fused feature can further consider the position information of the fruit in the two - dimensional image, further ensuring that the fused feature can more completely reflect the position information of the fruit, and further making the spatial position obtained by fruit detection based on the fused feature more accurate, that is, further improving the accuracy of fruit detection.
[0109] Based on any of the above - mentioned embodiments, Figure 4 This is the third flow - chart diagram of the fruit detection method provided by the present invention. As Figure 4 shown, the above - mentioned step 122 includes:
[0110] Step 1221: Based on the image features at at least one level among the at least two levels, determine the fusion weights corresponding to the image features at the at least one level.
[0111] Here, the fusion weight is the weight for feature enhancement of the point cloud features. The fusion weight can be a fusion probability. In one embodiment, the fusion weight is the fruit probability, which is obtained by performing fruit detection based on the corresponding image features and is used to represent the probability of the existence of fruits in the two-dimensional image.
[0112] In one specific embodiment, input the image features at at least one level into a classifier to obtain the fusion weights output by the classifier. Among them, the classifier can be a softmax classifier or an SVM (Support Vector Machine) classifier, etc.
[0113] It should be noted that the image features at one level correspond to one fusion weight.
[0114] For example, the fusion weight can be expressed by the following formula:
[0115] K j = softmax(R j );
[0116] In the formula, K j is the fusion weight corresponding to the image features at the j-th level, R j is the image feature at the j-th level, and softmax() is the softmax function.
[0117] Step 1222: Based on the fusion weights corresponding to the image features at the at least one level, perform feature enhancement processing on the point cloud features at the current level respectively to obtain the enhanced features corresponding to the image features at the at least one level.
[0118] Here, the enhanced feature is the enhanced partial feature for feature enhancement of the point cloud feature, and the enhanced feature is the feature corresponding to the enhanced representation ability added to the point cloud feature.
[0119] In one specific embodiment, based on the fusion weights corresponding to the image features at at least one level, perform weighted processing on the point cloud features at the current level respectively to obtain the enhanced features corresponding to the image features at at least one level.
[0120] It should be noted that the image features at one level correspond to one enhancement feature. The larger the fusion weight, the larger the scale of feature enhancement for the point cloud features; the smaller the fusion weight, the smaller the scale of feature enhancement for the point cloud features. More specifically, the larger the fruit probability, the larger the scale of feature enhancement for the point cloud features, that is, the point cloud features in the area where fruits may exist in the point cloud image are enhanced.
[0121] Fuse the image features of the two-dimensional image and the point cloud features of the point cloud image. Since the image features are used to represent the position information of the fruits in the two-dimensional image, the image features of the two-dimensional image can make up for the edge features and surface features of the fruits to enhance the representation ability of the point cloud features, so that the fusion features for subsequent fruit detection can consider the position information of the fruits in the two-dimensional image.
[0122] For example, the enhancement feature can be expressed by the following formula:
[0123] P ij = M i * softmax(R j ) ;
[0124] In the formula, P ij is the enhancement feature corresponding to the image feature at the j-th level in the case of fusing with the point cloud feature at the i-th level, M i is the point cloud feature at the i-th level, that is, the point cloud feature at the current level, R j is the image feature at the j-th level, and softmax() is the softmax function.
[0125] Step 1223: Respectively fuse the enhancement features corresponding to the image features at the at least one level with the point cloud features at the current level to obtain the cross-fusion features corresponding to the image features at the at least one level.
[0126] Here, the cross-fusion feature is a feature for enhancing the point cloud feature, and the representation ability of this cross-fusion feature is greater than that of the original point cloud feature.
[0127] Here, the enhancement feature and the point cloud feature can be fused by means such as addition, splicing, and weighted fusion.
[0128] In a specific embodiment, the enhancement features corresponding to the image features at the at least one level are respectively added to the point cloud features at the current level to obtain the cross-fusion features corresponding to the image features at the at least one level.
[0129] It should be noted that the image features at one level correspond to one cross-fusion feature.
[0130] For example, the cross-fusion feature can be represented by the following formula:
[0131] N ij = M i * softmax(R j ) + M i ;
[0132] In the formula, N ij is the cross-fusion feature corresponding to the image feature at the j-th level in the case of fusing with the point cloud feature at the i-th level, and M i is the point cloud feature at the i-th level, that is, the point cloud feature at the current level, and R j is the image feature at the j-th level, and softmax() is the softmax function.
[0133] Step 1224: Perform feature fusion on the cross-fusion features corresponding to the image features at the at least one level to obtain the cross-feature at the current level.
[0134] Here, the cross-fusion features can be fused by means such as addition, splicing, weighted fusion, etc.
[0135] In a specific embodiment, the cross-fusion features corresponding to the image features at the at least one level are added to obtain the cross-feature at the current level.
[0136] For example, when the current level is the first level and the image features at the at least one level include the image features at the third level, the cross-feature at the first level can be represented by the following formula:
[0137] N 13 = M1 * softmax(R3) + M1;
[0138] In the formula, N 13 is the cross-feature (which is also the cross-fusion feature) corresponding to the image feature at the third level in the case of fusing with the point cloud feature at the 1st level, M1 is the point cloud feature at the 1st level, that is, the point cloud feature at the current level, R3 is the image feature at the third level, and softmax() is the softmax function.
[0139] Another example, when the current level is the second level and the image features at the at least one level include the image features at the first level, the second level, and the third level, the cross-feature at the second level can be represented by the following formula:
[0140] A2 = N 21 + N 22 + N 23 ;
[0141] wherein, A2 is the cross feature at the second level, N 21 is the cross-fusion feature corresponding to the image feature at the first level in the case of fusing with the point cloud feature at the second level, N 22 is the cross-fusion feature corresponding to the image feature at the second level in the case of fusing with the point cloud feature at the second level, N 23 is the cross-fusion feature corresponding to the image feature at the third level in the case of fusing with the point cloud feature at the second level.
[0142] For another example, when the current level is the third level and the image features at at least one level include the image features at the first level and the image features at the second level, the cross feature at the third level can be expressed by the following formula:
[0143] A3 = N 31 + N 32 ;
[0144] wherein, A3 is the cross feature at the third level, N 31 is the cross-fusion feature corresponding to the image feature at the first level in the case of fusing with the point cloud feature at the third level, N 32 is the cross-fusion feature corresponding to the image feature at the second level in the case of fusing with the point cloud feature at the third level.
[0145] The method provided by the embodiments of the present invention determines at least one fusion weight through the image features at at least one level, and thus performs feature enhancement processing on the point cloud features at the current level based on at least one fusion weight, so as to enhance the representation ability of the point cloud features, so that the finally obtained fusion features can further consider the position information of the fruits in the two-dimensional image, thereby further ensuring that the fusion features can more completely reflect the position information of the fruits, and further making the spatial position obtained by fruit detection based on the fusion features more accurate, that is, further improving the accuracy of fruit detection.
[0146] Based on any of the above embodiments, the above steps 121-124 are the cases of fusing the image features of the two-dimensional image at at least two levels with the point cloud features of the point cloud image, and here is the case of fusing the image features of the two-dimensional image at one level with the point cloud features of the point cloud image. In this method, the above step 120 includes:
[0147] Step 125, obtaining the image features of the two-dimensional image at one level.
[0148] Here, the two-dimensional image is only subjected to feature extraction once, so as to obtain the image features at one level.
[0149] In a specific embodiment, a two-dimensional image is input into an image feature extraction layer to obtain image features at one level output by the image feature extraction layer. Among them, the image feature extraction layer may be a residual network layer, and the residual network layer may include 2 residual blocks, 4 residual blocks, 6 residual blocks, etc. The embodiments of the present invention do not make specific limitations on this.
[0150] Step 126: Fuse the point cloud features of the point cloud image at the current level with the image features at the one level to obtain cross features at the current level.
[0151] Here, the features corresponding to the point cloud image include point cloud features at at least one level.
[0152] If the point cloud features of the point cloud image at the current level are the point cloud features at the first level, then the point cloud features at the first level may be features obtained by further preprocessing the point cloud image; may also be features obtained by extracting features from the point cloud image; or may also be features obtained by extracting features from the preprocessed point cloud image.
[0153] If the point cloud features of the point cloud image at the current level are the point cloud features at other levels except the first level, then the point cloud features at the current level may be features obtained by extracting features from the cross features at the previous level.
[0154] Step 127: Extract features from the cross features at the current level to obtain point cloud features at the next level, and use the next level as the current level until the current level is the last level.
[0155] In an embodiment, the features corresponding to the point cloud image include point cloud features at multiple levels. In this case, after extracting features from the cross features at the current level to obtain point cloud features at the next level, use the next level as the current level and return to the above step 126 until the current level is the last level.
[0156] In this case, the image features at one level can be cross-fused with the point cloud features at at least two levels. That is to say, the image features at one level can be cross-fused with the point cloud features at different levels, and the cross-fusion method can be set according to actual needs.
[0157] Here, the feature extraction methods for the cross features at different levels may be the same or different.
[0158] In a specific embodiment, the features corresponding to the point cloud image include point cloud features at four levels. The point cloud features at the first level are the features obtained after preprocessing the point cloud image. The feature extraction layer for obtaining the point cloud features at the second level can be a residual network layer, and the residual network layer corresponding to the second level can include two residual blocks. The feature extraction layer for obtaining the point cloud features at the third level can be a residual network layer, and the residual network layer corresponding to the third level can include two residual blocks. The feature extraction layer for obtaining the point cloud features at the fourth level can be a residual network layer, and the residual network layer corresponding to the fourth level can include two residual blocks.
[0159] In another embodiment, the features corresponding to the point cloud image include point cloud features at one level. In this case, feature extraction is performed on the cross features at the current level to obtain the point cloud features at the last level.
[0160] In this case, the image features at one level can be fused with the point cloud features at one level, and feature extraction is performed on the fused features to obtain the point cloud features at the last level. That is to say, the image features at a single level can be fused with the point cloud features at a single level.
[0161] Step 128, determine the fusion feature based on the point cloud features at the last level.
[0162] Specifically, the point cloud features at the last level can be directly used as the fusion feature, or further feature extraction can be performed on the point cloud features at the last level to obtain the fusion feature.
[0163] The method provided by the embodiments of the present invention fuses the image features of a two-dimensional image at one level with the point cloud features of a point cloud image at at least one level. The edge features and surface features of the fruit can be compensated by the image features at one level, so as to enhance the representation ability of the point cloud features at each level, so that the finally obtained fusion feature can further consider the position information of the fruit in the two-dimensional image, thereby further ensuring that the fusion feature can more completely reflect the position information of the fruit, and further making the spatial position obtained by fruit detection based on the fusion feature more accurate, that is, further improving the accuracy of fruit detection.
[0164] Based on any of the above embodiments, in this method, the above step 126 includes:
[0165] Step 1261, determine the fusion weight corresponding to the image features at the one level based on the image features at the one level.
[0166] Here, the fusion weight is the weight for enhancing the point cloud features. The fusion weight can be a fusion probability. In one embodiment, the fusion weight is the fruit probability, that is, the probability of the existence of fruits in the two-dimensional image.
[0167] In a specific embodiment, the image features at one level are input into a classifier to obtain the fusion weight output by the classifier. Among them, the classifier can be a softmax classifier or an SVM (Support Vector Machine) classifier, etc.
[0168] For example, the fusion weight can be expressed by the following formula:
[0169] K = softmax(R);
[0170] In the formula, K is the fusion weight corresponding to the image features at one level, R is the image features at one level, and softmax() is the softmax function.
[0171] Step 1262: Based on the fusion weight corresponding to the image features at one level, perform feature enhancement processing on the point cloud features at the current level to obtain the enhanced features corresponding to the image features at one level.
[0172] Here, the enhanced features are the enhanced partial features for enhancing the point cloud features, and the enhanced features are the features corresponding to the increased representation ability of the point cloud features.
[0173] In a specific embodiment, based on the fusion weight corresponding to the image features at one level, perform weighted processing on the point cloud features at the current level to obtain the enhanced features corresponding to the image features at one level.
[0174] For example, the enhanced features can be expressed by the following formula:
[0175] P i = M i * softmax(R);
[0176] In the formula, P i is the enhanced feature corresponding to the image features at one level in the case of fusing with the point cloud features at the i-th level, M i is the point cloud features at the i-th level, that is, the point cloud features at the current level, R is the image features at one level, and softmax() is the softmax function.
[0177] Step 1263: Perform feature fusion on the enhanced features corresponding to the image features at one level and the point cloud features at the current level to obtain the cross-fusion features corresponding to the image features at one level.
[0178] Here, the cross-fusion feature is a feature for enhancing the point cloud feature, and the representation ability of this cross-fusion feature is greater than that of the original point cloud feature.
[0179] Here, the enhanced feature and the point cloud feature can be fused by means such as addition, splicing, weighted fusion, etc.
[0180] In a specific embodiment, the enhanced feature corresponding to the image feature at one level is added to the point cloud feature at the current level to obtain the cross-fusion feature corresponding to the image feature at one level.
[0181] For example, the cross-fusion feature can be represented by the following formula:
[0182] N i =M i *softmax(R)+M i ;
[0183] In the formula, N i is the cross-fusion feature corresponding to the image feature at one level in the case of fusing with the point cloud feature at the i-th level, M i is the point cloud feature at the i-th level, that is, the point cloud feature at the current level, R is the image feature at one level, and softmax() is the softmax function.
[0184] Step 1224, based on the cross-fusion feature corresponding to the image feature at one level, obtain the cross feature at the current level.
[0185] Here, the cross-fusion feature can be directly used as the cross feature, or further feature extraction can be performed on the cross-fusion feature to obtain the cross feature.
[0186] The method provided by the embodiments of the present invention determines a fusion weight through the image feature at one level, thereby performing feature enhancement processing on the point cloud feature at the current level based on a fusion weight, so as to enhance the representation ability of the point cloud feature, so that the finally obtained fusion feature can further consider the position information of the fruit in the two-dimensional image, thereby further ensuring that the fusion feature can more completely reflect the position information of the fruit, and further making the spatial position obtained by fruit detection based on the fusion feature more accurate, that is, further improving the accuracy of fruit detection.
[0187] Based on any of the above embodiments, in this method, the point cloud feature of the point cloud image at the first level is determined based on the following steps:
[0188] Based on the internal parameter matrix and the external parameter matrix of the acquisition device of the point cloud image, convert the point cloud image into a two-dimensional pseudo-image;
[0189] Based on the two-dimensional pseudo-image, determine the point cloud features of the point cloud image at the first level.
[0190] It should be noted that the image features of the two-dimensional image are two-dimensional features. To facilitate the fusion of the image features of the two-dimensional image and the point cloud features of the point cloud image, therefore, it is necessary to convert the point cloud image into a two-dimensional pseudo-image.
[0191] Here, the internal parameter matrix and the external parameter matrix are the inherent parameters of the acquisition device, and the embodiments of the present invention do not limit this.
[0192] Specifically, based on the internal parameter matrix, the external parameter matrix, and the conversion formula, convert the point cloud image into a two-dimensional pseudo-image.
[0193] For example, the conversion formula can be expressed by the following formula:
[0194] q = H·(K·p);
[0195] In the formula, q is the coordinate point of the two-dimensional pseudo-image; H is the external parameter matrix, which can be a matrix of size 4*4; K is the internal parameter matrix, which can be a matrix of size 4*4; p is the point of the point cloud image, which can be (x, y, z, l), (x, y) is the two-dimensional coordinate, z is the depth value, and l is the reflection intensity.
[0196] Specifically, the two-dimensional pseudo-image can be directly used as the point cloud feature at the first level, or further feature extraction can be performed on the two-dimensional pseudo-image to obtain the point cloud feature at the first level.
[0197] The method provided by the embodiments of the present invention converts the point cloud image into a two-dimensional pseudo-image to better fuse the image features of the two-dimensional image and the point cloud features of the point cloud image; in addition, based on the two-dimensional pseudo-image, the point cloud features are determined to provide support for the feature determination of the point cloud image at the first level.
[0198] Based on any of the above embodiments, in this method, in the above step 110, after obtaining the point cloud image of the fruit region to be recognized, the following steps are further included:
[0199] Perform downsampling processing on the point cloud image.
[0200] Specifically, first, count the density of the point cloud in different regions of the point cloud image. Based on the density of the point cloud, perform densification processing and sparsification processing on the point cloud image, so that the point cloud in the sparse part of the point cloud image is appropriately densified, and the point cloud in the dense part of the point cloud image is appropriately sparsified, so that the point cloud distributions in different regions of the point cloud image differ within a certain range; then, perform random downsampling processing on the point cloud image after densification processing and sparsification processing to obtain the downsampled point cloud image.
[0201] The method provided by the embodiment of the present invention downsamples the point cloud image to obtain a sparse point cloud, thereby reducing the computational amount of further processing the point cloud image subsequently, such as reducing the computational amount of feature extraction for the point cloud image subsequently, and further reducing the computational amount of fruit detection, and improving the fruit detection efficiency.
[0202] Based on any one of the above embodiments, in this method, in the above step 110, a two-dimensional image of the fruit region to be recognized is obtained, and then the following steps are further included:
[0203] Based on the rot region detection model, the rot region of the fruit is detected by detecting the image features of the two-dimensional image, and the fruit rot region is obtained.
[0204] The rot region detection model is trained based on the first sample two-dimensional image labeled with the fruit region and the second sample two-dimensional image labeled with the fruit region and the rot region.
[0205] Here, the two-dimensional image can be subjected to feature extraction multiple times to obtain image features at multiple levels. The embodiment of the present invention performs rot region detection based on the image features of the last level. How to perform feature extraction on the two-dimensional image has been specifically described in the above embodiments, and will not be elaborated here one by one.
[0206] Specifically, based on the rot region detection model, the rot region of the fruit is detected by detecting the image features of the two-dimensional image, so as to segment and obtain the fruit rot region from the two-dimensional image. More specifically, the image features of the two-dimensional image are input into the rot region detection model, and the fruit rot region output by the rot region detection model is obtained.
[0207] In a specific embodiment, the image features of the two-dimensional image are input into the semantic segmentation layer of the rot region detection model, and the fruit rot region output by the semantic segmentation layer is obtained.
[0208] Here, the first sample two-dimensional image is obtained by collecting the sample two-dimensional image and labeling the fruit region of the sample two-dimensional image. The first sample two-dimensional image has 1 fruit region label.
[0209] Here, the second sample two-dimensional image can be obtained by collecting the sample two-dimensional image with fruit rot and labeling the fruit region and the rot region of the sample two-dimensional image; or by selecting the sample two-dimensional image with fruit rot in the first sample two-dimensional image and labeling the rot region of the sample two-dimensional image with fruit rot. The second sample two-dimensional image has 1 fruit region label and 1 rot region label.
[0210] Further, the existing two-dimensional sample images of fruit rot may include fruit images with different degrees of rot, so as to better train the rot area detection model and improve the detection accuracy of the rot area detection model.
[0211] Among them, the sample two-dimensional image may be a fruit image, a fruit tree image, or an image of a larger area, and the embodiments of the present invention do not make specific limitations thereto. The sample two-dimensional image may include one fruit or multiple fruits. The sample two-dimensional image may be an RGB image, an HSV image, a grayscale image, etc., and the embodiments of the present invention do not make specific limitations thereto.
[0212] Among them, the fruit area label may be a point marked in the fruit or the contour points marking the fruit area. The rot area label may be the contour points marking the rot area.
[0213] It should be noted that the first sample two-dimensional image and the second sample two-dimensional image are used to jointly train the rot area detection model, that is, the first sample two-dimensional image and the second sample two-dimensional image are mixed together for model training.
[0214] In the method provided by the embodiments of the present invention, the detection of the spatial position of the fruit and the detection of the rot area of the fruit share the image features of the two-dimensional image, thereby reducing the computational amount of feature extraction, effectively compressing the model scale, reducing the computational amount required for fruit detection, and further improving the efficiency of fruit detection.
[0215] Based on any of the above embodiments, Figure 5 This is the fourth flowchart of the fruit detection method provided by the present invention. As Figure 5 shown, based on the rot area detection model, performing rot area detection on the image features of the two-dimensional image to obtain the fruit rot area, including:
[0216] Step 510, based on the multi-channel feature extraction layer in the rot area detection model, performing multi-channel feature extraction on the image features to obtain channel features of at least two channels.
[0217] Here, multi-channel feature extraction is performed on the image features, so as to split the image features into multiple groups of channel features.
[0218] In a specific embodiment, the multi-channel feature extraction layer is an FCN (Fully Convolutional Networks) layer. The image features are input into the FCN layer to obtain channel features of at least two channels output by the FCN layer. For example, the image features (h1, w1, c1) are input into the FCN layer to obtain K groups of channel features (f1, f2, f3... f k ) of size (h2, w2).
[0219] Step 520: Based on the feature fusion layer in the rotting area detection model, perform weighted fusion on the channel features of the at least two channels to obtain fused image features.
[0220] Specifically, assign a weight to the channel feature of each channel to perform weighted processing on the channel features of each channel to obtain at least two weighted features; then, perform feature fusion on the at least two weighted features to obtain fused image features. The weight corresponding to each channel can be learned during model training.
[0221] Among them, performing feature fusion on at least two weighted features can be to perform average operation processing on at least two weighted features, or can be feature fusion methods such as addition and splicing.
[0222] In a specific embodiment, perform average operation on at least two weighted features to obtain fused image features.
[0223] For example, the average operation can be expressed by the following formula:
[0224]
[0225] In the formula, G is the fused image feature, K is the number of channels of the multi-channel feature extraction layer, α i is the weight corresponding to the i-th channel feature, and f i is the channel feature of the i-th channel.
[0226] Step 530: Based on the semantic segmentation layer in the rotting area detection model, perform rotting area segmentation on the fused image features to obtain the fruit rotting area.
[0227] Here, the semantic segmentation layer can be a pixel-level semantic segmentation layer to perform pixel-level semantic segmentation.
[0228] In a specific embodiment, perform a softmax operation on each point in the fused image features and perform a segmentation operation based on a set threshold to obtain the fruit rotting area.
[0229] The method provided by the embodiment of the present invention extracts multi-channel features from image features, and performs weighted fusion on the channel features of multiple channels to assign different weights to the channel features of different channels to respectively focus on the information that needs to be focused on, so that the finally obtained fused image features are more accurate, and further the fruit rotting area obtained by performing rotting area segmentation based on the fused image features is more accurate, thereby improving the accuracy of fruit rotting area detection.
[0230] Based on any of the above embodiments, after detecting the rotted area of the two-dimensional image using the rotted area detection model to obtain the fruit rotted area, the following steps are further included:
[0231] Obtain the first confidence corresponding to the spatial position and the second confidence corresponding to the fruit rotted area;
[0232] When the first confidence is greater than the first threshold and the second confidence is greater than the second threshold, determine that the fruit is a rotted fruit;
[0233] When the first confidence is less than or equal to the first threshold and it is determined that the second confidence is greater than the second threshold, determine that the fruit exists;
[0234] When the first confidence is less than or equal to the first threshold and it is determined that the second confidence is less than or equal to the second threshold, determine that the fruit does not exist;
[0235] When the first confidence is greater than the first threshold and it is determined that the second confidence is less than or equal to the second threshold, determine that the fruit has no rotted area.
[0236] Here, the first threshold and the second threshold can be set according to actual needs, and the embodiments of the present invention do not make specific limitations on this.
[0237] It should be noted that when both the first confidence and the second confidence are greater than their respective thresholds, it indicates that both fruit detection and rotted area detection are relatively accurate. Therefore, it can be determined that the fruit is a rotted fruit. When the first confidence is not greater than the corresponding first threshold and the second confidence is greater than the second threshold, although the fruit detection is relatively inaccurate, the rotted area detection is relatively accurate. Therefore, it can still be determined that the fruit exists. When the first confidence is not greater than the corresponding first threshold and the second confidence is not greater than the second threshold, it indicates that both fruit detection and rotted area detection are relatively inaccurate. Therefore, it can be determined that the fruit does not exist. When the first confidence is greater than the first threshold and the second confidence is not greater than the second threshold, it indicates that the fruit detection is relatively accurate, but the rotted area detection is relatively inaccurate. Therefore, it can be determined that the fruit exists and has no rotted area.
[0238] The method provided by the embodiments of the present invention comprehensively considers the first confidence corresponding to the spatial position and the second confidence corresponding to the fruit rotted area to determine whether the fruit exists. If it exists, it determines whether the fruit has a rotted area, that is, it can distinguish normal fruits and rotted fruits, thereby improving the accuracy of fruit detection; in addition, it provides support for fruit picking robots to pick and classify fruits.
[0239] Based on any of the above embodiments, after detecting the image features of the two-dimensional image for the rotted area based on the rotted area detection model to obtain the fruit rotted area, the following steps are further included:
[0240] Based on the spatial position, determine the sphere corresponding to the fruit and determine the surface area of the sphere;
[0241] Based on the fruit rotted area, determine the rotted area of the fruit;
[0242] Based on the ratio of the rotted area to the surface area, determine the rotted degree of the fruit.
[0243] Specifically, the spatial position may include the three-dimensional information of the fruit. Thus, based on the spatial position, the sphere corresponding to the fruit can be fitted, and then the surface area of the sphere is calculated; the fruit rotted area may include the set of contour points of the rotted area. Thus, based on the fruit rotted area, the rotted area of the rotted area can be calculated; finally, based on the ratio of the rotted area of the fruit to the surface area of the fruit, the rotted degree of the fruit is determined.
[0244] It should be noted that the larger the ratio of the rotted area to the surface area, the higher the rotted degree; the smaller the ratio of the rotted area to the surface area, the lower the rotted degree. In one embodiment, the rotted degree can be divided into 10 levels at intervals of 0.1.
[0245] In addition, it should also be noted that the embodiments of the present invention only determine the rotted degree of the fruits that have been determined to be rotted.
[0246] The method provided by the embodiments of the present invention further determines the rotted degree of the rotted fruits through the spatial position obtained by fruit detection and the fruit rotted area obtained by rotted area detection, thereby improving the accuracy of fruit detection; in addition, it provides further support for fruit picking robots to pick and classify fruits.
[0247] Based on any of the above embodiments, in this method, the above step 130 includes:
[0248] Based on the fusion feature, perform fruit position detection to obtain the fruit area;
[0249] Based on the fruit area and the depth information of the point cloud image, determine the spatial position of the fruit in the fruit area to be recognized.
[0250] Specifically, based on the fusion feature, perform two-dimensional fruit position detection to obtain the two-dimensional fruit area. More specifically, input the fusion feature into the target detection layer to obtain the fruit area output by the target detection layer.
[0251] Specifically, the fruit region is a 2D detection box. Therefore, the 2D detection box is fused with the depth information of the point cloud image to obtain a 3D detection box, and then the spatial position of the fruit is obtained based on the 3D detection box.
[0252] Furthermore, based on the fruit region, the depth information of the point cloud image, and the rotation angle information of the point cloud image, the spatial pose of the fruit in the fruit region to be recognized is determined. Among them, the spatial pose includes the spatial position and the rotation angle, that is, it includes the length, width, height, and rotation angle.
[0253] It should be noted that since the point cloud image was previously converted into a two-dimensional pseudo-image, after the fruit region is determined, the coordinates of the fruit region need to be converted to the point cloud image to facilitate obtaining the spatial position by combining the depth information of the point cloud image subsequently. The specific conversion formula has been described above and will not be elaborated here.
[0254] The method provided by the embodiments of the present invention, based on the fusion features of the two-dimensional image and the point cloud image, first performs object detection on the fruit to obtain a two-dimensional fruit region, and then determines the spatial position of the fruit based on the fruit region in combination with the depth information of the point cloud image. Compared with directly performing three-dimensional object detection, the embodiments of the present invention can reduce the computational amount of object detection and further improve the efficiency of fruit detection.
[0255] Based on any of the above embodiments, after step 130, the method further includes
[0256] Determining the sphere corresponding to the fruit based on the spatial position;
[0257] Classifying the fruit according to the diameter size of the sphere.
[0258] Specifically, the spatial position may include the 3D information of the fruit, so that the sphere corresponding to the fruit can be fitted based on the spatial position. Then, the diameter of the sphere is calculated, and the fruit is classified according to the diameter size of the fruit.
[0259] It should be noted that the larger the diameter size of the fruit, the higher the size level of the fruit, and the smaller the diameter size of the fruit, the lower the size level of the fruit.
[0260] In addition, it should also be noted that the embodiments of the present invention only classify the size of the fruit when it has been determined that the fruit exists.
[0261] The method provided by the embodiments of the present invention further classifies the size of the fruit through the spatial position obtained by fruit detection, thereby improving the accuracy of fruit detection; in addition, it provides further support for fruit picking and fruit grading by the fruit picking robot.
[0262] Based on any of the above embodiments, a specific embodiment is described herein. First, a two-dimensional image and a point cloud image of the same fruit region to be recognized are obtained; then, feature extraction at different levels is performed on the two-dimensional image, the point cloud image is downsampled, and the downsampled point cloud image is converted into a two-dimensional pseudo-image, and then feature extraction at different levels is performed on the two-dimensional pseudo-image; thereafter, the image features of the two-dimensional image at different levels are cross-fused with the point cloud features of the point cloud image at different levels to obtain cross-features at different levels in the point cloud image. At the same time, feature extraction is performed on the cross-features at each level to obtain the point cloud features at each next level, and the point cloud features at the last level are determined as the fusion features; thereafter, based on the fusion features, fruit two-dimensional position detection is performed to obtain the fruit two-dimensional region, and then the fruit two-dimensional region is combined with the depth information of the point cloud image to obtain the three-dimensional spatial position of the fruit; at the same time, through the rot region detection model, rot region detection is performed on the image features of the two-dimensional image at the last level to obtain the fruit rot region; wherein, in the rot region detection, multi-channel feature extraction is performed on the image, and the multiple channel features after multi-channel feature extraction are weighted and fused to obtain the fused image features, and then semantic segmentation is performed on the fused image features to obtain the fruit rot region; after the spatial position of the fruit and the fruit rot region are both detected, based on the first confidence corresponding to the spatial position, the second confidence corresponding to the fruit rot region, and the thresholds corresponding to each confidence, it is determined whether the fruit exists, and whether the existing fruit is rotten. At the same time, based on the spatial position, a sphere corresponding to the fruit is determined, and then the fruit is sized based on the diameter size of the sphere, and the rot area is determined based on the fruit rot region, so as to determine the rot degree of the fruit based on the ratio of the rot area to the surface area of the sphere.
[0263] Based on any of the above embodiments, Figure 6 is a schematic structural diagram of the fruit detection device provided by the present invention, as Figure 6 shown, the device includes an acquisition module 610, a fusion module 620, and a detection module 630.
[0264] The acquisition module 610 is configured to acquire a two-dimensional image of the fruit region to be recognized, and a point cloud image of the fruit region to be recognized;
[0265] The fusion module 620 is configured to fuse the image features of the two-dimensional image and the point cloud features of the point cloud image to obtain fusion features, and the image features are used to characterize the position information of the fruit in the two-dimensional image;
[0266] The detection module 630 is configured to perform fruit detection based on the fusion features to obtain the spatial position of the fruit in the fruit region to be recognized.
[0267] In the device provided by the embodiment of the present invention, the image features of the two-dimensional image can be used to represent the position information of the fruits in the two-dimensional image. Based on this, by fusing the image features of the two-dimensional image with the point cloud features of the point cloud image, the edge features and surface features of the fruits can be compensated by the image features of the two-dimensional image, so as to enhance the representation ability of the point cloud features, enabling the fusion features for fruit detection to consider the position information of the fruits in the two-dimensional image, thereby ensuring that the fusion features can more completely reflect the position information of the fruits, and further making the spatial position obtained by fruit detection based on the fusion features more accurate, that is, improving the accuracy of fruit detection.
[0268] Based on any of the above embodiments, the fusion module 620 includes:
[0269] A feature acquisition unit, configured to acquire the image features of the two-dimensional image at at least two levels;
[0270] A feature fusion unit, configured to fuse the point cloud features of the point cloud image at the current level with the image features at at least one level among the at least two levels to obtain the cross features at the current level;
[0271] A feature extraction unit, configured to perform feature extraction on the cross features at the current level to obtain the point cloud features at the next level, and use the next level as the current level until the current level is the last level;
[0272] A feature determination unit, configured to determine the fusion features based on the point cloud features at the last level.
[0273] Based on any of the above embodiments, the feature fusion unit is further configured to:
[0274] Determine the fusion weights corresponding to the image features at at least one level among the at least two levels based on the image features at at least one level among the at least two levels;
[0275] Based on the fusion weights corresponding to the image features at at least one level, perform feature enhancement processing on the point cloud features at the current level respectively to obtain the enhanced features corresponding to the image features at at least one level;
[0276] Fuse the enhanced features corresponding to the image features at at least one level with the point cloud features at the current level respectively to obtain the cross fusion features corresponding to the image features at at least one level;
[0277] Fuse the cross fusion features corresponding to the image features at at least one level to obtain the cross features at the current level.
[0278] Based on any of the above embodiments, the fusion module 620 includes:
[0279] The feature acquisition unit is further configured to acquire the image features of the two-dimensional image at one level;
[0280] The feature fusion unit is further configured to fuse the point cloud features of the point cloud image at the current level with the image features at the one level to obtain the cross features at the current level;
[0281] The feature extraction unit is further configured to perform feature extraction on the cross features at the current level to obtain the point cloud features at the next level, and use the next level as the current level until the current level is the last level;
[0282] The feature determination unit is further configured to determine the fusion features based on the point cloud features at the last level.
[0283] Based on any of the above embodiments, the device further includes:
[0284] The image conversion module is configured to convert the point cloud image into a two-dimensional pseudo-image based on the internal parameter matrix and the external parameter matrix of the acquisition device of the point cloud image;
[0285] The feature determination module is configured to determine the point cloud features of the point cloud image at the first level based on the two-dimensional pseudo-image.
[0286] Based on any of the above embodiments, the device further includes:
[0287] The rot detection module is configured to detect the rot area of the image features of the two-dimensional image based on the rot area detection model to obtain the fruit rot area;
[0288] The rot area detection model is trained based on the first sample two-dimensional image labeled with the fruit area and the second sample two-dimensional image labeled with the fruit area and the rot area.
[0289] Based on any of the above embodiments, the rot detection module includes:
[0290] The multi-channel extraction unit is configured to perform multi-channel feature extraction on the image features based on the multi-channel feature extraction layer in the rot area detection model to obtain the channel features of at least two channels;
[0291] The weighted fusion unit is configured to perform weighted fusion on the channel features of the at least two channels based on the feature fusion layer in the rot area detection model to obtain the fused image features;
[0292] The semantic segmentation unit is configured to perform rot area segmentation on the fused image features based on the semantic segmentation layer in the rot area detection model to obtain the fruit rot area.
[0293] Based on any of the above embodiments, the apparatus further includes:
[0294] A confidence level acquisition module, configured to acquire a first confidence level corresponding to the spatial position and a second confidence level corresponding to the fruit rotting area;
[0295] A fruit determination module, configured to determine that the fruit is a rotten fruit when the first confidence level is greater than a first threshold and the second confidence level is greater than a second threshold;
[0296] The fruit determination module is further configured to determine that a fruit exists when the first confidence level is less than or equal to the first threshold and it is determined that the second confidence level is greater than the second threshold;
[0297] The fruit determination module is further configured to determine that no fruit exists when the first confidence level is less than or equal to the first threshold and it is determined that the second confidence level is less than or equal to the second threshold;
[0298] The fruit determination module is further configured to determine that the fruit has no rotting area when the first confidence level is greater than the first threshold and it is determined that the second confidence level is less than or equal to the second threshold.
[0299] Based on any of the above embodiments, the detection module 630 includes:
[0300] An area detection unit, configured to perform fruit position detection based on the fusion feature to obtain a fruit area;
[0301] A position determination unit, configured to determine the spatial position of the fruit in the fruit area to be recognized based on the fruit area and the depth information of the point cloud image.
[0302] Figure 7 An example of a schematic physical structure diagram of an electronic device is shown as Figure 7 shown. The electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communication interface 720, and the memory 730 complete mutual communication through the communication bus 740. The processor 710 may call logical instructions in the memory 730 to execute a fruit detection method, which includes: acquiring a two-dimensional image of the fruit area to be recognized and a point cloud image of the fruit area to be recognized; fusing the image feature of the two-dimensional image and the point cloud feature of the point cloud image to obtain a fusion feature, where the image feature is used to characterize the position information of the fruit in the two-dimensional image; performing fruit detection based on the fusion feature to obtain the spatial position of the fruit in the fruit area to be recognized.
[0303] In addition, when the logical instructions in the above-mentioned memory 730 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, external hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0304] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the fruit detection method provided by the above-mentioned various methods. The method includes: obtaining a two-dimensional image of the fruit region to be recognized, and a point cloud image of the fruit region to be recognized; fusing the image features of the two-dimensional image and the point cloud features of the point cloud image to obtain a fused feature, where the image features are used to characterize the position information of the fruit in the two-dimensional image; performing fruit detection based on the fused feature to obtain the spatial position of the fruit in the fruit region to be recognized.
[0305] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the fruit detection method provided by the above-mentioned various methods. The method includes: obtaining a two-dimensional image of the fruit region to be recognized, and a point cloud image of the fruit region to be recognized; fusing the image features of the two-dimensional image and the point cloud features of the point cloud image to obtain a fused feature, where the image features are used to characterize the position information of the fruit in the two-dimensional image; performing fruit detection based on the fused feature to obtain the spatial position of the fruit in the fruit region to be recognized.
[0306] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0307] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0308] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A fruit detection method, characterized in that, Including: Obtain a two-dimensional image of the fruit area to be recognized, and a point cloud image of the fruit area to be recognized; Obtain image features of the two-dimensional image at at least two levels, where the image features are used to characterize the position information of the fruit in the two-dimensional image; Based on the image features at at least one level among the at least two levels, determine the fusion weights corresponding to the image features at the at least one level; Based on the fusion weights corresponding to the image features at the at least one level, perform feature enhancement processing on the point cloud features of the point cloud image at the current level respectively to obtain the enhanced features corresponding to the image features at the at least one level, so as to enhance the point cloud features of the area where fruits may exist in the point cloud image; Perform feature fusion on the enhanced features corresponding to the image features at the at least one level and the point cloud features at the current level respectively to obtain the cross-fusion features corresponding to the image features at the at least one level; Perform feature fusion on the cross-fusion features corresponding to the image features at the at least one level to obtain the cross features at the current level; Perform feature extraction on the cross features at the current level to obtain the point cloud features at the next level, and use the next level as the current level until the current level is the last level; Based on the point cloud features at the last level, determine the fusion features; Perform fruit detection based on the fusion features to obtain the spatial position of the fruit in the fruit area to be recognized.
2. The fruit detection method according to claim 1, characterized in that The point cloud features of the point cloud image at the first level are determined based on the following steps: Based on the internal parameter matrix and external parameter matrix of the acquisition device of the point cloud image, convert the point cloud image into a two-dimensional pseudo-image; Based on the two-dimensional pseudo-image, determine the point cloud features of the point cloud image at the first level.
3. The fruit detection method according to claim 1, characterized in that After obtaining the two-dimensional image of the fruit area to be recognized, it further includes: Based on a rot detection model, perform rot detection on the image features of the two-dimensional image to obtain the fruit rot area; The rot detection model is trained based on a first sample two-dimensional image marked with a fruit area and a second sample two-dimensional image marked with a fruit area and a rot area.
4. The fruit detection method according to claim 3, wherein Based on the rot detection model, performing rot detection on the image features of the two-dimensional image to obtain the fruit rot area includes: Based on the multi-channel feature extraction layer in the rot detection model, perform multi-channel feature extraction on the image features to obtain channel features of at least two channels; Based on the feature fusion layer in the rot detection model, perform weighted fusion on the channel features of the at least two channels to obtain fused image features; Based on the semantic segmentation layer in the rot detection model, perform rot area segmentation on the fused image features to obtain the fruit rot area.
5. The fruit detection method according to claim 3, wherein, After based on the rot detection model, performing rot detection on the image features of the two-dimensional image to obtain the fruit rot area, it further includes: Obtain a first confidence corresponding to the spatial position, and a second confidence corresponding to the fruit rot area; When the first confidence level is greater than the first threshold and the second confidence level is greater than the second threshold, determine that the fruit is a rotten fruit; When the first confidence level is less than or equal to the first threshold and it is determined that the second confidence level is greater than the second threshold, determine that the fruit exists; When the first confidence level is less than or equal to the first threshold and it is determined that the second confidence level is less than or equal to the second threshold, determine that the fruit does not exist; When the first confidence level is greater than the first threshold and it is determined that the second confidence level is less than or equal to the second threshold, determine that the fruit has no rotten area.
6. The fruit detection method according to claim 1, wherein The fruit detection based on the fusion feature to obtain the spatial position of the fruit in the to-be-identified fruit region includes: Based on the fusion feature, perform fruit position detection to obtain the fruit region; Based on the fruit region and the depth information of the point cloud image, determine the spatial position of the fruit in the to-be-identified fruit region.
7. A fruit detection device, characterized in that, Includes: An acquisition module for acquiring a two-dimensional image of the to-be-identified fruit region and a point cloud image of the to-be-identified fruit region; A fusion module for: Acquire image features of the two-dimensional image at at least two levels, where the image features are used to characterize the position information of the fruit in the two-dimensional image; Based on the image features at at least one level among the at least two levels, determine the fusion weights corresponding to the image features at the at least one level; Based on the fusion weights corresponding to the image features at the at least one level, respectively perform feature enhancement processing on the point cloud features of the point cloud image at the current level to obtain the enhanced features corresponding to the image features at the at least one level, so as to enhance the point cloud features of the area where the fruit may exist in the point cloud image; Fusion the enhanced features corresponding to the image features at the at least one level with the point cloud features at the current level respectively to obtain the cross-fusion features corresponding to the image features at the at least one level; Fusion the cross-fusion features corresponding to the image features at the at least one level to obtain the cross features at the current level; Extract features from the cross features at the current level to obtain the point cloud features at the next level, and use the next level as the current level until the current level is the last level; Based on the point cloud features at the last level, determine the fusion feature; A detection module for performing fruit detection based on the fusion feature to obtain the spatial position of the fruit in the to-be-identified fruit region.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the fruit detection method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the fruit detection method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Tobacco leaf mildew detection system based on artificial intelligence
CN111986167A
Target detection network system and method applied to multi-sensor data fusion in rainy and snowy weather scene
CN114140672A
Citrus recognition and positioning method, device and equipment and storage medium
CN114332689A