Monocular vision-based grape fruit diameter and particle weight measurement model construction and detection method
The monocular vision-based grape diameter and weight measurement model solves the problem of low manual inspection efficiency, realizes the automated non-destructive measurement of grape diameter and weight, improves the inspection efficiency and accuracy, and is suitable for quality sorting in the grape industry.
Patent Information
- Application Number
- CN202510565611.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-09-23
AI Technical Summary
In existing technologies, the measurement of grape diameter and weight relies on manual inspection, which is inefficient and difficult to meet the quality sorting requirements during the fresh picking period. In addition, the monocular vision non-destructive inspection method is difficult to promote in the grape industry.
A monocular vision-based grape diameter and weight measurement model is adopted to achieve non-destructive and automatic measurement of fruit diameter and weight through data set preprocessing, feature extraction, network fusion deep model training and detection methods.
It improves the efficiency of grape testing, reduces labor and time costs, and achieves accurate sorting of grape quality, making it suitable for rapid non-destructive testing in large-scale production.
Smart Images

Figure CN120689862A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of grape non-destructive detection for agricultural fruit quality sorting, and in particular to a monocular vision-based grape diameter and weight measurement model construction and detection method. Background Art
[0002] The diameter and weight of a single grape are important indicators for evaluating grape quality. Currently, traditional grape quality sorting still relies on manual visual screening, which is inefficient. During the grape season, measuring the diameter and weight of each grape consumes a lot of manpower, material resources, and financial resources, and the manual efficiency cannot meet the quality sorting needs of freshly picked grapes. Therefore, the present invention can improve grape inspection efficiency and achieve accurate grape quality sorting by replacing manual inspection and screening with machine vision. Although there are many precision measurement methods and designs for small-scale objects such as workpieces in the field of monocular vision measurement, there are relatively few monocular vision non-destructive inspection methods and research for the diameter and weight of each grape in three-dimensional position space, making it difficult to promote single grape visual analysis technology in the industry. Summary of the Invention
[0003] In view of this, in order to solve the above-mentioned problems in the prior art, the present invention proposes a grape diameter and weight measurement model construction and detection method based on monocular vision, which can accurately and non-destructively detect the diameter and weight of each grape, and realize the automated measurement of grape diameter.
[0004] The present invention solves the above problems through the following technical means:
[0005] In one aspect, the present invention proposes a method for constructing a grape diameter and weight measurement model based on monocular vision, comprising the following steps:
[0006] S1. Obtain grape dataset images as samples and measure the diameter and weight of each grape in the sample.
[0007] S2. Use the preset data processing method to preprocess the grape dataset image and extract the single grape mask to obtain the enhanced processed image;
[0008] S3. Extracting features of individual grapes based on the enhanced image, wherein the features of individual grapes include the color of the individual grapes, the position of the center of the individual grapes in the image, the proportion of the individual grapes in the overall image, the diameter of the inscribed circle of the individual grapes, and the normalized diameter of the inscribed circle within the bunch;
[0009] S4. Using the obtained enhanced processed images, individual grape features, and the grape diameter and weight of each grape, a training set, a validation set, and a prediction set based on monocular vision grape diameter and weight are constructed;
[0010] S5. Constructing a network fusion deep model, wherein the network fusion deep model combines high-level features with low-level features of the basic visual algorithm to simultaneously complete the prediction of fruit diameter and grain weight;
[0011] S6. Use the mean squared error loss function to learn the fruit diameter and particle weight prediction tasks, add a dual-task attention distribution similarity constraint, and combine the mean squared error loss function and the constraint as the final model loss function;
[0012] S7. Input the training set into the network fusion deep model, and use the model loss function as the objective function of supervised learning. After each round of training, input the validation set into the model, and use the model loss function to obtain the validation set error. When the validation set error converges to the set threshold or reaches the maximum number of training steps, the network fusion deep model training is terminated to obtain a depth measurement model that integrates fruit diameter and particle weight.
[0013] Preferably, the method further comprises:
[0014] S8. Input the prediction set into the trained network fusion deep model, calculate the gap between the predicted data and the measured data to verify the accuracy of the model, and thus evaluate whether the model is suitable for grape fruit diameter and weight detection.
[0015] Preferably, the network fusion depth model in step S5 is a Transformer model.
[0016] Preferably, the Transformer model includes a fully connected layer, a batch normalization layer, a nonlinear operation layer, an embedding layer, a position encoder, a self-attention layer, a multi-head attention layer, a normalization layer, a residual connection layer and an output layer.
[0017] Preferably, the final model loss function in step S6 is:
[0018]
[0019] Among them, N is the total number of grape bunches, C is the number of grapes in each bunch, Grain weight label, f θ (x i ) t is the predicted value of grape weight for the t-th grape sample in the i-th bunch; Fruit diameter label, g θ (x i ) t is the predicted value of the grape diameter of the t-th grape sample in the i-th bunch, β is the constraint weight, JS is the Jensen-Shannon divergence, α 果径 is the attention distribution of the fruit path task, α 颗重 It is a heavy task attention distribution.
[0020] Preferably, step S7 includes:
[0021] S71, initializing the network fusion depth model;
[0022] S72, using mini-batch stochastic gradient descent for training, AdamW optimizer, learning rate of 0.0001, and weight decay of 0.0001;
[0023] S73, randomly collecting a sample subset from the training set and optimizing the loss function until the number of training steps reaches a first set threshold;
[0024] S74. Randomly collect a sample subset from the validation set to optimize the loss function until the number of training steps reaches a second set threshold.
[0025] Preferably, step S74 further includes:
[0026] A sample subset is randomly collected from the validation set to optimize the loss function until the mean square error between the quantitative analysis prediction value and the label value of the validation set converges to less than a third set threshold.
[0027] In another aspect, the present invention provides a method for detecting grape diameter and weight based on monocular vision, comprising the following steps:
[0028] A1. Obtain sample images of grapes for prediction;
[0029] A2. Smooth and reduce noise on the grape prediction sample image, filter the background and non-grape objects, and determine whether there are grapes in the foreground;
[0030] A3. Segment the whole bunch of grapes based on edge contours to obtain a segmented image of the whole bunch of grapes without background.
[0031] A4. Segment each grape in the whole grape segmentation image using the deep segmentation model and output a single grape segmentation mask.
[0032] A5. Extract individual grape features based on the grape bunch segmentation image and mask. The individual grape features include the grape's color, the location of the grape's center in the image, the proportion of the grape in the overall image, the diameter of the grape's inscribed circle, and the normalized diameter of the inscribed circle within the bunch.
[0033] A6. Input the whole grape bunch segmentation map and the individual grape features into the depth measurement model and predict the diameter and weight of the corresponding individual grapes.
[0034] Preferably, in order to eliminate the interference of non-measured targets, it is necessary to first convert the image from RGB space to HSV color space and filter the non-measured grape colors;
[0035] (H low ,S low ,V low )≤(h,s,v)≤(H up ,S up ,V up )
[0036] Among them, H low ,S low ,V low is the lower threshold of the HSV color space; H up ,S up ,V up is the upper threshold of the HSV color space; and h, s, v are the pixel-level HSV values. To further eliminate the influence of non-measurement targets, the edge contour of the whole bunch of grapes is segmented using the deep segmentation model YOLO to obtain a segmentation map of the whole bunch of grapes without background.
[0037] Preferably, in step A5, based on the segmentation map of the entire grape bunch and the mask, the respective means of the three channels in the effective pixels of the single grape image multiplied by the mask are calculated:
[0038]
[0039] Among them, R', G', B' are the mean values of the three channels, R i ,G i ,B i is the three-channel red, green and blue value of the effective pixel, and N is the number of effective pixels;
[0040] Extract the coordinate position (x, y) of the center of a single grape in the image, as well as the proportion of the single grape in the image;
[0041]
[0042] Where h and w are the length and width of the grape mask bounding box, and H and W are the length and width of the image;
[0043] For a single grape, the inscribed circle of the grape is extracted by Circle Hough Transform, and its inscribed circle diameter D is obtained based on the inscribed circle extraction;
[0044] Normalize the diameters of all inscribed circles in the whole bunch of grapes;
[0045]
[0046] Where D is the diameter of the inscribed circle of a single grape, D max is the diameter of the largest inscribed circle of a single grape in the whole bunch, D normalized is the ratio of the normalized diameters of a single grape.
[0047] Compared with the prior art, the beneficial effects of the present invention include at least:
[0048] The present invention adopts deep learning technology to build models and select model parameters, which can improve the accuracy and robustness of the model; the diameter and weight of each grape can be quickly and losslessly obtained, which greatly reduces the labor and time costs of measuring each grape in season, and ultimately achieves accurate sorting of grape quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0050] Figure 1 A flowchart of a method for constructing a grape diameter and weight measurement model based on monocular vision provided by an embodiment of the present invention;
[0051] Figure 2 A diagram of a grape diameter and weight measurement model based on monocular vision provided by an embodiment of the present invention;
[0052] Figure 3 A flowchart of a parameter adjustment model for measuring grape diameter and weight based on monocular vision provided by an embodiment of the present invention;
[0053] Figure 4 A flowchart of a method for detecting grape diameter and weight based on monocular vision provided by an embodiment of the present invention;
[0054] Figure 5 Schematic diagram of the true value and predicted result of the fruit diameter of a single grape in an embodiment of the present invention;
[0055] Figure 6 Schematic diagram of the true value and predicted result of the weight of a single grape in an embodiment of the present invention. DETAILED DESCRIPTION
[0056] To make the above-mentioned objectives, features, and advantages of the present invention more clearly understood, the technical solutions of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are also within the scope of protection of the present invention.
[0057] In the embodiments of the present application, the term "and / or" is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist at the same time, and B exists alone.
[0058] The terms "first" and "second" in the embodiments of the present application are only used for descriptive purposes and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of the present application, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a system, product or device comprising a series of components or units is not limited to the listed components or units, but may optionally also include components or units that are not listed, or may optionally also include other components or units that are inherent to these products or devices. In the description of the present application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0059] Example 1
[0060] like Figure 1 The figure shows a flow chart of a method for constructing a grape diameter and weight measurement model based on monocular vision, which includes the following steps:
[0061] Step S1: Obtain a grape dataset. The sample size should be sufficient. Samples can be collected before and after branch pruning, before and after pruning bad fruit, and after pruning bad fruit, covering multiple stages of the process, such as after picking and before sorting. The total number of grape bunches should be greater than 800. In this example, 1,262 bunches, totaling 21,892 grapes, were selected to form a single grape sample set.
[0062] Step S2: Smoothing and noise reduction processing are performed on the acquired image, and background and non-grape objects are filtered out.
[0063] Step S3: Extract the single grape image mask, the average of the three RGB channels of the single grape image, the coordinate position of the center of the single grape in the image, the proportion of the single grape in the image, the diameter of the inscribed circle, the normalized diameter of the inscribed circle within the bunch, and other features from the processed image.
[0064] Step S4: Combine the image, mask, features, true fruit diameter, and true fruit weight into a dataset, and randomly divide the dataset into training, validation, and prediction sets, with 80% used as the training set, 10% as the validation set, and the final 10% as the prediction set. A total of 1,262 grape bunches were collected, 1,000 of which were used as the training set for training the deep learning model, 131 as the validation set for adjusting model hyperparameters, and the remaining 131 as the prediction set for evaluating model performance.
[0065] Step S5: Figure 2 As shown in the figure, a network-fused deep model for fruit diameter and particle weight is constructed based on the dataset. It consists of a fully connected layer, a batch normalization layer, a nonlinear operation layer, an embedding layer, a position encoder, a self-attention layer, a multi-head attention layer, a normalization layer, a residual connection layer, and an output layer. Compared to convolutional neural networks, which implicitly acquire local position information through sliding convolution kernels and have difficulty modeling long-range dependencies, the position encoding enables the Transformer to explicitly process object size information.
[0066] The mask is multiplied by the image, and the 224x224 image is partitioned into 192 blocks using a 16x16 window. This is then fed into the pre-trained Vision Transformer network. Each self-attention module output is then connected to the extracted features, fusing the multi-scale deep features of a single grape with low-level features. Furthermore, through the attention mechanism of the Transformer decoder for fruit diameter and grape weight, both task-shared and task-independent multi-scale deep and low-level features are extracted, ultimately completing the fruit diameter and grape weight predictions.
[0067] Step S6: Construct a model loss function for training the Transformer network model. The loss function of the Transformer network adopts the mean square error loss function for learning the fruit diameter and particle weight prediction tasks, and adds a dual-task attention distribution similarity constraint. The above loss function and constraints are combined as the final model loss function. The model loss function is:
[0068]
[0069] Among them, N is the total number of grape bunches, C is the number of grapes in each bunch, Grain weight label, f θ (x i ) t is the predicted value of grape weight for the t-th grape sample in the i-th bunch; Fruit diameter label, g θ (x i ) t is the predicted value of the grape diameter of the t-th grape sample in the i-th bunch, β is the constraint weight, JS is the Jensen-Shannon divergence, α 果径 is the attention distribution of the fruit path task, α 颗重 It is a heavy task attention distribution.
[0070] S7. Input the training set into the network fusion deep model, and use the model loss function as the objective function of supervised learning. After each round of training, input the validation set into the model, and use the model loss function to obtain the validation set error. When the validation set error converges to the set threshold or reaches the maximum number of training steps, the network fusion deep model training is terminated to obtain a depth measurement model that integrates fruit diameter and particle weight.
[0071] like Figure 3 FIG. 1 is a flow chart of a method for training a Transformer network model according to an embodiment of the present invention, which specifically includes the following steps:
[0072] S71, initialize the Transformer network model;
[0073] S72, use the mini-batch stochastic gradient descent method to train the discriminator and generator, use the AdamW optimizer, the learning rate is 0.0001, and the weight decay is 0.0001;
[0074] S73, from the training set D tr Randomly collect sample subsets to optimize the loss function until the number of training steps reaches 100,000;
[0075] S74, from the validation set D val Randomly collect sample subsets to optimize the loss function until the number of training steps reaches a certain set value or until the mean square error loss between the quantitative analysis prediction value and the label value of the validation set converges to less than a certain set value.
[0076] S8. Use the trained Transformer model to input the prediction set into the model to obtain the regression results of fruit diameter and grain weight.
[0077] In this embodiment, after obtaining a grape diameter and weight detection model, the detection model must be verified in real time. The model, built based on the training samples, is used to predict the diameter and weight of newly collected grape samples. The difference between the predicted data and the measured data is calculated to verify the model's accuracy and assess its suitability for grape diameter and weight detection.
[0078] Example 2
[0079] like Figure 4 As shown, the present invention provides a method for detecting grape diameter and weight based on monocular vision, which specifically includes the following steps:
[0080] Step A1: Obtain a grape prediction sample and photograph the sample at a fixed angle and working distance using a monocular industrial vision camera;
[0081] Step A2: Smoothing and denoising the acquired image using Gaussian filtering, segmenting the foreground and background using the YOLO depth segmentation model, filtering out background and non-grape objects, and ultimately determining whether grapes exist in the foreground.
[0082] Step A3: First, consider the impact of non-measurement objects such as grape defects and foreign matter on grape diameter and weight. To eliminate interference from non-measurement objects, convert the image from RGB to HSV color space and filter the non-measurement grape colors.
[0083] (H low ,S low ,V low )≤(h,s,v)≤(H up ,S up ,V up )
[0084] Among them, H low ,S low ,V low is the lower threshold of the HSV color space, H up ,S up ,V up is the upper threshold of the HSV color space; and h, s, and v are pixel-level HSV values. Secondly, to further eliminate the influence of non-measured objects, the edge contours of the entire grape bunch are segmented using the deep segmentation model YOLO to obtain a segmentation map of the entire grape bunch without background.
[0085] Step A4: After obtaining the background-free segmentation image of the entire grape bunch, segment each grape in the image using the YOLO deep segmentation model fine-tuned for individual grapes, outputting a segmentation mask for each grape. Because fine-tuning YOLO is a conventional algorithm, it will not be described in detail here.
[0086] Step A5: Based on the image and the mask, calculate the mean of the three channels in the valid pixels of the single grape image multiplied by the mask;
[0087]
[0088] Among them, R', G', B' are the mean values of the three channels, R i ,G i ,B i It is the three-channel red, green and blue value of the effective pixel, and N is the total number of effective pixels.
[0089] Furthermore, the coordinate position (x, y) of the center of a single grape in the image and the proportion of the single grape in the image are extracted;
[0090]
[0091] Where h and w are the length and width of the grape mask bounding box, and H and W are the length and width of the image.
[0092] Furthermore, the inscribed circle of a single grape is extracted through Circle Hough Transform, and its inscribed circle diameter D is obtained based on the inscribed circle extraction.
[0093] Furthermore, all the inscribed circle diameters in the whole bunch of grapes were normalized;
[0094]
[0095] Where D is the diameter of the inscribed circle of a single grape, D max is the diameter of the largest inscribed circle of a single grape in the whole bunch, D normalized is the ratio of the normalized diameters of a single grape.
[0096] Step A6: Input the image and the features into the depth measurement model in Example 1 and predict the diameter and weight of the corresponding single grape.
[0097] Figure 5 Schematic diagram of the true value and predicted result of the fruit diameter of a single grape in an embodiment of the present invention; Figure 6 The figure below shows the true and predicted weight of a single grape in an embodiment of the present invention. As can be seen from the figures, the monocular vision and deep learning-based method achieves simultaneous detection of the diameter and weight of individual grapes in an entire bunch with high accuracy. In summary, the grape diameter and weight detection model constructed using the present invention can be used to detect grape samples of unknown size and weight with high accuracy, meeting the demand for rapid, non-destructive testing of grape diameter and weight in large-scale production.
[0098] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0099] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for constructing a grape diameter and weight measurement model based on monocular vision, characterized in that: The steps include: S1. Obtain grape dataset images as samples and measure the diameter and weight of each grape in the sample. S2. Use the preset data processing method to preprocess the grape dataset image and extract the single grape mask to obtain the enhanced processed image; S3. Extracting features of individual grapes based on the enhanced image, wherein the features of individual grapes include the color of the individual grapes, the position of the center of the individual grapes in the image, the proportion of the individual grapes in the overall image, the diameter of the inscribed circle of the individual grapes, and the normalized diameter of the inscribed circle within the bunch; S4. Using the obtained enhanced processed images, individual grape features, and the grape diameter and weight of each grape, a training set, a validation set, and a prediction set based on monocular vision grape diameter and weight are constructed; S5. Constructing a network fusion deep model, wherein the network fusion deep model combines high-level features with low-level features of the basic visual algorithm to simultaneously complete the prediction of fruit diameter and grain weight; S6. Use the mean squared error loss function to learn the fruit diameter and particle weight prediction tasks, add a dual-task attention distribution similarity constraint, and combine the mean squared error loss function and the constraint as the final model loss function; S7. Input the training set into the network fusion deep model, and use the model loss function as the objective function of supervised learning. After each round of training, input the validation set into the model, and use the model loss function to obtain the validation set error. When the validation set error converges to the set threshold or reaches the maximum number of training steps, the network fusion deep model training is terminated to obtain a depth measurement model that integrates fruit diameter and particle weight.
2. The method for constructing a grape diameter and weight measurement model based on monocular vision according to claim 1, characterized in that: The method further comprises: S8. Input the prediction set into the trained network fusion deep model, calculate the gap between the predicted data and the measured data to verify the accuracy of the model, and thus evaluate whether the model is suitable for grape fruit diameter and weight detection.
3. The method for constructing a grape diameter and weight measurement model based on monocular vision according to claim 1, characterized in that: In step S5, the network fusion deep model is a Transformer model.
4. The method for constructing a grape diameter and weight measurement model based on monocular vision according to claim 3, characterized in that: The Transformer model includes a fully connected layer, a batch normalization layer, a nonlinear operation layer, an embedding layer, a position encoder, a self-attention layer, a multi-head attention layer, a normalization layer, a residual connection layer, and an output layer.
5. The method for constructing a grape diameter and weight measurement model based on monocular vision according to claim 1, characterized in that: The final model loss function in step S6 is: Among them, N is the total number of grape bunches, C is the number of grapes in each bunch, Grain weight label, f θ (x i ) t is the predicted value of grape weight for the t-th grape sample in the i-th bunch; Fruit diameter label, g θ (x i ) t is the predicted value of the grape diameter of the t-th grape sample in the i-th bunch, β is the constraint weight, JS is the Jensen-Shannon divergence, α 果径 is the attention distribution of the fruit path task, α 颗重 It is a heavy task attention distribution.
6. The method for constructing a grape diameter and grape weight measurement model based on monocular vision according to claim 1, characterized in that: Step S7 includes: S71, initializing the network fusion depth model; S72, using mini-batch stochastic gradient descent for training, AdamW optimizer, learning rate of 0.0001, and weight decay of 0.0001; S73, randomly collecting a sample subset from the training set and optimizing the loss function until the number of training steps reaches a first set threshold; S74. Randomly collect a sample subset from the validation set to optimize the loss function until the number of training steps reaches a second set threshold.
7. The method for constructing a grape diameter and weight measurement model based on monocular vision according to claim 6, characterized in that: Step S74 further includes: A sample subset is randomly collected from the validation set to optimize the loss function until the mean square error between the quantitative analysis prediction value and the label value of the validation set converges to less than a third set threshold.
8. A method for detecting grape diameter and weight based on monocular vision, characterized in that: The steps include: A1. Obtain sample images of grapes for prediction; A2. Smooth and reduce noise on the grape prediction sample image, filter the background and non-grape objects, and determine whether there are grapes in the foreground; A3. Segment the whole bunch of grapes based on edge contours to obtain a segmented image of the whole bunch of grapes without background. A4. Segment each grape in the whole grape segmentation image using the deep segmentation model and output a single grape segmentation mask. A5. Extract individual grape features based on the grape bunch segmentation image and mask. The individual grape features include the grape's color, the location of the grape's center in the image, the proportion of the grape in the overall image, the diameter of the grape's inscribed circle, and the normalized diameter of the inscribed circle within the bunch. A6. Input the whole grape bunch segmentation image and the individual grape features into the depth measurement model described in any one of claims 1-7 and predict the diameter and weight of the corresponding individual grapes.
9. The method for detecting grape diameter and weight based on monocular vision according to claim 8, characterized in that: In step A3, to eliminate interference from non-measured objects, it is necessary to first convert the image from RGB space to HSV color space and filter the non-measured grape colors; (H low ,S low ,V low )≤(h,s,v)≤(H up ,S up ,V up ) Among them, H low ,S low ,V low is the lower threshold of the HSV color space; H up ,S up ,V up is the upper threshold of the HSV color space; and h, s, v are the pixel-level HSV values. To further eliminate the influence of non-measurement targets, the edge contour of the whole bunch of grapes is segmented using the deep segmentation model YOLO to obtain a segmentation map of the whole bunch of grapes without background.
10. The method for detecting grape diameter and weight based on monocular vision according to claim 8, characterized in that: In step A5, based on the grape bunch segmentation map and mask, the three channels of the valid pixels in the single grape image multiplied by the mask are averaged: Among them, R', G', B' are the mean values of the three channels, R i ,G i ,B i is the three-channel red, green and blue value of the effective pixel, and N is the number of effective pixels; Extract the coordinate position (x, y) of the center of a single grape in the image, as well as the proportion of the single grape in the image; Where h and w are the length and width of the grape mask bounding box, and H and W are the length and width of the image; For a single grape, the inscribed circle of the grape is extracted by Circle Hough Transform, and its inscribed circle diameter D is obtained based on the inscribed circle extraction; Normalize the diameters of all inscribed circles in the whole bunch of grapes; Where D is the diameter of the inscribed circle of a single grape, D max is the diameter of the largest inscribed circle of a single grape in the whole bunch, D normalized is the ratio of the normalized diameters of a single grape.