A fruit and vegetable grading method and system based on machine vision and deep learning
By using machine vision and deep learning-based methods, multi-view image processing and deep learning models, the problems of inaccurate segmentation and unstable recognition in traditional grading methods under complex environments are solved, thereby improving the accuracy and stability of fruit and vegetable grading and making it suitable for automated fruit and vegetable sorting.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AGRICULTURAL & RURAL AFFAIRS BUREAU OF SHANGYU DISTRICT SHAOXING CITY (RURAL REVITALIZATION BUREAU OF SHANGYU DISTRICT SHAOXING CITY)
- Filing Date
- 2026-03-31
- Publication Date
- 2026-06-26
AI Technical Summary
Traditional machine vision grading methods are easily affected by changes in posture, lighting fluctuations, background noise, and differences in defect scale when faced with complex defects on the surface of fruits and vegetables, resulting in inaccurate segmentation and unstable recognition.
A machine vision and deep learning-based approach is adopted. Multi-view images are acquired, standardized and preprocessed, and then segmented using an arbitrary target segmentation network with skip connections. The fruit and vegetable region features are extracted by combining a deep learning grading model, and the results are summarized from multiple perspectives to generate the final grading result.
It improves the accuracy and stability of fruit and vegetable grading, enabling more precise identification of complex defects and subtle differences, and achieving standardized, intelligent, and large-scale automated sorting of fruits and vegetables.
Smart Images

Figure CN122289778A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image recognition technology, and in particular to a method and system for grading fruits and vegetables based on machine vision and deep learning. Background Technology
[0002] Automated grading of fruits and vegetables is a key step in modern post-harvest processing. Its goal is to achieve rapid, objective, and standardized sorting based on the appearance integrity, surface defects, size, and quality grade of fruits and vegetables.
[0003] Manual grading relies primarily on workers' subjective judgment of the color, size, shape, and surface defects of fruits and vegetables. While simple to implement, it suffers from low efficiency, poor consistency, high labor intensity, and high long-term costs, making it difficult to meet the needs of large-scale sorting. Traditional machine vision grading methods typically complete target detection and grading through color threshold segmentation, edge extraction, texture feature analysis, or morphological processing.
[0004] However, traditional machine vision grading methods can achieve certain results in scenarios with stable lighting, simple backgrounds, and single defect types. But when faced with complex defects such as blemishes, bruises, cracks, and abrasions on the surface of fruits and vegetables, they are easily affected by changes in the posture of fruits and vegetables, lighting fluctuations, background noise, and differences in defect scale, resulting in inaccurate segmentation and unstable recognition. Summary of the Invention
[0005] In view of the shortcomings of the prior art, the purpose of this invention is to provide a fruit and vegetable grading method and system based on machine vision and deep learning, which can solve the technical problems of inaccurate segmentation and unstable recognition when facing complex defects such as fruit and vegetable lesions, bruises, cracks and abrasions on the surface of fruits and vegetables, which are easily affected by changes in fruit and vegetable posture, light fluctuations, background noise and differences in defect scale.
[0006] A first aspect of this invention proposes a fruit and vegetable grading method based on machine vision and deep learning, comprising: S1: Acquire original images of the fruits and vegetables to be graded from multiple perspectives; S2: Perform standardization preprocessing on the original image; S3: The original image after normalization preprocessing is segmented by an arbitrary target segmentation network based on skip connections to obtain the target fruit and vegetable regions; S4: Using a deep learning-based fruit and vegetable grading model, determine the candidate grade results and defect category results of the fruits and vegetables to be graded within the target fruit and vegetable area; S5: Summarize the candidate grade results and defect category results of the same fruit and vegetable from multiple perspectives to generate the final grade result of the fruit and vegetable to be graded.
[0007] In a second aspect, this invention proposes a fruit and vegetable grading system based on machine vision and deep learning, comprising: a processor and a memory; The memory stores programs or instructions that can run on the processor, which, when executed by the processor, implement the steps of the fruit and vegetable grading method based on machine vision and deep learning as described in the first aspect.
[0008] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this embodiment of the invention, an arbitrary target segmentation network based on skip connections is used to extract target fruit and vegetable regions, which can more accurately separate the main body of the fruit and vegetables from the background and avoid irrelevant regions interfering with the grading judgment. By using a deep learning grading model to identify candidate grades and defect categories for the target fruit and vegetable regions, the deep features such as texture, color, contour, and defect distribution on the fruit and vegetable surface can be fully explored, thereby improving the ability to identify complex defects and subtle differences. Finally, the candidate grade results and defect category results of the same fruit and vegetable from multiple perspectives are summarized, which can integrate information from different perspectives for unified judgment, reduce the impact of misjudgment from a single perspective, and make the final grading result more objective, reliable, and stable. The solution of this invention is particularly suitable for automated fruit and vegetable sorting scenarios, which can improve grading efficiency and facilitate standardized, intelligent, and large-scale application. Attached Figure Description
[0009] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.
[0010] Figure 1 This is a flowchart illustrating a fruit and vegetable grading method based on machine vision and deep learning, provided in an embodiment of the present invention.
[0011] Figure 2 This is a schematic diagram of the architecture of a deep learning-based fruit and vegetable grading model provided in an embodiment of the present invention.
[0012] Figure 3 This is a schematic diagram of the structure of a fruit and vegetable grading system based on machine vision and deep learning provided in an embodiment of the present invention. Detailed Implementation
[0013] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0014] The fruit and vegetable grading method based on machine vision and deep learning provided by the present invention will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0015] Reference manual attached Figure 1 The diagram shows a flowchart of a fruit and vegetable grading method based on machine vision and deep learning provided by an embodiment of the present invention.
[0016] This invention provides a fruit and vegetable grading method based on machine vision and deep learning, which may include the following steps: S1: Obtain original images of the fruits and vegetables to be graded from multiple perspectives.
[0017] S2: Perform standardization preprocessing on the original image.
[0018] In one possible implementation, S2 specifically involves performing size normalization, brightness equalization, color normalization, and background suppression on the original image.
[0019] It should be noted that performing size normalization, brightness equalization, color normalization, and background suppression on the original fruit and vegetable images can unify the basic specifications and visual features of the images, eliminate external interference such as differences in shooting equipment, uneven lighting, cluttered backgrounds, and color deviations. This ensures that fruit and vegetable images with different shooting conditions and shapes have a consistent input standard, significantly reducing the learning difficulty of subsequent target segmentation and grading models, improving the model's efficiency in extracting key features of fruits and vegetables and its recognition stability. It also effectively avoids segmentation deviations, defect misjudgments, and grade misclassifications caused by differences in image quality, ensuring the accuracy and consistency of fruit and vegetable grading results. At the same time, it simplifies the model preprocessing logic and improves the overall efficiency of the grading process. Size normalization, brightness equalization, color normalization, and background suppression are all mature existing technologies, and will not be elaborated upon here.
[0020] S3: The standardized preprocessed image is segmented using an arbitrary target segmentation network based on skip connections to obtain the target fruit and vegetable regions.
[0021] In one possible implementation, S3 specifically includes sub-steps S301 to S304: S301: Extract image embedding features from the standardized preprocessed image using an image encoder.
[0022] Optionally, the image encoder employs a multi-layer convolutional network architecture, with each layer consisting of two convolutional layers and one max-pooling layer. The convolutional layers use 3×3 convolutional kernels combined with the ReLU activation function, giving the model excellent non-linear feature processing capabilities. This is followed by max-pooling layers to reduce the dimensionality of the features while fully preserving core feature information. The number of output channels in this encoder increases progressively with each layer, starting at 64 channels and doubling after each pooling operation, up to a maximum of 1024 channels.
[0023] S302: The image embedding features are input to the decoder, and skip connections are used to directly connect the low-level features in the encoder to the corresponding layers in the decoder. The decoder then decodes the features to obtain the final decoded features. in, D l+1 Indicates the first l +1 layer decoding features, D l Indicates the first l Layer decoding features, F dec ( ) represents the convolution operation function in the decoder. U ( ) indicates an upsampling operation. F skip ( ) represents the feature fusion function in a skip connection. E l Indicates the first l Layer coding features.
[0024] It should be noted that the skip connection mechanism can directly transmit the low-level details such as fruit and vegetable edges and textures retained by the encoder to the decoder, avoiding the loss of details caused by upsampling and dimensional compression, accurately restoring the spatial shape and surface defect details of the fruit and vegetable targets, and greatly improving the integrity and accuracy of fruit and vegetable region segmentation.
[0025] Optionally, the decoder is mainly used to restore and reconstruct the spatial information of the high-dimensional features extracted by the encoder. Its structure consists of alternating upsampling layers and convolutional layers. After each upsampling, a 3×3 convolutional layer is connected to complete feature fusion. The convolutional layers in the decoder also use the ReLU activation function, which can achieve efficient non-linear feature transformation. The decoder is symmetrically designed with the encoder, and the number of channels decreases layer by layer, finally outputting a segmentation result with the same size as the input image.
[0026] S303: The final decoded features are sent to the mask decoder to generate a mask for the target fruit and vegetable region.
[0027] Specifically, the mask decoder performs pixel-by-pixel classification and prediction of the feature, directly outputting a pixel-accurate binary mask with the same size as the input image. This mask can clearly distinguish the target fruit / vegetable region from the background region, thus generating a complete and accurate mask for the target fruit / vegetable region, providing a reliable basis for subsequent fruit / vegetable cropping. The application of mask decoders is a very mature existing technology, and will not be elaborated further in this invention.
[0028] S304: Based on the target fruit and vegetable region mask, crop out the target fruit and vegetable region from the standardized preprocessed image.
[0029] Specifically, based on the generated target fruit and vegetable region mask, the pixel information marked as the fruit and vegetable foreground region by the mask is retained by multiplying the mask with the original image pixel by pixel, while invalid pixels in the background region are removed. Then, the image is cropped according to the minimum bounding rectangle of the mask or the outline boundary of the fruit and vegetable. Finally, a pure region containing only the fruit and vegetable target and free from background interference is accurately extracted from the original image, providing a standardized and effective input for subsequent grading and defect detection.
[0030] In this invention, image embedding features are extracted by an image encoder, and skip connection decoding is combined to preserve the details and spatial information of fruits and vegetables. Then, a mask decoder generates an accurate target fruit and vegetable region mask. Finally, the pure fruit and vegetable target region is cropped from the original image. This not only completely eliminates background redundancy interference and fully preserves key detection features such as surface texture, defects, and morphology of fruits and vegetables, but also improves the accuracy and completeness of target region segmentation. This provides clean and standard input data for subsequent fruit and vegetable grading and defect recognition models, effectively avoids model misjudgment caused by irrelevant information, and greatly improves the accuracy and stability of subsequent detection and grading.
[0031] In one possible implementation, after S301 and before S302, S3 further includes: Based on the prior knowledge of fruit and vegetable locations, bounding box cues, point cues, and / or text cues, cue information is constructed, and cue embedding features are extracted from the cue information using a cue encoder.
[0032] The image embedding features and the cue embedding features are input together into the decoder.
[0033] In this invention, a prompt information construction and prompt encoding stage is introduced. The prior information of fruit and vegetable location, bounding box, point or text prompts are transformed into prompt embedding features, which are then fused with image embedding features and input into the decoder. This provides accurate positioning guidance and prior information support for fruit and vegetable target segmentation, allowing the model to focus more on the core area of the fruit and vegetable, effectively avoiding interference from complex backgrounds, fruit and vegetable stacking, and local occlusion, significantly reducing the probability of segmentation misjudgment and missed detection, improving the accuracy, stability and scene adaptability of fruit and vegetable region segmentation, and ensuring that the target mask generated subsequently fits the real outline of the fruit and vegetable more closely.
[0034] In one possible implementation, S302 specifically includes: S3021: Calculate the attention weight of each layer of encoded features during the encoding process.
[0035] Optionally, a standard self-attention mechanism can be used to determine the attention weights of each feature. The standard self-attention mechanism involves constructing three types of vectors—query, key, and value—for each element in the sequence, calculating the relevance weights between that element and all other elements (usually obtained through dot product and Softmax normalization), and then using these weights to perform a weighted sum of the value vectors of all elements, thereby generating a feature representation that includes global contextual information. This mechanism can adaptively focus on the parts most important to the current task, capturing long-distance dependencies and dynamically adjusting the importance of features at different locations. Therefore, it is widely used in Transformers and various deep learning models, and is a very mature existing technology, which will not be elaborated upon in this invention.
[0036] Optionally, in practical applications, to reduce computational complexity and adapt to convolutional neural network structures, this invention innovatively employs convolutional mapping and activation functions to construct lightweight attention weights: in, A l Indicates the first l Attention weights for hierarchical features σ This represents the activation function, and Conv represents the convolution mapping. E l Indicates the first l Layer coding features, θ l Indicates the first l Convolutional layer parameters, Indicates use θ l right E l Perform convolution mapping.
[0037] It should be noted that by constructing lightweight attention weights through convolutional mapping and activation functions, compared to traditional standard self-attention mechanisms, the computational complexity and number of parameters of the model are greatly reduced, avoiding high computational consumption. This perfectly adapts to the native structure of convolutional neural networks without requiring significant modifications to the encoder architecture. Simultaneously, relying on convolutional mapping to effectively capture the local spatial correlations of fruit and vegetable features, and combining this with activation functions to complete weight normalization, the model can dynamically focus on core discriminative features such as fruit and vegetable outlines and defects. Furthermore, while ensuring feature enhancement effects, it makes the model more suitable for deployment on edge devices in agricultural grading scenarios, achieving a dual optimization of computational efficiency and segmentation accuracy.
[0038] S3022: Based on attention weights, skip connections are used to directly connect low-level features in the encoder to the corresponding layers in the decoder, and then the decoder performs decoding to obtain the final decoded features. in, This indicates element-wise multiplication.
[0039] In this invention, attention weights are introduced into the skip connections to perform weighted fusion of encoded features, which can dynamically highlight key features such as fruit and vegetable outlines, defects, and textures, while suppressing background noise and redundant features. This makes the low-level features transmitted by the skip connections more discriminative, effectively improving the decoder's accuracy in restoring the core information of fruits and vegetables, avoiding feature confusion and loss of details, significantly optimizing the quality of decoded features, and making the subsequently generated fruit and vegetable region masks more closely resemble the real shape, thus greatly improving the accuracy of target segmentation and scene robustness.
[0040] S4: Using a deep learning-based fruit and vegetable grading model, determine the candidate grade results and defect category results for the fruits and vegetables to be graded within the target fruit and vegetable area.
[0041] In one possible implementation, the deep learning-based fruit and vegetable grading model specifically includes: a backbone network, a neck network, and a detection head. S4 specifically includes sub-steps S401 to S403: S401: Extract the grading features of the target fruit and vegetable region through the backbone network.
[0042] S402: The hierarchical features are fused through the neck network to obtain fused features.
[0043] S403: Based on the fusion features, the detection head determines the candidate grade results and defect category results of the fruits and vegetables to be graded within the target fruit and vegetable area.
[0044] Reference manual attached Figure 2 The diagram shows a schematic of the architecture of a deep learning-based fruit and vegetable grading model provided by an embodiment of the present invention.
[0045] Figure 2 In this context, Conv represents the convolution module, C2f represents the C2f module, FE represents the feature enhancement module, DF represents the dynamic fusion module, UpSample represents the upsampling module, and C represents the concatenation process.
[0046] The backbone network includes: a first convolutional module, a second convolutional module, a first C2f module, a third convolutional module, a second C2f module, a fourth convolutional module, a first feature enhancement module, a fifth convolutional module, and a second feature enhancement module connected in series.
[0047] Among them, the C2f module is a lightweight feature extraction and fusion module based on the CSPNet idea. It divides the input features into two branches. One branch undergoes nonlinear feature transformation through a residual bottleneck structure, while the other branch directly retains the original features. Finally, the two feature paths are spliced and fused. This is a very mature existing technology, and will not be elaborated on in this invention.
[0048] Furthermore, low-level features are output through the third convolutional module, intermediate-level features are output through the fourth convolutional module, and high-level features are output through the second feature enhancement module.
[0049] It should be noted that by using a hierarchical and progressive feature extraction method, three complementary features—low-level texture edges, intermediate-level local structures, and high-level overall semantics—can be gradually extracted from the original image. This not only enhances feature expression and gradient propagation efficiency with the C2f module, but also focuses on key information such as fruit and vegetable defects and contours and suppresses redundant noise through the feature enhancement module. At the same time, the hierarchical output of multi-scale features provides a rich and complete feature foundation for the subsequent multi-scale feature fusion of the neck network, enabling the model to accurately adapt to defects and appearance features of different sizes and shapes of fruits and vegetables, and significantly improve the accuracy, robustness, and generalization ability of fruit and vegetable grading and defect identification.
[0050] Furthermore, both the first and second feature enhancement modules employ a single-head self-attention mechanism and a gating mechanism. Features input to the feature enhancement module are first divided into channels, and a single-head self-attention mechanism is applied to each channel to aggregate spatial features. The features obtained through the self-attention mechanism are then concatenated with the preserved original channel features, thereby achieving the fusion of global contextual information. The single-head self-attention mechanism can perform global relational modeling of features in the spatial dimension, effectively aggregating contextual features such as the global contour and overall morphology of fruit and vegetable targets, and capturing long-distance dependencies between features.
[0051] The features of the current stage are fused with those of the previous stage through residual connections and finally input into the convolutionally gated linear unit (CLU). The CLU introduces a 3×3 depthwise convolution before the activation function of the gated branch, transforming it into a gated channel attention mechanism based on neighborhood features. The gating mechanism in the CLU dynamically generates a gating signal for each channel, enhancing important features and suppressing redundant features, thereby improving the accuracy of feature selection.
[0052] In this invention, the combination of single-head self-attention and gating mechanism can achieve the complementary advantages of global semantic perception and local feature selection. It can grasp the overall characteristics and spatial correlation of fruits and vegetables through self-attention, and accurately purify local key details and eliminate interference with the help of gating mechanism. While keeping the model lightweight, it makes the feature expression more comprehensive and discriminative, and significantly improves the accuracy, robustness and scene adaptability of fruit and vegetable grading and defect detection.
[0053] Furthermore, the neck network employs a diffusion pyramid network based on feature selection focusing.
[0054] The feature selection-focused diffusion pyramid network takes low-level, intermediate-level, and high-level features from the backbone network as input.
[0055] The feature selection-focused diffusion pyramid network includes high-level fusion branches, intermediate fusion branches, and low-level fusion branches.
[0056] In the intermediate fusion branch, a dynamic fusion module is used to perform feature selection and focusing, taking low-level features, intermediate-level features, and high-level features as inputs, to obtain the feature selection and focusing results.
[0057] The feature selection focusing result is first used as the output of the intermediate fusion branch to obtain the intermediate fusion feature.
[0058] In the advanced fusion branch, the feature selection and focusing results from the intermediate fusion branch are convolved with the advanced features and then concatenated with them through the third C2f module to generate advanced fusion features. This process diffuses detailed features such as local defects and textures upwards into the advanced semantic features, allowing the advanced features, which originally focused on overall semantics, to accurately carry the key details of surface damage to fruits and vegetables, thus preventing the advanced features from losing information about subtle defects.
[0059] In the low-level fusion branch, the feature selection and focusing results from the intermediate fusion branch are upsampled and concatenated with the low-level features, and then generated into low-level fusion features through the fourth C2f module. High-level information such as global contour and target semantics is diffused into the low-level features, so that the low-level detail features that originally focused on edges and textures have a semantic orientation of the whole fruit and vegetable, avoiding the low-level features being messy and meaningless and mistaking the background as defects.
[0060] Through this bidirectional feature diffusion, the fused features of the high, medium, and low branches are no longer isolated, but simultaneously possess global semantics, local structure, and fine defect details. This not only preserves the advantages of features at different scales, but also achieves information complementarity and enhancement, ultimately greatly improving the model's ability to identify defects of different sizes and shapes in fruits and vegetables, as well as the accuracy and robustness of grading.
[0061] Furthermore, the dynamic fusion module aligns features at different levels through convolution and interpolation operations, and then divides the features at each level into multiple equal parts along the channel dimension. Attention is dynamically assigned to each part of the features at different levels, and features to be fused are adaptively selected based on this attention. in, α The fusion weights are represented by sigmoid, and the activation function is represented by sigmoid. ui The first intermediate feature i Each component b i The first low-level feature represents the i Each component h i The first high-level feature i Each component This represents the first intermediate feature after selective fusion. i Each component.
[0062] It should be noted that when the fusion weights α When the fusion weight is greater than 0.5, it tends to retain detailed information. α When the value is less than 0.5, there is a greater tendency to focus on high-level semantic information.
[0063] Furthermore, after the feature fusion components are spliced along the channel dimension, they are processed by convolution, batch normalization (BN), and ReLU activation function to generate a feature selection focusing result that integrates multi-scale information.
[0064] In this invention, the dynamic fusion module first completes the scale alignment of features at different levels through convolution and interpolation operations, and then divides the features into four parts along the channel for fine processing. Relying on the dynamic weight allocation mechanism, it adaptively balances the fusion ratio of detailed information and contextual semantic information, accurately focuses on key discriminative features such as fruit and vegetable defects and contours, and removes redundant information interference. Finally, after optimization by convolution, batch normalization and ReLU activation function, it generates feature selection focusing results that have both multi-scale information and high discriminativeness, which greatly improves the model's ability to identify defects in fruits and vegetables of different shapes and sizes, and effectively enhances the accuracy and scenario robustness of fruit and vegetable grading.
[0065] S5: Summarize the candidate grade results and defect category results of the same fruit and vegetable from multiple perspectives to generate the final grade result of the fruit and vegetable to be graded.
[0066] Optionally, the principle of majority rule can be adopted, and the candidate grade result and defect category result that appear most frequently from multiple perspectives can be used as the final grade result.
[0067] Optionally, based on the level judgment and defect classification data output by the model from each perspective, the frequency of occurrence of each level and defect category can be determined by majority voting statistics. At the same time, conflicting results can be calibrated by combining defect priority rules, abnormal misjudgment results caused by occlusion and angle deviation in a single perspective can be eliminated, high-frequency consistent judgment results can be accepted, and results with differences can be weighted and fused for correction. Finally, by integrating the effective detection information from multiple perspectives, a comprehensive, accurate and stable final classification result can be generated.
[0068] Reference manual attached Figure 3The diagram shows a structural schematic of a fruit and vegetable grading system based on machine vision and deep learning provided by an embodiment of the present invention.
[0069] This invention provides a fruit and vegetable grading system 20 based on machine vision and deep learning, including: a processor 201 and a memory 202; The memory 202 stores programs or instructions that can run on the processor 201. When the program or instructions are executed by the processor 201, they implement the steps of the above-mentioned fruit and vegetable grading method based on machine vision and deep learning, and can achieve the same technical effect. To avoid repetition, the present invention will not elaborate further.
Claims
1. A fruit and vegetable grading method based on machine vision and deep learning, characterized in that, include: S1: Acquire original images of the fruits and vegetables to be graded from multiple perspectives; S2: Perform standardization preprocessing on the original image; S3: The standardized preprocessed image is segmented using an arbitrary target segmentation network based on skip connections to obtain the target fruit and vegetable regions; S4: Using a deep learning-based fruit and vegetable grading model, determine the candidate grade results and defect category results of the fruits and vegetables to be graded within the target fruit and vegetable area; S5: Summarize the candidate grade results and defect category results of the same fruit and vegetable from multiple perspectives to generate the final grade result of the fruit and vegetable to be graded.
2. The fruit and vegetable grading method based on machine vision and deep learning according to claim 1, characterized in that, Specifically, S2 is: The original image is subjected to size normalization, brightness equalization, color normalization, and background suppression processing.
3. The fruit and vegetable grading method based on machine vision and deep learning according to claim 1, characterized in that, S3 specifically includes: S301: Extract image embedding features from the standardized preprocessed image using an image encoder; S302: Input the image embedding features into the decoder, and use skip connections to directly connect the low-level features in the encoder to the corresponding layers in the decoder, and then decode the images to obtain the final decoded features. S303: The final decoded features are sent to the mask decoder to generate a mask for the target fruit and vegetable region; S304: Based on the target fruit and vegetable region mask, crop out the target fruit and vegetable region from the standardized preprocessed image.
4. The fruit and vegetable grading method based on machine vision and deep learning according to claim 3, characterized in that, After S301 and before S302, S3 further includes: Based on the prior knowledge of fruit and vegetable location, bounding box cue, point cue and / or text cue, cue information is constructed, and cue embedding features are extracted from the cue information by the cue encoder; The image embedding features and the cue embedding features are input together into the decoder.
5. The fruit and vegetable grading method based on machine vision and deep learning according to claim 3, characterized in that, Specifically, S302 includes: S3021: Calculate the attention weight of each layer of encoded features during the encoding process; S3022: Based on the attention weights, skip connections are used to directly connect the low-level features in the encoder to the corresponding layers in the decoder, and the decoder is used to decode the features to obtain the final decoded features.
6. The fruit and vegetable grading method based on machine vision and deep learning according to claim 1, characterized in that, The deep learning-based fruit and vegetable grading model specifically includes: a backbone network, a neck network, and a detection head; S4 specifically includes: S401: Extract the grading features of the target fruit and vegetable region through the backbone network; S402: Through the neck network, feature fusion is performed on the hierarchical features to obtain fused features; S403: Based on the fusion features, the detection head determines the candidate grade results and defect category results of the fruits and vegetables to be graded within the target fruit and vegetable area.
7. The fruit and vegetable grading method based on machine vision and deep learning according to claim 6, characterized in that, The backbone network includes: a first convolutional module, a second convolutional module, a first C2f module, a third convolutional module, a second C2f module, a fourth convolutional module, a first feature enhancement module, a fifth convolutional module, and a second feature enhancement module connected in series; The third convolutional module outputs low-level features, the fourth convolutional module outputs intermediate-level features, and the second feature enhancement module outputs high-level features. Both the first feature enhancement module and the second feature enhancement module employ a single-head self-attention mechanism and a gating mechanism. The features input to the feature enhancement module are first divided into channels, and the single-head self-attention mechanism is applied to each channel to aggregate spatial features. The features of the current stage are fused with the features of the previous stage through residual connections and finally input into the convolutional gated linear unit. The gating mechanism in the convolutional gated linear unit enhances important features and suppresses redundant features by dynamically generating a gating signal for each channel.
8. The fruit and vegetable grading method based on machine vision and deep learning according to claim 7, characterized in that, The neck network employs a diffusion pyramid network based on feature selection focusing; The feature selection-focused diffusion pyramid network takes the low-level features, intermediate-level features, and high-level features from the backbone network as input; The feature selection-focused diffusion pyramid network includes a high-level fusion branch, an intermediate fusion branch, and a low-level fusion branch; In the intermediate fusion branch, a dynamic fusion module is used to perform feature selection and focusing, taking the low-level features, the intermediate-level features, and the high-level features as inputs, to obtain the feature selection and focusing result; The feature selection and focusing result is used as the output of the intermediate fusion branch to obtain the intermediate fusion feature; In the advanced fusion branch, the feature selection and focusing result from the intermediate fusion branch is convolved with the advanced features after convolution processing, and then the advanced fusion features are generated by the third C2f module. In the low-level fusion branch, the feature selection and focusing result from the intermediate fusion branch is upsampled and then concatenated with the low-level features, and the low-level fusion features are generated by the fourth C2f module.
9. The fruit and vegetable grading method based on machine vision and deep learning according to claim 8, characterized in that, The dynamic fusion module aligns features at different levels through convolution and interpolation operations, and then divides the features at each level into multiple equal parts along the channel dimension; it dynamically allocates attention to each part of the features at different levels, and adaptively selects the features to be fused based on the attention to perform feature fusion. After the feature fusion components are concatenated along the channel dimension, they are processed by convolution, batch normalization, and ReLU activation function to generate the feature selection focusing result that integrates multi-scale information.
10. A fruit and vegetable grading system based on machine vision and deep learning, characterized in that, include: Processor and memory; The memory stores programs or instructions that can run on the processor, which, when executed by the processor, implement the steps of the fruit and vegetable grading method based on machine vision and deep learning as described in any one of claims 1 to 9.