Apple Defect Classification Method Combining Sparse Channel Recognition and Neighborhood Enhancement Classification

By integrating the sparse channel individual recognition model and linear neighborhood enhancement classification model, the problems of rough classification of Apple surface defects and high computing resources in the prior art are solved, and the rapid and accurate identification of Apple surface defects and the classification of multiple types of defects are achieved.

CN116385782BActive Publication Date: 2025-06-27ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310345379.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-03
Publication Date
2025-06-27
Estimated Expiration
2043-04-03

AI Technical Summary

Technical Problem

The existing Apple surface defect recognition method has the problem of being too rough for defect classification and consumes high computing resources in deep learning models.

Method used

A method of apple surface defect classification that integrates sparse channel individual recognition model and linear neighborhood enhancement classification model is proposed. Through two processing modules of apple single fruit positioning and surface defect detection, Apple image segmentation and rapid detection of surface defects are realized.

Benefits of technology

It realizes rapid adaptive extraction of Apple and precise classification of multiple types of defects, reducing the consumption of computing resources and improving identification efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116385782B_ABST
    Figure CN116385782B_ABST
Patent Text Reader

Abstract

The present invention discloses an apple defect classification method that fuses sparse channel recognition and neighborhood enhancement classification. The present invention includes two processing modules: apple single fruit localization and surface defect detection. By establishing a sparse channel individual recognition model and a linear neighborhood enhancement apple defect classification model, the present invention realizes the fast segmentation of single fruits in an image containing multiple apples and the accurate classification of various types of defects. The present invention uses a dual-model chain structure asynchronous defect classification method for apple defect classification in multi-fruit images, improving the recognition efficiency and stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an apple defect classification method in the field of apple surface defect detection, and more particularly to an apple surface defect classification method that combines a sparse channel individual recognition model and a linear neighborhood enhancement classification model. Background Art

[0002] The surface defects of apples are important bases for evaluating apple quality and quality grading (GB / T 10651-2008 Fresh Apples). Compared with the time-consuming and laborious manual recognition methods that rely on the experience of fruit farmers and experts, the use of machine vision technology to detect apple surface defects through deep learning object detection methods has the advantages of being fast, accurate, and non-destructive, and has important practical application value.

[0003] Common methods for identifying apple surface defects using a machine vision system based on color images mainly include: Fan et al. (2010) (Fan S, Li J, Zhang Y, et al. On line detection of defective apples using computer vision system combined with deep learning methods[J]. Journal of Food Engineering, 2020, 286:110102.) segmented apple images using a multi-threshold method and judged whether there were defects in the apples by the positional relationship between the fruit stalk and the calyx and the number of regions of interest; Xin et al. (2021) (Xin Y, Ma S, Wei Y, et al. Detection of Apple Surface Defect Based on YOLOv3[C] / / . 2021 ASABE Annual International Virtual Meeting, 2021:1.) used YOLOv3 to identify the positions of apple surface defects; Wang et al. (2022) (Wang Z, Jin L, Wang S, et al. Apple stem / calyx real-time recognition using YOLO-v5 algorithm for fruit automatic loading system[J]. Postharvest Biology and Technology, 2022, 185:111808.) recognized the fruit stalk and calyx of apples based on an improved YOLOv5 object detection algorithm. Defective apples not only affect the product value from the perspective of quality, but apples with major defects such as rot will further infect the surrounding normal fruits, causing more serious losses. Different types of defects will have different impacts on the value of apples. Therefore, accurately identifying the categories of apple defects is of great significance.

[0004] Meanwhile, in the processing and sorting stage, to ensure that defects on multiple surfaces of apples can be detected completely, the images of apples need to be sampled multiple times at different working positions. However, the sampling positions for shooting are prone to uncertainty due to reasons such as unstable machine running speed, fruit body rotation, and fruit body rolling. In recent years, some scholars have used object detection algorithms to locate and identify multiple fruits in images and provide their position and category information: Zhang et al. (2022) (Zhang X, Xun Y, Chen Y. Automated identification of citrus diseases in orchards using deep learning[J]. Biosystems Engineering, 2022, 223: 249 - 258.) located and identified normal and damaged citrus fruits based on object detection and the SORT algorithm tracking algorithm.

[0005] The method for identifying apple surface defects based on a machine vision system has made some progress, but the classification of apple surface defects is too rough. In addition, existing deep learning models require a large amount of computing resources, etc. Summary of the Invention

[0006] To address the requirements and problems in the background technology, the present invention proposes an apple surface defect classification method that combines a sparse channel individual recognition model and a linear neighborhood enhancement classification model, including two processing modules: apple single fruit localization and surface defect detection, which respectively implement apple image segmentation and rapid detection of surface defects. The apple individual recognition speed is fast, and it can classify various types of defects such as apple rot, mechanical damage, fruit rust cracks, shrinkage and softening, extrusion bruising, and other damages.

[0007] The technical solution adopted by the present invention is as follows:

[0008] Step 1: Train a sparse channel individual recognition model according to the apple single fruit detection data set to obtain a trained sparse channel individual recognition model. The trained sparse channel individual recognition model outputs the corresponding original data set for apple single fruit surface defect detection.

[0009] Step 2: Train a single-fruit illumination correction model based on semantic segmentation using the apple semantic segmentation dataset corresponding to the apple single-fruit detection dataset to obtain a trained single-fruit illumination correction model based on semantic segmentation. Specifically, the single-fruit illumination correction model based on semantic segmentation consists of FastUnet semantic segmentation and adaptive fast illumination correction. The training set of the FastUnet model is the apple semantic segmentation dataset. The adaptive fast illumination correction uses the mask image generated by FastUnet to limit the correction area, and constructs a light field for the correction area using the least squares regression method to complete the illumination correction. After inputting the original dataset of apple single-fruit surface defect detection into the trained single-fruit illumination correction model based on semantic segmentation, the trained single-fruit illumination correction model based on semantic segmentation outputs the illumination correction dataset of apple single-fruit surface defect detection;

[0010] Step 3: Construct an enhanced dataset for apple single-fruit surface defect detection from the original dataset of apple single-fruit surface defect detection and the illumination correction dataset of apple single-fruit surface defect detection, and train a linear neighborhood enhanced apple defect classification model based on the enhanced dataset for apple single-fruit surface defect detection to obtain a trained linear neighborhood enhanced apple defect classification model;

[0011] Step 4: Input the multi-fruit image to be predicted into the trained sparse channel individual recognition model, and output the single-fruit images corresponding to each apple in the current multi-fruit image;

[0012] Step 5: Input all the single-fruit images corresponding to the multi-fruit image into the trained linear neighborhood enhanced apple defect classification model, and predict the defects and bounding boxes corresponding to each single-fruit image.

[0013] Perform data augmentation on the apple single-fruit detection dataset to obtain an augmented apple single-fruit detection dataset. The final apple single-fruit detection dataset consists of the apple single-fruit detection dataset and the augmented apple single-fruit detection dataset. For each multi-fruit image M in the apple single-fruit detection dataset, perform rotation, horizontal flipping, vertical flipping, shearing, scaling, displacement, color channel adjustment, and transparency adjustment on the current multi-fruit image M respectively to obtain eight first enhanced images E. Randomly select four from the eight first enhanced images E and perform mosaic data augmentation on them respectively to obtain four first second-enhanced images P1. The apple single-fruit detection dataset consists of multiple multi-fruit images M and their corresponding eight first enhanced images E and four first second-enhanced images P1.

[0014] Perform data augmentation on the apple semantic segmentation dataset to obtain the augmented apple semantic segmentation dataset, which is used as the final apple semantic segmentation dataset. For each apple semantic segmentation map in the apple semantic segmentation dataset, scale, horizontally flip, vertically flip, and shear the apple semantic segmentation map respectively to obtain four enhanced apple semantic segmentation maps. The apple semantic segmentation dataset is composed of the four enhanced apple semantic segmentation maps corresponding to each apple semantic segmentation map.

[0015] Perform data augmentation on the enhanced dataset for apple single fruit surface defect detection to obtain the augmented enhanced dataset for apple single fruit surface defect detection, which is used as the final enhanced dataset for apple single fruit surface defect detection. For each single fruit defect classification image A4 and the corresponding illumination correction image MA4 in the enhanced dataset for apple single fruit surface defect detection, perform rotation, horizontal flip, vertical flip, shear, scale, displacement, color channel adjustment, and transparency adjustment on each single fruit defect classification image A4 and the corresponding illumination correction image MA4 respectively to obtain 8 second enhanced images corresponding to the current single fruit defect classification image A4 and the corresponding illumination correction image MA4. Randomly select 4 images from the 8 second enhanced images corresponding to the current single fruit defect classification image A4 for mosaic data augmentation processing to obtain the corresponding four second secondary enhanced images P2. The final enhanced dataset for apple single fruit surface defect detection is composed of each single fruit defect classification image A4, the 8 second enhanced images corresponding to each illumination correction image MA4, and the four second secondary enhanced images P2.

[0016] The sparse channel individual recognition model in step 1 is composed of an individual recognition backbone network and an individual recognition detection network connected together. The individual recognition backbone network includes 2 convolutional modules, 3 pooling layers, and 4 sparse structure ELAN residual blocks. The individual recognition detection network includes a spatial pyramid pooling cross-stage local connection module, 9 convolutional modules, 4 sparse structure ELAN residual blocks, and 2 upsampling layers.

[0017] The first convolutional module, the second convolutional module, the first sparse structure ELAN residual block, the first pooling layer, the second sparse structure ELAN residual block, the second pooling layer, the third sparse structure ELAN residual block, the third pooling layer, and the fourth sparse structure ELAN residual block are connected in sequence. The second sparse structure ELAN residual block is connected to the sixth convolutional module in the individual recognition detection network. The third sparse structure ELAN residual block is connected to the fourth convolutional module in the individual recognition detection network. The fourth sparse structure ELAN residual block is connected to the spatial pyramid pooling cross-stage local connection module in the individual recognition detection network.

[0018] The spatial pyramid pooling cross-stage partial connection module is connected to the first upsampling layer after passing through the third convolution module. The output of the first upsampling layer and the output of the fourth convolution module are concatenated and then input into the fifth sparse structured ELAN residual block. The fifth sparse structured ELAN residual block is connected to the second upsampling layer after passing through the fifth convolution module. The output of the second upsampling layer and the output of the sixth convolution module are concatenated and then input into the sixth sparse structured ELAN residual block. The sixth sparse structured ELAN residual block is connected to the seventh convolution module, and after performing object detection on the output of the seventh convolution module, a third detection result is obtained;

[0019] The sixth sparse structured ELAN residual block is also connected to the eighth convolution module. The output of the eighth convolution module and the output of the fifth sparse structured ELAN residual block are concatenated and then input into the seventh sparse structured ELAN residual block. The seventh sparse structured ELAN residual block is connected to the ninth convolution module, and after performing object detection on the output of the ninth convolution module, a second detection result is obtained;

[0020] The seventh sparse structured ELAN residual block is also connected to the tenth convolution module. The output of the tenth convolution module and the output of the spatial pyramid pooling cross-stage partial connection module are concatenated and then input into the eighth sparse structured ELAN residual block. The eighth sparse structured ELAN residual block is connected to the eleventh convolution module, and after performing object detection on the output of the eleventh convolution module, a first detection result is obtained; The output of the individual recognition detection network is determined according to the confidence levels in the first - third detection results.

[0021] The linear neighborhood enhanced apple defect classification model in step 2 is composed of an apple defect classification backbone network and an apple defect classification network connected together. After replacing all the sparse structured ELAN residual blocks in the individual recognition backbone network with extended ELAN residual blocks, the apple defect classification backbone network is obtained; A first spatial-to-depth module is added between the sixth sparse structured ELAN residual block and the seventh convolution module in the individual recognition detection network, a second spatial-to-depth module is added between the seventh sparse structured ELAN residual block and the ninth convolution module, and a third spatial-to-depth module is added between the eighth sparse structured ELAN residual block and the eleventh convolution module. Then, after replacing all the sparse structured ELAN residual blocks in the individual recognition detection network with extended ELAN residual blocks, the apple defect classification backbone network is obtained.

[0022] The structures of the first sparse-structured ELAN residual block to the eighth sparse-structured ELAN residual block are the same, and each includes four convolutional modules. The input of the sparse-structured ELAN residual block serves as the input of the twelfth convolutional module and the input of the thirteenth convolutional module. The twelfth convolutional module is connected to the fourteenth convolutional module. The input of the sparse-structured ELAN residual block, the output of the twelfth convolutional module, the output of the thirteenth convolutional module, and the output of the fourteenth convolutional module are concatenated and then input into the fifteenth convolutional module. The output of the fifteenth convolutional module serves as the output of the sparse-structured ELAN residual block.

[0023] The extended ELAN residual block includes five convolutional modules. The input of the extended ELAN residual block serves as the input of the sixteenth convolutional module and the input of the seventeenth convolutional module. The sixteenth convolutional module is connected to the nineteenth convolutional module via the eighteenth convolutional module. The output of the sixteenth convolutional module, the output of the seventeenth convolutional module, the output of the eighteenth convolutional module, and the output of the nineteenth convolutional module are concatenated and then input into the twentieth convolutional module. The output of the twentieth convolutional module serves as the output of the extended ELAN residual block.

[0024] The spatial pyramid pooling cross-stage local connection module includes a spatial pyramid pooling module and four convolutional modules. The input of the spatial pyramid pooling cross-stage local connection module serves as the input of the twenty-first convolutional module and the input of the twenty-second convolutional module. The twenty-first convolutional module is connected to the twenty-third convolutional module via the spatial pyramid pooling module. The output of the twenty-third convolutional module is concatenated with the output of the twenty-second convolutional module and then input into the twenty-fourth convolutional module. The output of the twenty-fourth convolutional module serves as the output of the spatial pyramid pooling cross-stage local connection module.

[0025] The structures of the first space-to-depth module to the third space-to-depth module are the same, and each is composed of a space-to-depth layer connected to a non-strided convolutional layer.

[0026] During the multi-fruit image prediction process, the sparse channel individual recognition model and the linear neighborhood enhanced apple defect classification model adopt model asynchronous parallel operation. Specifically, the sparse channel individual recognition model is loaded on one GPU / CPU, and the remaining GPUs / CPUs are all loaded with the linear neighborhood enhanced apple defect classification model. The sparse channel individual recognition model and the linear neighborhood enhanced apple defect classification model perform asynchronous operation, and the image data transfer is completed through a linked list. The linear neighborhood enhanced apple defect classification models perform parallel operation.

[0027] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0028] 1. The present invention establishes a sparse channel individual recognition model, realizes the rapid and adaptive extraction of single apples from images containing multiple apples, obtains the position information of each single apple image in the scene, and can perform cropping according to the position information;

[0029] 2. The present invention establishes a linear neighborhood enhanced apple defect classification model, obtains the category and position information of defects on each single apple image, and automatically identifies and distinguishes the calyx and the fruit stalk;

[0030] 3. The present invention establishes a dual-model chain structure asynchronous defect classification method for classifying apple defects in multi-fruit images, dynamically plans the tasks to be calculated in the buffer area of the chain structure according to the computing resources, and realizes the efficient and stable operation of the recognition method when taking the picture data stream as the input. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is the network structure diagram of the sparse channel individual recognition model;

[0032] Figure 2 It is the network structure diagram of the linear neighborhood enhanced apple defect classification model;

[0033] Figure 3 It is the structural schematic diagram of the extended ELAN residual block;

[0034] Figure 4 It is the structural schematic diagram of the sparse structure ELAN residual block;

[0035] Figure 5 It is the structural schematic diagram of the space-to-depth module;

[0036] Figure 6 It is the comparison diagram of the original data of single fruit surface defect detection and the light-corrected data of single fruit surface defect detection;

[0037] Figure 7 It is the effect diagram of the dual-model chain structure asynchronous defect classification method;

[0038] Figure 8 It is the training flow chart of the dual-model chain structure asynchronous defect classification method;

[0039] Figure 9 It is the GradCAM activation heat map for different category defect recognition;

[0040] Figure 10 It is the structural schematic diagram of the spatial pyramid pooling cross-stage partial connection module;

[0041] Figure 11 It is the structural schematic diagram of the spatial pyramid pooling module;

[0042] Figure 12It is a schematic structural diagram of a convolution module.

[0043] Figure 13 It is a logic block diagram of the asynchronous parallel operation of the model in the method of the present invention. Detailed implementation manners

[0044] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0045] The present invention proposes a multi-model asynchronous processing method for apple single fruit positioning and surface defect detection. The following processes are included in the deep learning model training stage, as Figure 8 shown:

[0046] 1) Build a general apple imaging system, keeping the positions of the light source, camera, and the apple imaging area relatively fixed. Place an apple in the imaging area and collect images.

[0047] 2) Use labelImg to annotate single apples in the collected images to establish an apple single fruit detection data set. The annotated information includes the apples and their location information. Divide the training set: validation set: test set according to 3:1:1.

[0048] 3) First, cut out the single apple from the completed imaging area. Use labelme to annotate the apple semantic segmentation data set, and the annotated information is the foreground and background of the apple; use labelImg to annotate the original apple single fruit defect detection data set, and the annotated information includes the categories and locations of rot, mechanical damage, fruit rust cracks, shrinkage and softening, extrusion bruise, other damages, fruit stalks, calyces, etc.

[0049] 4) Use the apple single fruit detection data set. After performing data augmentation such as rotation, flipping, shearing, scaling, displacement, color channel adjustment, and transparency adjustment on the data set, then use Mosaic data augmentation, and train a sparse channel individual recognition model (Sparse Channel Apple Detection, SCADet) using the following hyperparameters.

[0050] Table 1 is the initialization parameter table of the SCADet network

[0051]

[0052] 5) Use the apple semantic segmentation data set. Perform data augmentation on the data set using four data augmentation algorithms: scaling, horizontal flipping, vertical flipping, and shearing, and train the FastUnet segmentation model.

[0053] Input the original dataset of apple single fruit surface defect detection into the FastUnet segmentation model to segment the foreground. Use the adaptive illumination correction algorithm to correct the illumination intensity of the foreground image, and generate the illumination-corrected dataset for apple single fruit surface defect detection. Merge the illumination-corrected dataset for apple single fruit surface defect detection with the original dataset of apple single fruit surface defect detection to generate the enhanced dataset for apple single fruit surface defect detection.

[0054] Use the enhanced dataset for apple single fruit surface defect detection. After performing data augmentation such as rotation, flipping, shearing, scaling, displacement, color channel adjustment, and transparency adjustment on the dataset, then use Mosaic data augmentation to train the linear neighborhood enhanced apple defect classification model (Neighborhood Fusion Apple Defect Detection, NFADDet).

[0055] 6) During the prediction process, there is no need to perform semantic segmentation and illumination correction of the fruit body. Input the multi-fruit image to be predicted into the trained sparse channel individual recognition model, and output the single-fruit images corresponding to each apple in the current multi-fruit image;

[0056] 7) Input all the single-fruit images corresponding to the multi-fruit image into the trained linear neighborhood enhanced apple defect classification model, and predict the defects and localization boxes corresponding to each single-fruit image.

[0057] As Figure 13 shown, deploy the trained SCADet and NFADDet models according to the content designed by the multi-algorithm asynchronous operation scheme. Input the complete scene image to be detected frame by frame into the apple surface defect recognition system, and obtain the position of each apple in the scene, as well as the type and position of the defects on the surface of the apple. The multi-algorithm asynchronous operation scheme is used to solve the problem that the number of apples in the input data stream is uneven. This situation is likely to cause that when the number of apples is large, it is difficult for the single-fruit surface defect fast detection algorithm to process in a timely manner, and when the number of apples is small, the single-fruit surface defect fast detection algorithm will be temporarily idle without operation tasks. The present invention adopts a first-in-first-out global double-ended queue with a linked list structure to store and cache the apple single-fruit surface defect images. When a series of apple single-fruit images are recognized by the single-fruit fast extraction algorithm in the complete imaging area of the complete scene picture, they are arranged in the order of the size of the coordinate axis in the running direction of the recognition object channel, and are input into the single-fruit surface defect fast detection algorithm one by one from the end of the queue, and the defect information of each apple is output. The next complete scene picture will continue to input the selected apple single-fruit pictures into the end of the queue after being recognized by the single-fruit fast extraction algorithm in the complete imaging area, and wait to be further processed after the pictures at the front end of the queue are processed, and finally complete the surface defect detection of all apple single fruits in the data stream.

[0058] AsFigure 1 As shown in the figure, the sparse channel individual recognition model is composed of an individual recognition backbone network and an individual recognition detection network connected together. The individual recognition backbone network includes 2 convolutional modules, 3 pooling layers, and 4 sparse structure ELAN residual blocks. The individual recognition detection network includes a spatial pyramid pooling cross-stage partial connection module, 9 convolutional modules, 4 sparse structure ELAN residual blocks, and 2 upsampling layers;

[0059] The first convolutional module, the second convolutional module, the first sparse structure ELAN residual block, the first pooling layer, the second sparse structure ELAN residual block, the second pooling layer, the third sparse structure ELAN residual block, the third pooling layer, and the fourth sparse structure ELAN residual block are connected in sequence. The second sparse structure ELAN residual block is connected to the sixth convolutional module in the individual recognition detection network. The third sparse structure ELAN residual block is connected to the fourth convolutional module in the individual recognition detection network. The fourth sparse structure ELAN residual block is connected to the spatial pyramid pooling cross-stage partial connection module in the individual recognition detection network;

[0060] Specifically:

[0061] The input of the first convolutional module is a 3x640x640 picture, and the output is an 8x320x320 feature map, using a 3x3 convolutional kernel and a stride of 2. The second convolutional module performs a 3x3 convolutional operation on the output of the previous layer, and the output is a 16x160x160 feature map, using a stride of 2. The first sparse structure ELAN residual block sets the internal channel number C ELAN to 16, and the output is a 16x160x160 feature map. The first pooling layer performs a 2x2 max pooling operation on the output of the previous layer, and the output is a 16x80x80 feature map. The second sparse structure ELAN residual block sets the internal channel number C ELAN to 32, and the output is a 32x80x80 feature map. The second pooling layer performs a 2x2 max pooling operation on the output of the previous layer, and the output is a 32x40x40 feature map. The third sparse structure ELAN residual block sets the internal channel number C-ELAN to 64, and the output is a 64x40x40 feature map. The third pooling layer performs a 2x2 max pooling operation on the output of the previous layer, and the output is a 64x20x20 feature map. The fourth sparse structure ELAN residual block sets the internal channel number C-ELAN to 128, and the output is a 128x20x20 feature map.

[0062] The Spatial Pyramid Pooling Cross-Stage Partial Connection Module is connected to the first upsampling layer after passing through the third convolutional module. The output of the first upsampling layer and the output of the fourth convolutional module are concatenated and then input into the fifth Sparse Structure ELAN Residual Block. The fifth Sparse Structure ELAN Residual Block is connected to the second upsampling layer after passing through the fifth convolutional module. The output of the second upsampling layer and the output of the sixth convolutional module are concatenated and then input into the sixth Sparse Structure ELAN Residual Block. The sixth Sparse Structure ELAN Residual Block is connected to the seventh convolutional module, and after performing object detection on the output of the seventh convolutional module, the third detection result is obtained;

[0063] The sixth Sparse Structure ELAN Residual Block is also connected to the eighth convolutional module. The output of the eighth convolutional module and the output of the fifth Sparse Structure ELAN Residual Block are concatenated and then input into the seventh Sparse Structure ELAN Residual Block. The seventh Sparse Structure ELAN Residual Block is connected to the ninth convolutional module, and after performing object detection on the output of the ninth convolutional module, the second detection result is obtained;

[0064] The seventh Sparse Structure ELAN Residual Block is also connected to the tenth convolutional module. The output of the tenth convolutional module and the output of the Spatial Pyramid Pooling Cross-Stage Partial Connection Module are concatenated and then input into the eighth Sparse Structure ELAN Residual Block. The eighth Sparse Structure ELAN Residual Block is connected to the eleventh convolutional module, and after performing object detection on the output of the eleventh convolutional module, the first detection result is obtained; The output of the individual recognition detection network is determined according to the confidence levels in the first - third detection results. In a specific implementation, the size of the feature map in the third detection result is 32x80x80, the size of the feature map in the second detection result is 32x40x40, and the size of the feature map in the first detection result is 128x20x20. The detection result with the highest confidence level among the first - third detection results is used as the output for each region in the multi-fruit image and is used as a single-fruit image.

[0065] Such as Figure 2As shown, the linear neighborhood enhanced apple defect classification model is composed of the apple defect classification backbone network and the apple defect classification network connected together. The apple defect classification backbone network is obtained by replacing all the sparse structure ELAN residual blocks in the individual recognition backbone network with extended ELAN residual blocks, and increasing the number of network convolution kernels to four times that of the individual recognition backbone network. In the individual recognition detection network, a first spatial-to-depth module is added between the sixth sparse structure ELAN residual block and the seventh convolution module, that is, the sixth sparse structure ELAN residual block is connected to the seventh convolution module after passing through the first spatial-to-depth module. A second spatial-to-depth module is added between the seventh sparse structure ELAN residual block and the ninth convolution module, that is, the seventh sparse structure ELAN residual block is connected to the ninth convolution module after passing through the second spatial-to-depth module. A third spatial-to-depth module is added between the eighth sparse structure ELAN residual block and the eleventh convolution module, that is, the eighth sparse structure ELAN residual block is connected to the eleventh convolution module after passing through the third spatial-to-depth module. Then, the apple defect classification backbone network is obtained by replacing all the sparse structure ELAN residual blocks in the individual recognition detection network with extended ELAN residual blocks. In the specific implementation, the detection result with the highest confidence in the first - third detection results is used as the output result of each single fruit image.

[0066] In the linear neighborhood enhanced apple defect classification model, first, a reparameterization operation is performed on the first feature map of 512x10x10. The number of channels is halved and a residual connection is added. Then, a 512x5x5 feature map is output through the SPD linear neighborhood fusion plus non-strided convolution module. Finally, a 1x1 convolution layer is used to output the number of anchor boxes * (the number of classes + 5) values. Secondly, a reparameterization operation is performed on the second feature map, and it is upsampled to a size of 128x20x20 and concatenated with the third feature map in the channel dimension. A reparameterization operation is performed again, and then a 256x10x10 feature map is output through the SPD linear neighborhood fusion plus non-strided convolution module, and a 1x1 convolution layer is used to output the number of anchor boxes * (the number of classes + 5) values. Finally, a reparameterization operation is performed on the third feature map, and it is upsampled to a size of 64x40x40 and concatenated with the previous two feature maps in the channel dimension. The last reparameterization operation is performed, and then a 128x20x20 feature map is output through the SPD linear neighborhood fusion plus non-strided convolution module, and a 1x1 convolution layer is used to output the number of anchor boxes * (the number of classes + 5) values.

[0067] Specifically:

[0068] The input of the first convolutional layer is a 3x320x320 image, and the output is a 32x160x160 feature map. A 3x3 convolutional kernel and a stride of 2 are used; the second convolutional layer performs a 3x3 convolution operation on the output of the previous layer, and the output is a 64x80x80 feature map, with a stride of 2; the first extended ELAN residual block sets the internal channel number C ELAN to 32, and the output is a 64x80x80 feature map; the first pooling performs a 2x2 max pooling operation on the output of the previous layer, and the output is a 64x40x40 feature map; the second extended ELAN residual block sets the internal channel number C ELAN to 64, and the output is a 128x40x40 feature map; the second pooling performs a 2x2 max pooling operation on the output of the previous layer, and the output is a 128x20x20 feature map; the third extended ELAN residual block sets the internal channel number C ELAN to 128, and the output is a 256x20x20 feature map; the third pooling performs a 2x2 max pooling operation on the output of the previous layer, and the output is a 256x10x10 feature map; the fourth extended ELAN residual block sets the internal channel number C ELAN to 256, and the output is a 512x10x10 feature map.

[0069] As Figure 4 shown, the structures of the first sparse structure ELAN residual block - the eighth sparse structure ELAN residual block are the same, all including four convolutional modules. The input of the sparse structure ELAN residual block serves as the input of the twelfth convolutional module and the input of the thirteenth convolutional module. The twelfth convolutional module is connected to the fourteenth convolutional module. After concatenating the input of the sparse structure ELAN residual block, the output of the twelfth convolutional module, the output of the thirteenth convolutional module, and the output of the fourteenth convolutional module, they are input into the fifteenth convolutional module, and the output of the fifteenth convolutional module serves as the output of the sparse structure ELAN residual block.

[0070] In the sparse structure ELAN residual block, the first convolutional layer performs a 1x1 convolution operation on the input feature map, and the output is a C ELAN xWxH feature map, with a stride of 1; the second convolutional layer performs a 1x1 convolution operation on the input feature map, and the output is a C ELAN xWxH feature map, with a stride of 1; the third convolutional layer performs a 3x3 convolution operation on the output of the previous layer, and the output is a C ELAN xWxH feature map, with a stride of 1; the fourth convolutional layer performs a 3x3 convolution operation on the output of the previous layer, and the output is a C RLAN xWxH feature map, with a stride of 1; the output layer adds the previous layer, the output of the third convolutional layer, the output of the second convolutional layer, and the output of the first convolutional layer to obtain (4xCRLAN ) A feature map of xWxH is subjected to a 1x1 convolution operation, and the output is (2xC ELAN ) xWxH feature map, which is equivalent to forming a residual block with a stride of 1.

[0071] As Figure 3 shown, the structures of the first extended ELAN residual block - the eighth extended ELAN residual block are the same, all including five convolution modules. The input of the extended ELAN residual block serves as the input of the sixteenth convolution module and the seventeenth convolution module. The sixteenth convolution module is connected to the nineteenth convolution module via the eighteenth convolution module. The outputs of the sixteenth convolution module, the seventeenth convolution module, the eighteenth convolution module, and the nineteenth convolution module are concatenated and then input into the twentieth convolution module. The output of the twentieth convolution module serves as the output of the extended ELAN residual block.

[0072] As Figure 10 shown, the spatial pyramid pooling cross-stage local connection module includes a spatial pyramid pooling module and four convolution modules. The input of the spatial pyramid pooling cross-stage local connection module serves as the input of the twenty-first convolution module and the twenty-second convolution module. The twenty-first convolution module is connected to the twenty-third convolution module via the spatial pyramid pooling module. The output of the twenty-third convolution module is concatenated with the output of the twenty-second convolution module and then input into the twenty-fourth convolution module. The output of the twenty-fourth convolution module serves as the output of the spatial pyramid pooling cross-stage local connection module.

[0073] As Figure 11 shown, the spatial pyramid pooling module includes three pooling layers. The input of the spatial pyramid pooling module serves as the input of the fourth - sixth pooling layers respectively. The outputs of the fourth - sixth pooling layers are concatenated with the input of the spatial pyramid pooling module and then serve as the output of the spatial pyramid pooling module.

[0074] As Figure 5 shown, the structures of the first spatial-to-depth module - the third spatial-to-depth module are the same, both consisting of a spatial-to-depth (SPD) layer and a non-strided convolution layer connected in sequence. Among them, the spatial-to-depth layer splits each 2×2×C sub-feature in the intermediate feature map of size S×S×C into 4×1×1×C and then reorganizes and slices them into Finally concatenated into to complete downsampling. The stride in the non-strided convolution layer operation is equal to 1, and the rest of the operations are the same as those of the convolution layer.

[0075] As Figure 12 shown, the structures of the first convolution module - the twenty-fourth convolution module are the same, all consisting of a first convolution layer, a first batch normalization layer, and a first activation layer connected in sequence.

[0076] In specific implementation, an image containing multiple apples is collected, and its resolution is scaled to 640×640 pixels, named as the multi-fruit image M; each apple is labeled with a rectangular box R in the original image M to obtain the apple individual labeled image A1.

[0077] Eight kinds of data augmentations, such as rotation, horizontal flipping, vertical flipping, shearing, scaling, displacement, color channel adjustment, and transparency adjustment, are performed on the original image M to obtain the enhanced image E. Four pictures are selected from the enhanced image E for mosaic data augmentation processing to obtain the secondary enhanced image P1. An apple individual recognition training image set is constructed from the original image and its corresponding enhanced image E and secondary enhanced image P1.

[0078] Multiple multi-fruit images M and their corresponding apple individual labeled images A1 are respectively used as the input parameters and output parameters of the sparse channel individual recognition model, and the sparse channel individual recognition model is trained to obtain the sparse channel individual recognition model.

[0079] The training hyperparameters of the sparse channel individual recognition model include a picture resolution of 640×640 pixels, a batch size of 8, a momentum factor of 0.937, an initial learning rate of 0.001, a weight decay coefficient of 0.0005, and 300 training iteration rounds.

[0080] The image regions labeled by each rectangular box R in the apple individual labeled image A1 are cropped to obtain apple images A2 containing only 1 apple, and the apple images A2 are scaled to 320×320 to obtain multiple single-fruit images A3. Various defects, calyces, and fruit stalks are labeled in the single-fruit images A3 to obtain the single-fruit defect classification labeled image A4.

[0081] For the single-fruit defect classification labeled image A4, a method combining FastUnet segmentation and adaptive illumination correction is used to obtain the illumination-corrected image MA4. The images A4 and MA4 are respectively subjected to eight kinds of data augmentations, such as rotation, horizontal flipping, vertical flipping, shearing, scaling, displacement, color channel adjustment, and transparency adjustment, to obtain the enhanced images EA4 and EMA4. Four pictures are randomly selected from the enhanced images EA4 and EMA4 for mosaic data augmentation processing to obtain the secondary enhanced image P2. An apple defect classification training image set is constructed from the single-fruit defect classification labeled image A4 and its corresponding enhanced images EA4 and EMA4 and secondary enhanced image P2.

[0082] The single-fruit images A3 and their corresponding single-fruit defect classification labeled images A4 are respectively used as the input parameters and output parameters of the linear neighborhood enhanced apple defect classification model, and the linear neighborhood enhanced apple defect classification model is trained to obtain the linear neighborhood enhanced apple defect classification model.

[0083] The training hyperparameters of the linear neighborhood enhancement apple defect classification model include an image resolution of 640×640 pixels, a batch size of 8, a momentum factor of 0.937, an initial learning rate of 0.001, a weight decay coefficient of 0.0005, and 300 training iterations.

[0084] During prediction, the resolution of the multi-fruit image collected is scaled to 640×640 pixels, and the obtained single-fruit image is scaled to 320×320. The single-fruit image is an apple image that contains only 1 apple. Connect multiple single-fruit images A7 with a linked list, and the length of the linked list is L; during prediction, calculate the available computing resources. When the number of available GPUs is N, if the number of GPUs N is greater than or equal to 2, load the sparse channel individual recognition model on one GPU, and parallelize the linear neighborhood enhancement apple defect classification model on the remaining n - 1 GPUs. Input the multiple single-fruit images A7 connected by the linked list into the linear neighborhood enhancement apple defect classification model in sequence, and each linear neighborhood enhancement apple defect classification model on the GPU processes images, and obtain the single-fruit defect classification annotation images A8 corresponding to each single-fruit image A7, as Figure 7 shown.

[0085] The present invention will be described below in conjunction with specific test experiments and data.

[0086] The computer used in the experiment is shown in Table 1:

[0087] Table 1 is the computer device parameter table

[0088]

[0089] First, according to the method described in the present invention, there are 30 each of Shaanxi Red Fuji, Golden Delicious, wild apples from Daliangshan, and Aksu Red Fuji apples, and 50 each of defective Red Fuji apples with frostbite, human damage, dry scars, chicken claw patterns, bitter pit disease, and soft rot disease, for a total of 420. Among them, each apple is photographed in 5 poses: frontal view of the fruit stalk, frontal view of the calyx, side view, oblique view of the fruit stalk, and oblique view of the calyx. A total of 2100 images are collected in the experiment, including 1500 defective fruit images and 600 healthy fruit images.

[0090] In the apple single-fruit detection dataset, the 2100 images are divided into 1260 training set images, 420 validation set images, and 420 test set images according to a ratio of 3:1:1. Data augmentation methods such as rotation, flipping, shearing, scaling, displacement, color channel adjustment, and transparency adjustment are amplified 8 times before Mosaic data augmentation, so the enhanced training set is 10080 images, enhancing the generalization ability of the data.

[0091] Use the training set of this dataset and the hyperparameters shown in Table 2 to train the single-fruit rapid extraction algorithm SCADet for the complete imaging area. During the training process, use the validation set for validation and optimize the model. Finally, use the test set images to evaluate the model's performance.

[0092] The horizontal comparison of the model used in this invention with mainstream deep learning models is shown in Table 3. The invention uses the Mean Average Precision (mAP) as the annotation for evaluating the detection performance of the model, and its calculation formula is as follows.

[0093]

[0094]

[0095]

[0096]

[0097] Among them, Precision represents the recognition accuracy. TP represents that the model determines a positive sample, and in fact it is also a positive sample, which means that the model predicts that there is an apple at a certain position, and there is indeed an apple at that position. FP represents that the model determines a positive sample, but in fact it is a negative sample, which means that the model predicts that there is an apple at a certain position, but there is no apple at that position. FN represents that the model determines a negative sample, but in fact it is a positive sample, which means that the background is misrecognized as an apple. Recall represents the recognition recall rate, AP represents the average precision, mAP represents the mean average precision, and AP i represents the average precision of the i-th class, and n represents the number of recognized classes.

[0098] According to the production application scenario and deployment requirements of this project, the model needs to be designed for lightweight. To evaluate the structure size and speed of the model, the following metrics are used: Floating Point Operations (FlOPs), parameters, and the image processing time per frame. Floating point operations refer to the number of floating point operations, which is used to measure the complexity of the algorithm.

[0099] Table 2 is the initialization parameter table of the SCADet network

[0100]

[0101] Table 3 is the table of the single-fruit rapid extraction effect of each model for the complete imaging area

[0102]

[0103]

[0104] The obtained training results are shown in Table 3. The mAP of the SCADet model for single apple recognition is 100% on the test set. At the same time, this model has achieved good lightweight effects during inference on both GPU and CPU, with inference times of 9.39 ms and 29.71 ms respectively. The number of model parameters and the required number of operations are one order of magnitude lower than those of mainstream models. This method is applicable to different varieties such as Shaanxi Red Fuji, Golden Delicious, wild apples from Daliangshan, and Aksu, and is also applicable to the single fruit recognition and localization of apples with different defects on the surface, such as frostbite, artificial damage, dry scars, chicken claw patterns, bitter pit disease, and soft rot disease, and has good generalization performance.

[0105] First, single apples are cropped from the completed imaging area. The labelme tool is used to annotate the apple semantic segmentation dataset, and the annotated information is the foreground and background of the apples. The apple semantic segmentation dataset consists of 1000 images in total, including 600 in the training set, 200 in the validation set, and 200 in the test set. Four data augmentation algorithms, namely scaling, horizontal flipping, vertical flipping, and shearing, are used to augment the training set. The FastUnet segmentation model is trained, and the validation set is used for evaluation during the training process. Finally, the model is selected to perform performance testing on the test set, and the average intersection over union of segmentation is 97.9%. The time taken to complete segmentation on the CPU is 13.9 ms.

[0106] Similarly, after single apples are cropped from the completed imaging area, 1576 images with surface defects are selected, and the labelImg tool is used to annotate the original dataset for apple single fruit defect detection. The annotated information includes the categories and locations of rot, mechanical damage, fruit rust cracks, shrinkage and softening, extrusion bruising, other damages, fruit stalks, calyces, etc.

[0107] The original dataset for apple single fruit defect detection is input into the FastUnet segmentation model to segment the foreground. The adaptive illumination correction algorithm is used to correct the illumination intensity of the foreground image, generating the illumination correction dataset for apple single fruit defect detection. The effect is as Figure 6 shown, Figure 6 (a) in it is the original image for single fruit defect detection, Figure 6 (b) in it is the illumination correction image for single fruit defect detection. The illumination correction dataset for apple single fruit defect detection and the original dataset for apple single fruit defect detection are merged to generate the enhanced dataset for apple single fruit defect detection. The size of the enhanced dataset for apple single fruit defect detection is 3152 images.

[0108] 3,152 images are divided into 1,892 training set images, 630 validation set images, and 630 test set images according to a ratio of 3:1:1. Data augmentation methods such as rotation, flipping, shearing, scaling, displacement, color channel adjustment, and transparency adjustment are used to amplify the dataset 8 times before Mosaic data augmentation. Therefore, the enhanced training set contains 15,136 images, improving the generalization ability of the augmented data.

[0109] Use the training set of this dataset and the hyperparameters shown in Table 4 to train the single-fruit fast extraction algorithm NFADDet for the complete imaging region. During training, use the validation set to verify and optimize the model. Finally, use the test set images to evaluate the model's performance and compare it with the mainstream lightweight recognition model YOLOv7-tiny.

[0110] Table 4 is the table of NFADDet network training initialization parameters

[0111]

[0112] Table 5 is the performance comparison table between NFADDet and YOLOv7-tiny

[0113]

[0114] Table 6 is the lightweight effect table for apple single-fruit surface defect recognition of each model

[0115]

[0116] The obtained training results are shown in Table 6 and Table 7. The mAP of the NFADDet model for identifying single apples is 0.771 on the test set. Except for the calyx being 0.2% lower than the comparison model, the recognition of other types of defects and the fruit stalk has been significantly improved. At the same time, the inference of this model on GPU and CPU ensures good lightweight effects, with inference times of 11.44ms and 43.70ms respectively. The number of parameters of the model is slightly more than that of the mainstream model due to the existence of the SPD module, but the number of floating-point operations required by the model is less than that of the comparison model. This method is suitable for the positioning and classification of various defects on the apple surface and the fruit stalk / calyx, and can accurately identify rot, fruit rust cracks, extrusion bruises, fruit stalk / calyx, as well as obvious mechanical damage, shrinkage and softening, and some other abnormal damages on the apple surface. Use GradCAM class activation to observe the attention area and recognition effect of the present invention on surface defects, as Figure 9 shown. The deep learning algorithm NFADDet has effectively identified the defect area and features during recognition. Therefore, the NFADDet model proposed in the present invention has good generalization performance in apple surface defect detection.

[0117] In theoretical calculations, it is assumed that the computer has N GPU cards (N ≥ 2). After an apple is included in the input image, the length of the global double-ended queue in the linked list structure is L at this time. The SCADet model is loaded on a separate GPU, and one NFADDet is loaded on each of the remaining N - 1 GPUs.

[0118] After this single image is input into the model, the operation time of the SCADet model is 9.39 ms, and the operation time of the NFADDet model is The time taken for the picture to be input into the model until all apple surface defects are detected is

[0119] Assume that the input image contains 9 apples. When the computer has 2 GPU cards, the total time is 112.35 ms; when the computer has 4 GPU cards, the total time is 43.71 ms; when the computer has 6 GPU cards, the total time is 32.29 ms; when the computer has 10 GPUs, the total time is 20.83 ms.

[0120] In actual operations, a set of 200 pictures collected by a high-throughput uniform light imaging system are used as the input for the trained SCADet and NFADDet models. Some images do not contain apples, and the average number of apples is 8.3. Processing the input data stream, the model solution has the best performance on a PC equipped with an Nvidia RTX 3060 GPU card. The model can process 87 apples per second in asynchronous operation and 72 apples per second in sequential operation. The asynchronous operation is 20.8% more efficient than the sequential operation. The processing performance of the model using the CPU directly on a PC without a GPU card is as follows: the model can process 26 apples per second in asynchronous operation and 24 apples per second in sequential operation. The asynchronous operation is 8.3% more efficient than the sequential operation.

[0121] Taking the multi-model asynchronous operation of a PC equipped with an Nvidia RTX 3060 GPU with the highest processing efficiency as an example for processing apples, assume that it takes 4 different surface defect detections to identify one apple. Then this device can complete the sorting of 21.75 apples per second and can process 78,300 apples per hour. This fully demonstrates the value of the present invention for real-time detection of apple surface defects in practical applications.

[0122] Finally, it should be noted that the above embodiments and descriptions are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced. Without departing from the spirit and scope of the disclosure of the technical solutions of the present invention, they should all be covered by the protection scope of the claims of the present invention.

Claims

1. An apple defect classification method that fuses sparse channel recognition and neighborhood enhanced classification, characterized in that It includes the following steps: Step 1: Train a sparse-channel individual recognition model based on the apple single-fruit detection dataset to obtain a trained sparse-channel individual recognition model. The trained sparse-channel individual recognition model outputs the original dataset for apple single-fruit surface defect detection. The sparse-channel individual recognition model in Step 1 is composed of an individual recognition backbone network and an individual recognition detection network connected. The individual recognition backbone network includes 2 convolutional modules, 3 pooling layers, and 4 sparse-structured ELAN residual blocks. The individual recognition detection network includes a spatial pyramid pooling cross-stage local connection module, 9 convolutional modules, 4 sparse-structured ELAN residual blocks, and 2 upsampling layers; Step 2: Train a single-fruit illumination correction model based on semantic segmentation according to the apple semantic segmentation dataset corresponding to the apple single-fruit detection dataset to obtain a trained single-fruit illumination correction model based on semantic segmentation. After inputting the original dataset for apple single-fruit surface defect detection into the trained single-fruit illumination correction model based on semantic segmentation, the trained single-fruit illumination correction model based on semantic segmentation outputs the illumination correction dataset for apple single-fruit surface defect detection. The single-fruit illumination correction model based on semantic segmentation consists of FastUnet semantic segmentation and adaptive fast illumination correction. Among them, the training set of the FastUnet model is the apple semantic segmentation dataset. Among them, the adaptive fast illumination correction uses the mask image generated by FastUnet to limit the correction area, and constructs a light field for the correction area using the least quadratic regression method to complete the illumination correction; Step 3: The original dataset for apple single-fruit surface defect detection and the illumination correction dataset for apple single-fruit surface defect detection constitute the enhanced dataset for apple single-fruit surface defect detection. Train a linear neighborhood enhanced apple defect classification model based on the enhanced dataset for apple single-fruit surface defect detection to obtain a trained linear neighborhood enhanced apple defect classification model; The linear neighborhood enhanced apple defect classification model is composed of an apple defect classification backbone network and an apple defect classification network connected. After replacing all the sparse-structured ELAN residual blocks in the individual recognition backbone network with extended ELAN residual blocks, the apple defect classification backbone network is obtained; Add a first space-to-depth module between the sixth sparse-structured ELAN residual block and the seventh convolutional module in the individual recognition detection network, add a second space-to-depth module between the seventh sparse-structured ELAN residual block and the ninth convolutional module, and add a third space-to-depth module between the eighth sparse-structured ELAN residual block and the eleventh convolutional module. Then, after replacing all the sparse-structured ELAN residual blocks in the individual recognition detection network with extended ELAN residual blocks, the apple defect classification backbone network is obtained; Step 4: Input the multi-fruit image to be predicted into the trained sparse-channel individual recognition model, and output the single-fruit images corresponding to each apple in the current multi-fruit image; Step 5: Input all the single-fruit images corresponding to the multi-fruit image into the trained linear neighborhood enhanced apple defect classification model, and predict the defects and bounding boxes corresponding to each single-fruit image.

2. The apple defect classification method integrating sparse channel recognition and neighborhood enhancement classification according to claim 1, characterized in that Perform data augmentation on the apple single-fruit detection dataset to obtain an augmented apple single-fruit detection dataset. The final apple single-fruit detection dataset consists of the apple single-fruit detection dataset and the augmented apple single-fruit detection dataset. For each multi-fruit image M in the apple single-fruit detection dataset, perform rotation, horizontal flipping, vertical flipping, shearing, scaling, displacement, color channel adjustment, and transparency adjustment on the current multi-fruit image M respectively to obtain eight first enhanced images E. Randomly select four from the eight first enhanced images E and perform mosaic data augmentation on them respectively to obtain four first second-enhanced images P1. The apple single-fruit detection dataset is composed of multiple multi-fruit images M and their corresponding eight first enhanced images E and four first second-enhanced images P1.

3. A method for classifying apple defects by fusing sparse channel recognition and neighborhood enhancement classification according to claim 1, characterized in that, Perform data augmentation on the apple semantic segmentation dataset to obtain an augmented apple semantic segmentation dataset, which is used as the final apple semantic segmentation dataset. For each apple semantic segmentation map in the apple semantic segmentation dataset, perform scaling, horizontal flipping, vertical flipping, and shearing on the apple semantic segmentation map respectively to obtain four enhanced apple semantic segmentation maps. The apple semantic segmentation dataset is composed of the four enhanced apple semantic segmentation maps corresponding to each apple semantic segmentation map. Perform data augmentation on the enhanced dataset for apple single-fruit surface defect detection to obtain an augmented enhanced dataset for apple single-fruit surface defect detection, which is used as the final enhanced dataset for apple single-fruit surface defect detection. For each single-fruit defect classification image A4 and the corresponding illumination correction image MA4 in the enhanced dataset for apple single-fruit surface defect detection, perform rotation, horizontal flipping, vertical flipping, shearing, scaling, displacement, color channel adjustment, and transparency adjustment on each single-fruit defect classification image A4 and the corresponding illumination correction image MA4 respectively to obtain eight second enhanced images corresponding to the current single-fruit defect classification image A4 and the corresponding illumination correction image MA4. Randomly select four from the eight second enhanced images corresponding to the current single-fruit defect classification image A4 and perform mosaic data augmentation on them respectively to obtain four corresponding second second-enhanced images P2. The final enhanced dataset for apple single-fruit surface defect detection is composed of each single-fruit defect classification image A4, the eight second enhanced images corresponding to each illumination correction image MA4, and the four second second-enhanced images P2.

4. The apple defect classification method integrating sparse channel recognition and neighborhood enhancement classification according to claim 1, characterized in that In the sparse channel individual recognition model, the first convolutional module, the second convolutional module, the first sparse structure ELAN residual block, the first pooling layer, the second sparse structure ELAN residual block, the second pooling layer, the third sparse structure ELAN residual block, the third pooling layer, and the fourth sparse structure ELAN residual block are connected in sequence. The second sparse structure ELAN residual block is connected to the sixth convolutional module in the individual recognition detection network. The third sparse structure ELAN residual block is connected to the fourth convolutional module in the individual recognition detection network. The fourth sparse structure ELAN residual block is connected to the spatial pyramid pooling cross-stage local connection module in the individual recognition detection network. The Spatial Pyramid Pooling Cross-Stage Partial Connection module is connected to the first upsampling layer after the third convolution module. The output of the first upsampling layer and the output of the fourth convolution module are concatenated and then input into the fifth Sparse Structure ELAN Residual Block. The fifth Sparse Structure ELAN Residual Block is connected to the second upsampling layer after the fifth convolution module. The output of the second upsampling layer and the output of the sixth convolution module are concatenated and then input into the sixth Sparse Structure ELAN Residual Block. The sixth Sparse Structure ELAN Residual Block is connected to the seventh convolution module, and after target detection on the output of the seventh convolution module, the third detection result is obtained; The sixth Sparse Structure ELAN Residual Block is also connected to the eighth convolution module. The output of the eighth convolution module and the output of the fifth Sparse Structure ELAN Residual Block are concatenated and then input into the seventh Sparse Structure ELAN Residual Block. The seventh Sparse Structure ELAN Residual Block is connected to the ninth convolution module, and after target detection on the output of the ninth convolution module, the second detection result is obtained; The seventh Sparse Structure ELAN Residual Block is also connected to the tenth convolution module. The output of the tenth convolution module and the output of the Spatial Pyramid Pooling Cross-Stage Partial Connection module are concatenated and then input into the eighth Sparse Structure ELAN Residual Block. The eighth Sparse Structure ELAN Residual Block is connected to the eleventh convolution module, and after target detection on the output of the eleventh convolution module, the first detection result is obtained; The output of the individual recognition detection network is determined according to the confidence levels in the first - third detection results.

5. The apple defect classification method integrating sparse channel recognition and neighborhood enhanced classification according to claim 4, characterized in that The structures of the first Sparse Structure ELAN Residual Block - the eighth Sparse Structure ELAN Residual Block are the same, and each includes four convolution modules. The input of the Sparse Structure ELAN Residual Block serves as the input of the twelfth convolution module and the input of the thirteenth convolution module. The twelfth convolution module is connected to the fourteenth convolution module. The input of the Sparse Structure ELAN Residual Block, the output of the twelfth convolution module, the output of the thirteenth convolution module, and the output of the fourteenth convolution module are concatenated and then input into the fifteenth convolution module. The output of the fifteenth convolution module serves as the output of the Sparse Structure ELAN Residual Block.

6. The apple defect classification method integrating sparse channel recognition and neighborhood enhancement classification according to claim 1, characterized in that The Extended ELAN Residual Block includes five convolution modules. The input of the Extended ELAN Residual Block serves as the input of the sixteenth convolution module and the input of the seventeenth convolution module. The sixteenth convolution module is connected to the nineteenth convolution module via the eighteenth convolution module. The output of the sixteenth convolution module, the output of the seventeenth convolution module, the output of the eighteenth convolution module, and the output of the nineteenth convolution module are concatenated and then input into the twentieth convolution module. The output of the twentieth convolution module serves as the output of the Extended ELAN Residual Block.

7. A method for classifying apple defects by fusing sparse channel recognition and neighborhood enhancement classification according to claim 4, characterized in that, The spatial pyramid pooling cross-stage local connection module includes a spatial pyramid pooling module and four convolutional modules. The input of the spatial pyramid pooling cross-stage local connection module serves as the input of the twenty-first convolutional module and the input of the twenty-second convolutional module. The twenty-first convolutional module is connected to the twenty-third convolutional module after passing through the spatial pyramid pooling module. The output of the twenty-third convolutional module is concatenated with the output of the twenty-second convolutional module and then input into the twenty-fourth convolutional module. The output of the twenty-fourth convolutional module serves as the output of the spatial pyramid pooling cross-stage local connection module.

8. A method for classifying apple defects by fusing sparse channel recognition and neighborhood enhanced classification according to claim 1, characterized in that The structures of the first spatial-to-depth module to the third spatial-to-depth module are the same, and each is composed of a spatial-to-depth layer and a non-strided convolutional layer connected together.

9. The apple defect classification method integrating sparse channel recognition and neighborhood enhanced classification according to claim 1, characterized in that, During the multi-fruit image prediction process, the sparse channel individual recognition model and the linear neighborhood enhanced apple defect classification model adopt asynchronous parallel operation of the models.