Real-time detection method for maturity and damage degree of apples

By generating adversarial network expansion data and combining model lightweighting, pruning and knowledge distillation technologies, the high equipment cost and poor real-time performance in Apple maturity and damage detection are solved, and efficient and accurate detection results are achieved in the actual production environment.

CN120219971APending Publication Date: 2025-06-27CHINA AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510380944.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing Apple maturity and damage detection technology has problems such as high equipment costs, poor real-time performance, and lack of accuracy, which is difficult to effectively apply in actual production environments.

Method used

Generative adversarial network expansion data is adopted to lighten the AppleLite model, and the model pruning, knowledge distillation and adding convolutional block attention modules are used to achieve real-time detection of Apple's maturity and damage degree through the target detection model.

Benefits of technology

The rapid detection of apple maturity and damage on the farm or sorting line is achieved, reducing the computing and storage requirements of the model, enabling it to run on resource-limited devices, and improving the accuracy and real-time detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219971A_ABST
    Figure CN120219971A_ABST
Patent Text Reader

Abstract

The invention discloses an apple maturity and damage degree real-time detection method. The method comprises the following steps: S1, expanding data by using a generative adversarial network; s2, the main body convolutional neural network of the AppleLite model is lightened; s3, compressing the model by using a model pruning technology; s4, utilizing knowledge to distill and compress the model; s5, a convolution block attention module is added in the AppleLite model; and S6, positioning the target in combination with the target detection model. According to the apple maturity and damage degree real-time detection method, on the basis of an AppleLite apple real-time detection model, detection requirements on the apple maturity and damage degree in an actual production environment are focused, real-time performance and calculation efficiency are considered, contradiction between data insufficiency and data precision is balanced, and the detection accuracy of the apple maturity and damage degree is improved. Through the targeted optimization combination of the prior art, the existing method is improved and integrated in combination with the actual demand, and the practical problem in the production scene is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision, computer neural networks, and object detection technology, and in particular to a method for real-time detection of apple maturity and damage degree. Background Art

[0002] Image detection technologies based on deep learning cover convolutional neural networks (CNNs), object detection models, generative adversarial networks (GANs), attention mechanisms, etc. These technologies are trained with a large number of image datasets, enabling the model to automatically learn features and then complete classification or object detection tasks. They each excel in different fields. For example, convolutional neural networks are suitable for image classification and feature extraction; object detection technologies have excellent real-time detection performance and can quickly identify surface defects and color features of objects in pictures; GANs are commonly used for data augmentation to generate realistic images, solve the problems of insufficient samples and class imbalance, and improve the generalization ability of the model to different maturity levels and damage types; the attention mechanism helps the model focus on key feature regions. However, the detection process of apple maturity and damage degree is complex, and it is difficult for a single technology to be competent independently.

[0003] Detection technologies based on traditional image processing mainly detect the maturity and damage degree of apples by analyzing features such as color, shape, texture, and size. However, this technology has many limitations. For example, color analysis relies on the change of the apple skin color from green to yellow or red to judge maturity and is extremely sensitive to lighting conditions; when detecting damage areas based on texture features such as skin smoothness or unevenness, it is difficult to handle complex backgrounds; when using algorithms such as Canny edge detection to identify the apple contour and judge depressions or skin breakages, due to relying on the change of image pixel intensity, once there are noise such as background noise points and tiny skin textures in the image, false edges are easily generated or important edges are missed. In short, in actual apple detection tasks, the robustness and accuracy of traditional image processing technologies are poor, and the limitations are significant.

[0004] Detection technologies based on multi-sensor fusion include hyperspectral imaging, near-infrared (NIR) spectroscopy, and 3D imaging technology, etc. Hyperspectral imaging detects maturity and internal defects by analyzing the spectral information of the apple skin and interior, but its cost is high and the real-time performance is poor; near-infrared spectroscopy measures components such as apple sugar and acidity using near-infrared light to indirectly judge maturity, but it has high requirements for optical equipment; 3D imaging technology generates the 3D shape of apples by means of laser scanning or stereo vision for detecting surface depressions or mechanical damage. However, the data volume generated by 3D imaging far exceeds that of ordinary images, and it is extremely sensitive to ambient light conditions, and the lighting conditions in the actual farm environment are usually difficult to control. Generally speaking, due to factors such as high cost and complex actual production environment, the application scope of detection technologies based on multi-sensor fusion is greatly limited. Summary of the Invention

[0005] The object of the present invention is to provide a method for real-time detection of apple maturity and damage degree, which solves a series of problems faced by existing detection means (such as detection technologies based on traditional image processing and detection technologies based on multi-sensor fusion) when dealing with apple detection tasks in a real background, including high equipment cost, poor real-time performance, and lack of guarantee of accuracy.

[0006] To achieve the above object, the present invention provides a method for real-time detection of apple maturity and damage degree, including the following steps:

[0007] S1. Use a generative adversarial network to augment data;

[0008] S2. Lightweight the main convolutional neural network of the AppleLite model;

[0009] S3. Compress the model using model pruning technology;

[0010] S4. Compress the model using knowledge distillation;

[0011] S5. Add a convolutional block attention module to the AppleLite model;

[0012] S6. Combine a target detection model to locate the target.

[0013] Preferably, S1 specifically includes the following steps:

[0014] S11. Construct a generator G and a discriminator G', where the generator is used to generate real new samples, and the discriminator is used to determine whether a sample is real or generated;

[0015] S12. Divide the original data set into two parts z and x, which are used as the training sets of the generator and the discriminator respectively. The generator learns to generate real apple images from random noise, and the discriminator learns to distinguish between generated images and real images;

[0016] S13. During the training process, the generator generates more realistic apple images, while the discriminator distinguishes the true and false of the images. Iterate like this until the generated images approximate real apple images to supplement the apple image data set.

[0017] Preferably, the specific process of S2 is as follows:

[0018] S21. By reducing the number of layers of the deep convolutional neural network and removing redundant layers in the network, initially reduce the number of parameters and the amount of calculation of the AppleLite model;

[0019] S22. Reduce the number of channels in each layer of convolution to further reduce the size and calculation overhead of the model;

[0020] S23. The AppleLite model uses depthwise separable convolutions, which decomposes the standard convolution operation into two steps: first, perform independent convolutions for each channel, and then perform pointwise convolutions to reduce the computational amount and the number of parameters.

[0021] S24. Quantize the AppleLite model by converting floating-point weights to low-precision integers to reduce the storage requirements of the model.

[0022] S25. Use lightweight activation functions to replace traditional ReLU activation functions.

[0023] Preferably, in S3, use model pruning techniques to compress the model, taking the absolute value of the weights as the pruning criterion. The specific process is as follows:

[0024] Sort each element in the weight matrix W by its absolute value, then select a threshold t, and set the weights below the threshold to 0. The pruned weight matrix W′ is expressed as:

[0025]

[0026] where W′[i, j] is the pruned weight matrix, W[i, j] is the original weight matrix, i and j represent the number of rows and columns of the matrix respectively. Replace the original weight matrix W in the model with the pruned weight matrix W′ to obtain a lightweight model.

[0027] Preferably, in S4, use knowledge distillation to transfer the knowledge of the teacher model to the student model. The specific process is as follows:

[0028] First, define a temperature parameter Td to smooth the output probability distribution of the model. Then, use the softmax function to calculate the output probability distributions of the teacher and student models:

[0029]

[0030] where, P T (x) is the output probability distribution of the teacher model, P S (x) is the output probability distribution of the student model, x is the input data, T(x) is the output of the teacher model, and S(x) is the output of the student model;

[0031] Minimize the probability distribution difference L between the teacher and student models. The measurement of the probability distribution difference L includes two parts, the divergence L KL and the cross-entropy L CE . Use a weight coefficient α to balance the two differences, then:

[0032] The calculation formula of L KL is:

[0033] L KL = KL(P T(x) ||P S(x) );

[0034] L CE The calculation formula of

[0035] L CE = CrossEntropy(PS (x) , y);

[0036] Then the calculation formula for the difference in probability distribution finally obtained is:

[0037] L = α × L KL + (1 - α) × L CE .

[0038] Preferably, in S5, a convolutional block attention module CBAM structure is added to the AppleLite model. The CBAM structure obtains the global statistical information of the feature map through the channel attention mechanism and the spatial attention mechanism, and calculates the weighted feature map;

[0039] Among them, the process of the channel attention mechanism is:

[0040] Input the feature map F', perform a pooling operation on the input feature map, model the channel importance through a shared perceptron, weight the channels according to the importance, and output the adjusted feature map Mc;

[0041] The process of the spatial attention mechanism is:

[0042] Input the feature map F', perform a pooling operation on the input feature map in the channel dimension to obtain the aggregated information of each spatial position, pass the aggregated features through a convolutional layer to generate a two-dimensional spatial attention map, multiply it element-wise with the original input feature map, and finally output the enhanced feature map Ms.

[0043] Preferably, the specific process of S6 is:

[0044] First, the object detection model extracts high-dimensional feature information from the input image and generates a feature map containing the abstract features of each region in the image;

[0045] Secondly, use the sliding window-based method or the region proposal network-based method to generate candidate boxes containing the target from the input image, and record the positions and sizes of the candidate boxes;

[0046] Then, enter feature pooling, combine the candidate boxes with the feature map, that is, crop and scale the feature map of each candidate box to the same size, and then use the pooling operation to extract the significant features of the region;

[0047] Next, feature vectors of a fixed size for each candidate box are output. The feature vectors are classified and regressed to determine whether the candidate box contains a certain target, the specific category of the target, and the position coordinate values of the target box are output;

[0048] Finally, non-maximum suppression is used to remove redundant candidate boxes, suppress other boxes with an overlap degree exceeding the threshold, select the best bounding box, and output the target category contained in the best bounding box and the coordinates of the target's bounding box.

[0049] Therefore, the present invention adopts the above real-time detection method for apple maturity and damage degree, and the beneficial effects are as follows:

[0050] (1) Through the lightweight design of the model and the optimized combination of various technologies, the AppleLite model of the present invention can quickly detect the maturity and damage degree of apples on the farm or sorting line, meeting the real-time requirements of the actual production scenario and adapting to the real-time detection requirements of the agricultural site and assembly line environment.

[0051] (2) The present invention adopts technologies such as reducing the network depth, reducing the number of convolutional channels per layer, depthwise separable convolution, quantization processing, using lightweight activation functions, as well as model pruning and knowledge distillation, significantly reducing the computational requirements and storage requirements of the model, enabling the model to run on devices with limited resources such as mobile devices and embedded systems, and facilitating low-cost hardware deployment in actual production.

[0052] (3) The present invention uses a generative adversarial network to expand data, solves the problems of insufficient data and class imbalance, and improves the generalization ability of the model; adds a convolutional block attention module to improve the recognition ability of apple maturity features and local damage; combines a target detection model to achieve accurate positioning. The comprehensive application of these technologies enables the AppleLite model to balance the detection speed while ensuring the detection accuracy, effectively balancing the relationship between accuracy, speed, and resource limitations.

[0053] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. Brief Description of the Drawings

[0054] Figure 1 is a flowchart of the operation of the generative adversarial network (GAN) in an embodiment of the real-time detection method for apple maturity and damage degree of the present invention;

[0055] Figure 2 is a schematic diagram of the knowledge distillation process in an embodiment of the real-time detection method for apple maturity and damage degree of the present invention;

[0056] Figure 3It is a schematic diagram of the channel attention mechanism process in an embodiment of the real-time detection method for apple maturity and damage degree of the present invention;

[0057] Figure 4 It is a schematic diagram of the spatial attention mechanism process in an embodiment of the real-time detection method for apple maturity and damage degree of the present invention. Specific implementation manners

[0058] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0059] Unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those of ordinary skill in the field to which the present invention belongs. The "first", "second" and similar terms used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", "right" are only used to represent relative positional relationships, and when the absolute position of the object to be described changes, the relative positional relationship may also change accordingly.

[0060] The detection based on the AppleLite model in the present invention is a lightweight (Light) apple (Apple) maturity and damage degree detection method based on a neural network (Net), AppleLite = Apple + Lit (Light) + e (Net). AppleLite needs to address the following main technical challenges to solve the problems of apple maturity and damage degree detection in actual production:

[0061] Real-time performance: It is necessary to achieve fast detection on farms or sorting lines.

[0062] Resource limitation: The computing resources of the device are limited, and the model is required to be lightweight.

[0063] Data imbalance and insufficiency: In the actually collected data, the sample quantity is insufficient and the class distribution is uneven.

[0064] Detection accuracy: It is required to achieve high-precision classification and positioning of apple maturity and damage degree.

[0065] To achieve the above object, the present invention provides a real-time detection method for apple maturity and damage degree, including the following steps:

[0066] S1, as Figure 1As shown, to address the problem of insufficient data, the AppleLite model uses a generative adversarial network (GAN) to augment data, generating high-quality realistic apple images, focusing on solving the problems of a small number of damaged samples and immature samples, and enhancing the generalization ability of the model. The specific steps are as follows:

[0067] S11. Construct a generator G and a discriminator G'. The generator is used to generate real new samples, while the discriminator is used to determine whether a sample is real or generated.

[0068] S12. Divide the original dataset into two parts, z and x, as the training sets for the generator and the discriminator respectively, enabling the generator to learn to generate real apple images from random noise, while the discriminator learns to distinguish between generated images and real images.

[0069] S13. During the training process, the generator and the discriminator compete with each other. The generator tries to generate more realistic apple images to "fool" the discriminator, while the discriminator endeavors to distinguish the true from the false of the images. Iterating like this until the generated images are increasingly better and approaching real apple images, the apple image dataset can be supplemented to solve the problems of insufficient samples and class imbalance.

[0070] S2. Lightweight the main convolutional neural network of the AppleLite model. Traditional deep convolutional neural networks contain many layers, resulting in very large computational and memory requirements. The specific process is as follows:

[0071] S21. In the present invention, by reducing the number of layers (depth) of the deep convolutional neural network and removing redundant layers in the network, the number of parameters and the computational amount of the AppleLite model are initially reduced.

[0072] S22. At each convolutional layer, the number of channels (i.e., the number of filters) directly affects the number of parameters and the computational amount of the network. Therefore, in the present invention, the number of channels in each layer of convolution is reduced, further reducing the size of the model and the computational overhead.

[0073] S23. At the same time, the AppleLite model adopts the depthwise separable convolution method, decomposing the standard convolution operation into two steps: first, performing independent convolution for each channel, and then performing pointwise convolution, reducing the computational amount and the number of parameters.

[0074] S24. Quantize the AppleLite model, converting the floating-point weights into low-bit integers (i.e., taking the integer part of the floating-point number and removing the decimal part and the decimal point) to reduce the storage requirements of the model and accelerate the inference process of the model.

[0075] S25. Use a lightweight activation function to replace the traditional ReLU activation function, aiming to improve the performance of the model without significantly increasing the computational cost. It has good adaptability for mobile devices and embedded systems.

[0076] To further compress the AppleLite model and further reduce parameter and resource requirements, the present invention adopts two model compression techniques: model pruning and knowledge distillation.

[0077] S3. Compress the model using the model pruning technique, taking the absolute value of the weights as the pruning criterion. The specific process is as follows:

[0078] First, sort each element in the weight matrix W by its absolute value.

[0079] Then, randomly select a threshold t based on empirical values and model complexity, and set the weights lower than the threshold to 0. The threshold will be adjusted during later experiments until a balance is achieved between model accuracy and complexity. In this way, the pruned weight matrix W′ can be expressed as:

[0080]

[0081] where W′[i, j] is the pruned weight matrix, W[i, j] is the original weight matrix, and i and j are the parameters of the matrix itself, representing the number of rows and columns of the matrix respectively. Then, replace the original weight matrix W in the model with the pruned weight matrix W′ to obtain a lightweight model.

[0082] S4. As Figure 2 shown, compress the model using knowledge distillation. The principle of knowledge distillation compression is to transfer the knowledge of a large and powerful "teacher model" to a small "student model". Here, the knowledge refers to the feature knowledge learned by the teacher model after training, such as a certain feature of an image and its corresponding label in a classification task. After knowledge transfer, the student model can learn the knowledge and features of the teacher model, so as to maintain high performance without increasing computing resources. Through distillation, a small student network can retain the accuracy of the large model as much as possible while reducing model parameters and computational cost, taking into account both the accuracy and lightweight of the model.

[0083] By pre-training the teacher model using a large-scale dataset, learn efficient feature representations for tasks such as classification and detection. The teacher model obtains soft labels representing the predicted probability distributions of each class for the input data. The student model is trained using the soft labels and optimizes the parameters of the student model using a loss function.

[0084] The specific process of knowledge distillation is as follows:

[0085] First, define a temperature parameter Td to smooth the output probability distribution of the model. Then, use the softmax function to calculate the output probability distributions of the teacher and student models:

[0086]

[0087] Among them, P T (x) is the output probability distribution of the teacher model, P S (x) is the output probability distribution of the student model, x is the input data, T(x) is the output of the teacher model, and S(x) is the output of the student model.

[0088] Next, minimize the probability distribution difference L between the teacher and student models so that the student model can learn the knowledge of the teacher model. The measurement of the probability distribution difference L includes two parts, L KL divergence and cross-entropy L CE . Balance the two differences with a weight coefficient α, then: KL divergence and cross-entropy L CE The calculation formula of L

[0089] is: KL The calculation formula of L

[0090] is: KL L T(x) =KL(P S(x) );

[0091] The calculation formula of L CE is:

[0092] L CE =CrossEntropy(P S(x) , y);

[0093] Then the final calculation formula of the probability distribution difference is:

[0094] L=α×L KL +(1-α)×C LE .

[0095] S5. Add the convolutional block attention module CBAM structure to the AppleLite model to help the model focus on key feature regions, integrate the attention mechanism into the lightweight model, improve the recognition ability of apple ripening features and local damage, so that the model can concentrate limited computing resources on key feature regions and effectively improve the model performance without additional computational load.

[0096] The CBAM structure mainly plays the role of optimizing the model through the channel attention mechanism and the spatial attention mechanism. The channel attention mechanism is used to make the network pay more attention to the importance between channels and eliminate redundant information; the spatial attention mechanism is used to make the network focus on the feature importance of different positions in the image, improving the local accuracy and global consistency of features. The CBAM structure obtains the global statistical information of the feature map through these two attention mechanisms and calculates the weighted feature map, thereby improving the representation ability of the feature map.

[0097] Among them, as Figure 3 shown, the process of the channel attention mechanism is as follows:

[0098] Input the feature map F’, perform pooling operations (global average pooling or global max pooling) on the input feature map, model the channel importance through a shared perceptron, weight the channels according to importance, enhance the attention to important channels, suppress the interference of useless channels, and output the adjusted feature map Mc.

[0099] As Figure 4 shown, the process of the spatial attention mechanism is as follows:

[0100] Input the feature map F’, perform pooling operations on the input feature map in the channel dimension, that is, global average pooling or global max pooling, to obtain the aggregated information of each spatial position. Pass this aggregated feature through a convolutional layer to generate a two-dimensional spatial attention map, and multiply it element-wise with the original input feature map to weight different spatial positions of the feature map, emphasize the important features of specific regions, and finally output the enhanced feature map Ms, realizing feature weighting in the spatial dimension.

[0101] S6. Combine the object detection model based on the convolutional neural network architecture to locate the object.

[0102] Complete the first-step feature extraction through the aforementioned lightweight convolutional neural network. First, the object detection model extracts high-dimensional feature information from the input image and generates a set of feature maps containing the abstract features of each region in the image, including the brightness of a certain pixel point in the image and the information on the brightness change between the upper and lower pixels, and converts them into digital expressions. The conversion process is called abstraction. Secondly, the original data enters the region proposal stage, and a set of candidate boxes that may contain the object are generated from the input image using the sliding window-based method or the region proposal network-based method, and the position and size of each candidate box are recorded.

[0103] Then, enter the feature pooling process, combine the candidate boxes with the feature maps, that is, crop and scale the feature maps of each candidate box to the same size, and then use pooling operations to extract the significant features of the region.

[0104] Next, feature vectors of a fixed size for each candidate bounding box are output, and these feature vectors are classified and regressed. By classification, it means that each candidate bounding box is classified through a classification layer (such as a fully connected layer) to determine whether the candidate bounding box contains a certain target and the specific category of the target. For example, to determine whether a certain area contains "apple" or "person". By regression, it means predicting the position of the candidate bounding box through a regression layer and outputting the 4 position coordinate values of the target bounding box (the x and y coordinates of the upper left corner, the width w, and the height h).

[0105] Finally, non-maximum suppression is used to remove redundant candidate bounding boxes, suppressing other boxes with an overlap degree exceeding the threshold, and selecting the best bounding box. Specifically, for each candidate bounding box, calculate its confidence score (i.e., the probability that the target is a specific category), and sort them in descending order of the score.

[0106] The formula for calculating the confidence score is as follows:

[0107]

[0108] Among them, Confidence Score refers to the confidence score, P(Object) is to predict whether there is a target within the bounding box through binary classification, is the intersection over union of the predicted bounding box and the ground truth bounding box.

[0109] The selected box should be the one with a high degree of overlap by other boxes and the highest confidence score at the same time. Because this means that the selected box covers the most comprehensive and reliable information, and other boxes will inevitably overlap with this box when covering similar information. For each box, compare it with all other boxes with a high overlap. If their overlap degree exceeds the threshold of 0.5, the overlapping boxes will be deleted, and only the box with the highest score will be retained.

[0110] Finally, the output of the best bounding box contains the target category and the bounding box coordinates of the target. After removing the redundancy, only the most accurate and representative target box position (i.e., the bounding box coordinates of the target) and the target category are retained.

[0111] Therefore, the present invention adopts the above-mentioned real-time detection method for apple maturity and damage degree, combines the requirements of the actual apple production scenario, and conducts modular improvement and integration. The optimized combination of AppleLite shows high practicability and adaptability, realizes real-time detection in the agricultural field, and through lightweight design, adapts to mobile devices and embedded systems, and meets the real-time requirements in the farm and pipeline environments.

[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements do not enable the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A real-time detection method for apple maturity and damage degree, characterized in that: The following steps are involved: S1. Use generative adversarial networks to expand data; S2. Lightweight the main convolutional neural network of the AppleLite model; S3, compress the model using model pruning technology; S4, compressing the model using knowledge distillation; S5. Add convolutional block attention module to AppleLite model; S6. Locate the target by combining the target detection model.

2. A method for real-time detection of apple maturity and damage according to claim 1, characterized in that: S1 specifically includes the following steps: S11. Construct a generator G and a discriminator G', where the generator is used to generate real new samples and the discriminator is used to determine whether a sample is real or generated; S12. Divide the original dataset into two parts z and x, which are used as training sets for the generator and discriminator respectively. The generator learns to generate real apple images from random noise, and the discriminator learns to distinguish between generated images and real images. S13. During the training process, the generator generates more realistic apple images, while the discriminator distinguishes whether the images are real or fake. This process is repeated until the generated images are close to real apple images, thus supplementing the apple image dataset.

3. A method for real-time detection of apple maturity and damage according to claim 2, characterized in that: The specific process of S2 is: S21. By reducing the number of layers of the deep convolutional neural network and removing redundant layers in the network, the number of parameters and the amount of calculation of the AppleLite model are initially reduced; S22, reduce the number of channels in each convolution layer to further reduce the size and computational overhead of the model; The S23 and AppleLite models use a depth-wise separable convolution method, which decomposes the standard convolution operation into two steps: first, independent convolution of each channel, and then point-by-point convolution, which reduces the amount of calculation and parameters; S24, quantizing the AppleLite model, converting floating point weights into low-order integers to reduce the storage requirements of the model; S25. Use a lightweight activation function to replace the traditional ReLU activation function.

4. A method for real-time detection of apple maturity and damage according to claim 3, characterized in that: In S3, model pruning technology is used to compress the model, and the absolute value of the weight is used as the pruning standard. The specific process is as follows: Sort each element in the weight matrix W by its absolute value, then select a threshold t and set the weights below the threshold to 0. The pruned weight matrix W′ is expressed as: Among them, W′[i, j] is the pruned weight matrix, W[i, j] is the original weight matrix, i and j represent the number of rows and columns of the matrix respectively. The original weight matrix W in the model is replaced with the pruned weight matrix W′ to obtain a lightweight model.

5. A method for real-time detection of apple maturity and damage according to claim 4, characterized in that: In S4, knowledge distillation is used to transfer the knowledge of the teacher model to the student model. The specific process is as follows: First, a temperature parameter Td is defined to smooth the output probability distribution of the model. Then, the softmax function is used to calculate the output probability distribution of the teacher and student models: Among them, P T (x) is the output probability distribution of the teacher model, P S (x) is the output probability distribution of the student model, x is the input data, T(x) is the output of the teacher model, and S(x) is the output of the student model; Minimize the probability distribution difference L between the teacher and student models. The measurement of the probability distribution difference L includes two parts, L KL Divergence and Cross Entropy L CE , using the weight coefficient α to balance the two differences, then: L KL The calculation formula is: L KL =KL(P T(x) ||P S(x) ); L CE The calculation formula is: L CE =CrossEntropy(P S(x) ,y); The final calculation formula for the probability distribution difference is: L=α×L KL +(1-a)×L CE 。 6. A method for real-time detection of apple maturity and damage according to claim 5, characterized in that: In S5, a convolutional block attention module (CBAM) structure is added to the AppleLite model. The CBAM structure obtains the global statistical information of the feature map through the channel attention mechanism and the spatial attention mechanism, and calculates the weighted feature map. Among them, the channel attention mechanism process is: Input feature map F', perform pooling operation on the input feature map, model the channel importance through the shared perceptron, weight the channels according to their importance, and output the adjusted feature map Mc; The spatial attention mechanism process is: Input feature map F', perform pooling operation on the input feature map in the channel dimension to obtain the aggregated information of each spatial position, pass the aggregated features through a convolutional layer to generate a two-dimensional spatial attention map, multiply it element-wise with the original input feature map, and finally output the enhanced feature map Ms.

7. A method for real-time detection of apple maturity and damage according to claim 6, characterized in that: The specific process of S6 is as follows: First, the object detection model extracts high-dimensional feature information from the input image and generates a feature map containing abstract features of each region in the image; Secondly, a sliding window-based approach or a region proposal network-based approach is used to generate a candidate box containing the target from the input image, and the position and size of the candidate box are recorded; Then, feature pooling is performed to combine the candidate boxes with the feature maps. The feature maps of each candidate box are cropped and scaled to the same size, and then the pooling operation is used to extract the salient features of the region. Next, a fixed-size feature vector of each candidate box is output, and the feature vector is classified and regressed to determine whether the candidate box contains a certain target and the specific category of the target, and the position coordinate value of the target box is output; Finally, non-maximum suppression is used to remove redundant candidate boxes, suppress other boxes whose overlap exceeds the threshold, select the best bounding box, and output the target category and bounding box coordinates of the target contained in the best bounding box.