Pineapple fruit maturity detection method in data scarcity scene
Through the improved CycleGAN model, high-quality pineapple image data is generated and combined with the YOLOv8m model pruning processing, the data scarcity and imbalance in the pineapple fruit maturity detection is solved, efficient and accurate maturity detection is achieved, and harvesting efficiency and fruit quality are improved.
Patent Information
- Application Number
- CN202510588946.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-19
AI Technical Summary
In the case of scarcity and imbalance of data, the accuracy of pineapple fruit maturity detection is difficult to improve, existing methods are difficult to capture subtle differences and resource limitations, and traditional deep learning lacks effective data augmentation methods.
The improved CycleGAN model is used to generate high-quality diversified image data, combined with the YOLOv8m model for training and pruning, forming a lightweight model for pineapple fruit maturity detection.
It significantly improves the accuracy and resource utilization efficiency of pineapple fruit maturity detection, can realize real-time detection on resource-constrained agricultural equipment, and improves harvesting efficiency and fruit quality.
Smart Images

Figure CN120510609A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision detection, and in particular to a method for detecting the maturity of pineapple fruit in a data-scarce scenario. Background Art
[0002] With the booming global pineapple industry, accurate pineapple fruit maturity detection plays a critical role in improving fruit quality, reducing harvest losses, and enhancing market competitiveness. In real-world cultivation environments, pineapple fruit ripening characteristics are subject to fluctuations in multiple environmental factors, such as light, temperature, and humidity, resulting in high variability. Furthermore, the limited amount of available data across categories and the uneven distribution of maturity types significantly limit the detection performance and generalizability of traditional deep learning techniques, such as convolutional neural networks. Acquiring pineapple maturity data in real-world scenarios presents numerous challenges. Firstly, pineapple plantations are complex environments, with fruit often obscured by branches and leaves, making some areas difficult to access, resulting in incomplete image collection. Secondly, the appearance of pineapples at different maturity stages can be subtle, especially in the early and mid-stages of ripeness, when color and texture changes are less pronounced, making accurate labeling and differentiation challenging. Furthermore, fruit farmers and researchers often lack specialized image acquisition equipment and skills, resulting in variable image quality, further exacerbating data scarcity and imbalance.
[0003] Existing detection methods have obvious shortcomings when dealing with data scarcity and imbalance. Traditional machine vision methods rely on manually designed feature extraction algorithms, such as color analysis and texture calculation. These methods have difficulty capturing subtle differences in fruit maturity when faced with complex natural environments and limited data samples, resulting in limited classification accuracy. Although deep learning methods have advantages in feature learning, most studies focus on single fruit category target detection and lightweight improvements. There are relatively few studies on fruit maturity classification detection, and there is a lack of effective means for data expansion and enhancement. For example, some studies use data augmentation techniques to change image attributes such as brightness and contrast, or perform operations such as rotation and flipping. However, these methods only change the image distribution in the pixel dimension, fail to deeply explore the intrinsic feature differences of fruits at different maturity stages, and cannot fundamentally solve the problems of data scarcity and imbalance. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for detecting the maturity of pineapple fruit in a data-scarce scenario.
[0005] To achieve the above objectives, the technical solutions provided by the present invention are:
[0006] A method for detecting pineapple fruit maturity in a data-scarce scenario, comprising:
[0007] Collect pineapple image data to obtain the original pineapple image dataset;
[0008] Create and train an improved CycleGAN model;
[0009] Generate a new pineapple image dataset using the trained improved CycleGAN model;
[0010] Integrate the original pineapple image dataset and the new pineapple image dataset to form an expanded pineapple image dataset;
[0011] Create a basic YOLOv8m model and train it on the expanded pineapple image dataset.
[0012] Prune the trained YOLOv8m model;
[0013] Adjust and resume training of the pruned YOLOv8m model to obtain a lightweight and improved YOLOv8m model;
[0014] The maturity of pineapple fruit is detected using a lightweight improved YOLOv8m model.
[0015] Furthermore, before collecting pineapple image data, the method also includes dividing the pineapple into maturity stages;
[0016] The maturity of pineapple is specifically divided into immature stage, early mature stage, mid-mature stage and fully mature stage;
[0017] Immature stage: the whole fruit is still mainly dark green; early maturity stage: the fruit color begins to change from dark green to light green; mid-maturity stage: the fruit begins to appear golden yellow in parts; fully mature stage: the fruit fully develops into all yellow.
[0018] Furthermore, when collecting pineapple image data, we avoided the midday strong light period and chose to collect and construct a pineapple image dataset containing multiple maturity stages, different lighting conditions and shooting angles under natural lighting conditions.
[0019] Furthermore, the generator of the improved CycleGAN model is created, including:
[0020] Multi-scale feature extraction module: This module is introduced in the initial downsampling stage of the generator. It uses convolution kernels of different sizes to perform convolution operations. The feature maps of different scales extracted by the convolution kernels of different sizes are fused through feature splicing to obtain a rich feature representation.
[0021] Improved residual block: It contains two convolutional layers and a skip link. The CoordATT attention mechanism module is inserted after the feature map processed by the second convolutional layer, and then the feature map after attention enhancement is added to the feature map of the skip link for output.
[0022] Modulated convolution module: Generates a modulation factor based on the statistical information of the input features, and then multiplies the modulation factor with the weight of the convolution kernel to obtain a dynamically adjusted convolution kernel.
[0023] Furthermore, the working process of the CoordATT attention mechanism module includes:
[0024] Perform global average pooling on the input feature map along the height direction to generate feature representation in the height direction; at the same time, perform global average pooling on the feature map along the width direction and transpose the result to generate feature representation in the width direction;
[0025] After concatenating the feature representations in the height and width directions, they are processed through a shared convolutional layer, normalization, and nonlinear activation functions to extract key spatial information and reduce redundancy.
[0026] The formula is as follows:
[0027] ReLU6(x)=min(max(0,x),6)
[0028]
[0029] Next, the processed features are separated into height and width attention weights, normalized by the Sigmoid function, and applied to the corresponding directions of the input feature map respectively. The formula is as follows:
[0030]
[0031] Finally, by multiplying the input feature map with the generated height and width attention weights, the model enhances the important spatial regions in the feature map while suppressing the unimportant parts, thereby improving the ability to focus on key information, enhancing the model's feature expression ability and final performance.
[0032] Furthermore, the process of training the improved CycleGAN model includes:
[0033] Perform image enhancement processing including random translation, rotation, cropping and local image distortion on the images in the original pineapple image dataset to obtain the preprocessed pineapple image dataset;
[0034] The improved CycleGAN model was trained using the preprocessed pineapple image dataset. Since there are multiple maturity categories, multiple pairs of data are used for mathematical permutations without regard to order. Each pair of data is trained once, and one training cycle is equivalent to 200 rounds of training. The hyperparameter settings are as follows: the learning rate is set to a cosine annealing decay strategy, the learning rate is set to 0.0002 for the first 100 rounds of training, and the learning rate is decayed according to the cosine annealing decay strategy for the last 100 rounds of training. The dataset loading mode is set to data misalignment mode, and the batch size is set to 1.
[0035] Furthermore, in the process of generating a new pineapple image dataset using the improved CycleGAN model, after the improved CycleGAN model generates pineapple images, the generated pineapple images also need to be screened:
[0036] By calculating the similarity index between the generated pineapple image and the real pineapple image, if the similarity index is lower than the set threshold, the generated pineapple image is a qualified image and is put into the new pineapple image dataset; otherwise, the generated pineapple image is an unqualified image and is not put into the new pineapple image dataset.
[0037] Furthermore, the trained YOLOv8m model is pruned, including:
[0038] Construct a parameter dependency graph of the trained YOLOv8m model to obtain the model's inter-layer and intra-layer dependencies.
[0039] Identify parameters in the trained YOLOv8m model that can be safely pruned using a sparse training strategy;
[0040] The trained YOLOv8m model is pruned based on its inter-layer dependencies, intra-layer dependencies, and the identified parameters that can be safely pruned.
[0041] Furthermore, the sparse training strategy is used to identify parameters in the trained YOLOv8m model that can be safely pruned, including:
[0042] For the parameter group Θ={θ1,θ2,...,θ m}, introduce the regularization term R Θ,j , used to motivate the model to learn the prunable parameters in the parameter group Θ;
[0043] Regularization term R Θ,j The expression is as follows:
[0044]
[0045] Where β is a scaling factor used to control the sparsification strength of the jth prunable dimension; S Θ,j S represents the importance score of the parameter θ of the j-th prunable dimension in the parameter group Θ; Θ,min and S Θ,max They represent the minimum and maximum importance scores in the parameter group Θ respectively; through sparse training, the trained model makes the weights of unimportant parameters gradually approach zero under the guidance of the regularization term, so that the parameters that can be pruned can be identified in the subsequent pruning process.
[0046] Furthermore, when pruning a convolutional layer, the dependencies of the batch normalization layer and activation function layer connected to the convolutional layer are considered at the same time to ensure that the pruning operation does not destroy the overall structure and performance of the model.
[0047] Compared with the existing technology, the principles and advantages of this technical solution are as follows:
[0048] 1. The improved CycleGAN model can generate high-quality and diverse synthetic images, effectively expanding the size of the data set and balancing the sample distribution. Compared with traditional data augmentation methods (such as rotation, flipping, etc.), the improved CycleGAN model not only changes the image distribution in the pixel dimension, but also can deeply explore the intrinsic feature differences of fruits at different maturity stages, generating images with more semantic consistency and detail authenticity. In this way, this technical solution fundamentally solves the problems of data scarcity and uneven sample distribution, provides richer and higher-quality training data for the training of the detection model, and significantly improves the generalization performance of the model. For example, during the model training process, the expanded data set enables the YOLOv8m model to better learn the characteristics of different maturity stages, thereby more accurately identifying fruits of various maturity levels in actual detection.
[0049] 2. A lightweight improvement to the YOLOv8m model was made. While maintaining high detection accuracy, this model significantly reduces the number of model parameters and computational complexity, enabling real-time and efficient detection applications on resource-constrained equipment such as agricultural robots. Compared with existing technologies, this model has significant advantages in detection speed and resource consumption, and can better meet the needs of actual agricultural production, improving the intelligence level and economic benefits of pineapple harvesting operations. For example, in actual deployment, the lightweight YOLOv8m (detection) model can quickly process image data and provide real-time feedback on maturity information, thereby improving harvesting efficiency.
[0050] 3. The pineapple maturity level is subdivided into four stages (immature, early mature, mid-mature, and fully mature) and successfully applied to the target detection network. This division method not only more accurately reflects the subtle changes in the pineapple ripening process, but also provides richer category information for the detection model, enabling the model to more accurately identify fruits at different maturity stages. Compared with existing technologies, this detailed maturity division and application significantly improves the accuracy of maturity detection, helping to improve the quality and market competitiveness of pineapple fruits. For example, in practical applications, more accurate maturity detection can reduce the decline in fruit quality caused by improper picking time, thereby increasing the market value of the fruit. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the services required for use in the embodiments or the prior art descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0052] Figure 1 This is a principle flow chart of a method for detecting pineapple fruit maturity in a data-scarce scenario according to an embodiment of the present invention;
[0053] Figure 2 This is a schematic diagram of a generator of an improved CycleGAN model used in a pineapple fruit maturity detection method in a data-scarce scenario of the present invention;
[0054] Figure 3 Schematic diagram of the multi-scale feature extraction module;
[0055] Figure 4 Schematic diagram of the improved residual block;
[0056] Figure 5 This is the workflow diagram of the CoordATT attention mechanism module;
[0057] Figure 6 This is the workflow diagram of the modulation convolution module;
[0058] Figure 7 Flowchart for perceptual loss calculation. DETAILED DESCRIPTION
[0059] The present invention will be further described below in conjunction with specific embodiments:
[0060] like Figure 1 As shown, the method for detecting pineapple fruit maturity in a data-scarce scenario described in this embodiment includes the following steps:
[0061] S1. Pineapple maturity stages are divided based on their color, texture, and other appearance characteristics, combined with field observations and the agricultural industry standard NY / T450-2001, to ensure that pineapples at different stages have clear and distinguishable characteristics.
[0062] Specifically, pineapple maturity is divided into the following stages: immature (A), early mature (B), mid-mature (C), and fully mature (D). In the immature stage, the entire fruit remains primarily dark green. In the early mature stage, the fruit begins to change color from dark green to light green. In the mid-mature stage, parts of the fruit begin to develop golden yellow. In the fully mature stage, the fruit fully develops and turns a solid yellow.
[0063] S2, collecting pineapple image data to obtain an original pineapple image dataset;
[0064] The specific process of this step is as follows:
[0065] Data collection was carried out in the pineapple plantation under natural lighting conditions, avoiding the midday strong sunlight period to reduce the impact of excessive light on image quality. The acquisition equipment used was an Intel Realsense D455f binocular depth camera, and was equipped with an adjustable angle tripod to facilitate stable shooting from multiple angles. During the collection process, pineapple plants at different maturity stages were randomly selected to ensure the randomness and representativeness of the samples. The camera photographed each fruit from multiple angles, including the front, side, top, etc., and at least 4 images were taken of each fruit to fully capture the appearance characteristics of the fruit. For fruits at the same maturity stage, different plants and locations were randomly selected for shooting to further increase the diversity and representativeness of the data.
[0066] This step results in a dataset of pineapple images at multiple maturity stages, under different lighting conditions, and shooting angles.
[0067] S3. Create and train an improved CycleGAN model, where Figure 2 As shown in Figure 1, the generator of the improved CycleGAN model has undergone multiple structural optimizations to improve the quality and diversity of generated images. The specific steps are as follows:
[0068] ① Multi-scale feature extraction module: In the initial downsampling stage of the generator, a multi-scale feature extraction module is introduced. Convolution kernels of different sizes (such as 3×3, 5×5, and 7×7) are used for convolution operations. Feature maps of different scales extracted by convolution kernels of different sizes are fused through feature splicing to obtain richer feature representations, such as Figure 3 The multi-scale feature extraction module can capture feature information at different scales, enhance the model's adaptability to geometric transformations, and thus achieve more accurate feature expression at different spatial resolutions.
[0069] ② Improved residual block: It contains two convolutional layers and a skip link. The CoordATT attention mechanism module is inserted after the feature map processed by the second convolutional layer, and then the feature map enhanced by attention is added to the feature map of the skip link and output, such as Figure 4 As shown;
[0070] like Figure 5 As shown in the figure, the working process of the CoordATT attention mechanism module includes:
[0071] Perform global average pooling on the input feature map along the height direction to generate feature representation in the height direction; at the same time, perform global average pooling on the feature map along the width direction and transpose the result to generate feature representation in the width direction;
[0072] After concatenating the feature representations in the height and width directions, they are processed through a shared convolutional layer, normalization, and nonlinear activation functions to extract key spatial information and reduce redundancy.
[0073] The formula is as follows:
[0074] ReLU6(x)=min(max(0,x),6)
[0075]
[0076] Next, the processed features are separated into height and width attention weights, normalized by the Sigmoid function, and applied to the corresponding directions of the input feature map respectively. The formula is as follows:
[0077]
[0078] Finally, by multiplying the input feature map with the generated height and width attention weights, the model enhances the important spatial regions in the feature map while suppressing the unimportant parts, thereby improving the ability to focus on key information, enhancing the model's feature expression ability and final performance.
[0079] ③Introduce modulated convolution in the convolution layer of the generator, such as Figure 6 The core idea of modulated convolution is to dynamically adjust the weight of the convolution kernel according to the input features. Specifically, modulated convolution generates a modulation factor based on the statistical information of the input features (such as mean, variance, etc.), and then multiplies this modulation factor with the weight of the convolution kernel to obtain a dynamically adjusted convolution kernel. In this way, the weight of the convolution kernel can be dynamically adjusted according to the changes in the input features, thereby better adapting to different input images and improving the generation effect of the generator.
[0080] ④ During the training process, in addition to using traditional adversarial loss and cycle consistency loss, perceptual loss is added. The calculation process of perceptual loss is as follows: Figure 7 As shown in Figure 2, perceptual loss is typically calculated based on a pre-trained deep convolutional network (ResNet). Specifically, the generated image and target image are fed into the pre-trained convolutional network to extract their feature representations. The difference (L1 distance) between the feature representation of the generated image and the feature representation of the target image is then calculated. This difference is the perceptual loss. By minimizing the perceptual loss, the image generated by the generator can be semantically closer to the target image, thereby improving the quality of the generated image.
[0081] The process of training the improved CycleGAN model includes:
[0082] 1. Data preprocessing: A series of preprocessing operations are performed on the images in the original dataset to improve data diversity and quality. Specific operations include traditional image enhancement techniques such as random translation, rotation, cropping, and local image warping. These operations can simulate image variations that may occur in real environments to a certain extent, enhancing the model's adaptability to different image conditions. Through these preprocessing operations, a preliminary expanded dataset is generated, providing a richer set of training samples for subsequent model training.
[0083] ② Model Training: The improved CycleGAN model was trained using the initially expanded dataset. Since there are four maturity categories, mathematically permuting and combining them without regard to order will result in six pairs of data, such as A and B, A and C, A and D, B and C, B and D, and C and D. Each pair of data was trained once, with each training cycle consisting of 200 epochs. Hyperparameter settings included: the learning rate was set to a cosine annealing decay strategy, the learning rate was set to 0.0002 for the first 100 epochs of training, and the learning rate was decayed using the cosine annealing decay strategy for the last 100 epochs of training. The dataset loading mode was set to data unaligned mode, and the batch size was set to 1.
[0084] S4. Generate a new pineapple image dataset using the trained improved CycleGAN model. By generating high-quality and diverse synthetic images, the dataset size is expanded and the sample distribution is balanced.
[0085] In this step, new pineapple images are generated by the trained improved CycleGAN. During the image generation process, the model generates a synthetic image similar to the target maturity stage based on the features of the input image. Specifically, an image of one category will generate new images converted from the other three categories. Therefore, the entire data set will be greatly expanded. In order to ensure the quality and diversity of the generated images, the present embodiment carries out strict quality assessment and screening of the generated images. By calculating the similarity index between the generated pineapple image and the real pineapple image, if the similarity index (FID value) is lower than the set threshold of 25, the generated pineapple image is qualified and placed in the new pineapple image data set. Otherwise, the generated pineapple image is unqualified and is not placed in the new pineapple image data set.
[0086] S5. Integrate the original pineapple image dataset and the new pineapple image dataset to form an expanded pineapple image dataset. The expanded pineapple image dataset is more balanced in sample size and distribution, providing rich and high-quality training data for the subsequent maturity detection model (YOLOv8m model) training.
[0087] S6. Create a basic YOLOv8m model and train it on the expanded pineapple image dataset.
[0088] The specific process of this step is as follows:
[0089] ① Dataset Annotation: The final expanded dataset is annotated using the Label img data annotation software. Use the mouse to draw a bounding box on the image, select the pineapple fruit, and select the corresponding maturity category for annotation. Each of the four categories is named, for example, A for the immature stage, B for the early mature stage, C for the mid-mature stage, and D for the fully mature stage. Each fruit in each image must be annotated to ensure accuracy and completeness. After annotation is complete, save the annotation information. Save the annotation information as a TXT file, which contains information such as the bounding box coordinates and category labels. This information will be used for subsequent model training.
[0090] ② Dataset division: Divide the labeled dataset into training set, validation set and test set, with proportions of 70%, 20% and 10% respectively.
[0091] ③ Train the basic network YOLOv8m: Hyperparameter settings: the learning rate is set to 0.001, a phased adjustment strategy is adopted, the batch size is set to 16, and the number of training rounds is set to 300 rounds.
[0092] S7. Prune the trained YOLOv8m model:
[0093] S7-1. Construct a parameter dependency graph of the trained YOLOv8m model to obtain the inter-layer dependency and intra-layer dependency of the model.
[0094] The YOLOv8m model has a complex network structure, including multiple convolutional layers, batch normalization layers, activation function layers, and residual connections. By analyzing the input-output relationship between each layer in the YOLOv8m model, we can identify inter-layer dependencies and intra-layer dependencies. For example, the output of the convolutional layer serves as the input of the next layer, forming an inter-layer dependency; while the input and output of the batch normalization layer have an intra-layer dependency. These dependencies are expressed as parameter groups, such as P = {p -1 ,p +1 ,p -2 ,p +2 ,...,p -n ,p +n}, where each p represents a parameterized layer or non-parametric operation, and the positive and negative superscripts represent the output and input of the layer, respectively. By constructing a dependency graph, we can clearly demonstrate the coupling relationship between the layers in the YOLOv8m model, providing a basis for subsequent pruning.
[0095] S7-2. Use the sparse training strategy to identify parameters in the trained YOLOv8m model that can be safely pruned, including:
[0096] For the parameter group Θ={θ1,θ2,...,θ m}, introduce the regularization term R Θ,j , used to motivate the model to learn the prunable parameters in the parameter group Θ;
[0097] Regularization term R Θ,j The expression is as follows:
[0098]
[0099] Where β is a scaling factor used to control the sparsification strength of the jth prunable dimension; S Θ,j S represents the importance score of the parameter θ of the j-th prunable dimension in the parameter group Θ; Θ,min and S Θ,max They represent the minimum and maximum importance scores in the parameter group Θ respectively; through sparse training, the trained model makes the weights of unimportant parameters gradually approach zero under the guidance of the regularization term, so that the parameters that can be pruned can be identified in the subsequent pruning process.
[0100] S7-3. Prune the trained YOLOv8m model based on the inter-layer dependencies and intra-layer dependencies of the trained YOLOv8m model and the identified parameters that can be safely pruned.
[0101] For each parameter group, the parameters with the smallest L2 norm are selected for pruning, as these parameters have the least impact on model performance. Since the dependency graph clearly defines the dependencies between layers, the pruning process can simultaneously consider both inter-layer and intra-layer dependencies, achieving consistent pruning across layers. When pruning a convolutional layer, the dependencies of the batch normalization layer and activation function layer connected to the convolutional layer are also considered to ensure that the pruning operation does not damage the overall structure and performance of the model. In this way, the number of parameters and model size of the YOLOv8m model can be effectively reduced, thereby accelerating the model's inference process while maintaining the model's detection accuracy as much as possible.
[0102] S8. Adjust and resume training of the pruned YOLOv8m model to obtain a lightweight and improved YOLOv8m model. The training hyperparameter settings are: the learning rate is set to 0.001, a phased adjustment strategy is adopted, the batch size is set to 16, and the number of training rounds is set to 200.
[0103] S9. Detect the maturity of pineapple fruit using the lightweight improved YOLOv8m model.
[0104] The embodiments described above are only preferred embodiments of the present invention and are not intended to limit the scope of implementation of the present invention. Therefore, any changes made based on the shape and principle of the present invention should be included in the scope of protection of the present invention.
Claims
1. A pineapple fruit maturity detection method in a data-scarce scenario, characterized in that: include: Collect pineapple image data to obtain the original pineapple image dataset; Create and train an improved CycleGAN model; Generate a new pineapple image dataset using the trained improved CycleGAN model; Integrate the original pineapple image dataset and the new pineapple image dataset to form an expanded pineapple image dataset; Create a basic YOLOv8m model and train it on the expanded pineapple image dataset. Prune the trained YOLOv8m model; Adjust and resume training of the pruned YOLOv8m model to obtain a lightweight and improved YOLOv8m model; The maturity of pineapple fruit is detected using a lightweight improved YOLOv8m model.
2. the pineapple fruit maturity detection method under a data scarcity scenario according to claim 1, is characterized in that, Before collecting pineapple image data, the process also involves dividing the pineapple into maturity stages; The maturity of pineapple is specifically divided into immature stage, early mature stage, mid-mature stage and fully mature stage; Immature stage: the whole fruit is still mainly dark green; early maturity stage: the fruit color begins to change from dark green to light green; mid-maturity stage: the fruit begins to appear golden yellow in parts; fully mature stage: the fruit fully develops into all yellow.
3. the pineapple fruit maturity detection method under a kind of data scarcity scenario according to claim 2, is characterized in that, When collecting pineapple image data, we avoided the midday strong light period and chose natural lighting conditions to collect and construct a pineapple image dataset that includes multiple maturity stages, different lighting conditions and shooting angles.
4. the pineapple fruit maturity detection method under a kind of data scarcity scenario according to claim 1, is characterized in that, The generator of the improved CycleGAN model created includes: Multi-scale feature extraction module: This module is introduced in the initial downsampling stage of the generator. It uses convolution kernels of different sizes to perform convolution operations. The feature maps of different scales extracted by the convolution kernels of different sizes are fused through feature splicing to obtain a rich feature representation. Improved residual block: It contains two convolutional layers and a skip link. The CoordATT attention mechanism module is inserted after the feature map processed by the second convolutional layer, and then the feature map after attention enhancement is added to the feature map of the skip link for output. Modulated convolution module: Generates a modulation factor based on the statistical information of the input features, and then multiplies the modulation factor with the weight of the convolution kernel to obtain a dynamically adjusted convolution kernel.
5. the pineapple fruit maturity detection method under a kind of data scarcity scenario according to claim 4, is characterized in that, The working process of the CoordATT attention mechanism module includes: Perform global average pooling on the input feature map along the height direction to generate feature representation in the height direction; at the same time, perform global average pooling on the feature map along the width direction and transpose the result to generate feature representation in the width direction; After concatenating the feature representations in the height and width directions, they are processed through a shared convolutional layer, normalization, and nonlinear activation functions to extract key spatial information and reduce redundancy. The formula is as follows: ReLU6(x)=min(max(0,x),6) Next, the processed features are separated into height and width attention weights, normalized by the Sigmoid function, and applied to the corresponding directions of the input feature map respectively. The formula is as follows: Finally, by multiplying the input feature map with the generated height and width attention weights, the model enhances the important spatial regions in the feature map while suppressing the unimportant parts, thereby improving the ability to focus on key information, enhancing the model's feature expression ability and final performance.
6. The pineapple fruit maturity detection method under a data scarcity scenario according to claim 1, wherein The process of training the improved CycleGAN model includes: Perform image enhancement processing including random translation, rotation, cropping and local image distortion on the images in the original pineapple image dataset to obtain the preprocessed pineapple image dataset; The improved CycleGAN model was trained using the preprocessed pineapple image dataset. Since there are multiple maturity categories, multiple pairs of data are used for mathematical permutations without regard to order. Each pair of data is trained once, and one training cycle is equivalent to 200 rounds of training. The hyperparameter settings are as follows: the learning rate is set to a cosine annealing decay strategy, the learning rate is set to 0.0002 for the first 100 rounds of training, and the learning rate is decayed according to the cosine annealing decay strategy for the last 100 rounds of training. The dataset loading mode is set to data misalignment mode, and the batch size is set to 1.
7. The pineapple fruit maturity detection method under a data scarcity scenario according to claim 1, wherein In the process of generating a new pineapple image dataset using the improved CycleGAN model, after the improved CycleGAN model generates pineapple images, the generated pineapple images also need to be screened: By calculating the similarity index between the generated pineapple image and the real pineapple image, if the similarity index is lower than the set threshold, the generated pineapple image is a qualified image and is put into the new pineapple image dataset; otherwise, the generated pineapple image is an unqualified image and is not put into the new pineapple image dataset.
8. The pineapple fruit maturity detection method under a data scarcity scenario according to claim 1, wherein Prune the trained YOLOv8m model, including: Construct a parameter dependency graph of the trained YOLOv8m model to obtain the model's inter-layer and intra-layer dependencies. Identify parameters in the trained YOLOv8m model that can be safely pruned using a sparse training strategy; The trained YOLOv8m model is pruned based on its inter-layer dependencies, intra-layer dependencies, and the identified parameters that can be safely pruned.
9. The pineapple fruit maturity detection method under a data scarcity scenario according to claim 8, wherein The sparse training strategy is used to identify parameters in the trained YOLOv8m model that can be safely pruned, including: For the parameter group Θ={θ1,θ2,...,θ m }, introduce the regularization term R Θ,j , used to motivate the model to learn the prunable parameters in the parameter group Θ; Regularization term R Θ,j The expression is as follows: Where β is a scaling factor used to control the sparsification strength of the jth prunable dimension; S Θ,j S represents the importance score of the parameter θ of the j-th prunable dimension in the parameter group Θ; Θ,min and S Θ,max They represent the minimum and maximum importance scores in the parameter group Θ respectively; through sparse training, the trained model makes the weights of unimportant parameters gradually approach zero under the guidance of the regularization term, so that the parameters that can be pruned can be identified in the subsequent pruning process.
10. The pineapple fruit maturity detection method under a data scarcity scenario according to claim 8, wherein When pruning a convolutional layer, the dependencies between the batch normalization layer and the activation function layer connected to the convolutional layer are considered to ensure that the pruning operation does not destroy the overall structure and performance of the model.