Intelligent Fruit Variety Recognition System and Method Based on Improved YOLO Model
By introducing the wavelet pooling layer and the CoTAttention attention mechanism layer in the YOLOv11 model, the calculation amount and speed problems of the existing fruit variety intelligent recognition system in real-time identification and complex scene processing are solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510206638.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-25
AI Technical Summary
The existing intelligent fruit variety identification system has problems such as huge calculations and slow processing speed when identifying and processing complex scenarios in real time, which is difficult to meet the needs of real-time identification.
Using the improved YOLO model, the wavelet pooling layer and the CoTAttention attention mechanism layer were introduced into the YOLOv11 model. The low-frequency information was retained through the wavelet pooling layer's denoising feature map, and the context information was captured through the CoTAttention attention mechanism to enhance the recognition ability of the model.
It improves the recognition accuracy and robustness of the model, and can quickly and accurately identify fruit varieties to meet the needs of real-time identification.
Smart Images

Figure CN119693940B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a fruit variety intelligent recognition system and method based on an improved YOLO model. Background Art
[0002] With the rapid development of artificial intelligence and computer vision technologies, fruit variety intelligent recognition systems are gradually being applied in fields such as agriculture and food detection. Traditional fruit variety recognition methods usually rely on manual observation and classification, which are both time-consuming and laborious, and are prone to errors. With the continuous progress of deep learning technologies, fruit variety intelligent recognition systems based on the YOLO algorithm are becoming an efficient and accurate solution.
[0003] Traditional fruit variety recognition systems usually adopt methods based on CNN and R-CNN for image classification and recognition. However, these methods have some disadvantages in practical applications, which limit their performance in efficient real-time recognition.
[0004] R-CNN is a deep learning method based on region extraction. It performs image segmentation by generating candidate regions and classifies each candidate region using a convolutional neural network to achieve object recognition. Although R-CNN can accurately identify fruit varieties, its computational cost is huge and the processing speed is slow because each candidate region needs to be separately subjected to CNN feature extraction and classification, which makes it perform poorly in real-time recognition tasks.
[0005] CNN is a basic network architecture in the field of deep learning and is widely used in tasks such as image classification and feature extraction. Traditional CNN methods usually require pre-extracting image features and inputting these features into a fully connected layer for classification. Although CNN performs excellently in terms of recognition accuracy, it lacks the ability to accurately predict the target position and quantity when dealing with complex scenes or multi-object detection. Therefore, CNN usually requires additional algorithms to handle multi-object detection tasks, which increases the computational cost and complexity.
[0006] Compared with R-CNN and CNN, YOLO has obvious advantages in fruit variety intelligent recognition systems. Its efficient and real-time characteristics enable it to quickly identify fruit varieties without sacrificing accuracy and adapt to various dynamic environments. While R-CNN and CNN perform well in static image recognition, due to their high computational complexity, they usually have difficulty meeting the requirements of real-time recognition. Therefore, YOLO has become a more ideal solution in the current field of fruit variety recognition.
[0007] The yolov11 adopted by this system is developed based on yolov8. On the basis of the yolov8 version, a C3k2 module is proposed, which can greatly improve the accuracy. In addition, two DWConv are inserted into the classification detection head, which can greatly reduce the number of parameters and the amount of calculation. Generally, yolov11 is superior to the previous versions. Summary of the Invention
[0008] In order to solve the problems existing in the prior art, the present invention provides an intelligent fruit variety recognition system and method based on an improved YOLO model to achieve accurate recognition of fruit varieties.
[0009] To achieve the above object, the present invention provides the following solutions:
[0010] An intelligent fruit variety recognition system based on an improved YOLO model, the system includes:
[0011] A data annotation module, used to annotate fruit images of each variety in the existing fruit dataset to obtain an annotated dataset;
[0012] An initial model construction module, used to introduce a wavelet pooling layer and a CoTAttention attention mechanism layer into the YOLOv11 model to obtain an initial model; wherein, the preset convolutional layer in the YOLOv11 model is replaced with the wavelet pooling layer, and the CoTAttention attention mechanism layer is introduced into the 11th layer of the YOLOv11 model;
[0013] A recognition model construction module, used to train the initial model based on the annotated dataset to obtain a fruit variety recognition model;
[0014] A fruit variety recognition module, used to recognize the fruit variety to be recognized based on the fruit variety recognition model to obtain a fruit variety recognition result.
[0015] Preferably, the data annotation module includes:
[0016] A data annotation unit, used to perform box marking on fruit images of each variety in the existing fruit dataset to obtain marked fruit images;
[0017] A data processing unit, used to flip and crop the marked fruit images to complete the preprocessing of the marked fruit images to obtain the annotated dataset.
[0018] Preferably, the initial model construction module includes a wavelet pooling layer introduction unit, used to replace the Conv convolutional layers in the 5th, 7th, and 9th layers of the Backbone network layer of the YOLOv11 model with wavelet pooling layers.
[0019] Preferably, the recognition model construction module includes:
[0020] A wavelet decomposition unit, configured to perform wavelet pooling on the fruit images in the labeled dataset based on the wavelet pooling layer in the initial model, to obtain a feature map of the fruit images; wherein the feature map includes a local feature map and a global feature map;
[0021] A context extraction unit, configured to mine context information from the feature map based on the CoTAttention attention mechanism layer in the initial model;
[0022] A variety recognition unit, configured to recognize the fruit variety based on the mined context information, and complete the construction of the fruit variety recognition model.
[0023] Preferably, the process of performing wavelet pooling on the fruit images in the labeled dataset includes:
[0024] Decompose the fruit images in the labeled dataset into sub-bands; wherein the sub-bands include a low-low frequency sub-band, a low-high frequency sub-band, a high-low frequency sub-band, and a high-high frequency sub-band;
[0025] Discard the sub-bands that do not meet the preset requirements to obtain the feature map of the fruit images.
[0026] The present invention also provides a fruit variety intelligent recognition method based on an improved YOLO model. Applying the system, the method includes:
[0027] Label the fruit images of each variety in the existing fruit dataset to obtain a labeled dataset;
[0028] Introduce a wavelet pooling layer and a CoTAttention attention mechanism layer into the YOLOv11 model to obtain an initial model; wherein, replace the preset convolutional layer in the YOLOv11 model with the wavelet pooling layer, and introduce the CoTAttention attention mechanism layer at the 11th layer of the YOLOv11 model;
[0029] Train the initial model based on the labeled dataset to obtain a fruit variety recognition model;
[0030] Recognize the fruit variety to be recognized based on the fruit variety recognition model to obtain a fruit variety recognition result.
[0031] Preferably, replace the Conv convolutional layers at the 5th, 7th, and 9th layers of the Backbone network layer of the YOLOv11 model with wavelet pooling layers.
[0032] Preferably, the method for constructing the fruit variety recognition model includes:
[0033] Perform wavelet pooling on the fruit images in the labeled dataset based on the wavelet pooling layer in the initial model to obtain the feature maps of the fruit images; wherein the feature maps include local feature maps and global feature maps;
[0034] Mine context information from the feature maps based on the CoTAttention attention mechanism layer in the initial model;
[0035] Identify fruit varieties based on the mined context information to complete the construction of the fruit variety recognition model.
[0036] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention is an improvement based on the latest version yolov11 of yolo end-to-end detection, and can have higher recognition accuracy compared with other object detections, and still has the fast detection ability of the yolo series. In addition, a wavelet pooling module is added. Through wavelet transform, the feature maps are better denoised, high-frequency information is removed, low-frequency information is retained, the main information of the feature maps is preserved, and recognizable features are extracted. It can effectively improve the robustness of the model. In addition, the CoTAttention attention mechanism is added to capture context information, strengthen the visual expression ability, and can more accurately pay attention to the capture of important information. Therefore, the added module can further improve the extraction and capture of feature information by yolov11 and improve the robustness of the entire model. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 It is a schematic structural diagram of a fruit variety intelligent recognition system based on an improved YOLO model in an embodiment of the present invention;
[0039] Figure 2 It is a schematic structural diagram of the network structure after the improvement of YOLOv11 in an embodiment of the present invention;
[0040] Figure 3 It is a schematic structural diagram of the original module of YOLOv11 in an embodiment of the present invention;
[0041] Figure 4 It is a schematic diagram of the CoTAttention attention mechanism in an embodiment of the present invention;
[0042] Figure 5Schematic diagram of the Wavlet Pooling algorithm in the embodiments of the present invention. Detailed implementation manners
[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0044] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0045] Embodiment 1
[0046] As Figure 1 shown, an intelligent fruit variety recognition system based on an improved YOLO model includes: a data annotation module, an initial model construction module, a recognition model construction module, and a fruit variety recognition module.
[0047] The data annotation module is used to annotate fruit images of each variety in the existing fruit dataset to obtain an annotated dataset. In this embodiment, fruit images of 14 categories including apples, avocados, bananas, cherries, grapes, kiwis, mangoes, oranges, peaches, pineapples, pomegranates, raspberries, strawberries, and watermelons open-sourced from the kaggle website are used as the dataset; the Labelimg image annotation tool is used to annotate fruit targets of various categories, and the annotated dataset is divided into a training set, a test set, and a validation set, which are distributed according to a ratio of 7:2:1.
[0048] A further implementation manner is that the data annotation module includes:
[0049] The data annotation unit is used to perform box marking on fruit images of each variety in the existing fruit dataset to obtain marked fruit images; in this embodiment, when using labelimg to mark the images of a training set with a batch-size, box marking is used, the center coordinates of the fruit are (x, y), the width of the annotation box is w, and the height is h.
[0050] The data processing unit is used to flip, crop, and rotate the marked fruit images to complete the preprocessing of the marked fruit images and obtain an annotated dataset.
[0051] The initial model construction module is used to introduce a wavelet pooling layer and a CoTAttention attention mechanism layer into the YOLOv11 model to obtain an initial model. Among them, the preset convolutional layer in the YOLOv11 model is replaced with a wavelet pooling layer (WaveletPool layer), and the CoTAttention attention mechanism layer is introduced into the 11th layer of the YOLOv11 model. A further implementation is as Figure 2 shown. The initial model construction module includes a wavelet pooling layer introduction unit, which is used to replace the Conv convolutional layers in the 5th, 7th, and 9th layers of the YOLOv11 model Backbone network layer with wavelet pooling layers. The wavelet pooling layers in the 5th, 7th, and 9th layers can greatly improve the aliasing detection for small targets. Especially in the market with a large number of fruits, the detection ability on fruit trees is more obvious. Introducing the CoTAttention attention mechanism layer into the 11th layer of the YOLOv11 model can enhance the inference ability of the model, make up for the information processing ability, and ignore unimportant information.
[0052] The recognition model construction module. As shown in the further implementation, the recognition model construction module includes:
[0053] The wavelet decomposition unit is used to perform wavelet pooling on the fruit images in the labeled dataset based on the wavelet pooling layer in the initial model to obtain the feature maps of the fruit images. Among them, the feature maps include local feature maps and global feature maps. A further implementation is that the process of performing wavelet pooling on the fruit images in the labeled dataset includes:
[0054] Decompose the fruit images in the labeled dataset into subbands. Among them, the subbands include low-low frequency subbands, low-high frequency subbands, high-low frequency subbands, and high-high frequency subbands.
[0055] Discard the non-low frequency subbands to obtain the feature maps of the fruit images.
[0056] In this embodiment, the wavelet transform used by Wavelet Pooling can minimize the artifacts generated by the neighborhood. During the recognition process, the image is very likely to take noise as part of the recognition, resulting in detection errors. Eliminating artifacts can restore the image to the state closest to the original truth, and can enhance the recognition ability of the image. In addition, for the commonly used max pooling and average pooling, they only perform downsampling in the spatial domain. Max pooling selects the maximum value in each pooling region, and average pooling selects the average value in each pooling region. These two poolings usually only retain the most significant features in a specific region. While wavelet pooling decomposes the picture into subbands of different frequencies, which can capture the regional and global features of the image. This makes wavelet pooling more effective in processing multi-dimensional and multi-frequency information, and is more capable of extracting more detailed local and global information.
[0057] Wavelet Pooling implements the wavelet transform for images as Figure 5 shown, where LL is the low-low frequency subband, LH is the low-high frequency subband, HL is the high-low frequency subband, and HH is the high-high frequency subband. After the first wavelet transform, the non-low frequency subbands are discarded, and then the wavelet transform is performed again on the low-low frequency subband, and finally a new feature map is generated to complete the pooling process.
[0058] The context extraction unit is used to mine context information from the feature map based on the CoTAttention attention mechanism layer in the initial model; specifically, as Figure 4 shown, after extracting the important information according to the feature information of the previous wavelet pooling process, an attention enhancement mechanism is implemented through a series of convolutions, fusions, segmentations, etc., which improves the model's attention to important channels and can improve the model's performance in processing complex reasoning tasks. The traditional self-attention mechanism can well trigger the feature interaction of different spatial positions according to the input itself. However, in the traditional self-attention mechanism, all pairwise query-key relationships are isolated, and the query-key pairs are independently learned without exploring the rich context between the two. This severely limits the self-learning ability on the 2D feature map. To alleviate this problem, CoTAttention integrates context information mining and self-attention learning into a unified architecture, which can fully mine the context between adjacent keys.
[0059] The variety recognition unit is used to recognize the fruit variety based on the mined context information to complete the construction of the fruit variety recognition model. The test set is input into the fruit variety recognition model to complete the recognition of the fruit type, and a series of evaluation indicators such as the corresponding precision rate and recall rate are obtained.
[0060] In summary, the present invention inputs the image to be detected into the network structure for forward propagation. First, it enters the key layer BackBone layer for feature extraction. Through a series of CBS, C3K2, WaveletPooling, SPPF, CoTAttention, C2PSA layers, the important information of the image is extracted through a series of convolutions and activation functions. The specific convolution processes of each layer are as Figure 3As shown. Then it enters the Neck layer, which is an important layer for feature fusion in the image recognition process. Through downsampling (Upsample layer) and a series of convolution fusion processes, the feature information of the BackBone is fused with the Neck layer. This part of the process can re-obtain the information that was previously ignored. Finally, it enters the Head layer, predicts the feature information obtained from the Neck layer, and gets the scores indicating whether the corresponding input picture belongs to this class and whether there is a corresponding object. Finally, the classification of fruit types and the determination of positions in the picture are completed.
[0061] Based on the existing yolov11, the present invention adds Wavelet Pooling, and the map50 (the average precision of the model when the intersection over union threshold of the ground truth box and the predicted box is 0.5) value is increased from 84.2% to 85.4%, and the map50-95 (the average precision of the model within the range of the intersection over union threshold of the ground truth box and the predicted box from 0.5 to 0.95) is increased from 62.2% to 63.5%. Then, the CoTAttention attention mechanism is added, and the map50 value is increased to 86.2%, and the map50-95 is increased to 64.1%. It can better identify the varieties of fruits according to the pictures.
[0062] The fruit variety recognition module is used to recognize the fruit variety to be recognized based on the fruit variety recognition model and obtain the fruit variety recognition result.
[0063] Embodiment 2
[0064] The present invention also provides a method for intelligent recognition of fruit varieties based on an improved YOLO model. Applying the system described in Embodiment 1, the method includes:
[0065] Annotate the fruit images of each variety in the existing fruit dataset to obtain an annotated dataset;
[0066] Introduce a wavelet pooling layer and a CoTAttention attention mechanism layer into the YOLOv11 model to obtain an initial model; among them, replace the preset convolutional layer in the YOLOv11 model with a wavelet pooling layer, and introduce a CoTAttention attention mechanism layer in the 11th layer of the YOLOv11 model;
[0067] Train the initial model based on the annotated dataset to obtain a fruit variety recognition model;
[0068] Recognize the fruit variety to be recognized based on the fruit variety recognition model to obtain the fruit variety recognition result.
[0069] A further implementation is to replace the Conv convolutional layers in the 5th, 7th, and 9th layers of the Backbone network layer of the YOLOv11 model with wavelet pooling layers.
[0070] A further embodiment lies in that the method for constructing a fruit variety recognition model includes:
[0071] Performing wavelet pooling on the fruit images in the labeled dataset based on the wavelet pooling layer in the initial model to obtain the feature maps of the fruit images; wherein the feature maps include local feature maps and global feature maps;
[0072] Mining context information from the feature maps based on the CoTAttention attention mechanism layer in the initial model;
[0073] Identifying the fruit variety based on the mined context information to complete the construction of the fruit variety recognition model.
[0074] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A fruit variety intelligent identification system based on an improved YOLO model, characterized in that: The system comprises: A data annotation module is used to annotate the images of various varieties of fruits in the existing fruit dataset to obtain annotated datasets; An initial model construction module, used to introduce a wavelet pooling layer and a CoTAttention attention mechanism layer into a YOLOv11 model to obtain an initial model; wherein the wavelet pooling layer is used to replace a preset convolution layer in the YOLOv11 model, and the CoTAttention attention mechanism layer is introduced into the 11th layer of the YOLOv11 model; the initial model construction module includes a wavelet pooling layer introduction unit, used to replace the 5th, 7th and 9th Conv convolution layers in the Backbone network layer of the YOLOv11 model with wavelet pooling layers, and the 5th, 7th and 9th wavelet pooling layers are used to realize aliasing detection for small targets; The recognition model construction module is used to train the initial model based on the labeled data set to obtain a fruit variety recognition model; the recognition model construction module includes: A wavelet decomposition unit is used to perform wavelet pooling on the fruit images in the labeled data set based on the wavelet pooling layer in the initial model to obtain a feature map of the fruit images; wherein the feature map includes a local feature map and a global feature map; A context extraction unit, configured to mine context information on the feature map based on the CoTAttention attention mechanism layer in the initial model; A variety identification unit, used to identify the fruit variety based on the mined context information, and complete the construction of the fruit variety identification model; The process of performing wavelet pooling on the fruit images in the labeled dataset includes: Decomposing the fruit images in the labeled data set into sub-bands; wherein the sub-bands include a low-low frequency sub-band, a low-high frequency sub-band, a high-low frequency sub-band and a high-high frequency sub-band; Discarding subbands that do not meet preset requirements to obtain a feature map of the fruit image; The fruit variety identification module is used to identify the fruit variety to be identified based on the fruit variety identification model to obtain a fruit variety identification result.
2. The system according to claim 1, characterized in that The data annotation module includes: A data labeling unit is used to mark the fruit images of various varieties in the existing fruit data set with boxes to obtain labeled fruit images; The data processing unit is used to flip and crop the marked fruit image, complete the preprocessing of the marked fruit image, and obtain the annotated data set.
3. A method for intelligent identification of fruit varieties based on an improved YOLO model, using the system described in any one of claims 1-2, characterized in that: The method comprises: Label the images of various varieties of fruits in the existing fruit dataset to obtain a labeled dataset; Introducing a wavelet pooling layer and a CoTAttention attention mechanism layer into the YOLOv11 model to obtain an initial model; wherein the wavelet pooling layer is used to replace the preset convolution layer in the YOLOv11 model, and the CoTAttention attention mechanism layer is introduced into the 11th layer of the YOLOv11 model; Training the initial model based on the labeled data set to obtain a fruit variety recognition model; The fruit variety to be identified is identified based on the fruit variety identification model to obtain a fruit variety identification result.
4. The method according to claim 3, characterized in that: Replace the 5th, 7th, and 9th Conv convolution layers in the Backbone network layer of the YOLOv11 model with wavelet pooling layers.
5. The method according to claim 3, characterized in that: The method for constructing the fruit variety identification model comprises: Based on the wavelet pooling layer in the initial model, wavelet pooling is performed on the fruit images in the labeled data set to obtain a feature map of the fruit images; wherein the feature map includes a local feature map and a global feature map; Performing context information mining on the feature map based on the CoTAttention attention mechanism layer in the initial model; The fruit variety is identified based on the mined context information to complete the construction of the fruit variety identification model.
Citation Information
Patent Citations
Citrus fruit identification method and system based on improved YOLO v7
CN116416613A