CHCG-YOLO11-based lightweight fruit identification method and system

By using the CHCG-YOLO11 lightweight fruit recognition method, combined with the coordinate attention mechanism and multi-scale feature fusion, the backbone and detection head networks are optimized, which solves the problems of high computational complexity and insufficient small target detection accuracy in fruit recognition, and realizes efficient and real-time multi-category fruit recognition.

CN120707556APending Publication Date: 2025-09-26SHANDONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510964916.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing fruit recognition technology has shortcomings in terms of high computational complexity, slow inference speed, insufficient small target detection accuracy, limited model generalization ability, and high computational overhead of detection head design. It is difficult to meet the needs of real-time, multi-target, and multi-category fruit recognition in smart agriculture.

Method used

A lightweight fruit recognition method based on CHCG-YOLO11 is adopted. By combining the coordinate attention mechanism, multi-scale feature fusion and DEGC lightweight detection head, the CHGNetV2 backbone network, CCFC neck network and DEGC detection head network are designed, and the model structure is optimized to achieve a balance between lightweight and high precision.

Benefits of technology

It significantly improves the detection accuracy and robustness of small-target fruits, reduces the number of model parameters and computing resource consumption, meets the needs of real-time detection in smart agriculture, and is suitable for deployment on mobile devices and resource-constrained embedded systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707556A_ABST
    Figure CN120707556A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of computer vision, and discloses a CHCG-YOLO11-based light-weight fruit identification method and a CHCG-YOLO11-based light-weight fruit identification system. According to the method, a CHCG-YOLO11 fruit recognition model is built, so that the requirements of real-time, multi-target and multi-category fruit recognition are met. According to the model, a CA mechanism is introduced into a backbone network, a CCFC module fusing ConvFormer and CGLU structures is introduced into a neck network, and a detection head network adopts a shared parameter strategy based on packet convolution to design a detection head to replace traditional standard convolution and depth separable convolution. According to the invention, through collaborative design and parameter optimization of the backbone network, the neck network and the detection head network, the bottlenecks of model parameter redundancy, large calculation overhead, insufficient small target detection performance and the like are effectively relieved, the objectives of model volume lightweight and real-time reasoning are realized, and the application requirements of intelligent agriculture and embedded systems are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and specifically provides a lightweight fruit recognition method and system based on CHCG-YOLO11. Background Art

[0002] With the rapid development of smart agriculture, vision-based object detection technology has been widely used in multiple applications, including fruit identification, quality inspection, grading and sorting, maturity assessment, and intelligent harvesting. Fruit identification, as a key component of this technology, directly impacts the efficiency and effectiveness of agricultural automation and intelligentization.

[0003] Early fruit recognition approaches based on traditional image processing and machine learning methods, such as those utilizing texture and color features and decision tree classification algorithms, achieved some success for single fruit categories or simple environments. However, due to their weak feature representation and poor anti-interference capabilities, they struggled to meet the complex multi-target, multi-category recognition requirements. In recent years, deep convolutional neural networks (CNNs), with their powerful automatic feature learning capabilities, have become the mainstream technology in the fruit recognition field. In particular, Faster R-CNN and SSD, based on region proposals and end-to-end detection frameworks, as well as the YOLO series of fast real-time detection frameworks, have significantly improved detection accuracy and real-time performance. However, they still have the following major shortcomings:

[0004] 1. The model has high computational complexity and is difficult to meet the real-time needs of resource-constrained environments

[0005] To achieve higher detection accuracy, many current deep learning models employ deep and parameter-intensive network structures, such as Faster R-CNN, SSD, and some YOLO models. These models often suffer from high parameter counts and computational resource consumption, resulting in slow inference speeds and making them difficult to deploy on mobile devices, embedded devices, or in intelligent agricultural applications requiring high real-time performance. In multi-category and multi-target recognition, computational complexity rises rapidly with the number of categories and detected targets, creating a bottleneck. This problem is primarily due to the complex model architectures that introduce numerous convolutional and fully connected layers and lack effective parameter and computational optimization strategies. Furthermore, there is a difficult trade-off between lightweight design and guaranteed accuracy. While attempts to mitigate computational overhead with methods such as model pruning, quantization, and distillation can partially mitigate this, they are often accompanied by a significant decrease in recognition accuracy and unstable performance in complex and changing real-world environments.

[0006] 2. Insufficient detection accuracy for small fruit targets

[0007] Most existing algorithms experience a significant drop in recognition accuracy when faced with small or occluded fruit objects. This is especially true in complex backgrounds, with varying lighting conditions and under occlusion, making it difficult for the models to effectively capture the detailed features of small objects, resulting in high rates of missed detections and false positives. This is primarily due to the inadequacy of traditional detection frameworks in feature extraction and multi-scale fusion, which prevents them from fully utilizing fine-grained information at low levels and high resolution, as well as the lack of effective attention mechanisms to enhance focus on key areas. While the introduction of multi-scale feature fusion and attention mechanisms has provided some improvements, the computational burden imposed by the increased complexity still limits further optimization.

[0008] 3. Limited model generalization capabilities

[0009] Existing models are mostly trained on limited, single datasets. Faced with the diverse fruit varieties, ambient lighting, and background variations found in real-world scenarios, their generalization performance is insufficient, resulting in poor model recognition stability. Especially when training data is scarce or the samples are unbalanced, the models are prone to overfitting or underfitting. This problem stems from the insufficient coverage of the training dataset and the model's limited adaptability to complex environments. Attempts have been made to expand sample diversity through data augmentation, but it remains difficult to fully simulate all environmental variations in real applications, and excessive augmentation can lead to training difficulties or decreased model performance.

[0010] 4. Insufficient feature expression capabilities, making it difficult to balance nonlinearity and global information fusion

[0011] Traditional convolutional networks perform well in extracting local features, but their ability to model global contextual information is limited, making it difficult to effectively distinguish targets from interference, especially in complex backgrounds. Furthermore, the limited use of nonlinear activation functions and gating mechanisms limits the model's ability to express complex feature patterns. Related research has attempted to introduce Transformer structures or gated convolutional units to enhance the model's ability to express nonlinear features and model long-range dependencies. However, this typically results in a significant increase in model structural complexity, reduced training and inference efficiency, and difficulty meeting real-time lightweight requirements.

[0012] 5. The design and calculation cost of the detection head is high

[0013] Some existing model detection heads utilize stacked layers of convolution, particularly multi-scale detection heads, which significantly increase the computational complexity and parameter requirements. While the detection head significantly impacts detection performance, traditional designs struggle to achieve both efficiency and accuracy under the demand for lightweight design. Designing lightweight detection heads often presents the challenge of maintaining detection performance while reducing computational costs using methods such as parameter sharing and grouped convolution. Ensuring accurate positioning and classification within a limited computational budget remains a technical challenge.

[0014] Therefore, it is necessary to propose a lightweight fruit recognition method and system based on CHCG-YOLO11 to solve the above technical problems existing in the prior art. Summary of the Invention

[0015] The purpose of the present invention is to provide a lightweight fruit recognition method and system based on CHCG-YOLO11, which achieves a balance between lightweight and high precision by combining coordinate attention mechanism, multi-scale feature fusion and DEGC lightweight detection head, and meets the needs of real-time, multi-target and multi-category fruit recognition in smart agriculture scenarios.

[0016] To achieve the above object, the present invention provides the following technical solutions:

[0017] The lightweight fruit recognition method based on CHCG-YOLO11 includes the following steps:

[0018] Step 1: Obtain fruit images, construct a dataset containing fruit images, and preprocess the dataset;

[0019] Step 2: Build the CHCG-YOLO11 fruit recognition model, which includes the backbone network, neck network, and detection head network;

[0020] The backbone network includes the HGStem module, a multi-stacked CHGBlock structure, an SPPF module, and a C2PSA module. Each CHGBlock module introduces the CA mechanism on the basis of HGBlock. The neck network uses the CCFC module to replace the C3k2 module in the original YOLO11 neck network. The CCFC module integrates the ConvFormer and CGLU structures in the C3k2 module. The detection head network includes the first, second, and third DEGC lightweight detection heads. Each DEGC lightweight detection head uses grouped convolution to replace traditional standard convolution and depth-separable convolution.

[0021] The processing of the preprocessed data in the CHCG-YOLO11 fruit recognition model is as follows:

[0022] The preprocessed data first passes through the backbone network, where it first passes through the HGStem module, then through a multi-stacked CHGBlock structure to improve feature extraction capabilities, and finally through the SPPF module and the C2PSA module. The multi-scale high-dimensional feature map output by the backbone network passes through the neck network and is input into the detection head network. Finally, the output module receives the output of the detection head network, merges the detection results at different scales, performs non-maximum suppression to remove overlapping boxes, and finally outputs the category and location of the fruit.

[0023] Step 3: Based on the data set of step 1, the CHCG-YOLO11 fruit recognition model built in step 2 is trained and optimized, and then the trained and optimized CHCG-YOLO11 fruit recognition model is used to identify fruit categories.

[0024] A lightweight fruit recognition system based on CHCG-YOLO11 includes an image acquisition device and a computer device; wherein the image acquisition device is used to capture fruit images and upload them to the computer device;

[0025] Computer equipment includes memory and one or more processors;

[0026] An executable code is stored in the memory; when the processor executes the executable code, it is used to implement the steps of the above-mentioned lightweight fruit recognition method based on CHCG-YOLO11.

[0027] Compared with the prior art, the present invention has the following beneficial effects:

[0028] (1) Through the innovative CHGNetV2 backbone network and coordinate attention mechanism, the present invention effectively enhances the feature expression of the fruit target area, especially greatly improves the detection ability of small target fruits, reduces missed detection and false detection, and significantly improves the recognition accuracy.

[0029] (2) This paper adopts depthwise separable convolution, multi-scale parallel convolution structure, and grouped convolution shared parameter design, which significantly reduces the model's parameter scale and computing resource consumption. Compared with the traditional YOLO11 model, the number of parameters and computational complexity are significantly reduced, greatly improving the model's operational efficiency.

[0030] (3) The present invention adopts an optimized lightweight network structure and a DEGC lightweight detection head, which not only improves the model inference speed, but also meets the strict requirements of real-time detection in smart agriculture and is suitable for deployment in mobile devices and resource-constrained embedded systems.

[0031] (4) The present invention integrates global information and nonlinear features through the CCFC module, and applies diversified data enhancement strategies. The model's adaptability to different lighting, occlusion and diverse backgrounds is significantly enhanced, thereby improving the recognition stability and generalization performance.

[0032] (5) Due to the small number of model parameters and high computational efficiency, the present invention not only significantly reduces the computational burden and storage requirements of the model, helps save hardware energy consumption, extends the battery life of embedded devices, and reduces the overall operating cost of the system, but also supports rapid deployment on multiple hardware platforms, is simple to operate, and facilitates rapid fruit classification and detection.

[0033] (6) To address the key challenges in lightweight fruit recognition, this paper constructs a closed-loop "feature optimization-fusion-prediction" system through three innovative modules: the CHGNetV2 module significantly reduces the number of model parameters and enhances feature extraction capabilities; the CCFC module effectively improves the efficiency of multi-scale feature fusion; and the DEGC module significantly reduces the computational complexity of the detection head. These three modules work together to systematically address core issues such as insufficient small object detection accuracy, excessive computational resource consumption, and poor adaptability to multi-scale scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments.

[0035] Figure 1 This is a flowchart of the lightweight fruit recognition method based on CHCG-YOLO11 in Example 1;

[0036] Figure 2 Schematic diagram of the structure of the backbone network, neck network and detection head network in Example 1;

[0037] Figure 3 Schematic diagram of the structure of the multi-stacked CHGBlock structure in Example 1;

[0038] Figure 4 Schematic diagram of the structure of the CHGBlock module in Example 1;

[0039] Figure 5 Schematic diagram of the structure of the CA attention layer in Example 1;

[0040] Figure 6 Schematic diagram of the structure of the CCFC module in Example 1;

[0041] Figure 7 Schematic diagram of the structure of the ConvFormer-CGLUBlock module in Example 1;

[0042] Figure 8 Schematic diagram of the structure of the CGLU module in Example 1;

[0043] Figure 9 It is a structural diagram of a traditional detection head;

[0044] Figure 10 This is a schematic diagram of the structure of the DEGC lightweight detection head in Example 1. DETAILED DESCRIPTION

[0045] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0046] Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative work shall fall within the scope of protection of the present invention.

[0047] In addition, the terms "first," "second," and so on, used in this disclosure are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referenced. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this disclosure, "plurality" means at least two, such as two or three, unless otherwise specifically defined.

[0048] In addition, the technical solutions between the various embodiments of the present invention can be combined with each other, but it must be based on the fact that ordinary technicians in this field can implement it. When the combination of technical solutions is mutually contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0049] Example 1

[0050] Existing deep learning-based fruit recognition technologies generally suffer from high model computational complexity, slow inference speed, insufficient small object detection accuracy, limited model generalization, and high computational overhead in detection head design. This embodiment describes a lightweight fruit recognition method based on CHCG-YOLO11, aiming to achieve the following goals through innovative network module design and structural optimization:

[0051] 1. Significantly reduce the number of model parameters and computational complexity, and improve the model's reasoning speed to meet the real-time detection needs of resource-constrained environments in smart agriculture.

[0052] 2. Improve the detection accuracy of multiple categories of fruits, especially small-target fruits, and enhance the model's robustness to complex environments, occlusions, and background interference.

[0053] 3. Strengthen feature expression capabilities and multi-scale information fusion to improve the recognition stability and generalization performance of the model.

[0054] 4. Design a lightweight and efficient detection head to balance detection accuracy and computational overhead to ensure the optimization of the overall performance of the model.

[0055] This example designs a lightweight and efficient CHCG-YOLO11 fruit recognition model based on the YOLO11 object detection architecture. Its overall structure comprises a backbone network (CHGNetV2), a neck network (including the CCFC module), a detection head network (the DEGC lightweight detection head), and a training optimization strategy. Data flow and computational collaboration between these modules enable rapid and accurate recognition of multiple fruit categories.

[0056] like Figure 1 As shown, the lightweight fruit recognition method based on CHCG-YOLO11 described in this embodiment includes the following steps:

[0057] Step 1: Obtain fruit images, build a dataset containing fruit images, and preprocess the dataset.

[0058] Step 2: Build the CHCG-YOLO11 fruit recognition model, which includes the backbone network, neck network, and detection head network.

[0059] like Figure 2 As shown in the figure, the backbone network (CHGNetV2 backbone network) includes the HGStem module, a multi-stacked CHGBlock structure, an SPPF module, and a C2PSA module. Each CHGBlock module introduces the CA mechanism on top of the HGBlock. The neck network uses the CCFC module to replace the C3k2 module in the original YOLO11 neck network. The CCFC module integrates the ConvFormer and CGLU structures in the C3k2 module. The detection head network includes the first, second, and third DEGC lightweight detection heads. Each DEGC lightweight detection head uses grouped convolution to replace traditional standard convolution and depthwise separable convolution.

[0060] The processing of the preprocessed data in the CHCG-YOLO11 fruit recognition model is as follows:

[0061] The preprocessed data (RGB fruit images of size 640×640) first passes through the backbone network, and then passes through the HGStem module within the backbone network. The HGStem module serves as the preprocessing layer of the backbone network and is composed of standard convolution modules; then it passes through a multi-stacked CHGBlock structure to improve feature extraction capabilities, and finally is processed by the SPPF module and the C2PSA module in sequence; the multi-scale high-dimensional feature map output by the backbone network passes through the neck network and is input into the detection head network. Finally, the output module receives the output of the detection head network, merges the detection results of different scales, performs non-maximum suppression to remove overlapping boxes, and finally outputs the category and position of the fruit.

[0062] The multi-stacked CHGBlock structure is improved by introducing the CA mechanism on the basis of HGBlock. Figure 3 As shown in FIG, the multi-stacked CHGBlock structure includes six CHGBlock modules and three depthwise convolution modules. The six CHGBlock modules are defined as the first, second, third, fourth, fifth, and sixth CHGBlock modules, and the four depthwise convolution modules are defined as the first, second, and third depthwise convolution modules.

[0063] The feature input is placed in a multi-stacked CHGBlock structure and processed in sequence by the first CHGBlock module, the first depth convolution module, the second CHGBlock module, the second depth convolution module, the third CHGBlock module, the fourth CHGBlock module, the fifth CHGBlock module, the third depth convolution module, and the sixth CHGBlock module.

[0064] like Figure 3 As shown in the figure, the feature processing process in the multi-stacked CHGBlock structure can be divided into four stages. The first stage contains only one CHGBlock module, while the second to fourth stages integrate the deep convolution module (DWConv module) and multiple CHGBlock modules to improve the feature extraction capability.

[0065] like Figure 4 As shown in the figure, the features first pass through multiple stacked 3×3 standard convolutions in each CHGBlock module to enhance feature expression. They are then compressed and fused through two 1×1 convolutions to reduce computation and optimize information flow.

[0066] Among them, first, multiple convolution transformations are performed on the input x1 ( Figure 4 Multiple 3×3 standard convolutions in the network) to generate multiple intermediate feature maps {y1,y2,…,y n}, and then perform channel compression through 1×1 squeeze convolution ( Figure 4 The first 1×1 convolution in , the formula is as follows:

[0067] y s =σ(W s Concat(x1,y1,y2,…,y n ));

[0068] Among them, W s Represents the weight parameter matrix of squeeze convolution, y sis the output feature of 1×1 squeeze convolution, σ is the nonlinear activation function Sigmoid, and Concat represents channel compression through 1×1 squeeze convolution.

[0069] Then through 1×1 excitation convolution ( Figure 4 The second 1×1 convolution in ( ) restores the channel dimension and obtains the fused feature map y. The formula is as follows:

[0070] y=σ(W e ·y s );

[0071] Among them, W e Represents the weight parameter matrix of the excitation convolution.

[0072] On this basis, the CA mechanism (CA attention layer) is introduced. Figure 5 As shown in the figure, spatial information is aggregated through adaptive average pooling in the X and Y directions. Specifically, the fused feature map y∈R C×H×W Enter the CA attention layer and first calculate the adaptive average pooling features in the X and Y directions. The formula is as follows:

[0073]

[0074] Among them, H represents the height of the feature map, W represents the width of the feature map, i represents the index in the height direction, j represents the index in the width direction, and Y x represents the adaptive average pooling feature along the X direction, Y y Represents the adaptive average pooling feature along the Y direction, and y(i,j) represents the fused feature map y∈R C×H×W .

[0075] Then use 1×1 convolution (two-dimensional convolution), BatchNorm normalization and h-swish activation function to fuse channel attention information and get the attention weight A x and A y , the formula is as follows:

[0076] A x =σ(W x ×Y x );

[0077] A y =σ(W y ×Y y );

[0078] Among them, W x and W yare the 1×1 convolution weights in the X and Y directions respectively, and σ is the nonlinear activation function Sigmoid.

[0079] Then, the channel attention mechanism is used to weight the input feature map y of the CA attention layer, and the enhanced features are calculated by channel-by-channel multiplication to improve the model's attention to the target area and the feature expression ability. The formula is as follows:

[0080] Y′=Y x ·A x +Y y ·A y ;

[0081] Where, the operation represents channel-by-channel multiplication, which scales each channel of the feature map by weight; Y′ is the output feature of the CA attention layer.

[0082] Finally, the output feature Y′ of the CA attention layer is added to the input feature x1 to obtain the final output feature of the CHGBlock module.

[0083] The CHGBlock module adopts a hierarchical data processing method, which enables the network to gradually learn high-level features from low-level features during the training process, and strengthens the modeling ability of global information with the help of the CA mechanism in the final stage, thereby improving the accuracy of feature expression. In addition, the introduction of the DWConv module not only reduces the spatial dimension of the feature map, but also expands the receptive field. Through the depth-wise separable convolution ( Figure 3 The efficient calculation of the depth convolution in the model further reduces the computational overhead and improves the operating efficiency of the model.

[0084] The final output of the backbone network is: multi-scale high-dimensional feature maps, with decreasing sizes and increasing number of channels.

[0085] Although the C3k2 module in the YOLO11 neck network has achieved significant success in reducing model parameters, its inference speed remains slow. This is primarily due to the fact that traditional convolution (Conv) relies on serial operations for feature extraction, which limits computational efficiency. The Gated Linear Unit (GLU) utilizes two linear projections, one of which regulates information flow through a gating mechanism, and performs feature fusion through element-by-element multiplication. This parallel computing approach effectively improves computational efficiency and reduces model complexity. The Convolutional GLU (CGLU), on the other hand, combines deep convolution with the GLU structure, further improving computational speed while enhancing feature extraction and position information modeling capabilities.

[0086] This embodiment integrates the convolutional feature mixer (ConvFormer) and CGLU structure in the neck C3k2 module of YOLO11 to optimize computational efficiency and enhance feature modeling capabilities.

[0087] like Figure 2 As shown, the neck network in this embodiment includes two upsampling modules, four dimension fusion modules, four CCFC modules and two convolution modules.

[0088] Define two sampling modules as the first and second sampling modules, four dimensional fusion modules as the first, second, third, and fourth dimensional fusion modules, four CCFC modules as the first, second, third, and fourth CCFC modules, and two convolution modules as the first and second convolution modules. The multi-scale high-dimensional feature maps output by the backbone network are input into the neck network and processed as follows:

[0089] The output features of the C2PSA module in the backbone network are input into the first dimension fusion module after passing through the first upsampling module, and are fused with the output features of the fifth CHGBlock module in the first dimension fusion module. The features fused by the first dimension fusion module first pass through the first CCFC module, and then pass through the second upsampling module and are input into the second dimension fusion module, and are fused with the output features of the second CHGBlock module in the second dimension fusion module. The features fused by the second dimension fusion module are processed by the second CCFC module and then input into the first DEGC lightweight detection head.

[0090] The output features of the second CCFC module are input into the third dimension fusion module after passing through the first convolution module, and are fused with the output features of the first CCFC module in the third dimension fusion module. The features fused by the third dimension fusion module are processed by the third CCFC module and then input into the second DEGC lightweight detection head.

[0091] The output features of the third CCFC module are input into the fourth dimension fusion module after passing through the second convolution module, and are fused with the output features of the C2PSA module in the fourth dimension fusion module. The features fused by the fourth dimension fusion module are processed by the fourth CCFC module and then input into the third DEGC lightweight detection head.

[0092] like Figure 6 As shown in Figure 1, the CCFC neck network integrates the ConvFormerCGLU structure to improve the efficiency of multi-scale feature fusion and effectively enhance the feature expression capability of the model. The feature processing process in each CCFC module is as follows:

[0093] Assume the input feature map is First, a standard convolutional layer (Conv) is used to perform preliminary feature extraction, the formula is: x′=Conv(x2); where x′∈R C′×H×W represents the initially extracted features, and C′ is the number of channels after adjustment.

[0094] Then, the initially extracted feature map x′ is split into channels to generate N sub-feature maps x i ′, to further extract key information, the formula is: i ′=Split(x′), i=1, 2,...,N.

[0095] Secondly, each sub-feature map x′ i Pass it to multiple ConvFormer-CGLUBlock modules in sequence for feature transformation to obtain all the transformed sub-features x″′ i , i=1,2,…,N。

[0096] like Figure 7 As shown, the sub-feature graph x′ after splitting i The processing process in the ConvFormer-CGLUBlock module is as follows: sub-feature map x′ i After normalization, the separable convolution layer is input, and then the output features of the separable convolution layer are combined with the sub-feature map x′ i The summed features are normalized and then input into the CGLU module. The output of the CGLU module is then summed with the features after the previous summation to finally obtain the output feature x″′ of the ConvFormer-CGLUBlock module. i , i=1,2,…,N。

[0097] Among them, ConvFormer( Figure 7 The part above the CGLU module in the figure is responsible for long-range dependency modeling and global information fusion. The formula is:

[0098] x″ i =ConvFormer(x′ i ) = Norm(x′ i )+SepConv(Norm(x′ i ));

[0099] In the formula, Norm(x′ i ) represents the sub-feature graph x′ i Normalization is performed, SepConv represents separable convolution processing, x″ i Represents the features processed by ConvFormer, that is, the features after global information fusion through normalization and separable convolution operations.

[0100] like Figure 8 As shown in Figure 2, the CGLU module enhances the feature expression capability through a gating mechanism. The formula is as follows:

[0101] x″′ i =σ(Wg *x″ i )⊙W v *x″ i ;

[0102] Among them, x″′ i is the output feature of the ConvFormer-CGLUBlock module, W g is the weight of the gating mechanism, W v is the eigenvalue mapping weight, σ(·) represents the sigmoid activation function, and ⊙ represents element-wise multiplication (Hadamard product).

[0103] Then, the output sub-feature x″′ of the ConvFormer-CGLUBlock module is i Concatenate (Concat) on the channel dimension to obtain the concatenated feature y C , the formula is: C =Concat(x″′1,x″′2,…,x″′ N ); where x″′1, x″′2, …, x″′ N are all sub-features after transformation by the ConvFormer-CGLUBlock module.

[0104] Finally, the concatenated features y C Channel integration is performed through a 1×1 convolution (Conv) to ensure the integrity and consistency of the output features. The formula is: y′ C =Conv(y C );where The final output feature of the CCFC module.

[0105] This example replaces the C3k2 module in the original YOLO11 neck network with the CCFC module, which reduces computational complexity while improving the model's feature representation capabilities. The neck network ultimately outputs a fused multi-scale feature map, which enhances background suppression and object recognition.

[0106] YOLO11 uses a decoupled head design and an anchor-free mechanism for object detection. Its detection head consists of two independent branches: the upper branch predicts the target box coordinates through 3×3 and 1×1 convolutions and is used to calculate the regression loss; the lower branch, consisting of depthwise separable convolution (DWConv) and standard convolution, is responsible for object classification and calculating the classification loss. However, when there are more feature layers, stacking multiple 3×3 convolutional layers significantly increases the number of parameters and computational overhead (FLOPs) of the detection head.

[0107] like Figure 9As shown in the figure, the computational complexity of the YOLO11n detection head reaches 1.870GFLOPs, accounting for 28.8% of the total computational complexity of the entire model. The computational complexity of the detection part is close to half of the total FLOPs of the model, becoming a bottleneck limiting the lightweight optimization of the model. To solve this problem, this embodiment proposes a shared parameter strategy based on the idea of ​​the Retina-Net detection head, which optimizes computational efficiency and model complexity while maintaining recognition accuracy. Figure 10 As shown in the figure. The specific strategy is to replace the standard convolutional layer and depthwise convolutional layer of the two branches in the detection head with two 3×3 grouped convolutions, which are used for classification and regression tasks respectively. In this embodiment, the detection head network includes three DEGC lightweight detection heads: the first, second, and third DEGC lightweight detection heads. Each DEGC lightweight detection head uses grouped convolution to replace traditional standard convolution and depthwise separable convolution.

[0108] like Figure 10 As shown, the number of input channels of each DEGC lightweight detection head is defined as ch[i]. The features in each DEGC lightweight detection head first pass through the stem structure, which contains two 3×3 group convolutions in series. The number of groups is It is used to reduce the amount of calculation and extract features at the same time. The formula is:

[0109]

[0110] Where X′ is the output feature of the stem structure, and X is the input feature of the DEGC lightweight detection head.

[0111] Then, the feature X′ is divided into two branches, each of which undergoes a 1×1 point-wise convolution ( Figure 10 The two branches output bounding box parameters and category scores respectively, and the formula is:

[0112] B=Conv(X′,1×1,4×reg max );

[0113] C = Conv(X′,1×1,nc);

[0114] Among them, reg max =16 represents the number of distribution fitting parameters for each regression channel, nc represents the number of categories, B represents the bounding box parameters, and C represents the category score.

[0115] The outputs of the CCFC modules of three different dimensions are respectively input into the classification and regression branches of the corresponding DEGC lightweight detection head. The classification branch outputs the category probability, and the regression branch outputs the bounding box coordinates. Finally, the final prediction result is obtained through NMS post-processing.

[0116] The computational complexity of grouped convolution is reduced by g times compared to standard convolution, while maintaining the ability to extract local features, which greatly reduces the amount of calculation. The formula is:

[0117]

[0118] Where FLOPs group is the computational complexity of group convolution, FLOPs standard is the computational complexity of standard convolution, and g is the number of groups.

[0119] Therefore, grouped convolution can reduce the amount of computation while maintaining the local feature extraction capability of convolution, thereby ensuring that detection performance is not affected while reducing computational overhead. Furthermore, the optimized detection head parameter count is reduced to 0.19M, and the computational complexity is reduced to 0.62GFLOPs, which are only 42.1% and 33.2% of the original YOLO11n structure, respectively. This improvement not only effectively reduces the parameters and FLOPs of the detection head, but also significantly improves the recognition accuracy of the model.

[0120] This embodiment innovatively designs the CHGNetV2 backbone network, organically combining multi-scale convolution with a coordinate attention mechanism. By stacking 3×3 convolutions and 1×1 channel compression and restoration operations in parallel, coupled with spatial coordinate attention weighting, this significantly reduces model parameters and computational overhead while effectively enhancing the feature representation and localization capabilities of the target region. This particularly improves the recognition accuracy and robustness of small fruit objects. A CCFC module, which fuses the ConvFormer and CGLU structures, is introduced into the neck network. This innovative combination of long-range dependency modeling and gated nonlinear activation strengthens the representation of multi-scale features and global information fusion, significantly improving the recognition stability and accuracy of multiple fruit categories in complex backgrounds while reducing computational complexity. Furthermore, a shared parameter strategy based on grouped convolutions is employed in the detection head design, replacing traditional standard and depthwise separable convolutions. This significantly reduces the number of parameters and computational overhead of the detection head. While maintaining detection accuracy, this significantly improves model inference speed, adapting to the needs of multi-scale object detection and achieving a good balance between accuracy and efficiency. This example effectively alleviates bottlenecks in existing technologies, such as redundant model parameters, high computational overhead, and insufficient small-target detection performance, through the coordinated design and parameter optimization of the backbone network, neck network, and detection head network. This enables the CHCG-YOLO11 fruit recognition model in this example to achieve lightweight model size and real-time inference while maintaining high recognition accuracy, meeting the application requirements of smart agriculture and embedded systems.

[0121] Step 3: Based on the data set of step 1, the CHCG-YOLO11 fruit recognition model built in step 2 is trained and optimized, and then the trained and optimized CHCG-YOLO11 fruit recognition model is used to identify the fruit category and location.

[0122] The lightweight fruit recognition method based on CHCG-YOLO11 in this embodiment uniformly scales the collected RGB fruit images to 640×640 pixels, normalizes them, and feeds them into the CHGNetV2 backbone network as model input. Within the backbone network, preliminary feature extraction is performed using the HGStem standard convolution (HGStem module). Then, through four stages, the CHGBlock module is used to perform multi-scale convolution feature extraction. At the same time, the coordinate attention mechanism is used to focus on key target areas, enhance target features, and reduce background interference. The multi-scale features output by the backbone network are then passed to the CCFC module, which first performs channel splitting, then uses a multi-layer ConvFormer long-range dependency modeling to capture global information, and then uses the CGLU module to strengthen nonlinear expression. The final features are fused and output in the channel dimension. The fused features are input to the DEGC lightweight detection head, which first uses grouped convolution to reduce computational overhead, and then generates bounding boxes and category predictions at each scale through point-by-point convolution. Finally, the multi-scale detection results are merged, and overlapping boxes are removed through post-processing such as NMS (non-maximum suppression) to output the final fruit category and location.

[0123] Example 2

[0124] This Example 2 describes a lightweight fruit recognition system based on CHCG-YOLO11. The system includes an image acquisition device and a computer. The image acquisition device is used to capture fruit images and upload them to the computer. The computer includes a memory and one or more processors. The memory stores executable code. When the processor executes the executable code, it implements the steps of the lightweight fruit recognition method based on CHCG-YOLO11 described in Example 1.

[0125] The embodiments of the present invention are only used to illustrate the technical solutions of the present invention rather than to limit the present invention. Those skilled in the art will appreciate that various changes, modifications, substitutions and variations may be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A lightweight fruit recognition method based on CHCG-YOLO11, characterized in that: The steps include: Step 1: Obtain fruit images, construct a dataset containing fruit images, and preprocess the dataset; Step 2: Build the CHCG-YOLO11 fruit recognition model, which includes the backbone network, neck network, and detection head network; The backbone network includes the HGStem module, a multi-stacked CHGBlock structure, an SPPF module, and a C2PSA module. Each CHGBlock module introduces the CA mechanism on the basis of HGBlock. The neck network uses the CCFC module to replace the C3k2 module in the original YOLO11 neck network. The CCFC module integrates the ConvFormer and CGLU structures in the C3k2 module. The detection head network includes the first, second, and third DEGC lightweight detection heads. Each DEGC lightweight detection head uses grouped convolution to replace traditional standard convolution and depth-separable convolution. The processing of the preprocessed data in the CHCG-YOLO11 fruit recognition model is as follows: The preprocessed data first passes through the backbone network, where it first passes through the HGStem module, then through a multi-stacked CHGBlock structure to improve feature extraction capabilities, and finally through the SPPF module and the C2PSA module. The multi-scale high-dimensional feature map output by the backbone network passes through the neck network and is input into the detection head network. Finally, the output module receives the output of the detection head network, merges the detection results at different scales, performs non-maximum suppression to remove overlapping boxes, and finally outputs the category and location of the fruit. Step 3: Based on the data set of step 1, the CHCG-YOLO11 fruit recognition model built in step 2 is trained and optimized, and then the trained and optimized CHCG-YOLO11 fruit recognition model is used to identify fruit categories.

2. The lightweight fruit recognition method based on CHCG-YOLO11 according to claim 1, wherein The multi-stacked CHGBlock structure includes six CHGBlock modules and three depth convolution modules; Define six CHGBlock modules as the first, second, third, fourth, fifth and sixth CHGBlock modules, and define four depth convolution modules as the first, second and third depth convolution modules; The feature input is placed in a multi-stacked CHGBlock structure and processed in sequence by the first CHGBlock module, the first depth convolution module, the second CHGBlock module, the second depth convolution module, the third CHGBlock module, the fourth CHGBlock module, the fifth CHGBlock module, the third depth convolution module, and the sixth CHGBlock module.

3. The lightweight fruit recognition method based on CHCG-YOLO11 according to claim 1 is characterized in that, The features are processed in each CHGBlock module as follows: The feature first undergoes multiple convolution transformations on the input x1 through multiple stacked 3×3 standard convolutions to generate multiple intermediate feature maps {y1,y2,…,y n }, and then perform channel compression through a 1×1 compressed convolution, the formula is: and s =σ(W s ·Concat(x1,y1,y2,…,y n )); Among them, W s Represents the weight parameter matrix of compressed convolution, y s is the output feature of 1×1 compressed convolution, σ is the nonlinear activation function Sigmoid, and Concat represents channel compression through 1×1 compressed convolution; Then restore the channel dimension through 1×1 excitation convolution to obtain the fused feature map y. The formula is as follows: y=σ(W e ·y s ); Among them, W e Represents the weight parameter matrix of the excitation convolution; Then, the fused feature map y is input into the CA attention layer to obtain the output feature Y′ of the CA attention layer; Finally, the output feature Y′ of the CA attention layer is added to the input feature x1 to obtain the final output feature of the CHGBlock module.

4. The lightweight fruit recognition method based on CHCG-YOLO11 according to claim 3 is characterized in that, The processing of the fused feature map y in the CA attention layer is as follows: For the input feature map y∈R C×H×W , first calculate the adaptive average pooling features in the X and Y directions, the formula is: Among them, H represents the height of the feature map, W represents the width of the feature map, i represents the index in the height direction, j represents the index in the width direction, and Y x represents the adaptive average pooling feature along the X direction, Y y Represents the adaptive average pooling feature along the Y direction, and y(i,j) represents the fused feature map y∈R C×H×W ; Then use 1×1 convolution, BatchNorm normalization and h-swish activation function to fuse channel attention information to obtain the attention weight A x and A y , the formula is as follows: A x =σ(W x ×Y x ); A y =σ(W y ×Y y ); Among them, W x and W y are the 1×1 convolution weights in the X and Y directions respectively, and σ is the nonlinear activation function Sigmoid; Then, the input feature map y is weighted using the channel attention mechanism, and the enhanced features are calculated by channel-by-channel multiplication, as follows: Y′=Y x ·TO x +Y y ·TO y ; Where, the operation represents channel-by-channel multiplication, which scales each channel of the feature map by weight; Y′ is the output feature of the CA attention layer.

5. The lightweight fruit recognition method based on CHCG-YOLO11 according to claim 2 is characterized in that, The neck network includes two upsampling modules, four dimension fusion modules, four CCFC modules and two convolution modules; Define two sampling modules as the first and second sampling modules respectively; define four dimension fusion modules as the first, second, third and fourth dimension fusion modules respectively; define four CCFC modules as the first, second, third and fourth CCFC modules respectively; define two convolution modules as the first and second convolution modules respectively; The output features of the C2PSA module are input into the first dimension fusion module after passing through the first upsampling module, and are fused with the output features of the fifth CHGBlock module in the first dimension fusion module. The features fused by the first dimension fusion module are first passed through the first CCFC module, and then through the second upsampling module and input into the second dimension fusion module, and are fused with the output features of the second CHGBlock module in the second dimension fusion module. The features fused by the second dimension fusion module are processed by the second CCFC module and then input into the first DEGC lightweight detection head; The output features of the second CCFC module are input into the third dimension fusion module after passing through the first convolution module, and are fused with the output features of the first CCFC module in the third dimension fusion module. The fused features of the third dimension fusion module are processed by the third CCFC module and then input into the second DEGC lightweight detection head; The output features of the third CCFC module are input into the fourth dimension fusion module after passing through the second convolution module, and are fused with the output features of the C2PSA module in the fourth dimension fusion module. The features fused by the fourth dimension fusion module are processed by the fourth CCFC module and then input into the third DEGC lightweight detection head.

6. The lightweight fruit recognition method based on CHCG-YOLO11 according to claim 1 is characterized in that, The features are processed in each CCFC module as follows: Assume the input feature map is First, a standard convolutional layer is used to perform preliminary feature extraction, the formula is: x′=Conv(x2); where x′∈R C′×H×W represents the initially extracted features, C′ is the number of channels after adjustment; Then, the initially extracted feature map x′ is split into channels to generate N sub-feature maps x′ i , to further extract key information, the formula is: x′ i =Split(x′), i=1, 2,...,N; Secondly, each sub-feature map x′ i Pass it to multiple ConvFormer-CGLUBlock modules in sequence for feature transformation to obtain all the transformed sub-features x″′ i , i=1,2,…,N; Then, the output sub-feature x″′ of the ConvFormer-CGLUBlock module is i Splicing is performed on the channel dimension to obtain the spliced ​​feature y C , the formula is: C =Concat(x″′1,x″′2,…,x″′ N ); where x″′1, x″′2, …, x″′ N are all sub-features after transformation by the ConvFormer-CGLUBlock module; Finally, the concatenated features y C Channel integration is performed through a 1×1 convolution, the formula is: y′ C =Conv(y C );where The final output feature of the CCFC module.

7. The lightweight fruit recognition method based on CHCG-YOLO11 according to claim 6 is characterized in that, The sub-feature graph x′ after splitting i The processing within the ConvFormer-CGLUBlock module is as follows: Sub-feature graph x′ i After normalization, the separable convolution layer is input, and then the output features of the separable convolution layer are combined with the sub-feature map x′ i The summed features are normalized and then input into the CGLU module. The output of the CGLU module is then summed with the features after the previous summation to finally obtain the output feature x″′ of the ConvFormer-CGLUBlock module. i , i=1,2,…,N; Among them, ConvFormer is responsible for long-range dependency modeling and global information fusion, the formula is: x″ i =ConvFormer(x′ i )=Norm(x′ i )+SepConv(Norm(x′ i )); In the formula, Norm(x′ i ) represents the sub-feature graph x′ i Normalization is performed, SepConv represents separable convolution processing, x′ i Represents the features processed by ConvFormer, that is, the features after global information fusion through normalization and separable convolution operations.

8. The lightweight fruit recognition method based on CHCG-YOLO11 according to claim 7 is characterized in that, The CGLU module enhances the feature expression capability through a gating mechanism, and the formula is: x″′ i =σ(W g *x″ i )⊙W v *x″ i ; Among them, x″′ i is the output feature of the ConvFormer-CGLUBlock module, W g is the weight of the gating mechanism, W v is the eigenvalue mapping weight, σ(·) represents the sigmoid activation function, and ⊙ represents element-wise multiplication.

9. The lightweight fruit recognition method based on CHCG-YOLO11 according to claim 1 is characterized in that, The feature processing process in the detection head network is as follows: The number of input channels of each DEGC lightweight detection head is defined as ch[i]. The features first pass through the stem structure in each DEGC lightweight detection head. The stem structure contains two serial 3×3 group convolutions with the number of groups being It is used to reduce the amount of calculation and extract features at the same time. The formula is: Where X′ is the output feature of the stem structure, and X is the input feature of the DEGC lightweight detection head; Then, the feature X′ is divided into two branches, each branch undergoes a 1×1 point-by-point convolution, and the two branches output the bounding box parameters and category scores respectively, as follows: B=Conv(X′,1×1,4×reg max ); C = Conv(X′,1×1,nc); Among them, reg max =16 represents the number of distribution fitting parameters for each regression channel, nc represents the number of categories, B represents the bounding box parameters, and C represents the category score.

10. A lightweight fruit recognition system based on CHCG-YOLO11, including image acquisition equipment and computer equipment; The image acquisition device is used to capture images of fruits and upload them to a computer device; Computer equipment includes memory and one or more processors; An executable code is stored in the memory; it is characterized in that when the processor executes the executable code, it is used to implement the steps of the lightweight fruit recognition method based on CHCG-YOLO11 according to any one of claims 1 to 9.