A method, system, electronic device, and storage medium for visually detecting impurities in tea leaves.

By enhancing the features of minute impurities using the ScharrStem module and the SGCB improvement module, and combining these with the channel sensing spatial modulation module for feature fusion, the problem of low accuracy in tea impurity detection was solved, achieving high-precision tea impurity detection.

CN121329970BActive Publication Date: 2026-03-10CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies suffer from low detection accuracy in tea impurity detection, mainly due to the high similarity between impurities and the tea background color and texture, the large variation in impurity size, and the lack of information, making it difficult to distinguish features and the model's insufficient ability to perceive impurities at multiple scales.

Method used

The ScharrStem module is used to enhance the contour features of small impurities, the SGCB module is used to prevent semantic decay, and the channel-aware spatial modulation module is used to achieve adaptive fusion of features at different levels, thereby improving the feature discrimination capability of multi-scale impurities.

Benefits of technology

It significantly improves the accuracy of tea impurity detection. Through multi-scale feature extraction and adaptive fusion, it ensures sensitivity and positioning accuracy for minute impurities, achieving high-quality detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121329970B_ABST
    Figure CN121329970B_ABST
Patent Text Reader

Abstract

This application discloses a visual detection method, system, electronic device, and storage medium for tea impurities. The method acquires an image of the tea to be detected and a training image dataset, preprocesses the image data in the training image dataset, and inputs the image of the tea to be detected into a trained tea impurity detection model. This yields multiple tea impurity detection results containing confidence scores and detection boxes for the target tea. The trained tea impurity detection model includes a backbone network, a neck network, and multiple detection heads. The backbone network includes a ScharrStem module, a first improved SGCB module, and a second improved SGCB module. The neck network includes a first improved SGCB module, a second improved SGCB module, and a channel-aware spatial modulation module. Based on the confidence scores and detection boxes in the multiple tea impurity detection results, the detection result for the target tea impurity is determined. This application can improve the detection accuracy of tea impurities.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of visual detection, in particular to a tea impurity visual detection method and system, electronic equipment and storage medium. BACKGROUND

[0002] With the rapid development of agricultural technology, intelligent and automated technology is gradually integrated into the modern agricultural production system, promoting the transformation of traditional agriculture to smart agriculture. As an important economic crop, tea has a wide range of consumer demand in the domestic and international market, and its quality control directly affects consumer experience and industry competitiveness. In the tea production process, impurity detection is a key link, which aims to identify and remove non-tea foreign matter such as straw, wax leaves and old stems, in order to protect product quality.

[0003] The advent of deep learning has brought revolutionary changes to the field of object detection, and the YOLO series algorithm is widely adopted in agricultural applications due to its outstanding speed and accuracy. However, directly using existing models for tea impurity detection still faces significant challenges. First, there is often a high degree of color and texture similarity between impurities and tea background, making it difficult to distinguish features. Second, impurity target size varies greatly, especially small impurities, which lack information in the feature map, making it easy to miss detection. Third, the feature fusion mechanism of existing models often fails to fully consider the complementarity of different levels of features in spatial details and semantic information, limiting the model's ability to perceive multi-scale, occluded or irregularly shaped impurities. Therefore, the detection accuracy of existing technology for tea impurities is relatively low. SUMMARY

[0004] The present application aims to provide a tea impurity visual detection method, system, electronic equipment and storage medium, which can improve the detection accuracy of tea impurities.

[0005] In a first aspect, the present application provides a tea impurity visual detection method, which comprises:

[0006] Obtaining a to-be-detected tea image and a training image data set, pre-processing the image data in the training image data set to obtain a pre-processed training image data set, and the to-be-detected tea image contains a plurality of to-be-detected targets;

[0007] input the tea leaf image to be detected into the trained tea leaf impurity detection model to obtain a plurality of tea leaf impurity detection results containing confidence and detection boxes of the target to be detected, the trained tea leaf impurity detection model is trained by the preprocessed training image data set, the trained tea leaf impurity detection model comprises a backbone network, a neck network and a plurality of detection heads, the backbone network comprises a ScharrStem module, a first SGCB improvement module and a second SGCB improvement module, the neck network comprises a first SGCB improvement module, a second SGCB improvement module and a channel perception spatial modulation module, wherein,

[0008] feature extraction is performed on the tea leaf image to be detected through the ScharrStem module to obtain a feature extraction result;

[0009] multi-scale feature extraction is performed on the feature extraction result through a plurality of stages containing the first SGCB improvement module or the second SGCB improvement module to obtain a multi-scale feature map;

[0010] the multi-scale feature map is processed through the first SGCB improvement module, the second SGCB improvement module and the channel perception spatial modulation module to obtain a plurality of target feature fusion results;

[0011] each target feature fusion result is detected by the plurality of detection heads to obtain a plurality of tea leaf impurity detection results;

[0012] a target tea leaf impurity detection result is determined according to the confidence and the detection box in the plurality of tea leaf impurity detection results.

[0013] Compared with the prior art, the first aspect of the present application has the following beneficial effects:

[0014] The method can strengthen the contour feature expression of small impurities through the ScharrStem module, provide a high-quality detail feature base for subsequent deep processing, effectively prevent semantic attenuation of small impurity features in multiple convolution transformations through the SGCB improvement module, and ensure that the sensitivity to fine targets can be maintained in deep networks, and adaptive fusion of different levels of features can be realized through the channel perception spatial modulation module, which significantly improves the feature discrimination ability and positioning accuracy of the model for multi-scale impurities. Therefore, the tea leaf impurity detection result is obtained through the trained tea leaf impurity detection model, which can improve the accuracy of the detection result, and the target tea leaf impurity detection result is determined according to the confidence and the detection box in the plurality of tea leaf impurity detection results, which further improves the detection accuracy of tea leaf impurities.

[0015] In a second aspect, the embodiments of the present application also provide a tea leaf impurity visual detection system, the system comprises:

[0016] a data acquisition unit, configured to acquire a to-be-detected tea image and a training image dataset, pre-process image data in the training image dataset to obtain a pre-processed training image dataset, and the to-be-detected tea image contains a plurality of to-be-detected targets;

[0017] a detection result obtaining unit, configured to input the to-be-detected tea image into a trained tea impurity detection model to obtain a plurality of tea impurity detection results containing confidence and detection boxes of the to-be-detected targets, the trained tea impurity detection model is obtained by training the pre-processed training image dataset, and the trained tea impurity detection model includes a backbone network, a neck network, and a plurality of detection heads, the backbone network includes a ScharrStem module, a first SGCB improvement module, and a second SGCB improvement module, the neck network includes the first SGCB improvement module, the second SGCB improvement module, and a channel-aware spatial modulation module, wherein

[0018] the ScharrStem module is configured to perform feature extraction on the to-be-detected tea image to obtain a feature extraction result;

[0019] a plurality of stages containing the first SGCB improvement module or the second SGCB improvement module are configured to perform multi-scale feature extraction on the feature extraction result to obtain a multi-scale feature map;

[0020] the multi-scale feature map is processed through the first SGCB improvement module, the second SGCB improvement module, and the channel-aware spatial modulation module to obtain a plurality of target feature fusion results;

[0021] the plurality of detection heads are configured to respectively detect each of the target feature fusion results to obtain a plurality of tea impurity detection results;

[0022] a target result obtaining unit, configured to determine a target tea impurity detection result according to the confidence and the detection boxes in the plurality of tea impurity detection results.

[0023] In a third aspect, an electronic device is provided, including at least one control processor and a memory in communication connection with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to perform the tea impurity visual detection method as described above.

[0024] In a fourth aspect, the embodiments of the present application further provide a computer readable storage medium, which stores computer executable instructions for causing a computer to execute the tea impurity visual detection method.

[0025] It can be understood that the beneficial effects of the second aspect to the fourth aspect compared with the related art are the same as the beneficial effects of the first aspect compared with the related art, which can be seen from the related description in the first aspect and will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0026] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the following description, including the appended drawings, wherein:

[0027] Figure 1 is a flowchart of an embodiment of the tea impurity visual detection method provided by the present application;

[0028] Figure 2 is a structural schematic diagram of the TSCM-Det model in the best embodiment of the tea impurity visual detection method provided by the present application;

[0029] Figure 3 is a structural schematic diagram of the ScharrStem module in the best embodiment of the tea impurity visual detection method provided by the present application;

[0030] Figure 4 is a structural schematic diagram of the SGCB improvement module in the best embodiment of the tea impurity visual detection method provided by the present application;

[0031] Figure 5 is a structural schematic diagram of the channel perception spatial modulation module in the best embodiment of the tea impurity visual detection method provided by the present application;

[0032] Figure 6 is a structural schematic diagram of an embodiment of the tea impurity visual detection system provided by the present application;

[0033] Figure 7 is a structural schematic diagram of an embodiment of the electronic device provided by the present application. DETAILED DESCRIPTION

[0034] Embodiments of the present application are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference signs represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as a limitation of the present application.

[0035] In the description of the present application, if the first, second, etc. are described, it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features or the order of the indicated technical features.

[0036] In the description of the present application, it should be understood that the orientation description, such as the orientation or position relationship indicated by up, down, etc. is based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present application and simplifying the description, and therefore cannot be understood as indicating or implying that the indicated device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present application.

[0037] In the description of the present application, it should be noted that, unless otherwise explicitly limited, the words such as setting, installing, connecting, etc. should be broadly understood, and the person skilled in the art can reasonably determine the specific meaning of the above words in the present application in combination with the specific content of the technical solution.

[0038] Since the detection accuracy of the existing technology for tea impurities is relatively low, in order to solve the problems existing in the prior art, the present application provides a tea impurity visual detection method, system, electronic device and storage medium.

[0039] Referring to Figure 1 , the flowchart of the tea impurity visual detection method provided by the embodiments of the present application. The tea impurity visual detection method is applied to an electronic device, which can be a server or a mobile terminal, etc. As Figure 1 shown, the tea impurity visual detection method can include the following steps:

[0040] Step S101, acquiring a to-be-detected tea image and a training image data set, pre-processing the image data in the training image data set to obtain a pre-processed training image data set, and the to-be-detected tea image containing a plurality of to-be-detected targets;

[0041] Step S102, inputting the to-be-detected tea image into a trained tea impurity detection model to obtain a plurality of tea impurity detection results containing confidence and detection boxes of to-be-detected targets, the trained tea impurity detection model being obtained by training the pre-processed training image data set, the trained tea impurity detection model including a backbone network, a neck network and a plurality of detection heads, the backbone network including a ScharrStem module, a first SGCB improvement module and a second SGCB improvement module, the neck network including the first SGCB improvement module, the second SGCB improvement module and a channel perception spatial modulation module, wherein,

[0042] extracting features of the to-be-detected tea image through the ScharrStem module to obtain a feature extraction result;

[0043] The multi-scale feature extraction result is obtained by multi-scale feature extraction on the feature extraction result through multiple stages including the first SGCB improvement module or the second SGCB improvement module;

[0044] The multi-scale feature map is processed through the first SGCB improvement module, the second SGCB improvement module and the channel perception spatial modulation module to obtain multiple target feature fusion results;

[0045] The multiple detection heads are used to detect each target feature fusion result respectively to obtain multiple tea impurity detection results;

[0046] In step S103, the target tea impurity detection result is determined according to the confidence and the detection frame in the multiple tea impurity detection results.

[0047] In the embodiment, the ScharrStem module can strengthen the contour feature expression of the micro impurities, providing a high-quality detailed feature base for subsequent deep processing; the SGCB improvement module can effectively prevent semantic attenuation of the micro impurity features in multiple convolution transformations, ensuring that the sensitivity to fine targets can be maintained in the deep network; and the channel perception spatial modulation module can realize adaptive fusion of different hierarchical features, significantly improving the feature discrimination ability and positioning accuracy of the model for multi-scale impurities. Therefore, the tea impurity detection result obtained by the trained tea impurity detection model can improve the accuracy of the detection result, and the target tea impurity detection result is determined according to the confidence and the detection frame in the multiple tea impurity detection results, further improving the detection accuracy of the tea impurities.

[0048] The image data in the training image data set can be preprocessed by rotation, cropping, illumination adjustment, scaling and noise addition.

[0049] The ScharrStem module can combine the Scharr operator in traditional image processing with deep learning feature extraction to establish an edge enhancement mechanism in the initial stage of feature extraction.

[0050] The SGCB improvement module can be a module obtained by integrating the SGCB sub-module into the original feature extraction unit C3K2 sub-module of YOLOv11.

[0051] The channel perception spatial modulation module can be a feature fusion module combining channel attention and spatial attention.

[0052] The target tea impurity detection result can be determined according to the confidence and the detection frame in the plurality of tea impurity detection results by using an empty scale perception soft non-maximum suppression algorithm to remove redundant frames.

[0053] In some embodiments, feature extraction is performed on the to-be-detected tea image by a Scharr Stem module to obtain a feature extraction result, including:

[0054] The to-be-detected tea image is subjected to a convolution operation to obtain an initial feature map;

[0055] The initial feature map is subjected to Scharr filtering to obtain an edge-enhanced feature map;

[0056] The edge-enhanced feature map is fused with the to-be-detected tea image to obtain a first fusion result;

[0057] The first fusion result is subjected to two convolution operations to obtain an intermediate feature representation;

[0058] The intermediate feature representation is subjected to Gaussian smoothing processing to obtain a smoothing result;

[0059] The smoothing result is fused with the intermediate feature representation to obtain a second fusion result;

[0060] The second fusion result is subjected to a convolution operation to obtain the feature extraction result.

[0061] In this embodiment, by combining the Scharr operator in traditional image processing with deep learning feature extraction, an edge enhancement mechanism is established in the initial stage of feature extraction, effectively strengthening the contour feature expression of micro impurities and providing a high-quality detailed feature base for subsequent deep processing.

[0062] The Scharr filtering of the initial feature map can be Scharr filtering of the initial feature map by using a Scharr operator.

[0063] The Gaussian smoothing processing of the intermediate feature representation can be Gaussian smoothing processing of the intermediate feature representation by using a 5x5 Gaussian filter.

[0064] In some embodiments, multi-scale feature extraction is performed on the feature extraction result by a plurality of stages containing the first SGCB improvement module or the second SGCB improvement module to obtain a multi-scale feature map, including:

[0065] The feature extraction result is input to a first stage containing the first SGCB improvement module for feature extraction to obtain a first-scale feature map;

[0066] The first scale feature map is down-sampled, batch normalized and processed by a SiLU activation function to obtain a first processing result;

[0067] The first processing result is input into a second stage containing a first SGCB improved module for feature extraction to obtain a second scale feature map;

[0068] The second scale feature map is down-sampled, batch normalized and processed by a SiLU activation function to obtain a second processing result;

[0069] The second processing result is input into a third stage containing a second SGCB improved module for feature extraction to obtain a third scale feature map;

[0070] The third scale feature map is down-sampled, batch normalized and processed by a SiLU activation function to obtain a third processing result;

[0071] The third processing result is input into a fourth stage containing a second SGCB improved module for feature extraction to obtain a fourth scale feature map.

[0072] In the embodiment, the feature extraction results are subjected to multi-scale feature extraction through multiple stages containing the first SGCB improved module or the second SGCB improved module, so that feature maps with different semantics and details in each stage are obtained, thereby laying a good data foundation for later tea impurity detection.

[0073] In some embodiments, the first SGCB improved module and the second SGCB improved module include an SGCB submodule and a C3K2 submodule, the SGCB submodule includes multiple convolution layers and an improved Scharr convolution layer; the SGCB submodule is added to the C3K2 submodule to obtain the SGCB improved module, when the C3K2 submodule in the SGCB improved module does not use the C3k submodule for feature enhancement, the SGCB improved module is the first SGCB improved module; when the C3K2 submodule in the SGCB improved module uses the C3k submodule for feature enhancement, the SGCB improved module is the second SGCB improved module; the improved Scharr convolution layer includes the following processing process:

[0074] The input feature map is subjected to a channel-by-channel Scharr gradient calculation to obtain an x-direction gradient map and a y-direction gradient map of each channel;

[0075] The x-direction gradient map and the y-direction gradient map are fused to obtain an original edge response map of each channel;

[0076] The input feature map is subjected to a channel-by-channel 3x3 neighborhood variance calculation and normalization processing to obtain a normalized variance map;

[0077] According to the normalized variance map, the weight of each channel is calculated;

[0078] The original edge response map corresponding to each channel is pixel-wise weighted by the weight of the channel to obtain an enhanced edge feature map of each channel;

[0079] The enhanced edge feature map of each channel is integrated in the original channel order to obtain a final output feature map.

[0080] In the embodiment, the SGCB sub-module includes multiple convolution layers and an improved Scharr convolution layer. By constructing a parallel processing path in the deep network, the accuracy of edge detection is preserved, the continuous preservation and enhancement of detailed features are realized, the semantic decay of small impurity features in multiple convolution transformations is effectively prevented, and the sensitivity to fine targets in the deep network is ensured.

[0081] In some embodiments, the multi-scale feature maps are processed by a first SGCB improvement module, a second SGCB improvement module, and a channel-aware spatial modulation module to obtain a plurality of target feature fusion results, including:

[0082] The fourth scale feature map is subjected to feature extraction to obtain a first feature image, and the first feature image is subjected to an upsampling operation to obtain an upsampled first feature image; the third scale feature map and the upsampled first feature image are subjected to dimension stacking to obtain a first stacking result, and the first stacking result is input into the first SGCB improvement module for feature fusion to obtain a feature fusion result;

[0083] The feature fusion result is subjected to an upsampling operation to obtain an upsampled second feature image, and the second scale feature map and the upsampled second feature image are subjected to dimension stacking to obtain a second stacking result;

[0084] The first scale feature map is subjected to a downsampling operation to obtain a downsampled third feature image, and the downsampled third feature image and the second stacking result are input into the channel-aware spatial modulation module for feature fusion to obtain a first enhanced feature map; the first enhanced feature map is input into the first SGCB improvement module for feature fusion to obtain a first target feature fusion result;

[0085] The first enhanced feature map is subjected to a downsampling operation to obtain a downsampled fourth feature image, and the downsampled fourth feature image and the feature fusion result are input into the channel-aware spatial modulation module for feature fusion to obtain a second enhanced feature map; the second enhanced feature map is input into the first SGCB improvement module for feature fusion to obtain a second target feature fusion result;

[0086] The second target feature fusion result is down-sampled to obtain a down-sampled fifth feature image, and the down-sampled fifth feature image and the first feature image are input into the channel perception spatial modulation module for feature fusion to obtain a third enhanced feature map; and the third enhanced feature map is input into the second SGCB improved module for feature fusion to obtain a third target feature fusion result.

[0087] In the embodiment, the SGCB improved module and the channel perception spatial modulation module are used for processing, which can effectively prevent semantic attenuation of micro impurity features in multiple convolutional transformations, ensure that the sensitivity to fine targets can be maintained in a deep network, and significantly improve the feature discrimination ability and positioning accuracy of the model for multi-scale impurities.

[0088] The plurality of target feature fusion results can be target feature fusion results corresponding to high-scale feature maps. For example, the embodiment includes a first-scale feature map, a second-scale feature map, a third-scale feature map, and a fourth-scale feature map, and the plurality of target feature fusion results can include one target feature fusion result corresponding to each of the second-scale feature map, the third-scale feature map, and the fourth-scale feature map, i.e., target feature fusion results corresponding to high-scale feature maps.

[0089] In some embodiments, the down-sampled third feature image and the second superimposed result are input into the channel perception spatial modulation module for feature fusion to obtain a first enhanced feature map, including:

[0090] The down-sampled third feature image is used as a low-level feature, and the second superimposed result is used as a high-level feature.

[0091] The low-level feature and the high-level feature are added element by element to obtain a first addition result.

[0092] The first addition result is input into the channel attention module for global average pooling and convolution operation to obtain a first convolution feature.

[0093] The first convolution feature is processed by an activation function and a convolution operation to obtain a channel attention weight.

[0094] The first addition result is maximum-pooled to obtain a first pooling result, and the first addition result is average-pooled to obtain a second pooling result.

[0095] The first pooling result and the second pooling result are spliced and then subjected to convolution operation to obtain a second convolution feature.

[0096] The first convolution feature is processed by convolution and an activation function to obtain a third convolution feature.

[0097] The third convolutional feature is expanded through a broadcast mechanism and multiplied with the second convolutional feature to obtain a collaborative spatial attention weight;

[0098] The collaborative spatial attention weight and the channel attention weight are element-level added to obtain a second addition result, and the first addition result and the second addition result are spliced to obtain a spliced result;

[0099] The spliced result is grouped and channel-arranged to obtain a fusion weight, and the fusion weight is convolved and activated to obtain a spatial modulation weight;

[0100] According to the spatial modulation weight, the low-level feature and the high-level feature, a weighted feature is determined;

[0101] The weighted feature, the low-level feature and the high-level feature are element-level added to obtain a third addition result;

[0102] The third addition result is subjected to a convolution operation to obtain a first enhanced feature map.

[0103] In the embodiment, through processing by the channel-aware spatial modulation module, adaptive fusion of different level features can be realized, and the feature discrimination ability and positioning accuracy of the model for multi-scale impurities are significantly improved.

[0104] In some embodiments, according to the confidence and the detection box in the plurality of tea impurity detection results, a target tea impurity detection result is determined, comprising:

[0105] According to the detection boxes in the plurality of tea impurity detection results, a spatial scale-aware overlap area ratio between two detection boxes is determined;

[0106] Based on the spatial scale-aware overlap area ratio and the confidence, a Gaussian confidence decay function is constructed, and the Gaussian confidence decay function is used to update the confidence;

[0107] Based on the detection box, the Gaussian confidence decay function and the confidence, redundant boxes are removed to obtain the target tea impurity detection result.

[0108] In the embodiment, by removing redundant boxes based on the detection box, the Gaussian confidence decay function and the confidence, the spatial redundancy can be accurately quantified on the basis of retaining the advantages of adapting to small targets, so as to realize "strong inhibition of redundant boxes and accurate retention of effective boxes", and meanwhile, combined with a spatial partition acceleration strategy, the detection accuracy and real-time performance are taken into account.

[0109] The above removal of redundant boxes based on the detection box, the Gaussian confidence decay function and the confidence to obtain the target tea impurity detection result can be removal of redundant boxes based on the detection box, the Gaussian confidence decay function and the confidence by using an improved spatial scale-aware soft non-maximum suppression algorithm to obtain the target tea impurity detection result.

[0110] To facilitate the understanding of those skilled in the art, a set of best embodiments is provided below:

[0111] The embodiment proposes a tea impurity visual detection method, which comprises the following steps:

[0112] S1, obtain the tea image after winnowing and perform pretreatment.

[0113] The tea impurity image dataset is constructed by collecting tea and impurity images from a certain production plant. An industrial camera with a resolution of 2448x2048 is used for shooting, simulating the state of tea impurities on the actual production line, covering tea and impurity photo images under different conditions, ensuring the diversity and representativeness of the data.

[0114] The image dataset consists of 439 images, including 6 categories: tea tender leaves, tender stems, wax leaves, old stems, weeds and plastic strips. Then, the dataset is labeled using the Labelme toolbox, and the dataset is divided into training set, test set and validation set according to the ratio of 7:2:1. The training set is preprocessed, and the data augmentation techniques such as rotation, cropping, light adjustment, scaling and noise addition are used to expand the training set data to 1305, improving the generalization ability of the model. Finally, the preprocessed tea impurity image dataset (i.e. the preprocessed training image dataset) is obtained.

[0115] Specifically, the preprocessing of the embodiment comprises the following steps:

[0116] S1.1, in the data acquisition stage, 439 high-resolution images of 2448x2048 are systematically collected under standard lighting conditions using an industrial camera, completely covering the typical morphological characteristics of the six major impurities (tea tender leaves, tender stems, wax leaves, old stems, weeds and plastic strips) on the tea production line. To ensure the quality and representativeness of the dataset, this embodiment establishes strict sample preparation specifications, controls the tea laying and stacking accurately, configures the impurity distribution according to the actual production ratio, and ensures that each image contains 3 to 5 different impurities to simulate the complex scene of the real production line.

[0117] S1.2, in the data annotation link, this embodiment organizes a professional annotation team to use the Labelme toolbox for fine polygon annotation of all images, uses visible part annotation strategy for some occluded targets, and performs three rounds of cross-checking on all annotation results to ensure annotation accuracy. Based on the rigorous data division principle, this embodiment divides the dataset into training set, test set and validation set according to the ratio of 7:2:1, and especially reserves some most challenging high-overlap, low-contrast samples to the validation set to fully evaluate the limit performance of the model. ​

[0118] S1.3、To improve the generalization ability of the model, a multi-level data augmentation pipeline is designed and implemented. In the geometric transformation layer, random rotation (±15° range), perspective transformation (0.1 affine amplitude) and random cropping (0.7 to 1.0 ratio) are systematically applied; in the photometric transformation layer, this embodiment simulates complex lighting environments through brightness adjustment (±20%), contrast adjustment (±15%), gamma correction (0.8 to 1.2 range) and hue saturation disturbance; in the noise simulation layer, this embodiment introduces Gaussian noise (σ value range is 0.01 to 0.05) and salt and pepper noise (density 0.1% to 0.5%) to enhance the robustness of the model to image degradation. In addition, this embodiment also uses advanced enhancement strategies such as CutMix mixing (20% to 40% mixing ratio) and Mosaic four-image splicing to significantly improve the model's detection ability for dense small targets.

[0119] S1.4、After this complete preprocessing process, the original training set is expanded from 307 to 1305 high-quality training samples. At the same time, this embodiment establishes a strict quality control mechanism, removes blurred images through Laplacian variance detection, removes invalid annotations with an area less than 25 pixels square, and ensures the visual reasonableness and annotation accuracy of the augmented images.

[0120] S2、Based on the YOLOv11 target detection model, an improved TSCM-Det model (i.e. tea leaf impurity detection model) is constructed.

[0121] This embodiment optimizes the network model in depth while maintaining high inference speed, aiming at the specific challenges of detecting tiny impurities in tea images. The overall architecture of the network model consists of three core parts: backbone network (Backbone), feature fusion neck (Neck) and detection head (Head). After preprocessing and data augmentation, the input image is sent to the backbone network for feature extraction.

[0122] Firstly, the original Stem downsampling module is replaced by a specially designed ScharrStem module. This module introduces a Scharr differential operator with rotation invariance and a multi-scale Gaussian fusion mechanism to construct a cognitive-driven paradigm based on differential geometry prior in the front end of feature extraction, strengthening edge information and suppressing noise in the initial stage of feature extraction, laying a foundation for subsequent detail detection.

[0123] Secondly, the original C3K2 module is replaced by the innovative SGCB improved module, and a deep feature enhancement mechanism based on double-path cognitive collaboration is constructed. Through the use of Scharr prior guidance with differential invariance in the deep network, the dynamic balance between detail perception and semantic understanding is achieved. The SGCB module and the preposed ScharrStem module jointly constitute a "detail enhancement pipeline" from input to deep layer. The ScharrStem module performs one-time, intensive detail extraction and retention at the front end of the network, while the SGCB module continuously maintains and enhances the details in subsequent multiple levels. This collaborative working mode ensures that the features sensitive to small tea impurities can be transmitted throughout the forward propagation process of the entire network, thereby significantly improving the detection robustness and accuracy of the model in complex tea background.

[0124] The feature fusion neck adopts a feature fusion module (CASM) based on channel-aware spatial modulation. Through the channel-aware mechanism, the semantic roles of different feature channels are understood, and then precise spatial modulation is implemented to achieve adaptive fusion of low-level detail features and high-level semantic features.

[0125] In the post-processing stage, in view of the problem of high overlap of target in tea image and fragmentation of prediction box at boundary, a spatial-scale-aware soft non-maximum suppression algorithm is used for redundant box removal. This method combines spatial-scale-aware intersection over area (SSAIoA) and soft non-maximum suppression strategy, optimizes the overlap degree metric by fusing spatial position-aware factor and scale matching-aware factor, and designs spatial-scale-aware soft non-maximum suppression algorithm (SSA-SoftNMS) based on this. This algorithm, while retaining the advantage of adapting to small targets, realizes "strong suppression of redundant boxes and accurate retention of effective boxes" by accurately quantifying spatial redundancy, and at the same time, combines spatial partitioning acceleration strategy, taking into account detection accuracy and real-time performance.

[0126] Specifically, referring to Figure 2 , the TSCM-Det model mainly includes a backbone network, a neck network and a detection head, and the specific processing flow includes:

[0127] 1、The input image (which can be the image of tea to be detected) is a three-channel RGB color image with a size of 640x640x3. The input image is first processed by the ScharrStem module, which mainly functions to process the input image in the initial stage of the model. By combining the scharr edge detection operator and Gaussian smoothing denoising, the edge and detail information are actively strengthened in the initial stage of the network. Instead of the traditional network violent downsampling operation, by actively strengthening the edge and detail information in the initial stage, the problem of noise and edge information degradation caused by layer transmission in the network layer is solved, and high-quality feature representation is provided for subsequent detection tasks. The ScharrStem module outputs a feature image with a size of 160x160x32 (i.e., the feature extraction result).

[0128] 2、The feature is then extracted by the first SGCB improvement module improved in stage 1, and a feature image with a size of 160x160x64 (i.e., the first scale feature map) is output.

[0129] 3、Then, a 3x3 convolution is performed for downsampling operation, and a feature image with a size of 80x80x64 is output. Then, a BN batch normalization operation and a SiLU (Sigmoid Linear Unit) activation function are performed to improve the nonlinear expression ability of the model.

[0130] 4、Then, the feature is extracted by the first SGCB improvement module improved in stage 2, and a feature image with a size of 80x80x128 (i.e., the second scale feature map) is output.

[0131] 5、Then, a 3x3 convolution is performed for downsampling operation, and a feature image with a size of 40x40x128 is output. Then, a BN batch normalization operation and a SiLU (Sigmoid Linear Unit) activation function are performed to improve the nonlinear expression ability of the model.

[0132] 6、Then, the feature is extracted by the second SGCB improvement module improved in stage 3, and a feature image with a size of 40x40x128 (i.e., the third scale feature map) is output.

[0133] 7、Then, a 3x3 convolution is performed for downsampling operation, and a feature image with a size of 20x20x256 is output. Then, a BN batch normalization operation and a SiLU (Sigmoid Linear Unit) activation function are performed to improve the nonlinear expression ability of the model.

[0134] 8、Then, the feature is extracted by the second SGCB improvement module improved in stage 4, and a feature image with a size of 20x20x128 (i.e., the fourth scale feature map) is output.

[0135] 9、For the high-level features obtained in stage 4 (i.e., the fourth scale feature map), spatial pyramid pooling (SPPF) and C2PSA modules are used to extract multi-scale features, making the network more robust to objects of different sizes, especially large objects. SPPF simulates the effect of a large kernel pooling layer by concatenating multiple identical small kernel (5x5) pooling layers. The C2PSA module is an extension of the C2f module, which incorporates a PSA (Pointwise Spatial Attention) block to enhance feature extraction and attention mechanisms. By introducing a PSA block into the standard C2f module, C2PSA achieves a more powerful attention mechanism, thereby improving the model's ability to capture important features for tea impurity detection. The output feature map size is 20x20x256 (i.e., the first feature image). The C2f module is a module in the YOLO series known to those skilled in the art, and is not described in detail in this embodiment.

[0136] 10、The output feature image in step 9 is upsampled, usually by bilinear interpolation or transposed convolution, to match the spatial dimensions of the output feature image in stage 3 (i.e., the third scale feature map).

[0137] 11、The output feature image in stage 3 is dimensionally stacked with the first feature image after upsampling in step 10 to obtain a first stacking result.

[0138] 12、The first stacking result is further fused by the first SGCB improvement module to obtain a feature fusion result.

[0139] 13、The first feature fusion result is upsampled to match the spatial feature dimensions of the output feature image in stage 2 (i.e., the second scale feature map).

[0140] 14、The output feature image in stage 2 is dimensionally stacked with the second feature image after upsampling in step 13 to obtain a second stacking result.

[0141] 15、The output feature image in stage 1 (i.e., the first scale feature map) is downsampled by 3x3 convolution to match the feature size of the second stacking result in step 14, i.e., the output feature image in stage 2. Figure 1 .

[0142] 16、The second stacking result and the third feature image after downsampling in step 15 are input into the CSAM module for feature fusion to obtain a first enhanced feature map.

[0143] 17、The first enhanced feature map is further fused by the first SGCB improvement module to obtain a first target feature fusion result.

[0144] 18, the first target feature fusion result is sent to the detection head for detection to obtain the final result (i.e. the first tea impurity detection result).

[0145] 19, the first enhanced feature map of step 16 is subjected to 3x3 convolution for downsampling operation, so that the feature size is consistent with the feature fusion result of step 12.

[0146] 20, the fourth feature image after downsampling of step 19 and the feature fusion result of step 12 are input into the CSAM module for feature fusion to obtain a second enhanced feature map.

[0147] 21, the second enhanced feature map is subjected to further feature fusion by the first SGCB improvement module to obtain a second target feature fusion result.

[0148] 22, the second target feature fusion result is sent to the detection head for detection to obtain the final result (i.e. the second tea impurity detection result).

[0149] 23, the second target feature fusion result of step 21 is subjected to 3x3 convolution for downsampling operation, so that the feature size is consistent with the output feature Figure 1 .

[0150] 24, the fifth feature image after downsampling of step 23 and the output feature map (i.e. the first feature image) of step 9 are input into the CSAM module for feature fusion to obtain a third enhanced feature map.

[0151] 25, the third enhanced feature map is subjected to further feature fusion by the second SGCB improvement module to obtain a third target feature fusion result.

[0152] 26, the third target feature fusion result is sent to the detection head for detection to obtain the final result (i.e. the third tea impurity detection result).

[0153] Among them, referring to Figure 3 , the Scharr Stem module specifically includes:

[0154] The input image is applied to 7x7 convolution operation to extract the initial feature map with rich semantic information, the Scharr edge enhancement branch, the initial feature map is respectively applied to the horizontal direction Scharr convolution kernel , and the vertical direction Scharr convolution kernel is respectively applied to the horizontal direction Scharr convolution kernel , and the vertical direction Scharr convolution kernel is respectively applied to the horizontal direction Scharr convolution kernel

[0155] , and the vertical direction Scharr convolution kernel ;

[0156] ;

[0157] wherein, denotes a convolution operation; the edge-enhanced feature map is fused with the input image to obtain a fusion result, and the fusion result is sequentially applied to two 3x3 convolution layers for feature transformation and down-sampling to obtain an intermediate feature representation ;

[0158] The Gaussian smoothing fusion branch applies a 5x5 Gaussian filter (using two sizes of 5x5 and 9x9, respectively) to the intermediate feature representation for multi-scale smoothing processing, wherein the Gaussian filter is defined as:

[0159] ;

[0160] wherein, denotes the size of the Gaussian kernel, denotes the standard deviation, denotes the pixel position.

[0161] The multi-scale smoothed feature is fused with the original feature through a normalization operation to obtain a smoothed feature :

[0162] ;

[0163] The smoothed feature is finally input into a 3x3 convolution down-sampling module to compress the spatial size to 1 / 4 of the input, and output the final feature map (i.e., the feature extraction result) of the ScharrStem module , denotes the height, denotes the width, denotes the number of channels.

[0164] wherein, referring to Figure 4 , the SGCB improvement module specifically includes:

[0165] The SGCB submodule is integrated into the original feature extraction unit C3K2 submodule of YOLOv11 through a structured embedding strategy to construct an SGCB improvement module with hierarchical detail preservation capability. The SGCB submodule is specifically as follows:

[0166] The input feature map is projected onto a high-dimensional manifold through a 1x1 convolution kernel, and then two orthogonal feature subspaces and , the input primitive of parallel processing path. The edge-aware branch computes the gradient magnitude response of the feature space through an improved Scharr convolution layer, generating detail-enhanced features ; the auxiliary path adopts the same topology, processed by a standard 3x3 convolution layer, aiming to learn high-level semantic representations, generating semantic features , which is responsible for maintaining the network's ability to understand complex scenes and different categories of impurities.

[0167] Subsequently, the output features and of the two branches are concatenated in the channel dimension and fused and dimensionally reduced through a 1x1 convolution layer. The 1x1 convolution acts as a feature fusioner, whose task is to adaptively learn how to optimally combine the precise edge information from the detail branch and the contextual information from the semantic branch. The fusion process can be represented as:

[0168] ;

[0169] Finally, the fused features are added to the input of the module through a shortcut connection to form a residual learning structure:

[0170] ;

[0171] where denotes the concatenation operation, denotes the 1x1 convolution operation.

[0172] After residual learning, the features are processed by convolution to obtain the final output results of the SGCB submodel .

[0173] For the SGCB sub-module, the original bottleneck module is improved by introducing an improved Scharr convolution layer for edge detail preservation and enhancement, ensuring that features sensitive to small tea impurities can be propagated throughout the entire network, thereby significantly improving the detection robustness and accuracy of the model in complex tea backgrounds. Then it is added to the original C3K2 sub-module. Referring to Figure 4 , there are two different modes for the original C3K2 sub-module. One is not to use the C3k sub-module for feature enhancement, i.e. SGCB module = False, which is the first SGCB improved module. One is to use the C3k sub-module for feature enhancement, i.e. SGCB module = True, which is the second SGCB improved module.

[0174] The C3K2 sub-module and the C3k sub-module are modules in the YOLO series known to those skilled in the art, and the present embodiment does not make specific descriptions.

[0175] The improved Scharr convolution layer specifically includes the following:

[0176] The traditional Scharr convolution detects edges by calculating the local gradient amplitude, but cannot distinguish between “true edge gradients” and “noise gradients”, and its weights are fixed values designed manually and cannot be adjusted adaptively, have weak generalization ability, and cannot adapt to the feature specificity of each channel.

[0177] To solve the above problems, the embodiment proposes a Scharr convolution based on local variance weighting (i.e., an improved Scharr convolution layer). The core goal is to solve the core pain points of traditional Scharr convolution (poor noise resistance and weak generalization ability), while retaining its edge detection accuracy, making edge feature extraction more suitable for complex scenarios (such as noise interference, abstract feature maps, and multi-channel features). The core improvements are as follows:

[0178] 1. Accurately distinguish edges and noise through local variance to suppress false responses.

[0179] For edge regions: the feature values in a 3x3 neighborhood differ greatly (such as the junction of an object and the background), their variance is large → the weight is large → the gradient response is enhanced (highlighting the true edge); for noise regions: the feature values in a 3x3 neighborhood are discrete but within a small range (such as isolated noise points), the variance is moderate → the weight is moderate → the gradient response is suppressed (filtering false edges); for smooth regions: the feature values in a 3x3 neighborhood are uniform, the variance is small → the weight is small → slight suppression (avoiding false edges).

[0180] Through this “variance-weight” mapping, adaptive adjustment of “edge enhancement and noise suppression” is achieved, solving the false edge problem of traditional Scharr from the source.

[0181] 2. Adaptively adapt to different scenarios and feature maps.

[0182] For different images: automatically adapt to noise intensity (more areas are suppressed for images with high noise; more areas are enhanced for clear images); for intermediate layer features in CNN: adapt to the structural patterns of abstract features (such as abstract edges of deep features, which can be accurately captured); for multi-channel features: calculate the variance and weight for each channel independently, adapting to the feature specificity of each channel (such as channel A suppressing noise and channel B enhancing edges).

[0183] 3. Introduce learnable parameters to improve end-to-end training results.

[0184] A learnable noise threshold is introduced (or set an independent noise threshold for each channel): The system will automatically optimize during training, learning the critical values ​​between "noise variance" and "marginal variance" in the dataset to further improve discrimination accuracy; the specific steps are as follows:

[0185] Phase 1: Channel-by-channel Scharr gradient calculation and amplitude fusion.

[0186] Let the input multi-channel feature map be... ,in Indicates the first aisle( The coordinates in the middle are ( The original pixel feature values. First, for The Scharr gradient calculation yields the following horizontal and vertical gradient kernel distributions:

[0187] ;

[0188] Then, gradient calculation is performed channel by channel (this can be done by directly convolving the original feature map), using grouped convolution (group=C) for each original channel. Separate convolutions yield gradient maps in both the x and y directions. and In the formula This represents a 2D convolution, with padding=1 to keep the gradient map the same size as the original map.

[0189] ;

[0190] ;

[0191] Then, gradient magnitude fusion is performed, and the L2 norm is calculated on the gradient maps in the x and y directions to obtain the original edge response map. :in for Avoid errors caused by a gradient of 0.

[0192] ;

[0193] Phase 2: Calculation and normalization of 3×3 neighborhood variance for each channel.

[0194] Local variance is a core indicator for measuring the degree of local dispersion in a feature map. Edge regions have high variance due to abrupt changes in feature values, noisy regions have moderate variance due to isolated fluctuations, and smooth regions have low variance due to uniform features. In this stage, the 3×3 neighborhood variance is calculated for each pixel in each channel, and normalization is used to eliminate numerical scale interference.

[0195] First, calculate the mean for the domain: define a 3x3 mean kernel. The convolution operation is performed on each channel to obtain the field mean map The mean of the coordinates is: the convolution also adopts the padding=1 strategy to ensure that the mean map matches the spatial dimensions of the original channel, providing a basis for subsequent variance calculation.

[0196] ;

[0197] Then the local variance derivation is performed: using the mathematical properties of variance (where represents the expectation) to derive the local variance, avoiding the inefficient calculation of pixel-by-pixel field traversal, the specific steps are as follows:

[0198] ①Calculate the eigenvalue square map: ;

[0199] ②Calculate the square field mean map: ;

[0200] ③Derive the local variance map : ;

[0201] ④Due to the possibility of negative variance caused by floating-point calculation error, perform a truncation operation on the result: , to ensure the non-negativity of the variance.

[0202] Then normalize the variance: in order to eliminate the influence of the difference in numerical scale of variance between different channels on subsequent weight calculation, normalize the variance map of each channel separately, mapping to the [0,1] interval:

[0203] ;

[0204] where and represent the minimum and maximum values of the th channel variance map, respectively. The normalized variance map preserves the relative distribution characteristics of the original variance.

[0205] Phase 3: Generation of per-channel weight map.

[0206] The core function of the weight map is to dynamically adjust the response strength of the amplitude according to the local variance, achieving the goal of "edge enhancement-noise suppression". Based on the normalized variance map, this phase generates a pixel-by-pixel adaptive weight combined with learnable parameters.

[0207] First, the design of the weight formula is as follows: a nonlinear mapping function is used to convert the normalized variance into a weight, and the formula is defined as:

[0208] ;

[0209] where, is a learnable noise threshold parameter (initial value set to 0.01), which functions to calibrate the discriminative sensitivity of variance. When is small (adapt to low-noise data), the region with slightly larger variance can make the weight quickly approach 1; when is large (adapt to high-noise data), the region with significantly larger variance can make the weight approach 1, avoiding noise false enhancement. Global sharing (single parameter) or per-channel sharing (c parameters) strategies can be used, which can be selected according to the degree of channel difference, and optimized to the optimal value through end-to-end training.

[0210] Stage 4: Weighted weighting and enhanced edge feature output.

[0211] This stage performs pixel-by-pixel multiplication operation on the generated per-channel weight map and gradient amplitude map, completes the adaptive enhancement of edge features, and finally outputs a multi-channel enhanced feature map. First, weighted fusion calculation is performed, and the weight of each channel is pixel-by-pixel weighted with the original edge response map:

[0212] ;

[0213] where, represents the enhanced edge feature map of the c-th channel. Through weight adjustment, the gradient response of the edge region is retained, and the gradient response of the noise and smooth region is suppressed.

[0214] Finally, the enhanced edge feature maps of the c channels are integrated in the original channel order to obtain the final output feature map , which has the same spatial dimension as the input feature map.

[0215] This design not only alleviates the gradient vanishing problem in deep networks, but also enables the SGCB sub-module to focus on learning the residual part of the input features, i.e., the incremental information related to tea impurity details. The SGCB sub-module and the preposed ScharrStem module jointly constitute a "detail enhancement pipeline" from input to deep layer. ScharrStem performs a one-time, intensive detail extraction and retention at the front end of the network, while the SGCB sub-module continuously maintains and enhances the details at multiple subsequent levels. This collaborative working mode ensures that the features sensitive to small tea impurities can be transmitted throughout the forward propagation process of the entire network, thereby significantly improving the detection robustness and accuracy of the model in complex tea background.

[0216] wherein, referring to Figure 5 , the channel-aware spatial modulation module (CASM module) specifically includes:

[0217] The design of the CASM module follows the idea of "guiding-modulating-refining", and its overall architecture is a multi-stage processing pipeline with a clear data flow. Let the two input feature maps to be fused be low-level detail features (i.e., low-level features) and high-level semantic features (high-level features).

[0218] Stage one, feature initialization and collaborative attention generation.

[0219] Feature initialization: first, a basic fusion feature is obtained through element-level addition as the input for subsequent attention calculation (i.e., the first addition result): This operation ensures that the calculation of subsequent attention can be carried out in a fused context. The generation of parallel channel attention also generates a guiding signal. The goal of this path is not only to produce channel weights, but also to generate a guiding signal for spatial attention calculation.

[0220] First, for the fusion feature along the spatial dimension, the embedding of global information is obtained by compressing it, generating a channel-level statistical description, which is a global average pooling along the channel. Among them is:

[0221] ;

[0222] Then use an MLP with dimension reduction function to learn the complex nonlinear relationship between channels. First, reduce the dimension through 1x1 convolution and refine the features.

[0223] ;

[0224] Among them, is the dimension reduction weight matrix, and the first convolution feature takes it as the channel guiding signal, which is a compact and abstract representation of the original channel information, represents the bias term.

[0225] Then is passed through the ReLU activation function and 1x1 convolution to increase the dimension to obtain the final channel attention weight :

[0226] ;

[0227] Among them, , .

[0228] Generation of parallel spatial attention. Firstly, global max pooling and average pooling are used to capture the extreme response and average response of the feature map, forming a preliminary description of the spatial features.

[0229] ;

[0230] ;

[0231] wherein, denotes the first pooling result, denotes the second pooling result, both of which are spliced in the channel dimension to obtain .

[0232] Then a 3x3 convolution is used to preliminarily model the spatial context of the spliced statistical features, embed local context information, and project it into a higher-dimensional feature space.

[0233] ;

[0234] wherein, is a 3x3 convolution layer, and the number of output channels is M, so the second convolution feature .

[0235] Then the guide signal generated in the channel attention path is projected into a dimension matching the feature by a transformation network composed of two 1x1 convolutions.

[0236] ;

[0237] ;

[0238] wherein, denotes the feature processed by the activation function, the first weight matrix , the second weight matrix , is a Sigmoid activation function, which ensures that the output value is in the interval [0, 1], and are bias terms. The third convolution feature , i.e., the spatial modulation gate.

[0239] Then the gate signal is expanded to the spatial dimension through the broadcast mechanism and multiplied element-wise with the primary spatial feature. The physical meaning of this step is that according to the importance of each channel, the activation value of each position in the spatial feature is dynamically recalibrated. The important feature channels corresponding to the spatial context information are enhanced, and the unimportant ones are suppressed.

[0240] ;

[0241] Finally, the modulated features are mapped to a single-channel spatial attention map:

[0242] ;

[0243] wherein, represents a broadcasting mechanism, is a 3x3 convolution with 1 output channel, is a channel-information-fused collaborative spatial attention weight.

[0244] Stage two: core modulation stage.

[0245] First, the spatial attention weight is element-wise added to the channel attention weight to achieve primary feature fusion. Then, a channel shuffling operation is introduced to promote information interaction between different channel groups by rearranging grouped channels, breaking the limitations of traditional grouped processing:

[0246] ;

[0247] wherein, represents a grouped channel rearrangement. Grouped channel rearrangement is a technique that constructs a new feature map by changing the arrangement order of channels in a tensor without changing the total data amount of the tensor (i.e., without increasing computation and parameters). It breaks the information isolation brought by grouped convolution and promotes information interaction between different feature groups. Grouped channel rearrangement employs a channel rearrangement known to those skilled in the art, which is not specifically described in this embodiment.

[0248] The fused weight after channel rearrangement is subjected to spatial refinement processing, two 3x3 convolutions are used for deep feature extraction, and finally a Sigmoid activation function is used for numerical normalization to generate spatial modulation weights with channel specificity:

[0249] ;

[0250] Based on the learned modulation weights, precise spatial modulation is performed on the input features. A complementary weighting strategy is adopted, in which low-level detailed features are multiplied by the modulation weights, while high-level semantic features are weighted by the weight complements:

[0251] ;

[0252] To preserve the integrity of the original features, a residual connection mechanism is introduced to combine the weighted features Add up with the original input features, and finally realize feature integration and dimension unification through 1x1 convolution to obtain enhanced feature map :

[0253] .

[0254] In the post-processing stage, a multi-dimensional evaluation system and an adaptive decay function are constructed to remove redundant boxes, which specifically includes:

[0255] In the target detection task, Non-Maximum Suppression (NMS) and its variants are the core post-processing technology for removing redundant detection boxes. The traditional NMS relies on the Intersection over Union (IoU) to measure the overlap, but it has significant limitations in tea impurities and other scenarios: first, IoU only focuses on the overlap area ratio of two boxes, and cannot distinguish between "center-aligned redundant overlap" and "edge-shifted effective neighbor"; second, it is sensitive to the "large box containing small box" scenario, and is prone to false suppression of small effective targets due to high IoU values; third, it does not consider the coupling effect of target spatial position and scale difference, resulting in a lack of targeted suppression strategy.

[0256] To solve the above problems, Spatial-Scale Aware Intersection over Area (SSAIoA) is proposed, which optimizes the overlap measure by fusing spatial position awareness factor and scale matching awareness factor, and designs Spatial-Scale Aware Soft-NMS (SSA-SoftNMS) based on it. This algorithm, while retaining the advantage of adapting to small targets, achieves "strong suppression of redundant boxes and precise preservation of effective boxes" by accurately quantifying spatial redundancy, and combines spatial partitioning acceleration strategy to balance detection accuracy and real-time performance. The specific algorithm process is as follows:

[0257] I. Algorithm input includes:

[0258] Detection box set Each box Uses YOLO output format: Where is the normalized center coordinate, is the normalized width and height, is the class ID; Confidence set corresponds to the detection box one by one, ; Image original size ; Preset hyperparameters: area threshold , SSAIoA parameter , hierarchical threshold , Gaussian decay parameter .

[0259] The empty scale-aware IoA (SSAIoA) calculation, the core of the SSAIoA is to modify the original IoA by a space-aware factor (SAF) ) and a scale-aware factor (SAF) ), which accurately quantifies the spatial redundancy of the two boxes. Specifically, it is defined as:

[0260] ;

[0261] Wherein is the original area coverage ratio, measures the spatial alignment of the centers of the two boxes, measures the scale matching degree of the two boxes, and finally , the larger the value, the higher the redundancy, and the more it needs to be suppressed.

[0262] 1. Basic IoA calculation:

[0263] ;

[0264] Wherein, represents the area function; represents the detection box currently being evaluated, which can be referred to as the candidate box; represents a detection box with a higher confidence ratio , which can be referred to as a "high-confidence box", represents the area of the intersection of the two boxes, represents the area of the candidate box itself.

[0265] 2. Calculation of the space-aware factor :

[0266] The space-aware factor is used to quantify the spatial alignment of the centers of the two boxes. Its core logic is "the closer the center distance, the higher the possibility of spatial redundancy", and a Gaussian decay function is used to achieve smooth awareness:

[0267] ;

[0268] Wherein: , directly reflects the spatial offset of the centers of the two boxes; is the normalized distance: , is the diagonal length of the image; is the position smoothing factor (recommended ), which controls the decay rate of the position offset. The smaller the position smoothing factor, the stronger the weakening effect of the center offset on , and only the center height alignment box will obtain a high factor value.

[0269] 3. Scale-aware factor The calculation of:

[0270] The scale perception factor is used to quantify the size matching degree of two boxes. The core logic is that the closer the size is, the more likely it is a repeated detection of the same target, and the greater the size difference is, the more likely it is an independent target. Precise perception is achieved through exponential enhancement.

[0271] ;

[0272] wherein, is the area similarity, and the value range is The closer the size is, the closer the value is to 1, and the greater the size difference is, the smaller the value is. is the scale enhancement factor, and the recommended value is to amplify the weight influence of size difference.

[0273] The perception effect of size difference is amplified through exponential operation. For example, the area similarity of two boxes is 0.9 (close in size), and after exponential operation, it is 0.81, while the similarity is 0.5 (large size difference) and the operation is 0.25. Therefore, the ability to distinguish the scene of “large box containing small box” can be enhanced.

[0274] Hierarchical dynamic threshold: based on the pixel-level area of the candidate box, adopt hierarchical threshold to adapt to the scale uneven characteristics:

[0275] ;

[0276] wherein, (recommended value is 0.85), (recommended value is 0.65), (recommended value is 0.45) are hierarchical thresholds for adapting to SSAIoA.

[0277] Gaussian confidence decay function: if , the confidence of is smoothly decayed to avoid effective target loss caused by hard deletion:

[0278] ;

[0279] wherein, is the decay smoothing factor (recommended ), is the decayed confidence; if , the confidence remains unchanged. This function replaces the traditional hard deletion with smooth confidence decay, effectively preserving the confidence of the actual effective detection target in the dense area while maintaining the detection accuracy.

[0280] The detailed execution steps of the SSASoftNMS algorithm include:

[0281] (1) Input and initialization:

[0282] Input detection box set and confidence set Create an empty list D to store the final retained detection results, and create a processing queue Q, where the results are processed according to confidence level. The indexes of all detection boxes are stored in descending order of their order.

[0283] (2) Loop processing:

[0284] When queue Q is not empty:

[0285] a) Select the highest-scoring detection box: Pop the detection box with the highest confidence level from Q. ;

[0286] b. Add it to the final result: and Add to list D;

[0287] c. Suppress remaining boxes by traversing them: For each remaining box in Q... :

[0288] i. Calculate the spatial scale-sensing overlap area ratio matrix: Calculate right coverage ;

[0289] ii. Determine the dynamic threshold: based on The area to determine its use threshold :

[0290] 1) If ,but ;

[0291] 2) If ,but ;

[0292] 3) If ,but ;

[0293] iii. Implement confidence decay based on Gaussian kernel:

[0294] if ,but ,otherwise It remains unchanged.

[0295] iv. Dynamically update the confidence distribution of the detection box: update the box confidence score .

[0296] d, clean up the queue: remove all bounding boxes in Q whose confidence S is lower than a preset final threshold (e.g. 0.01 or 0.001).

[0297] e, reorder the queue: according to the updated confidence, sort the queue Q from high to low.

[0298] (3) Output the optimized detection result through a double verification mechanism.

[0299] First verification, the above main loop itself is a confidence-based overlap verification.

[0300] Second verification, for the results in the final list D, a final confidence threshold (e.g. 0.5) can be applied. Only the bounding boxes with scores higher than this threshold will be finally output, which ensures that all output results have high reliability.

[0301] The final output result D contains the results that pass the verification and .

[0302] S3, train the TSCM-Det model through the pre-processed tea impurity image dataset to obtain a trained TSCM-Det model; input the tea image to be detected into the trained TSCM-Det model to obtain a tea impurity detection result.

[0303] Referring to Figure 6 , the embodiments of the present application also provide a tea impurity visual detection system, which comprises a data acquisition unit 601, a detection result obtaining unit 602 and a target result obtaining unit 603, wherein:

[0304] The data acquisition unit 601 is configured to acquire a tea image to be detected and a training image dataset, pre-process image data in the training image dataset to obtain a pre-processed training image dataset, and the tea image to be detected contains a plurality of targets to be detected.

[0305] The detection result obtaining unit 602 is configured to input the tea leaf image to be detected into the trained tea impurity detection model to obtain a plurality of tea impurity detection results containing a confidence and a detection box of a target to be detected. The trained tea impurity detection model is trained by the preprocessed training image data set. The trained tea impurity detection model includes a backbone network, a neck network and a plurality of detection heads. The backbone network includes a ScharrStem module, a first SGCB improvement module and a second SGCB improvement module. The neck network includes the first SGCB improvement module, the second SGCB improvement module and a channel-aware spatial modulation module. Wherein,

[0306] The ScharrStem module is configured to perform feature extraction on the tea leaf image to be detected to obtain a feature extraction result.

[0307] The plurality of stages containing the first SGCB improvement module or the second SGCB improvement module are configured to perform multi-scale feature extraction on the feature extraction result to obtain a multi-scale feature map.

[0308] The multi-scale feature map is processed by the first SGCB improvement module, the second SGCB improvement module and the channel-aware spatial modulation module to obtain a plurality of target feature fusion results.

[0309] The plurality of detection heads are configured to detect each target feature fusion result respectively to obtain a plurality of tea impurity detection results.

[0310] The target result obtaining unit 603 is configured to determine a target tea impurity detection result according to the confidence and the detection box in the plurality of tea impurity detection results.

[0311] It should be noted that, since the tea impurity visual detection system in the embodiment and the tea impurity visual detection method described above are based on the same inventive concept, the corresponding contents in the method embodiment are also applicable to the system embodiment, which will not be described in detail here.

[0312] Referring to Figure 7 The present application also provides an electronic device, which comprises:

[0313] at least one memory;

[0314] at least one processor;

[0315] at least one program;

[0316] The program is stored in the memory, and the processor executes the at least one program to implement the tea impurity visual detection method described above.

[0317] The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a vehicle-mounted computer, and the like.

[0318] The electronic device of the embodiment of the present application is described in detail below.

[0319] The processor 1600 can be implemented in a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute related programs to implement the technical solutions provided by the embodiments of the present application.

[0320] The memory 1700 can be implemented in a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1700 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present application are implemented by software or firmware, the related program codes are stored in the memory 1700 and are called and executed by the processor 1600 to implement the tea impurity visual detection method of the embodiments of the present application.

[0321] The input / output interface 1800 is configured to realize information input and output.

[0322] The communication interface 1900 is configured to realize the communication interaction between the device and other devices. The communication can be realized by a wired manner (for example, a USB, a network cable, and the like) or a wireless manner (for example, a mobile network, WIFI, Bluetooth, and the like).

[0323] The bus 2000 is configured to transmit information between the components (for example, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900) of the device.

[0324] The processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 are connected to each other through the bus 2000 to realize the communication connection between them in the device.

[0325] The embodiments of the present application also provide a storage medium, which is a computer readable storage medium and stores computer executable instructions for causing a computer to execute the above tea impurity visual detection method.

[0326] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0327] It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0328] The above is a detailed description of the preferred embodiments of this application. However, the embodiments of this application are not limited to the above-described implementation methods. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the embodiments of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of the embodiments of this application.

Claims

1. A method of visual inspection of tea leaf impurities, characterized in that, The method comprises: obtaining a to-be-detected tea image and a training image dataset, preprocessing image data in the training image dataset to obtain a preprocessed training image dataset, and the to-be-detected tea image containing a plurality of to-be-detected targets; inputting the to-be-detected tea image into a trained tea impurity detection model to obtain a plurality of tea impurity detection results containing confidence and detection boxes of the to-be-detected targets, the trained tea impurity detection model being trained by the preprocessed training image dataset, the trained tea impurity detection model comprising a backbone network, a neck network, and a plurality of detection heads, the backbone network comprising a ScharrStem module, a first SGCB improvement module, and a second SGCB improvement module, the neck network comprising the first SGCB improvement module, the second SGCB improvement module, and a channel-aware spatial modulation module, wherein the first SGCB improvement module and the second SGCB improvement module comprise an SGCB submodule and a C3K2 submodule, the SGCB submodule comprising a plurality of convolutional layers and an improved Scharr convolutional layer; the SGCB submodule is added to the C3K2 submodule to obtain an SGCB improvement module, when the C3K2 submodule in the SGCB improvement module does not use a C3k submodule for feature enhancement, the SGCB improvement module is a first SGCB improvement module; when the C3K2 submodule in the SGCB improvement module uses a C3k submodule for feature enhancement, the SGCB improvement module is a second SGCB improvement module; the improved Scharr convolutional layer comprises the following processing process: performing a channel-by-channel Scharr gradient calculation on an input feature map to obtain an x-direction gradient map and a y-direction gradient map of each channel; performing gradient amplitude fusion on the x-direction gradient map and the y-direction gradient map to obtain an original edge response map of each channel; performing channel-by-channel 3x3 neighborhood variance calculation and normalization processing on the input feature map to obtain a normalized variance map; calculating the weight of each channel according to the normalized variance map; performing pixel-by-pixel weighting on the original edge response map corresponding to each channel through the weight of each channel to obtain an enhanced edge feature map of each channel; integrating the enhanced edge feature map of each channel in the original channel order to obtain a final output feature map; extracting features from the to-be-detected tea image through the ScharrStem module to obtain a feature extraction result; performing multi-scale feature extraction on the feature extraction result through a plurality of stages containing the first SGCB improvement module or the second SGCB improvement module to obtain a multi-scale feature map; processing the multi-scale feature map through the first SGCB improvement module, the second SGCB improvement module, and the channel-aware spatial modulation module to obtain a plurality of target feature fusion results; detecting each target feature fusion result through the plurality of detection heads respectively to obtain a plurality of tea impurity detection results; and According to the confidence and the detection box in the plurality of tea impurity detection results, a target tea impurity detection result is determined.

2. The tea leaf impurity visual detection method according to claim 1, characterized in that, The feature extraction result is obtained by performing feature extraction on the to-be-detected tea image through the Scharr Stem module, including: performing convolution operation on the to-be-detected tea image to obtain an initial feature map; performing Scharr filtering on the initial feature map to obtain an edge-enhanced feature map; fusing the edge-enhanced feature map with the to-be-detected tea image to obtain a first fusion result; performing twice convolution operation on the first fusion result to obtain an intermediate feature representation; performing Gaussian smoothing processing on the intermediate feature representation to obtain a smoothing result; fusing the smoothing result with the intermediate feature representation to obtain a second fusion result; performing convolution operation on the second fusion result to obtain the feature extraction result.

3. The tea leaf impurity visual detection method according to claim 1, wherein, The multi-scale feature map is obtained by performing multi-scale feature extraction on the feature extraction result through a plurality of stages containing the first SGCB improvement module or the second SGCB improvement module, including: the feature extraction result is input into the first stage containing the first SGCB improvement module to perform feature extraction, and a first scale feature map is obtained; performing down-sampling, batch normalization and SiLU activation function processing on the first scale feature map to obtain a first processing result; the first processing result is input into the second stage containing the first SGCB improvement module to perform feature extraction, and a second scale feature map is obtained; performing down-sampling, batch normalization and SiLU activation function processing on the second scale feature map to obtain a second processing result; the second processing result is input into the third stage containing the second SGCB improvement module to perform feature extraction, and a third scale feature map is obtained; performing down-sampling, batch normalization and SiLU activation function processing on the third scale feature map to obtain a third processing result; the third processing result is input into the fourth stage containing the second SGCB improvement module to perform feature extraction, and a fourth scale feature map is obtained.

4. The tea leaf impurity visual detection method according to claim 1, characterized in that, The plurality of target feature fusion results are obtained by processing the multi-scale feature map through the first SGCB improvement module, the second SGCB improvement module and the channel perception spatial modulation module, including: performing feature extraction on the fourth scale feature map to obtain a first feature image, performing up-sampling operation on the first feature image to obtain an up-sampled first feature image; performing dimension stacking on the third scale feature map and the up-sampled first feature image to obtain a first stacking result, inputting the first stacking result into the first SGCB improvement module to perform feature fusion, and obtaining a feature fusion result; performing up-sampling operation on the feature fusion result to obtain an up-sampled second feature image, performing dimension stacking on the second scale feature map and the up-sampled second feature image to obtain a second stacking result; The first scale feature map is down-sampled to obtain a third down-sampled feature image, and the third down-sampled feature image and the second superposition result are input into the channel perception spatial modulation module for feature fusion to obtain a first enhanced feature map; the first enhanced feature map is input into the first SGCB improved module for feature fusion to obtain a first target feature fusion result; The first enhanced feature map is down-sampled to obtain a fourth down-sampled feature image, and the fourth down-sampled feature image and the feature fusion result are input into the channel perception spatial modulation module for feature fusion to obtain a second enhanced feature map; the second enhanced feature map is input into the first SGCB improved module for feature fusion to obtain a second target feature fusion result; The second target feature fusion result is down-sampled to obtain a fifth down-sampled feature image, and the fifth down-sampled feature image and the first feature image are input into the channel perception spatial modulation module for feature fusion to obtain a third enhanced feature map; the third enhanced feature map is input into the second SGCB improved module for feature fusion to obtain a third target feature fusion result.

5. The tea leaf impurity visual detection method according to claim 4, wherein, The third down-sampled feature image and the second superposition result are input into the channel perception spatial modulation module for feature fusion to obtain a first enhanced feature map, including: The third down-sampled feature image is used as a low-level feature, and the second superposition result is used as a high-level feature; The low-level feature and the high-level feature are element-level added to obtain a first addition result; The first addition result is input into a channel attention module for global average pooling and convolution operation to obtain a first convolution feature; The first convolution feature is processed through an activation function and a convolution operation to obtain a channel attention weight; The first addition result is maximum-pooled to obtain a first pooling result, and the first addition result is average-pooled to obtain a second pooling result; The first pooling result and the second pooling result are spliced and then convolved to obtain a second convolution feature; The first convolution feature is processed through a convolution and an activation function to obtain a third convolution feature; The third convolution feature is expanded through a broadcast mechanism and multiplied by the second convolution feature to obtain a collaborative spatial attention weight; The collaborative spatial attention weight and the channel attention weight are element-level added to obtain a second addition result, and the first addition result and the second addition result are spliced to obtain a spliced result; The spliced result is grouped and channel-arranged to obtain a fusion weight, and the fusion weight is processed through a convolution and an activation function to obtain a spatial modulation weight; A weighted feature is determined according to the spatial modulation weight, the low-level feature, and the high-level feature; The weighted feature, the low-level feature, and the high-level feature are element-level added to obtain a third addition result; The third addition result is convolved to obtain a first enhanced feature map.

6. The tea leaf impurity visual detection method according to claim 1, wherein, The target tea impurity detection result is determined according to the confidence and the detection frame in the plurality of tea impurity detection results, and the target tea impurity detection result is determined according to the confidence and the detection frame in the plurality of tea impurity detection results, comprising: According to the detection frame in the plurality of tea impurity detection results, the empty scale perception overlapping area ratio between two detection frames is determined; Based on the empty scale perception overlapping area ratio and the confidence, a Gaussian confidence decay function is constructed, which is used to update the confidence; Based on the detection frame, the Gaussian confidence decay function and the confidence, redundant frames are removed to obtain the target tea impurity detection result.

7. A visual tea leaf impurities detection system, characterized in that, The system comprises: A data acquisition unit is configured to acquire a to-be-detected tea image and a training image data set, pre-process image data in the training image data set to obtain a pre-processed training image data set, and the to-be-detected tea image contains a plurality of to-be-detected targets; A detection result obtaining unit is configured to input the to-be-detected tea image into a trained tea impurity detection model to obtain a plurality of tea impurity detection results containing the confidence and the detection frame of the to-be-detected target, the trained tea impurity detection model is trained by the pre-processed training image data set, the trained tea impurity detection model comprises a backbone network, a neck network and a plurality of detection heads, the backbone network comprises a ScharrStem module, a first SGCB improvement module and a second SGCB improvement module, the neck network comprises a first SGCB improvement module, a second SGCB improvement module and a channel perception spatial modulation module, wherein the first SGCB improvement module and the second SGCB improvement module comprise an SGCB submodule and a C3K2 submodule, the SGCB submodule comprises a plurality of convolutional layers and an improved Scharr convolutional layer; the SGCB submodule is added to the C3K2 submodule to obtain an SGCB improvement module, when the C3K2 submodule in the SGCB improvement module does not use a C3k submodule for feature enhancement, the SGCB improvement module is a first SGCB improvement module; when the C3K2 submodule in the SGCB improvement module uses a C3k submodule for feature enhancement, the SGCB improvement module is a second SGCB improvement module; the improved Scharr convolutional layer comprises the following processing process: The input feature map is subjected to a channel-by-channel Scharr gradient calculation to obtain an x-direction gradient map and a y-direction gradient map of each channel; the x-direction gradient map and the y-direction gradient map are subjected to gradient amplitude fusion to obtain an original edge response map of each channel; the input feature map is subjected to channel-by-channel 3*3 neighborhood variance calculation and normalization processing to obtain a normalized variance map; the weight of each channel is calculated according to the normalized variance map; the original edge response map corresponding to each channel is subjected to pixel-by-pixel weighting through the weight of each channel to obtain an enhanced edge feature map of each channel; the enhanced edge feature map of each channel is integrated in the original channel order to obtain a final output feature map; The Scharr Stem module is used for feature extraction on the tea image to be detected, and a feature extraction result is obtained; Multi-scale feature extraction is performed on the feature extraction result through multiple stages including the first SGCB improvement module or the second SGCB improvement module, and a multi-scale feature map is obtained; The multi-scale feature map is processed through the first SGCB improvement module, the second SGCB improvement module and the channel perception spatial modulation module, and multiple target feature fusion results are obtained; Each target feature fusion result is detected by using the multiple detection heads, and multiple tea impurity detection results are obtained; A target result obtaining unit is configured to determine a target tea impurity detection result according to the confidence and the detection frame in the multiple tea impurity detection results.

8. An electronic device, comprising: The memory is connected in communication with the at least one control processor, and stores instructions executable by the at least one control processor. The instructions are executed by the at least one control processor to enable the at least one control processor to perform the tea impurity visual detection method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer executable instructions for causing a computer to perform the tea impurity visual detection method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Steel surface defect detection method, device and equipment and storage medium

    CN121120578A

  • Image enhancement method and apparatus, device and medium

    US20250200718A1