Visual detection method and system for tea impurities, electronic equipment and storage medium
By introducing the ScharrStem module and the improved SGCB module into the tea impurity detection, combined with the channel sensing spatial modulation module, the problem of low accuracy in tea impurity detection was solved, and efficient identification and localization of impurities at multiple scales were achieved.
Patent Information
- Application Number
- CN202511882693.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2045-12-15
AI Technical Summary
Existing technologies suffer from low accuracy in tea impurity detection, particularly in their inability to identify minute and multi-scale impurities. Furthermore, the feature fusion mechanism fails to adequately consider the complementarity of features at different levels.
The ScharrStem module is used to enhance the contour feature representation of small impurities. The SGCB module is used to prevent the semantic decay of small impurity features in multiple convolution transformations. The channel-aware spatial modulation module is used to achieve adaptive fusion of features at different levels, thereby improving the model's feature discrimination ability and localization accuracy for multi-scale impurities.
It significantly improves the accuracy of tea impurity detection. The well-trained tea impurity detection model can more accurately identify and locate multi-scale impurities, thus improving the accuracy of the detection results.
Smart Images

Figure CN121329970A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of visual detection, in particular to a tea impurity visual detection method and system, an electronic device and a storage medium. BACKGROUND
[0002] With the rapid development of agricultural technology, intelligent and automated technology is gradually integrated into the modern agricultural production system, promoting the transformation of traditional agriculture to smart agriculture. As an important economic crop, tea has a wide range of consumer demand in the domestic and international markets, and its quality control directly affects consumer experience and industry competitiveness. In the tea production process, impurity detection is a key link, which aims to identify and remove non-tea foreign matter such as straw, wax leaves and old stems, in order to protect product quality.
[0003] The advent of deep learning has brought revolutionary changes to the field of object detection, and the YOLO series algorithm is widely adopted in agricultural applications due to its outstanding speed and accuracy. However, directly using existing models for tea impurity detection still faces significant challenges. First, there is often a high degree of color and texture similarity between impurities and tea background, making it difficult to distinguish features. Second, impurity target size varies greatly, especially small impurities, which lack information in the feature map, making it easy to miss detection. Third, the feature fusion mechanism of existing models often fails to fully consider the complementarity of different levels of features in spatial details and semantic information, limiting the model's ability to perceive multi-scale, occluded or irregularly shaped impurities. Therefore, the detection accuracy of existing technology for tea impurities is relatively low. SUMMARY
[0004] The present application aims to provide a tea impurity visual detection method and system, an electronic device and a storage medium, which can improve the detection accuracy of tea impurities.
[0005] In a first aspect, the present application provides a tea impurity visual detection method, which comprises: obtaining a to-be-detected tea image and a training image data set, preprocessing the image data in the training image data set to obtain a preprocessed training image data set, and the to-be-detected tea image contains a plurality of to-be-detected targets; input the tea leaf image to be detected into the trained tea leaf impurity detection model to obtain a plurality of tea leaf impurity detection results containing confidence and detection boxes of the target to be detected, the trained tea leaf impurity detection model is trained by the preprocessed training image data set, the trained tea leaf impurity detection model comprises a backbone network, a neck network and a plurality of detection heads, the backbone network comprises a ScharrStem module, a first SGCB improvement module and a second SGCB improvement module, the neck network comprises a first SGCB improvement module, a second SGCB improvement module and a channel perception spatial modulation module, wherein, feature extraction is performed on the tea leaf image to be detected through the ScharrStem module to obtain a feature extraction result; multi-scale feature extraction is performed on the feature extraction result through a plurality of stages containing the first SGCB improvement module or the second SGCB improvement module to obtain a multi-scale feature map; the multi-scale feature map is processed through the first SGCB improvement module, the second SGCB improvement module and the channel perception spatial modulation module to obtain a plurality of target feature fusion results; each target feature fusion result is detected by the plurality of detection heads respectively to obtain a plurality of tea leaf impurity detection results; a target tea leaf impurity detection result is determined according to the confidence and the detection box in the plurality of tea leaf impurity detection results.
[0006] Compared with the prior art, the first aspect of the present application has the following beneficial effects: The method can strengthen the contour feature expression of the micro impurities through the ScharrStem module, provide a high-quality detail feature base for subsequent deep processing, effectively prevent semantic attenuation of the micro impurity features in multiple convolution transformations through the SGCB improvement module, and ensure that the sensitivity to fine targets can be maintained in the deep network, and the adaptive fusion of different level features can be realized through the channel perception spatial modulation module, which significantly improves the feature discrimination ability and positioning accuracy of the model for multi-scale impurities. Therefore, the tea leaf impurity detection result is obtained through the trained tea leaf impurity detection model, which can improve the accuracy of the detection result, and the target tea leaf impurity detection result is determined according to the confidence and the detection box in the plurality of tea leaf impurity detection results, which further improves the detection accuracy of the tea leaf impurities.
[0007] In a second aspect, the embodiments of the present application also provide a tea leaf impurity visual detection system, which comprises: a data acquisition unit configured to acquire a tea leaf image to be detected and a training image data set, preprocess image data in the training image data set to obtain a preprocessed training image data set, and the tea leaf image to be detected contains a plurality of targets to be detected. The detection result obtaining unit is configured to input the tea leaf image to be detected into a trained tea impurity detection model to obtain a plurality of tea impurity detection results containing confidence and a detection box of the target to be detected, the trained tea impurity detection model being trained by the preprocessed training image data set, and the trained tea impurity detection model comprising a backbone network, a neck network and a plurality of detection heads, the backbone network comprising a ScharrStem module, a first SGCB improvement module and a second SGCB improvement module, the neck network comprising the first SGCB improvement module, the second SGCB improvement module and a channel-aware spatial modulation module, wherein The ScharrStem module is configured to perform feature extraction on the tea leaf image to be detected to obtain a feature extraction result. The plurality of stages containing the first SGCB improvement module or the second SGCB improvement module are configured to perform multi-scale feature extraction on the feature extraction result to obtain a multi-scale feature map. The first SGCB improvement module, the second SGCB improvement module and the channel-aware spatial modulation module are configured to process the multi-scale feature map to obtain a plurality of target feature fusion results. The plurality of detection heads are configured to respectively detect each of the target feature fusion results to obtain a plurality of tea impurity detection results. The target result obtaining unit is configured to determine a target tea impurity detection result according to the confidence and the detection box in the plurality of tea impurity detection results.
[0008] In a third aspect, an electronic device is provided, which includes at least one control processor and a memory in communication connection with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to perform the tea impurity visual detection method as described above.
[0009] In a fourth aspect, a computer readable storage medium is provided, which stores computer executable instructions for causing a computer to perform the tea impurity visual detection method as described above.
[0010] It can be understood that the beneficial effects of the second aspect to the fourth aspect compared with the related art are the same as the beneficial effects of the first aspect compared with the related art, and reference can be made to the related description in the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS
[0011] The above and / or additional aspects and advantages of the present application will become apparent and more readily appreciated from the following description, taken in conjunction with the following drawings in which: Figure 1 is a flowchart of an embodiment of the tea impurity visual detection method provided by the present application; Figure 2 is a structural schematic diagram of the TSCM-Det model in the best embodiment of the tea impurity visual detection method provided by the present application; Figure 3 is a structural schematic diagram of the ScharrStem module in the best embodiment of the tea impurity visual detection method provided by the present application; Figure 4 is a structural schematic diagram of the SGCB improvement module in the best embodiment of the tea impurity visual detection method provided by the present application; Figure 5 is a structural schematic diagram of the channel perception spatial modulation module in the best embodiment of the tea impurity visual detection method provided by the present application; Figure 6 is a structural schematic diagram of an embodiment of the tea impurity visual detection system provided by the present application; Figure 7 is a structural schematic diagram of an embodiment of the electronic device provided by the present application. DETAILED DESCRIPTION
[0012] Embodiments of the present application are described in detail below with reference to the accompanying drawings, in which the same or similar elements or elements having the same or similar functions are denoted by the same or similar reference numerals throughout. The embodiments described below by reference to the accompanying drawings are exemplary and are only used to explain the present application and cannot be understood as limiting the present application.
[0013] In the description of the present application, if there is a description to first, second, etc., it is only for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features or implicitly indicating the order of indicated technical features.
[0014] In the description of the present application, it is to be understood that the orientation description, such as the orientation or position relationship indicated by up, down, etc., is based on the orientation or position relationship shown in the drawings, only for the purpose of facilitating the description of the present application and simplifying the description, and is not to indicate or imply that the device or element indicated must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application.
[0015] In the description of the present application, it should be noted that, unless otherwise explicitly defined, the words such as setting, installing, connecting, etc. should be understood broadly, and the person skilled in the art can determine the specific meaning of the above words in the present application in combination with the specific content of the technical solution.
[0016] Since the detection accuracy of the existing technology for tea impurities is relatively low, in order to solve the problems existing in the prior art, the present application provides a tea impurity visual detection method, system, electronic device and storage medium.
[0017] Referring to Figure 1 , the flowchart of the tea impurity visual detection method provided by the embodiments of the present application. The tea impurity visual detection method is applied to an electronic device, which can be a server or a mobile terminal, etc. As Figure 1 shown, the tea impurity visual detection method can include the following steps: Step S101, acquiring a to-be-detected tea image and a training image data set, pre-processing the image data in the training image data set to obtain a pre-processed training image data set, and the to-be-detected tea image containing a plurality of to-be-detected targets; Step S102, inputting the to-be-detected tea image into a trained tea impurity detection model to obtain a plurality of tea impurity detection results containing confidence and detection boxes of to-be-detected targets, the trained tea impurity detection model being obtained by training the pre-processed training image data set, the trained tea impurity detection model including a backbone network, a neck network and a plurality of detection heads, the backbone network including a ScharrStem module, a first SGCB improvement module and a second SGCB improvement module, the neck network including the first SGCB improvement module, the second SGCB improvement module and a channel perception spatial modulation module, wherein, extracting features of the to-be-detected tea image through the ScharrStem module to obtain a feature extraction result; performing multi-scale feature extraction on the feature extraction result through a plurality of stages containing the first SGCB improvement module or the second SGCB improvement module to obtain a multi-scale feature map; processing the multi-scale feature map through the first SGCB improvement module, the second SGCB improvement module and the channel perception spatial modulation module to obtain a plurality of target feature fusion results; detecting each target feature fusion result using a plurality of detection heads to obtain a plurality of tea impurity detection results; Step S103, determining a target tea impurity detection result according to the confidence and the detection box in the plurality of tea impurity detection results.
[0018] In the embodiment, the ScharrStem module can enhance the expression of the contour features of the tiny impurities, provide a high-quality detail feature base for subsequent deep processing; the SGCB improvement module can effectively prevent semantic attenuation of the tiny impurity features in multiple convolution transformations, and ensure that the sensitivity to fine targets can still be maintained in the deep network; the channel perception spatial modulation module can realize adaptive fusion of different level features, and significantly improve the feature discrimination ability and positioning accuracy of the model for multi-scale impurities. Therefore, the tea impurity detection result obtained by the trained tea impurity detection model can improve the accuracy of the detection result, and the target tea impurity detection result is determined according to the confidence and the detection frame in the plurality of tea impurity detection results, further improving the detection accuracy of the tea impurities.
[0019] The above-mentioned pre-processing of image data in the training image data set can be pre-processing of image data in the training image data set by rotation, cropping, illumination adjustment, scaling, and noise addition.
[0020] The above-mentioned ScharrStem module can be a combination of the Scharr operator in traditional image processing and deep learning feature extraction, establishing an edge enhancement mechanism in the initial stage of feature extraction.
[0021] The above-mentioned SGCB improvement module can be a module obtained by integrating the SGCB sub-module into the original feature extraction unit C3K2 sub-module of YOLOv11.
[0022] The above-mentioned channel perception spatial modulation module can be a feature fusion module combining channel attention and spatial attention.
[0023] The above-mentioned determination of the target tea impurity detection result according to the confidence and the detection frame in the plurality of tea impurity detection results can be removing redundant frames by using a scale-aware soft non-maximum suppression algorithm to determine the target tea impurity detection result according to the confidence and the detection frame in the plurality of tea impurity detection results.
[0024] In some embodiments, the ScharrStem module is used for feature extraction of the tea image to be detected to obtain a feature extraction result, including: performing convolution operation on the tea image to be detected to obtain an initial feature map; performing Scharr filtering on the initial feature map to obtain an edge-enhanced feature map; fusing the edge-enhanced feature map with the tea image to be detected to obtain a first fusion result; performing twice convolution operation on the first fusion result to obtain an intermediate feature representation; performing Gaussian smoothing processing on the intermediate feature representation to obtain a smoothing result; Fuse the smoothing result with the intermediate feature representation to obtain a second fusion result; Perform a convolution operation on the second fusion result to obtain a feature extraction result.
[0025] In this embodiment, by combining the Scharr operator in traditional image processing with deep learning feature extraction, an edge enhancement mechanism is established in the initial stage of feature extraction, effectively strengthening the contour feature expression of micro impurities, and providing high-quality detailed feature base for subsequent deep processing.
[0026] The Scharr filtering on the initial feature map can be Scharr filtering on the initial feature map using a Scharr operator.
[0027] The Gaussian smoothing processing on the intermediate feature representation can be Gaussian smoothing processing on the intermediate feature representation using a 5x5 Gaussian filter.
[0028] In some embodiments, the feature extraction result is subjected to multi-scale feature extraction through multiple stages containing the first SGCB improvement module or the second SGCB improvement module, to obtain a multi-scale feature map, including: The feature extraction result is input to a first stage containing the first SGCB improvement module for feature extraction, to obtain a first scale feature map; The first scale feature map is subjected to down-sampling, batch normalization and SiLU activation function processing, to obtain a first processing result; The first processing result is input to a second stage containing the first SGCB improvement module for feature extraction, to obtain a second scale feature map; The second scale feature map is subjected to down-sampling, batch normalization and SiLU activation function processing, to obtain a second processing result; The second processing result is input to a third stage containing the second SGCB improvement module for feature extraction, to obtain a third scale feature map; The third scale feature map is subjected to down-sampling, batch normalization and SiLU activation function processing, to obtain a third processing result; The third processing result is input to a fourth stage containing the second SGCB improvement module for feature extraction, to obtain a fourth scale feature map.
[0029] In this embodiment, by performing multi-scale feature extraction on the feature extraction result through multiple stages containing the first SGCB improvement module or the second SGCB improvement module, a feature map with different semantics and details in each stage can be obtained, laying a good data foundation for later tea impurity detection.
[0030] In some embodiments, the first SGCB improvement module and the second SGCB improvement module comprise an SGCB submodule and a C3K2 submodule, the SGCB submodule comprises a plurality of convolutional layers and an improved Scharr convolutional layer; the SGCB submodule is added to the C3K2 submodule to obtain the SGCB improvement module, when the C3K2 submodule in the SGCB improvement module does not use the C3k submodule for feature enhancement, the SGCB improvement module is the first SGCB improvement module; when the C3K2 submodule in the SGCB improvement module uses the C3k submodule for feature enhancement, the SGCB improvement module is the second SGCB improvement module; the improved Scharr convolutional layer comprises the following processing process: performing a channel-by-channel Scharr gradient calculation on the input feature map to obtain an x-direction gradient map and a y-direction gradient map of each channel; performing gradient amplitude fusion on the x-direction gradient map and the y-direction gradient map to obtain an original edge response map of each channel; performing a channel-by-channel 3x3 neighborhood variance calculation and normalization processing on the input feature map to obtain a normalized variance map; calculating the weight of each channel according to the normalized variance map; performing a pixel-by-pixel weighting on the original edge response map corresponding to each channel through the weight of each channel to obtain an enhanced edge feature map of each channel; integrating the enhanced edge feature map of each channel in the original channel order to obtain a final output feature map.
[0031] In this embodiment, the SGCB submodule comprises a plurality of convolutional layers and an improved Scharr convolutional layer, by constructing a parallel processing path in the deep network, the accuracy of edge detection can be preserved, the continuous preservation and strengthening of detailed features can be realized, the semantic decay of small impurity features in multiple convolutional transformations can be effectively prevented, and the sensitivity to fine targets can be ensured in the deep network.
[0032] In some embodiments, the multi-scale feature maps are processed through the first SGCB improvement module, the second SGCB improvement module and the channel-aware spatial modulation module to obtain a plurality of target feature fusion results, comprising: performing feature extraction on the fourth-scale feature map to obtain a first feature image, performing an upsampling operation on the first feature image to obtain an upsampled first feature image; performing dimension stacking on the third-scale feature map and the upsampled first feature image to obtain a first stacking result, inputting the first stacking result into the first SGCB improvement module for feature fusion to obtain a feature fusion result; performing an upsampling operation on the feature fusion result to obtain an upsampled second feature image, performing dimension stacking on the second-scale feature map and the upsampled second feature image to obtain a second stacking result; The first scale feature map is down-sampled to obtain a third down-sampled feature image, and the third down-sampled feature image and the second superimposed result are input into the channel perception spatial modulation module for feature fusion to obtain a first enhanced feature map; the first enhanced feature map is input into the first SGCB improved module for feature fusion to obtain a first target feature fusion result; The first enhanced feature map is down-sampled to obtain a fourth down-sampled feature image, and the fourth down-sampled feature image and the feature fusion result are input into the channel perception spatial modulation module for feature fusion to obtain a second enhanced feature map; the second enhanced feature map is input into the first SGCB improved module for feature fusion to obtain a second target feature fusion result; The second target feature fusion result is down-sampled to obtain a fifth down-sampled feature image, and the fifth down-sampled feature image and the first feature image are input into the channel perception spatial modulation module for feature fusion to obtain a third enhanced feature map; the third enhanced feature map is input into the second SGCB improved module for feature fusion to obtain a third target feature fusion result.
[0033] In the embodiment, the SGCB improved module and the channel perception spatial modulation module are used for processing, which can effectively prevent semantic attenuation of micro impurity features in multiple convolution transformations, ensure that the sensitivity to fine targets in the deep network is maintained, and significantly improve the feature discrimination ability and positioning accuracy of the model for multi-scale impurities.
[0034] The above plurality of target feature fusion results can be target feature fusion results corresponding to high-scale feature maps. For example, the embodiment includes a first scale feature map, a second scale feature map, a third scale feature map, and a fourth scale feature map. The plurality of target feature fusion results can include a target feature fusion result corresponding to each of the second scale feature map, the third scale feature map, and the fourth scale feature map, i.e., a target feature fusion result corresponding to a high-scale feature map.
[0035] In some embodiments, the third down-sampled feature image and the second superimposed result are input into the channel perception spatial modulation module for feature fusion to obtain the first enhanced feature map, including: The third down-sampled feature image is taken as a low-level feature, and the second superimposed result is taken as a high-level feature; The low-level feature and the high-level feature are element-level added to obtain a first addition result; The first addition result is input into the channel attention module for global average pooling and convolution operation to obtain a first convolution feature; The first convolution feature is subjected to an activation function and a convolution operation to obtain a channel attention weight; pooling the first addition result to obtain a first pooling result, and average-pooling the first addition result to obtain a second pooling result; performing convolution operation on the first pooling result and the second pooling result after splicing to obtain a second convolution feature; performing convolution and activation function processing on the first convolution feature to obtain a third convolution feature; multiplying the third convolution feature after being expanded by a broadcast mechanism with the second convolution feature to obtain a collaborative spatial attention weight; performing element-level addition on the collaborative spatial attention weight and the channel attention weight to obtain a second addition result, splicing the first addition result and the second addition result to obtain a spliced result; performing grouped channel rearrangement on the spliced result to obtain a fusion weight, and performing convolution and activation function processing on the fusion weight to obtain a spatial modulation weight; determining a weighted feature according to the spatial modulation weight, the low-level feature and the high-level feature; performing element-level addition on the weighted feature, the low-level feature and the high-level feature to obtain a third addition result; performing convolution operation on the third addition result to obtain a first enhanced feature map.
[0036] In this embodiment, by processing through the channel-aware spatial modulation module, adaptive fusion of different level features can be achieved, and the feature discrimination ability and positioning accuracy of the model for multi-scale impurities can be significantly improved.
[0037] In some embodiments, determining a target tea impurity detection result according to a confidence and a detection box in a plurality of tea impurity detection results comprises: determining an empty scale-aware overlapping area ratio between two detection boxes according to the detection boxes in the plurality of tea impurity detection results; constructing a Gaussian confidence decay function based on the empty scale-aware overlapping area ratio and the confidence, the Gaussian confidence decay function being used to update the confidence; removing redundant boxes based on the detection boxes, the Gaussian confidence decay function and the confidence to obtain the target tea impurity detection result.
[0038] In this embodiment, by removing redundant boxes based on the detection boxes, the Gaussian confidence decay function and the confidence, the "redundant box strong suppression and effective box precise reservation" can be achieved through accurate quantification of spatial redundancy on the basis of retaining the advantages of adaptation to micro targets, and the detection accuracy and real-time performance can be taken into account by combining the spatial partitioning acceleration strategy.
[0039] The target tea impurity detection result can be obtained by removing redundant frames based on the detection frame, the Gaussian confidence attenuation function and the confidence.
[0040] For the convenience of those skilled in the art, a set of best embodiments is provided below: The embodiment provides a tea impurity visual detection method, which comprises the following steps: S1, obtaining a tea image after winnowing and performing pretreatment.
[0041] The tea impurity image dataset is constructed by collecting tea and impurity images of a production plant. An industrial camera with a resolution of 2448*2048 is used for shooting, simulating the state of tea impurities on the actual production line, covering tea and impurity photo images under different conditions, and ensuring the diversity and representativeness of the data.
[0042] The image dataset consists of 439 images, including 6 categories: tea tender leaves, tender stems, wax leaves, old stems, weeds and plastic strips. Then, the dataset is labeled using the Labelme toolbox, and the dataset is divided into a training set, a test set and a validation set according to a ratio of 7:2:1. The training set is preprocessed, and the embodiment adopts data enhancement techniques such as rotation, cropping, light adjustment, scaling and noise addition to expand the training set data to 1305, thereby improving the generalization ability of the model. Finally, the preprocessed tea impurity image dataset (i.e., the preprocessed training image dataset) is obtained.
[0043] Specifically, the preprocessing of the embodiment comprises the following steps: S1.1, in the data acquisition stage, 439 high-resolution images of 2448*2048 are systematically collected under standard lighting conditions using an industrial camera, which completely covers the typical morphological characteristics of six categories of impurities (tea tender leaves, tender stems, wax leaves, old stems, weeds and plastic strips) on the tea production line. To ensure the quality and representativeness of the dataset, the embodiment establishes strict sample preparation specifications, accurately controls the tea laying and stacking, configures the impurity distribution according to the actual production ratio, and ensures that each image contains 3 to 5 different impurities to simulate the complex scene of the real production line.
[0044] S1.2 In the data annotation stage, this embodiment organized a professional annotation team to perform refined polygon annotations on all images using the Labelme toolbox. For partially occluded targets, a visible portion annotation strategy was adopted, and all annotation results underwent three rounds of cross-validation to ensure annotation accuracy. Based on rigorous data partitioning principles, this embodiment randomly divided the dataset into training, testing, and validation sets in a 7:2:1 ratio. In particular, some of the most challenging high-overlap, low-contrast samples were reserved for the validation set to fully evaluate the model's extreme performance.
[0045] S1.3 To enhance the model's generalization ability, this embodiment designed and implemented a multi-level data augmentation pipeline. At the geometric transformation level, random rotation (±15° range), perspective transformation (0.1 affine amplitude), and random cropping (0.7 to 1.0 scale) were systematically applied. At the photometric transformation level, this embodiment simulated complex lighting environments through brightness adjustment (±20%), contrast adjustment (±15%), gamma correction (0.8 to 1.2 range), and hue and saturation perturbation. At the noise simulation level, this embodiment introduced Gaussian noise (σ range of 0.01 to 0.05) and salt-and-pepper noise (density of 0.1% to 0.5%) to enhance the model's robustness to image degradation. In addition, this embodiment also adopted advanced augmentation strategies such as CutMix blending (20% to 40% blending ratio) and Mosaic four-image stitching, significantly improving the model's ability to detect dense small targets.
[0046] S1.4 After this complete preprocessing process, the original training set was expanded from 307 valid images to 1305 high-quality training samples. At the same time, this embodiment established a strict quality control mechanism, which eliminated blurry images through Laplacian variance detection, removed invalid annotations with an area of less than 25 square pixels, and ensured the visual rationality and annotation accuracy of the enhanced images.
[0047] S2. Based on the YOLOv11 target detection model, construct an improved TSCM-Det model (i.e., tea impurity detection model).
[0048] This embodiment maintains high inference speed while deeply optimizing for the specific challenges of detecting minute impurities in tea images. The overall network model architecture consists of three core parts: the backbone, the feature fusion neck, and the detection head. After preprocessing and data augmentation, the input image is fed into the backbone network for feature extraction.
[0049] Firstly, the original Stem down-sampling module is replaced by a specially designed Scharr Stem module, which introduces the Scharr differential operator with rotational invariance and a multi-scale Gaussian fusion mechanism to construct a cognitive-driven paradigm based on differential geometric prior at the front end of feature extraction, strengthening edge information and suppressing noise in the initial stage of feature extraction, laying a foundation for subsequent detail detection.
[0050] Secondly, the original C3K2 module is replaced by an innovative SGCB improved module, which constructs a deep feature enhancement mechanism based on dual-path cognitive collaboration. By using the Scharr prior guide with differential invariance in the deep network, the dynamic balance between detail perception and semantic understanding is achieved. The SGCB module and the preposed Scharr Stem module jointly constitute a "detail enhancement pipeline" from input to deep layer. The Scharr Stem module performs a one-time, intensive detail extraction and reservation at the front end of the network, while the SGCB module continuously maintains and enhances the details at multiple subsequent levels. This collaborative working mode ensures that features sensitive to small tea impurities can be transmitted throughout the forward propagation process of the entire network, thereby significantly improving the detection robustness and accuracy of the model in complex tea backgrounds.
[0051] The feature fusion neck adopts a feature fusion module (CASM) based on channel perception spatial modulation, which understands the semantic roles of different feature channels through channel perception mechanism, and then implements precise spatial modulation to achieve adaptive fusion of low-level detail features and high-level semantic features.
[0052] In the post-processing stage, to address the problem of target high overlap and fragmented prediction boxes at the boundary in tea images, a spatial-scale-aware soft non-maximum suppression algorithm is used for redundant box removal. This method combines Spatial-Scale Aware Intersection over Area (SSAIoA) and soft non-maximum suppression strategy, optimizes the overlap degree metric by fusing spatial position perception factor and scale matching perception factor, and designs a Spatial-Scale Aware Soft-NMS (SSA-SoftNMS) based on this. This algorithm, while retaining the advantages of small target adaptation, achieves "strong suppression of redundant boxes and precise preservation of effective boxes" by accurately quantifying spatial redundancy, while combining spatial partitioning acceleration strategy to balance detection accuracy and real-time performance.
[0053] Specifically, referring to Figure 2 , the TSCM-Det model mainly includes a backbone network, a neck network, and a detection head, and the specific processing flow includes: 1、The input image (which can be the image of tea to be detected) is a three-channel RGB color image with a size of 640x640x3. The input image is first processed by the ScharrStem module, which mainly functions to process the input image in the initial stage of the model. By combining the scharr edge detection operator and Gaussian smoothing denoising, the edge and detail information are actively strengthened in the initial stage of the network. Instead of the traditional network violent downsampling operation, by actively strengthening the edge and detail information in the initial stage, the problem of noise and edge information degradation caused by layer transmission in the network layer is solved, and high-quality feature representation is provided for subsequent detection tasks. The ScharrStem module outputs a feature image with a size of 160x160x32 (i.e., the feature extraction result).
[0054] 2、The feature is then extracted by the first SGCB improvement module improved in stage 1, and a feature image with a size of 160x160x64 (i.e., the first scale feature map) is output.
[0055] 3、Then, a 3x3 convolution is performed for downsampling operation, and a feature image with a size of 80x80x64 is output. Then, a BN batch normalization operation and a SiLU (Sigmoid Linear Unit) activation function are performed to improve the nonlinear expression ability of the model.
[0056] 4、Then, the feature is extracted by the first SGCB improvement module improved in stage 2, and a feature image with a size of 80x80x128 (i.e., the second scale feature map) is output.
[0057] 5、Then, a 3x3 convolution is performed for downsampling operation, and a feature image with a size of 40x40x128 is output. Then, a BN batch normalization operation and a SiLU (Sigmoid Linear Unit) activation function are performed to improve the nonlinear expression ability of the model.
[0058] 6、Then, the feature is extracted by the second SGCB improvement module improved in stage 3, and a feature image with a size of 40x40x128 (i.e., the third scale feature map) is output.
[0059] 7、Then, a 3x3 convolution is performed for downsampling operation, and a feature image with a size of 20x20x256 is output. Then, a BN batch normalization operation and a SiLU (Sigmoid Linear Unit) activation function are performed to improve the nonlinear expression ability of the model.
[0060] 8、Then, the feature is extracted by the second SGCB improvement module improved in stage 4, and a feature image with a size of 20x20x128 (i.e., the fourth scale feature map) is output.
[0061] 9、For the high-level features obtained in stage 4 (i.e., the fourth scale feature map), spatial pyramid pooling (SPPF) and C2PSA modules are used to extract multi-scale features, making the network more robust to objects of different sizes, especially large objects. SPPF simulates the effect of a large kernel pooling layer by concatenating multiple identical small kernel (5x5) pooling layers. The C2PSA module is an extension of the C2f module, which incorporates a PSA (Pointwise Spatial Attention) block to enhance feature extraction and attention mechanisms. By introducing a PSA block into the standard C2f module, C2PSA achieves a more powerful attention mechanism, thereby improving the model's ability to capture important features for tea impurity detection. The output feature map size is 20x20x256 (i.e., the first feature image). The C2f module is a module in the YOLO series known to those skilled in the art, and is not described in detail in this embodiment.
[0062] 10、The output feature image in step 9 is upsampled, usually by bilinear interpolation or transposed convolution, to match the spatial dimensions of the output feature image in stage 3 (i.e., the third scale feature map).
[0063] 11、The output feature image in stage 3 is dimensionally stacked with the first feature image after upsampling in step 10 to obtain a first stacking result.
[0064] 12、The first stacking result is further fused by the first SGCB improvement module to obtain a feature fusion result.
[0065] 13、The first feature fusion result is upsampled to match the spatial feature dimensions of the output feature image in stage 2 (i.e., the second scale feature map).
[0066] 14、The output feature image in stage 2 is dimensionally stacked with the second feature image after upsampling in step 13 to obtain a second stacking result.
[0067] 15、The output feature image in stage 1 (i.e., the first scale feature map) is downsampled by a 3x3 convolution to match the feature size of the second stacking result in step 14, i.e., the output feature image in stage 2. Figure 1 .
[0068] 16、The second stacking result and the third feature image after downsampling in step 15 are input into the CSAM module for feature fusion to obtain a first enhanced feature map.
[0069] 17、The first enhanced feature map is further fused by the first SGCB improvement module to obtain a first target feature fusion result.
[0070] 18, the first target feature fusion result is sent to the detection head for detection to obtain the final result (i.e. the first tea impurity detection result).
[0071] 19, the first enhanced feature map of step 16 is subjected to 3x3 convolution for downsampling operation, so that the feature size is consistent with the feature fusion result of step 12.
[0072] 20, the fourth feature image after downsampling of step 19 and the feature fusion result of step 12 are input into the CSAM module for feature fusion to obtain a second enhanced feature map.
[0073] 21, the second enhanced feature map is subjected to further feature fusion by the first SGCB improvement module to obtain a second target feature fusion result.
[0074] 22, the second target feature fusion result is sent to the detection head for detection to obtain the final result (i.e. the second tea impurity detection result).
[0075] 23, the second target feature fusion result of step 21 is subjected to 3x3 convolution for downsampling operation, so that the feature size is consistent with the output feature Figure 1 .
[0076] 24, the fifth feature image after downsampling of step 23 and the output feature map (i.e. the first feature image) of step 9 are input into the CSAM module for feature fusion to obtain a third enhanced feature map.
[0077] 25, the third enhanced feature map is subjected to further feature fusion by the second SGCB improvement module to obtain a third target feature fusion result.
[0078] 26, the third target feature fusion result is sent to the detection head for detection to obtain the final result (i.e. the third tea impurity detection result).
[0079] Wherein, referring to Figure 3 , the Scharr Stem module specifically includes: The input image is applied to 7x7 convolution operation to extract the initial feature map with rich semantic information, the Scharr edge enhancement branch, the initial feature map is respectively applied to horizontal direction Scharr convolution kernel , and vertical direction Scharr convolution kernel Convolution calculation is performed, and the edge enhancement feature map is obtained by gradient amplitude calculation as follows: ; ; wherein, denotes a convolution operation; the edge-enhanced feature map is obtained after fusing with the input image, and the fusion result is applied to two 3x3 convolution layers in sequence for feature transformation and down-sampling to obtain an intermediate feature representation ; The Gaussian smoothing fusion branch applies a 5x5 Gaussian filter (using two sizes of 5x5 and 9x9, respectively) to the intermediate feature representation for multi-scale smoothing processing, wherein the Gaussian filter is defined as: ; wherein, denotes the size of the Gaussian kernel, denotes the standard deviation, denotes the pixel position.
[0080] The multi-scale smoothed feature is fused with the original feature through a normalization operation to obtain a smoothed feature : ; The smoothed feature is finally input into a 3x3 convolution down-sampling module to compress the spatial size to 1 / 4 of the input, and output the final feature map (i.e., the feature extraction result) of the ScharrStem module , denotes the height, denotes the width, denotes the number of channels.
[0081] wherein, referring to Figure 4 , the SGCB improvement module specifically includes: The SGCB submodule is integrated into the original feature extraction unit C3K2 submodule of YOLOv11 through a structured embedding strategy to construct an SGCB improvement module with hierarchical detail preservation capability. The SGCB submodule is specifically as follows: The input feature map is projected onto a high-dimensional manifold through a 1x1 convolution kernel, and then two orthogonal feature subspaces are generated through a symmetric partition operation in the channel dimension , forming the input elements of the parallel processing path. The edge-aware branch calculates the gradient amplitude response of the feature space through an improved Scharr convolution layer to generate a detail-enhanced feature ; the auxiliary path adopts the same topology structure and is processed through a standard 3x3 convolution layer, aiming to learn high-level semantic representations to generate semantic features . This branch is responsible for maintaining the network's understanding ability of complex scenes and different types of impurities.
[0082] The output features of the two branches are then concatenated and The concatenation is performed along the channel dimension and fused and down-sampled by a 1x1 convolutional layer. The 1x1 convolution acts as a feature fuser whose task is to learn how to optimally combine the precise edge information from the detail branch and the contextual information from the semantic branch adaptively. The fusion process can be represented as: ; Finally, the fused features are added to the input of the module through a shortcut connection, forming a residual learning structure: ; wherein, denotes the concatenation operation, denotes the 1x1 convolution operation.
[0083] After the residual learning, the features are processed by convolution to obtain the final output result of the SGCB submodel .
[0084] For the SGCB sub-module, the original bottleneck module is improved by introducing an improved Scharr convolution layer for edge detail preservation and enhancement, to ensure that the features sensitive to small tea impurities can be transmitted throughout the forward propagation process of the entire network, thereby significantly improving the detection robustness and accuracy of the model in complex tea background. Then it is added to the original C3K2 sub-module. Referring to Figure 4 , for the original C3K2 sub-module, there are two different modes, one is not to use the C3k sub-module for feature enhancement, i.e. SGCB module=False, which is the first SGCB improved module; one is to use the C3k sub-module for feature enhancement, i.e. SGCB module=True, which is the second SGCB improved module.
[0085] The C3K2 sub-module and the C3k sub-module are modules in the YOLO series known to those skilled in the art, and the present embodiment does not make specific description.
[0086] The improved Scharr convolution layer specifically includes the following contents: The traditional Scharr convolution detects edges by calculating the local gradient amplitude, but cannot distinguish between "true edge gradient" and "noise gradient", and its weights are fixed values designed by hand, cannot be adaptively adjusted, have weak generalization ability, and cannot adapt to the feature specificity of each channel.
[0087] To solve the above problems, the embodiment proposes a Scharr convolution based on local variance weighting (i.e. improved Scharr convolution layer). The core goal is to solve the core pain points of traditional Scharr convolution (anti-noise difference and weak generalization ability), while retaining the accuracy of edge detection, making edge feature extraction more suitable for complex scenes (such as noise interference, abstract feature maps, and multi-channel features). The core improvements are as follows: 1. Accurately distinguish edges and noise by local variance to suppress false responses.
[0088] For edge areas: the feature value difference in a 3x3 neighborhood is large (such as the junction of an object and the background), the variance is large → the weight is large → the gradient response is enhanced (highlighting the real edge); noise area: the feature values in a 3x3 neighborhood are discrete but small in range (such as isolated noise points), the variance is moderate → the weight is moderate → the gradient response is suppressed (filtering false edges); smooth area: the feature values in a 3x3 neighborhood are uniform, the variance is small → the weight is small → slight suppression (avoiding false edges).
[0089] Through this "variance-weight" mapping, adaptive adjustment of "edge enhancement and noise suppression" is achieved, solving the false edge problem of traditional Scharr from the source.
[0090] 2. Adaptively adapt to different scenes and feature maps.
[0091] For different images: automatically adapt to noise intensity (more areas are suppressed for images with high noise; more areas are enhanced for clear images); for intermediate layer features in CNN: adapt to the structure pattern of abstract features (such as abstract edges of deep features, which can be accurately captured); for multi-channel features: calculate the variance and weight for each channel independently, and adapt to the feature specificity of each channel (such as channel A suppressing noise and channel B enhancing edges).
[0092] 3. Introduce learnable parameters to improve end-to-end training results.
[0093] Introduce a learnable noise threshold (or set an independent noise threshold for each channel): It will be automatically optimized during training, learning the critical values of "noise variance" and "edge variance" in the data set, further improving the discrimination accuracy; the specific steps are as follows: Stage 1: Channel-by-channel Scharr gradient calculation and amplitude fusion.
[0094] Let the input multi-channel feature map be , where represents the channel (C) ) in the coordinate ( original pixel feature value of the input image. First, for each pixel Scharr gradient calculation is performed, and the horizontal gradient and vertical gradient kernel distribution is: ; Then, channel-wise gradient calculation is performed (convolution can be directly performed on the original feature map), and group convolution (group = C) is used to individually convolve each original channel to obtain the gradient map in the x and y directions and , where represents two-dimensional convolution, and padding = 1 keeps the gradient map consistent with the original map size.
[0095] ; ; Then, gradient amplitude fusion is performed, and L2 norm calculation is performed on the x and y direction gradient maps to obtain the original edge response map : where is to avoid errors caused by a gradient of 0.
[0096] ; Stage 2: Channel-wise 3x3 neighborhood variance calculation and normalization.
[0097] Local variance is a core indicator for measuring the local dispersion degree of a feature map. The edge region has a large variance due to the sudden change of feature values, the noise region has a moderate variance due to isolated fluctuations, and the smooth region has a small variance due to uniform features. In this stage, the 3x3 neighborhood variance of each pixel in each channel is calculated, and the numerical scale interference is eliminated through normalization.
[0098] First, the mean value is calculated for each domain: define a 3x3 mean value kernel , and perform convolution operation on each channel to obtain the domain mean value map , where the mean value of the coordinate is: the convolution also adopts the padding = 1 strategy to ensure that the mean value map matches the spatial dimension of the original channel, providing a basis for subsequent variance calculation.
[0099] ; Then, the local variance derivation is performed: the local variance is derived using the mathematical properties of variance (where represents expectation) to avoid inefficient calculation of pixel-by-pixel domain traversal, and the specific steps are as follows: ① Calculate the square of the feature value map: ; ② Calculate the square of the domain mean value map: ; ③ Derive local variance map : ; ④ Perform truncation operation on the result to ensure non-negativity of variance due to floating-point calculation errors: , ensuring the non-negativity of variance.
[0100] Then normalize the variance: To eliminate the influence of the difference in the numerical scale of variance between different channels on subsequent weight calculation, normalize the variance map of each channel separately, mapping to the [0, 1] interval: ; where, and represent the minimum and maximum values of the variance map of the channel, respectively. The normalized variance map preserves the relative distribution characteristics of the original variance.
[0101] Stage 3: Generation of per-channel weight map.
[0102] The core role of the weight map is to dynamically adjust the response strength of the amplitude according to the local variance, achieving the goal of "edge enhancement-noise suppression". In this stage, based on the normalized variance map, a per-pixel adaptive weight is generated combined with learnable parameters.
[0103] First, calculate the weight formula as follows: use a nonlinear mapping function to convert the normalized variance to weight, and the formula is defined as: ; where, is a learnable noise threshold parameter (initial value set to 0.01), which serves to calibrate the discriminant sensitivity of the variance. When is small (adapt to low-noise data), the area with slightly larger variance can quickly make the weight tend to 1; when is large (adapt to high-noise data), the area with significantly larger variance can make the weight tend to 1, avoiding noise false enhancement. Global sharing (single parameter) or per-channel sharing (c parameters) strategies can be used, and the specific choice can be made according to the degree of channel difference, and optimized to the optimal value through end-to-end training.
[0104] Stage 4: Weighted weighting and enhanced edge feature output.
[0105] In this stage, the generated per-channel weight map is multiplied pixel by pixel with the gradient amplitude map to complete the adaptive enhancement of edge features, and finally the multi-channel enhanced feature map is output. First, perform weighted fusion calculation, and perform pixel-by-pixel weighting of the weight and the original edge response map for each channel: ; wherein, represents the enhanced edge feature map of the c-th channel, through weight adjustment, the gradient response of the edge region is retained, and the gradient response of the noise and smooth region is inhibited.
[0106] Finally, the enhanced edge feature maps of the c channels are integrated in the original channel order to obtain the final output feature map , which has the same spatial dimension as the input feature map.
[0107] This design not only alleviates the gradient vanishing problem in deep networks, but also enables the SGCB sub-module to focus on learning the residual part of the input feature, i.e. the incremental information related to the details of tea impurities. The SGCB sub-module and the preposed ScharrStem module jointly constitute a "detail enhancement pipeline" from input to deep layer. ScharrStem performs a one-time, intensive detail extraction and retention at the front end of the network, while the SGCB sub-module continuously maintains and enhances the details at multiple subsequent levels. This collaborative working mode ensures that the features sensitive to small tea impurities can be transmitted throughout the forward propagation process of the entire network, thereby significantly improving the detection robustness and accuracy of the model in complex tea background.
[0108] wherein, referring to Figure 5 , the channel-aware spatial modulation module (CASM module) specifically comprises: The design of the CASM module follows the idea of "guiding-modulating-refining", and its overall architecture is a multi-stage processing pipeline with clear data flow. Let the two input feature maps to be fused be low-level detail features (i.e. low-level features) and high-level semantic features high-level semantic features (i.e. high-level features).
[0109] Stage 1: Feature initialization and collaborative attention generation.
[0110] Feature initialization: first, a basic fusion feature is obtained through element-level addition as the input for subsequent attention calculation (i.e. the first addition result): This operation ensures that the calculation of subsequent attention can be carried out in a fused context. The generation of parallel channel attention generates a guiding signal. The goal of this path is not only to produce channel weights, but also to generate a guiding signal for spatial attention calculation.
[0111] First, for the fusion feature , perform compression along the spatial dimension to obtain the embedding of global information and generate a channel-level statistical description, i.e. global average pooling along the channel. Wherein is: ; Then a MLP with dimension reduction function is used to learn the complex nonlinear relationship between channels. First, dimension reduction is performed by 1x1 convolution, and the features are refined.
[0112] ; wherein, is the dimension reduction weight matrix, the first convolution feature It is used as a channel guide signal, which is a compact and abstract representation of the original channel information, represents the bias term.
[0113] Then is dimensioned by ReLU activation function and 1x1 convolution to obtain the final channel attention weight : ; wherein, , .
[0114] Generation of parallel spatial attention. First, global max pooling and average pooling are used to capture the extreme response and average response of the feature map, forming a preliminary description of the spatial features.
[0115] ; ; wherein, represents the first pooling result, represents the second pooling result, and the two are spliced in the channel dimension to obtain .
[0116] Then a 3x3 convolution is used to preliminarily model the spatial context of the spliced statistical features, embed local context information, and project it into a higher dimensional feature space.
[0117] ; wherein, is a 3x3 convolution layer, and the output channel number is M, so the second convolution feature .
[0118] Then the guide signal generated in the channel attention path is projected into a feature space matching the feature by a transformation network composed of two 1x1 convolutions.
[0119] ; ; wherein, represents the feature after activation function processing, the first weight matrix , the second weight matrix , is a Sigmoid activation function, which ensures that the output value is in the interval [0, 1], and are bias terms. The third convolutional feature , i.e. spatial modulation gate.
[0120] Then the gating signal is expanded to the spatial dimension through a broadcast mechanism, and is element-wise multiplied with the primary spatial feature. The physical meaning of this step is that according to the importance of each channel, the activation value of each position in the spatial feature is dynamically recalibrated. The important feature channel is reset, and its corresponding spatial context information is enhanced, and the unimportant one is suppressed.
[0121] ; Finally, the modulated feature is refined to map it to a single-channel spatial weight map: ; wherein, represents a broadcast mechanism, is a 3x3 convolution with an output channel of 1, is a synergistic spatial attention weight that integrates channel information.
[0122] Stage two: core modulation stage.
[0123] First, the spatial attention weight is element-wise added to the channel attention weight to achieve primary feature fusion, and then a channel shuffling operation is introduced to promote information exchange between different channel groups by grouping channel rearrangement, breaking the limitations of traditional grouping processing: ; wherein, represents grouping channel rearrangement. Grouping channel rearrangement is a technique that constructs a new feature map by changing the arrangement order of channels in a tensor without changing the total data amount of the tensor (i.e. without increasing the calculation and parameters). It breaks the information isolation brought by grouped convolution and promotes information exchange between different feature groups. Grouping channel rearrangement uses channel rearrangement known to those skilled in the art, which is not specifically described in this embodiment.
[0124] The fused weight after channel rearrangement is subjected to spatial refinement processing, two 3x3 convolutions are used for deep feature extraction, and finally a Sigmoid activation function Numerical normalization is performed to generate spatial modulation weights specific to the channel: ; Based on the learned modulation weights, precise spatial modulation is performed on the input features. A complementary weighting strategy is used, where low-level detail features are multiplied by the modulation weights, while high-level semantic features are weighted by the weight complements: ; To preserve the integrity of the original features, a residual connection mechanism is introduced to add the weighted features to the original input features. Finally, a 1x1 convolution is used to integrate the features and unify the dimensions, resulting in an enhanced feature map : .
[0125] In the post-processing stage, a multi-dimensional evaluation system and an adaptive decay function are constructed to remove redundant boxes, which includes: In the target detection task, Non-Maximum Suppression (NMS) and its variants are the core post-processing techniques for removing redundant detection boxes. Traditional NMS relies on the Intersection over Union (IoU) to measure overlap, but it has significant limitations in tea impurities and other scenarios: first, IoU only focuses on the overlap area ratio of two boxes, and cannot distinguish between "center-aligned redundant overlap" and "edge-shifted effective neighbor"; second, it is sensitive to "large box containing small box" scenarios, and is prone to false suppression of small effective targets due to high IoU values; third, it does not consider the coupling effect of target spatial position and scale difference, resulting in a lack of targeted suppression strategy.
[0126] To solve the above problems, Spatial-Scale Aware Intersection over Area (SSAIoA) is proposed, which optimizes the overlap measure by integrating spatial position awareness factors and scale matching awareness factors, and designs a Spatial-Scale Aware Soft-NMS (SSA-SoftNMS) algorithm based on this. This algorithm, while retaining the advantages of adapting to small targets, achieves "strong suppression of redundant boxes and precise preservation of effective boxes" by accurately quantifying spatial redundancy, while combining spatial partitioning acceleration strategies to balance detection accuracy and real-time performance. The specific algorithm flow is as follows: I. Algorithm inputs include: A set of detection boxes Each box Uses the YOLO output format: Where is the normalized center coordinate, is the normalized width and height, Category ID; confidence set Each detection box corresponds to a separate detection box. Original image size Preset hyperparameter: area threshold SSAIoA parameters Stratification threshold Gaussian decay parameters .
[0127] Spatial Scale Perceived Overlap Area Ratio (SSAIoA) calculation, the core of SSAIoA is through spatial perception factors ( ) and scale-sensing factor ( The original IoA was corrected, and the spatial redundancy of the two boxes was precisely quantified, specifically defined as: ; in This is the original area coverage ratio. To measure the spatial alignment of the centers of the two frames, The degree of scale matching between the two frames is measured, and ultimately The larger the value, the higher the redundancy, and the more it needs to be suppressed.
[0128] 1. Basic IoA Calculation: ; in, This represents a function for finding the area. The bounding box currently being evaluated can be called a candidate box; Represents a confidence ratio A higher-confidence bounding box can be called a "high-confidence bounding box". This represents the area of the intersection of the two boxes. This represents the area of the candidate box itself.
[0129] 2. Spatial perception factors Calculation: The spatial perception factor is used to quantify the spatial alignment of the centers of two frames. Its core logic is that "the closer the center distance, the higher the possibility of spatial redundancy," and a Gaussian decay function is used to achieve smooth perception. ; in: It directly reflects the spatial offset between the centers of the two frames; Normalized distance: , The length of the image diagonal; Location smoothing factor (recommended) ), which controls the decay rate of position offset. The smaller the value, the greater the center offset. The stronger the weakening effect, the higher the factor value will be for boxes that are center-height aligned.
[0130] 3. Scale-sensing factor Calculation: The scale-aware factor is used to quantify the degree of size matching between two bounding boxes. Its core logic is that "the closer the sizes are, the more likely they are repeated detections of the same target; the greater the size difference, the more likely they are independent targets." It achieves accurate perception through exponential enhancement.
[0131] ; in, For area similarity, the value range is... The closer the sizes are, the closer the value is to 1; the greater the size difference, the smaller the value. For the scale enhancement factor, a recommended value is [value]. The weighting effect of amplified size differences.
[0132] The perceptual effect of size differences is amplified through exponential operations. For example, two boxes with an area similarity of 0.9 (close in size) will appear larger after exponential operations. The exponential calculation yields a similarity of 0.81, while the similarity of 0.5 (due to large size differences) yields a similarity of 0.25. Therefore, this enhances the ability to distinguish scenes with "large frames containing smaller frames".
[0133] Hierarchical dynamic threshold: based on candidate boxes pixel-level area The hierarchical threshold is used to adapt to the scale non-uniformity characteristics: ; in, (Recommended value: 0.85) (Recommended value: 0.65) (The recommended value is 0.45) is the stratification threshold for adapting to SSAIOA.
[0134] Gaussian confidence decay function: if Then for The confidence level is smoothly decayed to avoid the loss of valid targets due to hard deletion: ; in, For the decay smoothing factor (recommended) ), The confidence level after attenuation; if If the confidence level remains unchanged, the function replaces the traditional hard deletion with a smooth confidence decay, effectively preserving the second-highest confidence but actually effective detection targets in dense regions while maintaining detection accuracy.
[0135] The detailed execution steps of the specific SSASoft NMS algorithm include: (1) Input and initialization: Input the detection frame set and the confidence set , create an empty list D for storing the final retained detection results, and create a processing queue Q, which stores all the indexes of the detection frames in order from high to low according to the confidence .
[0136] (2) Loop processing: When the queue Q is not empty: a. Select the current highest confidence detection frame: pop out the detection frame with the highest confidence from Q ; b. Add it to the final result: add and to the list D; c. Traverse the remaining frames for suppression: for each frame remaining in Q : i. Calculate the empty scale perception overlap area ratio matrix: calculate the coverage of on ; ii. Determine the dynamic threshold: determine the threshold used by according to the area of : 1) If , then ; 2) If , then ; 3) If , then ; iii. Implement confidence decay based on Gaussian kernel: If , then , otherwise remain unchanged.
[0137] iv. Dynamically update the confidence distribution of the detection frame: update the confidence score of the frame .
[0138] d. Clean up the queue: remove all detection frames in Q whose confidence S is lower than the preset final threshold (such as 0.01 or 0.001).
[0139] e. Reorder the queue: reorder the queue Q according to the updated confidence from high to low.
[0140] (3) The optimized detection result is output through a double verification mechanism.
[0141] The first re-verification and the main loop itself are a confidence-based overlap verification.
[0142] The second re-verification, for the results in the final list D, can impose a final confidence threshold (such as 0.5), and only the bounding boxes with scores higher than the threshold will be finally output, which ensures that all output results have high reliability.
[0143] The final output result D contains the bounding boxes that pass the verification and .
[0144] S3, training the TSCM-Det model through the pre-processed tea leaf impurity image dataset to obtain a trained TSCM-Det model; inputting the tea leaf image to be detected into the trained TSCM-Det model to obtain a tea leaf impurity detection result.
[0145] Referring to Figure 6 , the embodiment of the application also provides a tea leaf impurity visual detection system, which comprises a data acquisition unit 601, a detection result obtaining unit 602 and a target result obtaining unit 603, wherein: The data acquisition unit 601 is configured to acquire a tea leaf image to be detected and a training image dataset, pre-process image data in the training image dataset to obtain a pre-processed training image dataset, and the tea leaf image to be detected contains a plurality of detection targets to be detected. The detection result obtaining unit 602 is configured to input the tea leaf image to be detected into a trained tea leaf impurity detection model to obtain a plurality of tea leaf impurity detection results containing confidence and detection boxes of the detection targets to be detected, wherein the trained tea leaf impurity detection model is trained through the pre-processed training image dataset, and the trained tea leaf impurity detection model comprises a backbone network, a neck network and a plurality of detection heads, the backbone network comprises a ScharrStem module, a first SGCB improvement module and a second SGCB improvement module, the neck network comprises the first SGCB improvement module, the second SGCB improvement module and a channel perception spatial modulation module, and wherein The ScharrStem module is configured to perform feature extraction on the tea leaf image to be detected to obtain a feature extraction result. The plurality of stages comprising the first SGCB improvement module or the second SGCB improvement module are configured to perform multi-scale feature extraction on the feature extraction result to obtain a multi-scale feature map. The multi-scale feature map is processed through the first SGCB improvement module, the second SGCB improvement module and the channel perception spatial modulation module to obtain a plurality of target feature fusion results. A plurality of detection heads are used to detect each target feature fusion result to obtain a plurality of tea impurity detection results. The target result obtaining unit 603 is configured to determine a target tea impurity detection result according to the confidence and the detection frame in the plurality of tea impurity detection results.
[0146] It should be noted that, since the tea impurity visual detection system in the embodiment and the tea impurity visual detection method described above are based on the same inventive concept, the corresponding contents in the method embodiment are also applicable to the system embodiment, which will not be described in detail here.
[0147] Referring to Figure 7 The electronic device provided by the embodiment of the present application comprises: at least one memory; at least one processor; at least one program; The program is stored in the memory, and the processor executes the at least one program to implement the tea impurity visual detection method described above.
[0148] The electronic device can be any intelligent terminal, including a mobile phone, a tablet computer, a personal digital assistant (PDA), a vehicle-mounted computer, etc.
[0149] The electronic device of the embodiment of the present application will be described in detail below.
[0150] The processor 1600 can be implemented in the form of a general central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute related programs to implement the technical solutions provided by the embodiments of the present application. The memory 1700 can be implemented in the form of a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1700 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present application are implemented by software or firmware, the related program codes are saved in the memory 1700 and are called and executed by the processor 1600 to implement the tea impurity visual detection method of the embodiments of the present application.
[0151] The input / output interface 1800 is configured to realize information input and output. The communication interface 1900 is configured to realize the communication interaction between the device and other devices, and the communication can be realized through a wired manner (for example, a USB, a network cable and the like) or a wireless manner (for example, a mobile network, WIFI, Bluetooth and the like). The bus 2000 is configured to transmit information between various components (for example, the processor 1600, the memory 1700, the input / output interface 1800 and the communication interface 1900) of the device. The processor 1600, the memory 1700, the input / output interface 1800 and the communication interface 1900 are connected to each other through the bus 2000 to realize the communication connection between the device.
[0152] The disclosure further provides a storage medium, which is a computer readable storage medium, and stores computer executable instructions for causing a computer to execute the tea impurity visual detection method.
[0153] The memory is a non-transitory computer readable storage medium, and can be used to store a non-transitory software program and a non-transitory computer executable program. In addition, the memory can include a high-speed random access memory, and can also include a non-transitory memory, for example, at least one magnetic disk storage device, a flash memory device or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory remotely arranged relative to the processor, and the remote memory can be connected to the processor through a network. Examples of the network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.
[0154] As will be appreciated by one of ordinary skill in the art, all or some of the steps, systems, and techniques disclosed herein can be embodied in software, firmware, hardware, and / or suitable combination thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a micro-processing unit, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on computer readable media, which can comprise computer storage media (or non-transitory media), and communication media (or transitory media). As is well known to those of ordinary skill in the art, the term computer storage media includes both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer. Further, as is well known to those of ordinary skill in the art, communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism, and includes any information delivery media.
[0155] The above is the specific description of the preferred implementation of the embodiments of the present application, but the embodiments of the present application are not limited to the above implementation, and those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the embodiments of the present application, and these equivalent modifications or replacements are all included in the scope defined by the claims of the embodiments of the present application.
Claims
1. A method for visually detecting impurities in tea leaves, characterized in that, The method includes: Acquire a tea image to be detected and a training image dataset. Preprocess the image data in the training image dataset to obtain a preprocessed training image dataset. The tea image to be detected contains multiple targets to be detected. The tea leaf image to be detected is input into a trained tea impurity detection model to obtain multiple tea impurity detection results containing the confidence scores and detection boxes of the target. The trained tea impurity detection model is obtained through the preprocessed training image dataset. The trained tea impurity detection model includes a backbone network, a neck network, and multiple detection heads. The backbone network includes a ScharrStem module, a first improved SGCB module, and a second improved SGCB module. The neck network includes a first improved SGCB module, a second improved SGCB module, and a channel-aware spatial modulation module. The ScharrStem module is used to extract features from the tea image to be detected, and the feature extraction results are obtained. Multi-scale feature extraction is performed on the feature extraction results through multiple stages that include the first SGCB improvement module or the second SGCB improvement module to obtain a multi-scale feature map; The multi-scale feature map is processed by the first SGCB improvement module, the second SGCB improvement module and the channel-aware spatial modulation module to obtain the fusion result of multiple target features; The multiple detection heads are used to detect each of the target feature fusion results to obtain multiple tea impurity detection results; The target tea impurity detection result is determined based on the confidence level and detection frame in the multiple tea impurity detection results.
2. The method for visually detecting tea impurities according to claim 1, characterized in that, The step of extracting features from the tea image to be detected using the ScharrStem module to obtain feature extraction results includes: The tea leaf image to be detected is subjected to a convolution operation to obtain an initial feature map; The initial feature map is subjected to Scharr filtering to obtain an edge-enhanced feature map; The edge enhancement feature map is fused with the tea leaf image to be detected to obtain a first fusion result; The first fusion result is subjected to two convolution operations to obtain the intermediate feature representation; The intermediate feature representation is Gaussian smoothed to obtain a smoothed result; The smoothing result is fused with the intermediate feature representation to obtain a second fusion result; The second fusion result is then subjected to a convolution operation to obtain the feature extraction result.
3. The visual detection method for tea impurities according to claim 1, characterized in that, The step of performing multi-scale feature extraction on the feature extraction results through multiple stages including the first SGCB improvement module or the second SGCB improvement module to obtain a multi-scale feature map includes: The feature extraction results are input into the first stage containing the first SGCB improvement module for feature extraction to obtain a first-scale feature map. The first scale feature map is downsampled, batch normalized, and processed using the SiLU activation function to obtain the first processing result; The first processing result is input into the second stage, which includes the first SGCB improvement module, for feature extraction to obtain a second-scale feature map. The second-scale feature map is downsampled, batch normalized, and processed using the SiLU activation function to obtain the second processing result. The second processing result is input into the third stage, which includes the second SGCB improvement module, for feature extraction to obtain a third-scale feature map; The third-scale feature map is downsampled, batch normalized, and processed using the SiLU activation function to obtain the third processing result; The third processing result is input into the fourth stage, which includes the second SGCB improvement module, for feature extraction to obtain the fourth scale feature map.
4. The visual detection method for tea impurities according to claim 3, characterized in that, The first and second SGCB improvement modules each include an SGCB submodule and a C3K2 submodule. The SGCB submodule includes multiple convolutional layers and an improved Scharr convolutional layer. Adding the SGCB submodule to the C3K2 submodule yields the SGCB improvement module. When the C3K2 submodule in the SGCB improvement module does not use the C3K submodule for feature enhancement, the SGCB improvement module is the first SGCB improvement module. When the C3K2 submodule in the SGCB improvement module uses the C3K submodule for feature enhancement, the SGCB improvement module is the second SGCB improvement module. The improved Scharr convolutional layer includes the following processing steps: The input feature map is subjected to channel-by-channel Scharr gradient calculation to obtain the gradient map in the x-direction and the gradient map in the y-direction for each channel; The gradient magnitudes of the x-direction gradient map and the y-direction gradient map are fused to obtain the original edge response map of each channel; The input feature map is subjected to channel-by-channel 3×3 neighborhood variance calculation and normalization to obtain a normalized variance map; The weight of each channel is calculated based on the normalized variance plot. The original edge response map corresponding to each channel is weighted pixel by pixel by the weight of each channel to obtain the enhanced edge feature map of each channel; The enhanced edge feature maps of each channel are integrated in the original channel order to obtain the final output feature map.
5. The method for visually detecting tea impurities according to claim 1, characterized in that, The multi-scale feature map is processed through a first SGCB improvement module, a second SGCB improvement module, and a channel-aware spatial modulation module to obtain multiple target feature fusion results, including: Feature extraction is performed on the fourth-scale feature map to obtain a first feature image. The first feature image is then upsampled to obtain an upsampled first feature image. The third-scale feature map is then superimposed with the upsampled first feature image to obtain a first superposition result. The first superposition result is then input into the first SGCB improvement module for feature fusion to obtain a feature fusion result. The feature fusion result is upsampled to obtain an upsampled second feature image. The second scale feature image is then superimposed with the upsampled second feature image to obtain a second superposition result. The first-scale feature map is downsampled to obtain a third-scale feature image. The third-scale feature image and the second superposition result are then input into the channel-aware spatial modulation module for feature fusion to obtain a first-enhanced feature map. The first-enhanced feature map is then input into the first SGCB improvement module for feature fusion to obtain a first-target feature fusion result. The first enhanced feature map is downsampled to obtain a downsampled fourth feature image. The downsampled fourth feature image and the feature fusion result are input to the channel-aware spatial modulation module for feature fusion to obtain a second enhanced feature map. The second enhanced feature map is input to the first SGCB improvement module for feature fusion to obtain a second target feature fusion result. The second target feature fusion result is downsampled to obtain a downsampled fifth feature image. The downsampled fifth feature image and the first feature image are input to the channel sensing spatial modulation module for feature fusion to obtain a third enhanced feature map. The third enhanced feature map is input to the second SGCB improvement module for feature fusion to obtain a third target feature fusion result.
6. The visual detection method for tea impurities according to claim 5, characterized in that, The step of inputting the downsampled third feature image and the second superposition result into the channel-aware spatial modulation module for feature fusion to obtain the first enhanced feature map includes: The downsampled third feature image is used as a low-level feature, and the second overlay result is used as a high-level feature; The low-level features and the high-level features are added element-wise to obtain the first addition result; The first summation result is input into the channel attention module for global average pooling and convolution operations to obtain the first convolutional feature. The first convolutional feature is processed through an activation function and a convolution operation to obtain the channel attention weights; The first summation result is subjected to max pooling to obtain a first pooling result, and the first summation result is subjected to average pooling to obtain a second pooling result. The first pooling result and the second pooling result are concatenated and then convolutional to obtain the second convolutional feature; The first convolutional feature is processed by convolution and activation function to obtain the third convolutional feature; The third convolutional feature is expanded by a broadcast mechanism and then multiplied with the second convolutional feature to obtain the collaborative spatial attention weights; The collaborative spatial attention weights and the channel attention weights are added element-wise to obtain a second addition result. The first addition result and the second addition result are then concatenated to obtain a concatenated result. The splicing result is grouped and rearranged to obtain fusion weights, and the fusion weights are then processed by convolution and activation functions to obtain spatial modulation weights. The weighted features are determined based on the spatial modulation weights, the low-level features, and the high-level features; The weighted features, the low-level features, and the high-level features are added element-wise to obtain a third addition result; The third summation result is then subjected to a convolution operation to obtain the first enhanced feature map.
7. The method for visually detecting tea impurities according to claim 1, characterized in that, The step of determining the target tea impurity detection result based on the confidence level and detection frame in the multiple tea impurity detection results includes: Based on the detection frames in the multiple tea impurity detection results, determine the ratio of the spatial scale sensing overlap area between two detection frames; Based on the spatial scale-aware overlap area ratio and the confidence level, a Gaussian confidence decay function is constructed, which is used to update the confidence level. Based on the detection box, the Gaussian confidence decay function, and the confidence level, redundant boxes are removed to obtain the target tea impurity detection result.
8. A visual inspection system for tea impurities, characterized in that, The system includes: The data acquisition unit is used to acquire the tea image to be detected and the training image dataset, and to preprocess the image data in the training image dataset to obtain the preprocessed training image dataset. The tea image to be detected contains multiple targets to be detected. The detection result acquisition unit is used to input the tea image to be detected into a trained tea impurity detection model to obtain multiple tea impurity detection results containing the confidence scores and detection boxes of the target to be detected. The trained tea impurity detection model is trained using the preprocessed training image dataset. The trained tea impurity detection model includes a backbone network, a neck network, and multiple detection heads. The backbone network includes a ScharrStem module, a first SGCB improvement module, and a second SGCB improvement module. The neck network includes a first SGCB improvement module, a second SGCB improvement module, and a channel-aware spatial modulation module. The ScharrStem module is used to extract features from the tea image to be detected, and the feature extraction results are obtained. Multi-scale feature extraction is performed on the feature extraction results through multiple stages that include the first SGCB improvement module or the second SGCB improvement module to obtain a multi-scale feature map; The multi-scale feature map is processed by the first SGCB improvement module, the second SGCB improvement module and the channel-aware spatial modulation module to obtain the fusion result of multiple target features; The multiple detection heads are used to detect each of the target feature fusion results to obtain multiple tea impurity detection results; The target result acquisition unit is used to determine the target tea impurity detection result based on the confidence level and detection frame in the multiple tea impurity detection results.
9. An electronic device, characterized in that, It includes at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, which, when executed by the at least one control processor, enable the at least one control processor to perform the visual detection method for tea impurities as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the visual inspection method for tea impurities as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Steel surface defect detection method, device and equipment and storage medium
CN121120578A
Image enhancement method and apparatus, device and medium
US20250200718A1
Cited By
Sonar image continuous frame target detection method and system based on image reconstruction fusion
CN121937854A
Tea bud visual detection method, system and equipment based on layered feature enhancement
CN122336570A