Tea leaf impurity detection method, system and equipment based on multi-level treatment and medium
By employing a multi-level processing method, combining feature extraction and fusion of depthwise separable convolution and dilated convolution, the problems of missed detection of small-scale impurities and inaccurate edge segmentation in tea impurity detection are solved, achieving high-precision impurity identification and positioning, and meeting the needs of automated removal and mechanical grasping.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CENT SOUTH UNIV
- Filing Date
- 2026-04-13
- Publication Date
- 2026-05-12
AI Technical Summary
Existing tea impurity detection technologies struggle to simultaneously retain global contextual information and local edge details, leading to missed detection of small-scale impurities and inaccurate edge segmentation, failing to meet the precise contour requirements for automated removal and mechanical grasping.
A multi-level processing method is adopted, which obtains initial feature maps at different levels through layer-by-layer convolution and downsampling. Combined with feature enhancement and step-by-step fusion, high-precision pixel-level segmentation results are generated. Depthwise separable convolution and dilated convolution are used to expand the receptive field and perform multi-scale feature extraction and adaptive fusion.
It improves the accuracy of tea impurity detection, enables precise identification and positioning of tea leaves and impurities, provides accurate contour information to meet the needs of automated removal and mechanical grasping, and improves the recognition rate of small-scale impurities and edge segmentation accuracy.
Smart Images

Figure CN122023949A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of tea impurity detection technology, and in particular to tea impurity detection methods, systems, equipment and media based on multi-level processing. Background Technology
[0002] During the harvesting and processing of tea, various endogenous or exogenous impurities, such as old stems, waxy leaves, weeds, and plastic ropes, are often mixed in. These impurities not only affect the quality and taste of the tea but also impact food safety. Therefore, impurity detection is a crucial step in ensuring product quality and food safety during the tea processing. Existing detection technologies mainly include manual sorting, spectral-based physical detection, traditional machine vision, and deep learning-based object detection or semantic segmentation methods.
[0003] However, because tea impurities are generally small in size, irregular in shape, have blurred edges, and resemble the appearance of the tea leaves themselves, conventional semantic segmentation networks in existing technologies are prone to losing details of small targets during downsampling. Furthermore, existing feature fusion methods often employ simple concatenation or addition, making it difficult to simultaneously preserve global contextual information and local edge details. In addition, existing methods mostly output bounding boxes or coarse segmentation maps, failing to provide the precise contour information required for automated removal and mechanical grasping, resulting in missed detection of small-scale impurities and inaccurate edge segmentation. Summary of the Invention
[0004] The main objective of this disclosure is to propose a method, system, device, and storage medium for detecting tea impurities based on multi-level processing, which can solve the technical problem of low accuracy in tea impurity detection in the prior art.
[0005] A first aspect of this application provides a method for detecting tea impurities based on multi-level processing, the method comprising: In response to a detection command, acquire an image of the tea leaves to be detected; The image of the tea leaves to be detected is subjected to layer-by-layer convolution and downsampling to obtain... Initial feature maps at different levels; It is a positive integer; The initial feature maps at different levels have different resolutions; Regarding the Initial feature maps at different levels are executed The second feature processing operation yields the first... The first enhanced feature map output after the second feature processing; Based on the first enhanced feature map, the detection result of the tea image to be detected is obtained; Wherein, the to the Initial feature maps at different levels are executed Any feature processing operation in a sub-feature processing operation includes: For the The initial feature map is enhanced to obtain the first feature map. The enhanced feature map; and the first enhanced feature map; and the first The enhanced feature map and its corresponding first enhancement feature map The initial feature maps are fused to obtain the first... The enhanced feature map; will the first The enhanced feature map and its corresponding first enhancement feature map The initial feature maps are fused to obtain the first... Each enhanced feature map is processed sequentially until the first enhanced feature map is fused with the corresponding second initial feature map to obtain the first enhanced feature map; where, .
[0006] The first aspect of this application provides a tea impurity detection method based on multi-level processing, which acquires an image of the tea to be detected in response to a detection command; and performs layer-by-layer convolution and downsampling on the tea image to be detected to obtain... Initial feature maps at different levels; It is a positive integer; The initial feature maps at different levels have different resolutions; for Initial feature maps at different levels are executed The second feature processing operation yields the first... The first enhanced feature map is output after the second feature processing. Based on the first enhanced feature map, the detection result of the tea image to be detected is obtained. It can generate high-precision pixel-level segmentation results by combining multi-level feature extraction with feature enhancement and step-by-step fusion, thereby improving the accuracy of tea impurity detection.
[0007] To achieve the above objectives, a second aspect of the present invention provides a tea impurity detection system based on multi-level processing, the system comprising: The response module is used to respond to detection commands and acquire images of the tea leaves to be detected. The first module is used to perform layer-by-layer convolution and downsampling on the tea image to be detected, to obtain... Initial feature maps at different levels; It is a positive integer; The initial feature maps at different levels have different resolutions; The second module is used for the... Initial feature maps at different levels are executed The second feature processing operation yields the first... The first enhanced feature map output after the second feature processing; The detection module is used to obtain the detection result of the tea image to be detected based on the first enhanced feature map; Wherein, the to the Initial feature maps at different levels are executed Any feature processing operation in a sub-feature processing operation includes: For the The initial feature map is enhanced to obtain the first feature map. The enhanced feature map; and the first enhanced feature map; and the first The enhanced feature map and its corresponding first enhancement feature map The initial feature maps are fused to obtain the first... The enhanced feature map; will the first The enhanced feature map and its corresponding first enhancement feature map The initial feature maps are fused to obtain the first... Each enhanced feature map is processed sequentially until the first enhanced feature map is fused with the corresponding second initial feature map to obtain the first enhanced feature map; where, .
[0008] To achieve the above objectives, a third aspect of the present invention provides an electronic device, comprising: at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, the instructions being executed by the at least one control processor to enable the at least one control processor to perform the above-described method for detecting tea impurities based on multi-level processing.
[0009] To achieve the above objectives, a fourth aspect of the present invention provides a computer-readable storage medium storing computer-executable instructions for causing a computer to perform the above-described method for detecting tea impurities based on multi-level processing.
[0010] It is understood that the beneficial effects of the second to fourth aspects compared with the related technologies are the same as the beneficial effects of the first aspect compared with the related technologies. Please refer to the relevant description in the first aspect above, which will not be repeated here. Attached Figure Description
[0011] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is a schematic flowchart of a tea impurity detection method based on multi-level processing provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a tea impurity detection system based on multi-level processing provided in an embodiment of this application; Figure 3 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0012] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0013] In the description of this application, the use of terms such as "first," "second," etc., is for the purpose of distinguishing technical features only and should not be construed as indicating or implying relative importance or implicitly indicating the number of technical features indicated or the order of the technical features indicated.
[0014] In the description of this application, it should be understood that the orientation descriptions, such as up, down, etc., are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application.
[0015] In the description of this application, it should be noted that, unless otherwise explicitly defined, terms such as "setup," "installation," and "connection" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this application in conjunction with the specific content of the technical solution.
[0016] During the harvesting and processing of tea, various endogenous or exogenous impurities, such as old stems, waxy leaves, weeds, and plastic ropes, are often mixed in. These impurities not only affect the quality and taste of the tea but also impact food safety. Therefore, impurity detection is a crucial step in ensuring product quality and food safety during the tea processing. Existing detection technologies mainly include manual sorting, spectral-based physical detection, traditional machine vision, and deep learning-based object detection or semantic segmentation methods.
[0017] However, because tea impurities are generally small in size, irregular in shape, have blurred edges, and resemble the appearance of the tea leaves themselves, conventional semantic segmentation networks in existing technologies are prone to losing details of small targets during downsampling. Furthermore, existing feature fusion methods often employ simple concatenation or addition, making it difficult to simultaneously preserve global contextual information and local edge details. In addition, existing methods mostly output bounding boxes or coarse segmentation maps, failing to provide the precise contour information required for automated removal and mechanical grasping, resulting in missed detection of small-scale impurities and inaccurate edge segmentation.
[0018] Based on this, embodiments of this application provide a tea impurity detection method, system, electronic device, and medium based on multi-level processing, aiming to generate high-precision pixel-level segmentation results by combining multi-level feature extraction with feature enhancement and step-by-step fusion, thereby improving the accuracy of tea impurity detection.
[0019] The tea impurity detection method, system, electronic device and medium based on multi-level processing provided in this application are specifically described through the following embodiments. First, the tea impurity detection method based on multi-level processing in this application embodiment is described.
[0020] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0021] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0022] The tea impurity detection method based on multi-level processing provided in this application relates to the field of tea impurity detection technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the tea impurity detection method based on multi-level processing, but is not limited to the above forms.
[0023] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0024] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0025] Therefore, referring to Figure 1 This application provides a method for detecting tea impurities based on multi-level processing. This method is applied to a central controller, which can be a server, an electronic device, or a mobile terminal, etc. There are no specific limitations here. The method includes the following steps S110 to S140.
[0026] Step S110: In response to the detection command, acquire an image of the tea leaves to be detected; Step S120: Perform layer-by-layer convolution and downsampling on the image of the tea leaves to be detected, and obtain... Initial feature maps at different levels; It is a positive integer; The initial feature maps at different levels have different resolutions; Step S130, for Initial feature maps at different levels are executed The second feature processing operation yields the first... The first enhanced feature map output after the second feature processing; Step S140: Based on the first enhanced feature map, obtain the detection result of the tea image to be detected; Among them, for Initial feature maps at different levels are executed Any feature processing operation in a sub-feature processing operation includes: For the The initial feature map is enhanced to obtain the first feature map. The enhanced feature map; and the first enhanced feature map; and the first The enhanced feature map and its corresponding first enhancement feature map The initial feature maps are fused to obtain the first... The enhanced feature map; will the first The enhanced feature map and its corresponding first enhancement feature map The initial feature maps are fused to obtain the first... Each enhanced feature map is processed sequentially until the first enhanced feature map is fused with the corresponding second initial feature map to obtain the first enhanced feature map; where, .
[0027] First, let's analyze the key terms used in this application: The tea image to be inspected refers to the original image containing tea leaves and any impurities that may be mixed in, captured by an industrial camera in the tea refining production line. For example, the tea image to be inspected can be an RGB color image. To reduce the impact of ambient light fluctuations and improve image quality, the acquisition process is preferably completed in a closed darkroom environment with a ring light source and a vibrating feeding mechanism, and preprocessing operations such as scaling and normalization are performed before inputting the image into the network.
[0028] Layer-by-layer convolution and downsampling are fundamental operations in deep neural networks (such as residual convolutional neural networks) for extracting image features. Layer-by-layer convolution is used to extract local features such as texture and edges in an image; downsampling is used to reduce the spatial resolution of the feature map, thereby expanding the network's receptive field and reducing computational cost, allowing features to gradually transition from shallow, concrete details to deep, abstract semantics.
[0029] The initial feature map refers to the intermediate data representation output at different network depths after the image of tea leaves to be detected has been processed layer by layer by the backbone network. Initial feature maps at different levels have different resolutions. High-resolution (shallow) initial feature maps mainly retain detailed information such as the spatial location and edge contours of the target, while low-resolution (deep) initial feature maps mainly contain high-level semantic information used to distinguish the target category.
[0030] In this step, upon receiving a detection command, an image of the tea leaves to be detected is acquired. For example, on a tea refining production line, when a detection start signal or trigger signal is received, an image acquisition device acquires an image containing tea leaves and impurities. The image acquisition device is set up in a closed darkroom environment, and works with a ring light source and a vibrating feeding mechanism to ensure that the tea leaves and impurities are evenly dispersed within the imaging area, guaranteeing stable image quality.
[0031] Furthermore, feature extraction is performed on the tea leaf image to be detected, progressing from shallow to deep. Taking a value of 4 as an example, after performing layer-by-layer convolution and downsampling on the tea image to be detected, four initial feature maps of different levels will be output, with resolutions of the input image respectively. , , and Among them, higher resolution (such as...) The shallow feature map of a target mainly preserves spatial location details such as the target's edges and texture; lower resolution (e.g., The deep feature map of the target mainly contains high-level abstract semantic information used to determine the target category.
[0032] Furthermore, a multi-stage, progressively refined architecture is employed to process the multi-level feature maps, executing a total of [number] steps. Each feature processing operation follows the same processing logic, and the feature representation is gradually optimized through multiple iterations.
[0033] This embodiment provides... One feature processing operation in the sub-feature processing operation specifically includes: First, processing the deepest initial feature map of the current stage (i.e., the first...) The initial feature map is subjected to feature enhancement processing to obtain the first feature map. The enhanced feature map; then, the first... The enhanced feature map and its corresponding first enhancement feature map The initial feature maps are fused to obtain the first one. The enhanced feature map; then, the first... The enhanced feature map and its corresponding first enhancement feature map The initial feature maps are fused to obtain the first... This process is repeated until the first enhanced feature map is fused with the corresponding second initial feature map to obtain the first enhanced feature map. .
[0034] Therefore, through the above processing, each feature processing operation achieves a gradual fusion and refinement from deep semantic features to shallow detailed features. After the feature processing operation, the first enhanced feature map output at the end integrates multi-scale semantic information and detail information.
[0035] Furthermore, the first enhanced feature map obtained after multiple feature processing operations is input into the subsequent processing stage to extract the detection result of the tea image to be detected. This detection result is used for subsequent impurity identification and localization. Specifically, from the first... Multiple local and global context features are extracted from the initial feature map, including features based on depthwise separable convolution and dilated convolution. Multiple local features and global contextual features are extracted from each initial feature map.
[0036] Depthwise separable convolution decomposes standard convolution into channel-wise convolution and pointwise convolution, effectively controlling the number of parameters and computational cost while expanding the receptive field. Dilated convolution expands the receptive field without increasing the number of parameters by inserting gaps between convolution kernel elements. Combining the two can, on the one hand, expand the receptive field to capture a wider range of contextual information, and on the other hand, reduce model complexity to meet the real-time detection needs of industrial scenarios.
[0037] In some embodiments, step S140 includes the following steps S210 to S240: Step S210, from the first Multiple local features and global contextual features are extracted from each initial feature map; the receptive fields of the multiple local features are different. Step S220: Fuse the local features and global context features to obtain the first... One enhanced feature map.
[0038] In this embodiment, multiple local features and global contextual features are extracted from the Nth initial feature map. The receptive fields of these local features differ, achieved by setting multiple local branches with different receptive fields. Specifically, dilated convolutions with different dilation rates can be used to extract fine-grained local texture information, mid-scale regional structure information, and broader contextual information to accommodate the multi-scale target features present in tea impurities. Simultaneously, global semantic information of the entire feature map is extracted through a global contextual branch to enhance the understanding of complex backgrounds and similarly categorized targets.
[0039] Furthermore, the local features and global context features are fused to obtain the Mth enhanced feature map, which involves fusing the multiple local features extracted above with the global context features. Preferably, the output results of each branch are merged using channel concatenation, and then channel compression and feature recombination are performed through convolution operations to finally output the enhanced deep feature map. This enhanced feature map has both local detail representation ability and global context awareness ability.
[0040] In some embodiments, in step S140, Initial feature maps at different levels are executed Any image fusion operation in the sub-feature processing operation includes the following steps S310 to S330: Step S310, the first The enhanced feature map and its corresponding first enhancement feature map The initial feature map is scale-aligned and channel-aligned to obtain the aligned feature map. Step S320: Generate soft attention weights based on each aligned feature map; Step S330: Based on the soft attention weights, for the first... The enhanced feature map and its corresponding first enhancement feature map The initial feature map is dynamically weighted and fused to obtain the first... One enhanced feature map; among which... It is a positive integer. Less than or equal to , .
[0041] In this embodiment, the first The enhanced feature map and its corresponding first enhancement feature map The initial feature maps are scale-aligned and channel-aligned to obtain aligned feature maps. Since the two feature maps involved in the fusion have differences in resolution and number of channels, the low-resolution feature needs to be upsampled first to make its spatial size consistent with the high-resolution feature; at the same time, the high-resolution feature is channel-aligned through convolution operation to make its channel dimension match the low-resolution feature, ensuring that the two feature maps have consistent size and number of channels before fusion.
[0042] Further, soft attention weights are generated based on each aligned feature map. Specifically, in some embodiments, generating soft attention weights based on each aligned feature map in step S320 includes the following steps S410 to S430: Step S410: Add each aligned feature map element by element to generate an intermediate feature map; Step S420: Extract global channel context information and local channel context information from the intermediate feature map respectively; Step S430: Combine the global channel context information with the local channel context information, and map them to the interval between 0 and 1 through the activation function to obtain the soft attention weights.
[0043] In this embodiment, the aligned feature maps are added element-wise, that is, the aligned high-resolution features and low-resolution features are added element-wise to obtain an intermediate feature map that fuses the information of both, which serves as the basis for generating attention weights. Then, global channel context information and local channel context information are extracted from the intermediate feature map.
[0044] Specifically, global channel context information is extracted through global branches: global average pooling is performed on the intermediate feature map to obtain global statistical information of the entire feature map, and then one-dimensional convolution is used to realize the interaction between channels to obtain global channel context features that reflect the importance of different channels in the global scope. At the same time, local channel context information is extracted through local branches: channel compression convolution, depthwise separable convolution, and channel recovery convolution are performed on the intermediate feature map in sequence to extract the channel relationships in the local spatial neighborhood, which is used to preserve the detailed information of edges, weak textures, and small objects.
[0045] Furthermore, the global branch output and the local branch output are combined, and then a soft attention weight with a value between 0 and 1 is generated by the Sigmoid activation function. This not only takes into account the requirements of global semantic discrimination, but also takes into account the requirements of local detail preservation, thereby realizing adaptive fusion control of multi-resolution features.
[0046] Furthermore, the generated soft attention weights are used to adaptively weight and combine the two input features, thereby highlighting the feature components that contribute more to the segmentation task and suppressing redundant background information, thus improving the quality of multi-scale feature fusion.
[0047] In some embodiments, step S140 includes the following steps S510 to S520: Step S510: Perform pixel-by-pixel classification on the first enhanced feature map to obtain the probability of the category to which each pixel belongs; Step S520: Determine the final category label of each pixel according to the category probability of each pixel, so as to generate a pixel-level segmentation result map with the same resolution as the tea image to be detected.
[0048] In this embodiment, the first enhanced feature map obtained after multi-stage refinement is classified pixel by pixel. By predicting the category of each pixel, the probability value of the pixel belonging to each category is output. Then, based on the category probability of each pixel, the final category label of each pixel is determined. Preferably, the final category label of each pixel can be determined by the maximum probability principle, and the classification result is restored to the same resolution as the original input image by upsampling.
[0049] The final output is a pixel-level segmentation result image, where each pixel is labeled as one of the preset categories, including background, tender leaves, tender stems, waxy leaves, old stems, weeds, and plastic ropes, thereby obtaining the precise outline area of tea leaves and various impurities.
[0050] In some embodiments, after step S140, the following steps S610 to S640 are further included: Step S610: Based on the pixel-level segmentation result image, determine the impurities in the tea image to be detected, and extract the contour boundary, center coordinates, area information and anchor points of the impurities. Step S620: Send the outline boundary, center coordinates, area information and grab anchor points of the impurities to the execution mechanism so that the execution mechanism can perform the corresponding removal and sorting operations.
[0051] In this embodiment, based on the pixel-level segmentation result image, the contour boundary, center point coordinates, area, and anchor point position suitable for mechanical grasping of various impurity targets can be further calculated, thereby accurately reflecting the true shape and spatial position of the impurities and meeting the accuracy requirements of automated rejection and sorting.
[0052] Furthermore, the extracted precise location information is sent to execution equipment such as automatic sorting devices, robotic arm systems, or air-blowing rejection mechanisms to complete the automatic rejection and sorting of tea impurities, realizing closed-loop control of the entire process from image detection to production line execution.
[0053] In one embodiment, the above-mentioned tea impurity detection method based on multi-level processing is applied to construct a tea impurity detection network based on multi-level processing. This network performs multi-level feature extraction, multi-scale feature enhancement, cross-resolution feature adaptive fusion, and multi-stage progressive refinement on the acquired images to output pixel-level segmentation results, thereby achieving accurate identification, localization, and contour extraction of tea impurities. This tea impurity detection network based on multi-level processing includes steps S1 to S7.
[0054] Step S1: Acquire the image of the tea leaves to be detected: Acquire an image containing tea leaves and impurities. This image is preferably an RGB color image, acquired by an industrial camera installed on the tea refining production line and pre-processed to adjust the acquired image to a preset resolution before being sent to the subsequent detection network. To improve image quality, the acquisition process can be carried out in a closed darkroom environment, using a ring light source and a vibrating feeding mechanism to ensure that the tea leaves and impurities are evenly dispersed within the imaging area. The pre-processing includes operations such as size scaling, normalization, and data augmentation to meet the network input requirements.
[0055] Step S2: Extracting multi-level feature maps using the backbone network: The image to be detected is input into the backbone network to extract feature maps at multiple different levels and resolutions. Specifically, the backbone network is a residual convolutional neural network (ResNet50). The backbone network performs layer-by-layer convolution and downsampling operations on the input image to output... Feature maps at each level.
[0056] by Taking 4 as an example, the multi-level feature maps include at least: a first-level feature map, whose resolution is 1 / 3 of the input image resolution. The second-level feature map has a resolution equal to that of the input image. The third layer feature map has a resolution equal to that of the input image. The fourth-level feature map has a resolution equal to that of the input image. .
[0057] Among them, the higher resolution shallow feature maps are mainly used to preserve the target's edges, texture, and spatial location information; the lower resolution deep feature maps are mainly used to represent high-level semantic information related to the target's category.
[0058] Step S3: Construct the first-stage segmentation sub-network and enhance deep features: Input the deep feature map output in step S2 into the first-stage segmentation sub-network. Preferably, the fourth-level feature map with the lowest resolution is used as the deep feature to be enhanced in the first stage and input into the multi-scale perceptual enhancement module MSPEM, wherein the first-stage segmentation sub-network is used to perform preliminary semantic segmentation on the input image and output the first-stage refined features.
[0059] Step S3.1: Input the deep feature map into the multi-scale perception enhancement module. This multi-scale perception enhancement module includes multiple parallel branches for extracting features and fusing global context information under different receptive fields. Specifically, the multi-scale perception enhancement module includes a residual connection branch, a first local branch, a second local branch, a third local branch, and a global context branch.
[0060] Step S3.2: Preserve the original deep semantic information through residual connection branches: The residual connection branches directly pass the input features to the subsequent fusion operation to preserve the original deep semantic representation and enhance the stability of network training.
[0061] Step S3.3: Extract local features from different receptive fields through multiple local branches: Input the input features into multiple local branches respectively. Each local branch preferably includes the following processing steps: 1. Use 1×1 convolution to compress the input features through channels; 2. Depthwise separable convolution is used to extract local features from the compressed features; 3. Use dilated convolution to further expand the receptive field; 4. The channel dimension is restored by 1×1 convolution, and then normalized and activated.
[0062] Preferably, multiple local branches employ different void ratios to form receptive fields of different scales. For example, the void ratios corresponding to the three local branches can be 1, 3, and 5, respectively. Thus, through the above settings, multiple local branches can extract fine-grained local texture information, mesoscale regional structural information, and broader context-related information, respectively.
[0063] Step S3.4: Extract global semantic information through the global context branch: Input the input features into the global context branch. The global context branch preferably extracts global statistical information of the entire feature map through global average pooling, and then performs feature transformation through convolution, normalization and activation operations to obtain a global semantic representation.
[0064] Step S3.5: Fuse the outputs of each branch and generate enhanced features: Fuse the outputs of the residual connection branch, multiple local branches, and the global context branch. The preferred fusion method is channel concatenation, followed by channel compression and feature recombination using a 1×1 convolution to obtain the enhanced deep feature map.
[0065] Thus, through step S3, multi-scale enhancement of the deep features in the first stage is completed, so that the enhanced deep features have both the ability to express local details and the ability to express global context.
[0066] Step S4: Perform cross-resolution adaptive feature fusion in the first stage: The enhanced deep feature map obtained in step S3 is fused step by step with a higher-resolution shallow feature map to restore spatial resolution and form the first-stage refined features. For example, the output representation on the backbone network in the first stage is ResBlock3—36*32*256, where 36*32 represents the feature resolution and 256 represents the number of feature channels. The output obtained by MSPEM in step S3 is 18*16*256, where 36*32 represents the feature resolution and 256 represents the number of feature channels. Therefore, ResBlock3—36*32*256 (a higher-resolution shallow feature map) and the output feature map obtained by MSPEM (18*16*256) are fused in SAFFM. Similarly, in this embodiment, there are 6 SAFFMs, and the fusion of each SAFFM is the fusion of a higher-resolution feature with a lower-resolution feature. Cross-resolution feature fusion is preferably performed through the soft attention feature fusion module SAFFM.
[0067] Step S4.1: Obtain the two features to be fused. The features to be fused include: one high-resolution feature, preferably from a shallower feature map of the backbone network; and one low-resolution feature, preferably from the enhanced deep feature map obtained in step S3 or the previous fusion output feature map.
[0068] Step S4.2: Perform scale alignment and channel alignment on the two features: First, upsample the low-resolution features to make their spatial dimensions consistent with the high-resolution features. Second, perform channel alignment on the high-resolution features through 1×1 convolution to give the two features a channel dimension that can be fused.
[0069] Step S4.3: Construct intermediate feature map: Add the two feature paths that have undergone scale alignment and channel alignment element-wise to generate an intermediate feature map. The intermediate feature map is used for subsequent local and global attention modeling.
[0070] Step S4.4: Extract global channel context through global branches: Input the intermediate feature map into the global branches. Preferably, the global branches include: 1. Perform global average pooling on the intermediate feature maps; 2. Achieve inter-channel interaction through one-dimensional convolution; 3. The global channel representation is further enhanced through convolution, normalization, and activation operations.
[0071] Specifically, the global branch outputs global channel context features, which are used to represent the importance of different channels in the global scope.
[0072] Step S4.5: Extract local channel context through local branches: Input the intermediate feature map into the local branches. Preferably, the local branches include: 1. Compress the channels using 1×1 convolution; 2. Extract detailed information from the local spatial neighborhood through depthwise separable convolution; 3. Then use 1×1 convolution to restore the channel dimension.
[0073] Local branch outputs local channel context features to preserve details of edges, weak textures, and small objects.
[0074] Step S4.6: Generate soft attention weights: Combine the global branch output with the local branch output, and generate soft attention weights with values between 0 and 1 by mapping through the Sigmoid function. These soft attention weights are used to characterize the contribution of each channel feature in the current fusion process.
[0075] Step S4.7: Perform two-way feature fusion based on soft attention weights: Use soft attention weights to dynamically weight and fuse high-resolution features and low-resolution features, and output the fused feature map of the current level. This will enable adaptive highlighting of feature components that contribute more to the current segmentation task and suppress redundant background information.
[0076] Step S4.8 Repeatedly perform step-by-step fusion: In the first stage, repeat steps S4.1 to S4.7 to fuse the enhanced deep features with multiple shallow feature maps of higher resolution in sequence, gradually restore the spatial resolution, and finally obtain the first stage refined feature map.
[0077] Step S5: Construct subsequent stage segmentation sub-networks and further refine features: Input the first-stage refined feature map obtained in step S4 into the second-stage segmentation sub-network for further optimization. The second-stage segmentation sub-network can adopt the same or similar structure as the first stage, including: a deep feature enhancement module, a soft attention feature fusion module, and a multi-level progressive decoding structure.
[0078] Specifically, in the second stage, the blurred boundary regions, small-scale target regions, and easily confused category regions in the output of the first stage are further refined to obtain the second-stage refined feature map. Furthermore, the second-stage refined feature map is input into the third-stage segmentation sub-network, and deep enhancement, cross-resolution fusion, and progressive recovery operations are repeatedly performed to output the third-stage refined feature map. Through multi-stage continuous refinement, the segmentation features gradually approach the true target region from a coarse semantic representation, improving the accuracy of segmentation boundary and small target recognition.
[0079] Step S6: Output the final pixel-level segmentation result: Input the refined feature map from step S5 into the classification head to predict the category for each pixel. Preferably, the classification head includes a convolutional classification layer and a SoftMax probability output layer. The probability value of each pixel belonging to each category is obtained through SoftMax, and then the final category label for each pixel is determined according to the maximum probability principle.
[0080] Furthermore, the classification results are upsampled to restore the resolution to match the input image, resulting in the final pixel-level segmentation image. In the pixel-level segmentation image, each pixel is labeled as one of the preset categories. Preferably, the categories include background, tender leaves, tender stems, waxy leaves, old stems, weeds, and plastic rope.
[0081] Step S7: Extract impurity location information based on segmentation results and use it for subsequent execution: Based on the pixel-level segmentation result map obtained in step S6, the region contour, center coordinates, area information, boundary information and circumscribed rectangle information of each type of impurity target can be further extracted.
[0082] In one implementation of this embodiment, mechanical gripping points, rejection anchor points, or execution area information can be generated based on the segmented area, and the information can be sent to an automatic sorting device, a robotic arm system, an air-blowing rejection mechanism, or other execution equipment to complete the automatic rejection and sorting of tea impurities.
[0083] Therefore, this embodiment effectively alleviates the problem of insufficient one-time modeling of small-scale targets and complex boundaries by constructing a multi-stage progressively refined architecture and performing multiple feature processing operations on multi-level feature maps, thereby improving the recognition rate of small-scale impurities and the accuracy of edge segmentation. At the same time, by extracting and fusing local features and global context features from multiple different receptive fields from deep feature maps, multi-scale perception enhancement is achieved, which can simultaneously take into account the overall semantic information of large-scale targets and the detailed information of small-scale targets, thus enhancing the adaptability to multi-scale impurities. In addition, by generating soft attention weights to dynamically weight and fuse the enhanced feature map and the initial feature map, it can adaptively highlight the feature components that contribute more to the segmentation task, suppress redundant background information, and improve the quality of multi-scale feature fusion. Finally, the pixel-level segmentation result map is output, which can obtain the precise contour regions of tea leaves and various impurities, providing high-precision positional information for subsequent mechanical grasping, automatic sorting, and foreign object removal, thus meeting the practical application needs of automated tea refining production lines.
[0084] like Figure 2 As shown in some embodiments of this application, a tea impurity detection system based on multi-level processing is provided. The system includes a response module 210, a first module 220, a second module 230, and a detection module 240. Specifically: Response module 210 is used to acquire an image of the tea leaves to be detected in response to a detection command; The first module 220 is used to perform layer-by-layer convolution and downsampling on the image of the tea leaves to be detected, resulting in... Initial feature maps at different levels; It is a positive integer; The initial feature maps at different levels have different resolutions; The second module 230 is used for... Initial feature maps at different levels are executed The second feature processing operation yields the first... The first enhanced feature map output after the second feature processing; The detection module 240 is used to obtain the detection result of the tea image to be detected based on the first enhanced feature map; Among them, for Initial feature maps at different levels are executed Any feature processing operation in a sub-feature processing operation includes: For the The initial feature map is enhanced to obtain the first feature map. The enhanced feature map; and the first enhanced feature map; and the first The enhanced feature map and its corresponding first enhancement feature map The initial feature maps are fused to obtain the first... The enhanced feature map; will the first The enhanced feature map and its corresponding first enhancement feature map The initial feature maps are fused to obtain the first... Each enhanced feature map is processed sequentially until the first enhanced feature map is fused with the corresponding second initial feature map to obtain the first enhanced feature map; where, .
[0085] It should be noted that the tea impurity detection system based on multi-level processing provided in this embodiment is based on the same inventive concept as the tea impurity detection method based on multi-level processing described above. Therefore, the relevant content of the tea impurity detection method based on multi-level processing described above also applies to the content of the tea impurity detection system based on multi-level processing. Therefore, it will not be repeated here.
[0086] In this embodiment, the system responds to a detection command by acquiring an image of the tea leaves to be detected; the image of the tea leaves to be detected is then subjected to layer-by-layer convolution and downsampling to obtain... Initial feature maps at different levels; It is a positive integer; The initial feature maps at different levels have different resolutions; for Initial feature maps at different levels are executed The second feature processing operation yields the first... The first enhanced feature map is output after the second feature processing. Based on the first enhanced feature map, the detection result of the tea image to be detected is obtained. In this way, high-precision pixel-level segmentation results can be generated by combining multi-level feature extraction with feature enhancement and step-by-step fusion, thereby improving the accuracy of tea impurity detection.
[0087] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described tea impurity detection method based on multi-level processing.
[0088] like Figure 3 , Figure 3 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. The electronic device includes: At least one battery; At least one memory; At least one processor; At least one program; The program is stored in memory, and the processor executes at least one program to implement the tea impurity detection method based on multi-level processing described above in this disclosure.
[0089] This electronic device can be any smart terminal, including mobile phones, tablets, personal digital assistants (PDAs), and in-vehicle computers.
[0090] The electronic devices according to embodiments of this application will now be described in detail.
[0091] The processor 1600 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this disclosure. The memory 1700 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1700 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1700 and is called and executed by the processor 1600 to execute a tea impurity detection method based on multi-level processing according to an embodiment of this disclosure.
[0092] The input / output interface 1800 is used to implement information input and output. The communication interface 1900 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 2000 transmits information between various components of the device (e.g., processor 1600, memory 1700, input / output interface 1800, and communication interface 1900); The processor 1600, memory 1700, input / output interface 1800 and communication interface 1900 are connected to each other within the device via bus 2000.
[0093] This disclosure also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-described tea impurity detection method based on multi-level processing.
[0094] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0095] The embodiments described in this disclosure are for the purpose of more clearly illustrating the technical solutions of this disclosure and do not constitute a limitation on the technical solutions provided by this disclosure. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by this disclosure are also applicable to similar technical problems.
[0096] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this disclosure, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0097] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0098] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0099] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any related variations, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0100] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0101] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0102] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0103] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0104] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0105] The above is a detailed description of the preferred embodiments of this application. However, the embodiments of this application are not limited to the above-described implementation methods. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the embodiments of this application. All such equivalent modifications or substitutions are included within the scope defined by the claims of the embodiments of this application.
[0106] The embodiments of this application have been described in detail above with reference to the accompanying drawings. However, this application is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of this application.
Claims
1. A method for detecting tea impurities based on multi-stage processing, characterized in that, The method includes: In response to a detection command, acquire an image of the tea leaves to be detected; The image of the tea leaves to be detected is subjected to layer-by-layer convolution and downsampling to obtain... Initial feature maps at different levels; It is a positive integer; The initial feature maps at different levels have different resolutions; Regarding the Initial feature maps at different levels are executed The second feature processing operation yields the first... The first enhanced feature map output after the second feature processing; Based on the first enhanced feature map, the detection result of the tea image to be detected is obtained; Wherein, the to the Initial feature maps at different levels are executed Any feature processing operation in a sub-feature processing operation includes: For the The initial feature map is enhanced to obtain the first feature map. The enhanced feature map; and the first enhanced feature map; and the first The enhanced feature map and its corresponding first enhancement feature map The initial feature maps are fused to obtain the first... The enhanced feature map; the first... The enhanced feature map and its corresponding first enhancement feature map The initial feature maps are fused to obtain the first... Each enhanced feature map is processed sequentially until the first enhanced feature map is fused with the corresponding second initial feature map to obtain the first enhanced feature map; where, .
2. The tea impurity detection method based on multi-stage processing according to claim 1, characterized in that, The first The initial feature map is enhanced to obtain the first feature map. One enhanced feature map, including: From the first Multiple local features and global contextual features are extracted from each initial feature map; the receptive fields of these multiple local features are different. The local features and the global context features are fused to obtain the first... One enhanced feature map.
3. The tea impurity detection method based on multi-stage processing according to claim 2, characterized in that, The from the first Multiple local and global context features are extracted from the initial feature maps, including features based on depthwise separable convolution and dilated convolution. Multiple local features and global contextual features are extracted from each initial feature map.
4. The tea impurity detection method based on multi-stage processing according to claim 1, characterized in that, The above Initial feature maps at different levels are executed Any image fusion operation in the sub-feature processing operation includes: The first The enhanced feature map and its corresponding first enhancement feature map The initial feature map is scale-aligned and channel-aligned to obtain the aligned feature map. Based on the aligned feature maps, soft attention weights are generated. Based on the soft attention weights, the first The enhanced feature map and its corresponding first enhancement feature map The initial feature map is dynamically weighted and fused to obtain the first... One enhanced feature map; among which... It is a positive integer. Less than or equal to , .
5. The tea impurity detection method based on multi-stage processing according to claim 4, characterized in that, The generation of soft attention weights based on the aligned feature maps includes: The aligned feature maps are added element by element to generate an intermediate feature map; Global channel context information and local channel context information are extracted from the intermediate feature map, respectively. The global channel context information and the local channel context information are combined and mapped to the interval of 0 to 1 through an activation function to obtain the soft attention weights.
6. The tea impurity detection method based on multi-stage processing according to claim 1, characterized in that, The step of obtaining the detection result of the tea image to be detected based on the first enhanced feature map includes: The first enhanced feature map is classified pixel by pixel to obtain the probability of the category to which each pixel belongs; Based on the category probability of each pixel, the final category label of each pixel is determined to generate a pixel-level segmentation result map with the same resolution as the tea image to be detected.
7. The tea impurity detection method based on multi-stage processing according to claim 6, characterized in that, After obtaining the detection result of the tea image to be detected based on the first enhanced feature map, the method further includes: Based on the pixel-level segmentation result image, impurities in the tea image to be detected are determined, and the contour boundary, center coordinates, area information and anchor points of the impurities are extracted. The outline boundary of the impurity, the center coordinates, the area information, and the grasping anchor point are sent to the execution mechanism so that the execution mechanism can perform the corresponding removal and sorting operations.
8. A tea impurity detection system based on multi-stage processing, characterized in that, The system includes: The response module is used to respond to detection commands and acquire images of the tea leaves to be detected. The first module is used to perform layer-by-layer convolution and downsampling on the tea image to be detected, to obtain... Initial feature maps at different levels; It is a positive integer; The initial feature maps at different levels have different resolutions; The second module is used for the... Initial feature maps at different levels are executed The second feature processing operation yields the first... The first enhanced feature map output after the second feature processing; The detection module is used to obtain the detection result of the tea image to be detected based on the first enhanced feature map; Wherein, the to the Initial feature maps at different levels are executed Any feature processing operation in a sub-feature processing operation includes: For the The initial feature map is enhanced to obtain the first feature map. The enhanced feature map; and the first enhanced feature map; and the first The enhanced feature map and its corresponding first enhancement feature map The initial feature maps are fused to obtain the first... The enhanced feature map; the first... The enhanced feature map and its corresponding first enhancement feature map The initial feature maps are fused to obtain the first... Each enhanced feature map is processed sequentially until the first enhanced feature map is fused with the corresponding second initial feature map to obtain the first enhanced feature map; where, .
9. An electronic device, characterized in that, It includes at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, which, when executed by the at least one control processor, enable the at least one control processor to perform a multi-level processing-based tea impurity detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the tea impurity detection method based on multi-level processing as described in any one of claims 1 to 7.