Method and device for identifying tender tea stems based on multi-scale feature decoupling

Through the multi-scale feature decoupling method of young tea stem recognition, the backbone network, attention mechanism and feature decoupling module are used to separate the features of young buds and young tea stems, which solves the problem of false detection of young tea stems and improves the recognition accuracy and equipment reliability.

CN120747889APending Publication Date: 2025-10-03ZHEJIANG SCI-TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510700892.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In the existing tea bud detection model, the misdetection of young tea stems leads to incorrect operation and damage of picking equipment, affecting equipment life and picking efficiency.

Method used

A young tea stem recognition method based on multi-scale feature decoupling is adopted. Multi-scale features are extracted through the backbone network, and the attention mechanism and multi-scale interaction structure are introduced. The feature decoupling module is used to separate the features of young buds and young tea stems into different feature subspaces to eliminate redundant information.

Benefits of technology

It improves the accuracy of identifying young tea stems, reduces the problem of misidentification of picking equipment, reduces the risk of equipment damage, and improves picking efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747889A_ABST
    Figure CN120747889A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for identifying tender tea stems based on multi-scale feature decoupling, which are used for enhancing global semantics on multiple scales by utilizing a multi-scale interactive structure and improving the quality of visual representation through real-time evaluation. Firstly, a global space channel attention module is combined to extract multi-scale features, and adaptive adjustment of feature mapping weights is promoted. Secondly, a multi-scale interaction structure is designed, and a new attention area is generated through information learned from a shallow layer and a deep layer; and finally, the category targets are projected to respective feature subspaces to solve the feature information redundancy problem. The recognition algorithm obtained through attention, multi-scale interaction and feature decoupling fusion can achieve high accuracy, the problem that the tender tea stems are wrongly recognized during picking operation of the tea picking robot is effectively solved, and damage to the picking tail end can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of image recognition technology, and in particular relates to a method and device for identifying young tea stems based on multi-scale feature decoupling. Background Art

[0002] Tea picking is a cyclical task, especially in scenarios where intelligent robots are used to identify and pick tea buds. This process usually requires multiple repetitions to complete. During the first picking task, due to the fixed camera perspective and environmental complexity (such as buds obscured by leaves, densely distributed, and changeable weather), the target detection algorithm cannot identify all buds at once, so multiple rounds of picking are required. However, during this process, bright green tea stems are left behind in previously picked locations. These tea stems are similar in shape and color to buds, interfering with subsequent recognition tasks. The remaining tea stems are easily misidentified as buds, which not only causes the end effector of the picking device to operate incorrectly, but may also damage the device, affecting its service life and picking efficiency. Therefore, in tea bud detection methods, how to effectively distinguish between tea stems and tea buds has become one of the key challenges in improving model accuracy and equipment reliability. Summary of the Invention

[0003] The purpose of this application is to provide a method and device for identifying young tea stems based on multi-scale feature decoupling, and to provide a high-performance false detection identification method to address the problem of false detection of young tea stems in existing tea bud detection models.

[0004] In order to achieve the above objectives, the technical solutions of this application are as follows: A method for identifying young tea stems based on multi-scale feature decoupling, comprising: The collected tea images are input into the backbone network to extract multi-scale features; The multi-scale features extracted by the backbone network are fed into their respective attention modules to obtain the attention features corresponding to each scale feature; Feed the attention features into the multi-scale interaction module to obtain the highlight image; The highlighted image is re-input into the backbone network to obtain the features output by the last layer of the backbone network, and then sent to the feature decoupling module for decoupling to obtain the decoupled features; The decoupled features are input into the classification module to obtain the classification results.

[0005] Furthermore, the multi-scale features extracted by the backbone network are fed into respective attention modules to obtain attention features corresponding to the respective scale features, including: For each scale feature, the maximum pooling and average pooling operations are performed respectively, and then they are added together after passing through the multi-layer perceptron, and then the channel attention weight is obtained after passing through the activation function; Multiply the channel attention weight by the input scale feature to obtain the channel feature; The channel features undergo depth-wise separable convolution and activation function to generate spatial attention weights; Multiply the spatial attention weights by the channel features to obtain the final attention features. Multiply the spatial attention weights by the channel features to obtain the final attention features.

[0006] Furthermore, the attention features are fed into the multi-scale interaction module to obtain a highlight map, including: The attention features are refined through the basic convolution module; The refined features are upsampled using bilinear interpolation to keep the scale of the features consistent and obtain the upsampled attention features; Determine the range of high attention areas in the upsampled attention features, and then crop the corresponding attention map from the input tea image based on the range of high attention areas; Calculate the overlapping area of ​​each attention map and obtain the image of the overlapping area as the highlight map.

[0007] Furthermore, the feature decoupling module performs the following operations: The input features are extracted through a dual-branch convolutional layer, a normalization layer, and a linear layer, and then flattened through a fully connected layer to obtain the first and second features. The first feature and the second feature are then passed through a multi-layer perceptron to obtain the decoupled features.

[0008] Furthermore, the step of re-inputting the obtained highlighted image into the backbone network includes: According to the scale of the highlighted image, the highlighted image is input into the backbone network after extracting the corresponding scale feature.

[0009] The present application also proposes a device for identifying young tea stems based on multi-scale feature decoupling, which includes a processor and a memory storing a plurality of computer instructions. When the computer instructions are executed by the processor, the steps of the above method are implemented.

[0010] The present application proposes a method and device for identifying young tea stems based on multi-scale feature decoupling, and introduces an attention mechanism into the multi-scale features extracted by the backbone network to minimize the interference of low-quality information in the multi-scale features on the network. Secondly, a multi-scale interaction structure is proposed to mine the inactivated information hidden in the global features by forming multi-scale areas of interest for the image. Finally, by introducing a feature decoupling module, the features of the young buds and young tea stems in the image are separated into different feature subspaces, thereby eliminating redundant information. The recognition algorithm obtained by the present application through attention, multi-scale interaction, and feature decoupling fusion can achieve a high accuracy rate, effectively reduce the problem of misidentification of young tea stems in tea-picking robots during picking operations, and can reduce damage to the picking end. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 This is a flow chart of the young tea stem identification method based on multi-scale feature decoupling in this application.

[0012] Figure 2 This is the network structure diagram of this application.

[0013] Figure 3 This is the structural diagram of the characteristic decoupling module of this application. DETAILED DESCRIPTION

[0014] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0015] In one embodiment, as shown in FIG1 , a method for identifying young tea stems based on multi-scale feature decoupling is provided, comprising: Step S1: Input the collected tea leaves image into the backbone network to extract multi-scale features.

[0016] Specifically, the recognition network model constructed in this application is as follows Figure 2 As shown, this application inputs the collected tea leaves image into the constructed recognition network model to obtain the recognition result.

[0017] The backbone network adopts CNN architecture, for example, ShuffleNet V2 network can be used. The backbone network outputs feature maps of the input image through a five-stage (stage 1-5) CNN architecture and selects the last three layers of features ( x 3 , x 4 , x 5) is used as the basic feature information for subsequent feature processing, and the pixel sizes of the output feature maps are 56×56, 28×28, and 14×14 from shallow to deep.

[0018] Step S2: The multi-scale features extracted by the backbone network are sent to their respective attention modules to obtain the attention features corresponding to each scale feature.

[0019] Specifically, the attention module ( Figure 2 GSCA in

[15] includes channel attention and spatial attention. The multi-scale features extracted by the backbone network are fed into their respective attention modules to obtain the attention features corresponding to each scale feature, including: Step S2.1: For each scale feature, perform maximum pooling and average pooling operations respectively, then add them together after passing through a multi-layer perceptron, and then pass through an activation function to obtain the channel attention weight.

[0020] Max pooling and average pooling are performed on any of the three input scale features. Max pooling selects the maximum value within the pooling window, while average pooling calculates the average value within the pooling window. Both pooling results are passed through a multi-layer perceptron (MLP) to map the features to a low-dimensional space and then back to the original channel dimension, thereby capturing inter-channel dependencies.

[0021] The two feature vectors after MLP processing are added together and then activated by the Sigmoid function to obtain the channel attention weight vector.

[0022] Step S2.2: Multiply the channel attention weight by the input scale feature to obtain the channel feature.

[0023] The input scale feature is multiplied by the channel attention weight to obtain the channel feature of the channel attention output.

[0024] Step S2.3: The channel features undergo depth-wise separable convolution and activation function to generate spatial attention weights.

[0025] The channel features output by the channel attention are used as input and then passed through a depthwise separable convolution (DW Conv). The first depthwise separable convolution reduces the number of channels from C to C / r (the number of groups g equals the reduction rate r). The convolution kernel sizes are 1×7 and 7×1, respectively. Preliminary feature extraction is performed on the spatial dimensions, resulting in a feature map of size C / r × H × W. Another depthwise separable convolution is then performed, this time restoring the number of channels from C / r to C, with the convolution kernel sizes also being 1×7 and 7×1. Further spatial feature extraction is performed, resulting in a new feature map of size C × H × W. The resulting feature map is then subjected to a sigmoid function to generate spatial attention weights.

[0026] Step S2.4: Multiply the spatial attention weight with the channel feature to obtain the final attention feature.

[0027] In this embodiment, steps 2.3 and 2.4 are to perform spatial attention operations. The output of the spatial attention is the attention feature output by the entire attention module. Each scale feature is passed through its own attention module to obtain the corresponding attention feature ( a 3 , a 4 , a 5 ).

[0028] In this embodiment, spatial attention features are extracted using multi-scale depth-shared 1D convolutions (1×7 and 7×1 separable convolutions) to capture multi-semantic spatial information. Compared to traditional spatial attention mechanisms, this extracts spatial features from multiple scales rather than processing at a single scale, capturing richer local and global feature representations. When detecting objects of varying sizes, multi-scale approaches can balance the global characteristics of large objects with the local details of small objects.

[0029] Step S3: Send the attention features into the multi-scale interaction module to obtain a highlight image.

[0030] Specifically, the multi-scale interaction module performs the following operations: Step 3.1: Perform feature refinement on the attention features through the basic convolution module.

[0031] The basic convolution module (BC) contains a convolution layer, a pooling layer, a ReLU activation layer, and a normalization layer connected in sequence to further refine the features.

[0032] Step 3.2: Use bilinear interpolation to upsample the refined features to keep the scale of the features consistent and obtain the upsampled attention features.

[0033] In this process, the low-resolution attention map needs to be upsampled to high resolution, with a feature size of 56×56 as the target size, and 14×14 and 28×28 upsampled by 4 times and 2 times respectively.

[0034] Step 3.3: Determine the range of the high attention area in the upsampled attention feature, and then crop the corresponding attention map from the input tea image based on the range of the high attention area.

[0035] Find the boundary points of the high-attention region in the upsampled attention feature, obtain the coordinates of the upper left and lower right corners of the rectangle, and determine the coordinate range of the cropping box. The high-attention region is determined by traversing each position in the feature. When the attention value of a position changes from above the threshold (0.8) to below the threshold (0.2), the position can be considered a high-attention region boundary point.

[0036] For the input tea image, the original 448×448 image can be reduced to a scale of 56×56 through maximum pooling, so that the upsampled attention features can be overlaid on the tea image, facilitating subsequent cropping.

[0037] For the three attention features, we perform cropping operations on the reduced-scale tea image according to the cropping frame to obtain the cropped attention maps ( l 3 , l 4 , l 5 ).

[0038] Step 3.4: Calculate the overlapping area of ​​each attention map and obtain the image of the overlapping area as the highlight map.

[0039] Calculate three attention maps ( l 3 , l 4 , l 5 ) to obtain the final range of interest of the multi-scale interaction module, which is then cropped to obtain the highlight image. Here, cropping can be performed on any of the three attention maps to obtain the highlight image.

[0040] In order to ensure that the cropped image is the same size as the attention map, the highlight map needs to be scaled to ensure that it is consistent with the size of the third-layer backbone features (56×56).

[0041] Step S4: re-input the obtained highlighted image into the backbone network to obtain the features output by the last layer of the backbone network, and send it to the feature decoupling module for decoupling to obtain the decoupled features.

[0042] In this embodiment, the highlight image is re-input into the backbone network, and the size of the highlight image can be resized to the same size as the original input tea image, and then input into the backbone network for feature extraction.

[0043] Preferably, according to the scale of the highlight image, the highlight image is input into the backbone network to extract the corresponding scale feature. Figure 2 As shown in the figure, the scale of the highlight image is 56×56, while the feature map extracted by the third layer of the backbone is 56×56. Therefore, the highlight image is sent back to the fourth layer of the backbone (stage4), and further feature extraction is performed by stage4-5 to obtain a 14×14 feature map F.

[0044] For the further extracted feature F, it is sent to the feature decoupling module ( Figure 2 FD in), feature decoupling module such as Figure 3 As shown in Figure 2, it is used to separate different features in the image into different feature subspaces and eliminate redundant features.

[0045] The feature decoupling module decomposes the given feature map F into two orthogonal subspaces of the latent space. Non-redundancy is achieved by imposing a soft orthogonality constraint between the two representations. Specifically, the input feature map F is first subjected to a dual-branch convolutional layer, a normalization layer, and a linear layer for feature extraction, and then flattened by a fully connected layer to obtain Fc and , and then use a multi-layer perceptron (MLP) to decouple the latent space into two orthogonal subspaces. The two different sets of convolutional layers used here act as the decoupling backbone, and the decoupled features are obtained after passing through the shared MLP layer. The decoupling process can be expressed as:

[0046] in Represents the feature vectors of tender buds and tender tea stems obtained after feature decoupling operation.

[0047] It should be noted that Figure 3 The fully connected layer is omitted in the figure. That is, it represents the features obtained by the fully connected layer. In the above formula, Conv is directly used to represent all operations before MLP, which will not be repeated here.

[0048] This step can separate the features of tender buds and tender tea stems in the image into different feature subspaces. The tender buds and tender tea stems samples will form clearly separated clusters based on the texture features. The tender buds have full texture and rounded contours; the tender tea stems have slender texture.

[0049] Step S5: Input the decoupled features into the classification module to obtain the classification results.

[0050] Specifically, in the classification module, the two features output by the feature decoupling module are subjected to global maximum pooling operations respectively, and each feature map is compressed into a vector of fixed length to reduce the feature dimension while retaining the most important features. The two vectors are then used as the input of the classifier to obtain the classification result. The classifier calculates the probability that each feature vector belongs to a different category based on the feature information learned from the existing training data. Here, the thresholds of tender buds and tender tea stems are set before model training (the default setting is 0.5), and each feature vector will calculate the cross loss with the true label category. Only when the calculated probability score is greater than the set confidence level will the classification result be given. And since the feature space of tender buds and tender tea stems has been separated by the feature decoupling module in step S4, the classification module only needs to judge The recognition result can be determined by whether the calculated probability result of each channel is greater than the threshold, and whether it is a young bud or a young tea stem.

[0051] The present application utilizes a multi-scale interaction structure to strengthen the global semantics at multiple scales, and improves the quality of visual representation through real-time evaluation. Specifically, first, an attention mechanism is introduced into the multi-scale features extracted by the backbone network to minimize the interference of low-quality information in the multi-scale features on the network. Secondly, a multi-scale interaction structure is proposed to mine the inactivated information hidden in the global features by forming multi-scale areas of interest for the image. Finally, by introducing a feature decoupling module, the features of the tender buds and tender tea stems in the image are separated into different feature subspaces, thereby eliminating redundant information. The recognition algorithm obtained by the present application through attention, multi-scale interaction, and feature decoupling fusion can achieve a higher accuracy rate, effectively reduce the problem of misidentification of tender tea stems in tea picking robots during picking operations, and can reduce damage to the picking ends.

[0052] To explore the recognition accuracy of young tea stems using different lightweight network models, we selected SuffleNet V2, ResNet18, and MobileNet V3, three representative models in terms of performance and lightweightness, for comparative experiments. The scales of 0.5×, 1.0×, 1.5×, and 2.0× in the table represent different scale versions of SuffleNet V2. Table 1 shows the experimental data: Table 1

[0053] The comparison results in Table 1 demonstrate that ShuffleNet V2, the feature extraction backbone used in this application, demonstrates significant advantages over other mainstream backbones, ResNet-18 and MobileNet-V3. The ShuffleNet V2 2.0× model has a parameter size of 5.42M, significantly lower than ResNet-18's 11.24M. However, its accuracy reaches 94.99%, exceeding both MobileNet V3 Large (92.76%) and ResNet-18 (93.52%). This high accuracy achieved with relatively few parameters demonstrates the superiority of its architectural design in balancing model complexity and performance.

[0054] This application also provides the ablation experiment results as shown in Table 2: Table 2

[0055] To validate the superior model performance of the feature decoupling module, the GSCA attention module, and the multi-scale interaction module, ablation experiments were conducted. The experimental results show that the addition of FD, GSCA, and MSI improves model performance (accuracy and F1 score) to varying degrees. This demonstrates the applicability and effectiveness of these modules across diverse network architectures, enabling them to optimize models and improve performance on specific tasks. In practical applications, optimal performance can be achieved by selecting an appropriate combination of the backbone network and innovative modules based on specific requirements and resource constraints. These experimental results also provide valuable insights and foundation for further research and model improvement.

[0056] In another embodiment, the present application also provides a young tea stem identification device based on multi-scale feature decoupling, including a processor and a memory storing a plurality of computer instructions, wherein the computer instructions implement the steps of the above method when executed by the processor.

[0057] Regarding the specific limitations of the young tea stem identification device based on multi-scale feature decoupling, please refer to the limitations of the young tea stem identification method based on multi-scale feature decoupling above, which will not be repeated here. The above-mentioned young tea stem identification device based on multi-scale feature decoupling can be implemented in whole or in part by software, hardware and a combination thereof. It can be embedded in or independent of the processor in the computer device in the form of hardware, or it can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0058] The memory and processor are electrically connected, directly or indirectly, to enable data transmission or interaction. For example, these components may be electrically connected via one or more communication buses or signal lines. The memory stores a computer program executable on the processor, and the processor executes the computer program stored in the memory to implement the method for identifying young tea stems based on multi-scale feature decoupling in an embodiment of the present invention.

[0059] The memory may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. The memory is used to store a program, and the processor executes the program after receiving an execution instruction.

[0060] The processor may be an integrated circuit chip with data processing capabilities. The processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor.

[0061] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A method for identifying young tea stems based on multi-scale feature decoupling, characterized in that: The method for identifying young tea stems based on multi-scale feature decoupling includes: The collected tea images are input into the backbone network to extract multi-scale features; The multi-scale features extracted by the backbone network are fed into their respective attention modules to obtain the attention features corresponding to each scale feature; Feed the attention features into the multi-scale interaction module to obtain the highlight image; The highlighted image is re-input into the backbone network to obtain the features output by the last layer of the backbone network, and then sent to the feature decoupling module for decoupling to obtain the decoupled features; The decoupled features are input into the classification module to obtain the classification results.

2. The method for identifying young tea stems based on multi-scale feature decoupling according to claim 1, characterized in that: The multi-scale features extracted by the backbone network are fed into respective attention modules to obtain attention features corresponding to each scale feature, including: For each scale feature, the maximum pooling and average pooling operations are performed respectively, and then they are added together after passing through the multi-layer perceptron, and then the channel attention weight is obtained after passing through the activation function; Multiply the channel attention weight by the input scale feature to obtain the channel feature; The channel features undergo depth-wise separable convolution and activation function to generate spatial attention weights; Multiply the spatial attention weights by the channel features to obtain the final attention features. Multiply the spatial attention weights by the channel features to obtain the final attention features.

3. The method for identifying young tea stems based on multi-scale feature decoupling according to claim 1, characterized in that: The attention features are fed into the multi-scale interaction module to obtain a highlight map, including: The attention features are refined through the basic convolution module; The refined features are upsampled using bilinear interpolation to keep the scale of the features consistent and obtain the upsampled attention features; Determine the range of high attention areas in the upsampled attention features, and then crop the corresponding attention map from the input tea image based on the range of high attention areas; Calculate the overlapping area of ​​each attention map and obtain the image of the overlapping area as the highlight map.

4. The method for identifying young tea stems based on multi-scale feature decoupling according to claim 1, characterized in that: The feature decoupling module performs the following operations: The input features are extracted through a dual-branch convolutional layer, a normalization layer, and a linear layer, and then flattened through a fully connected layer to obtain the first and second features. The first feature and the second feature are then passed through a multi-layer perceptron to obtain the decoupled features.

5. The method for identifying young tea stems based on multi-scale feature decoupling according to claim 1, characterized in that: The step of re-inputting the obtained highlighted image into the backbone network includes: According to the scale of the highlighted image, the highlighted image is input into the backbone network after extracting the corresponding scale feature.

6. A device for identifying young tea stems based on multi-scale feature decoupling, comprising a processor and a memory storing a plurality of computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.