Wheat imperfect grain identification method based on double-branch network and feature fusion

Through the method of fusion of dual-branch networks and feature, combined with improved Swin Transformer and multi-scale dynamic convolution, the accuracy and efficiency problems in wheat imperfect grain recognition are solved, and high-precision and efficient recognition effects are achieved.

CN120279545APending Publication Date: 2025-07-08HENAN UNIVERSITY OF TECHNOLOGY

Patent Information

Application Number
CN202510299615.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art has low accuracy and poor efficiency in wheat imperfect grain recognition, making it difficult to take into account feature extraction of global semantics and local details, and the multi-scale feature fusion is not dynamic enough, resulting in information loss or noise interference.

Method used

Using a dual-branch network architecture, combined with the improved Swin Transformer and multi-scale dynamic convolutional branches, multi-stage feature extraction and fusion are achieved through the cross-attention fusion module and the progressive feature fusion mechanism, and the feature weight allocation is optimized using the window-level and head-level sparse attention mechanisms.

Benefits of technology

It significantly improves the recognition accuracy and robustness of imperfect wheat grains, improves the recognition efficiency, reduces calculation overhead, and meets the real-time detection needs of agricultural scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279545A_ABST
    Figure CN120279545A_ABST
Patent Text Reader

Abstract

The invention discloses a wheat imperfect grain identification method based on a double-branch network and feature fusion, and the method comprises the following steps: S1, inputting a wheat imperfect grain image, and generating initial features through image segmentation and linear embedding; s2, performing multi-stage feature extraction by adopting a double-branch network architecture, specifically, improving a Swin Transform and a multi-scale dynamic convolution branch; s3, after each level of feature extraction, performing interactive fusion on the double-branch features through a cross attention fusion module; s4, aggregating multi-scale features through a four-stage feature fusion module; and S5, finally outputting an identification result through a pooling linear layer. According to the method, the accuracy rate in the wheat imperfect grain multi-classification task is remarkably improved and reaches 98.82%, a high-precision and low-cost automatic solution is provided for grain quality detection, and the problems that an existing method is insufficient in recognition precision and limited in feature expression ability are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical fields of computer vision and deep learning, and particularly to a method for identifying imperfect wheat kernels based on a dual-branch network and feature fusion. This method combines an improved Swin Transformer network with a dynamic multi-scale convolution branch, and through a progressive feature fusion mechanism, it realizes high-precision automatic detection of imperfect wheat kernels, and can be widely applied to agricultural grain quality detection, intelligent sorting equipment, and the field of agricultural product processing. Background Art

[0002] In agricultural production, the rapid and accurate identification of imperfect wheat kernels (such as mildewed kernels, broken kernels, insect-eaten kernels, etc.) is crucial for grain quality grading and storage management. Traditional detection methods mainly rely on manual visual sorting or image processing techniques based on hand-designed features, which have problems such as low efficiency, strong subjectivity, and poor generalization ability. Although deep learning techniques have made significant progress in image classification tasks in recent years, in practical applications, they still face the following challenges: on the one hand, due to the subtle differences in the morphology and texture of imperfect wheat kernels, it is difficult for a single convolutional neural network or Transformer model to simultaneously capture local details and global context information, resulting in insufficient feature extraction ability; on the other hand, existing deep learning models have defects in multi-scale feature fusion. Traditional methods use simple splicing or weighted summation, and cannot dynamically adjust the weights of cross-modal features, easily leading to information loss or noise interference.

[0003] The patent document with the publication number CN116704187A discloses a semantic alignment real-time semantic segmentation method, system, and storage medium, belonging to the field of real-time semantic segmentation, including: training stage: training a real-time semantic segmentation model using a training set, where each sample in the training set includes an image related to the semantic segmentation task and a corresponding label for indicating the segmentation result; the real-time semantic segmentation model includes: a convolutional network, a semantic extraction branch, and a semantic alignment network; the convolutional network includes a local convolutional module, a convolutional attention module, and a semantic decoding module; application stage: inputting the image to be processed into the trained convolutional network to obtain the segmentation result. The present invention can improve the accuracy of the real-time semantic segmentation task, and by discarding the semantic extraction branch and the semantic alignment network during inference, simplifies the backbone network during inference to a single-branch convolutional network, and with the characteristics of high segmentation accuracy and low inference latency, improves the deployment prospect of the real-time semantic segmentation network in practical applications.

[0004] The patent document with the publication number CN114627081A discloses a method for identifying imperfect wheat grains based on Mask R-CNN. In view of the problems of the existing manual detection being cumbersome, the imaging equipment being too expensive, the accuracy being low, and the lack of practicality, the following solutions are proposed, which include the following steps: S1: Collect image samples of imperfect wheat grains, perform pixel-level annotation on them, and obtain the labeled tag image data; S2: Construct a deep learning network model suitable for imperfect wheat grain images; S3: Use the tag image obtained in S1 as a training sample to train and fine-tune the deep learning network model built in S2 to obtain the best model; S4: Use the trained model to detect the test set.

[0005] The patent document with the publication number CN116704188A discloses an algorithm for segmenting wheat grain images with different bulk densities based on an improved U-Net network. The residual stacking module is used as the main backbone structure in the downsampling process. In the feature fusion part, after the primary feature map, the CBAM attention mechanism module is embedded to adaptively adjust the feature fusion weights of different pixel points from the channel and spatial positions; in the decoder part, the self-attention module is embedded to enhance the correlation between different targets.

[0006] The patent document with the publication number CN113869251A discloses an online detection method for imperfect wheat grains based on an improved ResNet. This method builds an image acquisition platform for imperfect wheat grains, establishes an algorithm for segmenting two adhered wheat grains based on a concave point mask, combines the dual advantages of the device and the algorithm to single-grain the mixed wheat; enhances the image in three ways, namely, increasing the brightness after rotating the image by 180°, adding noise, and reducing the brightness after mirroring, to establish a data set; constructs a detection model for imperfect wheat grains based on an improved ResNet by reducing the number of network layers, introducing dense connections, adding an attention mechanism, adding dilated convolutions, and increasing the receptive field; uses the pixels, bulk density, etc. of wheat images to establish a detection method for the imperfect grain rate of wheat based on the national standard mass ratio, providing a scientific basis for the quality evaluation in the wheat market circulation.

[0007] In the prior art, no solution that takes into account both high accuracy and high efficiency has been proposed. Therefore, providing a new network architecture to improve the accuracy and robustness of imperfect wheat grain recognition is a problem worthy of research. Summary of the Invention

[0008] The purpose of the present invention is to propose a method for identifying imperfect wheat grains based on a dual-branch network and feature fusion, so as to solve the problems of low accuracy and poor efficiency in traditional imperfect wheat grain recognition, and improve the recognition accuracy and recognition efficiency of imperfect wheat grains in actual production.

[0009] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0010] A method for identifying imperfect wheat kernels based on a dual-branch network and feature fusion, comprising the following steps:

[0011] S1, input the image of imperfect wheat kernels, and generate initial features through image segmentation and linear embedding;

[0012] S2, adopt a dual-branch network architecture for multi-stage feature extraction, specifically an improved Swin Transformer and a multi-scale dynamic convolution branch;

[0013] S3, after each stage of feature extraction, interactively fuse the dual-branch features through a cross-attention fusion module;

[0014] S4, aggregate multi-scale features through a four-stage feature fusion module;

[0015] S5, finally output the recognition result through a pooling linear layer.

[0016] In step S1, the input image is divided into image patches of size 4x4 through convolution operations, and each image patch, that is, each patch, is used as an input sample for subsequent feature extraction and processing.

[0017] In step S1, each image patch is mapped to an embedding space using a convolutional layer, and through flattening and transposing operations, the features of each image patch are converted into a sequence element.

[0018] In step S2, a dual-branch network architecture is adopted, combined with an improved Swin Transformer module, and global features are extracted using a dual sparse attention mechanism at the window level and the head level. Among them, the window-level sparse attention calculates the window importance score through a gating network, obtains the maximum value along the spatial dimension, and then weights the attention weights within the window with the importance score.

[0019] In step S2, for the head-level sparse attention, define the learnable parameter α h Scale the value vector V: V′ = V ⊙ α h , α h is initialized as a vector of all 1s, and the importance weights of each head are learned during training. The composite attention output is: Output = Concat(Attn′1V′,..., Attn′ H V′ H )W o , where H is the number of attention heads, and W o is the projection matrix.

[0020] Step S2, the multi-scale dynamic convolution branch processes the input feature map, extracts multi-scale features through multiple parallel branches. For each branch, channel compression is first performed, and then dilated convolutions are used to extract local, medium, and global features respectively. Subsequently, the number of channels is restored to the original. The weights are adaptively generated through global average pooling and convolution. Finally, a channel recalibration module is used to enhance the features and output multi-scale features.

[0021] Step S3, the cross-attention fusion module receives the sequence features of the Swin Transformer and the spatial features of the CNN, generates a query vector q using a linear transformation, generates key vector k and value vector v through convolution, calculates the similarity between q and k to generate attention weights and weights v, and the output is the sum of the original features and the weighted features, which is adjusted by learnable parameters.

[0022] Step S4, the four-stage feature fusion module is independently configured for feature maps of different resolutions. In the first two stages, an early weak fusion strategy is adopted, and the fusion intensity is reduced through fixed weights. In the last two stages, a late adaptive fusion strategy is adopted, and weights are generated through dynamic gating to achieve adaptive allocation of feature importance.

[0023] Step S5, through global average pooling and spatial compression, the global features and spatial features of each stage are concatenated to retain multi-scale information. A linear layer is used to map the compressed features to the category space, and the recognition result of wheat imperfect grains is output.

[0024] Positive and beneficial effects:

[0025] 1. The present invention adopts a dual-branch collaborative architecture, combines the global modeling ability of the improved Swin Transformer with the local feature extraction ability of multi-scale dynamic convolution, and solves the problem that traditional single-branch models are difficult to balance global semantics and local details through complementary learning of the dual branches, significantly improving the fine-grained recognition accuracy of wheat imperfect grains, reaching 98.82% on the test set.

[0026] 2. The present invention designs a progressive dynamic feature fusion mechanism, realizes multi-level feature optimization through cross-attention and channel recalibration, and uses the cross-attention module to dynamically allocate feature weights, solving the problem of noise interference caused by traditional feature concatenation or simple addition, and enhancing the expression ability of key features.

[0027] 3. The present invention introduces a dual sparse attention mechanism at the window level and the head level in the Swin Transformer, suppresses redundant calculations through dynamic gating, solves the problem of high computational complexity of traditional self-attention, reduces the computational overhead by 20% while ensuring accuracy, significantly improves the model inference efficiency, and meets the real-time detection requirements of agricultural scenarios.

[0028] 4. The present invention proposes a multi-scale dynamic convolution module and an adaptive downsampling strategy. By using dilated convolution to capture spatial information at different granularities, the problem of poor adaptability of single-scale convolution to wheat morphological changes is solved, and the robustness of the model to imperfect grains such as damaged and mildewed grains is significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 FIG. is an overall system block diagram of a wheat imperfect grain recognition method based on a dual-branch network and feature fusion proposed by the present invention;

[0030] Figure 2 FIG. is a flowchart of the multi-scale dynamic convolution module of the present invention;

[0031] Figure 3 FIG. is a flowchart of the cross-attention fusion module of the present invention;

[0032] Figure 4 FIG. is a flowchart of the head-level sparse attention module of the present invention;

[0033] Figure 5 FIG. is a flowchart of the window-level sparse attention module of the present invention;

[0034] Figure 6 FIG. is an evaluation diagram of the model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] The present invention will be further described in detail below with reference to the drawings and embodiments, but the embodiments of the present invention are not limited thereto.

[0036] A wheat imperfect grain recognition method based on a dual-branch network and feature fusion includes the following steps:

[0037] S1. Input an image of wheat imperfect grains, and generate initial features through image segmentation and linear embedding;

[0038] S2. Adopt a dual-branch network architecture for multi-stage feature extraction, specifically an improved Swin Transformer and a multi-scale dynamic convolution branch;

[0039] S3. After each stage of feature extraction, interact and fuse the dual-branch features through a cross-attention fusion module;

[0040] S4. Aggregate multi-scale features through a four-stage feature fusion module;

[0041] S5. Finally, output the recognition result through a pooling linear layer.

[0042] Embodiment 1

[0043] As Figure 1 shown, the implementation process of the method of the present invention specifically includes the following steps:

[0044] Input image preprocessing and feature embedding. Input the image of imperfect wheat kernels. Use image preprocessing to adjust the image to a size of 224×224. Divide the image into 4×4 image patches through convolutional operations. Each image patch, i.e., each patch, is regarded as an independent input sample for subsequent feature extraction and processing. In a specific implementation, a convolutional layer with a kernel size of 4×4 and a stride of 4 is used to map the input image from 3 channels to a 96-dimensional embedding space. The output dimension is [B, 96, 56, 56], where B is the batch size. Through flattening and transposing operations, the features are converted into a sequence form [B, 56×56, 96], which serves as the input to the Swin Transformer branch. At the same time, a normalization layer is used to normalize the embedded features to ensure training stability.

[0045] As Figure 2 , Figure 4 , Figure 5 shown, the dual branch includes an improved Swin Transformer branch and a multi-scale dynamic convolutional branch. The specific implementation is as follows:

[0046] Swin Transformer branch: Based on the improved Swin Transformer module, global features of wheat images are extracted through window-level and head-level dual sparse attention mechanisms. The window-level sparse attention mechanism calculates the window importance score through a gating network. The formula is as follows:

[0047] g w =σ(W2*GELU*(W1*MaxPool(W i ))), where W1 and W2 are learnable parameters, σ is the Sigmoid function, and the maximum value is taken along the spatial dimension through MaxPool. The weighted formula for the attention weight Attn within the window and the importance score is: Among them, Q and K represent the query matrix respectively, is the scaling factor, where d is the dimension of the query and the key. The scaling factor is used to prevent gradient vanishing or explosion. B is the relative position encoding, and α w is a learnable scaling factor. By suppressing the attention weights of low-importance windows, redundant calculations are reduced. The head-level sparse attention mechanism dynamically scales the value vector through learnable parameters, and automatically optimizes the importance weights of each attention head during the training process to suppress irrelevant features.

[0048] Multi-scale dynamic convolution branch. The input feature map is x, with dimensions [B, C, H, W], where B is the batch size, C is the number of channels, and H and W are the height and width of the feature map respectively. Multiple parallel convolution branches are used to perform multi-scale feature extraction on the input feature map. Each branch contains the following processes: use a 1x1 convolution to compress the number of input channels from C to C / 4, use a 3x3 dilated convolution to extract local features at different scales, with dilation rates of 1, 2, and 3 respectively, corresponding to local, medium, and global scale feature extraction, use a 1x1 convolution to restore the number of channels from C / 4 to C, and the weights are adaptively adjusted according to the content of the input feature map. Global average pooling and 1x1 convolution are used to generate the weights to ensure efficient and lightweight weight calculation. The fused features are enhanced in the channel dimension through a channel recalibration module, and the final multi-scale features are output.

[0049] Cross-attention feature fusion, as Figure 3 shown, the feature input is a two-branch input. Swin Transformer branch: The input sequence feature is denoted as swin_feat ∈ R B*L*C , where B is the batch size, L is the sequence length, and C is the feature dimension. The CNN branch inputs spatial features, denoted as cnn_feat ∈ R B*H*W*C , where H and W are the height and width of the feature map respectively. A query vector q is generated from the sequence feature of the Swin Transformer branch through a linear transformation, and the formula is: q = Linear(swin_feat). A key vector k and a value vector v are generated from the spatial features of the CNN branch through a convolution operation, and the formulas are: k = Conv2d(cnn_feat), v = Conv2d(cnn_feat). The similarity between the query vector q and the key vector k is calculated using matrix multiplication to generate an attention weight matrix attn, and the formula is attn = softmax(q · k T ). The value vector v is weighted using the attention weight matrix to obtain the fused feature fused, and the formula is: fused = attn · v. The fused feature is weighted and summed with the feature of the original Swin Transformer branch to obtain the final fused feature, and the formula is: output = swin_feat + γ · fused, where γ is a learnable weight parameter and swin_feat is the original input sequence feature. Through the cross-attention mechanism, the local detailed features of the CNN branch can be effectively fused with the global features of the Swin Transformer branch, thereby improving the accuracy and robustness of wheat imperfect grain recognition.

[0050] Progressive multi-stage feature fusion. The four-stage feature fusion module includes stage1 to stage4, which correspond to feature maps of different resolutions respectively. Each stage independently configures a fusion module to achieve progressive feature enhancement from local details to global semantics. In the first and second stages, namely the high-resolution feature stage, an early weak fusion strategy is adopted. By fixing the weights, the fusion intensity is reduced to retain more original feature details and avoid overfitting. In the third and fourth stages, namely the low-resolution feature stage, a late adaptive fusion strategy is adopted. Weights are generated through dynamic gating. Among them, the gating network generates spatial attention weights through global average pooling and a fully connected layer to achieve adaptive allocation of feature importance.

[0051] Multi-scale feature aggregation and classification. Through global average pooling and spatial compression, the feature dimension is reduced, and the computational complexity is lowered. The global features and spatial features of each stage are concatenated to retain multi-scale information and improve the classification accuracy. A linear layer is used to map the compressed features to the category space to output the recognition result of wheat imperfect grains.

[0052] Such as Figure 6 , experiments show that the present invention achieves an accuracy of 98.82% on the test set, which is significantly improved compared with traditional models, and significantly improves the classification accuracy and computational efficiency.

[0053] The present invention adopts a dual-branch collaborative architecture, combines the global modeling ability of the improved Swin Transformer with the local feature extraction ability of multi-scale dynamic convolution, and solves the problem that traditional single-branch models are difficult to balance global semantics and local details through dual-branch complementary learning. The fine-grained recognition accuracy of wheat imperfect grains is significantly improved, reaching 98.82% on the test set.

[0054] The present invention designs a progressive dynamic feature fusion mechanism to achieve multi-level feature optimization through cross-attention and channel recalibration. The cross-attention module is used to dynamically allocate feature weights, solving the problem of noise interference caused by traditional feature concatenation or simple addition, and enhancing the expression ability of key features.

[0055] The present invention introduces a dual sparse attention mechanism at the window level and head level in the Swin Transformer. By dynamically gating to suppress redundant calculations, the problem of high computational complexity of traditional self-attention is solved. While ensuring the accuracy, the computational overhead is reduced by 20%, significantly improving the model inference efficiency and meeting the real-time detection requirements of agricultural scenarios.

[0056] The above embodiments are only preferred embodiments of the present invention, and the protection scope of the present invention is not limited thereto. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

Claims

1. A method for identifying imperfect wheat kernels based on a dual-branch network and feature fusion, characterized in that It includes the following steps: S1. Input the image of wheat imperfect grains, and generate initial features through image segmentation and linear embedding. S2. Adopt a dual-branch network architecture for multi-stage feature extraction, specifically an improved Swin Transformer and a multi-scale dynamic convolution branch. S3. After each stage of feature extraction, interactively fuse the dual-branch features through a cross-attention fusion module. S4. Aggregate multi-scale features through a four-stage feature fusion module. S5. Finally, output the recognition result through a pooling linear layer.

2. The wheat imperfect grain recognition method based on a dual-branch network and feature fusion according to claim 1, wherein: In step S1, the input image is divided into image patches of size 4x4 through convolution operations. Each image patch, that is, each patch, serves as an input sample for subsequent feature extraction and processing.

3. The wheat imperfect grain recognition method based on a dual-branch network and feature fusion according to claim 1, characterized in that: In step S1, a convolutional layer is used to map each image patch to an embedding space. Through flattening and transposing operations, the features of each image patch are converted into a sequence element.

4. A method for identifying imperfect kernels of wheat based on a dual-branch network and feature fusion according to claim 1, characterized in that: In step S2, a dual-branch network architecture is adopted, combined with an improved Swin Transformer module, to extract global features using a dual sparse attention mechanism at the window level and the head level. Among them, the window-level sparse attention calculates the window importance score through a gating network, obtains the maximum value along the spatial dimension, and then weights the attention weights within the window with the importance score.

5. A wheat imperfect grain recognition method based on a dual-branch network and feature fusion according to claim 1, characterized in that: Step S2, head-level sparse attention, define learnable parameter α h Scale the value vector V: V′ = V ⊙ α h , α h is initialized as a vector of all 1s, and the importance weights of each head are learned during training. The composite attention output is: Output = Concat(Attn′1V′,......, Attn″ H V′ H )W o , where H is the number of attention heads, and W o is the projection matrix.

6. The wheat imperfect grain recognition method based on a dual-branch network and feature fusion according to claim 1, characterized in that: In step S2, the multi-scale dynamic convolution branch processes the input feature map, extracts multi-scale features through multiple parallel branches. Each branch first performs channel compression, and then uses dilated convolution to extract local, medium, and global features respectively, and then restores to the original number of channels. The weights are adaptively generated through global average pooling and convolution. Finally, the channel recalibration module is used to enhance the features and output multi-scale features.

7. A wheat imperfect grain recognition method based on a dual-branch network and feature fusion according to claim 1, characterized in that: In step S3, the cross-attention fusion module receives the sequence features of the Swin Transformer and the spatial features of the CNN, generates a query vector q using a linear transformation, generates a key vector k and a value vector v through convolution, calculates the similarity between q and k to generate attention weights and weights v, and the output is the sum of the original features and the weighted features, which is adjusted through learnable parameters.

8. A wheat imperfect grain recognition method based on a dual-branch network and feature fusion according to claim 1, characterized in that: In step S4, the four-stage feature fusion module is independently configured for feature maps of different resolutions. Among them, the early weak fusion strategy is adopted in the first and second stages, and the fusion intensity is reduced through fixed weights. The late adaptive fusion strategy is adopted in the third and fourth stages, and weights are generated through dynamic gating to achieve adaptive allocation of feature importance.

9. The wheat imperfect grain recognition method based on a dual-branch network and feature fusion according to claim 1, characterized in that: In step S5, through global average pooling and spatial compression, the global features and spatial features of each stage are concatenated to retain multi-scale information. A linear layer is used to map the compressed features to the class space, and the recognition result of wheat imperfect grains is output.

Citation Information

Patent Citations

  • Improved ResNet-based wheat imperfect grain online detection method

    CN113869251A

  • Mask R-CNN-based wheat imperfect grain identification method

    CN114627081A

  • Real-time semantic segmentation method and system for semantic alignment and storage medium

    CN116704187A

  • Different volume weight wheat grain image segmentation algorithm based on improved U-Net network

    CN116704188A

Cited By

  • Wild animal identification method and device based on image

    CN121305623A

  • Infrared and visible light image fusion method and system based on progressive attention

    CN121458557A