Wheat imperfect grain detection method based on local attention and learnable confrontation
By introducing parallel convolutional branch structure, multi-branch feature weighted network model and deformable attention DAS module in wheat imperfect grain detection, combined with the learning adversarial training strategy, the shortcomings of existing detection methods in recognition accuracy and robustness are solved, and more efficient and stable detection effects are achieved.
Patent Information
- Application Number
- CN202510181121.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-10
AI Technical Summary
The existing wheat imperfect grain detection methods have shortcomings in recognition accuracy and robustness, especially in complex environments that are susceptible to noise or external interference, resulting in a degradation of recognition performance.
A wheat imperfect grain detection method based on local attention and learning-can-adversarial adversariality is proposed, and multi-scale features are captured through parallel convolutional branch structures and multi-branch feature-weighted network model (MBCWNet), and deformable attention DAS module is embedded in each feature extraction unit. At the same time, a learnable adversarial training strategy (LAS-AWP) is introduced, and the policy network is used to automatically generate adversarial samples for model training, enhancing the anti-interference ability and robustness of the model.
By introducing local attention and learning adversarial training strategies, the recognition accuracy and robustness of wheat imperfect grain detection model is significantly improved, and the recognition accuracy can be maintained in complex environments, which is suitable for practical industrial applications.
Smart Images

Figure CN120125519A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of machine learning, and particularly relates to a method for detecting imperfect wheat kernels based on local attention and learnable adversarial mechanism. Background Art
[0002] During the storage and transportation of wheat after harvesting, some wheat kernels are prone to form imperfect kernels due to damage or disease infection. The existence of these imperfect kernels not only affects the overall market value of wheat, but also poses a potential threat to the health of ultimate consumers. Therefore, how to accurately identify imperfect kernels in wheat is of great significance for ensuring food quality and safety, improving the efficiency of the food supply chain, and enhancing the economic benefits of the industry.
[0003] Currently, the identification of imperfect wheat kernels has become a research hotspot in both academia and industry. Traditionally, people mainly rely on manual inspection to identify imperfect kernels. However, manual inspection is not only inefficient but also highly subjective, and it is prone to inconsistent identification results due to differences in the experience of inspectors. Therefore, this method is difficult to meet the requirements of large-scale and efficient detection. To solve this problem, more and more researchers have started to use machine learning and deep learning technologies to explore more automated identification methods.
[0004] Traditional machine learning methods for detecting imperfect wheat kernels usually rely on manually designed and extracted features. Although these methods can improve the identification accuracy to a certain extent, the generalization ability of their models is poor, and it is difficult to adapt to complex and changing actual environments. In addition, feature engineering often requires a large amount of domain knowledge, resulting in a high cost for model development. On the other hand, hyperspectral imaging technology, as a new detection method, can provide more detailed spectral information, thus improving the accuracy of identification. However, hyperspectral imaging equipment is expensive, and its image processing process is complex and cumbersome, which limits its popularization in practical applications. With the rapid development of deep learning technology, convolutional neural networks (CNNs) have gradually become the mainstream method in the field of identifying imperfect wheat kernels. Compared with traditional machine learning, deep learning-based models can automatically learn features from a large amount of data, reducing the dependence on manual feature extraction, and showing significant advantages in identification accuracy. However, although existing deep learning models have achieved high accuracy in the identification of imperfect wheat kernels, there is still room for improvement in the identification accuracy of the models. In addition, these methods are not robust enough in complex environments and are easily affected by noise or external interference, resulting in a significant decline in identification performance.
[0005] Therefore, further improving the accuracy and robustness of the wheat imperfect grain recognition model and developing new methods that can operate stably in the actual environment are still the key issues that current researchers urgently need to solve. To meet this demand, proposing a more efficient and stable wheat imperfect grain recognition method has important application prospects for ensuring food security and improving agricultural production efficiency. Summary of the Invention
[0006] The present invention proposes a wheat imperfect grain detection method based on local attention and learnable adversarial training. First, through a parallel convolutional branch structure, the multi-branch feature weighted network model (MBCWNet) can capture feature information at different scales and uses a channel weighting strategy for feature segmentation, thereby enhancing the model's multi-scale perception ability. Secondly, a deformable attention DAS module is embedded in each feature extraction unit, effectively improving the attention weight of the model to significant feature regions and further improving the recognition accuracy. Finally, a learnable adversarial training strategy (LAS-AWP) is introduced, and the policy network is used to automatically generate the most targeted adversarial samples for model training, thereby significantly enhancing the anti-interference ability and robustness of the model. The invention can be widely applied to the fields of wheat quality detection and food security assurance.
[0007] The object of the present invention is achieved by the following technical solutions:
[0008] In the first aspect, the present invention provides a wheat imperfect grain detection method based on local attention and learnable adversarial training, including the following steps:
[0009] Step 1: Construct the detection model MBCWNet-DAS;
[0010] The detection model MBCWNet-DAS is a multi-branch model MBCWNet with a parallel structure constructed based on the ConvNeXt network structure, and deformable attention is inserted into each Block of the multi-branch model MBCWNet;
[0011] Step 2: Construct a learnable adversarial training framework LAS-AWP to generate adversarial samples AE;
[0012] The learnable adversarial training framework LAS-AWP includes a policy network and a sample generator;
[0013] The policy network, based on the clean samples and the attention weights backpropagated by the detection model MBCWNet-DAS, generates targeted attack strategies by perturbing the attention weights;
[0014] The sample generator, based on the attack strategy, automatically generates attack samples AE according to the characteristics of the current task stage and the clean samples;
[0015] Step 3: Mix the generated adversarial examples AE with the clean examples and input them into the multi-branch model MBCWNet for joint training;
[0016] Step 4: Use the MBCWNet-DAS model after adversarial training for the wheat imperfect grain detection task in actual industrial applications.
[0017] Based on the above, the construction method of the detection model MBCWNet-DAS is as follows:
[0018] (1) First, select appropriate numbers of convolutional branches and convolutional kernel sizes, and construct a multi-branch model MBCWNet with a parallel structure based on the ConvNeXt network structure for preliminary feature extraction to obtain a primary feature map R b×c×h×w :
[0019] f 1 : x 1 → x’ 1 ∈ R b×c×h×w
[0020] f 2 : x 2 → x’ 2 ∈ R b×c×h×w
[0021] ……
[0022] f n : x n → x’ n ∈ R b×c×h×w
[0023] Among them, f 1 , f 2 ,…f n are all depth convolutions with convolutional kernels of different sizes;
[0024] Then, calculate the weights of each channel of the primary feature map and input them into each branch of the parallel structure in turn according to the weight magnitudes; each branch uses a different convolutional kernel size to capture image detail information at different scales, forming a multi-scale feature vector set V:
[0025] x 1 , x 2 ,…x n = Split(max(W 1 , W 2 ,…W n ))
[0026] W 1 , W 2 ,…W n = ECA(x)
[0027] ECA is a channel weight calculation method, and W 1 , W 2 , …W n are the weights of each channel of the primary feature map;
[0028] Next, fuse the information flows of each branch and retain the most representative feature information of the wheat imperfect grains:
[0029] x' = fuse(x' 1 , x' 2 , …x' n )
[0030] Fuse is the concat operation;
[0031] Finally, introduce the Gelu activation function and Batch Normalization to further compact the features:
[0032] y = Gelu(Batch Normalization(x′))
[0033] (2) First, insert the deformable attention DAS module into each Block of the multi-branch model MBCWNet;
[0034] The deformable attention DAS module predicts the offset of each pixel in the feature map through the deformable convolution mechanism, and geometrically deforms the feature map according to these offsets to capture the morphological features of the imperfect grains;
[0035]
[0036] In the formula, def(q) represents the output result of the deformable convolution at position q; W i represents the weight of the convolution kernel at position i; W q represents the weight of the offset position, which is used to adjust the weight of the upsampling position of the input feature map; q ref,k represents the fixed sampling position of the convolution kernel relative to the reference position q; Δq i represents the learned offset in the deformable convolution, which is used to adjust the sampling point at position i of the convolution kernel; q ref,k +Δq i represents the adjusted sampling position;
[0037] Then, apply a Sigmoid activation function and layer normalization after the deformable attention DAS module to obtain the attention result:
[0038] M = Sigmoid(Layer Normalization(def(X c )))
[0039] Where X c represents the feature map at position c, and def(X c ) represents the output result of the deformable convolution at position c, and M represents the attention result;
[0040] Finally, multiply the obtained attention result by the original input x:
[0041] x' = x ⊙ M.
[0042] Based on the above, the total adversarial training loss function of the learnable adversarial training framework LAS-AWP is:
[0043]
[0044] is the loss function for evaluating the quality of the policy network:
[0045]
[0046] is the loss function for evaluating the attack strategy:
[0047]
[0048] Among them, α and β are used to balance and a represents the new attack strategy generated by the policy network, p(a|e; s) represents the automatically generated sample-dependent strategy, λ is the step size, e is the clean sample, is the perturbation, is the newly generated adversarial sample, t represents the parameters of MBCWNet-DAS, is the new parameter of MBCWNet-DAS, f(·) represents a basic attack, is a basic loss function, y is the label of the clean sample, s represents the parameters of the policy network,
[0049] In a second aspect, the present invention provides a wheat imperfect grain detection system based on local attention and learnable adversarial, including:
[0050] A construction module for constructing a wheat imperfect grain detection network; the wheat imperfect grain detection network includes: an MBCWNet-DAS model and a learnable adversarial training framework LAS-AWP;
[0051] The detection model MBCWNet-DAS is a multi-branch model MBCWNet with a parallel structure constructed based on the ConvNeXt network structure, and deformable attention is inserted into each Block of the multi-branch model MBCWNet;
[0052] The learnable adversarial training framework LAS-AWP includes a policy network and a sample generator;
[0053] The policy network generates targeted attack strategies by perturbing the attention weights based on clean samples and the attention weights backpropagated by the detection model MBCWNet-DAS;
[0054] The sample generator automatically generates attack samples AE based on the attack strategies and the characteristics of the clean samples according to the current task stage;
[0055] An acquisition module for acquiring clean samples;
[0056] A training module for training the wheat imperfect grain detection network based on clean samples and the generated adversarial samples AE to obtain a trained wheat imperfect grain detection network;
[0057] An identification module for detecting wheat imperfect grains through the trained wheat imperfect grain detection network in the wheat imperfect grain detection task in actual industrial applications.
[0058] In a third aspect, the present invention provides an electronic device, including:
[0059] A memory and a processor, where the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the wheat imperfect grain detection method based on local attention and learnable adversaries.
[0060] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the wheat imperfect grain detection method based on local attention and learnable adversaries as described above.
[0061] In a fifth aspect, the present invention provides a computer program product including computer programs / instructions, and when the computer programs / instructions are executed by a processor, they implement the wheat imperfect grain detection method based on local attention and learnable adversaries as described above.
[0062] The beneficial effects of the present invention are:
[0063] (1) The MBCWNet-DAS model of the present invention can capture multi-scale features while maintaining the computational efficiency of the model by introducing the multi-branch structure MBCWNet based on channel weight segmentation, effectively reducing the complexity of the model and improving the overall recognition speed.
[0064] The DAS local attention mechanism is added inside the MBCWNet model, enabling the MBCWNet-DAS model to adaptively dynamically adjust the feature weights according to the saliency of the input data, thereby enhancing the recognition accuracy and stability of the model when dealing with complex backgrounds.
[0065] (2) The present invention adopts the learnable adversarial training framework LAS-AWP. Based on the clean samples and the attention weights backpropagated by the detection model MBCWNet-DAS, targeted attack strategies are generated by perturbing the attention weights, enabling the MBCWNet-DAS model to automatically identify its potential weaknesses and generate highly targeted adversarial samples. By training with these adversarial samples, the robustness of the MBCWNet-DAS model is significantly improved, enabling it to maintain a high recognition accuracy in the face of external interference. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 It is the network structure diagram of the present invention.
[0067] Figure 2 It is the structure diagram of MBCWNet-Block based on channel weight segmentation (the darker the color in the gradient color, the greater the weight).
[0068] Figure 3 It is the structure diagram of the MBCWNet-DAS model (different colors in the figure represent different degrees of attention). DETAILED DESCRIPTION OF THE INVENTION
[0069] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the protection scope of the present invention is not limited to the following.
[0070] AWP is a method to improve the robustness of the model by perturbing the model parameters. AWP introduces perturbations in the model parameter space. The core idea of AWP is that if a model is robust to the weight parameters, then it will also be more robust to the adversarial samples of the input.
[0071] Such as Figure 1As shown, after the detection model MBCPNEXeXt-DAS and the learnable adversarial training framework LAS-AWP are constructed in the present invention, after the policy network in the learnable adversarial training framework LAS-AWP learns the characteristics of the data samples, it automatically generates an attack strategy. Subsequently, the adversarial samples generated according to this attack strategy will have stronger sample dependence, enabling them to better target the weaknesses of the target network (MBCPNEXeXt-DAS) for attack, so as to better optimize the robustness of MBCPNEXeXt-DAS.
[0072] In the first aspect, as Figure 2 shown, this embodiment provides a method for detecting imperfect grains of wheat based on local attention and learnable adversarial, which includes the following steps:
[0073] S1: Collect the received image features through a multi-branch model based on channel weight segmentation to form a multi-scale feature vector set V;
[0074] S11: First, select appropriate numbers of convolutional branches and convolutional kernel sizes, and construct a multi-branch model MBCWNet with a parallel structure based on the ConvNeXt network structure. The structure of MBCWNet is as Figure 2 shown, which is used for preliminary feature extraction to obtain a primary feature map R b×c×h×w :
[0075] f 1 : x 1 →x’ 1 ∈R b×c×h×w
[0076] f 2 : x 2 →x’ 2 ∈R b×c×h×w
[0077] ……
[0078] f n : x n →x’ n ∈R b×c×h×w
[0079] Among them, f 1 , f 2 ,…f n are all depth convolutions with convolutional kernels of different sizes; x 1 , x 2 ,…x n are the original feature maps after segmentation; x’ 1 , x’ 2 ,…x’ n are the new feature maps obtained after convolution;
[0080] S12: Calculate the weights of each channel of the primary feature map and sequentially input them into each branch of the parallel structure according to the weight magnitudes. Each branch uses a different convolutional kernel size to capture image detail information at different scales, forming a multi-scale feature vector set V:
[0081] x 1 ,x 2 ,…x n =Split(max(W 1 ,W 2 ,…W n ))
[0082] W 1 ,W 2 ,…W n =ECA(x)
[0083] ECA is a channel weight calculation method, and W 1 ,W 2 ,…W n are the weights of each channel of the primary feature map, that is, the weights of each channel obtained after the primary feature map passes through the ECA attention mechanism;
[0084] S13: Fuse the information flows of each branch, retain the most representative feature information of wheat imperfect grains, and obtain a more representative new feature map:
[0085] x'=fuse(x' 1 ,x' 2 ,…x' n )
[0086] Fuse is Figure 2 the concat operation in;
[0087] Then, further compact the features by introducing the Gelu activation function and BatchNormalization:
[0088] y=Gelu(Batch Normalization(x′)).
[0089] S2: Introduce deformable attention in the multi-branch model MBCWNet to adapt to significant information of different scales and shapes by predicting offsets and feature deformations;
[0090] S21: As Figure 3 shown, insert a deformable attention DAS module in each Block of the multi-branch model MBCWNet. This deformable attention DAS module dynamically adjusts the attention weights according to the feature importance of the local region;
[0091] S22: Through the deformable convolution mechanism, the deformable attention DAS module predicts the offset of each pixel in the feature map and geometrically deforms the feature map according to these offsets to capture the morphological features of imperfect grains;
[0092]
[0093] where def(q) represents the output result of the deformable convolution at position q; W i represents the weight of the convolutional kernel at position i; W q represents the weight of the offset position, which is used to adjust the weight of the upsampling position on the input feature map; q ref,k represents the fixed sampling position of the convolutional kernel relative to the reference position q; Δq i represents the learned offset in the deformable convolution, which is used to adjust the sampling point at position i of the convolutional kernel; q ref,k +Δq i represents the adjusted sampling position;
[0094] After the deformable attention DAS module, a Sigmoid activation function and layer normalization are applied to obtain the attention result:
[0095] M = Sigmoid(Layer Normalization(def(X c )))
[0096] where X c represents the feature map at position c, def(X c ) represents the output result of the deformable convolution at position c, and M represents the attention result;
[0097] Multiply the obtained attention result by the original input x:
[0098] x' = x ⊙ M
[0099] S23: By continuously adjusting and learning these deformation parameters, the model can adapt to wheat imperfect grains of different shapes and scales, improve the focusing ability on significant regions, and thus enhance the recognition accuracy.
[0100] S3: Construct a learnable adversarial training framework LAS-AWP based on adversarial weight perturbation, and use the learning of the policy network to generate attack strategies;
[0101] S31: Introduce LAS-AWP into the adversarial training framework, which generates targeted attack strategies by perturbing network weights. Under the learnable adversarial training framework, the generation process of the attack sample AE is as follows:
[0102]
[0103] Among them, e represents the clean sample, and e adv represents the generated attack sample AE, is the perturbation, a represents the new attack strategy generated by the policy network, t represents the parameters of MBCPNEXeXt-DAS, and f(·) represents the AWP attack;
[0104] S32: The policy network automatically generates the attack sample AE according to the current task stage and the characteristics of the sample. These attack samples AE are adaptively adjusted according to the weak points of the target model to achieve the maximum attack effect. Generating samples using the policy network depends on the policy, that is, p(a|e; s). Where s represents the parameters of the policy network.
[0105] S33: Through the generated attack sample AE, continuously adjust the model weights to find the optimal adversarial strategy of the MBCPNEXeXt-DAS model under different task conditions, thereby improving the robustness of the MBCPNEXeXt-DAS model against external attacks;
[0106] S4: Use the generated adversarial sample AE for adversarial training to optimize the robustness of the model to identify imperfect grains of wheat;
[0107] S41: Mix the generated adversarial sample AE with the clean sample for joint training. Through continuous confrontation and learning, strengthen the robustness of the MBCWNet-DAS model when facing complex tasks and attacks;
[0108] S42: During the adversarial training process, adjust the loss function of the MBCWNet-DAS model so that it can further improve its recognition accuracy of imperfect grains and reduce the misrecognition rate when facing the adversarial sample AE. In this process, parameter optimization of the policy network and the target network is involved. If the target network can correctly predict the label of the adversarial sample, it means that the attack strategy is effective. The loss function for evaluating the quality of the policy network can be expressed as:
[0109]
[0110] Where λ is the step size, e is the clean sample, is the perturbation, e adv represents the generated attack sample AE, is the newly generated adversarial sample, t represents the parameters of MBCWNet-DAS, is the new parameter of MBCWNet-DAS, f(·) represents a basic attack, is a basic loss function, y is the label of the clean sample, s represents the parameters of the policy network,
[0111]
[0112] Meanwhile, the performance of the new target model in predicting clean samples should also be considered. Another loss function for evaluating the attack strategy is defined as:
[0113]
[0114] Combining the above two loss functions, the total adversarial training loss function of LAS-AWP is:
[0115]
[0116] where α and β are used to balance and a represents the new attack strategy generated by the policy network, and p(a|e;s) represents the automatically generated sample-dependent strategy;
[0117] S43: Finally, the MBCWNet-DAS model after adversarial training not only has higher accuracy under normal conditions but also can maintain high robustness when facing adversarial attacks, and is suitable for the wheat imperfect grain detection task in practical industrial applications.
[0118] Step 5: Use the MBCWNet-DAS model after adversarial training for the wheat imperfect grain detection task in practical industrial applications.
[0119] In a second aspect, the present embodiment provides a wheat imperfect grain detection system based on local attention and learnable adversarial, which is characterized by including:
[0120] A construction module for constructing a wheat imperfect grain detection network; the wheat imperfect grain detection network includes: an MBCWNet-DAS model and a learnable adversarial training framework LAS-AWP;
[0121] The detection model MBCWNet-DAS is a multi-branch model MBCWNet with a parallel structure constructed based on the ConvNeXt network structure, and deformable attention is inserted into each Block of the multi-branch model MBCWNet;
[0122] The learnable adversarial training framework LAS-AWP includes a policy network and a sample generator;
[0123] The policy network generates a targeted attack strategy by perturbing the attention weights based on the clean samples and the attention weights backpropagated by the detection model MBCWNet-DAS;
[0124] The sample generator automatically generates an attack sample AE based on the attack strategy according to the characteristics of the current task stage and the clean samples;
[0125] An acquisition module for acquiring clean samples;
[0126] A training module for training the wheat imperfect grain detection network based on the clean samples and the generated adversarial samples AE to obtain a trained wheat imperfect grain detection network;
[0127] An identification module for detecting wheat imperfect grains through the trained wheat imperfect grain detection network in the wheat imperfect grain detection task in actual industrial applications.
[0128] It should be noted that the system embodiment of this embodiment is similar to the above method embodiment, so the description is relatively simple. For related parts, please refer to the above method embodiment.
[0129] In a third aspect, this embodiment of the present application also provides an electronic device, including:
[0130] One or more processors;
[0131] A memory for storing one or more programs,
[0132] When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the steps of the wheat imperfect grain detection method based on local attention and learnable adversarial as described in the embodiment.
[0133] This embodiment of the present application also provides a computer-readable storage medium, on which computer programs / instructions are stored. When the computer programs / instructions are executed by a processor, the steps of the wheat imperfect grain detection method based on local attention and learnable adversarial disclosed in this embodiment of the present application are implemented.
[0134] This embodiment of the present application also provides a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps in the wheat imperfect grain detection method based on local attention and learnable adversarial disclosed in this embodiment of the present application are implemented.
[0135] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0136] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, apparatus, or computer program product. Therefore, the embodiments of the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0137] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, systems, devices, storage media, and program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0138] The above has introduced in detail the methods, systems, devices, media, and products provided by the present application. Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for detecting imperfect wheat grains based on local attention and learnable adversarial learning, characterized in that: The steps include: Step 1: Build the detection model MBCWNet-DAS; The detection model MBCWNet-DAS is a multi-branch model MBCWNet with a parallel structure built based on the ConvNeXt network structure, and deformable attention is inserted into each Block of the multi-branch model MBCWNet; Step 2: Construct the learnable adversarial training framework LAS-AWP to generate adversarial samples AE; The learnable adversarial training framework LAS-AWP includes a policy network and a sample generator; The strategy network generates a targeted attack strategy by perturbing the attention weights based on the clean samples and the back-propagated attention weights of the detection model MBCWNet-DAS; The sample generator automatically generates attack samples AE based on the attack strategy, according to the current task stage and the characteristics of the clean samples; Step 3: Mix the generated adversarial sample AE with the clean sample and input it into the multi-branch model MBCWNet for joint training; Step 4: Use the adversarially trained MBCWNet-DAS model for the wheat imperfect grain detection task in actual industrial applications.
2. The method for detecting imperfect wheat grains based on local attention and learnable adversarial according to claim 1, characterized in that: The construction method of the detection model MBCWNet-DAS is: (1) First, we select the appropriate number of convolution branches and convolution kernel size, and build a parallel multi-branch model MBCWNet based on the ConvNeXt network structure for preliminary feature extraction to obtain the primary feature map R b×c×h×w : f1: x1→x'1∈R b×c×h×w <h2 style=";text-align:left;direction:ltr">f2:x2→x'2∈R<h2 style=";text-align:left;direction:ltr"> b×c×h×w …… f n :x n →x’ n ∈R b×c×h×w Among them, f1,f2,…f n They are all depthwise convolutions with kernels of different sizes; Then, the weights of each channel of the primary feature map are calculated and passed to each branch of the parallel structure in turn according to the weight size; each branch uses a different convolution kernel size to capture image details of different scales to form a multi-scale feature vector set V: x1,x2,…x n =Split(max(W1,W2,…W n )) W1,W2,…W n =ECA(x) ECA is a channel weight calculation method, W1, W2, ... W n is the weight of each channel of the primary feature map; Next, the information streams of each branch are merged to retain the most representative characteristic information of imperfect wheat grains: x′=fuse(x′1,x′2,…x′ n ) Fuse is a concat operation; Finally, the Gelu activation function and Batch Normalization are introduced to further compact the features: y=Gelu(Batch Normalization(x′)) (2) First, a deformable attention DAS module is inserted into each block of the multi-branch model MBCWNet; The deformable attention DAS module predicts the offset of each pixel in the feature map through a deformable convolution mechanism, and geometrically deforms the feature map according to these offsets to capture the imperfect particle morphology characteristics; Where def(q) represents the output of the deformable convolution at position q; W i Represents the weight of the convolution kernel at position i; W q Represents the weight of the offset position, which is used to adjust the weight of the sampling position on the input feature map; q ref,k represents the fixed sampling position of the convolution kernel relative to the reference position q; Δq i Represents the offset learned in the deformable convolution, which is used to adjust the sampling point of the convolution kernel position i; q ref,k +Δq i represents the adjusted sampling position; Then, after the deformable attention DAS module, a sigmoid activation function and layer normalization are applied to obtain the attention result: M=Siamoid(Layer Normalization(def(X c ))) Where X c represents the feature map at position c, def(X c ) represents the output result of the deformable convolution at position c, and M represents the attention result; Finally, the attention result is multiplied with the original input x: x′=x⊙M.
3. The method for detecting imperfect wheat grains based on local attention and learnable adversarial learning according to claim 2, characterized in that: The total loss function of the adversarial training of the learnable adversarial training framework LAS-AWP is: The loss function for evaluating the quality of the policy network is: The loss function for evaluating the attack strategy is: in, Alpha and beta are used for balance and a represents the new attack strategy generated by the policy network, p(a|e;s) represents the automatically generated sample-dependent strategy, λ is the step size, e is the clean sample, For disturbance, is the newly generated adversarial sample, t represents the parameters of MBCWNet-DAS, is the new parameter of MBCWNet-DAS, f(·) represents a basic attack, is a basic loss function, y is the label of the clean sample, s represents the parameters of the policy network, 4. A wheat imperfect grain detection system based on local attention and learnable adversarial learning, characterized in that: include: Building module, used to construct a wheat imperfect grain detection network; The wheat imperfect grain detection network includes: MBCWNet-DAS model and learnable adversarial training framework LAS-AWP; The detection model MBCWNet-DAS is a multi-branch model MBCWNet with a parallel structure built based on the ConvNeXt network structure, and deformable attention is inserted into each Block of the multi-branch model MBCWNet; The learnable adversarial training framework LAS-AWP includes a policy network and a sample generator; The strategy network generates a targeted attack strategy by perturbing the attention weights based on the clean samples and the back-propagated attention weights of the detection model MBCWNet-DAS; The sample generator automatically generates attack samples AE based on the attack strategy, according to the current task stage and the characteristics of the clean samples; Acquisition module, used to obtain clean samples; A training module, used for training the wheat imperfect grain detection network based on clean samples and generated adversarial samples AE to obtain a trained wheat imperfect grain detection network; The recognition module is used to detect imperfect wheat grains in the task of detecting imperfect wheat grains in actual industrial applications by using a trained imperfect wheat grain detection network.
5. The wheat imperfect grain detection system based on local attention and learnable adversarial according to claim 4, characterized in that: The construction method of the detection model MBCWNet-DAS is: (1) First, we select the appropriate number of convolution branches and convolution kernel size, and build a parallel multi-branch model MBCWNet based on the ConvNeXt network structure for preliminary feature extraction to obtain the primary feature map R b×c×h×w : f1: x1→x'1∈R b×c×h×w <h2 style=";text-align:left;direction:ltr">f2:x2→x'2∈R<h2 style=";text-align:left;direction:ltr"> b×c×h×w …… f n :x n →x’ n ∈R b×c×h×w Among them, f1,f2,…f n They are all depthwise convolutions with kernels of different sizes; Then, the weights of each channel of the primary feature map are calculated and passed to each branch of the parallel structure in turn according to the weight size; each branch uses a different convolution kernel size to capture image details of different scales to form a multi-scale feature vector set V: x1,x2,…x n =Split(max(W1,W2,…W n )) W1,W2,…W n =ECA(x) ECA is a channel weight calculation method, W1, W2, ... W n is the weight of each channel of the primary feature map; Next, the information streams of each branch are merged to retain the most representative characteristic information of imperfect wheat grains: x′=fuse(x′1,x′2,…x’ n ) Fuse is a concat operation; Finally, the Gelu activation function and Batch Normalization are introduced to further compact the features: y=Gelu(Batch Normalization(x′)) (2) First, a deformable attention DAS module is inserted into each block of the multi-branch model MBCWNet; The deformable attention DAS module predicts the offset of each pixel in the feature map through a deformable convolution mechanism, and geometrically deforms the feature map according to these offsets to capture the imperfect particle morphology characteristics; Where def(q) represents the output of the deformable convolution at position q; W i Represents the weight of the convolution kernel at position i; W q Represents the weight of the offset position, which is used to adjust the weight of the sampling position on the input feature map; w ref,k represents the fixed sampling position of the convolution kernel relative to the reference position q; Δq i Represents the offset learned in the deformable convolution, which is used to adjust the sampling point of the convolution kernel position i; q ref,k +Δq i represents the adjusted sampling position; Then, after the deformable attention DAS module, a sigmoid activation function and layer normalization are applied to obtain the attention result: M=Siamoid(Layer Normalization(def(X c ))) Where X c represents the feature map at position c, def(X c ) represents the output result of the deformable convolution at position c, and M represents the attention result; Finally, the attention result is multiplied with the original input x: x′=x⊙M.
6. The wheat imperfect grain detection system based on local attention and learnable adversarial according to claim 5, characterized in that: The total loss function of the adversarial training of the learnable adversarial training framework LAS-AWP is: The loss function for evaluating the quality of the policy network is: The loss function for evaluating the attack strategy is: in, Alpha and beta are used for balance and a represents the new attack strategy generated by the policy network, p(a|e;s) represents the automatically generated sample-dependent strategy, λ is the step size, e is the clean sample, For disturbance, is the newly generated adversarial sample, t represents the parameters of MBCWNet-DAS, is the new parameter of MBCWNet-DAS, f(·) represents a basic attack, is a basic loss function, y is the label of the clean sample, s represents the parameters of the policy network, 7. An electronic device, characterized in that: include: A memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the wheat imperfect grain detection method based on local attention and learnable adversarial according to any one of claims 1-3.
8. A computer-readable storage medium, characterized in that: It stores a computer program, which, when executed by a processor, implements the wheat imperfect grain detection method based on local attention and learnable adversarial as described in any one of claims 1 to 3.
9. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the method for detecting imperfect wheat grains based on local attention and learnable adversarial methods as described in any one of claims 1 to 3 is implemented.