Wheat variety identification method based on VIT and spatial attention mechanism

The wheat variety identification method based on VIT and spatial attention mechanism solves the problem of intra-species variety identification, realizes the learning of fine-grained feature differences of wheat varieties, and improves the accuracy and efficiency of variety identification, adaptability and production efficiency.

CN120656059APending Publication Date: 2025-09-16SUZHOU VOCATIONAL UNIVERSITY (SUZHOU OPEN UNIVERSITY)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510740454.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In existing technologies, wheat variety identification mainly focuses on inter-species classification, and lacks effective learning and identification of fine-grained characteristic differences between different varieties within a species. It is difficult to quickly and reliably identify excellent varieties that adapt to specific environments during the critical stage of wheat growth.

Method used

A wheat variety identification method based on VIT and spatial attention mechanism is adopted to quantify the subtle differences of wheat varieties through image acquisition, WMVD dataset construction, ViT+SPA model framework design and feature extraction, combined with T-SNE visualization and texture trait extraction.

Benefits of technology

The accuracy and efficiency of wheat variety identification have been improved, and it is possible to identify varieties with strong adaptability, high quality and high yield at different growth stages, providing scientific support and laying the foundation for genetic breeding and agricultural cultivation production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656059A_ABST
    Figure CN120656059A_ABST
Patent Text Reader

Abstract

The invention discloses a wheat variety identification method based on VIT and a spatial attention mechanism. The wheat variety identification method comprises the following steps: S1, image acquisition: acquiring a wheat canopy image and a side plant type image by using camera equipment; s2, constructing a wheat variety data set: constructing a WMVD data set, setting a window with a fixed size on an original image, and gradually sliding to generate a plurality of sub-images; s3, designing a wheat variety identification and selection model framework and constructing a feature extraction network; s4, high-dimensional feature visualization based on T-SNE: carrying out dimension reduction on the high-dimensional features extracted by the ViT + SPA model and realizing feature visualization; and S5, texture character extraction based on the wheat feature map. According to the method, the feature extraction capability of the wheat canopy image and the side plant type image is enhanced, the ViT + SPA model is constructed to pay attention to important areas in the images, and the recognition precision of the deep learning model on different wheat varieties is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wheat identification, and specifically relates to a wheat variety identification method based on the combination of VIT and spatial attention mechanism. Background Art

[0002] Variety selection across diverse wheat varieties under multiple habitats facilitates the rapid selection of wheat varieties with superior agronomic traits adapted to specific environmental conditions in the context of global climate change. This can improve wheat yield, quality, and stress resistance, thereby ensuring market demand and food security (Rempelos et al., 2020). Because different wheat varieties exhibit distinct morphological and physiological differences during key growth stages (particularly from booting to flowering), the wheat flowering prediction research presented in Chapter 3 enables faster and more reliable selection of superior varieties at these critical stages of wheat growth and development, providing a broader genetic basis and quantitative evidence for breeding wheat varieties with greater climate adaptability (Zhao et al., 2023). Furthermore, variety selection based on biochemical characteristics of different wheat varieties (such as gluten quality, starch content, and kernel hardness) can help maintain the diversity of wheat genetic resources, enhance the stability and risk resilience of agricultural ecosystems, and achieve standardized production and consistent quality (Jamiel et al., 2019). Furthermore, developing agronomic management measures tailored to different varieties at different growth stages (for example, irrigation and fertilization plans tailored to their water and fertilizer needs) can reduce production risks and improve efficiency (Bijay-SinghandCraswell, 2021). Therefore, scientific variety selection not only facilitates genetic breeding research and agricultural cultivation production but also lays the foundation for future variety improvement and agricultural technology innovation.

[0003] Currently, most wheat variety identification research focuses on classification between species, with relatively little research on identifying varieties within a species. Due to the limited morphological differences between varieties within a species, existing variety identification tasks are typically limited to the identification of a few representative varieties. Given the distinct morphological and physiological differences between wheat varieties at different critical growth stages, leveraging deep learning techniques to learn and identify multiple wheat varieties, in particular, is a challenging and significant research area.

[0004] The information disclosed in this background section is only intended to enhance understanding of the overall background of the invention and should not be considered as an admission or any form of suggestion that the information constitutes the prior art already known to a person skilled in the art. Summary of the Invention

[0005] The purpose of the present invention is to provide a wheat variety identification method based on VIT and spatial attention mechanism, which can solve the above problems.

[0006] In order to achieve the above object, a specific embodiment of the present invention provides the following technical solutions: A wheat variety identification method based on VIT and spatial attention mechanism includes the following steps: S1. Image acquisition: Use camera equipment to capture images of wheat canopy and side plant types; S2. Construction of wheat variety dataset: Construct the WMVD dataset by setting a fixed-size window on the original image and sliding it step by step to generate multiple sub-images; S3. Design of wheat variety identification model framework and construction of feature extraction network: S31 and ViT use the self-attention mechanism of the Transformer layer to extract global features; S32, add the spatial pyramid attention module to enhance the feature extraction capability of the basic network; S4. High-dimensional feature visualization based on T-SNE: Dimensionality reduction of high-dimensional features extracted by the ViT+SPA model and visualization of features: S5. Texture trait extraction based on wheat feature map: The gray-level co-occurrence matrix is ​​used to extract the areas of interest in the model in the wheat feature map to quantify the subtle differences in two-dimensional morphology and structure of different varieties.

[0007] In one or more embodiments of the present invention, when collecting wheat canopy images, the camera equipment is parallel to the ground and maintains a distance of 1.8m from the canopy. When collecting wheat side plant type images, the shooting rod of the camera equipment is 1m away from the cell, and the camera of the camera equipment is maintained at 45° to the cell to ensure that the entire side of the cell is covered.

[0008] In one or more embodiments of the present invention, the WMVD dataset includes awned type, unawned type, dense canopy type, and sparse canopy type.

[0009] In one or more embodiments of the present invention, the specific steps of S31 are: S311, first divide the input image into fixed-size image blocks through ViT, then flatten each image block and linearly map it to a high-dimensional space. The input image is , divide it into N image blocks, the size of each image block is , then each image block is flattened into a vector and mapped to a high-dimensional space through linear transformation, where Respectively represent the length, width, and number of channels of the image; The specific formula is as follows:

[0010] S312, retaining position information of different image blocks through position coding; Among them, the position encoding representation of each image block is shown by the following formula

[0011] in, and are the weight and bias of the linear transformation, Represents the code of the i-th position.

[0012] S313: Before entering the Transformer layer, a classification label is added to the input sequence. The input sequence is extracted from the wheat image features through the multi-head attention mechanism and feedforward neural network in multiple Transformer layers. The specific extraction formula for extracting image features is as follows:

[0013]

[0014] Where, Each patch block represented by is linearly transformed by the formula of step S311 to obtain the input matrix, Represents the parameter matrix to be learned; is the attention function, represents the i-th attention head, represents the parameter matrix to be learned for the i-th attention head, The dimension of the formula key, is a scaling factor used to prevent certain areas from having too much attention weight; MSA stands for Multi-Head Attention. Represents the linear change matrix of multi-head attention output, represents the output of the l-th layer of multi-head attention, LN represents the layer normalization, is the final output of layer l.

[0015] In one or more embodiments of the present invention, the specific steps of S32 are: S321, using the output of the ViT layer as input, passes through two adaptive average pooling layers to automatically adjust the pooling window and stride; S322, resize the two outputs into two-dimensional vectors and generate a 1D attention map through the Concat function; S322. Attention weights are obtained through calculations of a series of FC layers and BN layers.

[0016] In one or more embodiments of the present invention, the specific step of S4 is: S41. Similarity modeling in high-dimensional space. t-SNE uses the Gaussian kernel function to calculate the pairwise similarity between all data points in high-dimensional space. For a given pair of high-dimensional data points, , whose similarity measure is expressed as conditional probability Indicates that in a given In the case of The probability of being selected; data points that are closer are given a higher probability of selection, while data points that are farther away are given a lower probability of selection, so that all points are selected. Symmetrically transform to obtain the joint probability distribution ; S42. Similarity modeling in low-dimensional space, Calculate the pairwise similarity between data points in low-dimensional space through t distribution; S43, By minimizing the Kullback-Leibler divergence between the high-dimensional probability distribution P and the low-dimensional probability distribution Q; t-SNE iteratively optimizes the objective function through gradient descent:

[0017] In one or more embodiments of the present invention, the specific steps of S5 are: GLCM generates a matrix describing the spatial relationship of gray levels by counting the co-occurrence frequency of gray levels at specific distances and directions, and multiple texture traits are calculated based on GLCM.

[0018] In one or more embodiments of the present invention, the specific distance in step S5 is 1 pixel and the direction is 45°.

[0019] In one or more embodiments of the present invention, the plurality of texture traits include entropy, homogeneity, angular second moment, variance, contrast, variability, mean, and standard deviation.

[0020] In one or more embodiments of the present invention, The specific calculation formula of the texture properties is as follows: Assume that p(i,j) represents the element in the i-th row and j-th column of the gray-level co-occurrence matrix, N is the total number of gray levels (N=255), is the mean of the GLCM.

[0021] The formula for calculating entropy is:

[0022] The formula for calculating homogeneity is:

[0023] The calculation formula for the angular second moment is:

[0024] The formula for calculating the mean is:

[0025] The formula for calculating variance is:

[0026] The formula for calculating contrast is:

[0027] The formula for calculating the difference is:

[0028] The formula for calculating the standard deviation is:

[0029] Compared with the existing technology, the wheat variety identification method based on VIT and spatial attention mechanism of the present invention has the following advantages: ViT demonstrates significant advantages in extracting fine-grained image features. By segmenting the image into fixed-size blocks and treating each block as a sequence input to the transformer, ViT learns global image features. ViT excels at handling long-range dependencies and capturing complex patterns, making it particularly suitable for detecting subtle differences. ViT is therefore suitable for identifying multiple crop varieties.

[0030] The feature extraction capability of wheat canopy images and side plant shape images has been enhanced, and a ViT+SPA model has been constructed to focus on important areas in the image (such as canopy density, ear shape, and leaf stem features), thereby improving the recognition accuracy of the deep learning model for different wheat varieties.

[0031] Texture descriptions were used to extract features such as entropy, contrast, and differences from the characteristic regions learned by the deep learning model. This was then used to cluster different wheat varieties, providing an important quantitative basis for further understanding variety classification. Finally, the ViT+SPA model was used to classify wheat organ-level features (awned, awnless, sparse canopy, and dense canopy), providing scientific support for the subsequent large-scale selection of wheat varieties with strong adaptability, high quality, high yield, and good stress resistance. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0033] Figure 1 The China Field Trial Base and the field distribution of 42 tested wheat varieties (2023-24 wheat growing season); Figure 2 This is a wheat image acquired using a smartphone; Figure 3 Crop images for sliding window based images; Figure 4 Wheat dataset images for different organ types; Figure 5 This is a schematic diagram of the research framework for wheat variety selection; Figure 6 It is a feature extraction network structure diagram; Figure 7 Schematic diagram of the multi-head attention mechanism; Figure 8 Schematic diagram of the spatial pyramid attention module structure; Figure 9 This is a diagram showing the accuracy and loss of the Chinese wheat variety recognition model; Figure 10 This is a schematic diagram comparing the variety feature maps extracted by different models; Figure 11 Schematic diagram of texture feature extraction; Figure 12 Schematic diagram of texture feature extraction; Figure 13 Schematic diagram of correlation analysis of texture properties of training set and test set; Figure 14 Schematic diagram of T-SNE visualization of high-dimensional features of different models; Figure 15 Schematic diagram of comprehensive performance comparison between different models. DETAILED DESCRIPTION

[0034] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0035] like Figures 1-8 As shown, a wheat variety identification method based on VIT and spatial attention mechanism in one embodiment of the present invention includes the following steps: Field trial materials and experimental design: The China Wheat Experimental Field is located at the Baima Experimental Base of Nanjing Agricultural University (31°61'N, 119°18'). The experiment will be conducted from 2023 to 2024. Figure 4 The experimental field was divided into 126 plots, covering 42 wheat varieties representative of the middle and lower reaches of the Yangtze River, with three replicates for each variety. S1. Image acquisition: Camera equipment is used to collect images of wheat canopy and side plant types; like Figure 2 As shown in the figure, a. is the acquisition of multi-view image data, b. is the canopy image, and c. is the side plant type image. Specifically, in the process of wheat variety selection, since the canopy image and the side image can provide multi-level and multi-angle phenotypic information (for example, the canopy image of wheat during the flowering period can reflect the flag leaf morphology, canopy structure, density, color and spike type differences of different wheat varieties), the side plant type image can reflect the stem morphology (thickness, uprightness), plant tillering layer degree, spike distribution, etc. (Shavrukov et al., 2017). In order to obtain clearer images of wheat plant type and canopy type from the jointing stage to the flowering stage, a smartphone is mounted on a handheld adjustable imaging pole to collect wheat canopy images and side plant type images (as shown in the figure). Figure 2 -a). When collecting wheat canopy images, the mobile phone camera was parallel to the ground and kept 1.8m away from the canopy (GSD = 0.04cm·pixel-1). The mobile phone was controlled by a Bluetooth module to shoot. The maximum resolution of each image was 9248×6944 pixels ( Figure 4 When collecting wheat plant type images from the side, the shooting pole should be 1m away from the plot, and the mobile phone camera should be kept at a 45° angle to the plot to ensure that the entire side of the plot is covered (e.g. Figure 2 -c).

[0036] S2. Construction of wheat variety dataset: Construct a WMVD dataset by setting a fixed-size window on the original image and sliding it step by step to generate multiple sub-images; the WMVD dataset includes awned type, no awn type, dense canopy type, and sparse canopy type. Specifically, such as Figure 3 As shown in Figure 1 (a is a cropped canopy image; b is a cropped side image), a sliding window-based image cropping method was used when constructing the wheat variety dataset (WMVD). By setting a fixed-size window on the original image and sliding it step by step to generate multiple sub-images, the detailed information of the canopy or side plant type is retained while reducing the interference of the global background. The core of the sliding window method is to select a suitable window size, which affects the accuracy of image cropping and the retention of details. The following factors were comprehensively considered: (1) image resolution, that is, the adaptation of the resolution to the sliding window to ensure that sufficient image area is covered; (2) canopy and side plant type characteristics, that is, the window size needs to ensure the acquisition of detailed features such as wheat leaf texture, distribution density, and stem morphology. After multiple experiments and adjustments, the window size was finally selected as 2048×2048 pixels, which is one-third of the original image width.

[0037] The step size setting also affects the efficiency and coverage of image cropping. A smaller step size can generate more sub-images, covering more detailed areas, but it increases processing time and storage requirements; a larger step size can improve efficiency, but may miss some detailed information. To ensure that the sliding window can continuously cover the entire image area without leaving obvious gaps, the step size is set to 64 pixels, that is, the window slides 64 pixels horizontally or vertically. This setting can improve computational efficiency and reduce repeated calculations while ensuring continuous coverage of the image area. To ensure model training, the cropped sub-images are also expanded using image enhancement methods.

[0038] In the process of wheat variety selection, the classification and induction of organ-level (ear morphology, canopy structure) images are of great significance. Organ-level characteristics are subject to genetic regulation (such as awn length, ear length and canopy density, etc.), so they are relatively stable under different growth conditions. These characteristics can generally be applied to the standardized identification of wheat varieties. In addition, organ characteristics are generally suitable for field observation and recording. Different organ characteristics can reflect the adaptability of wheat varieties to different environmental conditions, thereby improving the efficiency of variety selection. For example, the ear morphology is correlated with the variety's tolerance to drought or waterlogging conditions, and the density of the canopy structure can show resistance to diseases and pests. Based on the multi-variety dataset obtained in step S3, the WMVD dataset was manually divided into four major categories (such as awn type, no awn type, dense canopy type, and sparse canopy type) with the combined objectives. Figure 4 This data provides a rich training dataset for the subsequent construction of wheat variety selection models.

[0039] S3. Design of wheat variety identification model framework and construction of feature extraction network: Wheat variety selection research mainly includes two tasks: variety identification and classification of wheat organ-level characteristics (i.e., with awns, without awns, dense canopy, sparse canopy). Figure 5 The overall framework of the model is shown in Figure 1. Subfigure a illustrates the method for variety recognition based on canopy and side profile plant images. First, a multi-layer Transformer is used to extract features between different wheat varieties. Next, an attention module is used to weight these high-dimensional features. The attention module uses weighted weights to highlight key morphological and structural features of different wheat varieties, reducing the impact of noise on the model and thus improving variety recognition accuracy.

[0040] Secondly, feature map visualization can demonstrate the model's attention distribution, reflecting the key areas the model focuses on when processing different images. An effective feature map should have higher weights in key areas, indicating that the model captures the differences between varieties. Using fully connected modules and the SoftMax function, multi-variety classification can be achieved. To verify the applicability of the Transformer- and attention-based approach, a comparison was also conducted with several traditional feature extraction deep learning models, including VGG-16, ResNet-101, and Inception-v3.

[0041] Sub-figure b shows the process of wheat organ-level feature classification: First, based on the ear image data, we combine the Transformer and spatial attention modules to achieve the classification of wheat ears with and without awn features. In the canopy density classification task, a deep learning model is established based on the canopy image to achieve the classification of canopy density and sparseness.

[0042] S31 and ViT use the self-attention mechanism of the Transformer layer to extract global features; In the wheat multi-variety recognition task, ViT uses the self-attention mechanism of the Transformer layer to extract global features (such as Figure 6-b), can capture subtle phenotypic differences between wheat varieties. Compared to traditional convolutional neural networks (such as ResNet101 and VGG16), this mechanism allows the model to focus on global information when processing the input image, without the need to accumulate local features layer by layer. This effectively addresses the limitations of traditional convolutional neural networks in dealing with long-range dependencies and complex morphologies, and makes up for the shortcomings of convolutional neural networks in change analysis and accumulated feature learning. In addition, in convolutional networks, due to the fixed size of the convolution kernel and the same feature extraction ability for the input image, it is difficult to dynamically adjust the size of the focus area. ViT, on the other hand, can dynamically adjust the weights according to the input image through the self-attention mechanism, focusing on important feature areas in the wheat variety image.

[0043] S311. First, the input image is divided into fixed-size image blocks through ViT. Each image block is then flattened and linearly mapped into a high-dimensional space. The input image is (representing the length, width, and number of channels of the image, respectively). It is divided into N image blocks, each of which is of size . Then, each image block is flattened into a vector and mapped to the high-dimensional space through linear transformation. The specific formula is as follows:

[0044] S312. Position encoding preserves the position information of different image blocks, ensuring that when the model processes wheat variety identification, it can not only focus on the overall morphological characteristics but also accurately locate local detail changes.

[0045] Among them, the position encoding representation of each image block is shown by the following formula

[0046]

[0047] in, and are the weight and bias of the linear transformation, Represents the code of the i-th position.

[0048] S313: Before entering the Transformer layer, a classification label is added to the input sequence. Then, the input sequence is extracted from the wheat image features through the multi-head attention mechanism and feedforward neural network in multiple Transformer layers. The specific extraction formula for extracting image features is as follows:

[0049] Where, Each patch block represented by is linearly transformed by the formula of step S311 to obtain the input matrix, Represents the parameter matrix to be learned; is the attention function, represents the i-th attention head, represents the parameter matrix to be learned for the i-th attention head, The dimension of the formula key, Scaling factor, used to prevent certain areas from having too much attention weight; MSA stands for Multi-Head Attention. Represents the linear change matrix of multi-head attention output, represents the output of the l-th layer of multi-head attention, LN represents the layer normalization, is the final output of layer l.

[0050] Furthermore, the multi-head attention mechanism, through multiple independent attention heads, can extract features from different subspaces. Each attention head can focus on different parts of the input image, thereby capturing richer and more diverse feature information. This is particularly important for extracting subtle feature differences in wheat variety recognition, helping the model better distinguish between different varieties. Compared to a single attention mechanism, the multi-head attention mechanism, through multiple independent attention heads, can effectively reduce information bottlenecks. Each attention head focuses on different information channels, ensuring that more feature information is retained and utilized, which helps improve the performance of wheat variety recognition.

[0051] S32. Adding a spatial pyramid attention module to enhance the feature extraction capability of the basic network Although the ViT model performs well in capturing global features, it has some shortcomings in extracting local detail features. Figure 8 As shown in Figure 5, the feature extraction capability of the base network is enhanced by adding a spatial pyramid attention module, which aims to guide the model to focus on important areas in the feature map and ignore irrelevant parts.

[0052] S321: This module takes the output of the ViT layer as input and passes it through two adaptive average pooling layers to automatically adjust the pooling window and stride. The 4×4 scale Adaptive Average Pooling (AAP) captures more spatial and structural information, while the 2×2 scale average pooling strikes a balance between structural information and structural regularization.

[0053] S322. Resize the two outputs into two-dimensional vectors and generate a 1D attention map through the Concat function.

[0054] S322. Attention weights are calculated through a series of FC and BN layers. SPA is used in conjunction with the multi-head self-attention mechanism of the ViT model, combining global and local features to enhance the model's ability to learn image details, thereby improving variety classification performance.

[0055] S4. High-dimensional feature visualization based on T-SNE: We use t-distributed stochastic neighbor embedding (t-SNE) to reduce the dimensionality of the high-dimensional features extracted by the ViT+SPA model and visualize the features. t-SNE is a nonlinear dimensionality reduction method that is particularly suitable for mapping high-dimensional data into a low-dimensional space.

[0056] S41. Similarity modeling in high-dimensional space. t-SNE uses the Gaussian kernel function to calculate the pairwise similarity between all data points in high-dimensional space. For a given pair of high-dimensional data points, , whose similarity measure is expressed as conditional probability , indicating that in a given In the case of The probability of being selected; data points that are closer are given a higher probability of selection, while data points that are farther away are given a lower probability of selection, so that all points are selected. Symmetrically transform to obtain the joint probability distribution ; S42, Similarity modeling in low-dimensional space, t-SNE calculates the pairwise similarity between data points in low-dimensional space through t distribution. The similarity measure is expressed as the joint probability ,The t distribution is used to better deal with the crowding phenomenon of data, so that the point pairs that are far apart in the low-dimensional space have non-zero probability; S43. Optimization for embedding. To map high-dimensional data to a low-dimensional space while maintaining the structure of pairwise similarity, t-SNE achieves this goal by minimizing the Kullback-Leibler divergence between the high-dimensional probability distribution P and the low-dimensional probability distribution Q. t-SNE iteratively optimizes the objective function through gradient descent:

[0057] This optimization process continues until the low-dimensional embedding reaches a stable state. The final low-dimensional embedding not only maintains the local structure of the original high-dimensional data, but also can display clusters and subclusters of data points in two-dimensional or three-dimensional space, making it easier to understand the structure and relationships in high-dimensional data. In the specific implementation process, the t-SNE implementation in the Scikit-learn library is used to reduce the high-dimensional features (64 dimensions) extracted by the Vit+SPA model to 2 dimensions. The learning rate is set to 200 and the number of iterations is 1000 to ensure the stability and reliability of the results. The reduced data is visualized using the Matplotlib library. In two-dimensional space, data points of different wheat varieties are marked with different colors or shapes, which can intuitively show the differences between varieties.

[0058] S5. Texture trait extraction based on wheat feature map: The gray-level co-occurrence matrix is ​​used to extract the areas of interest in the model from the wheat feature map to quantify the subtle differences in two-dimensional morphology and structure of different varieties (such as leaf surface roughness, vein distribution, wheat ear morphology, etc.).

[0059] GLCM generates a matrix describing the spatial relationship of gray levels by counting the co-occurrence frequency of gray levels at a specific distance (usually 1 pixel) and direction (45°). Based on GLCM, a total of eight texture traits are calculated, including: Entropy, Homogeneity, Angular Second Moment (ASM), Variance, Contrast, Dissimilarity, Mean, and Standard Deviation.

[0060] The specific calculation formula of texture properties is as follows; Assumptions represents the element in the i-th row and j-th column of the gray-level co-occurrence matrix, N is the total number of gray levels (N=255), is the mean of the GLCM.

[0061] The formula for calculating entropy is:

[0062] The entropy value reflects the complexity of the image: the higher the entropy value, the more complex the grayscale distribution of the image and the greater the amount of information. For the wheat feature map, a high entropy value corresponds to complex leaf texture details.

[0063] The formula for calculating homogeneity is:

[0064] Homogeneity measures the uniformity of the grayscale distribution in an image: higher homogeneity indicates that the grayscale levels of adjacent pixels in the image are closer. High homogeneity indicates a smooth image texture, reflecting the uniformity of the leaf surface.

[0065] The calculation formula for the angular second moment is:

[0066] ASM reflects the texture consistency of an image: higher ASM values ​​indicate more consistent texture in the image. For a wheat image, a high ASM value might indicate consistent leaf arrangement and a uniform texture structure.

[0067] The formula for calculating the mean is:

[0068] The mean represents the average grayscale value of an image: it reflects the overall brightness level of the image. For wheat images, the mean provides brightness information for the leaves and background, helping to distinguish different structural features.

[0069] The formula for calculating variance is:

[0070] Variance measures the dispersion of the grayscale values ​​in an image: a larger variance indicates a wider distribution of grayscale values. High variance may reflect texture variations and the diversity of leaf structures in the image.

[0071] The formula for calculating contrast is:

[0072] Contrast measures the local variations between gray levels: higher contrast indicates greater gray level variations in the image and more visible details. For the wheat image, high contrast indicates greater clarity of leaf veins, edges, and texture.

[0073] The formula for calculating the difference is:

[0074] The dissimilarity reflects the degree of difference in the grayscale of the image: the greater the dissimilarity, the more significant the difference in grayscale between adjacent pixels. High dissimilarity corresponds to changes in leaf texture and edge features.

[0075] The formula for calculating the standard deviation is:

[0076] The standard deviation measures the degree of fluctuation in grayscale values: a larger standard deviation indicates more dramatic changes in grayscale values. A high standard deviation indicates significant texture detail and variation in the image.

[0077] Experimental Example 1 1.1. Construction of intelligent identification model for Chinese wheat varieties like Figure 9 Figure 2 shows the evolution of the model's accuracy and loss on the training and validation sets during training on a dataset of Chinese wheat varieties. At the start of training, the model's accuracy reached 50%. As the number of training rounds increased, the model continuously adjusted its parameters through gradient descent and backpropagation algorithms, gradually adapting to the new dataset. Throughout training, the accuracy showed a gradual upward trend, while the loss gradually decreased. After 70 rounds of training, the model achieved a good fit, with an accuracy of 96.9% on the validation set and a loss of 0.14. In the test set, the model achieved a classification accuracy of 93.8% for 20 Chinese varieties. This result demonstrates that through the pre-training and retraining approach, the model not only effectively utilizes existing feature representations but also demonstrates excellent performance on new classification tasks.

[0078] 1.2. Model performance in extracting features from multiple wheat varieties like Figure 10 Figure 2 shows the model's feature extraction and visualization results for multiple wheat varieties. Taking Yangfumai No. 6, Baoji 0601, Xumai 32, and Yangmai No. 16 as examples, the first and third rows list the original wheat images, while the second and fourth rows show the model's extracted feature heatmaps. The color variations in the feature heatmaps reflect the model's attention to features in different regions. Red and yellow areas indicate regions of high model attention, typically with significant characteristics or differences, while blue and green areas indicate regions of lesser model attention. These heatmaps demonstrate the model's ability to effectively extract salient features across varieties, such as leaf shape, ear structure, and texture. Specifically, for these four varieties, the model displays a high degree of attention to the edges of the ear and leaves. The concentrated color variations in these regions indicate that the model accurately identifies and focuses on discriminative features, enabling effective differentiation between wheat varieties. Furthermore, the feature heatmaps clearly show that soil and other noisy areas typically appear blue and green, indicating lower model attention. This demonstrates that the model effectively filters out irrelevant information during feature extraction, focusing on key features relevant to wheat variety classification. This capability not only improves the model's classification accuracy but also reduces the impact of background noise on the model's decision-making process, further validating the model's robustness and adaptability in handling complex environments.

[0079] 1.3 Texture Characteristic Analysis Based on Feature Map The texture traits of the feature maps extracted based on the model can reflect the fine-grained differences in the characteristics of different wheat varieties, such as Figure 11-12As shown in the figure (test dataset, n=1000), eight box plots illustrate the distribution of texture traits of feature maps for different wheat varieties in the test set. Fifty images were randomly selected from each variety in the test set. Based on the feature maps extracted by the model, eight texture traits were calculated: entropy, homogeneity, ASM, variance, contrast, dissimilarity, mean, and standard deviation. This systematic analysis of these texture features reveals differences in textural properties among different wheat varieties. In the entropy box plot, higher entropy values ​​indicate greater image complexity, indicating richer and more diverse textural features. For example, the entropy values ​​for Yangmai 25 (YM25) are most widely distributed between the lower and upper quartiles (25%-75%) (0.3-0.84), while the entropy values ​​for Zhenmai 8 (ZM8) are more tightly distributed between the upper and lower quartiles (0.4-0.64). The highest median entropy value was for Yangmai 20 (YM20), reaching 0.72, while the lowest median was for Wanyu 2 (WY2), at 0.33. This indicates that entropy values ​​vary among different varieties. Homogeneity reflects the similarity of image textures, and a box plot shows the distribution of homogeneity across wheat varieties. Wanyu 2 has a high homogeneity, with the 25%-75% data set ranging from 0.38 to 0.76, indicating a more uniform texture. Yangmai 16 has a low homogeneity, with the 25%-75% data set ranging from 0.31 to 0.51. The angular second moment (ASM) feature measures the uniformity of the canopy structure in the texture image. Different varieties have obvious differences in the performance of canopy structure characteristics. The ASM median of Siskin and Paragon varieties is the largest at 0.7, indicating that their canopy structures are relatively uniform. The ASM medians of Yangfumai No. 8 (0.39) and Sukomai No. 1 (SKM1) are lower, indicating that the canopy structures of these varieties are quite different.

[0080] In order to verify the reliability of the model in extracting texture traits on the Chinese wheat variety test set, the correlation between the texture features of the feature maps of the training set and the test set was analyzed. Specifically, for each wheat variety, 50 feature maps were randomly selected from the training set and the test set, and different texture traits were calculated. The correlation analysis was performed using Origin software (e.g. Figure 13 ). For the specific texture trait extraction results of the training set feature map, please refer to Appendix F. The analysis results show that the similarity of different texture traits between the test set and the training set is high. Among them, the contrast and variance traits have the highest similarity, reaching 0.87 and 0.884 respectively. The similarity of other texture traits is also quite significant, including the angular second moment (ASM, 0.803), homogeneity (0.81), entropy (0.821), mean (0.831), dissimilarity (0.805) and standard deviation (0.771). These results show that the improved model can stably extract texture features between the training set and the test set and has high reliability.

[0081] 1.4 Wheat organ-level feature classification results The results of the study based on the ViT+SPA model showed that the accuracy of the model in the classification of Chinese wheat spike type (with awns, without awns) reached 92.3%, and the accuracy in the classification of canopy density (tight canopy, sparse canopy) reached 91.8%. Figure 14 Figure 2 shows the distribution of image features for two classification tasks. In the spike type classification task, green scatter points represent 912 samples of wheat varieties without awns, while purple scatter points represent 908 samples of wheat varieties with awns. The T-SNE dimensionality reduction results clearly separate the two types of samples, with a clear boundary between awned and awnless wheat types, demonstrating the superiority of the ViT+SPA model in capturing wheat spike type characteristics. In the wheat canopy compactness classification task, red scatter points represent 921 samples of wheat with compact canopies, while blue scatter points represent 928 samples of wheat with sparse canopies. Similarly, these samples show a clear clustering effect in the two-dimensional space after dimensionality reduction, verifying the reliability of the model in extracting canopy compactness features.

[0082] like Figure 15 As shown in the figure, the accuracy of the six models in identifying 13 wheat varieties is shown. It can be seen from the figure that the prediction accuracy of the ViT+SPA model on all varieties is higher than that of the other five models, showing a clear advantage.

[0083] Specifically, the ViT+SPA model achieved the highest prediction accuracy for the Lili and Robigs varieties, reaching 94.8% and 94.6%, respectively. The model's prediction accuracy for the Kerrin and Skyfall varieties was lower, at 90.1% and 91.1%, respectively. The ViT+SPA model's recognition accuracy for 13 wheat varieties ranged from 90.1% to 94.8%, with a fluctuation of only 4.7%, demonstrating its good prediction stability across varieties. Other models not only had lower accuracy but also exhibited significant fluctuations across varieties. For example, the VGG-16 and ResNet101 models exhibited accuracy fluctuations of 9.8% and 7.1% across varieties, respectively, making it difficult to provide standardized predictions. These conclusions demonstrate the superiority and stability of the ViT+SPA model in the wheat variety identification task.

[0084] A wheat field trial was conducted in China between 2022 and 2024. During the flowering period, smartphones were used to capture multi-scale images of wheat plants, capturing canopy and plant profiles from various angles. Using a sliding window technique, the original images were cropped and preprocessed to obtain more refined local features. Based on this, a wheat multi-variety dataset (WMVD) was constructed, encompassing 20 commonly used varieties in the middle and lower reaches of the Yangtze River in China and containing a total of 102,130 high-resolution images. These images were collected under diverse climatic and geographical conditions, ensuring data diversity and representativeness. Furthermore, a training set (71,491), validation set (20,426), and test set (10,212) were divided in a 7:2:1 ratio, providing the data foundation for the construction of the wheat variety recognition model ViT+SPA-Wheat.

[0085] The dataset covers different wheat morphological characteristics, including awned and awnless varieties, dense and sparse canopies. By analyzing these wheat morphological characteristics, researchers can reveal the physiological characteristics of different wheat varieties. The construction of the WMVD dataset not only provides rich data support for wheat variety identification and phenotypic analysis, but also lays an important foundation for intelligent agricultural management, pest and disease detection and control, and interdisciplinary research. It has significant value and application prospects.

[0086] In the wheat variety identification task, the ViT model, through the integration of self-attention mechanism and global features, can better cope with the subtle differences and complex backgrounds between different wheat varieties and organ characteristics, thus having stronger generalization ability. Experiments have shown that the ViT model performs well on different test datasets and has high robustness and stability. Because different varieties of wheat may have subtle changes in morphology, which may be difficult to capture in traditional CNN models, the overall accuracy of the ViT model is 83.2% (e.g. Figure 15 ), indicating that the ViT model has advantages in processing feature extraction tasks of different wheat varieties.

[0087] While the ViT model excels in capturing global features, it struggles with extracting local, detailed features. Therefore, the feature extraction capabilities of the base network are enhanced by adding a spatial pyramid attention module (SPA). This module aims to direct the model to focus on important regions of the input image and ignore irrelevant or background portions. Incorporating the SPA module into the ViT model resulted in an accuracy of 92.2% in wheat variety recognition, a 9% improvement over the ViT model alone. In the organ-level feature classification task, the ViT+SPA model achieved an accuracy of 94.8% for ear type classification, compared to only 84.6% for the ViT model alone. Furthermore, the model achieved an accuracy of 96.4% for canopy compactness, also outperforming the ViT model alone's 85.6%. This demonstrates that the ViT+SPA model is capable of distinguishing wheat varieties from both local and global features. This combination not only enhances the model's ability to capture local details but also retains the ViT model's strengths in global feature extraction.

[0088] In addition, the ViT+SPA model also provides a certain degree of model interpretability. The attention weights can be used to intuitively observe the image areas that the model focuses on.

[0089] To quantify subtle morphological and structural differences between different varieties, such as leaf surface roughness, vein distribution, and ear morphology, eight texture features were extracted from the model's target regions using the GLCM. These features provide additional information, complementing traditional morphological and spectral features, enabling better differentiation between wheat varieties with similar morphology but distinct textures. A systematic analysis of these texture features allows quantification of morphological differences in the target regions across different trained deep learning models. For example, higher entropy values ​​in entropy box plots indicate greater image complexity, meaning larger target regions and richer and more diverse image features. Therefore, higher entropy values ​​indicate that the feature maps extracted from certain varieties contain more information and detail, exhibiting complex and diverse textural structures. Homogeneity is an important metric for measuring the uniformity of grayscale variations within an image. High homogeneity values ​​indicate minimal pixel value variation and relatively consistent grayscale levels within the image region, indicating that images of different varieties do not differ significantly in leaf and ear color, canopy density, and other aspects. Figure 4The data demonstrates the variability of texture traits across varieties. For example, Yangfumai No. 6 exhibits a relatively uniform contrast distribution, with the 25%-75% of the data falling between 0.1 and 0.80. This is followed by Yangmai No. 16, with the data ranging from 0.22 to 0.78. Other varieties exhibit a more concentrated distribution of contrast traits. Xumai No. 32 exhibits a high degree of homogeneity, with the 25%-75% of the data falling between 0.58 and 0.80, indicating a more uniform texture. Baoji 0601, on the other hand, exhibits a lower degree of homogeneity, with the 25%-75% of the data falling between 0.2 and 0.44. Differences in texture traits within the model's target region often reflect differences in physiological structure and genetic background across varieties. For example, some varieties (such as Yangfumai No. 8 and Sukemai No. 1) may exhibit more complex texture structures, reflecting the development of more cellular layers and tissue structures during growth. Analyzing these texture features can reveal deeper physiological and genetic differences between varieties, providing important insights for cultivar improvement and genetic research.

[0090] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

[0091] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

Claims

1. A wheat variety identification method based on VIT and spatial attention mechanism, characterized in that: The steps include: S1. Image acquisition: Use camera equipment to capture images of wheat canopy and side plant types; S2. Construction of wheat variety dataset: Construct the WMVD dataset by setting a fixed-size window on the original image and sliding it step by step to generate multiple sub-images; S3. Design of wheat variety identification model framework and construction of feature extraction network: S31 and ViT use the self-attention mechanism of the Transformer layer to extract global features; S32, add the spatial pyramid attention module to enhance the feature extraction capability of the basic network; S4. High-dimensional feature visualization based on T-SNE: Dimensionality reduction of high-dimensional features extracted by the ViT+SPA model and visualization of features; S5. Texture trait extraction based on wheat feature map: The gray-level co-occurrence matrix is ​​used to extract the areas of interest in the model in the wheat feature map to quantify the subtle differences in two-dimensional morphology and structure of different varieties.

2. The wheat variety identification method based on VIT and spatial attention mechanism according to claim 1, characterized in that: When collecting wheat canopy images, the camera equipment was parallel to the ground and kept 1.8m away from the canopy. When collecting wheat side plant type images, the camera equipment's shooting pole was 1m away from the plot, and the camera of the camera equipment was kept at 45° to the plot to ensure that the entire side of the plot was covered.

3. The wheat variety identification method based on VIT and spatial attention mechanism according to claim 2, characterized in that: The WMVD dataset includes awned type, non-awned type, dense canopy type, and sparse canopy type.

4. The wheat variety identification method based on VIT and spatial attention mechanism according to claim 3, characterized in that: The specific steps of S31 are: S311, through ViT, first split the input image into fixed-size image blocks, then flatten each image block and linearly map it to a high-dimensional space, the input image is X∈R H×W×C , split it into N image blocks, each of which is P×P×C in size. Each image block is flattened into a vector and mapped to a high-dimensional space through linear transformation, where H, W, and C represent the length, width, and number of channels of the image, respectively; The specific formula is as follows: x p =W p flatten(x p )+b p ,p=1,2,…N S312, retaining position information of different image blocks through position coding; Among them, the position encoding representation of each image block is shown by the following formula Among them, W p and b p are the weight and bias of the linear transformation, E i Represents the code of the i-th position. S313: Before entering the Transformer layer, a classification label is added to the input sequence. The input sequence is extracted from the wheat image features through the multi-head attention mechanism and feedforward neural network in multiple Transformer layers. The specific extraction formula for extracting image features is as follows: Q=XW Q ,K=XW K ,V=XW V MSA(Q,K,V)=Concat(head1,head2,…,head h )W o z′ l =MSA(LN(z l-1 ))+z l-1 With l =MLP(LN(z′ l ))+z′ l Where, X represents the input matrix obtained by linear transformation of each patch block in step S311, and W Q 、W K 、W V Represents the parameter matrix to be learned; Attention is the attention function, head i represents the i-th attention head, represents the parameter matrix to be learned for the i-th attention head, d k The dimension of the formula key, is a scaling factor used to prevent certain areas from having too much attention weight; MSA stands for multi-head attention, W o Represents the linear change matrix of multi-head attention output, z ′ l represents the output of the l-th layer multi-head attention, LN represents the layer normalization, z l is the final output of layer l.

5. The wheat variety identification method based on VIT and spatial attention mechanism according to claim 3 or 4, characterized in that: The specific steps of S32 are: S321, using the output of the ViT layer as input, passes through two adaptive average pooling layers to automatically adjust the pooling window and stride; S322, resize the two outputs into two-dimensional vectors and generate a 1D attention map through the Concat function; S322. Attention weights are obtained through calculations of a series of FC layers and BN layers.

6. The wheat variety identification method based on VIT and spatial attention mechanism according to claim 1, characterized in that: The specific steps of S4 are: S41. Similarity modeling in high-dimensional space. t-SNE uses the Gaussian kernel function to calculate the pairwise similarity between all data points in high-dimensional space. For a given high-dimensional data point pair (x i ,x j ), whose similarity measure is expressed as conditional probability P j|i , which means that given x i In the case of x j The probability of being selected; data points that are closer are given a higher probability of selection, while data points that are farther away are given a lower probability of selection, so that all point pairs (x i ,x j ) is symmetrized to obtain the joint probability distribution P ij ; S42, Similarity modeling in low-dimensional space, t-SNE calculates the pairwise similarity between data points in low-dimensional space through t distribution; S43, t-SNE minimizes the Kullback-Leibler divergence between the high-dimensional probability distribution P and the low-dimensional probability distribution Q; t-SNE iteratively optimizes the objective function through gradient descent:

7. The wheat variety identification method based on VIT and spatial attention mechanism according to claim 1, characterized in that: The specific steps of S5 are: GLCM generates a matrix describing the spatial relationship of gray levels by counting the co-occurrence frequency of gray levels at specific distances and directions, and multiple texture traits are calculated based on GLCM.

8. The wheat variety identification method based on VIT and spatial attention mechanism according to claim 7, characterized in that: The specific distance in the step S5 is 1 pixel and the direction is 45°.

9. The wheat variety identification method based on VIT and spatial attention mechanism according to claim 8, characterized in that: The plurality of texture traits include entropy, homogeneity, angular second moment, variance, contrast, variability, mean and standard deviation.

10. The wheat variety identification method based on VIT and spatial attention mechanism according to claim 9, characterized in that: The specific calculation formula of the texture properties is as follows: Assume that p(i, j) represents the element in the i-th row and j-th column of the gray-level co-occurrence matrix, N is the total number of gray levels (N=255), and u is the mean of GLCM. The formula for calculating entropy is: The formula for calculating homogeneity is: The calculation formula for the angular second moment is: The formula for calculating the mean is: The formula for calculating variance is: The formula for calculating contrast is: The formula for calculating the difference is: The formula for calculating the standard deviation is: