Winter wheat growth stage classification method based on improved YOLOv8

Through the improvement of the YOLOv8n-cls model, the dual-branch downsampling module, full-dimensional dynamic convolution and triple attention mechanism were introduced, which solved the identification problem in complex environments in winter wheat growth stage, and achieved more efficient classification and identification effects.

CN120298884APending Publication Date: 2025-07-11XINJIANG UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510348209.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify the winter wheat growth stage in complex agricultural environments, especially when light and background factors change, the performance of the deep learning model has significantly decreased, resulting in inefficient classification and waste of resources.

Method used

Structural improvements were made to the YOLOv8n-cls model, dual-branch downsampling module, full-dimensional dynamic convolution and triple attention mechanism were introduced to enhance feature extraction and classification capabilities.

Benefits of technology

It improves the accuracy of winter wheat growth stage identification and agricultural production efficiency, improves resource utilization, and enhances the robustness and recognition accuracy of the model in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298884A_ABST
    Figure CN120298884A_ABST
Patent Text Reader

Abstract

The invention provides a winter wheat growth stage classification method based on improved YOLOv8, and the method comprises the steps: carrying out the multi-angle and multi-scale omnibearing shooting of a crop, and obtaining an original image covering a plurality of growth stages of the crop; performing data preprocessing on the original image to obtain a data set; the YOLOv8n-cls model is improved to obtain an improved model, and the improved model comprises the steps that a down-sampling convolution kernel in the YOLOv8n-cls model is replaced with a double-branch down-sampling module combining convolution and pooling; standard convolution in a Bottleneck structure of a C2f feature extraction module in the YOLOv8n-cls model is replaced with full-dimensional dynamic convolution, and a C2f-ODConv module is obtained; a triple attention module is added to the output end of the C2f-ODConv module; and performing classification operation on the data set through the improved model to obtain a classification result. According to the method, the YOLOv8n-cls model is structurally improved, and a triple attention mechanism is introduced, so that the winter wheat growth stage identification and classification accuracy, the agricultural production efficiency and the resource utilization rate are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent agriculture, and particularly to a method for classifying the growth stages of winter wheat based on improved YOLOv8. Background Art

[0002] Wheat is one of the most important food crops globally and plays a crucial role in food security due to its extensive economic and nutritional value. Accurately classifying the growth stages of winter wheat is of great significance for agricultural producers to formulate precise planting management strategies, reduce production costs, increase yields, and achieve sustainable agricultural development. The traditional method for obtaining information on the growth stages of winter wheat mainly relies on manual observation. Agricultural workers analyze the morphological characteristics of winter wheat through on-site inspections and combined with personal experience. However, this method is inefficient, easily affected by subjective judgment, and difficult to obtain dynamic growth information under large-scale planting conditions in a timely manner, and it is no longer able to meet the requirements of modern agricultural production.

[0003] Early classification of crop growth stages mainly relied on manually designed feature extraction methods, such as color, texture, shape, and their combined descriptors. These methods distinguish different growth stages by statistically analyzing the color distribution, texture patterns, or morphological characteristics in images. Although these methods have achieved certain results in specific scenarios, they have significant limitations, such as: it is difficult for manually extracted features to fully capture the high-dimensional information of crop appearance, requiring a large amount of domain knowledge and resource support, high development costs, being easily interfered with in complex natural environments (such as light changes, background interference), and having poor generalization ability, etc.

[0004] With the rapid development of deep learning technology, especially the excellent performance of convolutional neural networks (CNNs) in image classification tasks. CNNs can automatically extract information features from low-level textures to high-level semantics through multi-level convolutional operations, getting rid of the dependence on manually designed features. However, existing deep learning methods still face the following problems: most studies only use small-scale datasets of 3 - 5 growth stages, which are difficult to comprehensively cover the crop growth cycle; the feature differences between adjacent stages are weak, especially in complex field environments, where background factors such as light and soil exacerbate the difficulty of feature extraction; existing models are mostly trained in controlled environments, and their performance drops significantly when facing actual agricultural scenarios (such as low light or complex backgrounds). Therefore, it is very necessary to design a method for classifying the growth stages of winter wheat based on improved YOLOv8. Summary of the Invention

[0005] The object of the present invention is to provide a method for classifying the growth stages of winter wheat based on improved YOLOv8. By improving the structure of the YOLOv8n-cls model and introducing a triple attention mechanism, the accuracy of identifying and classifying the growth stages of winter wheat, agricultural production efficiency, and resource utilization rate are improved.

[0006] To achieve the above object, the present invention provides the following solutions:

[0007] A winter wheat growth stage classification method based on improved YOLOv8, comprising the following steps:

[0008] Take all-round photos of the crops from multiple angles and scales to obtain original images covering multiple growth stages of the crops;

[0009] Perform data preprocessing operations on the original images to obtain a dataset;

[0010] Improve the structure of the YOLOv8n-cls model to obtain an improved model. The steps for improving the structure of the YOLOv8n-cls model include:

[0011] Replace the downsampling convolutional kernel in the YOLOv8n-cls model with a dual-branch downsampling module that combines convolution and pooling. The dual-branch downsampling module includes a pooling branch and a convolutional branch;

[0012] Replace the standard convolution in the Bottleneck structure of the C2f feature extraction module in the YOLOv8n-cls model with a full-dimensional dynamic convolution to obtain a C2f-ODConv module;

[0013] Add a triple attention module at the output end of the C2f-ODConv module. The triple attention module is composed of three parallel branches;

[0014] Perform classification operations on the dataset through the improved model to obtain classification results.

[0015] Optionally, the data preprocessing operations include: resolution adjustment, data augmentation, image scaling, and dataset partitioning.

[0016] Optionally, the downsampling steps of the dual-branch downsampling module include:

[0017] Perform a splitting operation on the crop growth stage feature map in the channel dimension to obtain a first sub-feature map and a second sub-feature map;

[0018] In the pooling branch, perform max pooling and convolution operations on the first sub-feature map in sequence to obtain a pooled output feature map;

[0019] In the convolutional branch, perform a convolution operation on the second sub-feature map to obtain an intermediate feature map;

[0020] Divide the intermediate feature map into m groups, and perform convolution operations on each group respectively to obtain m second-constructed feature maps;

[0021] Perform a concatenation operation on the intermediate feature map and the second-constructed feature map to obtain a convolutional output feature map;

[0022] Perform a concatenation operation on the pooling output feature map and the convolutional output feature map in the channel dimension to obtain a downsampled feature map.

[0023] Optionally, the convolutional steps of the full-dimensional dynamic convolution include:

[0024] Perform global average pooling, fully connected, ReLU activation, and branching operations on the input features in sequence to obtain four branch features;

[0025] After inputting the branch features into a fully connected layer, obtain attention maps of different dimensions; the attention maps include: the first attention map, the second attention map, the third attention map, and the fourth attention map; the different dimensions include: space, number of input channels, number of output channels, and convolutional kernel;

[0026] Normalize the first attention map, the second attention map, and the third attention map respectively through the Sigmoid function, and normalize the fourth attention map through the Softmax function to obtain attention coefficients of different dimensions;

[0027] Dynamically adjust the convolutional kernel weights for the attention coefficients of different dimensions through the multi-head attention mechanism to obtain an optimized convolutional kernel;

[0028] Perform a convolution operation on the input features through the optimized convolutional kernel to obtain multi-dimensional output features.

[0029] Optionally, the feature filtering steps of the triple attention module include:

[0030] Perform rotation, dimensionality reduction, convolution, normalization, attention weight generation, and rotation reset operations on the input tensor along the height axis in sequence to obtain a first branch tensor;

[0031] Perform rotation, dimensionality reduction, convolution, normalization, attention weight generation, and rotation reset operations on the input tensor along the width axis in sequence to obtain a second branch tensor;

[0032] Reduce the input tensor to two channels and then perform convolution, normalization, sigmoid activation, and attention weight generation operations in sequence to obtain a third branch tensor;

[0033] Perform an average aggregation operation on the first branch tensor, the second branch tensor, and the third branch tensor to obtain a triple attention output.

[0034] Optionally, the expression of the convolutional output feature map is: Y2 = Concat(X'2,(x1,x2,......,x m)); where Y2 is the convolutional output feature map, Concat(·) is the concatenation operation, X'2 is the intermediate feature map, and x m is the m-th second-constructed feature map.

[0035] Optionally, the expression for the multi-dimensional output feature is: where y c is the multi-dimensional output feature, is the optimized convolutional kernel, * is the convolution operation, x is the input feature, and n is the number of dimensions.

[0036] Optionally, the expression for the triple-attention output is: where y s is the triple-attention output, and are the first-branch tensor, the second-branch tensor, and the third-branch tensor respectively, and ω1, ω2, and ω3 are the attention weights of the first-branch tensor, the second-branch tensor, and the third-branch tensor respectively.

[0037] According to the specific embodiments provided by the present invention, the following technical effects are disclosed: The winter wheat growth stage classification method based on the improved YOLOv8 provided by the present invention includes: taking all-round pictures of crops from multiple angles and scales to obtain original images covering multiple growth stages of the crops; performing data preprocessing operations on the original images to obtain a data set; improving the structure of the YOLOv8n-cls model to obtain an improved model, including: replacing the downsampling convolutional kernel in the YOLOv8n-cls model with a double-branch downsampling module that combines convolution and pooling; the double-branch downsampling module includes a pooling branch and a convolutional branch; replacing the standard convolution in the Bottleneck structure of the C2f feature extraction module in the YOLOv8n-cls model with a full-dimensional dynamic convolution to obtain a C2f-ODConv module; adding a triple-attention module at the output end of the C2f-ODConv module; classifying the data set through the improved model to obtain a classification result. This method improves the accuracy of winter wheat growth stage recognition and classification, agricultural production efficiency, and resource utilization rate by improving the structure of the YOLOv8n-cls model and introducing a triple-attention mechanism. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0039] Figure 1Flow chart of the winter wheat growth stage classification method of the present invention;

[0040] Figure 2 Structural diagram of the EOT - YOLO framework of the present invention;

[0041] Figure 3 Structural diagram of the EDown module of the present invention;

[0042] Figure 4 Structural diagram of the C2f - ODConv module framework of the present invention;

[0043] Figure 5 Structural diagram of the TA module framework of the present invention. Detailed implementation manners

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0045] To make the above - mentioned objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0046] As Figure 1 shown, the present invention provides a winter wheat growth stage classification method based on improved YOLOv8, including the following steps:

[0047] Step 100: Take all - around photos of the crops from multiple angles and scales to obtain original images covering multiple growth stages of the crops;

[0048] In this embodiment, the data source is a wheat demonstration base in a certain county of a certain autonomous region. Nikon 50D camera and vivo S16 mobile phone are used for shooting, ensuring image quality and diversity. The shooting angles include vertical 90°, horizontal and vertical 30° and 45°, covering different scales from single plant, local to canopy. Finally, 25,319 original images covering ten growth stages of winter wheat are obtained.

[0049] Step 200: Perform data pre - processing operations on the original images to obtain a data set;

[0050] Specifically, the data pre - processing operations include:

[0051] Resolution adjustment: The original image resolutions are 4752×3168 and 3456×4608 pixels. The images are uniformly cropped to 1500×1000 pixels to reduce computational costs, remove irrelevant background information, and retain effective features.

[0052] Data augmentation: Perform operations of vertical flipping, horizontal flipping, random rotation, brightness enhancement, and Gaussian noise on the images after resolution adjustment to improve the generalization ability of the model in complex field environments.

[0053] Image scaling: Scale all images to 640×640 pixels and perform normalization to accelerate model convergence.

[0054] Dataset division: Divide the dataset into a training set and a test set in a ratio of 8:2.

[0055] Step 300: Improve the structure of the YOLOv8n-cls model to obtain an improved model.

[0056] Specifically, based on the YOLOv8n-cls model, this embodiment proposes an improved model - EOT-YOLO, whose framework is as Figure 2 shown. The specific ways to improve the structure of the YOLOv8n-cls model include:

[0057] Step 301: Replace the downsampling convolution kernel in the YOLOv8n-cls model with a dual-branch downsampling module (EDown) that combines convolution and pooling; the dual-branch downsampling module includes a pooling branch and a convolutional branch.

[0058] Step 302: Replace the standard convolution in the Bottleneck structure of the C2f feature extraction module in the YOLOv8n-cls model with a full-dimensional dynamic convolution to obtain a C2f-ODConv module.

[0059] Step 303: Add a triple attention module at the output end of the C2f-ODConv module; the triple attention module consists of three parallel branches.

[0060] Specifically, the EDown module is as Figure 3 shown. The downsampling steps implemented through this module include:

[0061] Perform a splitting operation on the winter wheat growth stage feature map X in the channel dimension to obtain a first sub-feature map X1 and a second sub-feature map X2; the expression for the splitting process is: X1, X2 = Split(X); where Split(·) represents the splitting operation.

[0062] In the pooling branch, first perform a 3×3 max pooling operation with a stride of 2 on the feature map to reduce its spatial resolution, and then perform a convolution operation to adjust the number of channels to obtain the pooled output feature map Y1. The expression for this process is: Y1 = Conv(MaxPool2D(X1)); where Conv(·) is the convolution operation and MaxPool2D(·) is the max pooling operation.

[0063] In the convolution branch, perform a 3×3 convolution operation with a stride of 2 on the second sub-feature map X2 to obtain the intermediate feature map X'2. The expression for this process is: X'2 = X2 * f; where * is the convolution operation, and f is the convolution filter used.

[0064] Divide the intermediate feature map X'2 into m groups, and perform a 5×5 convolution operation on each group to obtain m secondary structure feature maps x i ; The expression for this process is: where x' 2,i represents the i-th group of feature maps in X'2.

[0065] Concatenate the intermediate feature map X'2 and the secondary structure feature maps x i to obtain the convolution output feature map Y2. The expression for this process is: Y2 = Concat(X'2, (x1, x2,......, x m )); where Concat(·) is the concatenation operation, and x m is the m-th secondary structure feature map.

[0066] Concatenate the pooled output feature map Y1 and the convolution output feature map Y2 along the channel dimension to obtain the downsampled feature map.

[0067] It should be noted that in the YOLOv8n-cls model, using a 3×3 convolution kernel with a stride of 2 for the downsampling operation will cause the loss of important detailed information in the image. This lack of detail restricts the accurate capture of subtle features in adjacent growth stages of winter wheat, thus affecting the classification performance. In this embodiment, the original convolution kernel is replaced with a dual-branch downsampling module. On the one hand, this module extracts key local features through convolution, and on the other hand, it captures global features through pooling, balancing the extraction of local and global information. In addition, grouped convolution is used to reduce the computational complexity and enhance feature diversity, making the model more flexible in extracting local features. And the dual-branch structure of the EDown module effectively improves the representation ability of detailed features, reduces the loss of detailed information during downsampling, and effectively retains the features of different growth stages.

[0068] Specifically, when the convolutional model of the YOLOv8n-cls model processes wheat canopy images, its ability to capture key features is limited. Especially in complex scenarios with large variations in lighting and background, it is difficult to effectively extract and distinguish the subtle differences in different growth stages. In this embodiment, the standard convolution in the C2f feature extraction module of the YOLOv8n-cls model is replaced with Omni-Dimensional Dynamic Convolution (ODConv) to obtain a new C2f-ODConv module, so as to enhance the feature extraction ability of the model. The specific structure of the C2f-ODConv module is as Figure 4 shown.

[0069] More specifically, the Omni-Dimensional Dynamic Convolution (ODConv) consists of four independent attention branches, and its convolution steps include:

[0070] The input feature is first subjected to global average pooling (GAP), fully connected (FC), and ReLU activation operations in sequence, and then divided into four branch features by four independent attention branches. Then each branch feature passes through a fully connected layer to generate attention maps of different dimensions. Among them, the attention maps include: the first attention map, the second attention map, the third attention map, and the fourth attention map. Then, the first attention map, the second attention map, and the third attention map are respectively normalized by the Sigmoid function, and the fourth attention map is normalized by the Softmax function to obtain four types of attention coefficients α si , α ci , α fi , and α wi , corresponding to four different dimensions of space, the number of input channels, the number of output channels, and the convolution kernel respectively.

[0071] The convolution kernel weights are dynamically adjusted for the attention coefficients of different dimensions through the multi-head attention mechanism to obtain an optimized convolution kernel. The expression of the dynamic adjustment process is: Among them, is the convolution kernel after attention weight adjustment (optimized convolution kernel), ⊙ represents element-wise multiplication, and are the attention coefficients on different dimensions, respectively calculated through the multi-head attention mechanism. Finally, the input feature is convolved by the optimized convolution kernel to obtain multi-dimensional output features. The calculation formula of the output y c of ODConv is:

[0072]

[0073] Among them, x is the input feature, n is the number of dimensions, and in this embodiment, n = 4.

[0074] It should be noted that ODConv dynamically adjusts the weights of the convolutional kernel in four dimensions: space, input channels, output channels, and convolutional kernels through a multi-dimensional attention mechanism, enhancing the model's adaptability to input features, optimizing the comprehensiveness of feature extraction and context awareness capabilities. This enables the C2f-ODConv module to adaptively adjust according to the features of images at different growth stages, improving the model's robustness and recognition accuracy in complex environments, and also significantly enhancing the model's ability to capture key features of winter wheat at different growth stages and its performance in dealing with subtle feature changes.

[0075] Specifically, the triple attention (TA) module consists of three independent parallel branches. The detailed structure of the TA module is as Figure 5 shown, and the feature filtering steps of the three branches are as follows:

[0076] For the first branch to handle the interaction between height (H) and channels (C), the input tensor x is rotated counterclockwise by 90° along the height (H) axis to generate tensor x1, and then dimensionality reduction is performed through Z-pool to form Then, it is processed through a 5×5 convolutional layer and batch normalization, and finally, an attention weight is generated through the sigmoid function, and the generated attention weight is applied to and it is rotated clockwise by 90°, making return to the original dimension, and finally the first branch tensor

[0077] The processing process of the second branch is similar to that of the first branch, except that the input tensor x is rotated counterclockwise along the width (W) axis to form tensor x2.

[0078] The third branch directly uses Z-pool to reduce the dimensionality of the input tensor x to two channels to obtain the simplified tensor x3, and then through the same operations of convolution, normalization, sigmoid activation, and attention weight generation, the third branch tensor

[0079] Finally, through the formula the first branch tensor, the second branch tensor, and the third branch tensor are averaged and aggregated to obtain the triple attention output y s , where ω1, ω2, and ω3 are the attention weights of the first branch tensor, the second branch tensor, and the third branch tensor respectively.

[0080] It should be noted that when the YOLOv8n-cls model processes canopy images, it is often disturbed by irrelevant information such as soil and weeds in the background, lacking an efficient redundant information filtering mechanism, unable to focus on the key areas of wheat, and weakening the reliability and accuracy of classification. The TA module introduced in this embodiment realizes the interactive modeling of spatial features and channel features by capturing the dependency relationships between dimensions through information interaction across space (H and W dimensions) and channels (C dimension). Through the TA module, when the model processes the complex background of the winter wheat growth stage, it can more accurately focus on the key areas closely related to the growth stage, effectively reducing the interference of environmental noise.

[0081] Step 400: Classify the dataset through the improved model to obtain the classification result.

[0082] In this embodiment, the classification results of the improved model EOT-YOLO are compared with 13 other existing models. The same self-built winter wheat growth stage dataset and experimental environment settings are used, and evaluation indicators such as precision (Pre), recall (Rec), accuracy (ACC), and F1 score are used for testing. The results are shown in Table 1.

[0083] Table 1 Comparison table of classification result evaluation indicators

[0084]

[0085]

[0086] The EOT-YOLO model of the present invention improves the accuracy from 93.08% to 96.11%, the precision from 93.26% to 96.23%; the recall from 93.50% to 96.27%; and the F1 value from 93.31% to 96.23%.

[0087] The beneficial effects of the present invention are as follows:

[0088] 1) The double-branch structure of the EDown module reduces the loss of detailed information in the downsampling process and effectively retains the features of different growth stages;

[0089] 2) By replacing the standard convolution in the Bottleneck structure of the C2f feature extraction module in the YOLOv8n-cls model with full-dimensional dynamic convolution, the robustness and recognition accuracy of the model in complex environments are improved;

[0090] 3) The introduction of the triple attention mechanism improves the accuracy of feature extraction and reduces the interference of environmental noise.

[0091] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.

[0092] In the present invention, specific examples are used to illustrate the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A winter wheat growth stage classification method based on improved YOLOv8, characterized in that, It includes the following steps: Take all-round photos of crops from multiple angles and scales to obtain original images covering multiple growth stages of the crops; Perform data preprocessing operations on the original images to obtain a dataset; Improve the structure of the YOLOv8n-cls model to obtain an improved model. The steps for improving the structure of the YOLOv8n-cls model include: Replace the downsampling convolution kernel in the YOLOv8n-cls model with a dual-branch downsampling module that combines convolution and pooling. The dual-branch downsampling module includes a pooling branch and a convolution branch; Replace the standard convolution in the Bottleneck structure of the C2f feature extraction module in the YOLOv8n-cls model with a full-dimensional dynamic convolution to obtain a C2f-ODConv module; Add a triple attention module to the output end of the C2f-ODConv module. The triple attention module is composed of three parallel branches; Perform classification operations on the dataset through the improved model to obtain classification results.

2. The winter wheat growth stage classification method based on improved YOLOv8 according to claim 1, wherein The data preprocessing operations include: resolution adjustment, data augmentation, image scaling, and dataset division.

3. The winter wheat growth stage classification method based on improved YOLOv8 according to claim 1, wherein, The downsampling steps of the dual-branch downsampling module include: Perform a splitting operation on the crop growth stage feature map in the channel dimension to obtain a first sub-feature map and a second sub-feature map; In the pooling branch, perform max pooling and convolution operations on the first sub-feature map in sequence to obtain a pooled output feature map; In the convolution branch, perform a convolution operation on the second sub-feature map to obtain an intermediate feature map; Divide the intermediate feature map into m groups, and perform convolution operations on each group respectively to obtain m two-construct feature maps; Perform a splicing operation on the intermediate feature map and the two-construct feature maps to obtain a convolution output feature map; Perform a splicing operation on the pooled output feature map and the convolution output feature map in the channel dimension to obtain a downsampled feature map.

4. The winter wheat growth stage classification method based on the improved YOLOv8 according to claim 1, wherein, The convolution steps of the full-dimensional dynamic convolution include: Perform global average pooling, fully connected, ReLU activation, and branching operations on the input features in sequence to obtain four branch features; Input the branch features into a fully connected layer to obtain attention maps of different dimensions. The attention maps include: a first attention map, a second attention map, a third attention map, and a fourth attention map. The different dimensions include: space, number of input channels, number of output channels, and convolution kernel; Normalize the first attention map, the second attention map, and the third attention map respectively through the Sigmoid function, and normalize the fourth attention map through the Softmax function to obtain attention coefficients of different dimensions; Dynamically adjust the convolution kernel weights of the attention coefficients of different dimensions through a multi-head attention mechanism to obtain an optimized convolution kernel; Perform a convolution operation on the input features through the optimized convolution kernel to obtain multi-dimensional output features.

5. The winter wheat growth stage classification method based on the improved YOLOv8 according to claim 1, characterized in that, The feature filtering steps of the triple attention module include: Perform rotation, dimensionality reduction, convolution, normalization, attention weight generation, and rotation reset operations on the input tensor along the height axis in sequence to obtain a first branch tensor; Perform the operations of the rotation, the dimensionality reduction, the convolution, the normalization, the attention weight generation, and the rotation reset on the input tensor sequentially along the width axis to obtain a second branch tensor; Reduce the input tensor to two channels and then perform the operations of the convolution, the normalization, the sigmoid activation, and the attention weight generation sequentially to obtain a third branch tensor; Perform an average aggregation operation on the first branch tensor, the second branch tensor, and the third branch tensor to obtain a triple attention output.

6. The winter wheat growth stage classification method based on improved YOLOv8 according to claim 3, characterized in that, The expression of the convolutional output feature map is: Y2 = Concat(X'2, (x1, x2,......, x m )); where Y2 is the convolutional output feature map, Concat(·) is the concatenation operation, X'2 is the intermediate feature map, and x m is the m-th second-constructed feature map.

7. The winter wheat growth stage classification method based on improved YOLOv8 according to claim 4, characterized in that, The expression of the multi-dimensional output feature is as follows: Among them, y c is the multi-dimensional output feature, is the optimized convolution kernel, * represents the convolution operation, x is the input feature, and n is the number of dimensions.

8. The winter wheat growth stage classification method based on the improved YOLOv8 according to claim 5, characterized in that, The expression of the triple attention output is as follows: where y s is the triple attention output, and are the first branch tensor, the second branch tensor, and the third branch tensor respectively, and ω1, ω2, and ω3 are the attention weights of the first branch tensor, the second branch tensor, and the third branch tensor respectively.

Citation Information

Cited By

  • Ludisia discolor harvest time prediction method and system fused with multi-source data

    CN121144966A

  • A method and system for predicting the harvest period of *Hymenochloa crus-galli* by integrating multi-source data

    CN121144966B