Attention mechanism and capsule layer fused multi-class crop cultivation area classification method
By integrating attention mechanisms with capsule layers, this study solves the problem of fine classification of various crop cultivation areas in complex agricultural landscapes, achieving high-precision and refined classification of cultivation areas and improving the model's discriminative power and generalization performance.
Patent Information
- Application Number
- CN202511468132.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-10-15
AI Technical Summary
Existing technologies struggle to achieve efficient and accurate fine classification of multi-crop cultivated areas in complex agricultural landscapes. In particular, under conditions of farmland fragmentation, diverse crop types, and weed interference, traditional models are unable to effectively model spatial-spectral features, resulting in insufficient discriminative power and classification generalization performance.
A multi-crop cultivation area classification method integrating attention mechanism and capsule layer is proposed. Through principal component analysis dimensionality reduction, spatial and channel attention weighting, lightweight convolution module construction and capsule layer construction, the dynamic routing of capsule network and vectorized neuron representation of spatial spectral features are utilized, and the model is optimized by combining classification and reconstruction loss functions.
The network's ability to extract and represent the spatial-spectral features of hyperspectral remote sensing images has been enhanced, improving the model's classification and generalization performance and enabling high-precision and refined classification of complex cultivated areas.
Smart Images

Figure CN120932113A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fine classification technology for multi-crop cultivation areas in complex agricultural scenarios, and in particular to a multi-crop cultivation area classification method that integrates attention mechanisms and capsule layers. Background Technology
[0002] Efficient and accurate fine classification of multi-crop cultivated areas is a key link in ensuring food security and promoting sustainable agricultural development. However, in multi-crop agricultural landscapes, the fragmentation of cultivated land, the diversity of crop types and the complexity of intercropping patterns, coupled with factors such as soil background differences and weed disturbance, result in highly complex spatial and spectral characteristics of land features. This increases intra-class variability and reduces inter-class separability, thus posing a significant challenge to the fine classification of cultivated areas.
[0003] Hyperspectral imagery can accurately characterize the spectral properties of ground features, providing a unique advantage for the fine classification of complex agricultural scenes. In recent years, with the development of deep learning technology, hyperspectral image classification has made significant progress in the field of agricultural remote sensing. Existing models include Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and Deep Belief Networks (DBNs). Among them, CNNs excel in spatial feature extraction, and their variants, such as 1D-CNN, 2D-CNN, and 3D-CNN, have achieved good classification results. However, RNNs and DBNs require the spatially structured image to be unfolded into a one-dimensional vector as input, which can easily lead to the loss of key spatial-spectral joint information. At the same time, traditional CNNs are built based on scalar neurons, which has limitations in representing complex features.
[0004] To overcome these bottlenecks, scholars have proposed the capsule concept and developed capsule networks. This method models hierarchical spatial relationships through dynamic routing mechanisms and vectorized neurons, and can represent instantiation parameters of object features. Although existing studies have explored the application of capsule networks in hyperspectral image classification, their performance in complex agricultural landscapes remains limited. Specifically, traditional capsule networks are not designed for agricultural scenarios and struggle to effectively model spatial-spectral features under conditions of mixed pixels, multiple crop types, and fragmented fields, resulting in insufficient discriminative power and classification generalization performance in complex environments. Summary of the Invention
[0005] The purpose of this invention is to provide a multi-crop cultivation area classification method that integrates attention mechanisms and capsule layers, so as to achieve high-precision and refined classification of cultivation areas under complex agricultural landscapes.
[0006] This invention provides a method for classifying multiple crop cultivation areas by integrating attention mechanisms and capsule layers, including: Step 1: Acquire hyperspectral images of various crop cultivation areas, perform principal component analysis for dimensionality reduction and image segmentation on the hyperspectral images to obtain image blocks; Step 2: Perform spatial attention weighting on the image block; Step 3: Construct a lightweight convolutional module based on the output of Step 2; Step four: Perform channel attention weighting on the output of step three; Step 5: Construct the capsule layer based on the output of Step 4. The capsule layer includes primary capsules, dynamic routing, and digital capsules. Vector-based neurons represent capsules, and the magnitude and direction of the vectors are used to represent the probability and pose attributes of the existence of instances, respectively. The primary capsules use convolution operations to convert features into vector capsules. The digital capsules work with the dynamic routing to process the output of the primary capsules and generate high-level capsules that represent different classes of semantic attributes. Step 6: Classify and reconstruct the digital capsule results. The classification uses the L2 norm to process the digital capsule results, forming a vector with n elements. The index corresponding to the maximum value element in the vector is the category code, thereby determining the target class to which the classification result belongs. The reconstruction is achieved by a fully connected layer. Step 7: Accuracy assessment of the fine classification results of cultivated areas.
[0007] Further, in step one, the principal component analysis dimensionality reduction includes: performing principal component analysis dimensionality reduction on the hyperspectral image to compress its band number to 1 / 4 of the original, and removing high redundancy and noise between hyperspectral image bands; the image segmentation includes: dividing the data after principal component analysis dimensionality reduction into image blocks with a length × width × number of channels of 11 × 11 × 64, which are used as input to the classification network.
[0008] Further, in step two, the image patches are subjected to max pooling and mean pooling in the channel dimension, respectively. The pooling results are then concatenated in a stacked manner to form a concatenated feature of size 11×11×2. Convolution and activation operations are then performed sequentially on this feature to generate spatial attention matrix weights with values ranging from 0 to 1. These spatial attention matrix weights are then multiplied element-wise with the original image patches to complete the weighting operation. The spatial attention weighting process is expressed by the following formula:
[0009] In the formula, F represents the input feature data, AvgPool(·) and MaxPool(·) represent average pooling and max pooling, respectively, Conv(·) represents the convolution operation with a kernel size of 7×7, σ(·) represents the Sigmoid activation function, and O s The spatial attention matrix weights are obtained; the convolution operations involved are represented as follows:
[0010] In the formula, f(·) represents the activation output of the β-th channel in convolutional layer l; f(·) represents the activation function, which is Sigmoid. It is the feature to be activated in the β-th channel of convolutional layer l, derived from the feature map of the previous layer. With the corresponding convolution kernel Perform convolution summation operations and, by adjusting the bias... Obtained by addition; It is used to obtain The input feature map subset; "*" represents the convolution operator.
[0011] Furthermore, in step three, based on the output of step two, a convolution operation is first performed to generate 11×11×128 intrinsic features. Then, a depthwise convolution operation is performed on these features, calculated as follows:
[0012] In the formula, Y represents the intrinsic features after regular convolution, u is the input feature, and the convolution kernel is... ∈ m×k×k×c;
[0013] In the formula, i,j represents performing the j-th linear operation on the i-th intrinsic feature yi of the intrinsic feature set Y, thereby generating the corresponding j-th enhancement feature y. ij The last layer i,s This indicates that by performing an identity mapping on the intrinsic features, the final set of enhanced features obtained is {y}. 11 , y 12 , …, y ms The relationship between the number of intrinsic features m, the number of linear operations s, and the number of final enhanced feature channels n is described as n = m·s (m ≤ n). Ignoring the last identity transformation, this relationship is further expressed as m(s-1) = n / s·(s-1).
[0014] Furthermore, in step four, channel attention operations are performed on the enhanced features output in step three. First, spatial dimension compression is performed on the feature data, and then average pooling is performed on the spatial dimension scale. Considering the information exchange between the current channel and its k neighboring channels, one-dimensional convolution is used to generate channel-dimensional attention weight values. These weights are then multiplied element-wise with the enhanced features to complete the channel-dimensional weighting operation. The calculation process is as follows:
[0015] In the formula, y represents the input channel feature data, C1D(·) represents one-dimensional convolution, and σ(·) represents the Sigmoid activation function. The channel attention weights are obtained, where k represents the size of the one-dimensional convolution kernel, adaptively determined by the following formula:
[0016] In the formula, C represents the number of channels of the input data, k represents the size of the calculated one-dimensional convolution kernel, odd indicates that the obtained value is odd, b=1, and γ=2.
[0017] Furthermore, in step five, the primary capsule uses convolution operations to convert features into capsules in vector form. It contains multiple parallel small convolution kernels with a kernel size of 4×4×8, totaling 32. Each convolution kernel extracts a set of features to form a capsule vector. The information transmission process of vector neurons in the capsule layer includes three parts: affine transformation, weighted summation, and nonlinear activation. Affine transformation encodes the abstract spatial relationship between the lower-level capsules and higher-level features. The input lower-level capsules are processed through the weight matrix W. ij Multiplying (i=1, 2, 3, …) yields high-level features; weighted summation utilizes the coupling coefficient c ij (i=1, 2, 3, …) The high-level features are weighted and summed to obtain the input vector required for the high-level capsule. The calculation formula is as follows:
[0018] In the formula, u i W represents the underlying capsule of the input. ij U is the weight matrix. j|i Represents high-level features, c ij It is a coupling coefficient that has the properties of non-negativity and summing to 1. j This is the input vector required for the high-level capsule; Nonlinear activation utilizes the squash function (·) to transform the vector The magnitude of j is normalized to the interval (0, 1) while maintaining the vector direction, so that the magnitude of the final output high-level capsule vj can be used to characterize the probability of the existence of the object entity:
[0019] In the formula, ||·|| represents the L2 norm, which is the magnitude of the vector, and ||·||2 represents the square of the L2 norm. The first part of the calculation formula is to compress the vector, and the second part is the direction vector. The whole formula achieves a normalization operation on the vector to be consistent with the original direction. Dynamic routing is used to update and determine the coupling coefficient c. ij The dynamic routing process includes: the input underlying capsule u i With weight matrix W i After multiplication, the high-level feature U is obtained. i Based on this, softmax normalization, vector weighted summation, vector squash (·), and coupling coefficient update operations are performed, where the process variable b is initialized to 0; through multiple iterations, the high-level capsule is finally obtained.
[0020] Furthermore, in step six, the expression for the reconstruction operation is:
[0021] In the formula, x is the input vector data, W is the weight matrix, Reshape(·) represents the reshaping operation, and O r This indicates data reconstruction; The loss function is used to constrain the training process of the classification network. A joint loss function is constructed that includes classification and reconstruction losses. The reconstruction loss plays a regularization role and calculates the sum of squares of the differences between the reconstruction output and the original input during the decoding stage.
[0022] In the formula, L c This represents the classification loss for a single capsule neuron, where c is the number of categories, and T is the classification loss. c It is an indicator function, T when the category is c. c The value is 1 otherwise; m+ and m- represent the upper and lower edge thresholds, respectively. The former penalizes the classifier's prediction of a class that does not actually exist, while the latter penalizes the classifier's prediction of a class that does not actually exist. m+ = 0.9, m- = 0.1; λ is the sparsity coefficient, used to adjust the weight, with a value of 0.5; Ltotal represents the total loss, I is the original input, and Or is the decoder's reconstructed output. The weighting factor has a value of 2 × 10. -5 .
[0023] Furthermore, in step seven, the evaluation indicators include overall accuracy, average classification accuracy, and the Kappa coefficient κ; the classification results are compared with ground reference data to obtain the confusion matrix M, and the evaluation indicators are calculated using the following formulas:
[0024] In the formula, n is the number of categories, N is the total number of test samples, and M is the total number of samples. uu M represents the number of samples in which class u is correctly identified as u. uv This represents the number of samples whose category v is identified as u, where u∈[1, n], v∈[1, n].
[0025] The present invention has the following beneficial effects: The multi-crop cultivation area classification method of the present invention, which integrates attention mechanism and capsule layer, enhances the network's ability to extract and represent the spatial and spectral features of hyperspectral remote sensing images by incorporating lightweight convolution, spatial and channel attention, and capsule modules into the encoding and decoding structure. Simultaneously, it employs joint classification and reconstruction loss to improve the model's classification generalization performance, thereby achieving high-precision and refined classification of complex cultivation areas. Attached Figure Description
[0026] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative effort.
[0027] Figure 1 This is a flowchart of the technology of this invention; Figure 2 These are hyperspectral images and land cover labels for the implementation area; Figure 3 This is the dynamic routing execution process; Figure 4 It is the result of refined classification of cultivated areas. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention. The technical solutions provided by various embodiments of this invention will be described in detail below with reference to the accompanying drawings.
[0029] Please see Figures 1 to 4This invention provides a method for classifying multiple crop cultivation areas by integrating attention mechanisms and capsule layers. The technical process is as follows: Figure 1 As shown, it includes: Step 1: Obtain hyperspectral images of various crop cultivation areas, perform principal component analysis for dimensionality reduction and image segmentation on the hyperspectral images to obtain image blocks.
[0030] The hyperspectral imagery was acquired from open-source data in the remote sensing community. This data was collected by a full-spectrum multimodal imaging spectrometer sensor in a multi-crop cultivated area. The image spectral range is 400–1000 nm, with 256 bands; the image size is 1580 × 3750 pixels, and the spatial resolution is 0.5 m. The data contains 21 land cover categories, including 14 categories of seedlings, economic crops, and food crops. During model training, 1% of each category was stratified and sampled. The hyperspectral imagery covers the following area: Figure 2 As shown, the category names and detailed sample information are listed in Table 1.
[0031] Table 1. Land cover categories and sample information in hyperspectral images of a multi-crop cultivated area.
[0032] Principal component analysis (PCA) dimensionality reduction includes: performing PCA on the hyperspectral image to compress its band count to one-quarter of the original, removing high redundancy and noise between hyperspectral image bands, and improving the signal-to-noise ratio and subsequent classification efficiency. Image segmentation includes: dividing the PCA-reduced data into image blocks with a length × width × number of channels of 11 × 11 × 64, which are used as input to the classification network.
[0033] Step 2: Perform spatial attention weighting on the image block.
[0034] Image patches are subjected to both max pooling and mean pooling in the channel dimension. The pooling results are then concatenated in a stacked manner to form a concatenated feature of size 11×11×2 (taking a single image patch as an example). Convolution and activation operations are then performed sequentially on this feature to generate spatial attention matrix weights (11×11×1) with values ranging from 0 to 1. These spatial attention matrix weights are then element-wise multiplied with the original image patch to complete the weighting operation. The spatial attention weighting process is expressed by the following formula:
[0035] In the formula, F represents the input feature data, AvgPool(·) and MaxPool(·) represent average pooling and max pooling, respectively, Conv(·) represents the convolution operation with a kernel size of 7×7, σ(·) represents the Sigmoid activation function, and O sThe spatial attention matrix weights are obtained; the convolution operations involved are represented as follows:
[0036] In the formula, f(·) represents the activation output of the β-th channel in convolutional layer l; f(·) represents the activation function, which is Sigmoid. It is the feature to be activated in the β-th channel of convolutional layer l, derived from the feature map of the previous layer. With the corresponding convolution kernel Perform convolution summation operations and, by adjusting the bias... Obtained by addition; It is used to obtain The input feature map subset; "*" represents the convolution operator.
[0037] Step 3: Construct a lightweight convolutional module based on the output of Step 2.
[0038] Based on the output of step two, a convolution operation is first performed to generate 11×11×128 intrinsic features. Then, a depthwise convolution operation is performed on these features, which can be regarded as a linear operation. The calculation process is as follows:
[0039] In the formula, Y represents the intrinsic features after regular convolution, u is the input feature, and the convolution kernel is... ∈ m×k×k×c;
[0040] In the formula, i,j Represents the i-th eigenfeature y of the eigenfeature set Y. i Perform the j-th linear operation to generate the corresponding j-th enhanced feature y. ij The last layer i,s This indicates that by performing an identity mapping on the intrinsic features, the final set of enhanced features obtained is {y}. 11 , y 12 , …, y ms The relationship between the number of intrinsic features m, the number of linear operations s, and the number of final enhanced feature channels n is described as n = m·s (m ≤ n). Ignoring the last identity transformation, this relationship is further expressed as m(s-1) = n / s·(s-1).
[0041] Step four: Perform channel attention weighting on the output of step three.
[0042] Channel attention operations are performed on the enhanced features output in step three. First, spatial compression is performed on the feature data, followed by average pooling on the spatial scale. Considering the information exchange between the current channel and its k neighboring channels, one-dimensional convolution is used to generate channel-dimensional attention weight values. These weights are then multiplied element-wise with the enhanced features to complete the channel-dimensional weighting operation. The calculation process is as follows:
[0043] In the formula, y represents the input channel feature data, C1D(·) represents one-dimensional convolution, and σ(·) represents the Sigmoid activation function. The channel attention weights are obtained, where k represents the size of the one-dimensional convolution kernel, adaptively determined by the following formula:
[0044] In the formula, C represents the number of channels of the input data, k represents the size of the calculated one-dimensional convolution kernel, odd indicates that the obtained value is odd, b=1, and γ=2.
[0045] Step 5: Construct the capsule layer based on the output of Step 4. The capsule layer includes primary capsules, dynamic routing, and digital capsules. Vector-based neurons represent capsules, and the magnitude and direction of the vectors are used to represent the probability and pose attributes of the instance, respectively. The primary capsule uses convolution operations to convert features into vector capsules. The digital capsule, together with the dynamic routing, processes the output of the primary capsule to generate high-level capsules representing different classes of semantic attributes.
[0046] The primary capsule uses convolution operations to transform features into vector-like capsules. It contains multiple parallel small convolutional kernels, each 4×4×8 in size, for a total of 32 kernels. Each kernel extracts a set of features to form a capsule vector, with an example size of [32(×16), 8]. The digital capsule, in conjunction with dynamic routing, further processes the output of the primary capsule to generate higher-level capsules representing semantic attributes of different classes, with an example size of [n, 16], where n represents the number of classes.
[0047] The information transmission process of vector neurons in the capsule layer includes three parts: affine transformation, weighted summation, and nonlinear activation. Affine transformation encodes the abstract spatial relationship between the lower-level capsules and higher-level features. The input lower-level capsules are processed through the weight matrix W. ij Multiplying (i=1, 2, 3, …) yields high-level features; weighted summation utilizes the coupling coefficient c ij (i=1, 2, 3, …) The high-level features are weighted and summed to obtain the input vector required for the high-level capsule. The calculation formula is as follows:
[0048] In the formula, u i W represents the underlying capsule of the input. ij U is the weight matrix. j|i Represents high-level features, c ij It is a coupling coefficient that has the properties of non-negativity and summing to 1. j This is the input vector required for the high-level capsule; Nonlinear activation utilizes the squash function (·) to transform the vector The magnitude of j is normalized to the interval (0, 1) while keeping the vector direction unchanged, so that the final output high-level capsule v j The modulus is used to characterize the probability of the existence of an object entity:
[0049] In the formula, ||·|| represents the L2 norm, which is the magnitude of the vector, and ||·||2 represents the square of the L2 norm. The first part of the calculation formula compresses the vector, and the second part is the direction vector. The whole formula normalizes the vector to be consistent with the original direction. The main function of dynamic routing is to update and determine the coupling coefficient c. ij This addresses how to selectively transfer target feature information from lower-level capsules to higher-level capsules with appropriate weights. The dynamic routing operation includes: input lower-level capsule u i With weight matrix W i After multiplication, the high-level feature U is obtained. i Based on this, softmax normalization, vector weighted summation, vector squash (·), and coupling coefficient updates are performed, with the process variable b initialized to 0. The dynamic routing execution process is as follows: Figure 3 As shown, the inputs are u1 and u2; the high-level features are U1 and U2; the required input vector for the high-level capsule is: Process variables: a, b; Coupling coefficient: Output: v. Through multiple iterations, the high-level capsule is finally obtained. It can be seen that the larger the inner product of the process vector ai and the high-level feature Ui, the larger the coupling coefficient value obtained from the update. At this point, the lower-level capsule will transmit richer information to the higher-level capsule.
[0050] Step six involves classifying and reconstructing the digital capsule results. The classification process utilizes the L2 norm to process the digital capsule results, forming a vector with n elements. The index corresponding to the maximum value element in the vector is the category code, thereby determining the target class to which the classification result belongs. The reconstruction is achieved using a fully connected layer.
[0051] Steps two through five constitute the encoding stage, and this step is the decoding stage, mainly including classification and reconstruction. Reconstruction is primarily achieved using fully connected layers, with the corresponding data size change being
[512] ->
[1024] ->[11×11×C / 4]. The expression for the reconstruction operation is:
[0052] In the formula, x is the input vector data, W is the weight matrix, Reshape(·) represents the reshaping operation, and Or represents the reconstructed data; The loss function is used to constrain the training process of the classification network. A joint loss function is constructed that includes classification and reconstruction losses. The reconstruction loss plays a regularization role and calculates the sum of squares of the differences between the reconstruction output and the original input during the decoding stage.
[0053] In the formula, L c This represents the classification loss for a single capsule neuron, where c is the number of categories, and T is the classification loss. c It is an indicator function, T when the category is c. c The value is 1 otherwise; m+ and m- represent the upper and lower edge thresholds, respectively. The former penalizes the classifier's prediction of a class that does not actually exist, while the latter penalizes the classifier's prediction of a class that does not actually exist. m+ = 0.9, m- = 0.1; λ is the sparsity coefficient, used to adjust the weight, with a value of 0.5; Ltotal represents the total loss, I is the original input, and Or is the decoder's reconstructed output. The weighting factor has a value of 2 × 10. -5 .
[0054] Step 7: Accuracy assessment of the fine classification results of cultivated areas.
[0055] The detailed classification results of a certain multi-crop cultivation area are as follows: Figure 4 As shown. The evaluation metrics include Overall Accuracy (OA), Average Accuracy (AA), and Kappa coefficient (κ). The classification results are compared with ground reference data to obtain the confusion matrix M, and the evaluation metrics are calculated using the following formulas:
[0056] In the formula, n is the number of categories, N is the total number of test samples, and M is the total number of samples. uu M represents the number of samples in which class u is correctly identified as u. uvThis represents the number of samples where category v is identified as u, where u∈[1, n], v∈[1, n]. Based on the method proposed in this invention, the accuracy indicators OA of the refined classification results for cultivated areas are 99.19%, AA is 96.41%, and κ is 0.99.
[0057] In summary, this invention proposes a refined classification method for multi-crop cultivated areas that integrates attention mechanisms and capsule layers. By introducing lightweight convolutions in the encoding stage and combining spatial and channel attention mechanisms, the method effectively enhances the network's ability to extract and filter joint spatial and spectral features from hyperspectral remote sensing images, mitigating interference from complex backgrounds and fragmented plots. In the feature representation stage, vectorized neurons in capsule layers are used to model hierarchical spatial relationships, and classification and reconstruction losses are combined in the decoding stage. This enhances the model's discriminative power and classification generalization performance for multi-crop cultivated areas, achieving high-precision and refined classification.
[0058] This invention also provides a storage medium storing a computer program. When executed by a processor, the computer program implements some or all of the steps in various embodiments of the multi-crop cultivation area classification method integrating attention mechanism and capsule layer provided by this invention. The storage medium may be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0059] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or certain parts of the embodiments of the present invention.
[0060] The embodiments of the present invention described above do not constitute a limitation on the scope of protection of the present invention.
Claims
1. A method for classifying multi-crop cultivated areas by integrating attention mechanisms and capsule layers, characterized in that, include: Step 1: Acquire hyperspectral images of various crop cultivation areas, perform principal component analysis for dimensionality reduction and image segmentation on the hyperspectral images to obtain image blocks; Step 2: Perform spatial attention weighting on the image block; Step 3: Construct a lightweight convolutional module based on the output of Step 2; Step four: Perform channel attention weighting on the output of step three; Step 5: Construct the capsule layer based on the output of Step 4. The capsule layer includes primary capsules, dynamic routing, and digital capsules. Vector-based neurons represent capsules, and the magnitude and direction of the vectors are used to represent the probability and pose attributes of the existence of instances, respectively. The primary capsules use convolution operations to convert features into vector capsules. The digital capsules work with the dynamic routing to process the output of the primary capsules and generate high-level capsules that represent different classes of semantic attributes. Step 6: Classify and reconstruct the digital capsule results. The classification uses the L2 norm to process the digital capsule results, forming a vector with n elements. The index corresponding to the maximum value element in the vector is the target class. The reconstruction is achieved by a fully connected layer. Step 7: Accuracy assessment of the fine classification results of cultivated areas.
2. The method for classifying multi-crop cultivated areas by integrating attention mechanisms and capsule layers as described in claim 1, characterized in that, In step one, the principal component analysis dimensionality reduction includes: performing principal component analysis dimensionality reduction on the hyperspectral image to compress its band number to 1 / 4 of the original, and removing high redundancy and noise between hyperspectral image bands; the image segmentation includes: dividing the data after principal component analysis dimensionality reduction into image blocks with a length × width × number of channels of 11 × 11 × 64, which are used as input to the classification network.
3. The method for classifying multi-crop cultivated areas by integrating attention mechanisms and capsule layers as described in claim 1, characterized in that, In step two, the image patches are subjected to max pooling and mean pooling in the channel dimension, respectively. The pooling results are then concatenated in a stacked manner to form a concatenated feature of size 11×11×2. Convolution and activation operations are then performed sequentially on these features to generate spatial attention matrix weights with values ranging from 0 to 1. The spatial attention matrix weights are then element-wise multiplied with the original image patches to complete the weighting operation. The spatial attention weighting process is expressed by the following formula: In the formula, F represents the input feature data, AvgPool(·) and MaxPool(·) represent average pooling and max pooling, respectively, Conv(·) represents the convolution operation with a kernel size of 7×7, σ(·) represents the Sigmoid activation function, and O s The spatial attention matrix weights are obtained; the convolution operations involved are represented as follows: In the formula, f(·) represents the activation output of the β-th channel in convolutional layer l; f(·) represents the activation function, which is Sigmoid. It is the feature to be activated in the β-th channel of convolutional layer l, derived from the feature map of the previous layer. With the corresponding convolution kernel Perform convolution summation operations and, by adjusting the bias... Obtained by addition; It is used to obtain The input feature map subset; "*" represents the convolution operator.
4. The method for classifying multi-crop cultivated areas by integrating attention mechanisms and capsule layers as described in claim 1, characterized in that, In step three, based on the output of step two, a convolution operation is first performed to generate 11×11×128 intrinsic features. Then, a depthwise convolution operation is performed on these features. The calculation process is as follows: In the formula, Y represents the intrinsic features after regular convolution, u is the input feature, and the convolution kernel is... ∈ m×k×k×c; In the formula, i,j Represents the i-th eigenfeature y of the eigenfeature set Y. i Perform the j-th linear operation to generate the corresponding j-th enhanced feature y. ij The last layer i,s This indicates that by performing an identity mapping on the intrinsic features, the final set of enhanced features obtained is {y}. 11 , y 12 , …, y ms The relationship between the number of intrinsic features m, the number of linear operations s, and the number of final enhanced feature channels n is described as n = m·s (m ≤ n). Ignoring the last identity transformation, this relationship is further expressed as m(s-1) = n / s·(s-1).
5. The method for classifying multi-crop cultivated areas by integrating attention mechanisms and capsule layers as described in claim 1, characterized in that, In step four, channel attention operations are performed on the enhanced features output in step three. First, spatial compression is performed on the feature data, followed by average pooling on the spatial scale. Considering the information exchange between the current channel and its k neighboring channels, one-dimensional convolution is used to generate channel-dimensional attention weight values. These weights are then multiplied element-wise with the enhanced features to complete the channel-dimensional weighting operation. The calculation process is as follows: In the formula, y represents the input channel feature data, C1D(·) represents one-dimensional convolution, and σ(·) represents the Sigmoid activation function. The channel attention weights are obtained, where k represents the size of the one-dimensional convolution kernel, adaptively determined by the following formula: In the formula, C represents the number of channels of the input data, k represents the size of the calculated one-dimensional convolution kernel, odd indicates that the obtained value is odd, b=1, and γ=2.
6. The method for classifying multi-crop cultivated areas by integrating attention mechanisms and capsule layers as described in claim 1, characterized in that, In step five, the primary capsule uses convolution operations to convert features into capsules in vector form. It contains multiple parallel small convolution kernels with a kernel size of 4×4×8, totaling 32 kernels. Each convolution kernel extracts a set of features to form a capsule vector. The information transmission process of vector neurons in the capsule layer includes three parts: affine transformation, weighted summation, and nonlinear activation. Affine transformation encodes the abstract spatial relationship between low-level capsules and high-level features. The input low-level capsules are transformed through the weight matrix W. ij (i=1, 2, 3, …) are multiplied to obtain higher-level features; Weighted summation using coupling coefficient c ij (i=1,2,3,…) The high-level features are weighted and summed to obtain the input vector required for the high-level capsule. The calculation formula is as follows: In the formula, u i W represents the underlying capsule of the input. ij U is the weight matrix. j|i c represents a high-level feature. ij It is a coupling coefficient that has the properties of non-negativity and summing to 1. j This is the input vector required for the high-level capsule; Nonlinear activation utilizes the squash function (·) to transform the vector The magnitude of j is normalized to the interval (0, 1) while maintaining the vector direction, so that the magnitude of the final output high-level capsule vj can be used to characterize the probability of the existence of the object entity: In the formula, ||·|| represents the L2 norm, which is the magnitude of the vector, and ||·||2 represents the square of the L2 norm. The first part of the calculation formula is to compress the vector, and the second part is the direction vector. The whole formula achieves a normalization operation on the vector to be consistent with the original direction. Dynamic routing is used to update and determine the coupling coefficient c. ij The dynamic routing process includes: the input underlying capsule u i With weight matrix W i After multiplication, the high-level feature U is obtained. i Based on this, softmax normalization, vector weighted summation, vector squash (·), and coupling coefficient update operations are performed, where the process variable b is initialized to 0; through multiple iterations, the high-level capsule is finally obtained.
7. The method for classifying multi-crop cultivated areas by integrating attention mechanisms and capsule layers as described in claim 1, characterized in that, In step six, the expression for the reconstruction operation is: In the formula, x is the input vector data, W is the weight matrix, Reshape(·) represents the reshaping operation, and Or represents the reconstructed data; The loss function is used to constrain the training process of the classification network. A joint loss function is constructed that includes classification and reconstruction losses. The reconstruction loss plays a regularization role and calculates the sum of squares of the differences between the reconstruction output and the original input during the decoding stage. In the formula, L c This represents the classification loss for a single capsule neuron, where c is the number of categories, and T is the classification loss. c It is an indicator function, T when the category is c. c The value is 1 otherwise; m+ and m- represent the upper and lower edge thresholds, respectively. The former penalizes the classifier's prediction of a class that does not actually exist, while the latter penalizes the classifier's prediction of a class that does not actually exist. m+ = 0.9, m- = 0.1; λ is the sparsity coefficient, used to adjust the weight, with a value of 0.5; Ltotal represents the total loss, I is the original input, and Or is the decoder's reconstructed output. The weighting factor has a value of 2 × 10. -5 .
8. The method for classifying multi-crop cultivated areas by integrating attention mechanisms and capsule layers as described in claim 1, characterized in that, In step seven, the evaluation metrics include overall accuracy, average classification accuracy, and the Kappa coefficient κ. The classification results are compared with ground reference data to obtain the confusion matrix M, and the evaluation metrics are calculated using the following formulas: In the formula, n is the number of categories, N is the total number of test samples, and M is the total number of samples. uu M represents the number of samples in which class u is correctly identified as u. uv This represents the number of samples whose category v is identified as u, where u∈[1, n], v∈[1, n].
Citation Information
Patent Citations
Low-illumination image classification method based on attention mechanism and capsule network
CN111950649A
Hyperspectral image classification method combining deep capsule network and Markov random field
CN113128370A
Hyperspectral image classification method based on spectral space attention fusion and deformable convolutional residual network
CN113361485A
Sugarcane disease identification method based on attention mechanism residual capsule network
CN115565168A
Image classification method based on improved capsule network
CN116246110A