A method for constructing a crop planting structure extraction model based on a Kolmogorov-Arnold network

CN122799262APending Publication Date: 2026-09-22ZHENGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610854223.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-12
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

传统水体分割方法,例如归一化差异水体指数(NDWI),以及支持向量机(SVM)和随机森林等机器学习模型,主要依赖光谱特征的统计分析,难以有效处理地物间的复杂同谱异物与混合像元问题,尤其在面对高分辨率影像中的细节模糊、小型水体漏分以及复杂背景干扰时,分割精度显著下降

Benefits of technology

[0020]本发明的有益效果是:1、本发明通过KAN网络增强非线性表征能力、RBF 核激活卷积强化特征提取,在公开高光谱数据集上取得极具竞争力的分类精度。相较于原始HybridSN 模型实现明显提升,能够为农作物类型识别、种植结构提取提供高置信度、高鲁棒性的识别结果,降低误分、漏分概率,满足农业遥感高精度监测需求;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799262A_ABST
    Figure CN122799262A_ABST
Patent Text Reader

Abstract

This invention discloses a method for constructing a crop planting structure extraction model based on the Kolmogorov–Arnold network, comprising the following steps: step (1) preprocessing hyperspectral remote sensing images; step (2) constructing a three-dimensional kernel activation convolution module activated by radial basis function (RBF); step (3) constructing a two-dimensional spatial convolution fusion module; and step (4) constructing a classifier based on KANLinear layers. This invention enhances nonlinear representation capabilities through the KAN network and strengthens feature extraction through RBF kernel activation convolution, achieving highly competitive classification accuracy on publicly available hyperspectral datasets. Compared to the original HybridSN model, it achieves a significant improvement, providing high-confidence and robust identification results for crop type recognition and planting structure extraction, reducing the probability of misclassification and omission, and meeting the high-precision monitoring needs of agricultural remote sensing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of hyperspectral remote sensing image processing and precision agriculture, specifically to a method for constructing a crop planting structure extraction model based on the Kolmogorov–Arnold network. Background Technology

[0002] With the deepening development of refined water resource management, high-precision and high-efficiency water body segmentation has become a core requirement for water body segmentation tasks in high-resolution remote sensing images. Traditional water body segmentation methods, such as the Normalized Difference Water Index (NDWI) and machine learning models such as Support Vector Machines (SVM) and Random Forests, mainly rely on statistical analysis of spectral features. They are difficult to effectively handle complex issues of homonymous and heteronymous objects and mixed pixels among ground features, especially when faced with detail blurring, missing small water bodies, and complex background interference in high-resolution images, resulting in a significant decrease in segmentation accuracy. Although deep learning models such as U-Net and SegFormer have improved segmentation performance in general scenarios in recent years by mining deep semantic features, their inherent static multi-scale fusion strategies still have shortcomings such as insufficient adaptive perception capabilities and inadequate integration of detailed features and contextual information, failing to meet the requirements of large-scale monitoring for refined and robust water body information extraction.

[0003] While existing improved models based on multi-scale feature fusion have made progress in complex scene segmentation, they still have limitations in water body extraction applications: First, they struggle to effectively mitigate the loss of detail and blurred boundaries caused by low-resolution source data or complex interference; second, segmentation results often result in the omission or misclassification of small water bodies; and third, they largely rely on preset or static fusion strategies, making it difficult to dynamically and adaptively integrate features based on the actual scale distribution and morphological characteristics of the water bodies in the input image. These problems severely restrict the reliability of water body segmentation results in key scenarios such as water resource management and climate change research.

[0004] Therefore, developing a novel segmentation model that can adaptively fuse multi-scale features and enhance detail preservation is a pressing technical challenge in the field of remote sensing image processing and water segmentation.

[0005] With the rapid development of precision agriculture and smart farmland, the rapid and accurate extraction of crop planting structure has become a key technological support for farmland monitoring, yield estimation, arable land management, and food security early warning. Hyperspectral remote sensing, with its advantages of high spectral resolution, continuous bands, and rich information on ground texture and biochemical characteristics, has become the mainstream technology for fine classification of crops and extraction of planting structure.

[0006] Currently, crop classification methods based on hyperspectral imagery mainly rely on traditional machine learning and deep learning models. Among them, the HybridSN model, a classic spectral-spatial joint classification network, extracts spectral features through 3D convolution and spatial features through 2D convolution, and has been applied in hyperspectral image classification tasks. However, in complex real-world farmland scenarios, existing methods still have significant technical limitations:

[0007] Hyperspectral data exhibits complex nonlinear spectral distributions and class aliasing. Traditional convolutional and fully connected layers are insufficient in representing nonlinear boundaries, making it difficult to accurately distinguish crops with similar spectra. To improve accuracy, existing models generally introduce attention mechanisms, leading to complex model structures, increased parameter counts, and reduced training and inference efficiency. Classification results commonly suffer from coarse boundaries and poor spatial consistency. Traditional fully connected layers, with their fixed linear transformations, cannot adaptively fit the fine decision boundaries of complex farmland plots, resulting in classification accuracy and visualization effects that fail to meet the needs of precision agricultural management.

[0008] This paper proposes a method for constructing a crop planting structure extraction model based on the Kolmogorov–Arnold network to address the aforementioned problems. Summary of the Invention

[0009] The purpose of this invention is to provide a method for constructing a crop planting structure extraction model based on the Kolmogorov-Arnold network. The resulting crop planting structure extraction model does not require the introduction of an attention mechanism. Instead, it enhances the model's ability to express and discriminate high-order nonlinear features through nonlinear activation functions and learnable basis functions, thus more effectively addressing the challenges of crop classification under complex spectral spatial structures.

[0010] The objective of this invention is achieved as follows:

[0011] A method for constructing a crop planting structure extraction model based on the Kolmogorov–Arnold network includes a crop planting structure extraction model based on the Kolmogorov–Arnold network. The Kolmogorov–Arnold network-based crop planting structure extraction model includes a three-dimensional kernel activation convolution module activated by radial basis function (RBF), a two-dimensional spatial convolution fusion module, and a classifier built based on KANLinear layers. First, the spectral-spatial joint features are extracted by embedding the three-dimensional convolution module activated by radial basis function. Then, the spatial context representation is enhanced by the two-dimensional spatial convolution fusion module. Finally, the classification decision is completed by the classifier built based on KANLinear layers.

[0012] The method for constructing a crop planting structure extraction model based on the Kolmogorov–Arnold network includes the following steps: (1) Preprocessing the hyperspectral remote sensing image, and using the preprocessed data as the input to the subsequent three-dimensional kernel activation convolution module; Since hyperspectral remote sensing data usually contains hundreds of continuous spectral bands, there is strong band correlation and information redundancy. Directly inputting it into the model will lead to increased computational complexity and is prone to the "curse of dimensionality" problem. Therefore, this method uses principal component analysis (PCA) to reduce the dimensionality of the hyperspectral data. In the PCA dimensionality reduction process, the original three-dimensional hyperspectral data array is first reshaped into a two-dimensional array to meet the input format requirements of the PCA algorithm; then, covariance analysis and eigenvalue decomposition are performed on the spectral vector of each pixel, and the top 30 principal components are selected as new feature representations based on the size of the eigenvalues, thereby effectively removing redundant bands and noise interference while retaining the main spectral information;

[0013] Step (2) Construct a three-dimensional kernel activation convolution module activated by radial basis function (RBF). The three-dimensional kernel activation convolution module activated by radial basis function (RBF) is used to enhance the discriminativeness and adaptability of hyperspectral data feature extraction.

[0014] Step (3) Construct a two-dimensional spatial convolutional fusion module. While retaining the ability to perceive spatial structure, the two-dimensional spatial convolutional fusion module realizes the gradual transformation of feature representation from three-dimensional spectral space to two-dimensional deep semantic space, providing a highly discriminative feature foundation for the subsequent classification decision layer.

[0015] Step (4) Construct a classifier based on KANLinear layers. The classifier based on KANLinear layers projects high-dimensional spectral-spatial features to a low-dimensional discriminative space through hierarchical nonlinear transformation. In each layer of the network, the features are used to extract discriminative information through KAN linear layers and Swish activation functions. Finally, the processed features are input into the classifier to complete the classification task.

[0016] The specific operation of step (1) is as follows: Principal component analysis is used for dimensionality reduction. In the PCA dimensionality reduction process, the original three-dimensional hyperspectral data array is first reshaped into a two-dimensional array to meet the input format requirements of PCA transformation; then PCA transformation is performed on the spectral information of each pixel, and the first 30 principal components are selected as new feature vectors. Before performing PCA transformation, the data is standardized to ensure that the mean of each feature is 0 and the standard deviation is 1; let the original data matrix be... Where m is the number of samples and n is the number of features; the standardization formula is as follows: (1) In the formula, is the standardized value of the element in the i-th row and j-th column, X ij μ is the original data value. j Let σ be the mean of the j-th feature. j Let be the standard deviation of the j-th feature; After data standardization, the covariance matrix C is calculated using the following formula: (2) In the formula, X′ is the standardized dataset matrix; Perform eigenvalue decomposition on the covariance matrix C to obtain the eigenvalues ​​λ. i and its corresponding eigenvector v i : (3) In the formula, λ i For the i-th eigenvalue, v i These are the corresponding eigenvectors; the magnitude of the eigenvalues ​​reflects the variance of the data along the direction of the corresponding eigenvector. The goal of PCA dimensionality reduction is to map the original n-dimensional high-dimensional data to a lower-dimensional k-dimensional space. Specifically, the eigenvalues ​​of the covariance matrix are first sorted in descending order, and the eigenvectors corresponding to the k largest eigenvalues ​​are selected to construct the projection matrix. ; The normalized data matrix X′ is then projected onto the projection matrix W to obtain the dimensionality-reduced data matrix Xk: (4) In hyperspectral remote sensing image processing, fixed-size image patches are extracted around each pixel to capture the local spatial structure information of the pixel. A two-pixel edge is filled around each pixel to ensure that each pixel is contained within a 25×25 window, which simultaneously covers the target pixel and its neighboring pixel information. Based on this method, the image patches extracted from the entire image contain both the spectral information of each pixel and its local spatial neighborhood information, providing rich input data for subsequent deep learning models.

[0017] According to the method for constructing a crop planting structure extraction model based on Kolmogorov-Arnold network as described in claim 1, the specific operation of step (2) is as follows: a three-dimensional kernel-activated convolutional module activated by radial basis function (RBF) is used as input, which is the hyperspectral feature data obtained in step (1) after standardization, mean removal preprocessing and PCA dimensionality reduction. , where B is the number of spectral bands after PCA dimensionality reduction, and H and W are the spatial height and width of the hyperspectral image, respectively; First, the dimensionality-reduced hyperspectral data is input into a three-dimensional convolutional structure, using... A three-dimensional convolutional kernel of a certain size is used to simultaneously extract local neighborhood features in the spectral dimension B, spatial height dimension H, and spatial width dimension W, capturing spectral-spatial joint correlation information and obtaining initial spectral-spatial joint features. The formula for 3D convolution is: (5) In the formula: k represents the weights of the 3D convolution kernel. s k h k w These represent the dimensions of the convolution kernel in the spectral, height, and width dimensions, respectively; Cout is the number of output feature channels; b 3D is the 3D convolution bias term, i and j are the row and column coordinates of the spatial features, respectively, and c is the output feature channel index; Secondly, a radial basis function (RBF) activation mechanism is introduced during the 3D convolution process. By replacing the traditional fixed linear mapping with a learnable kernel function, the Hybrd-KANet model can construct flexible nonlinear feature mapping relationships in the spectral-spatial joint domain. The RBF activation formula is as follows: (6) In the formula: This is the learnable bandwidth parameter for the RBF kernel function. For the learnable center parameters of the RBF kernel, For the Euclidean norm, this formula is adjusted... and Adaptive fitting of the spectral-spatial nonlinear characteristic distributions of different crops; Then, after each level of 3D convolution, a batch normalization (BN) layer and a Swish activation function are sequentially set. Batch normalization is used to stabilize the feature distribution and accelerate the model convergence speed. The batch normalization calculation formula is as follows: (7) In the formula: The mean of the features, The variance of the feature To prevent tiny constants with a denominator of 0, , The learnable scaling and offset parameters are batch-normalized; the Swish activation function is used to enhance the nonlinear representation of features and improve gradient propagation, and its expression is: (8) In the formula: The Sigmoid activation function is used to map feature values ​​to the [0,1] interval, achieving non-linear activation while alleviating the gradient vanishing problem; Finally, a progressive spectral dimension compression strategy is adopted, gradually reducing the kernel size along the spectral axis in different 3D convolutional layers. This involves starting with a larger spectral receptive field to capture coarse-grained global spectral responses, gradually transitioning to a smaller spectral receptive field to mine fine-grained spatial texture features, forming a multi-level feature abstraction system, and ultimately outputting deep spectral-spatial joint features. C 3D To determine the final output feature channel number, the deep spectral-spatial joint features are directly used as the input to the two-dimensional spatial convolution fusion module.

[0018] The method for constructing a crop planting structure extraction model based on the Kolmogorov-Arnold network according to claim 1 is characterized in that the specific operation of step (3) is as follows: the input of the two-dimensional spatial convolution fusion module is the high-dimensional spectral-spatial joint features output by the three-dimensional kernel-activated convolution module activated by radial basis function (RBF). C 3D H represents the number of feature channels output by the 3D kernel-activated convolutional module activated by the radial basis function (RBF), and H and W represent the spatial dimensions. First, the high-dimensional features output by the 3D kernel-activated convolution module are reorganized according to the channel dimension, merging the spectral dimension with the channel dimension to eliminate the dimensional limitation of the 3D features on the 2D convolution. The feature dimension transformation formula after reorganization is as follows: (9) In the formula: This represents the number of spectral dimensions remaining after progressive compression, for example, reducing the original channel dimensions to... The characteristics, recombine and unfold into This makes it a high-dimensional channel feature that can be processed by two-dimensional convolution; 32 is the number of channels and 18 is the remaining spectral dimension. Subsequently, a 3×3 two-dimensional convolution kernel is used to perform local neighborhood modeling in the spatial dimension, while simultaneously completing cross-channel feature fusion, compressing high-dimensional channel features to 64 channels. The two-dimensional convolution operation formula is as follows: (10) In the formula: b represents the weights of the two-dimensional convolution kernel. 2D For two-dimensional convolution bias terms, For a two-dimensional convolution operator, the final output is... ; Next, batch normalization (BN) and nonlinear activation functions are used to further stabilize the feature distribution and enhance spatial semantic representation. The computation process is consistent with the batch normalization and Swish activation in the 3D kernel activation convolution module with radial basis function (RBF) activation, i.e.: (11) (12) Finally, output a two-dimensional deep semantic feature map. This process preserves the spatial structure, edge texture, and local contextual information of crop regions while converting three-dimensional spectral-spatial features into two-dimensional deep semantic features, providing highly discriminative input features for the subsequent KANLinear classifier.

[0019] The method for constructing a crop planting structure extraction model based on the Kolmogorov-Arnold network according to claim 1 is characterized in that the specific operation of step (4) is as follows: a classifier based on the KANLinear layer is used to perform nonlinear classification mapping on the deep semantic features output by the two-dimensional spatial convolution fusion module to achieve accurate determination of crop categories. First, the two-dimensional feature map output by the two-dimensional spatial convolution fusion module is... Flatten the vector to convert it into a one-dimensional feature vector. ,in The flattened feature dimensions are represented by the following formula: (13) Subsequently, this one-dimensional feature vector is input into the KANLinear layer. Through the learnable one-dimensional function mapping mechanism in the Kolmogorov-Arnold network (KAN), the input features are nonlinearly combined and approximated by an adaptive function. The core operation formula of the KANLinear layer is as follows: (14) In the formula: M is the number of learnable basis functions in the KANLinear layer. For the one-dimensional learnable basis functions of the KAN network, The learnable weight vector corresponding to each basis function For the bias term of the KANLinear layer, this formula achieves a complex nonlinear mapping of input features through a linear combination of multiple sets of learnable basis functions; Then, the KANLinear layer is used to replace the traditional fully connected layer, so that the classifier can not only perform linear weighting, but also learn more complex nonlinear decision boundaries, in order to overcome the shortcomings of the traditional fully connected layer in linear fitting ability and improve the accuracy of crop category discrimination. Finally, the features output by the KANLinear layer are... The input / output layer obtains classification scores for each crop category. Where K is the total number of crop categories, and is converted into a category probability distribution using the Softmax function. The Softmax calculation formula is: (15) In the formula: S k Let P(k) be the classification score for the k-th crop category, and P(k) be the probability that the pixel belongs to the k-th crop category. The category with the highest probability is selected as the final category for the pixel, i.e.: (16) By determining the category pixel by pixel, the final result of the extraction of crop planting structure in the whole area is generated, realizing the identification of crop types and the accurate division of planting areas.

[0020] The beneficial effects of this invention are: 1. This invention enhances nonlinear representation capabilities through KAN networks and strengthens feature extraction through RBF kernel-activated convolutions, achieving highly competitive classification accuracy on publicly available hyperspectral datasets. Compared to the original HybridSN model, it achieves a significant improvement, providing high-confidence and robust identification results for crop type recognition and planting structure extraction, reducing the probability of misclassification and omission, and meeting the high-precision monitoring needs of agricultural remote sensing;

[0021] 2. To address the pain points of high spectral dimensionality, complex category distribution, and difficulty in characterizing nonlinear relationships in hyperspectral data, this invention adopts a combination of KAN network and RBF kernel function, which enables the model to have stronger nonlinear function fitting and boundary approximation capabilities. It can accurately capture subtle spectral differences and spatial texture differences between different crops, significantly alleviate the problem of spectral variation of the same type of crop and spectral confusion of different types of crops, and enable the model to maintain stable and reliable recognition performance in complex farmland environments and fragmented plots.

[0022] 3. This invention completely replaces the traditional fully connected layer with a KANLinear layer. Through a dual-branch structure of a baseline path and a spline path, it achieves global feature mapping and adaptive fitting of local features, effectively overcoming the limitations of traditional fully connected layers, such as limited expressive power, susceptibility to overfitting, and coarse decision boundaries. The model exhibits stronger generalization ability for new regions, new seasons, and data acquired from different sensors, demonstrates stronger cross-scene transfer capabilities, and has a wider range of practical applications.

[0023] 4. Existing hyperspectral classification models generally rely on attention mechanisms to improve accuracy, resulting in a large number of parameters, unstable training, and increased inference time. This invention does not rely on attention mechanisms; it achieves high-precision classification solely through kernel function activation and KAN structure improvements. This simplifies the network structure, reduces redundant computation, makes model training more stable, converges faster, and inference speed more controllable, facilitating subsequent embedded deployment and real-time interpretation on UAVs.

[0024] 5. This invention can be used for crop planting structure extraction, precise crop type identification, and planting area monitoring. The output results can provide accurate, reliable, and stable data support for agricultural management departments, planting entities, and research institutions, promoting the digital, intelligent, and efficient development of precision agriculture, and has significant practical value and socio-economic benefits. Attached Figure Description

[0025] Figure 1 This is a diagram showing the overall structure of Hybrid-KANet, the crop planting structure extraction model based on the Kolmogorov–Arnold network, as described in this invention.

[0026] Figure 2 This is a diagram of the three-dimensional kernel-activated convolutional layer structure of the present invention;

[0027] Figure 3 This is a schematic diagram of the KAN linear layer structure of the present invention;

[0028] Figure 4 A comparison of the overall average boundary curvature of different classification models on the IP and LK datasets;

[0029] Figure 5 The region classification error distribution plots for the IP and LK datasets are shown; (a) IP dataset; (b) LK dataset. Detailed Implementation

[0030] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0031] like Figures 1 to 3 As shown, a method for constructing a crop planting structure extraction model based on the Kolmogorov-Arnold network is presented. This includes a Hybrid-KANet model for crop planting structure extraction based on the Kolmogorov-Arnold network. The model comprises a 3D kernel-activated convolutional module activated by radial basis functions (RBF), a 2D spatial convolutional fusion module, and a classifier built based on KANLinear layers. First, spectral-spatial joint features are extracted by embedding the 3D convolutional module activated by radial basis functions. Then, the 2D spatial convolutional fusion module enhances the spatial context representation. Finally, the classifier built based on KANLinear layers completes the classification decision. Evaluation of the Kolmogorov-Arnold network-based crop planting structure extraction model:

[0032] To comprehensively evaluate the performance of the proposed crop extraction method, this invention employs multiple standard evaluation metrics, covering overall classification accuracy, inter-class accuracy, model stability, and adaptability to complex scenarios. Specifically, the evaluation metrics include overall accuracy (OA), average accuracy (AA), Kappa coefficient (Kappa), and mean intersection-over-union ratio (mIoU). These metrics effectively reflect the model's performance in hyperspectral remote sensing data classification tasks from multiple perspectives.

[0033] First, OA is a commonly used metric for evaluating classifier performance, defined as the proportion of correctly classified samples. Its calculation formula is as follows:

[0034] (17)

[0035] In the formula, C represents the total number of categories in the hyperspectral remote sensing image. For the i-th category, TP i FP is the true number of cases in this class (the number of samples correctly predicted as class i). i TN is the number of false positives in this class (the number of samples that actually belong to other classes but are incorrectly predicted as class i). i FN is the true negative number of this class (the number of samples that do not actually belong to class i but are correctly predicted as not belonging to class i). i This represents the number of false negatives in this class (the number of samples that actually belong to class i but are incorrectly predicted as other classes).

[0036] To further evaluate the classification accuracy for each category, single-class accuracy and average (AA) are introduced. Single-class accuracy is used to quantify the classification accuracy of a single category, and the calculation formula is as follows:

[0037] (18)

[0038] In the formula, Acc i This represents the classification accuracy of the i-th class.

[0039] AA is defined as the mean of the precision across all categories, and is calculated using the following formula:

[0040] (19)

[0041] In addition, the Kappa coefficient κ is used to measure the consistency between the classification result and random classification, and its calculation formula is as follows:

[0042] (20)

[0043] In the formula, P e The expected accuracy in the case of random guessing is calculated using the following formula:

[0044] (twenty one)

[0045] Where N is the total number of samples. The Kappa coefficient ranges from [−1, 1]: when κ = 1, the model's prediction is completely consistent with the true label. When κ = 0, the model's prediction is no different from random guessing. When κ < 0, the model's prediction is worse than random guessing.

[0046] Intersection over Union (IoU) is a core metric in image classification tasks, used to measure the spatial accuracy of predictions for each class. It involves first calculating the IoU for each class, then averaging the IoUs across all classes. This metric is particularly important in hyperspectral remote sensing land cover classification tasks. First, the IoU for the i-th class is calculated. i :

[0047] (twenty two)

[0048] Therefore, the average intersection-union ratio (mIoU) is defined as follows:

[0049] (twenty three)

[0050] The above evaluation metrics construct a comprehensive framework for assessing model performance in hyperspectral crop extraction tasks, which can effectively reveal the advantages and disadvantages of the proposed methods, and are particularly suitable for performance analysis under class imbalance and complex background conditions.

[0051] After training and prediction, based on the model evaluation metrics mentioned above, the experimental results of this model and the current mainstream models on the IndianPines (IP) dataset are shown in Table 1.

[0052] Table 1. Performance comparison of different classification algorithms on the IP dataset

[0053]

[0054] Table 1 shows a comparison of evaluation metrics for different algorithms on the IP dataset. This invention outperforms other comparative methods in Kappa coefficient, OA, and AA. Specifically, the Kappa coefficient of this invention reaches 99.11%, the overall accuracy reaches 99.22%, and the average accuracy reaches 98.16%, all higher than models such as U-Net, 3D-CNN, and HybridSN, highlighting its significant advantages in classification accuracy and robustness. Regarding the spatial consistency evaluation metric, the mIoU of this invention reaches 97.03%, better than U-Net's 95.85% and other comparative methods, indicating that this invention has a significant effect on spatial segmentation quality. In terms of training time, the training time of this invention is 8.17 minutes. In comparison, DGFNet and U-Net have significantly longer training times and perform poorly in accuracy and spatial metrics.

[0055] The experimental results comparing the model of this invention with the current mainstream models on the WHU-Hi-LongKou (LK) dataset are shown in Table 2.

[0056] Table 2 Performance comparison of different classification algorithms on the LK dataset

[0057]

[0058] Table 2 shows a comparison of evaluation metrics for different algorithms on the LK dataset. This invention demonstrates superior performance across all evaluation metrics, showing significant improvements over other benchmark and state-of-the-art methods, fully proving its superiority in comprehensive classification accuracy, class balance, and spatial consistency. Furthermore, the training time of this invention is 9.22 minutes, achieving an effective balance between performance and computational cost.

[0059] Compared to this invention, the HybridSN model performs slightly worse, revealing its limitations in capturing fine-grained spatial patterns. The 3D-CNN and U-Net models are insufficient in overall accuracy. Due to limited spatial modeling capabilities, the MLP and 1D-CNN models perform significantly worse. The DGFNet model has the longest training time of all models. The MDvT model performs relatively evenly, but its mean intersection-over-union score is still low. The DiffFormer model has moderate overall performance, but its overall accuracy is still slightly lower than the optimal model.

[0060] (2) Average boundary curvature analysis of the model

[0061] This invention explores the geometric characteristics of classification boundaries generated by different hyperspectral image classification models by analyzing the average boundary curvature. For example... Figure 4 As shown, on both the IP and LK datasets, the overall curvature of the proposed Hybrid-KANet model is consistently lower than that of the benchmark models. Specifically, on the IP dataset, Hybrid-KANet achieves the lowest curvature value of 0.0063, outperforming traditional models such as MLP and 1D-CNN. Similarly, on the LK dataset, due to the more complex scene structure, the overall curvature of all models is generally higher, but Hybrid-KANet still exhibits strong performance with a curvature value of 0.0250, lower than MLP and 1D-CNN models. These results indicate that the kernel adaptive function in Hybrid-KANet helps generate smoother and more coherent decision boundaries.

[0062] Furthermore, the performance of the CNN-Transformer and DiffFormer models deserves special attention. As shown in Figure 4, their overall boundary curvatures on the IP dataset reach 0.0320 and 0.0515, respectively, with DiffFormer exhibiting the highest curvature among all models. These results indicate that although both models incorporate the Transformer architecture to enhance feature representation capabilities, they still have significant shortcomings in terms of boundary smoothness.

[0063] To further explore the model's characteristics, this invention conducted classification curvature analysis, and the results are shown in Table 3. Hybrid-KANet's boundary curvature values ​​across multiple categories approached the minimum level, indicating superior smoothness of its classification boundaries. In complex categories such as Corn and Corn-notill, Hybrid-KANet significantly outperformed the MLP and HybridSN models. In relatively simple or spatially continuous categories such as Wheat and Broad-leaf Soybean, Hybrid-KANet also maintained its competitiveness. For challenging crop categories such as Cotton and Sesame in the LK dataset, Hybrid-KANet also exhibited stable curvature values.

[0064] Table 3. Quantitative Comparison of Boundary Mean Curvature of Eight Classification Models on Typical Categories of the IP and LK Datasets

[0065]

[0066] The above results demonstrate that the network architecture enhanced by KAN can effectively capture spectral-spatial features. Smoother classification boundaries imply better visual coherence, a characteristic that enhances the model's application value in downstream tasks requiring spatial continuity, such as land cover mapping and agricultural monitoring. Hybrid-KANet's continued advantage in class boundary curvature metrics confirms the technical advantages of embedding prior adaptive features within a hyperspectral deep classification framework.

[0067] (3) Regional error statistics based on local grid division

[0068] To more comprehensively evaluate the classification performance of the proposed model at spatial scales, this invention proposes a regional error statistical method based on local grid partitioning. For the IP dataset with a spatial resolution of 20m, the method divides the image into grid cells of approximately 100m² and calculates the classification error within each cell. Experimental results show that the average classification error rate on this dataset is 0.0080, meaning that only about 0.8% of pixels are misclassified on average within each 100m² region. Furthermore, for the LK dataset with a spatial resolution of 0.463m, the method further employs a finer scale (approximately 10m² per cell, corresponding to 21×21 pixels) to conduct error analysis. The results show that the average classification error rate of the proposed model on this dataset is only 0.0016, meaning that only about 0.16% of pixels are misclassified within each 10m² region.

[0069] like Figure 5 As shown, this invention uses heatmaps to visualize the classification errors of the two datasets, intuitively presenting the spatial distribution characteristics of classification performance in different regions. The above results effectively verify the adaptability and accuracy of the Hybrid-KANet model of this invention in processing multi-source heterogeneous hyperspectral data.

[0070] (4) Ablation experiment of basis functions in Hybrid-KANet model

[0071] To verify the effectiveness of kernel function selection in the model and the classification performance of nonlinear kernel functions, this invention conducted a systematic comparative experiment on the IP dataset and the LK hyperspectral dataset, testing various kernel functions. The experiment employed a controlled variable method: fixing the network depth, regularization strategy, and training process, only changing the kernel function type of the 3D kernel activation convolution. Five core metrics were selected for performance evaluation: Kappa coefficient, overall accuracy (OA), average accuracy (AA), and mean intersection-over-union ratio (mIoU).

[0072] Table 4 presents the performance evaluation results of different kernel functions on the IP and LK hyperspectral datasets, showcasing the conclusions of the kernel function ablation experiments. In the complex crop distribution scenario of the IP dataset, RBF exhibits a significant advantage: its Kappa coefficient reaches 99.11%, 4.59% higher than the Fourier kernel. The average accuracy is 98.16%, slightly lower than Matern (v=1.5). The most significant performance improvement is in mIoU, with RBF achieving 97.03%, higher than the Matern (v=1.5) kernel function.

[0073] In large-scale structured farmland scenarios on the LK dataset, the performance gap between various kernel functions has narrowed. Nevertheless, RBF still maintains its leading position, with a Kappa coefficient of 99.83%, which is 0.2% higher than B-Spline. Notably, the Fourier kernel achieves an mIoU of 98.98% on the LK dataset, demonstrating its adaptability to high-resolution agricultural landscapes.

[0074] Table 4 Performance metrics of different kernel functions on the IP and LK datasets

[0075]

[0076] In summary, the advantages of this invention are as follows: (1) Addressing the problems of insufficient nonlinear expression and limited decision boundary fitting ability in traditional hyperspectral crop classification models, this invention breaks through the limitations of the existing HybridSN framework and introduces the Kolmogorov-Arnold network (KAN) into the field of hyperspectral remote sensing crop classification for the first time, forming a Hybrid-KANet integrated network structure. This architecture breaks away from the dependence of traditional neural networks on fixed activation functions and linear weights, and achieves stronger global function fitting ability through learnable continuous functions, fundamentally improving the model's modeling level of complex farmland spectral features, and providing a new network foundation for high-precision crop classification.

[0077] (2) In the spectral-spatial feature extraction stage, this invention designs a three-dimensional fast kernel activation convolution module with RBF radial basis function activation, which directly enhances the ability of the convolutional layer to express complex nonlinear relationships in hyperspectral data through nonlinear kernel mapping. This module can adaptively learn local feature changes in the spectral dimension, significantly improve feature discriminativeness without introducing an attention mechanism, and reduce model complexity and computational overhead, thus solving the problem of slow inference speed caused by existing models relying on attention.

[0078] (3) To improve the spatial continuity and boundary smoothness of crop classification, this invention adds a two-dimensional spatial convolution fusion module, which flattens the high-dimensional spectral-spatial features output by three-dimensional convolution along the spatial dimension and performs cross-channel fusion, thereby enhancing structural information such as spatial texture and regional connectivity while reducing dimensionality. This module can effectively suppress noise, making the final classification results more consistent with the spatial distribution characteristics of farmland plots, and significantly improving the mapping effect and practicality.

[0079] (4) To address the shortcomings of traditional fully connected layers, such as their inability to fit fine decision boundaries and their tendency to cause classification confusion, this invention uses KANLinear layers to replace all fully connected layers, constructing a KAN-based kernel activation classifier. This classifier adopts a dual-path structure of baseline path and spline path: the baseline path achieves global nonlinear feature mapping through Swish activation, ensuring global consistency. The spline path achieves local feature adaptive calibration through learnable B-spline basis functions, improving the ability to distinguish edge pixels and confused categories, making the classification decision more consistent with the actual distribution patterns of farmland features.

Claims

1. A method for constructing a crop planting structure extraction model based on the Kolmogorov–Arnold network, characterized in that, This includes a crop planting structure extraction model based on the Kolmogorov–Arnold network. The model comprises a 3D kernel-activated convolutional module activated by radial basis function (RBF), a 2D spatial convolutional fusion module, and a classifier built based on KANLinear layers. First, the spectral-spatial joint features are extracted by embedding the 3D convolutional module activated by radial basis function. Then, the spatial context representation is enhanced by the 2D spatial convolutional fusion module. Finally, the classification decision is made by the classifier built based on KANLinear layers. The method for constructing a crop planting structure extraction model based on Kolmogorov–Arnold network includes the following steps: Step (1) Preprocessing the hyperspectral remote sensing image, and using the preprocessed data as the input to the subsequent three-dimensional kernel activation convolution module; Step (2) Construct a three-dimensional kernel activation convolution module activated by radial basis function (RBF). The three-dimensional kernel activation convolution module activated by radial basis function (RBF) is used to enhance the discriminativeness and adaptability of hyperspectral data feature extraction. Step (3) Construct a two-dimensional spatial convolutional fusion module. While retaining the ability to perceive spatial structure, the two-dimensional spatial convolutional fusion module realizes the gradual transformation of feature representation from three-dimensional spectral space to two-dimensional deep semantic space, providing a highly discriminative feature foundation for the subsequent classification decision layer. Step (4) Construct a classifier based on KANLinear layers. The classifier based on KANLinear layers projects high-dimensional spectral-spatial features to a low-dimensional discriminative space through hierarchical nonlinear transformation. In each layer of the network, the features are used to extract discriminative information through KAN linear layers and Swish activation functions. Finally, the processed features are input into the classifier to complete the classification task.

2. The method for constructing a crop planting structure extraction model based on the Kolmogorov–Arnold network according to claim 1, characterized in that, The specific operation of step (1) is as follows: Principal component analysis is used for dimensionality reduction. In the PCA dimensionality reduction process, the original three-dimensional hyperspectral data array is first reshaped into a two-dimensional array to meet the input format requirements of PCA transformation; then PCA transformation is performed on the spectral information of each pixel, and the first 30 principal components are selected as new feature vectors. Before performing PCA transformation, the data is standardized to ensure that the mean of each feature is 0 and the standard deviation is 1; let the original data matrix be... Where m is the number of samples and n is the number of features; the standardization formula is as follows: (1) In the formula, is the standardized value of the element in the i-th row and j-th column, X ij μ is the original data value. j Let σ be the mean of the j-th feature. j Let be the standard deviation of the j-th feature; After data standardization, the covariance matrix C is calculated using the following formula: (2) In the formula, X′ is the standardized dataset matrix; Perform eigenvalue decomposition on the covariance matrix C to obtain the eigenvalues ​​λ. i and its corresponding eigenvector v i : (3) In the formula, λ i For the i-th eigenvalue, v i These are the corresponding eigenvectors; the magnitude of the eigenvalues ​​reflects the variance of the data along the direction of the corresponding eigenvector. The goal of PCA dimensionality reduction is to map the original n-dimensional high-dimensional data to a lower-dimensional k-dimensional space. Specifically, the eigenvalues ​​of the covariance matrix are first sorted in descending order, and the eigenvectors corresponding to the k largest eigenvalues ​​are selected to construct the projection matrix. ; The normalized data matrix X′ is then projected onto the projection matrix W to obtain the dimensionality-reduced data matrix Xk: (4) In hyperspectral remote sensing image processing, fixed-size image patches are extracted around each pixel to capture the local spatial structure information of the pixel. A two-pixel edge is filled around each pixel to ensure that each pixel is contained within a 25×25 window, which simultaneously covers the target pixel and its neighboring pixel information. Based on this method, the image patches extracted from the entire image contain both the spectral information of each pixel and its local spatial neighborhood information, providing rich input data for subsequent deep learning models.

3. The method for constructing a crop planting structure extraction model based on the Kolmogorov-Arnold network according to claim 1, characterized in that, The specific operation of step (2) is as follows: a three-dimensional kernel-activated convolutional module activated by radial basis function (RBF) is used as input, which is the hyperspectral feature data obtained in step (1) after standardization, mean removal preprocessing and PCA dimensionality reduction. , where B is the number of spectral bands after PCA dimensionality reduction, and H and W are the spatial height and width of the hyperspectral image, respectively; First, the dimensionality-reduced hyperspectral data is input into a three-dimensional convolutional structure, using... A three-dimensional convolutional kernel of a certain size is used to simultaneously extract local neighborhood features in the spectral dimension B, spatial height dimension H, and spatial width dimension W, capturing spectral-spatial joint correlation information and obtaining initial spectral-spatial joint features. The formula for 3D convolution is: (5) In the formula: k represents the weights of the 3D convolution kernel. s k h k w These represent the dimensions of the convolution kernel in the spectral, height, and width dimensions, respectively; Cout is the number of output feature channels; b 3D is the 3D convolution bias term, i and j are the row and column coordinates of the spatial features, respectively, and c is the output feature channel index; Secondly, a radial basis function (RBF) activation mechanism is introduced during the 3D convolution process. By replacing the traditional fixed linear mapping with a learnable kernel function, the Hybrd-KANet model can construct flexible nonlinear feature mapping relationships in the spectral-spatial joint domain. The RBF activation formula is as follows: (6) In the formula: This is the learnable bandwidth parameter for the RBF kernel function. For the learnable center parameters of the RBF kernel, For the Euclidean norm, this formula is adjusted... and Adaptive fitting of the spectral-spatial nonlinear characteristic distributions of different crops; Then, after each level of 3D convolution, a batch normalization (BN) layer and a Swish activation function are sequentially set. Batch normalization is used to stabilize the feature distribution and accelerate the model convergence speed. The batch normalization calculation formula is as follows: (7) In the formula: The mean of the features, The variance of the feature To prevent tiny constants with a denominator of 0, , The learnable scaling and offset parameters are batch-normalized; the Swish activation function is used to enhance the nonlinear representation of features and improve gradient propagation, and its expression is: (8) In the formula: The Sigmoid activation function is used to map feature values ​​to the [0,1] interval, achieving non-linear activation while alleviating the gradient vanishing problem; Finally, a progressive spectral dimension compression strategy is adopted, gradually reducing the kernel size along the spectral axis in different 3D convolutional layers. This involves starting with a larger spectral receptive field to capture coarse-grained global spectral responses, gradually transitioning to a smaller spectral receptive field to mine fine-grained spatial texture features, forming a multi-level feature abstraction system, and ultimately outputting deep spectral-spatial joint features. C 3D To determine the final output feature channel number, the deep spectral-spatial joint features are directly used as the input to the two-dimensional spatial convolution fusion module.

4. The method for constructing a crop planting structure extraction model based on the Kolmogorov-Arnold network according to claim 1, characterized in that, The specific operation of step (3) is as follows: The input of the two-dimensional spatial convolution fusion module is the high-dimensional spectral-spatial joint features output by the three-dimensional kernel-activated convolution module activated by radial basis function (RBF). C 3D H represents the number of feature channels output by the 3D kernel-activated convolutional module activated by the radial basis function (RBF), and H and W represent the spatial dimensions. First, the high-dimensional features output by the 3D kernel-activated convolution module are reorganized according to the channel dimension, merging the spectral dimension with the channel dimension to eliminate the dimensional limitation of the 3D features on the 2D convolution. The feature dimension transformation formula after reorganization is as follows: (9) In the formula: This represents the number of spectral dimensions remaining after progressive compression, for example, reducing the original channel dimensions to... The characteristics, recombine and unfold into This makes it a high-dimensional channel feature that can be processed by two-dimensional convolution; 32 is the number of channels and 18 is the remaining spectral dimension. Subsequently, a 3×3 two-dimensional convolution kernel is used to perform local neighborhood modeling in the spatial dimension, while simultaneously completing cross-channel feature fusion, compressing high-dimensional channel features to 64 channels. The two-dimensional convolution operation formula is as follows: (10) In the formula: b represents the weights of the two-dimensional convolution kernel. 2D For two-dimensional convolution bias terms, For a two-dimensional convolution operator, the final output is... ; Next, batch normalization (BN) and nonlinear activation functions are used to further stabilize the feature distribution and enhance spatial semantic representation. The computation process is consistent with the batch normalization and Swish activation in the 3D kernel activation convolution module with radial basis function (RBF) activation, i.e.: (11) (12) Finally, output a two-dimensional deep semantic feature map. This process preserves the spatial structure, edge texture, and local contextual information of crop regions while converting three-dimensional spectral-spatial features into two-dimensional deep semantic features, providing highly discriminative input features for the subsequent KANLinear classifier.

5. The method for constructing a crop planting structure extraction model based on the Kolmogorov-Arnold network according to claim 1, characterized in that, The specific operation of step (4) is as follows: The classifier built based on the KANLinear layer is used to perform nonlinear classification mapping on the deep semantic features output by the two-dimensional spatial convolution fusion module to achieve accurate determination of crop categories. First, the two-dimensional feature map output by the two-dimensional spatial convolution fusion module is... Flatten the vector to convert it into a one-dimensional feature vector. ,in The flattened feature dimensions are represented by the following formula: (13) Subsequently, this one-dimensional feature vector is input into the KANLinear layer. Through the learnable one-dimensional function mapping mechanism in the Kolmogorov-Arnold network (KAN), the input features are nonlinearly combined and approximated by an adaptive function. The core operation formula of the KANLinear layer is as follows: (14) In the formula: M is the number of learnable basis functions in the KANLinear layer. For the one-dimensional learnable basis functions of the KAN network, The learnable weight vector corresponding to each basis function For the bias term of the KANLinear layer, this formula achieves a complex nonlinear mapping of input features through a linear combination of multiple sets of learnable basis functions; Then, the KANLinear layer is used to replace the traditional fully connected layer, so that the classifier can not only perform linear weighting, but also learn more complex nonlinear decision boundaries, in order to overcome the shortcomings of the traditional fully connected layer in linear fitting ability and improve the accuracy of crop category discrimination. Finally, the features output by the KANLinear layer are... The input / output layer obtains classification scores for each crop category. Where K is the total number of crop categories, and is converted into a category probability distribution using the Softmax function. The Softmax calculation formula is: (15) In the formula: S k Let P(k) be the classification score for the k-th crop category, and P(k) be the probability that the pixel belongs to the k-th crop category. The category with the highest probability is selected as the final category for the pixel, i.e.: (16) By determining the category pixel by pixel, the final result of the extraction of crop planting structure in the whole area is generated, realizing the identification of crop types and the accurate division of planting areas.