Lightweight single backbone soil salinity continuous value estimation method for small sample agricultural remote sensing
Patent Information
- Application Number
- CN202610926112.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-08-28
AI Technical Summary
该方法具有较强的多源特征融合能力,但模型结构和数据组织方式较复杂,对多源数据完整性、样本规模、计算资源和训练调参过程要求较高
[0026] Preferably, in step (S6), the Huber loss function uses a squared error term when the absolute value of the prediction error is less than or equal to the threshold δ, and uses a linear error term when the absolute value of the prediction error is greater than the threshold δ.
Smart Images

Figure CN122657764A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for estimating continuous values of soil salinity using a lightweight single-framework system, particularly a method for estimating continuous values of soil salinity using a lightweight single-framework system for small-sample agricultural remote sensing. Background Technology
[0002] Soil salinization is a significant factor restricting farmland productivity, crop growth, and sustainable land use. Traditional soil salinity monitoring typically relies on manual sampling, laboratory testing, or portable conductivity meters. While these methods offer high accuracy, they suffer from drawbacks such as high sampling costs, limited spatial coverage, and difficulty in rapidly obtaining continuous spatial distribution data. With the development of UAV remote sensing, multispectral imaging, and deep learning technologies, utilizing crop canopy spectra, color, texture, and their derived indices to invert soil salinity has become an effective technological approach.
[0003] With the development of satellite remote sensing, UAV remote sensing, multispectral imaging, and machine learning technologies, soil salinity retrieval from remote sensing images has become an important technical direction. Existing methods typically extract features such as band reflectance, vegetation index, and salinity index, and combine them with models such as multiple linear regression, partial least squares regression, support vector machine, random forest, and backpropagation neural network to estimate soil salinity. These methods are simple to implement, but they mainly rely on manually constructed or screened features, making it difficult to fully explore the local spatial structure, texture changes, and nonlinear coupling relationships between multi-channel features in remote sensing images.
[0004] For example, CN113125383A discloses a method for monitoring and early warning of secondary salinization of farmland based on UAV remote sensing. This method utilizes UAV multispectral and thermal infrared imagery, combined with correlation analysis and machine learning models, to estimate soil salinity and provide early warning of salinization. While this method improves the efficiency of field image acquisition, its modeling process mainly relies on spectral values, spectral indices, or statistical features, lacking end-to-end image feature learning of spatially aligned feature sub-maps. Furthermore, it places greater emphasis on salinization level identification or early warning, making it difficult to directly meet the needs for continuous numerical estimation of soil salinity in precision irrigation and drainage and salt damage risk assessment.
[0005] In recent years, deep learning methods have also been increasingly used for remote sensing inversion of soil salinity. For example, CN114062439A discloses a method for estimating soil profile salinity, which utilizes Sentinel-2 time-series imagery, combined with random forest independent variable selection and a spatiotemporal convolutional network regression model to estimate soil profile salinity. This method can improve inversion capabilities by utilizing time-series information, but it relies on multiple satellite imagery periods, and the data acquisition, preprocessing, and time-series organization processes are relatively complex. It also lacks adaptability for rapid modeling under conditions of agricultural experimental fields, plot-scale data, and small sample measured labels.
[0006] CN118916634B discloses a method for soil salinity inversion, which extracts remote sensing image features, temporal geoscientific features, and spatial geoscientific features from multi-source remote sensing data, and achieves soil salinity inversion through band-sensitive weights, time-sensitive weights, and attention maps. This method has strong multi-source feature fusion capabilities, but its model structure and data organization are complex, requiring high levels of multi-source data integrity, sample size, computational resources, and efficient training and parameter tuning. In scenarios with small sample sizes, low computational resources, or rapid field estimation, it is prone to problems such as unstable training, high risk of overfitting, and high deployment costs.
[0007] In summary, existing soil salinity remote sensing estimation techniques still have the following shortcomings: First, traditional index methods and machine learning methods mainly rely on manual features, making it difficult to fully utilize the spatial structure and cross-channel coupling information in multi-channel remote sensing images; Second, some methods focus on salinization level identification or early warning, making it difficult to directly output continuous soil salinity values; Third, complex deep learning or multi-source attention models are highly dependent on sample size and computational resources, and are prone to overfitting and deployment difficulties in small-sample agricultural remote sensing scenarios; Fourth, existing technologies lack a single-backbone CNN regression method that is suitable for small-sample agricultural remote sensing scenarios and can directly output continuous soil salinity estimates while maintaining model lightweightness. Summary of the Invention
[0008] Purpose of the invention: The purpose of this invention is to provide a method for estimating continuous values of lightweight, single-framework soil salinity from small-sample agricultural remote sensing data.
[0009] Technical Solution: The present invention provides a lightweight single-backbone soil salinity continuous value estimation method for small-sample agricultural remote sensing, comprising the following steps: (S1) acquiring agricultural remote sensing images of the target farmland area and measured soil salinity data corresponding to the agricultural remote sensing images in time and space, wherein the agricultural remote sensing images include RGB images and / or multispectral images; (S2) preprocessing the agricultural remote sensing images to generate multi-channel salinity sensitive feature maps based on the preprocessed agricultural remote sensing images; (S3) spatially cropping the multi-channel salinity sensitive feature maps according to the spatial location information of field sampling points, experimental plots, or plot boundaries to obtain feature sub-map samples that correspond one-to-one with the measured soil salinity data, and constructing a small-sample agricultural remote sensing soil salinity estimation dataset including measured soil salinity data, feature sub-map samples, and continuous salinity labels.
[0010] Preferably, the measured soil salinity data includes continuous numerical salinity labels; the multi-channel salinity-sensitive feature map includes at least one of the following: original band features, vegetation index features, color index features, salinity-sensitive index features, and texture features.
[0011] Preferably, after step (S3), the following steps are also included:
[0012] (S4) The small sample agricultural remote sensing soil salinity estimation dataset is divided into training set, validation set and test set according to the salinity numerical distribution.
[0013] (S5) Construct a lightweight single-backbone convolutional neural network model, using the feature sub-map samples as input to the lightweight single-backbone convolutional neural network model and the continuous estimated value of soil salinity as output to construct the mapping relationship between measured soil salinity data and feature sub-map samples; the lightweight single-backbone convolutional neural network model includes a CNN backbone network, a global average pooling layer and a continuous value regression head connected in sequence.
[0014] This invention designs a lightweight single-backbone CNN model for small-sample scenarios in agricultural remote sensing. Compared with directly using large deep CNNs, it can significantly reduce the number of parameters and training complexity.
[0015] Preferably, after step (S5), the following steps are also included:
[0016] (S6) The lightweight single-backbone convolutional neural network model is trained using a training strategy oriented towards small-sample continuous regression to obtain the trained lightweight single-backbone convolutional neural network model. The training strategy includes one of Huber loss function, regularization constraint, cross-validation and hyperparameter optimization. This invention reduces the bias caused by uneven salt distribution in small-sample datasets on model training and evaluation through salt stratified sampling and cross-validation mechanism, and improves the model's adaptability to samples with different salt levels.
[0017] (S7) Input the agricultural remote sensing feature sub-map of the area to be estimated into the trained lightweight single-backbone convolutional neural network model, and output the continuous estimated value of soil salinity at the corresponding spatial location. This invention outputs continuous soil salinity values, rather than salinity level categories, which can more finely characterize the changes in salinity gradient and provide quantitative basis for precision irrigation and drainage, saline-alkali land management, crop salt stress assessment, and farmland water and salt regulation.
[0018] Preferably, the spatial location information of the field sampling points, experimental plots, or plot boundaries includes geographic coordinates, sampling point number, experimental plot number, or plot boundary number;
[0019] The continuous salinity label includes soil soluble salt content, soil electrical conductivity, apparent electrical conductivity converted salinity value or equivalent salinity index;
[0020] The continuous salinity label is determined by measured soil salinity data from multiple soil layers. The average salinity values of the 10-20cm, 20-30cm, and 30-40cm soil layers are taken as the continuous salinity label for the root zone.
[0021] Preferably, the CNN backbone network in step (S5) is one of the following: lightweight VGG type backbone, lightweight residual type backbone or lightweight dense connection type backbone, and the lightweight single backbone convolutional neural network model does not contain feature-level fusion structures of two or more parallel heterogeneous backbones.
[0022] When the CNN backbone is a lightweight VGG backbone, the lightweight VGG backbone includes three convolutional blocks connected in sequence. The number of convolutional kernels in the three convolutional blocks are 64, 128 and 256 respectively. Each convolutional block includes a first 3×3 convolutional layer, a first batch normalization layer, a first ReLU activation layer, a second 3×3 convolutional layer, a second batch normalization layer, a second ReLU activation layer and a pooling layer connected in sequence.
[0023] Preferably, when the CNN backbone is a lightweight residual backbone, the lightweight residual backbone includes an initial 3×3 convolutional layer and three residual stages connected in sequence, each residual stage including at least one residual block; the residual block includes a main branch and a shortcut connection branch, the main branch including two 3×3 convolutional layers, a batch normalization layer and a ReLU activation layer; when the spatial size or number of channels of the residual block input feature map is inconsistent with the output feature map of the main branch, a 1×1 convolutional layer is set in the shortcut connection branch to achieve matching of spatial size and number of channels.
[0024] Preferably, when the CNN backbone is a lightweight densely connected backbone, the lightweight densely connected backbone includes an initial 3×3 convolutional layer, three dense blocks and a transition layer. A transition layer is set between every two dense blocks. Each dense block includes multiple bottleneck units. The bottleneck unit includes a 1×1 convolutional layer and a 3×3 convolutional layer connected in sequence. The transition layer includes a 1×1 convolutional layer, a Dropout layer and an average pooling layer connected in sequence.
[0025] Preferably, the continuous value regression head includes at least one fully connected layer, a Dropout layer, and a linear output layer connected in sequence, wherein the linear output layer has an output neuron for outputting a continuous estimate of soil salinity.
[0026] Preferably, in step (S6), the Huber loss function uses a squared error term when the absolute value of the prediction error is less than or equal to the threshold δ, and uses a linear error term when the absolute value of the prediction error is greater than the threshold δ.
[0027] This invention employs the Huber loss function, which enables the model to maintain the sensitivity of squared loss within a small error range and reduce the impact of outliers within a large error range, thereby improving the regression robustness of agricultural measured data in the presence of noise.
[0028] An electronic device according to the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of any of the methods described herein.
[0029] The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described herein.
[0030] Beneficial Effects: Compared with existing technologies, this invention has the following significant advantages: 1. This invention designs a lightweight single-backbone CNN model for small-sample agricultural remote sensing scenarios. Compared with directly using large deep CNNs, it can significantly reduce the number of parameters and training complexity, reduce the risk of overfitting, and improve training stability under limited samples. This invention does not rely on complex multi-backbone fusion structures, and only runs one backbone network during deployment, resulting in low computational overhead. It is suitable for rapid soil salinity estimation in agricultural experimental scenarios, edge computing devices, UAV ground stations, or low-computing-power workstations. 2. This invention uses global average pooling to replace large-scale flattening layers or high-dimensional fully connected layers, enabling the model to reduce the parameter scale while preserving high-level spatial semantic features. It is suitable for continuous regression modeling of high-resolution, small-sized agricultural remote sensing feature sub-maps. 3. This invention outputs continuous soil salinity values, rather than... The salinity classification allows for a more refined characterization of salinity gradient changes, providing quantitative basis for precision irrigation and drainage, saline-alkali land management, crop salt stress assessment, and farmland water and salt regulation; 4. This invention reduces the bias caused by uneven salinity distribution in small sample datasets on model training and evaluation through salinity stratification sampling and cross-validation mechanisms, improving the model's adaptability to samples with different salinity levels; 5. The use of the Huber loss function maintains the sensitivity of the squared loss within a small error range and reduces the impact of outliers within a large error range, thereby improving the regression robustness of agricultural measured data in the presence of noise; 6. This invention is compatible with lightweight VGG-type backbones, lightweight residual-type backbones, and lightweight densely connected backbones, and selects the most suitable single-backbone model for the target region and sample conditions through a unified training strategy, exhibiting good adaptability and scalability. Attached Figure Description
[0031] Figure 1 This is a flowchart of the present invention;
[0032] Figure 2 This is a flowchart of agricultural remote sensing images, measured salinity labels, and spatially aligned feature sub-image samples in this invention;
[0033] Figure 3 This is a structural diagram of the lightweight single-backbone CNN continuous regression model in this invention;
[0034] Figure 4 This is a structural diagram of the lightweight visual geometry network-type single backbone network in this invention;
[0035] Figure 5 This is a schematic diagram of the lightweight residual single backbone network structure in this invention;
[0036] Figure 6 This is a schematic diagram of the lightweight, densely connected single-backbone network structure in this invention;
[0037] Figure 7 This is a schematic diagram of the model training process based on Huber loss, cross-validation, and Bayesian optimization in this invention.
[0038] Figure 8 Comparison charts of measured and predicted soil salinity values for lightweight densely connected single-backbone convolutional neural network models, lightweight residual single-backbone convolutional neural network models, and lightweight visual geometry group single-backbone convolutional neural network models.
[0039] Figure 9 This is a schematic diagram illustrating the generation of a spatial distribution map of soil salinity from continuous estimates in this invention. Detailed Implementation
[0040] like Figure 1 As shown, this invention proposes a lightweight single-backbone continuous soil salinity estimation method for small-sample agricultural remote sensing, including: Step S1: Acquire remote sensing image data and soil salinity label data.
[0041] Acquire agricultural remote sensing images of the target farmland area, and simultaneously or quasi-simultaneously acquire measured soil salinity data corresponding to the agricultural remote sensing images. The agricultural remote sensing images include RGB images, multispectral images, near-infrared images, or red-edge images. The measured soil salinity data are continuous numerical labels and may include soil soluble salt content, soil electrical conductivity, apparent electrical conductivity converted to salinity, saturated paste extract conductivity, or their equivalent salinity indicators.
[0042] Agricultural remote sensing imagery, including crop canopy imagery, was acquired using an unmanned aerial vehicle (UAV) platform. The UAV platform was equipped with an RGB camera and a multispectral camera, simultaneously acquiring RGB and multispectral images. Radiometric, color, and geometrical corrections were performed before or during image acquisition. The multispectral camera included at least blue, green, red, red-edge, and near-infrared bands. The UAV flew under clear, low-wind, and stable lighting conditions, preferably between 11:00 and 14:00, at an altitude of 20–60 m, more preferably 25 m; the forward overlap was 70%–85%, more preferably 75%. Radiometric calibration was performed using a standard reflectivity plate before acquiring multispectral images; color correction was performed using a color calibration plate before acquiring RGB images to reduce the impact of lighting variations, sensor drift, and camera exposure differences on subsequent salinity estimation.
[0043] Within the same or similar time window as the UAV image acquisition, measured soil salinity data were obtained. For each sampling point, test plot, soil box, or plot unit, its geographical coordinates, sampling number, or boundary number were recorded, and the corresponding soil salinity measurement values were collected. The apparent electrical conductivity (ECa) of the soil layers at depths of 0–10 cm, 10–20 cm, 20–30 cm, and 30–40 cm was measured using a soil conductivity meter. Subsequently, soil samples were collected, and after natural air drying, grinding, sieving, extraction, centrifugation, or filtration, the soluble salt content (SSC) was determined. If the salinity value is to be converted from electrical conductivity, the following empirical conversion formula can be used:
[0044] ,
[0045] Wherein, SSC represents the soluble salt content or equivalent salt content of the soil, ECa represents the apparent electrical conductivity, and a and b are regression coefficients obtained by fitting measured soil samples from the target area, with a value of 0.0013 and b value of 0.0008.
[0046] For root zone soil salinity estimation tasks, salinity values from multiple soil layers can be aggregated to obtain a continuous label for root zone soil salinity. For example, for the 10–40 cm root zone, the continuous label for the i-th sample's root zone salinity can be calculated using the following formula. :
[0047] ,
[0048] in, , and These represent the measured soil salinity data for the i-th sample in the 10–20 cm, 20–30 cm, and 30–40 cm soil layers, respectively. For surface salinity estimation tasks, the measured soil salinity data in the 0–10 cm layer can be directly used as continuous labels.
[0049] The above method yields a set of continuous soil salinity tags with spatial location information. :
[0050] ,
[0051] Among them, P i y represents the spatial location or number of the i-th sampling point, test plot, or land parcel unit. i This indicates that the position P is related to i The corresponding continuous label for soil salinity, where N represents the total number of samples.
[0052] like Figure 1 As shown, the "Data Acquisition" section employs parallel acquisition of agricultural remote sensing images and measured soil salinity data. For example... Figure 2 The diagram shows the correspondence between drone imagery, field sampling, conductivity measurements, and laboratory conversion labels.
[0053] Step S2: Preprocess the agricultural remote sensing images.
[0054] The agricultural remote sensing images acquired in step S1 are preprocessed to ensure that images acquired from different bands, sensors, and times have unified spatial coordinates, unified spatial resolution, and comparable radiometric scales. The preprocessing includes one or more of the following: image stitching, aerial triangulation, orthorectification, radiometric correction, color normalization, band overlay, image registration, resampling, and spatial resolution unification.
[0055] First, aerial triangulation, 3D reconstruction, and orthorectification are performed on the raw single-band images collected by the UAV to generate orthorectified images covering the target farmland area. Then, the RGB images and / or multispectral images are overlaid to obtain RGB composite images and multispectral composite images. Next, based on ground control points, sensor positioning information, or image feature points, geometric registration is performed on the RGB composite images and multispectral composite images to make different image data coincide in the same geographic coordinate system. Finally, all images are resampled to a uniform spatial resolution.
[0056] For multispectral images, radiometric correction can be performed using a standard reflectivity plate. Let the original pixel value of the b-th band be DN. b The dark current or black level value is D. b The standard reflectivity plate pixel value is P b The true reflectance of the standard reflectance plate is The reflectance after correction for the b-th band is... It can be represented as:
[0057] ,
[0058] For RGB images, color normalization or white balance correction can be performed to reduce the differences in lighting conditions between different voyages.
[0059] Geometric registration between different images can be represented as:
[0060] ,
[0061] Where (x,y) are the original image coordinates, (x′,y′) are the registration coordinates in the target coordinate system, and T is the coordinate transformation function obtained based on ground control points, sensor attitude, or image feature matching. After registration, different bands, different indices, and different texture features have consistent spatial positions at the pixel level.
[0062] Step S3: Construct a multi-channel salt-sensitive feature map.
[0063] After image preprocessing, features related to soil salinity response are extracted from RGB and / or multispectral images to generate a multichannel salinity-sensitive feature map. The multichannel salinity-sensitive feature map includes at least one of the following: original band features, vegetation index features, color index features, salinity sensitivity index features, and texture features.
[0064] Let the preprocessed image region be Ω, and the c-th feature channel be F. c Then the multi-channel salt-sensitive feature map It can be represented as:
[0065] ,
[0066] in, The total number of feature channels, and each feature channel , and These represent the spatial height and width of the feature map, respectively. After stacking the two-dimensional feature channels along the channel dimension, a three-dimensional feature tensor is obtained. ,in It represents a number composed of real numbers with a size of The three-dimensional tensor space.
[0067] Raw band features include the red channel of RGB images. Green light channel and Blu-ray channel And the blue band of multispectral images Green light band Red light band Red-edge band and near-infrared band .
[0068] Spectral indices constructed based on the reflectance of the above-mentioned bands include the Normalized Difference Vegetation Index (NDVI), Soil-Adjusted Vegetation Index (SAVI), Optimized Soil-Adjusted Vegetation Index (OSAVI), Green Chlorophyll Index (GCI), Normalized Green-Red Difference Index (NGRDI), and Modified Green-Red Vegetation Index (MGRVI), and their calculation formulas are as follows:
[0069] , , ,
[0070] , , ,
[0071] in, , and These represent the reflectance of the target crop canopy in the near-infrared band, the red band, and the green band, respectively; when using the above multispectral imagery, , and They can correspond to the near-infrared band respectively. Red light band and green light band . This represents the adjustment factor used to reduce the influence of soil background; in this embodiment, it is set to 0.5. 0.16 is the soil background adjustment constant in OSAVI. and Let represent the squares of the reflectance in the green light band and the squares of the reflectance in the red light band, respectively.
[0072] The reflectance values for the aforementioned bands are dimensionless values obtained through preprocessing. This preprocessing includes radiometric calibration and reflectance conversion, and may further include atmospheric correction, geometric correction, image registration, orthorectification, image stitching, and outlier removal, depending on the image type. When calculating each spectral index, if the denominator of the formula is zero or less than a preset effective threshold, the corresponding pixel is marked as an invalid pixel and removed.
[0073] Texture features can be extracted based on the gray-level co-occurrence matrix (GLCM). For the first... For each image channel, within a preset sliding window, the co-occurrence relationship of gray-level pixel pairs is statistically analyzed according to preset pixel spacing and orientation, and a normalized gray-level co-occurrence matrix is constructed. Furthermore, the contrast, homogeneity, entropy, and second angular moment of this channel are calculated and denoted as follows: , , and .
[0074] ,
[0075] ,
[0076] ,
[0077] ,
[0078] in, Indicates the image channel number; , , and They represent the first The contrast, homogeneity, entropy, and second moment of each image channel are used to characterize the degree of gray-level difference, gray-level distribution concentration, random complexity, and uniformity of the local texture of that channel. This indicates the number of gray levels in the image after grayscale measurement. and Indicates the grayscale level index, and ; Indicates the first The gray levels in each image channel are respectively and The normalized probability of pixel pairs co-occurring under the preset pixel spacing and orientation satisfies: , Represents the natural logarithm function; This represents a preset positive stable term, used to avoid... Logarithmic operations are undefined.
[0079] To reduce the dimensionality of candidate features and the risk of overfitting with small samples, candidate features are screened based on their contribution and redundant features are removed. First, a random forest model is used to evaluate the contribution of each candidate feature to reducing the salt prediction error. Then, a recursive feature elimination method is used to iteratively remove the features with the lowest contribution, resulting in candidate feature subsets corresponding to different feature counts. Subsequently, K-fold cross-validation is used to calculate the average error of each candidate feature subset. Finally, the one-standard-error criterion is used to select the simplest feature subset within the acceptable range of prediction performance.
[0080] Let the K-fold cross-validation error be m when the number of features is m. The standard deviation of the K fold error is s m Then the standard error for: Let the minimum average error be... for: ,
[0081] The corresponding standard error is SE min Then the number of candidate features m under the one standard error criterion ∗It can be represented as:
[0082] ,
[0083] In the formula, min is the minimum value function; This represents the standard error corresponding to the minimum average cross-validation error.
[0084] To further reduce strong linearity and collinearity, the variance inflation factor (VIF) can be calculated for the retained features. For the j-th feature, an auxiliary regression model is established using the remaining features as independent variables to obtain the coefficient of determination. Then the feature for: .
[0085] When the VIF of a feature exceeds a preset threshold, it is determined that the feature has strong linear redundancy with other features. The preset threshold can be 30. When there is high collinearity among multiple features, features with higher importance ranking in the random forest are retained first, while features with lower importance and strong collinearity are removed. In this way, the original high-dimensional candidate features are compressed into a more compact multi-channel salt-sensitive feature map, improving the training stability of the subsequent lightweight single-backbone model.
[0086] To reduce the differences in numerical scale and value range among different feature channels, standardization can be performed on each feature channel:
[0087] Alternatively, use minimum-maximum normalization:
[0088] ,
[0089] in, and They represent the first Feature values of each feature channel before and after processing; Indicates the feature channel number; and They represent the first The mean and standard deviation of each feature channel in the training set; and These represent the minimum and maximum values of the feature channel in the training set, respectively. This represents a preset, extremely small positive number to prevent the denominator from being zero. During model validation and application, the corresponding statistical parameters determined from the training set are used to process the input features.
[0090] Figure 1 The "Construction of Multi-channel Salt-Sensitive Features" module corresponds to this step. Figure 2 The “Feature Map Generation” module in the document illustrates the process of generating multi-channel feature maps from the original bands, vegetation indices, and texture features.
[0091] Step S4: Construct spatially aligned feature submap samples.
[0092] Based on the spatial location information of field sampling points, experimental plots, soil box boundaries, or plot boundaries, the multi-channel salinity sensitive feature map is spatially cropped to obtain feature sub-map samples that correspond one-to-one with the measured soil salinity labels.
[0093] Let the first The spatial positioning information or boundary area corresponding to each sampling point or test area is: The multi-channel salt sensitivity feature map is as follows The target cutting size is Then the first Each feature sub-map sample is represented as:
[0094] ,
[0095] in, Indicates the first Each sampling point or test cell corresponds to a feature sub-map sample, and ,
[0096] In the formula, Indicates the serial number of the sampling point or test area; This refers to a cropping operation that extracts the target region from a multi-channel salt-sensitive feature map based on its spatial location or boundary range and adjusts it to a preset size. This represents a multichannel salt-sensitive feature map composed of one or more of the original band features, spectral index features, and texture features. Indicates the first The spatial coordinates of the sampling point and its neighborhood range, or the... Vector boundaries of each test cell; and These represent the preset height and width of the feature sub-map, respectively; Indicates the number of feature channels; Represents a set of real numbers with a size of A three-dimensional feature array.
[0097] When extracting feature sub-images centered on sampling points, the coordinates of the sampling points and the preset window radius are determined. When extracting feature sub-maps on a per-test-plot or per-soil-box basis, the corresponding vector boundaries are determined. For irregularly shaped boundary regions, their minimum bounding rectangle can be extracted, and the extracted region can be resampled to... Alternatively, a mask can be generated based on the vector boundary, setting the pixels outside the boundary to preset fill values or invalid values to reduce the impact of non-target regions on subsequent model training.
[0098] Each feature submap sample X i Its corresponding continuous salt label y i Form a sample pair:
[0099] ,
[0100] Where D represents a small-sample agricultural remote sensing soil salinity estimation dataset, and N represents the number of samples. To ensure spatial consistency between samples and labels, sample pairs can be established through geographic coordinate matching, sampling point number matching, experimental plot number matching, soil box number matching, or plot boundary number matching. If multiple sampling points exist in the same plot, the mean, median, or weighted average of the salinity values from multiple sampling points can be used as the label for that plot. If multiple candidate clipping windows exist around the same sampling point, the window with the closest spatial distance to the sampling point can be selected, or a fixed-size window can be constructed centered on the sampling point.
[0101] Based on 12 salinity treatment experimental plots and their multiple UAV observations, several spatially aligned samples were constructed; each sample consists of a multi-channel feature sub-map and a continuous soil salinity value. In this way, remote sensing image data is transformed from large-scale orthophotos into small-sized multi-channel image samples suitable for CNN input.
[0102] Figure 2 In this process, UAV imagery is preprocessed to form multi-channel feature maps; field sampling points, experimental plots, or soil bin boundaries provide the basis for spatial cropping; the cropped feature sub-maps are matched one-to-one with continuous salinity labels through numbering or coordinates, ultimately forming... Small sample datasets in the form of...
[0103] Step S5: Divide the data into layers according to the distribution of salinity values.
[0104] To avoid training bias caused by uneven distribution of low-salt, high-salt, or intermediate-salt samples under small sample conditions, this invention divides the small sample dataset into training, validation, and test sets according to the salt content distribution.
[0105] Let the set of salt labels for all samples be . The salinity can be divided into Q salinity intervals based on the statistical distribution of salinity values. :
[0106] ,
[0107] Where, τ qThe threshold can be determined based on quantiles or based on the agronomic salt damage level. For example, it can be divided into three intervals: low salt, medium salt, and high salt, or into four intervals: no salinization, slight salinization, moderate salinization, and severe salinization.
[0108] For each salinity interval I q Training, validation, and test samples are randomly selected proportionally from the sample set belonging to this interval, ensuring that the proportion of samples with different salinity levels is the same or approximately the same in each data subset. Let the training set, validation set, and test set be D, respectively. train D val and D test Then the following conditions are met:
[0109] , ,
[0110] The ratio of training set to test set is 7:3 or 6:4 to 8:2. The validation set can be further partitioned from the training set, or it can be generated cyclically within the training set using K-fold cross-validation. This invention uses salt stratification to ensure that low-salt, medium-salt, and high-salt samples all participate in model training and evaluation, reducing model evaluation instability caused by skewed salt label distribution.
[0111] Figure 1 The "salt stratification dataset partitioning" module corresponds to this step, where the training set, validation set, and test set are used for model parameter learning, hyperparameter selection, and independent performance evaluation, respectively.
[0112] Step S6: Construct a lightweight single-backbone convolutional neural network model.
[0113] The lightweight single-backbone convolutional neural network model includes an input layer, a single convolutional neural network backbone, a global average pooling layer, a continuous value regression head, and a linear output layer. During a single forward inference process, the lightweight single-backbone convolutional neural network model extracts features using only one convolutional neural network backbone, without containing two or more parallel backbones, nor does it contain a feature-level fusion structure composed of output features from multiple heterogeneous backbones.
[0114] Input to the model The input feature tensor of each model is represented as:
[0115] ,
[0116] in, Indicates the sample sequence number; Indicates the first Each model inputs a feature tensor; Represents the real number field; and These represent the height and width of the model input feature tensor, respectively. This indicates the total number of input feature channels.
[0117] The lightweight single-backbone convolutional neural network model performs deep feature extraction on the model input feature tensor to obtain:
[0118] ,
[0119] and: ,
[0120] in, The parameter set is Mapping function for lightweight single-backbone convolutional neural network models; This represents the set of trainable parameters in the lightweight single-backbone convolutional neural network model. Indicates the first Deep feature maps corresponding to each sample; and These represent the height and width of the deep feature map, respectively. This represents the total number of channels in the deep feature map.
[0121] By averaging the spatial features of each channel in the deep feature map using a global average pooling layer, we obtain:
[0122] ,
[0123] The corresponding global feature vector is represented as:
[0124] ,
[0125] in, Indicates the first The deep feature map of the nth sample The feature values obtained by global average pooling of each channel; Indicates the first The global feature vector corresponding to each sample; Indicates the first The Dth global feature vector corresponding to each sample; The channel number representing the deep feature map; and These represent the row position number and column position number in the deep feature map, respectively; Represents a D-dimensional real vector space; Representing deep feature maps In the line, number Column and number Feature values at each channel; superscript This represents the transpose of a vector.
[0126] The Global Average Pooling Layer (GAP) is used to pool cells with a size of [size missing]. The deep feature map is compressed into a length of One-dimensional global feature vectors are used to reduce the number of model parameters generated by flattening operations and large-scale fully connected layers.
[0127] The continuous value regression head maps the global feature vector to obtain the hidden feature vector. :
[0128] ,
[0129] And obtain the first through the linear output layer Continuous estimates of soil salinity for each sample:
[0130] ,
[0131] in, This represents a nonlinear activation function; in one embodiment, the modified linear unit activation function ReLU is used. and These represent the weight matrix and bias vector of the first fully connected layer in the continuous value regression head, respectively. and These represent the weight vector and bias scalar of the linear output layer, respectively. Indicates the first Continuous estimates of soil salinity for each sample; symbol " "" indicates the model estimate. Indicates the first The global feature vector is obtained by global average pooling of the deep feature maps of each sample, and ,in This represents the number of channels in the deep feature map.
[0132] The nonlinear activation function is:
[0133] ,
[0134] in, The input variables represent the non-linear activation function; This indicates that the maximum value among the values within the parentheses is selected.
[0135] The linear output layer contains one output neuron and does not have a Softmax activation function. Here, Softmax represents a normalized exponential function used to convert multiple output values into a class probability distribution. In this embodiment, the output is a continuous value of soil salinity, rather than discrete class probabilities.
[0136] Hidden feature vectors A random deactivation layer or a regularization layer is placed between the linear output layer and the model to reduce the risk of overfitting under small sample training conditions. The random deactivation layer, abbreviated as Dropout, is used to randomly block some neuron outputs according to a preset deactivation probability during model training; random blocking is not performed during model inference.
[0137] like Figure 3 As shown, the model input feature tensor passes sequentially through a single-backbone convolutional neural network, a global average pooling layer, a continuous value regression head, and a linear output layer to obtain continuous estimates of soil salinity. Figure 3 The term "single backbone" in this context means that the model has only one convolutional neural network backbone path for deep feature extraction, which distinguishes it from models with dual or multiple backbone parallel feature extraction and fusion structures.
[0138] In this context, CNN stands for Convolutional Neural Network, GAP stands for Global Average Pooling, ReLU stands for Modified Linear Unit Activation Function, and Dropout represents random deactivation.
[0139] The backbone of a convolutional neural network can adopt a lightweight visual geometrical network structure, namely a lightweight VGG-type single-backbone network, such as... Figure 4 As shown. The network consists of three convolutional blocks connected in sequence, each convolutional block including two convolutional kernels with a size of [missing information]. The system consists of three convolutional layers and one max-pooling layer. The output channels of the three convolutional blocks are 64, 128, and 256, respectively. Each convolutional layer is followed by a batch normalization layer and a modified linear unit activation function.
[0140] make: ,
[0141] No. The computation process of each convolutional block is represented as follows:
[0142] ,
[0143] ,
[0144] ,
[0145] in: ,
[0146] The output of the third convolutional block serves as the deep feature map extracted by a lightweight VGG-type single backbone network:
[0147] ,
[0148] in, Indicates the sample sequence number; Indicates the first Each model inputs a feature tensor; Indicates the convolution block number; This represents the input to a lightweight VGG-type single backbone network; and They represent the first Intermediate feature maps output by the first and second convolutional layers in a convolutional block; Indicates the first The feature map output after max pooling of each convolutional block; Indicates the first Number of output channels per convolutional block; Indicates the first The first convolutional block Each convolutional layer performs a convolution operation with a kernel size of . The number of output channels is ,in ; This indicates the batch normalization operation following the corresponding convolutional layer; This represents the modified linear unit activation function; Indicates the first Maximum pooling operation in each convolutional block; Indicates the first The deep feature map corresponding to each sample.
[0149] The maximum pooling layer adopts The pooling window and stride are set to 2 to progressively reduce the spatial size of the feature maps. The deep feature map output by the third convolutional block... After passing through a global average pooling layer, a fully connected layer containing 128 neurons, a random inactivation layer, and a single-neuron linear output layer in sequence, the first... Continuous estimates of soil salinity for each sample.
[0150] The lightweight VGG-type single backbone network utilizes continuous... Convolutional processing extracts local spatial features of the crop canopy, and by reducing the number of convolutional blocks and using a global average pooling layer, the number of network parameters and the risk of overfitting are reduced.
[0151] In another embodiment, the convolutional neural network backbone can be a lightweight residual single-backbone network, such as... Figure 5 As shown. The network includes an initial... The convolutional layer consists of three sequentially connected residual stages, each residual stage comprising at least one residual block, and each residual block comprising a main branch and a shortcut connection branch.
[0152] Shallow feature map output from the initial convolutional layer Represented as:
[0153] ,
[0154] No. In the residual stage, the first The output feature map of each residual block is represented as follows:
[0155] ,
[0156] Residual mapping of the main branch Represented as:
[0157] ,
[0158] When the spatial dimensions and number of channels of the residual block input and output are the same, the shortcut connection branch uses identity mapping:
[0159] ,
[0160] When the spatial dimensions or number of channels of the residual block input and output are inconsistent, the quick connection branch adopts... Convolution for dimension matching:
[0161] ,
[0162] in, This indicates the number of output channels of the initial convolutional layer; Indicates the residual stage number, and ; Indicates the first The residual block number within each residual stage; Indicates the first The input of the first sample In the residual stage, the first Feature map of each residual block; The input feature map represents the residual mapping function; This represents the residual mapping of the main branch execution; This represents the set of trainable parameters in the main branch; Indicates a shortcut mapping; This represents the set of trainable parameters in the shortcut connection branch; when using an identity mapping, this set is empty. Indicates the use of space size or channel number matching. Convolution operation; This indicates the batch normalization operation performed after the initial convolutional layer; Indicates the kernel size as The number of output channels is The initial convolution operation; and They represent the first In the residual stage, the first The first and second main branches of the residual block Convolution operation; and They represent the first and second mentioned above, respectively. Batch normalization is performed after the convolution operation; Represents the input feature map The output feature map is obtained after the shortcut connection branch mapping.
[0163] When performing spatial downsampling, the first branch of the main branch Convolution and shortcut connection branches All convolutions use a stride of 2 to ensure consistent spatial dimensions between the outputs of the two branches. The output of the last residual block in the third residual stage is used as the deep feature map. .
[0164] The deep feature map The model sequentially passes through a global average pooling layer, a first fully connected layer, a second fully connected layer, a random deactivation layer, and a single-neuron linear output layer. The number of neurons in the first fully connected layer is determined based on the model size or hyperparameter optimization results, and the second fully connected layer preferably includes 32 neurons.
[0165] The lightweight residual single-backbone network improves gradient propagation through fast connections, reduces the risk of training degradation or gradient instability when the network deepens, and maintains the network's lightweight nature by limiting the number of residual stages and the size of the regression head.
[0166] The backbone of a convolutional neural network can also be a lightweight, densely connected single-backbone network, such as... Figure 6 As shown. The network includes an initial... The system comprises a convolutional layer, three dense blocks, and a transition layer between adjacent dense blocks. The initial convolutional layer preferably has 32 output channels, each dense block preferably includes three bottleneck units, and the growth rate is preferably 16.
[0167] The output of the initial convolutional layer Represented as:
[0168] ,
[0169] No. In the th dense block, the th New feature maps generated by each bottleneck unit Represented as:
[0170] ,
[0171] Nonlinear transformation of bottleneck unit Represented as:
[0172] ,
[0173] No. Output feature map of a dense block Represented as:
[0174] ,
[0175] Transition layer between adjacent dense blocks Represented as:
[0176] ,
[0177] in, Indicates the first The input of the first sample Feature maps of dense blocks; Indicates the dense block number, and ; Indicates the index of the bottleneck unit within the dense block; Indicates the first In one specific implementation, the total number of bottleneck units contained in a dense block is... ; Indicates the first In the th dense block, the th New feature maps generated for each bottleneck unit; This indicates that the feature maps within the brackets are stitched together along the feature channel dimension; Indicates the first In the th dense block, the th Nonlinear transformations performed by each bottleneck unit; This represents the spliced feature map received by the bottleneck unit. This represents the growth rate, i.e., the rate at which each bottleneck unit passes through. In one specific implementation, the number of new feature channels added by convolution... . Indicates the kernel size as An initial convolution operation with 32 output channels; and They represent the first In the th dense block, the th The first and second batch normalization operations within each bottleneck unit; This indicates the number of channels used for adjustment or compression in the bottleneck unit. Convolution operation; This indicates that the number of output channels is of Convolution operation; Indicates the first The number of channels is compressed in each transition layer. Convolution operation; The probability of inactivation is expressed as Random deactivation operation; Indicates the first Average pooling operations are used in each transition layer to reduce the spatial size of the feature map.
[0178] The channel compression factor of the transition layer is 0.5, meaning the number of output channels in the transition layer is half the number of its input channels; the average pooling layer uses... The pooling window and stride 2 are used to reduce the spatial size of the feature map.
[0179] The output of the third dense block is used as a deep feature map. The deep feature map The layers sequentially pass through a global average pooling layer, a first fully connected layer, a second fully connected layer, and a single-neuron linear output layer. Random deactivation layers can be set between fully connected layers or before the linear output layer. The number of neurons in the first fully connected layer is determined through hyperparameter optimization, and the second fully connected layer preferably includes 32 neurons.
[0180] The lightweight, densely connected single backbone network achieves the transfer and reuse of features at different levels by splicing feature maps generated by bottleneck units within the same dense block along the channel dimension; and reduces redundant features and model parameters by compressing feature channels and reducing spatial size through transition layers.
[0181] In the above, VGG represents the visual geometric group network structure; BN represents batch normalization; ReLU represents the modified linear unit activation function, and its corresponding function is:
[0182] ,
[0183] in, This represents the input value of the activation function. This indicates the operation of retrieving the maximum value.
[0184] , , , and These represent convolution, max pooling, average pooling, random deactivation, and channel concatenation operations, respectively; GAP represents global average pooling; a single-neuron linear output layer represents an output layer containing only one output neuron without a Softmax classification activation function.
[0185] Figure 4 , Figure 5 and Figure 6The lightweight VGG, lightweight residual, and lightweight densely connected single-backbone continuous regression models shown respectively output continuous estimates of soil salinity in the order of "single backbone feature extraction - global average pooling - continuous value regression head - single neuron linear output", and each model uses only one backbone structure.
[0186] Step S7: Train the model using a training strategy oriented towards small sample continuous regression.
[0187] The lightweight single-backbone convolutional neural network continuous regression model constructed in step S6 was trained using a training set and a validation set. The training set was used to update the model's trainable parameters, while the validation set was used for hyperparameter selection, cross-validation evaluation, and early stopping detection. The training objective of the model was to reduce the error between the continuously estimated soil salinity values and the measured soil salinity values. One or more methods were employed, including Huber loss function, L2 regularization, random deactivation, cross-validation, early stopping, and hyperparameter optimization, to improve the model's training stability and generalization ability under small-sample agricultural remote sensing conditions.
[0188] Let the first The measured soil salinity values for each training sample are The model outputs a continuous estimate of soil salinity. Then the first Estimation error per training sample Represented as:
[0189] ,
[0190] The regression loss for a single training sample is calculated using the Huber loss function. :
[0191] ,
[0192] in, Indicates the training sample number; Indicates the first The estimation error of each training sample; and They represent the first Measured soil salinity values and continuous estimates of soil salinity for each training sample; This represents the preset Huber loss threshold, and ; This represents the absolute value of the estimation error.
[0193] when Not greater than When the Huber loss is used, squared loss is employed to improve the model's ability to fit smaller estimation errors; when Greater than In this case, the Huber loss is linearly increased to reduce the impact of outliers, measured noise, or local image errors on model parameter updates.
[0194] By adding L2 regularization constraints to the training objective, we obtain the regularized training objective function of the model. :
[0195] ,
[0196] in, This represents the set of all trainable parameters in the lightweight single-backbone convolutional neural network continuous regression model. This represents the total number of samples in the training set; Denotes the L2 regularization coefficient, and ; This represents the squared L2 norm of all trainable parameters of the model. By minimizing this regularized training objective function, excessive increase in model parameters can be suppressed while fitting the training samples, thereby reducing the risk of overfitting under small sample conditions.
[0197] During model training, a random deactivation layer can also be set in the continuous value regression head. (The first...) Hidden feature vectors of each sample For example, the hidden feature vector after random deactivation. Represented as:
[0198] ,
[0199] The random deactivation mask satisfies:
[0200] ,
[0201] in, This represents the hidden feature vector before random deactivation. This represents the hidden feature vector after random deactivation. Indicates and Random deactivation mask vectors with the same dimensions; This indicates element-wise multiplication; Indicates the proportion of random inactivation, and ; Indicates the probability of retention is The Bernoulli distribution. During the training phase, some hidden features are randomly masked and distributed according to... Scale compensation is performed to reduce co-adaptation among neurons; random masking is not performed during model validation and inference phases.
[0202] During model training, early stopping can be implemented based on the validation set loss. When the validation set loss does not decrease further within a preset number of consecutive training epochs, model training is stopped, and the model parameters corresponding to the lowest validation set loss are retained. The preset number of consecutive training epochs represents the number of epochs for early stopping, and its specific value can be determined based on the number of training samples and the model convergence.
[0203] The model hyperparameters are determined using Bayesian optimization. The hyperparameters to be optimized may include the learning rate. The number of samples contained in each mini-batch Random inactivation rate L2 regularization coefficient Huber loss threshold One or more of the following: maximum number of training rounds, number of neurons in a fully connected layer, and number of output channels in each convolutional layer.
[0204] Represent a set of candidate hyperparameters as a hyperparameter vector. ,use During folded cross-validation, the hyperparameter vector is calculated. The corresponding cross-validation average error :
[0205] ,
[0206] Among them, the Root mean square error corresponding to the validation set Represented as:
[0207] ,
[0208] in, This represents a hyperparameter vector consisting of a set of candidate hyperparameters; This represents the total number of folds in the cross-validation. This indicates the cross-validation fold number, and ; Indicates the use of hyperparameter vectors At that time, the first The root mean square error on the validation set; RMSE represents the root mean square error. Indicates the first The number of samples included in the validation set; Indicates the first The sample number in the validation set; and They represent the first The first fold verification set Measured soil salinity values for each sample and using hyperparameter vectors The obtained continuous estimates of soil salinity.
[0209] The Bayesian optimizer builds a surrogate model based on the evaluated hyperparameter vectors and their cross-validation average error, and uses the acquisition function to determine the Bayesian optimization order. Candidate hyperparameter vectors selected in rounds :
[0210] ,
[0211] in, This indicates the iteration round of the Bayesian optimization; Represent the candidate hyperparameter space; Indicates based on the previous The surrogate model was established based on the hyperparameter evaluation results of the round; In the proxy model Under the given conditions, candidate hyperparameter vector The corresponding acquisition function value; This represents the candidate hyperparameter vector that maximizes the acquisition function.
[0212] Through multiple rounds of Bayesian optimization and cross-validation, a combination of hyperparameters with a smaller average cross-validation error is selected, and the selected hyperparameters are used to complete model training, resulting in a lightweight single-backbone convolutional neural network continuous regression model for outputting continuous estimates of soil salinity.
[0213] In K-fold cross-validation, K can be 3 or 5, with 3 being the preferred choice. For small sample data, too high a number of folds may result in too few samples being validated per fold, thus increasing evaluation volatility; using 3-fold or 5-fold cross-validation can balance evaluation stability and computational efficiency.
[0214] Early stopping can be employed during model training. Training is stopped when the validation set loss or validation set RMSE fails to improve within P consecutive training epochs, and the model parameters that best perform on the validation set are saved. The model optimizer can be Adam, AdamW, RMSprop, or SGD, with Adam or AdamW being preferred.
[0215] like Figure 7 As shown, the training process of this invention includes small sample data input, candidate hyperparameter setting, Bayesian optimization, K-fold cross-validation, Huber loss calculation, regularization constraints, Dropout to suppress overfitting, validation metric feedback, and optimal model output. Through this closed-loop training strategy, the training stability of a lightweight single-backbone CNN continuous regression model can be improved in agricultural remote sensing scenarios with limited sample size, measurement noise in salt labels, and high image feature dimensionality.
[0216] After model training is complete, the accuracy of the lightweight single-backbone convolutional neural network continuous regression model is evaluated using an independent test set or retained fold samples that did not participate in the corresponding fold model training during cross-validation. The model's output continuous estimates of soil salinity are compared with the corresponding measured soil salinity values, and the coefficient of determination is used. Root mean square error and mean absolute error Evaluate model performance. This includes evaluating the average measured soil salinity values of the evaluation samples. for:
[0217] ,
[0218] Based on this, the coefficient of determination Root mean square error The mean absolute error (MAE) is calculated using the following formulas:
[0219] ,
[0220] ,
[0221] ,
[0222] in, The coefficient of determination is used to characterize the degree to which the model fits the measured changes in soil salinity. This represents the root mean square error, used to characterize the overall level of model estimation error; Mean absolute error is used to characterize the average level of the absolute deviation between the model estimate and the measured value. This represents the total number of samples participating in the accuracy evaluation; Indicates the sample number being evaluated; Indicates the first Measured soil salinity values for each evaluation sample; Indicates the first Continuous estimates of soil salinity for each evaluation sample; This represents the arithmetic mean of the measured soil salinity values for all evaluation samples.
[0223] Coefficient of determination The larger the value, the better the model fits the measured changes in soil salinity; root mean square error and mean absolute error The smaller the value, the smaller the deviation between the model estimation result and the measured result.
[0224] like Figure 8As shown, scatter regression analysis can be performed to compare the continuous soil salinity estimation results of lightweight densely connected single-backbone networks, lightweight VGG single-backbone networks, and lightweight residual single-backbone networks. The horizontal axis represents the measured soil salinity value, and the vertical axis represents the model-estimated soil salinity value. The regression line and its confidence interval are used to characterize the consistency between the predicted and measured values. By comparing the R², RMSE, and MAE of different single-backbone models, a single-backbone CNN with high prediction accuracy, low error, and suitable structural complexity can be selected as the final estimation model under the condition of small-sample agricultural remote sensing data in the target area. This accuracy evaluation process is used for model selection and effect verification, and does not change the technical characteristics of this invention, which uses a lightweight single-backbone CNN for continuous soil salinity estimation.
[0225] Step S8: Output continuous estimates of soil salinity and generate a spatial distribution map.
[0226] The agricultural remote sensing image of the area to be estimated is preprocessed in the same or similar manner as in steps S2 and S3, and a multi-channel salinity-sensitive feature map is constructed. Subsequently, following the processing method in step S4, the area to be estimated is cropped based on spatial partitioning methods such as sliding windows, sampling units, plot boundaries, or regular grids to obtain several sub-maps of features to be predicted. These sub-maps are then input into the lightweight single-backbone CNN model trained in step S7 to obtain continuous soil salinity estimation results for each spatial unit.
[0227] For the r-th spatial unit to be predicted, its feature sub-map is: The model outputs a continuous estimate of soil salinity in the r-th spatial unit. for:
[0228] ,
[0229] Among them, f θ ∗ This represents the lightweight single-backbone CNN continuous regression model after training, θ ∗ This represents the optimal model parameters.
[0230] If a sliding window method is used to estimate the area, with a window size of H×W and a sliding step size of s, when multiple windows produce estimates for the same pixel or grid cell, an average fusion method can be used to generate the estimated soil salinity value at pixel or grid cell q. :
[0231] ,
[0232] Where q represents a pixel or grid cell in the target region, Ω rThis represents the spatial extent covered by the r-th prediction window. The final spatial distribution of salinity can also be generated using distance-weighted averaging, median fusion, parcel mean fusion, or other spatial aggregation methods.
[0233] The continuous estimated values of soil salinity output by the model can be written into a raster file according to the corresponding geographic coordinates to form a continuous spatial distribution map of soil salinity. Furthermore, the continuous estimated values can be further mapped into hierarchical spatial expressions such as low salinity, medium salinity, and high salinity according to salinity evaluation standards, management zoning requirements, or preset grading thresholds, so as to facilitate spatial interpretation and farmland management applications.
[0234] Figure 9 This paper presents a spatial visualization representation of soil salinity estimation results. The agricultural remote sensing image of the area to be estimated is preprocessed to construct a multi-channel salinity-sensitive feature map, including salinity-sensitive indices such as the Visible-band Difference Vegetation Index (VDVI), Salinity Ratio Index (SAIO), Salinity Index (SI), and Normalized Difference Salinity Index (NDSI), along with related remote sensing features. The cropped feature map is then input into a trained lightweight single-backbone CNN continuous regression model. The model performs nonlinear fusion and continuous regression estimation on the multi-channel features to obtain the estimated soil salinity value for each spatial unit. Figure 9 Taking the estimation results after the typical salt sensitivity index was used to construct the feature as an example, the hierarchical spatial distribution of low-salt, medium-salt and high-salt areas is shown. It can be used to characterize the spatial heterogeneity of soil salt in farmland and the distribution of potential high-salt risk areas, and provide spatial reference for saline-alkali land management, irrigation and drainage regulation and field management decisions.
Claims
1. A lightweight, single-skeleton continuous soil salinity estimation method for small-sample agricultural remote sensing, comprising the following steps: (S1) Acquire agricultural remote sensing images of the target farmland area and measured soil salinity data corresponding to the agricultural remote sensing images in time and space, wherein the agricultural remote sensing images include RGB images and / or multispectral images; (S2) Preprocess the agricultural remote sensing images to generate a multi-channel salinity-sensitive feature map based on the preprocessed agricultural remote sensing images, characterized by further including the following steps: (S3) Based on the spatial location information of field sampling points, experimental plots or plot boundaries, the multi-channel salinity sensitive feature map is spatially cropped to obtain feature sub-map samples that correspond one-to-one with the measured soil salinity data, and a small sample agricultural remote sensing soil salinity estimation dataset including measured soil salinity data, feature sub-map samples and continuous salinity labels is constructed.
2. The lightweight single-skeleton continuous soil salinity estimation method for small-sample agricultural remote sensing according to claim 1, characterized in that, The measured soil salinity data includes continuous numerical salinity labels; the multi-channel salinity-sensitive feature map includes at least one of the following: original band features, vegetation index features, color index features, salinity-sensitive index features, and texture features.
3. The lightweight single-skeleton continuous soil salinity estimation method for small-sample agricultural remote sensing according to claim 1, characterized in that, Following step (S3), the following is also included: (S4) The small sample agricultural remote sensing soil salinity estimation dataset is divided into training set, validation set and test set according to the salinity numerical distribution. (S5) Construct a lightweight single-backbone convolutional neural network model, using the feature sub-map samples as input to the lightweight single-backbone convolutional neural network model and the continuous estimated value of soil salinity as output to construct the mapping relationship between measured soil salinity data and feature sub-map samples; the lightweight single-backbone convolutional neural network model includes a CNN backbone network, a global average pooling layer and a continuous value regression head connected in sequence.
4. The lightweight single-skeleton continuous soil salinity estimation method for small-sample agricultural remote sensing according to claim 3, characterized in that, Following step (S5), the following is also included: (S6) The lightweight single-backbone convolutional neural network model is trained using a training strategy oriented towards small sample continuous regression to obtain the trained lightweight single-backbone convolutional neural network model. The training strategy includes one of Huber loss function, regularization constraint, cross-validation and hyperparameter optimization. (S7) Input the agricultural remote sensing feature sub-map of the area to be estimated into the trained lightweight single-backbone convolutional neural network model, and output the continuous estimated value of soil salinity at the corresponding spatial location.
5. The lightweight single-skeleton continuous soil salinity estimation method for small-sample agricultural remote sensing according to claim 1, characterized in that, The spatial location information of the field sampling points, experimental plots, or plot boundaries includes geographic coordinates, sampling point number, experimental plot number, or plot boundary number; The continuous salinity label includes soil soluble salt content, soil electrical conductivity, apparent electrical conductivity converted salinity value or equivalent salinity index; The continuous salinity label is determined by measured soil salinity data from multiple soil layers. The average salinity values of the 10-20cm, 20-30cm, and 30-40cm soil layers are taken as the continuous salinity label for the root zone.
6. The lightweight single-skeleton continuous soil salinity estimation method for small-sample agricultural remote sensing according to claim 3, characterized in that, In step (S5), the CNN backbone network is one of the following: lightweight VGG type backbone, lightweight residual type backbone, or lightweight dense connection type backbone, and the lightweight single-backbone convolutional neural network model does not contain feature-level fusion structures of two or more parallel heterogeneous backbones. When the CNN backbone is a lightweight VGG backbone, the lightweight VGG backbone includes three convolutional blocks connected in sequence. The number of convolutional kernels in the three convolutional blocks are 64, 128 and 256 respectively. Each convolutional block includes a first 3×3 convolutional layer, a first batch normalization layer, a first ReLU activation layer, a second 3×3 convolutional layer, a second batch normalization layer, a second ReLU activation layer and a pooling layer connected in sequence.
7. The lightweight single-skeleton continuous soil salinity estimation method for small-sample agricultural remote sensing according to claim 6, characterized in that, When the CNN backbone is a lightweight residual backbone, the lightweight residual backbone includes an initial 3×3 convolutional layer and three residual stages connected in sequence. Each residual stage includes at least one residual block. The residual block includes a main branch and a shortcut connection branch. The main branch includes two 3×3 convolutional layers, a batch normalization layer, and a ReLU activation layer. When the spatial size or number of channels of the input feature map of the residual block is inconsistent with the output feature map of the main branch, a 1×1 convolutional layer is set in the shortcut connection branch to achieve matching of spatial size and number of channels.
8. The lightweight single-skeleton continuous soil salinity estimation method for small-sample agricultural remote sensing according to claim 6, characterized in that, When the CNN backbone is a lightweight densely connected backbone, the lightweight densely connected backbone includes an initial 3×3 convolutional layer, three dense blocks and a transition layer. A transition layer is set between every two dense blocks. Each dense block includes multiple bottleneck units. The bottleneck unit includes a 1×1 convolutional layer and a 3×3 convolutional layer connected in sequence. The transition layer includes a 1×1 convolutional layer, a Dropout layer and an average pooling layer connected in sequence.
9. The lightweight single-skeleton continuous soil salinity estimation method for small-sample agricultural remote sensing according to claim 4, characterized in that, The continuous value regression head includes at least one fully connected layer, a Dropout layer, and a linear output layer connected in sequence. The linear output layer has an output neuron for outputting a continuous estimate of soil salinity.
10. The lightweight single-skeleton continuous soil salinity estimation method for small-sample agricultural remote sensing according to claim 4, characterized in that, In step (S6), when the absolute value of the prediction error is less than or equal to the threshold δ, the Huber loss function uses a squared error term; when the absolute value of the prediction error is greater than the threshold δ, the Huber loss function uses a linear error term.
Citation Information
Patent Citations
Agricultural land secondary salinization monitoring and early warning method and system based on remote sensing
CN113125383A