An image recognition-based farmland quality monitoring method
Patent Information
- Application Number
- CN202610772410.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-08-18
AI Technical Summary
在复杂地形区域耕地边界通常呈现出连续变化且具有自然曲率特征,而现有插值方法在重构过程中容易引入棋盘效应或锯齿状边界伪影,使得生成的耕地质量分布图在边界位置出现明显失真,不仅影响耕地面积测算的精度,还会降低耕地地块级分析结果的可靠性,难以满足精细化管理与空间决策的需求
本发明通过构建Spectral-AwareUniRepLKNet主干特征提取机制,实现了分光谱组、自适应大感受野加方向感知卷积的协同建模方式,在保持广域耕地宏观结构建模能力的同时,有效抑制了深层特征中的高频信息过度平滑问题,通过基于耕地质量敏感度系数对输入通道进行分光谱组重构,使不同光谱通道按照其对盐渍化微斑、局部干旱胁迫斑块及病虫害异常区域的敏感程度进行差异化处理,结合耕地纹理复杂度系数对卷积核边长和扩张率进行动态选择,使网络在纹理复杂区域自动采用更大有效感受野,在纹理平稳区域保持适度感受野,在同一网络结构内实现多尺度感知能力的自适应分配,同时通过方向感知卷积对耕地结构主方向进行对齐建模使梯田走向、田垄分布及水系结构在特征提取阶段得到一致性表达,提升了微小退化病灶的响应强度,使仅占少量像素的高频异常区域在深层特征中仍可保持可分辨性。
Smart Images

Figure CN122597992A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of farmland quality monitoring technology, and in particular to a farmland quality monitoring method based on image recognition. Background Technology
[0002] With the rapid development of remote sensing imaging technology, UAV aerial photography technology, and multi-source geographic information acquisition methods, farmland quality monitoring is gradually shifting from traditional manual inspection to automated and intelligent analysis based on multi-source remote sensing images. In existing technologies, semantic segmentation and target recognition of wide-area remote sensing images typically employ feature extraction based on deep convolutional neural networks or self-attention mechanisms. To obtain greater spatial contextual information, the understanding of the overall geographic structure is often enhanced by increasing network depth, expanding the receptive field of convolution, or introducing multi-scale feature fusion strategies. However, in practical applications, as the number of network layers increases and the receptive field expands, features are prone to excessive smoothing of spatial information during multiple downsampling and convolution operations, leading to significant attenuation of high-frequency detail information. This makes it difficult to effectively identify small, scattered farmland degradation lesions. Small patches formed in the early stages of soil salinization, subtle texture changes caused by localized drought stress, and localized spectral anomalies caused by early pests and diseases are often masked by low-frequency background information in deep feature representations, resulting in serious missed detections and failing to meet the needs of refined farmland quality monitoring.
[0003] In existing multi-scale feature reconstruction processes, bilinear interpolation, nearest neighbor interpolation, or transposed convolution methods are commonly used to upsample and restore low-resolution semantic feature maps. These methods are essentially local, fixed-rule interpolation operations, lacking the ability to model the continuity of spatial structures and surface topology. In complex terrain areas, farmland boundaries typically exhibit continuous variations and natural curvature. However, existing interpolation methods are prone to introducing checkerboard effects or jagged boundary artifacts during reconstruction, resulting in significant distortion of the generated farmland quality distribution map at the boundary locations. This not only affects the accuracy of farmland area calculations but also reduces the reliability of farmland plot-level analysis results, making it difficult to meet the needs of refined management and spatial decision-making. Summary of the Invention
[0004] One objective of this invention is to propose an image recognition-based method for monitoring farmland quality. This invention can simultaneously improve both the consistency of macroscopic structure and the sensitivity to microscopic anomalies, significantly enhancing the precision of farmland quality monitoring results and the reliability of practical applications.
[0005] A method for monitoring farmland quality based on image recognition according to an embodiment of the present invention includes: Remote sensing datasets covering the target farmland monitoring area are collected and preprocessed to obtain a standardized remote sensing input dataset. This dataset is then jointly encoded to obtain a multimodal fusion data tensor, which is then divided into blocks to generate a farmland quality monitoring sample block sequence. The sequence of farmland quality monitoring sample blocks is input into the Spectral-AwareUniRepLKNet backbone feature extraction network, which outputs multi-scale deep semantic feature representations. We apply structural reparameterization fusion to multi-scale deep semantic feature representations, folding the multi-branch convolutional structure in the training stage into a single-path large-kernel convolutional structure in the inference stage, and outputting reparameterized semantic feature representations. The reparameterized semantic feature representation is subjected to hierarchical aggregation, cross-scale alignment and sliding gated residual fusion. Digital terrain elevation data is injected into the reparameterized semantic feature representation to obtain a fused feature map of farmland quality. The fused feature map of farmland quality is input into the Topo-AwareFeatUp high-resolution feature reconstruction module to perform dimensionality restoration on the fused feature map of farmland quality and output the final high-resolution semantic feature map. Farmland quality status classification is performed on high-resolution semantic feature maps at the pixel level, block level, and farmland plot level, and the multi-label segmentation results of farmland quality are output. The multi-label segmentation results of cultivated land quality are subjected to micro-region consistency screening and checkerboard artifact suppression to generate a smooth cultivated land quality distribution map. The smooth cultivated land quality distribution map is spatially overlaid with cultivated land plot boundary data and digital terrain elevation data to calculate the cultivated land quality grade, degradation area ratio and abnormal patch density of each cultivated land plot, and generate cultivated land plot-level cultivated land quality diagnosis results.
[0006] Optionally, the process of collecting remote sensing datasets covering the target cultivated land monitoring area and performing preprocessing operations includes: Multispectral remote sensing images, high-resolution orthophotos, digital topographic elevation data, historical temporal images, and farmland plot boundary data covering the target farmland monitoring area were collected. Spatial registration, radiometric correction, geometric correction, cloud and fog removal, noise suppression, resolution unification, and coordinate system unification were performed to obtain a standardized remote sensing input dataset. The spectral information, texture information, topographic gradient information, water system distribution information, and farmland plot morphology information in the standardized remote sensing input dataset were jointly encoded to obtain a multimodal fusion data tensor covering the target farmland monitoring area. The multimodal fusion data tensor was then divided into blocks according to a preset sliding window strategy to generate a farmland quality monitoring sample block sequence.
[0007] Optionally, the step of inputting the farmland quality monitoring sample block sequence into the Spectral-AwareUniRepLKNet backbone feature extraction network includes: For each input channel in each farmland quality monitoring sample block, the farmland quality sensitivity coefficient of the input channel is calculated. All input channels are rearranged and grouped according to the farmland quality sensitivity coefficient to form a sub-spectral group input representation. Calculate the farmland texture complexity coefficient for each spectral group in the sub-spectral group input representation. Based on the comparison between the farmland texture complexity coefficient and the preset threshold, select the actual convolution kernel side length and the actual convolution dilation rate from the corresponding candidate set. The effective receptive field size is obtained by multiplying the convolution dilation rate by the result of subtracting 1 from the convolution kernel side length and adding 1. Based on each farmland quality monitoring sample block, the main orientation angle of farmland structure is constructed. Based on the main orientation angle of farmland structure, the convolution sampling coordinates are rotated and transformed. Under the combined effect of the rotational sampling coordinates and the orientation unfolding modulation coefficient, orientation sensing large receptive field channel-by-channel convolutional encoding is independently performed on each channel in each spectral group to obtain the convolutional response within the orientation sensing group. A high-frequency anomaly residual map is constructed on the spectral group for the farmland quality monitoring sample block. Based on the high-frequency anomaly residual map, a farmland quality anomaly fidelity gating map is generated and weighted and combined with the reciprocal of the effective receptive field size to obtain the fidelity fusion coefficient. Based on the farmland quality anomaly fidelity gating map and the fidelity fusion coefficient, the convolutional response within the direction perception group and the original input features are weighted and fused to obtain the high-frequency preservation group convolutional response. The high-frequency intra-group convolutional responses of each spectral group in the same coding layer are concatenated to obtain an inter-group fusion feature tensor. The features of each spectral group are weighted by the contribution weight of the spectral group and then concatenated to obtain a weighted inter-group fusion feature tensor. Channel mapping and stage compression are performed on the weighted inter-group fusion feature tensor to obtain stage coding features. When the current coding layer is not the last coding layer, scale-progressive mapping is performed on the stage coding features to obtain scale-progressive features. When the current coding layer is the last coding layer, the stage coding features are directly output as scale-progressive features. The scale-progressive features output by each coding layer are defined as multi-scale deep semantic features with progressively decreasing spatial resolution. The outputs of each coding layer are then aggregated in hierarchical order to obtain the multi-scale deep semantic feature representation corresponding to each farmland quality monitoring sample block.
[0008] Optionally, the application of structure reparameterization fusion to the multi-scale deep semantic feature representation, folding the multi-branch convolutional structure in the training phase into a single-path large-kernel convolutional structure in the inference phase, includes: For each scale-progressive feature in the multi-scale deep semantic feature representation, a multi-branch channel-wise convolutional structure is constructed for the training phase. During the training phase, each branch of the multi-branch channel-wise convolutional structure is subjected to folding processing of convolution parameters and batch normalization parameters. The scaling and translation effects on the feature distribution in the batch normalization mapping are converted into equivalent corrections to the convolution kernel weights and convolution bias terms, resulting in independent equivalent convolution kernels and equivalent convolution bias terms for each branch. Perform convolution kernel size alignment processing to center-align the equivalent convolution kernels of the large convolution kernel main branch, the small convolution kernel compensation branch, and the identity mapping preservation branch under the actual convolution kernel side length dynamically configured by the farmland texture complexity coefficient. Then, perform element-wise tensor accumulation on the aligned equivalent convolution kernels and equivalent convolution bias terms to obtain the global fusion convolution kernel and global fusion convolution bias terms corresponding to the single-path large kernel convolution structure in the inference stage. Using global fusion convolution kernels and global fusion convolution bias terms as kernel weight basis, single-path large kernel convolution calculation is performed on the corresponding scale-progressive features under the combined effect of the sampling coordinates after orientation rotation and the orientation unfolding modulation coefficients to obtain reparameterized semantic feature representation; The reparameterized semantic feature representations corresponding to each coding layer are aggregated in hierarchical order to obtain the reparameterized semantic feature representation corresponding to each farmland quality monitoring sample block, and then aggregated to obtain the set of reparameterized semantic feature representations corresponding to the sequence of farmland quality monitoring sample blocks.
[0009] Optionally, the step of performing hierarchical aggregation, cross-scale alignment, and sliding-gated residual fusion on the reparameterized semantic feature representation to inject digital terrain elevation data into the reparameterized semantic feature representation includes: Hierarchical aggregation basis construction is performed on the reparameterized semantic feature representations at all levels, and spatial clipping and scale mapping are performed on the digital terrain elevation data corresponding to the nth cultivated land quality monitoring sample block to obtain the corresponding terrain injection features; Perform top-down hierarchical aggregation processing on the aggregated base features at each level to obtain the hierarchical aggregated features of the current coding layer; Perform cross-scale alignment processing on the aggregated features at each level to obtain cross-scale aligned features; Sliding gated residual fusion processing is performed on cross-scale aligned features and terrain-injected features to inject digital terrain elevation data into reparameterized semantic feature representations, thereby obtaining fused features of cultivated land quality at all levels. The fusion features of cultivated land quality at all levels are aggregated hierarchically into shallow, medium and deep layers to obtain a fusion feature map of cultivated land quality that includes shallow edge response, medium texture response and deep semantic response.
[0010] Optionally, the step of inputting the fused feature map of cultivated land quality into the Topo-AwareFeatUp high-resolution feature reconstruction module to perform dimensionality-upgrading recovery of the fused feature map of cultivated land quality includes: Construct high-resolution query coordinates consistent with the spatial resolution of the standardized remote sensing input dataset, and map the fused feature map of cultivated land quality into local implicit reconstruction conditional features in the continuous coordinate query space; The local implicit reconstruction conditional features are input into the task-independent high-resolution implicit reconstruction operator, and implicit semantic reconstruction is performed on each high-resolution query coordinate to obtain the initial high-resolution semantic feature map. A multi-view mapping result is constructed from the fused feature map of cultivated land quality, and a synchronous spatial transformation is performed on the high-resolution query coordinates. Based on the high-resolution implicit reconstruction result corresponding to each multi-view mapping result, a multi-view mapping is constructed. Figure 1 Consistency constraints; Constructing curvature responses and establishing curvature continuity constraints on the initial high-resolution semantic feature maps, and then applying them to multi-view... Figure 1 Consistency constraints and curvature continuity constraints are weighted and summed according to preset weights to obtain high-resolution implicit reconstruction constraints. Based on the high-resolution implicit reconstruction constraints, the task-independent high-resolution implicit reconstruction operator is updated online using the backpropagation algorithm. Using a task-independent high-resolution implicit reconstruction operator updated online with the current samples, full-resolution upscaling recovery is performed on the fused feature map of farmland quality, outputting a final high-resolution semantic feature map with the same spatial resolution as the standardized remote sensing input dataset.
[0011] Optionally, the step of performing farmland quality status classification on the high-resolution semantic feature map at the pixel granularity, block granularity, and farmland plot granularity respectively includes: Shared feature mapping is performed on the final high-resolution semantic feature map to obtain shared predictive basis features for multi-granularity farmland quality status classification; Pixel-level farmland quality status classification is performed on the shared prediction basis features. The independent classification probability of each abnormal category is calculated for each pixel location, and the classification probability of normal farmland is calculated based on the classification probability of all abnormal categories, resulting in a pixel-level multi-label classification probability map. Perform block-scale farmland quality status classification on the shared prediction basis features, and map the block-scale classification results back to the pixel space to obtain a block-scale multi-label classification probability map; Based on the boundary data of cultivated land plots, the cultivated land quality status classification at the granular level is performed on the shared prediction basis features. The granular classification results of cultivated land plots are mapped back to the pixel space to obtain a multi-label classification probability map of cultivated land plot granularity. Perform hybrid granularity multi-label fusion classification on pixel-level multi-label classification probability maps, block-level multi-label classification probability maps, and cultivated land plot-level multi-label classification probability maps to obtain cultivated land quality multi-label segmentation results.
[0012] Optionally, the process of performing micro-region consistency screening and checkerboard artifact suppression on the multi-label segmentation results of cultivated land quality to generate a smoothed cultivated land quality distribution map, and then performing spatial overlay analysis on the smoothed cultivated land quality distribution map with cultivated land parcel boundary data and digital terrain elevation data, includes: The multi-label segmentation results of cultivated land quality were subjected to micro-region consistency screening according to the abnormality category. Isolated micro-abnormal regions that were inconsistent with the surrounding spatial distribution were removed to obtain the multi-label segmentation results after screening. The checkerboard artifact suppression was performed on the screened multi-label segmentation results to obtain a smoothed multi-label classification probability map. A smoothed farmland quality distribution map was generated based on the smoothed multi-label classification probability. Spatial overlay analysis was performed on the smoothed farmland quality distribution map, farmland plot boundary data, and digital topographic elevation data to calculate the degradation area ratio, abnormal patch density, and topographic disturbance coefficient of each farmland plot. The farmland quality grade of each farmland plot is calculated based on the proportion of degraded area, density of abnormal patches, and topographic disturbance coefficient, generating farmland quality diagnosis results at the plot level.
[0013] Optionally, the farmland quality grade is determined based on the proportion of degraded area, the density of abnormal patches, and the topographic disturbance coefficient. When the proportion of degraded area is less than or equal to the preset high-quality threshold, the density of abnormal patches is less than or equal to the preset high-quality density threshold, and the topographic disturbance coefficient is less than or equal to the preset high-quality disturbance threshold, the corresponding cultivated land area will be identified as a high-quality cultivated land area. When the proportion of degraded area is within the preset good range or the density of abnormal patches is within the preset good range and the topographic disturbance coefficient is less than or equal to the preset good disturbance threshold, the corresponding cultivated land area will be judged as a good cultivated land area. When the proportion of degraded area is within the preset general range or the density of abnormal patches is within the preset general range and the topographic disturbance coefficient is within the preset general range, the corresponding cultivated land area will be identified as a general cultivated land area. When the proportion of degraded area is greater than the preset poor threshold, or the density of abnormal patches is greater than the preset poor density threshold, or the topographic disturbance coefficient is greater than the preset poor disturbance threshold, the corresponding cultivated land area will be identified as poor cultivated land.
[0014] The beneficial effects of this invention are: This invention constructs a Spectral-AwareUniRepLKNet backbone feature extraction mechanism, achieving a collaborative modeling approach combining spectral grouping, adaptive large receptive field, and direction-aware convolution. While maintaining the ability to model the macroscopic structure of cultivated land over a wide area, it effectively suppresses the problem of excessive smoothing of high-frequency information in deep features. By reconstructing the input channels into spectral groups based on the cultivated land quality sensitivity coefficient, different spectral channels are differentiated according to their sensitivity to salinization micro-pollutants, local drought stress patches, and abnormal areas caused by pests and diseases. Combined with the cultivated land texture complexity coefficient, the convolution kernel side length and expansion rate are dynamically selected, enabling the network to automatically adopt a larger effective receptive field in textured areas and maintain an appropriate receptive field in textured areas. This achieves adaptive allocation of multi-scale perception capabilities within the same network structure. At the same time, by aligning the main direction of cultivated land structure through direction-aware convolution, the terrace orientation, ridge distribution, and water system structure are consistently expressed during the feature extraction stage, improving the response intensity of small degenerative lesions and ensuring that high-frequency abnormal areas occupying only a few pixels remain distinguishable in deep features.
[0015] This invention achieves continuous coordinate implicit representation and multi-view functionality by introducing the Topo-AwareFeatUp high-resolution implicit reconstruction mechanism. Figure 1 A high-fidelity feature upscaling method based on consistency constraints and curvature continuity constraints is employed by constructing a continuous mapping relationship between high-resolution query coordinates and low-resolution features, through multi-view... Figure 1 Consistency constraints ensure that the reconstruction results under different spatial transformations remain consistent after inverse mapping, effectively suppressing the accumulation of spatial artifacts during the reconstruction process. Furthermore, curvature continuity constraints constrain the second-order changes of farmland boundaries and anomalous area edges, making the boundary transition smoother and more continuous. In complex terrain areas, continuous topological restoration of farmland boundaries can be achieved, improving the accuracy of farmland plot-level area calculation and boundary determination.
[0016] This invention constructs an end-to-end farmland quality monitoring system by coupling Spectral-AwareUniRepLKNet and Topo-AwareFeatUp across modules. It achieves a unified capability of global context awareness and local high-frequency detail recovery. Spectral-AwareUniRepLKNet performs efficient semantic compression on a wide area of farmland at the front end, extracting deep semantic features including irrigation patterns, plot connectivity, and regional environmental context. Topo-AwareFeatUp implicitly reconstructs these compressed features at the back end, recovering high-resolution structural details from the low-resolution semantic space. This makes high-frequency anomaly information that was weakened during the encoding process explicit again, ensuring both the sensitivity to detect minor anomalies and maintaining the stability of the overall spatial distribution. In complex farmland scenarios, it can simultaneously achieve a unified improvement in macroscopic structural consistency and microscopic anomaly sensitivity, significantly enhancing the refinement of farmland quality monitoring results and the reliability of practical applications. Attached Figure Description
[0017] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a method for monitoring farmland quality based on image recognition proposed in this invention; Figure 2 This is the processing flow of Spectral-AwareUniRepLKNet in the image recognition-based farmland quality monitoring method proposed in this invention. Detailed Implementation
[0018] Example 1: Reference Figures 1-2 A method for monitoring farmland quality based on image recognition, comprising: Remote sensing datasets covering the target farmland monitoring area are collected and preprocessed to obtain a standardized remote sensing input dataset. This dataset is then jointly encoded to obtain a multimodal fusion data tensor, which is then divided into blocks to generate a farmland quality monitoring sample block sequence. In this embodiment, remote sensing datasets covering the target cultivated land monitoring area are collected and preprocessed, including: Multispectral remote sensing images, high-resolution orthophotos, digital topographic elevation data, historical temporal images, and farmland plot boundary data covering the target farmland monitoring area were collected. Spatial registration, radiometric correction, geometric correction, cloud and fog removal, noise suppression, resolution unification, and coordinate system unification were performed to obtain a standardized remote sensing input dataset. The spectral information, texture information, topographic gradient information, water system distribution information, and farmland plot morphology information in the standardized remote sensing input dataset were jointly encoded to obtain a multimodal fusion data tensor covering the target farmland monitoring area. The multimodal fusion data tensor was then divided into blocks according to a preset sliding window strategy to generate a farmland quality monitoring sample block sequence.
[0019] In Example 1, the standardized remote sensing input dataset is subjected to spectral band mapping, texture feature extraction, digital terrain elevation data gradient solution, water system distribution feature extraction, and farmland plot boundary morphology feature extraction, respectively, to obtain mutually registered spectral feature layers, texture feature layers, terrain gradient feature layers, water system distribution feature layers, and farmland plot morphology feature layers. The spectral feature layers, texture feature layers, terrain gradient feature layers, water system distribution feature layers, and farmland plot morphology feature layers are then overlaid and jointly encoded according to channel dimension to obtain a multimodal fusion data tensor covering the target farmland monitoring area. The multimodal fusion data tensor is divided into blocks according to a preset sliding window strategy to generate a farmland quality monitoring sample block sequence.
[0020] Based on the spatial resolution of the standardized remote sensing input dataset, the minimum apparent size of microdegenerative lesions, the scale of changes in farmland plot boundaries, the plot size distribution characteristics of the target farmland monitoring area, and the input size requirements of the Spectral-AwareUniRepLKNet backbone feature extraction network, the sliding window size, horizontal step size, vertical step size, window overlap ratio, and edge patching method are predetermined to form a pre-defined sliding window strategy. The pre-defined sliding window strategy is used to divide the multimodal fusion data tensor into multiple local sample blocks with fixed spatial ranges, and by controlling the overlap area between adjacent local sample blocks, microdegenerative lesions, plot boundary transition areas, and complex terrain boundary areas in the target farmland monitoring area are repeatedly covered in adjacent local sample blocks.
[0021] The sequence of farmland quality monitoring sample blocks is input into the Spectral-AwareUniRepLKNet backbone feature extraction network, which outputs multi-scale deep semantic feature representations. In this embodiment, the sequence of farmland quality monitoring sample blocks is input into the Spectral-AwareUniRepLKNet backbone feature extraction network, including: For each input channel in each farmland quality monitoring sample block, the farmland quality sensitivity coefficient of the input channel is calculated. All input channels are rearranged and grouped according to the farmland quality sensitivity coefficient to form a sub-spectral group input representation. In Example 1, the sensitivity coefficient of cultivated land quality is obtained by weighted summation of the spatial spectral gradient response, historical phase difference response, and boundary bending coupling response corresponding to the input channel.
[0022] Differential operations are performed on the input channels in the horizontal and vertical directions, and the difference results are synthesized by amplitude. The gradient amplitude is averaged over the entire cultivated land quality monitoring sample block to obtain the spatial spectral gradient response, which is used to measure the high-frequency abrupt change intensity in space of salinization micro-patch, local drought stress patch, and early disease and pest spots.
[0023] The pixel values of the current time phase and the historical time phase at the same spatial location are differentially analyzed point by point. After the absolute value of the difference results is processed, they are averaged within the sample block of cultivated land quality monitoring to obtain the historical time phase difference response, which is used to measure the intensity of degradation and evolution of the same cultivated land area in the time dimension.
[0024] Boundary curves are extracted from farmland plot boundary data, and the directional change rate of the boundary curves at each location is calculated to obtain the boundary bending intensity distribution. The boundary bending intensity distribution is mapped to the spatial location corresponding to the c-th input channel. The pixel responses of the input channel within the boundary neighborhood are weighted and statistically analyzed to obtain the boundary bending coupling response, which is used to measure the influence intensity of farmland boundary fragmentation, plot fragmentation, and abnormal transition zones in the boundary neighborhood on the input channel.
[0025] The sub-spectral group input representation includes high-frequency anomaly sensitive spectral groups, structural continuity sensitive spectral groups, or background stability sensitive spectral groups, so that the sub-spectral group input representation directly carries the differences in the types of farmland quality anomalies.
[0026] Calculate the farmland texture complexity coefficient for each spectral group in the sub-spectral group input representation. Based on the comparison between the farmland texture complexity coefficient and the preset threshold, select the actual convolution kernel side length and the actual convolution dilation rate from the corresponding candidate set. The effective receptive field size is obtained by multiplying the convolution dilation rate by the result of subtracting 1 from the convolution kernel side length and adding 1. In Example 1, the farmland texture complexity coefficient is obtained by weighted summation of local spectral texture undulation intensity, abnormal patch aggregation intensity, and boundary tortuosity intensity; A sliding window is used in the spatial dimension of the feature tensor within the group. The mean value of the pixel values within the sliding window is calculated, and the squared deviation between each pixel value and the mean value within the sliding window is calculated. The squared deviations are summed and normalized to obtain the local variance value. The local variance values of all sliding windows are averaged to obtain the local spectral texture undulation intensity, which is used to measure the spatial undulation of soil salt crust fine texture, crop growth uneven texture, and local bare land texture.
[0027] Pixels exceeding a preset threshold are marked as abnormal pixels. Connected region marking is performed on the abnormal pixels. The number of pixels in each connected region is counted. The product of the average area of all connected regions and the number of regions is calculated to obtain the abnormal plaque aggregation intensity, which is used to measure the degree of spatial aggregation of microdegenerative lesions.
[0028] Discrete sampling is performed on the boundary curves of cultivated land plots. The directional change angle between adjacent sampling points is calculated, and the absolute values of the directional change angles are accumulated. The results are then normalized based on the boundary length to obtain the boundary tortuosity intensity, which is used to measure the degree of boundary tortuosity of hilly land, sloping cultivated land, and fragmented cultivated land.
[0029] When the spectral group category is labeled as a high-frequency anomaly sensitive spectral group, the convolution kernel candidate set is set to a discrete set containing convolution kernels with side lengths of 17 and 31, and the dilation rate candidate set is set to a discrete set containing dilation rates of 1 and 2; when the spectral group category is labeled as a structure continuity sensitive spectral group, the convolution kernel candidate set is set to a discrete set containing convolution kernels with side lengths of 31 and 41, and the dilation rate candidate set is set to a discrete set containing dilation rates of 2 and 3; when the spectral group category is labeled as a background stability sensitive spectral group, the convolution kernel candidate set is set to a single-element set containing only a convolution kernel with side length of 41, and the dilation rate candidate set is set to a single-element set containing only a dilation rate of 3.
[0030] Based on the comparison between the farmland texture complexity coefficient and the preset farmland texture complexity threshold, a specific convolution kernel side length is selected from the corresponding convolution kernel candidate set, and a specific dilation rate is selected from the corresponding dilation rate candidate set. When the farmland texture complexity coefficient is less than the farmland texture complexity threshold, the smaller value in the candidate set is selected; when the farmland texture complexity coefficient is greater than or equal to the farmland texture complexity threshold, the larger value in the candidate set is selected. The effective receptive field size is calculated based on the selected convolution kernel side length and dilation rate. The effective receptive field size is obtained by multiplying the dilation rate by the result of subtracting 1 from the convolution kernel side length and adding one. It is used to represent the spatial coverage of the current spectral group in the current coding layer.
[0031] Based on each farmland quality monitoring sample block, the main orientation angle of farmland structure is constructed. Based on the main orientation angle of farmland structure, the convolution sampling coordinates are rotated and transformed. Under the combined effect of the rotational sampling coordinates and the orientation unfolding modulation coefficient, orientation sensing large receptive field channel-by-channel convolutional encoding is independently performed on each channel in each spectral group to obtain the convolutional response within the orientation sensing group. In Example 1, the principal orientation angle of the cultivated land structure is obtained by calculating the structural tensor of the cultivated land quality monitoring sample block, and is used to represent the orientation of terraces, ridges, irrigation ditches, and the main extension direction of the plot boundary:
[0032] in, This represents the principal orientation angle of the farmland structure in the nth farmland quality monitoring sample block. and This represents the structure tensor component corresponding to the nth farmland quality monitoring sample block. This represents the stability constant to prevent the denominator from being zero.
[0033] The convolutional sampling coordinates are rotated based on the principal orientation angle of the cultivated land structure, and the direction unfolding modulation parameter corresponding to the spectral group category is linearly combined with the normalized ratio of the effective receptive field size in the current spectral group to control the degree of unfolding of convolutional sampling in the principal orientation of the cultivated land structure.
[0034] in, This represents the modulation coefficients of the nth farmland quality monitoring sample block in the direction corresponding to the rth spectral group at the I-th coding layer. Indicates the category label of the spectral group. The corresponding network can learnable direction of the modulation parameters. This represents the stability constant to prevent the denominator from being zero. This represents the effective receptive field size of the nth farmland quality monitoring sample block in the first-level coding layer and the rth spectral group.
[0035] Under the combined effect of the orientation-rotated sampling coordinates and the orientation-unfolded modulation coefficients, orientation-aware large receptive field channel-wise convolutional encoding is performed on each spectral group to obtain the orientation-aware group intra-convolutional response:
[0036] in, This indicates the position of the nth farmland quality monitoring sample block on the first-level coding layer and the rth spectral group. The convolutional response value within the orientation-aware group at that location. Represents a nonlinear activation mapping function. This represents the discrete sampling coordinates of the convolution kernel within its domain. Indicates the length of the convolution kernel For the defined set of convolution kernel sampling coordinates, This indicates that the convolution kernel of the nth farmland quality monitoring sample block in the I-level coding layer and the r-th spectral group is located at the sampling coordinates. The corresponding convolution weight parameters, This represents the feature value of the feature tensor within the input group of the nth farmland quality monitoring sample block at the specified coordinate position in the (I-1)th level coding layer and the rth spectral group. When I=1, it represents the original sub-spectral group input feature. This represents the convolution dilation rate of the nth farmland quality monitoring sample block in the I-level coding layer and the r-th spectral group. This represents the lateral rotation sampling coordinate component obtained by rotating the original sampling u-coordinates of the nth farmland quality monitoring sample block under the constraint of the main direction of farmland structure. This represents the longitudinal rotational sampling coordinate component obtained by rotating the original sampling coordinates v of the nth farmland quality monitoring sample block under the constraint of the main direction of farmland structure. This represents the convolution bias term of the nth farmland quality monitoring sample block on the first-level coding layer and the rth spectral group.
[0037] Each channel within each spectral group is independently encoded using a direction-aware, large receptive field channel-wise convolutional encoding. This ensures that the spectral features of different channels do not interfere with each other during spatial encoding, and that the convolutional response within the direction-aware group maintains spatial and channel dimensions consistent with the original input feature tensor.
[0038] A high-frequency anomaly residual map is constructed on the spectral group for the farmland quality monitoring sample block. Based on the high-frequency anomaly residual map, a farmland quality anomaly fidelity gating map is generated and weighted and combined with the reciprocal of the effective receptive field size to obtain the fidelity fusion coefficient. Based on the farmland quality anomaly fidelity gating map and the fidelity fusion coefficient, the convolutional response within the direction perception group and the original input features are weighted and fused to obtain the high-frequency preservation group convolutional response. In Example 1, to suppress excessive smoothing of high-frequency information of small degenerative lesions during the large receptive field convolutional coding process, a high-frequency anomaly residual map is constructed for each cultivated land quality monitoring sample block on each coding layer and each spectral group. The high-frequency anomaly residual map is obtained by calculating the absolute value of the difference between the input feature of the current spectral group and its three-by-three neighborhood average result, and is used to measure the high-frequency differences of the edges of salinized micro-patch, drought stress crack edges, early-stage small patch edges of pests and diseases, and narrow strip areas of boundary fracture zones.
[0039] Based on the high-frequency anomaly residual map, a fertile land quality anomaly fidelity gated map is obtained by convolutional mapping and weighted modulation by combining the fertile land texture complexity coefficient and the intra-group average fertile land quality sensitivity coefficient.
[0040] By weighting and combining the farmland quality anomaly fidelity gating map with the inverse of the effective receptive field size, a fidelity fusion coefficient is calculated. Based on this fidelity fusion coefficient, the orientation-aware intra-group convolutional response and the original input features with consistent dimensionality due to the use of channel-wise convolution are weighted and fused using residuals to obtain the high-frequency intra-group convolutional response:
[0041] in, This represents the high-frequency hold-up intra-group convolution response of the nth farmland quality monitoring sample block in the I-level coding layer and the r-th spectral group. This represents element-wise multiplication. This represents the fidelity fusion coefficient.
[0042] The high-frequency intra-group convolutional responses of each spectral group in the same coding layer are concatenated to obtain an inter-group fusion feature tensor. The features of each spectral group are weighted by the contribution weight of the spectral group and then concatenated to obtain a weighted inter-group fusion feature tensor. Channel mapping and stage compression are performed on the weighted inter-group fusion feature tensor to obtain stage coding features. When the current coding layer is not the last coding layer, scale-progressive mapping is performed on the stage coding features to obtain scale-progressive features. When the current coding layer is the last coding layer, the stage coding features are directly output as scale-progressive features. In Example 1, the average farmland quality sensitivity coefficient, farmland texture complexity coefficient, and effective receptive field size within the group are normalized respectively. The normalized average farmland quality sensitivity coefficient, normalized farmland texture complexity coefficient, and normalized effective receptive field size are then weighted and summed according to a preset weighting coefficient to obtain the comprehensive response value of the spectral group.
[0043] The scale-progressive features output by each coding layer are defined as multi-scale deep semantic features with progressively decreasing spatial resolution. The outputs of each coding layer are then aggregated in hierarchical order to obtain the multi-scale deep semantic feature representation corresponding to each farmland quality monitoring sample block.
[0044] This implementation constructs a Spectral-Aware UniRepLKNet backbone architecture, achieving deep decoupling and reconstruction of remote sensing physical priors and deep learning underlying operators. Compared with the existing UniRepLKNet, it evolves fixed-dimensional channel grouping into physical sensitivity-based spectral group reorganization driven by spectral gradient, temporal differences, and boundary bending response. This enables directional perception of farmland quality anomalies from the initial feature extraction stage. It introduces a direction-aware, channel-wise large-kernel convolution explicitly guided by the structural tensor, combined with an adaptive kernel growth mechanism and directional unfolding modulation coefficients, solving the problem that traditional large-kernel convolution is difficult to converge and easily loses topological continuity under complex geographic textures. It also designs a high-frequency anomaly fidelity-gated residual structure coupled with the inverse of the effective receptive field size, effectively reversing the smoothing effect of large receptive field encoding on small lesion signals.
[0045] The Spectral-AwareUniRepLKNet backbone architecture not only significantly improves the model's collaborative recognition accuracy of macroscopic connectivity patterns and microscopic fragmented lesions in wide-area farmland monitoring, but also reduces inference power consumption on edge computing devices through deep separable convolution and structural reparameterization techniques.
[0046] We apply structural reparameterization fusion to multi-scale deep semantic feature representations, folding the multi-branch convolutional structure in the training stage into a single-path large-kernel convolutional structure in the inference stage, and outputting reparameterized semantic feature representations. In this embodiment, structural reparameterization fusion is applied to multi-scale deep semantic feature representations, folding the multi-branch convolutional structure in the training phase into a single-path large-kernel convolutional structure in the inference phase, including: For each scale-progressive feature in the multi-scale deep semantic feature representation, a multi-branch channel-wise convolutional structure is constructed for the training phase. In Example 1, the multi-branch channel-wise convolutional structure during the training phase includes a dynamic large convolutional kernel main branch, a small convolutional kernel compensation branch, and an identity mapping preservation branch.
[0047] The scale-progressive features output by the nth farmland quality monitoring sample block at the l-th level coding layer are used as the input features of the multi-branch channel-wise convolutional structure during the training phase. Channel-wise convolution operations are performed on the input features, including the main branch with a large convolution kernel and the same number of input feature channels, the compensation branch with a small convolution kernel, and the identity mapping-preserving branch. Batch normalization mapping is then performed in each branch to obtain the main branch response features, the compensation branch response features, and the preservation branch response features.
[0048] During the training phase, each branch of the multi-branch channel-wise convolutional structure is subjected to folding processing of convolution parameters and batch normalization parameters. The scaling and translation effects on the feature distribution in the batch normalization mapping are converted into equivalent corrections to the convolution kernel weights and convolution bias terms, resulting in independent equivalent convolution kernels and equivalent convolution bias terms for each branch. In Example 1, the scaling parameters, translation parameters, mean parameters, standard deviation parameters, and stability constants in the batch normalization mapping of each branch are analyzed. The operations of normalizing, scaling, and translating the input features in the batch normalization mapping are transformed into equivalent forms of scaling the convolutional kernel weights and linearly correcting the convolutional bias term. The convolutional kernel weight scaling process and convolutional bias term correction process are performed on the main branch with large convolutional kernel and the compensation branch with small convolutional kernel, respectively, to obtain the equivalent convolutional kernel and equivalent convolutional bias term of the main branch with large convolutional kernel, as well as the equivalent convolutional kernel and equivalent convolutional bias term of the compensation branch with small convolutional kernel. The same scaling and correction process is performed on the identity mapping-preserving branch, which is equivalent to a convolutional kernel with the center as the unit response and the rest as zero. Batch normalization folding is then performed on it to obtain the equivalent convolutional kernel and equivalent convolutional bias term of the identity mapping-preserving branch.
[0049] Perform convolution kernel size alignment processing to center-align the equivalent convolution kernels of the large convolution kernel main branch, the small convolution kernel compensation branch, and the identity mapping preservation branch under the actual convolution kernel side length dynamically configured by the farmland texture complexity coefficient. Then, perform element-wise tensor accumulation on the aligned equivalent convolution kernels and equivalent convolution bias terms to obtain the global fusion convolution kernel and global fusion convolution bias terms corresponding to the single-path large kernel convolution structure in the inference stage. Using global fusion convolution kernels and global fusion convolution bias terms as kernel weight basis, single-path large kernel convolution calculation is performed on the corresponding scale-progressive features under the combined effect of the sampling coordinates after orientation rotation and the orientation unfolding modulation coefficients to obtain reparameterized semantic feature representation; In Example 1, the scale-progressive features output from the l-th level coding layer of the nth farmland quality monitoring sample block are input into the global fusion convolution kernel. Convolution is performed in the sampling coordinate system after the orientation rotation, combined with the main orientation angle of the farmland structure. A global fusion convolution bias term is superimposed on the convolution result. Feature transformation is performed through a nonlinear activation mapping function to obtain a reparameterized semantic feature representation. The reparameterized semantic feature representation is numerically equivalent to the sum of the channel-wise output results of the main branch of the large convolution kernel, the compensation branch of the small convolution kernel, and the identity mapping-preserving branch under the same orientation constraint during the training phase. This achieves a lossless conversion from a multi-branch structure in the training phase to a single-path large kernel convolution structure that takes into account orientation awareness in the inference phase.
[0050] The reparameterized semantic feature representations corresponding to each coding layer are aggregated in hierarchical order to obtain the reparameterized semantic feature representation corresponding to each farmland quality monitoring sample block, and then aggregated to obtain the set of reparameterized semantic feature representations corresponding to the sequence of farmland quality monitoring sample blocks.
[0051] The reparameterized semantic feature representation is subjected to hierarchical aggregation, cross-scale alignment and sliding gated residual fusion. Digital terrain elevation data is injected into the reparameterized semantic feature representation to obtain a fused feature map of farmland quality. In this embodiment, hierarchical aggregation, cross-scale alignment, and sliding-gated residual fusion are performed on the reparameterized semantic feature representation, and digital terrain elevation data is injected into the reparameterized semantic feature representation, including: Hierarchical aggregation basis construction is performed on the reparameterized semantic feature representations at all levels, and spatial clipping and scale mapping are performed on the digital terrain elevation data corresponding to the nth cultivated land quality monitoring sample block to obtain the corresponding terrain injection features; In Example 1, channel mapping processing is performed on the reparameterized semantic feature representations corresponding to each level of coding layer for the nth farmland quality monitoring sample block to complete the construction of the hierarchical aggregated base and obtain the hierarchical aggregated base features at each level.
[0052] The digital terrain elevation data corresponding to the nth farmland quality monitoring sample block is cropped according to the spatial coverage of the farmland quality monitoring sample block to obtain a local digital terrain elevation block. The elevation change rate in the horizontal and vertical directions of the local digital terrain elevation block is calculated in the spatial dimension, and the elevation change rate is synthesized and averaged within the farmland quality monitoring sample block to obtain the terrain slope response.
[0053] The second-order rate of change of the local digital terrain elevation block is calculated in the spatial dimension and normalized to obtain the terrain curvature response. The local digital terrain elevation block, terrain slope response and terrain curvature response are concatenated according to the channel dimension to obtain the terrain representation tensor. The terrain representation tensor is downsampled with the same spatial resolution as the l-th level coding layer, and the downsampled result is channel mapped to obtain the terrain injection feature corresponding to the l-th level coding layer, which is strictly consistent with the spatial resolution and channel dimension of each level coding layer.
[0054] Perform top-down hierarchical aggregation processing on the aggregated base features at each level to obtain the hierarchical aggregated features of the current coding layer; In Example 1, channel mapping processing is performed on the reparameterized semantic feature representation corresponding to the l-th level coding layer of the nth farmland quality monitoring sample block to obtain the hierarchical aggregated base features.
[0055] The hierarchical aggregation base features corresponding to the highest level coding layer are directly used as the highest level aggregation features. The remaining coding layers are aggregated step by step in the order from the highest level to the lowest level. The hierarchical aggregation features of the previous level are mapped to the spatial resolution of the current coding layer through the scale upsampling function, and then added element by element with the hierarchical aggregation base features of the current coding layer to obtain the hierarchical aggregation features of the current coding layer.
[0056] By using top-down hierarchical aggregation, the global irrigation pattern response, farmland connectivity response, and regional environmental context response in the high-level reparameterized semantic feature representation are progressively transmitted to the low-level features, while maintaining consistency in spatial location correspondence.
[0057] Perform cross-scale alignment processing on the aggregated features at each level to obtain cross-scale aligned features; In Example 1, the hierarchical aggregation features corresponding to the l-th level coding layer and the hierarchical aggregation features corresponding to the l+1-th level coding layer of the nth farmland quality monitoring sample block are concatenated, and convolutional mapping is performed on the concatenated features to obtain the cross-scale alignment offset.
[0058] Based on the cross-scale alignment offset, spatial resampling is performed on the upper-level aggregated features after scale upsampling mapping, so that the upper-level aggregated features are aligned with the current coding layer's level aggregated features in spatial position. The current coding layer's level aggregated features are then added element-wise with the cross-scale aligned upper-level aggregated features to obtain the cross-scale aligned features. For the highest level coding layer, its level aggregated features are directly output as the cross-scale aligned features.
[0059] Sliding gated residual fusion processing is performed on cross-scale aligned features and terrain-injected features to inject digital terrain elevation data into reparameterized semantic feature representations, thereby obtaining fused features of cultivated land quality at all levels. In Example 1, the cross-scale aligned features corresponding to the l-th level coding layer of the nth farmland quality monitoring sample block are concatenated with the terrain injection features, and convolutional mapping is performed on the concatenated features to generate a sliding gating map through the Sigmoid mapping function.
[0060] The terrain injection features are weighted element-wise based on a sliding gating map, and the reparameterized semantic feature representation after channel dimension alignment is weighted element-wise based on the complement of the sliding gating map. This ensures that all feature components participating in the residual overlay are consistent in the channel dimension. The two are then residually overlaid with the cross-scale aligned features to obtain the fused farmland quality features.
[0061] in, This represents the fused farmland quality feature corresponding to the nth farmland quality monitoring sample block at the I-level coding layer. This represents the cross-scale alignment feature corresponding to the nth farmland quality monitoring sample block at the 1st level coding layer. This represents the sliding gating graph corresponding to the nth farmland quality monitoring sample block in the I-level coding layer. This represents the terrain injection feature corresponding to the nth farmland quality monitoring sample block in the I-level coding layer. This represents the reparameterized semantic feature representation of the nth farmland quality monitoring sample block obtained after structural reparameterization fusion at the I-level coding layer. This represents the element-wise multiplication operator.
[0062] By using sliding gated residual fusion processing, the digital terrain elevation data responses of areas with abrupt slope changes, undulating areas at the boundaries of land parcels, and gully erosion micro-topography areas are selectively injected into the reparameterized semantic feature representation, while maintaining the structural continuity of the original semantic features.
[0063] The fusion features of cultivated land quality at all levels are aggregated hierarchically into shallow, medium and deep layers to obtain a fusion feature map of cultivated land quality that includes shallow edge response, medium texture response and deep semantic response.
[0064] In Example 1, the farmland quality fusion feature corresponding to the first-level coding layer is used as the shallow edge response feature, the farmland quality fusion feature corresponding to the intermediate coding layer between the first-level coding layer and the Lth-level coding layer is used as the middle texture response feature, and the farmland quality fusion feature corresponding to the Lth-level coding layer is used as the deep semantic response feature. Upsampling processing is performed on the middle texture response feature and the deep semantic response feature to make their spatial resolution consistent with that of the first-level coding layer, and they are concatenated with the shallow edge response feature in the channel dimension to obtain the farmland quality fusion feature map.
[0065] The fused feature map of farmland quality is input into the Topo-AwareFeatUp high-resolution feature reconstruction module to perform dimensionality restoration on the fused feature map of farmland quality and output the final high-resolution semantic feature map. In this embodiment, the fused feature map of cultivated land quality is input into the Topo-AwareFeatUp high-resolution feature reconstruction module to perform dimensionality-upgrading recovery of the fused feature map of cultivated land quality, including: Construct high-resolution query coordinates consistent with the spatial resolution of the standardized remote sensing input dataset, and map the fused feature map of cultivated land quality into local implicit reconstruction conditional features in the continuous coordinate query space; In Example 1, a high-resolution query coordinate set with the same spatial resolution as the standardized remote sensing input dataset is constructed. The high-resolution query coordinate set consists of all pixel positions within the spatial range of the cultivated land quality monitoring sample block, and each high-resolution query coordinate corresponds to a target pixel position.
[0066] Each high-resolution query coordinate is projected onto the continuous coordinate space of the fused feature map of farmland quality according to the spatial scale mapping relationship to obtain low-resolution projected coordinates. A fixed number of local support coordinates are selected in the local neighborhood of each low-resolution projected coordinate. The corresponding local support features are extracted for each local support coordinate, and the relative displacement vector between the high-resolution query coordinate and the local support coordinate is calculated. All local support features and the corresponding relative displacement vectors are concatenated and weighted according to normalized weights to obtain the local implicit reconstruction condition features corresponding to the high-resolution query coordinates. The local implicit reconstruction condition features are used to establish the correspondence between the shallow edge response, the middle texture response and the deep semantic response in the fused feature map of farmland quality in the continuous coordinate space.
[0067] The local implicit reconstruction conditional features are input into the task-independent high-resolution implicit reconstruction operator, and implicit semantic reconstruction is performed on each high-resolution query coordinate to obtain the initial high-resolution semantic feature map. In Example 1, the task-independent high-resolution implicit reconstruction operator is constructed as a continuous function approximation model based on multi-layer nonlinear mapping. The continuous function approximation model takes local implicit reconstruction condition features as input and is constructed by combining multi-layer pointwise fully connected mapping and nonlinear activation mapping functions. The structure includes an input mapping layer, an implicit feature transformation layer and an output mapping layer.
[0068] The local implicit reconstruction conditional features corresponding to the a-th high-resolution query coordinates are input into the input mapping layer. Channel compression and feature standardization are performed on the local implicit reconstruction conditional features to obtain the initial implicit representation features. The initial implicit representation features are then sequentially input into multiple implicit feature transformation layers. Each implicit feature transformation layer enhances the continuous spatial representation of the features through point-by-point linear mapping and nonlinear activation mapping functions, so that the local implicit reconstruction conditional features form a stable functional expression relationship in the continuous coordinate space. The features processed by the implicit feature transformation layers are input into the output mapping layer, which maps the implicit representation features into the corresponding high-resolution semantic feature vectors. All high-resolution semantic feature vectors are rearranged according to their spatial position order to obtain the initial high-resolution semantic feature map.
[0069] The task-agnostic high-resolution implicit reconstruction operator does not introduce farmland quality category label information during the construction process. The parameters are learned only through the mapping relationship between local implicit reconstruction condition features and high-resolution query coordinates, enabling the task-agnostic high-resolution implicit reconstruction operator to have a unified reconstruction capability for different farmland quality anomaly types.
[0070] A multi-view mapping result is constructed from the fused feature map of cultivated land quality, and a synchronous spatial transformation is performed on the high-resolution query coordinates. Based on the high-resolution implicit reconstruction result corresponding to each multi-view mapping result, a multi-view mapping is constructed. Figure 1 Consistency constraints; In Example 1, each multi-view mapping result is obtained by performing rotation, scaling, or flipping transformations on the fused feature map of cultivated land quality, and the same spatial transformation operation is performed on the high-resolution query coordinate set to maintain strict spatial alignment between the query coordinates and the feature map. Local implicit reconstruction conditional feature construction processing and high-resolution implicit reconstruction processing are performed on each multi-view mapping result and the aligned high-resolution query coordinates respectively to obtain the high-resolution semantic feature map corresponding to each view.
[0071] The high-resolution semantic feature maps corresponding to each view are mapped back to the original view coordinate system through the corresponding inverse spatial transformation. For each high-resolution query coordinate position, the difference value between the reconstruction results under different views is calculated. The difference values of all query coordinate positions are then averaged to obtain the multi-view result. Figure 1 Consistency constraints are used to measure the degree of difference in farmland boundaries, anomalous lesion outlines, and plot topological transition regions obtained from reconstruction under different views after inverse mapping.
[0072] in, This represents the multiview corresponding to the nth farmland quality monitoring sample block. Figure 1 Consistency constraints Indicates the number of multiple views. This represents the spatial height of the nth farmland quality monitoring sample block. The width of the nth farmland quality monitoring sample block is represented in pixels, and v represents the view index variable, used to identify the vth multi-view mapping result. This represents the high-resolution query coordinate index variable, used to identify the position of the a-th high-resolution query coordinate. This represents the high-resolution semantic feature vector corresponding to the a-th high-resolution query coordinate in the n-th cultivated land quality monitoring sample block. This represents the high-resolution semantic feature vector corresponding to the th high-resolution query coordinate a of the nth farmland quality monitoring sample block under the vth view. This represents the inverse spatial mapping function corresponding to the nth farmland quality monitoring sample block in the vth view. This represents the norm operation.
[0073] Constructing curvature responses and establishing curvature continuity constraints on the initial high-resolution semantic feature maps, and then applying them to multi-view... Figure 1 Consistency constraints and curvature continuity constraints are weighted and summed according to preset weights to obtain high-resolution implicit reconstruction constraints. Based on the high-resolution implicit reconstruction constraints, the task-independent high-resolution implicit reconstruction operator is updated online using the backpropagation algorithm. In Example 1, to prevent overflow of high-resolution query coordinates during differential calculation, edge mirroring is performed on the initial high-resolution semantic feature map. Second-order differential calculations are then performed on the padded initial high-resolution semantic feature map in both the horizontal and vertical directions. For each high-resolution query coordinate, the second-order differential value in the horizontal and vertical directions is calculated, and the absolute values are summed to obtain the curvature response corresponding to the high-resolution query coordinates. This response is used to measure the degree of change in the farmland boundary curve and the degree of edge bending in abnormal areas.
[0074] in, This represents the curvature response value corresponding to the a-th high-resolution query coordinate in the n-th cultivated land quality monitoring sample block. This represents the high-resolution semantic feature vector corresponding to the a-th high-resolution query coordinate in the n-th cultivated land quality monitoring sample block. This represents the high-resolution semantic feature vector of the nth farmland quality monitoring sample block, after being positively offset by one unit neighborhood in the horizontal direction based on the ath high-resolution query coordinates. This represents the high-resolution semantic feature vector of the nth farmland quality monitoring sample block, after being negatively offset by one unit neighborhood in the horizontal direction based on the ath high-resolution query coordinates. This represents the high-resolution semantic feature vector of the nth farmland quality monitoring sample block, after being positively offset by one unit neighborhood in the vertical direction based on the ath high-resolution query coordinates. It represents the high-resolution semantic feature vector of the nth farmland quality monitoring sample block after being negatively offset by one unit neighborhood in the vertical direction based on the ath high-resolution query coordinates.
[0075] The curvature response difference between adjacent high-resolution query coordinates is statistically analyzed, and the difference between all adjacent query coordinate pairs is averaged to obtain the curvature continuity constraint, which is then applied to multi-view queries. Figure 1 Consistency constraints and curvature continuity constraints are weighted and summed according to preset weights to obtain high-resolution implicit reconstruction constraints. During the forward inference monitoring phase of the current farmland quality monitoring sample block, based on the high-resolution implicit reconstruction constraints, the task-independent high-resolution implicit reconstruction operator is updated online in multiple rounds using the backpropagation algorithm and gradient descent optimizer to ensure that the reconstruction result simultaneously satisfies multi-view constraints in the high-resolution space. Figure 1 Consistency and boundary curvature continuity.
[0076] Using a task-independent high-resolution implicit reconstruction operator updated online with the current samples, full-resolution upscaling recovery is performed on the fused feature map of farmland quality, outputting a final high-resolution semantic feature map with the same spatial resolution as the standardized remote sensing input dataset.
[0077] In Example 1, the task-independent high-resolution implicit reconstruction operator, updated online by the current sample, is used to re-perform implicit semantic reconstruction on each high-resolution query coordinate, thereby obtaining the final high-resolution semantic feature vector corresponding to each high-resolution query coordinate.
[0078] The final high-resolution semantic feature vectors corresponding to all high-resolution query coordinates are rearranged according to spatial location order to restore them to a two-dimensional structure consistent with the spatial resolution of the standardized remote sensing input dataset, while maintaining consistent feature representation in the channel dimension, thus obtaining the final high-resolution semantic feature map.
[0079] Farmland quality status classification is performed on high-resolution semantic feature maps at the pixel level, block level, and farmland plot level, and the multi-label segmentation results of farmland quality are output. In this embodiment, farmland quality status classification is performed on the high-resolution semantic feature map at the pixel level, block level, and farmland plot level, respectively, including: Shared feature mapping is performed on the final high-resolution semantic feature map to obtain shared predictive basis features for multi-granularity farmland quality status classification; Pixel-level farmland quality status classification is performed on the shared prediction basis features. The independent classification probability of each abnormal category is calculated for each pixel location, and the classification probability of normal farmland is calculated based on the classification probability of all abnormal categories, resulting in a pixel-level multi-label classification probability map. The pixel-level classification of farmland quality status is as follows: farmland quality status categories are predefined, and categories other than normal farmland are defined as abnormal categories. Channel-wise classification mapping is performed on the shared prediction basis features at each pixel location to obtain the pixel-level classification response value corresponding to each abnormal category. Independent Sigmoid mapping is performed on the pixel-level classification response value of each abnormal category to obtain the pixel-level classification probability of the corresponding abnormal category.
[0080] By subtracting the probability from each of the abnormal categories and multiplying them together, the pixel-level classification probability of normal farmland is obtained. The classification probabilities of each category corresponding to all pixel positions are rearranged according to spatial location to obtain a pixel-level multi-label classification probability map, so that the same pixel position corresponds to multiple abnormal categories at the same time.
[0081] Perform block-scale farmland quality status classification on the shared prediction basis features, and map the block-scale classification results back to the pixel space to obtain a block-scale multi-label classification probability map; The classification of farmland quality status at the block-level granularity is as follows: The farmland quality monitoring sample block is divided into multiple non-overlapping block regions according to a preset fixed size. Average pooling is performed on the shared prediction basis features corresponding to all pixels in each block region to obtain block-level granularity aggregate features. Classification mapping is performed on the block-level granularity aggregate features to obtain the block-level granularity classification response value corresponding to each anomaly category. Independent Sigmoid mapping is performed on the block-level granularity classification response value of each anomaly category to obtain the block-level granularity classification probability of the corresponding anomaly category.
[0082] For each abnormal category probability, subtract the probability from the given probability and multiply them together to obtain the block-level classification probability of normal farmland. Then, fill the block-level classification probability back into all pixel positions within the corresponding block area, so that each pixel position inherits the classification probability of its block area, thus obtaining a block-level multi-label classification probability map.
[0083] Based on the boundary data of cultivated land plots, the cultivated land quality status classification at the granular level is performed on the shared prediction basis features. The granular classification results of cultivated land plots are mapped back to the pixel space to obtain a multi-label classification probability map of cultivated land plot granularity. The classification of arable land quality status at the granular level is as follows: Based on the boundary data of arable land plots, the nth arable land quality monitoring sample block is divided into multiple arable land plot regions. For background pixels not included in the boundary of arable land plots, the classification probability of each anomaly category is set to zero by default. Average pooling is performed on the shared prediction basis features corresponding to all pixels in each arable land plot region to obtain the granular aggregate features of arable land plots. Classification mapping is performed on the granular aggregate features of arable land plots to obtain the granular classification response value of arable land plots corresponding to each anomaly category. Independent Sigmoid mapping is performed on the granular classification response value of arable land plots for each anomaly category to obtain the granular classification probability of arable land plots for the corresponding anomaly category.
[0084] The probability of each abnormal category is subtracted from the probability and then multiplied together to obtain the granular classification probability of normal cultivated land. The granular classification probability of cultivated land is then filled back into all pixel positions within the corresponding cultivated land area, so that each pixel position inherits the classification probability of its respective cultivated land area, thus obtaining a multi-label classification probability map of cultivated land granularity.
[0085] Perform hybrid granularity multi-label fusion classification on pixel-level multi-label classification probability maps, block-level multi-label classification probability maps, and cultivated land plot-level multi-label classification probability maps to obtain cultivated land quality multi-label segmentation results.
[0086] In Example 1, linear mapping and Softmax normalization are performed on the shared prediction basis features to generate pixel-level fusion weights, block-level fusion weights, and farmland-level fusion weights for each anomaly category, so that the sum of the fusion weights of the three granularities is one.
[0087] For each anomaly category, the corresponding pixel-level classification probability, block-level classification probability, and farmland-level classification probability are extracted at the same pixel location. Then, based on the pixel-level fusion weight, block-level fusion weight, and farmland-level fusion weight, a weighted summation process is performed on the three granularity classification probabilities to obtain the final multi-label classification probability of the anomaly category at the current pixel location.
[0088] Based on the final multi-label classification probabilities corresponding to all anomaly categories, the final multi-label classification probability of normal cultivated land is calculated for the current pixel position. The determination of normal cultivated land and all anomaly categories maintain a logical complementary relationship.
[0089] For each anomaly category, a corresponding preset empirical multi-label judgment threshold is set, and an independent threshold judgment is performed for each anomaly category. When the final multi-label classification probability of the corresponding anomaly category is greater than or equal to the multi-label judgment threshold of the anomaly category, the multi-label judgment result of the corresponding anomaly category is set to 1; otherwise, it is set to 0, so that each anomaly category has independent activation capability at the same pixel position.
[0090] When all the multi-label judgment results corresponding to the abnormal categories are zero, the multi-label judgment result corresponding to the normal cultivated land is set to 1. When the multi-label judgment result corresponding to at least one abnormal category is 1, the multi-label judgment result corresponding to the normal cultivated land is set to 0.
[0091] The multi-label judgment results corresponding to all cultivated land quality status categories are rearranged according to spatial location to obtain the cultivated land quality multi-label segmentation result corresponding to each cultivated land quality monitoring sample block.
[0092] The multi-label segmentation results of cultivated land quality are subjected to micro-region consistency screening and checkerboard artifact suppression to generate a smooth cultivated land quality distribution map. The smooth cultivated land quality distribution map is spatially overlaid with cultivated land plot boundary data and digital terrain elevation data to calculate the cultivated land quality grade, degradation area ratio and abnormal patch density of each cultivated land plot, and generate cultivated land plot-level cultivated land quality diagnosis results.
[0093] In this embodiment, the multi-label segmentation results of cultivated land quality are subjected to micro-region consistency screening and checkerboard artifact suppression processing to generate a smoothed cultivated land quality distribution map. The smoothed cultivated land quality distribution map is then spatially overlaid with cultivated land parcel boundary data and digital terrain elevation data, including: The multi-label segmentation results of cultivated land quality were subjected to micro-region consistency screening according to the abnormality category. Isolated micro-abnormal regions that were inconsistent with the surrounding spatial distribution were removed to obtain the multi-label segmentation results after screening. In Example 1, for each abnormal category of the multi-label segmentation results of cultivated land quality, the corresponding pixel set is extracted, the connected component labeling process is performed on the pixel set, the pixels with spatial connectivity are divided into multiple abnormal regions, and the number of pixels in each abnormal region is counted.
[0094] For each abnormal region, expand the neighborhood range outward by one pixel to obtain the region neighborhood set. Sum the multi-label classification probabilities of the abnormal categories corresponding to all pixels in the region neighborhood set and divide by the number of pixel positions contained in the region neighborhood set to obtain the region neighborhood consistency coefficient.
[0095] When the number of pixels in a region is less than the area threshold of the small region corresponding to the anomaly category and the consistency coefficient of the region's neighborhood is less than the consistency threshold corresponding to the anomaly category, the multi-label judgment result of the anomaly category corresponding to all pixel positions in the anomaly region is set to zero; otherwise, the original multi-label judgment result remains unchanged, and the small region consistency screening is completed.
[0096] The checkerboard artifact suppression was performed on the screened multi-label segmentation results to obtain a smoothed multi-label classification probability map. A smoothed farmland quality distribution map was generated based on the smoothed multi-label classification probability. In Example 1, a corresponding smoothed multi-label judgment threshold is set for each anomaly category. When the smoothed multi-label classification probability of the corresponding anomaly category is greater than or equal to the smoothed multi-label judgment threshold of the anomaly category, the multi-label judgment result of the corresponding anomaly category is set to 1; otherwise, it is set to 0, thus obtaining the smoothed anomaly category multi-label judgment result.
[0097] For each pixel location, the smoothed multi-label judgment result corresponding to all anomaly categories is counted. When the smoothed multi-label judgment result corresponding to all anomaly categories is zero, the multi-label judgment result corresponding to normal cultivated land is set to one. When the smoothed multi-label judgment result corresponding to at least one anomaly category is one, the multi-label judgment result corresponding to normal cultivated land is set to zero, thus obtaining the smoothed multi-label judgment result corresponding to normal cultivated land.
[0098] The smoothed multi-label judgment results corresponding to all cultivated land quality status categories are expanded into channels according to the category dimension, so that each cultivated land quality status category corresponds to an independent spatial distribution channel. The corresponding multi-label judgment results are filled into each channel according to the pixel spatial position, thus constructing a smooth cultivated land quality distribution map with a multi-channel spatial distribution structure.
[0099] Spatial overlay analysis was performed on the smoothed farmland quality distribution map, farmland plot boundary data, and digital topographic elevation data to calculate the degradation area ratio, abnormal patch density, and topographic disturbance coefficient of each farmland plot. In Example 1, the coverage area of the farmland quality monitoring sample block is divided into multiple farmland plot areas based on the farmland plot boundary data. The number of all pixel locations in each farmland plot area is counted, and the actual ground area corresponding to a single pixel in the standardized remote sensing input dataset is defined as the ground area per unit pixel.
[0100] For each cultivated land plot, a degradation determination is performed on all pixel locations. When the smoothed multi-label determination result corresponding to any anomaly category is one, the pixel location is determined as a degraded pixel location. All degraded pixel locations are counted and divided by the total number of pixel locations in the cultivated land plot area to obtain the degradation area ratio. The degradation area ratio is used to measure the proportion of degraded areas in the cultivated land plot area.
[0101] Within each cultivated land plot area, connected component statistics are performed on the smoothed multi-label judgment results corresponding to all anomaly categories to obtain multiple abnormal patches. The number of abnormal patches is counted, and the actual ground area corresponding to the cultivated land plot area is obtained by multiplying the total number of pixel positions within the cultivated land plot area with the ground area per unit pixel. The abnormal patch density is obtained by multiplying the number of abnormal patches by 10,000 and dividing by the actual ground area. The abnormal patch density is used to measure the distribution density of abnormal patches within a unit area.
[0102] For each cultivated land plot, the average slope response and average curvature response are calculated from the digital topographic elevation data. The results are then weighted and summed according to a preset weighting coefficient to obtain the topographic disturbance coefficient. The topographic disturbance coefficient is used to measure the degree of influence of topographic relief on the quality status of cultivated land.
[0103] The farmland quality grade of each farmland plot is calculated based on the proportion of degraded area, density of abnormal patches, and topographic disturbance coefficient, generating farmland quality diagnosis results at the plot level.
[0104] In this embodiment, the quality grade of cultivated land is determined based on the proportion of degraded area, the density of abnormal patches, and the topographic disturbance coefficient: When the proportion of degraded area is less than or equal to the preset high-quality threshold, the density of abnormal patches is less than or equal to the preset high-quality density threshold, and the topographic disturbance coefficient is less than or equal to the preset high-quality disturbance threshold, the corresponding cultivated land area will be identified as a high-quality cultivated land area. When the proportion of degraded area is within the preset good range or the density of abnormal patches is within the preset good range and the topographic disturbance coefficient is less than or equal to the preset good disturbance threshold, the corresponding cultivated land area will be judged as a good cultivated land area. When the proportion of degraded area is within the preset general range or the density of abnormal patches is within the preset general range and the topographic disturbance coefficient is within the preset general range, the corresponding cultivated land area will be identified as a general cultivated land area. When the proportion of degraded area is greater than the preset poor threshold, or the density of abnormal patches is greater than the preset poor density threshold, or the topographic disturbance coefficient is greater than the preset poor disturbance threshold, the corresponding cultivated land area will be identified as poor cultivated land.
[0105] Example 2: During a routine farmland quality monitoring cycle, the implementers conducted a quality inspection of a continuously cultivated area. This area was divided into 48 farmland plots, with a total monitoring area of 612.8 hectares. The plots exhibited characteristics such as small topographic relief, densely interwoven field ridges, obvious irrigation and drainage traces, localized enhanced surface reflection, and strip-like uneven growth. In the preliminary manual survey, the implementers found that although some plots only showed slight color differences and discontinuous small patches to the naked eye, the field sampling results already showed a superimposed situation of soil compaction, mild salinization, micro-erosion, and nutrient deficiency.
[0106] During this monitoring cycle, the implementers first acquired standardized remote sensing input data for the region. The raw data collected included high-resolution visible light imagery, multispectral imagery, land parcel boundary vector data, and digital terrain elevation data. A total of 4680 raw visible light imagery images, 1560 multispectral imagery images, and digital terrain elevation data with a spatial resolution of 0.25 meters were collected. After the raw data entered the preprocessing stage, 126 images with trailing shadows, 53 images with duplicate overlays, and 38 images with high reflectance saturation were removed, leaving 4463 valid visible light imagery images. Subsequently, brightness normalization, shadow suppression, texture-preserving filtering, and interspectral correction were performed on the retained images. During this process, the average brightness dispersion coefficient between different flight strips in the monitoring area decreased from 0.213 to 0.067, the average local contrast of the blurred boundary area increased from 12.3 to 21.7, and the mean square error at the image stitching point decreased from 1.84 pixels to 0.41 pixels. When the implementer browsed the sample block numbered "BK-12", they found that before preprocessing, a light-colored strip less than 1 meter wide in this sample block was almost indistinguishable from the shadow of the field ridge. The system's initial abnormal response to this area was only 0.34. After completing color constancy correction and texture-aware filtering, the reflection difference of the strip was preserved while the shadow response was suppressed, and the abnormal response of the same area increased to 0.62, providing a stable input for subsequent segmentation.
[0107] The implementers established a training sample library. The training samples were derived from verified plots in historical inspection cycles and the current cycle, forming a total of 15,360 image block samples, each with a uniform size of 512×512 pixels. Among them, there were 5440 normal cultivated land samples, 2740 compacted samples, 1980 salinized samples, 1520 eroded samples, 2860 nutrient-deficient samples, and 820 composite abnormal samples. To ensure the authenticity of the labels, the sample labeling was not based solely on the appearance of the image, but was completed by combining soil penetration resistance, electrical conductivity, organic matter content, slope micro-ditch density, and field verification records. In Example 2, the training sample block numbered "TR-043" appeared in the image as having slightly weak crop growth, hard surface texture, and indistinct color difference between rows. It could easily be judged as normal based on the RGB image alone. However, the measured penetration resistance of the sample block was 2.58 MPa, and the surface bulk density was 1.57 g / cm³, which was significantly higher than the 1.16 MPa and 1.31 g / cm³ of the normal samples in the same batch. Therefore, it was labeled as a compacted sample. Training sample block “TR-087” exhibits an irregular grayish-white reflective surface, a measured electrical conductivity of 2.36 mS / cm, and a pH of 8.48, and is labeled as a salinized sample. Training sample block “TR-114” is located in a slope transition zone, with a microgroove density of 0.82 m / m², a slope response value of 0.127, and a curvature response value of 0.091, and is labeled as an erosion sample. Training sample block “TR-169”, although possessing intact surface texture, has pale leaf color and a decreased near-red edge response, a measured organic matter content of 13.1 g / kg, and a leaf SPAD value of 29.4, and is labeled as a nutrient-deficient sample. Training sample block “TR-203” exhibits both high penetration resistance and low organic matter content, and the system has labeled it with both compaction and nutrient deficiency tags.
[0108] During the model training phase, the implementers divided the aforementioned samples into training, validation, and test sets in a 3:1:1 ratio. The Spectral-Aware UniRepLKNet module was then used to perform spectrally sensitive modeling of the differences in blue, green, red, near-red edges, and near-infrared responses, while preserving fine-grained texture within a large kernel receptive field. On the test sample "TS-031," the average response value of the ordinary convolutional backbone to the mildly salinized area was only 0.46, easily confused with bare, reflective areas in the field. After introducing Spectral-Aware UniRepLKNet, the salinization response in this area increased to 0.79, while the erroneous response of the adjacent bare field ridges decreased from 0.39 to 0.18. On sample "TS-056", the compacted area originally appeared as a small, fragmented block, and the output of the ordinary network was not very continuous, with the largest connected region of a single block being only 214 pixels. After the method of the present invention uses large kernel convolution and inter-spectral attention enhancement, the compacted area is identified as a continuous band-like area, and the largest connected region is expanded to 691 pixels, which is more consistent with the actual compaction distribution on the ground.
[0109] The implementer fed the feature map into the Topo-Aware FeatUp module for topology-aware enhancement. In the plot numbered "DK-07", the original output identified a real gully as three discontinuous short lines, with a maximum break distance of 11 pixels between the segments. After processing with Topo-Aware FeatUp, the gully was restored to a continuous anomalous structure, the average boundary offset decreased from 3.4 pixels to 1.1 pixels, and the boundary F1 value increased from 0.68 to 0.84. During further verification of this plot, the implementer discovered that the gully was accompanied by mild nutrient deficiency on both sides. Ordinary methods, due to detail loss during the upsampling stage, mixed the degradation zones on both sides with the background. However, Topo-Aware FeatUp, when restoring the spatial topology, distinguished the gully erosion line from the mild degradation zones on both sides, allowing both erosion and nutrient deficiency labels to be output stably and simultaneously.
[0110] After feature extraction and topology enhancement, the system generates a pixel-level multi-label classification probability map. When the implementer viewed plot number "GK-12" on the monitoring platform, they found obvious but not completely continuous anomalous responses in the central and eastern parts of the plot. The pixel-level probability results provided by the system showed that the central area had a compaction probability of 0.81 and a nutrient deficiency probability of 0.74, belonging to a composite anomaly; the eastern strip area had a salinization probability of 0.77 and a compaction probability of only 0.29, belonging to a single anomaly. At the same time, an area in the southwest corner affected by the shadow of the field ridge, which was previously given an anomaly judgment of 0.58 by ordinary methods before preprocessing, was reduced to a maximum anomaly probability of 0.21 after spectral correction and topology discrimination in the method of this invention, and was no longer misjudged. The initial multi-label segmentation results of the entire monitoring area output a total of 3842 anomalous connected regions, including 1243 compaction types, 931 salinization types, 618 erosion types, and 1050 nutrient deficiency types. Due to the presence of bright bare ground, watermarks, and work marks between plots, this initial result still contains some isolated, minor false detection areas.
[0111] When entering the micro-area consistency screening stage, the system establishes minimum area constraints and neighborhood consistency constraints for each anomaly category. Taking the parameters in Example 2 as an example, the area threshold for compaction is set to 36 pixels, for salinization to 49 pixels, for erosion to 25 pixels, and for nutrient deficiency to 42 pixels; the corresponding consistency thresholds are 0.43, 0.47, 0.40, and 0.44, respectively. When inspecting the “GK-12” plot, the implementer found 17 scattered small salinization candidate patches in the southeast corner, each patch less than 30 pixels in area, with neighborhood consistency coefficients ranging from 0.19 to 0.33. These patches did not have corresponding salt spots during ground inspection and were false positives caused by reflection and texture noise. Therefore, the system removed all 17 small patches according to the rules. Conversely, an erosion anomaly in the center, with an area of only 28 pixels, although small, had a neighborhood consistency coefficient of 0.58 and was connected to a continuous gully, so it was retained. After the entire monitoring area underwent micro-region consistency screening, a total of 612 isolated micro-abnormal regions were cleared, accounting for 15.93% of the initial number of abnormal regions, and the number of false detection pixels decreased by 21.4%.
[0112] During the stages of chessboard artifact suppression and smoothing of the farmland quality distribution map generation, the system continued to perform category-level smoothing on the screened multi-label results. Implementers noted that the ordinary network exhibited chessboard-like probability fluctuations with "interlaced bright and dark pixels" in certain areas, particularly noticeable in the transition areas between composite anomalies and the normal background. The system set smoothing thresholds for each anomaly category and suppressed and reconstructed the multi-label classification probability map. In the "GK-12" plot, the central composite anomaly area originally contained 43 interlaced voids, resulting in a jagged boundary between compaction and nutrient deficiency. After smoothing, 37 of these voids were corrected, and the boundary irregularity index decreased from 1.38 to 1.11, significantly improving the overall continuity of the area. For the entire monitoring area, the total area of the anomaly region after smoothing was slightly adjusted from 132.46 hectares to 129.83 hectares. While the area change was small, artifact noise was significantly reduced, and the spatial coherence of the anomaly patches was closer to the results of manual verification.
[0113] In the plot-level spatial overlay analysis phase, the system overlays the smoothed farmland quality distribution map with farmland plot boundary data and digital terrain elevation data block by block. The implementers focused on three plots: “GK-12,” “DK-07,” and “BM-04.” Taking “GK-12” as an example, this plot has a total area of 8.64 hectares and a total of 78,400 pixels, of which 21,964 pixels were classified as degraded, representing a degraded area ratio of 0.280. A total of 37 anomalous patches were identified within this plot, resulting in an anomalous patch density of 42.8 patches per 10,000 square meters. Digital terrain elevation data shows that the average slope response of this plot is 0.118, and the average curvature response is 0.162. After weighting by 0.6 and 0.4, the terrain disturbance coefficient is 0.136. Based on these three indicators, the system classifies “GK-12” as a poor farmland plot. Looking at plot "DK-07," with a total area of 6.12 hectares, a degraded area ratio of 0.173, an abnormal patch density of 28.4 per 10,000 square meters, and a topographic disturbance coefficient of 0.121, it was classified as general arable land. As for plot "BM-04," although there were slight color differences in some areas, the system ultimately identified a degraded area ratio of only 0.051, an abnormal patch density of 9.6 per 10,000 square meters, and a topographic disturbance coefficient of 0.047, thus classifying it as high-quality arable land. When the implementers compared these results with the on-site verification records, they found that the classification of the three plots was completely consistent with the soil samples and inspection records.
[0114] To verify the effectiveness of the method of this invention, the implementers conducted a comparative experiment using a traditional method on the same batch of data. The traditional method uses a process of ordinary U-Net backbone + bilinear upsampling + global threshold segmentation + manual rule repair. The comparison results show that on the test set, the overall mIoU of the traditional method is 0.821, the F1 value is 0.874, the boundary F1 value is 0.703, and the accuracy rate of parcel-level classification is 78.4%; the overall mIoU of the method of this invention is 0.903, the F1 value is 0.932, the boundary F1 value is 0.851, and the accuracy rate of parcel-level classification is 92.6%. Looking at the anomaly types, the IoU of the traditional method for the four anomalies of compaction, salinization, erosion, and nutrient deficiency are 0.79, 0.76, 0.69, and 0.73, respectively, while the method of this invention achieves IoUs of 0.88, 0.91, 0.86, and 0.87, respectively. When dealing with slender, weakly bounded targets like erosion, traditional methods are most prone to fracture identification and missed detection. The method of this invention maintains spatial topological continuity through Topo-Aware FeatUp, thus achieving the most significant improvement.
[0115] In terms of actual work efficiency, the traditional method takes an average of 1.84 seconds to process a single 2048×2048 pixel sample block, and due to the large number of false positives, subsequent manual revisions require an average of 0.63 seconds. The method of this invention takes an average of 1.12 seconds to process a sample block of the same size, and subsequent manual revisions require an average of only 0.17 seconds. In other words, across the entire 612.8-hectare monitoring area, the traditional method takes approximately 98 minutes to complete a single process, while the method of this invention takes only approximately 61 minutes, reducing the workload of manual review by approximately 64.5%. Regarding the quality of the map patches, the traditional method produces an average of 31.6 isolated false positive micro-patterns per plot, while the method of this invention reduces this to 8.7. Table 1 below provides the overall comparison data between the method of this invention and the traditional method in this embodiment: Table 1. Overall comparison data between the method of the present invention and the conventional method in this embodiment. Overall mIoU 0.821 0.903 Overall F1 score 0.874 0.932 Boundary F1 value 0.703 0.851 IoU of slab 0.79 0.88 Salting IoU 0.76 0.91 Erosion IoU 0.69 0.86 Nutrient Deficiency IoU 0.73 0.87 Accuracy of land parcel grading 78.4% 92.6% Average number of isolated false positive microplasms / plot 31.6 8.7 Average processing time per block 1.84 seconds 1.12 seconds Single block manual revision time 0.63 seconds 0.17 seconds Total processing time for the entire region at one time Approximately 98 minutes Approximately 61 minutes To further illustrate the correspondence between the training samples and the identification results, the implementers extracted six representative samples from the test set for individual verification. Sample "TS-011" corresponds to normal cultivated land, with a measured organic matter content of 28.4 g / kg and a penetration resistance of 1.12 MPa. The system gave a normal judgment probability of 0.94, and no abnormal label was generated; the traditional method, however, falsely reported a small area of nutrient deficiency at the edge of the field ridge. Sample "TS-031" corresponds to mild salinization, with a measured electrical conductivity of 2.08 mS / cm. The method of this invention gave a salinization probability of 0.79, while the traditional method only gave 0.52, failing to reach its fixed threshold and thus missing the detection. Sample "TS-056" corresponds to compaction, with a measured penetration resistance of 2.63 MPa. The method of this invention identified it as a continuous compaction zone, while the traditional method only identified three broken patches. Sample "TS-074" corresponds to erosion, with surface microgrooves only 0.7 meters wide. The method of this invention successfully maintained the continuity of the grooves, while the traditional method broke off a section in the middle. Sample "TS-091" corresponds to nutrient deficiency, with measured organic matter of 12.7 g / kg. The method of this invention outputs a nutrient deficiency probability of 0.83, while the traditional method outputs 0.49 due to shading. Sample "TS-118" corresponds to a coexistence of compaction and nutrient deficiency. The method of this invention provides dual labels simultaneously, while the traditional method only identifies compaction, missing the nutrient deficiency.
[0116] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for monitoring farmland quality based on image recognition, characterized in that, include: Remote sensing datasets covering the target farmland monitoring area are collected and preprocessed to obtain a standardized remote sensing input dataset. This dataset is then jointly encoded to obtain a multimodal fusion data tensor, which is then divided into blocks to generate a farmland quality monitoring sample block sequence. The sequence of farmland quality monitoring sample blocks is input into the Spectral-AwareUniRepLKNet backbone feature extraction network, which outputs multi-scale deep semantic feature representations. We apply structural reparameterization fusion to multi-scale deep semantic feature representations, folding the multi-branch convolutional structure in the training stage into a single-path large-kernel convolutional structure in the inference stage, and outputting reparameterized semantic feature representations. The reparameterized semantic feature representation is subjected to hierarchical aggregation, cross-scale alignment and sliding gated residual fusion. Digital terrain elevation data is injected into the reparameterized semantic feature representation to obtain a fused feature map of farmland quality. The fused feature map of farmland quality is input into the Topo-AwareFeatUp high-resolution feature reconstruction module to perform dimensionality restoration on the fused feature map of farmland quality and output the final high-resolution semantic feature map. Farmland quality status classification is performed on high-resolution semantic feature maps at the pixel level, block level, and farmland plot level, and the multi-label segmentation results of farmland quality are output. The multi-label segmentation results of cultivated land quality are subjected to micro-region consistency screening and checkerboard artifact suppression to generate a smooth cultivated land quality distribution map. The smooth cultivated land quality distribution map is spatially overlaid with cultivated land plot boundary data and digital terrain elevation data to calculate the cultivated land quality grade, degradation area ratio and abnormal patch density of each cultivated land plot, and generate cultivated land plot-level cultivated land quality diagnosis results.
2. The method for monitoring farmland quality based on image recognition according to claim 1, characterized in that, The process of collecting remote sensing datasets covering the target cultivated land monitoring area and performing preprocessing operations includes: Multispectral remote sensing images, high-resolution orthophotos, digital topographic elevation data, historical temporal images, and farmland plot boundary data covering the target farmland monitoring area were collected. Spatial registration, radiometric correction, geometric correction, cloud and fog removal, noise suppression, resolution unification, and coordinate system unification were performed to obtain a standardized remote sensing input dataset. The spectral information, texture information, topographic gradient information, water system distribution information, and farmland plot morphology information in the standardized remote sensing input dataset were jointly encoded to obtain a multimodal fusion data tensor covering the target farmland monitoring area. The multimodal fusion data tensor was then divided into blocks according to a preset sliding window strategy to generate a farmland quality monitoring sample block sequence.
3. The method for monitoring farmland quality based on image recognition according to claim 2, characterized in that, The step of inputting the farmland quality monitoring sample block sequence into the Spectral-AwareUniRepLKNet backbone feature extraction network includes: For each input channel in each farmland quality monitoring sample block, the farmland quality sensitivity coefficient of the input channel is calculated. All input channels are rearranged and grouped according to the farmland quality sensitivity coefficient to form a sub-spectral group input representation. Calculate the farmland texture complexity coefficient for each spectral group in the sub-spectral group input representation. Based on the comparison between the farmland texture complexity coefficient and the preset threshold, select the actual convolution kernel side length and the actual convolution dilation rate from the corresponding candidate set. The effective receptive field size is obtained by multiplying the convolution dilation rate by the result of subtracting 1 from the convolution kernel side length and adding 1. Based on each farmland quality monitoring sample block, the main orientation angle of farmland structure is constructed. Based on the main orientation angle of farmland structure, the convolution sampling coordinates are rotated and transformed. Under the combined effect of the rotational sampling coordinates and the orientation unfolding modulation coefficient, orientation sensing large receptive field channel-by-channel convolutional encoding is independently performed on each channel in each spectral group to obtain the convolutional response within the orientation sensing group. A high-frequency anomaly residual map is constructed on the spectral group for the farmland quality monitoring sample block. Based on the high-frequency anomaly residual map, a farmland quality anomaly fidelity gating map is generated and weighted and combined with the reciprocal of the effective receptive field size to obtain the fidelity fusion coefficient. Based on the farmland quality anomaly fidelity gating map and the fidelity fusion coefficient, the convolutional response within the direction perception group and the original input features are weighted and fused to obtain the high-frequency preservation group convolutional response. The high-frequency intra-group convolutional responses of each spectral group in the same coding layer are concatenated to obtain an inter-group fusion feature tensor. The features of each spectral group are weighted by the contribution weight of the spectral group and then concatenated to obtain a weighted inter-group fusion feature tensor. Channel mapping and stage compression are performed on the weighted inter-group fusion feature tensor to obtain stage coding features. When the current coding layer is not the last coding layer, scale-progressive mapping is performed on the stage coding features to obtain scale-progressive features. When the current coding layer is the last coding layer, the stage coding features are directly output as scale-progressive features. The scale-progressive features output by each coding layer are defined as multi-scale deep semantic features with progressively decreasing spatial resolution. The outputs of each coding layer are then aggregated in hierarchical order to obtain the multi-scale deep semantic feature representation corresponding to each farmland quality monitoring sample block.
4. The method for monitoring farmland quality based on image recognition according to claim 3, characterized in that, The application of structure reparameterization fusion to multi-scale deep semantic feature representation folds the multi-branch convolutional structure in the training phase into a single-path large-kernel convolutional structure in the inference phase, including: For each scale-progressive feature in the multi-scale deep semantic feature representation, a multi-branch channel-wise convolutional structure is constructed for the training phase. During the training phase, each branch of the multi-branch channel-wise convolutional structure is subjected to folding processing of convolution parameters and batch normalization parameters. The scaling and translation effects on the feature distribution in the batch normalization mapping are converted into equivalent corrections to the convolution kernel weights and convolution bias terms, resulting in independent equivalent convolution kernels and equivalent convolution bias terms for each branch. Perform convolution kernel size alignment processing to center-align the equivalent convolution kernels of the large convolution kernel main branch, the small convolution kernel compensation branch, and the identity mapping preservation branch under the actual convolution kernel side length dynamically configured by the farmland texture complexity coefficient. Then, perform element-wise tensor accumulation on the aligned equivalent convolution kernels and equivalent convolution bias terms to obtain the global fusion convolution kernel and global fusion convolution bias terms corresponding to the single-path large kernel convolution structure in the inference stage. Using global fusion convolution kernels and global fusion convolution bias terms as kernel weight basis, single-path large kernel convolution calculation is performed on the corresponding scale-progressive features under the combined effect of the sampling coordinates after orientation rotation and the orientation unfolding modulation coefficients to obtain reparameterized semantic feature representation; The reparameterized semantic feature representations corresponding to each coding layer are aggregated in hierarchical order to obtain the reparameterized semantic feature representation corresponding to each farmland quality monitoring sample block, and then aggregated to obtain the set of reparameterized semantic feature representations corresponding to the sequence of farmland quality monitoring sample blocks.
5. The method for monitoring farmland quality based on image recognition according to claim 4, characterized in that, The process of performing hierarchical aggregation, cross-scale alignment, and sliding-gated residual fusion on the reparameterized semantic feature representation, and injecting digital terrain elevation data into the reparameterized semantic feature representation, includes: Hierarchical aggregation basis construction is performed on the reparameterized semantic feature representations at all levels, and spatial clipping and scale mapping are performed on the digital terrain elevation data corresponding to the nth cultivated land quality monitoring sample block to obtain the corresponding terrain injection features; Perform top-down hierarchical aggregation processing on the aggregated base features at each level to obtain the hierarchical aggregated features of the current coding layer; Perform cross-scale alignment processing on the aggregated features at each level to obtain cross-scale aligned features; Sliding gated residual fusion processing is performed on cross-scale aligned features and terrain-injected features to inject digital terrain elevation data into reparameterized semantic feature representations, thereby obtaining fused features of cultivated land quality at all levels. The fusion features of cultivated land quality at all levels are aggregated hierarchically into shallow, medium and deep layers to obtain a fusion feature map of cultivated land quality that includes shallow edge response, medium texture response and deep semantic response.
6. The method for monitoring farmland quality based on image recognition according to claim 5, characterized in that, The step of inputting the fused feature map of cultivated land quality into the Topo-AwareFeatUp high-resolution feature reconstruction module to perform dimensionality-upgrading recovery of the fused feature map of cultivated land quality includes: Construct high-resolution query coordinates consistent with the spatial resolution of the standardized remote sensing input dataset, and map the fused feature map of cultivated land quality into local implicit reconstruction conditional features in the continuous coordinate query space; The local implicit reconstruction conditional features are input into the task-independent high-resolution implicit reconstruction operator, and implicit semantic reconstruction is performed on each high-resolution query coordinate to obtain the initial high-resolution semantic feature map. A multi-view mapping result is constructed from the fused feature map of cultivated land quality, and a synchronous spatial transformation is performed on the high-resolution query coordinates. Multi-view consistency constraints are constructed based on the high-resolution implicit reconstruction results corresponding to each multi-view mapping result. A curvature response is constructed from the initial high-resolution semantic feature map and a curvature continuity constraint is established. The multi-view consistency constraint and the curvature continuity constraint are weighted and summed according to a preset weight to obtain the high-resolution implicit reconstruction constraint. Based on the high-resolution implicit reconstruction constraint, the task-independent high-resolution implicit reconstruction operator is updated online with adaptive parameters through the backpropagation algorithm. Using a task-independent high-resolution implicit reconstruction operator updated online with the current samples, full-resolution upscaling recovery is performed on the fused feature map of farmland quality, outputting a final high-resolution semantic feature map with the same spatial resolution as the standardized remote sensing input dataset.
7. The method for monitoring farmland quality based on image recognition according to claim 6, characterized in that, The process of classifying farmland quality status on high-resolution semantic feature maps at the pixel, block, and farmland plot levels includes: Shared feature mapping is performed on the final high-resolution semantic feature map to obtain shared predictive basis features for multi-granularity farmland quality status classification; Pixel-level farmland quality status classification is performed on the shared prediction basis features. The independent classification probability of each abnormal category is calculated for each pixel location, and the classification probability of normal farmland is calculated based on the classification probability of all abnormal categories, resulting in a pixel-level multi-label classification probability map. Perform block-scale farmland quality status classification on the shared prediction basis features, and map the block-scale classification results back to the pixel space to obtain a block-scale multi-label classification probability map; Based on the boundary data of cultivated land plots, the cultivated land quality status classification at the granular level is performed on the shared prediction basis features. The granular classification results of cultivated land plots are mapped back to the pixel space to obtain a multi-label classification probability map of cultivated land plot granularity. Perform hybrid granularity multi-label fusion classification on pixel-level multi-label classification probability maps, block-level multi-label classification probability maps, and cultivated land plot-level multi-label classification probability maps to obtain cultivated land quality multi-label segmentation results.
8. The method for monitoring farmland quality based on image recognition according to claim 7, characterized in that, The process involves performing micro-region consistency screening and checkerboard artifact suppression on the multi-label segmentation results of cultivated land quality to generate a smoothed cultivated land quality distribution map. This smoothed map is then spatially overlaid with cultivated land parcel boundary data and digital terrain elevation data, including: The multi-label segmentation results of cultivated land quality were subjected to micro-region consistency screening according to the abnormality category. Isolated micro-abnormal regions that were inconsistent with the surrounding spatial distribution were removed to obtain the screened multi-label segmentation results. The checkerboard artifact suppression was performed on the screened multi-label segmentation results to obtain a smoothed multi-label classification probability map. A smoothed farmland quality distribution map was generated based on the smoothed multi-label classification probability. Spatial overlay analysis was performed on the smoothed farmland quality distribution map, farmland plot boundary data, and digital topographic elevation data to calculate the degradation area ratio, abnormal patch density, and topographic disturbance coefficient of each farmland plot. The farmland quality grade of each farmland plot is calculated based on the proportion of degraded area, density of abnormal patches, and topographic disturbance coefficient, generating farmland quality diagnosis results at the plot level.
9. A method for monitoring farmland quality based on image recognition according to claim 8, characterized in that, The quality grade of cultivated land is determined based on the proportion of degraded area, the density of abnormal patches, and the topographic disturbance coefficient. When the proportion of degraded area is less than or equal to the preset high-quality threshold, the density of abnormal patches is less than or equal to the preset high-quality density threshold, and the topographic disturbance coefficient is less than or equal to the preset high-quality disturbance threshold, the corresponding cultivated land area will be identified as a high-quality cultivated land area. When the proportion of degraded area is within the preset good range or the density of abnormal patches is within the preset good range and the topographic disturbance coefficient is less than or equal to the preset good disturbance threshold, the corresponding cultivated land area will be judged as a good cultivated land area. When the proportion of degraded area is within the preset general range or the density of abnormal patches is within the preset general range and the topographic disturbance coefficient is within the preset general range, the corresponding cultivated land area will be identified as a general cultivated land area. When the proportion of degraded area is greater than the preset poor threshold, or the density of abnormal patches is greater than the preset poor density threshold, or the topographic disturbance coefficient is greater than the preset poor disturbance threshold, the corresponding cultivated land area will be identified as poor cultivated land.