Crop all-weather monitoring and disaster assessment method based on cross-modal self-distillation

CN122090284BActive Publication Date: 2026-09-29SANYA SCI & EDUCATION INNOVATION PARK WUHAN UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610541526.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-23
Publication Date
2026-09-29
Estimated Expiration
2046-04-23

AI Technical Summary

Technical Problem

[0006]为解决现有遥感影像变化检测方法中存在的上下文建模能力局限、跨尺度特征融合机制不足以及边界分割精度瓶颈等技术问题,本发明提出基于跨模态自蒸馏的农作物全天候监测与灾害评估方法

Benefits of technology

(1)实现全天候、跨源的农情连续精准监测:本设备依托跨模态自蒸馏核心机制,驱动模型深度挖掘并学习光学与 SAR 影像中共享的作物语义原型,从特征层面彻底消除不同传感器成像机理差异带来的模态鸿沟,有效规避传感器差异被误判为农情变化的问题。即便光学影像因云雾、阴雨等天气遮挡缺失时,亦可通过 SAR 影像无缝接续完成农作物生长状态监测,突破天气条件对农业遥感监测的制约,实现全时段、无间断的农情监测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122090284B_ABST
    Figure CN122090284B_ABST
Patent Text Reader

Abstract

The application discloses a crop all-weather monitoring and disaster assessment method based on cross-modal self-distillation, and relates to the technical field of intelligent agriculture and agricultural remote sensing monitoring. The method collects remote sensing images of a farmland to be monitored to generate multiple perspective sample graphs, constructs a cross-modal self-distillation crop condition monitoring network, the network comprises a teacher network and a student network, and both are composed of a backbone encoder and a projection head; then, a target probability distribution of the teacher network, a prediction probability of the student network and a total loss function are calculated, and the teacher network parameters and the student network parameters are updated; finally, a Euclidean distance is calculated by using the trained backbone encoder to generate a change intensity graph, and a crop growth distribution graph or a disaster damage mask graph is output through threshold segmentation. The application effectively solves the problems of optical image loss caused by cloud and fog interference and high annotation cost of farmland plots in agricultural monitoring, and realizes all-weather, automatic crop growth monitoring and accurate disaster assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of smart agriculture and agricultural remote sensing monitoring technology, specifically to a method for all-weather monitoring and disaster assessment of crops based on cross-modal self-distillation. Background Technology

[0002] With the rapid development of Earth observation technology, acquiring multi-source heterogeneous remote sensing images of the same area (such as optical images and synthetic aperture radar SAR images) has become a core data support for crop growth monitoring and disaster assessment. In precision agriculture and food security monitoring, timely and accurate acquisition of crop growth status and disaster damage is crucial. Currently, using multi-temporal remote sensing images for agricultural monitoring is the mainstream method. However, existing agricultural remote sensing monitoring faces the following severe challenges, seriously restricting the realization of all-weather, automated monitoring: 1) Difficulty in all-weather monitoring: Key growth periods for crops, such as the jointing and heading stages, are often accompanied by cloudy and rainy weather, resulting in missing or obscured optical remote sensing images. Although synthetic aperture radar (SAR) has all-weather imaging capabilities, its imaging mechanism is completely different from that of optical images (e.g., SAR reflects crop structure and dielectric properties, while optical images reflect chlorophyll and spectral properties). This huge modal gap can easily lead traditional methods to misinterpret the imaging differences between different sensors as changes in crop growth or disasters, making it impossible to achieve effective fusion of cross-source data.

[0003] 2) High cost of sample annotation: Due to the fragmented nature of agricultural plots, the variety of crops, and the rapid changes in growth characteristics, it is extremely costly and time-consuming to create a large-scale, high-precision multimodal crop change detection dataset. Moreover, it is difficult to cover the entire scenario of different crops, different growth stages, and different types of disasters, resulting in poor model generalization ability and difficulty in promoting it in different agricultural production areas.

[0004] 3) Weak anti-interference ability: The farmland environment is complex. Non-essential differences caused by changes in the angle of sunlight and soil moisture are often misidentified as changes in crops, resulting in a high false alarm rate in monitoring. The accuracy of disaster assessment is difficult to meet the actual application needs of agricultural production, agricultural insurance loss assessment and other applications.

[0005] Therefore, there is an urgent need for an unsupervised change detection method that does not require manual annotation and can effectively integrate the advantages of optics and SAR to achieve all-weather agricultural monitoring, overcome the dual limitations of weather and sensors, and realize all-weather, precise monitoring of crop growth and disasters. Summary of the Invention

[0006] To address the limitations of existing remote sensing image change detection methods, such as insufficient context modeling capabilities, inadequate cross-scale feature fusion mechanisms, and bottlenecks in boundary segmentation accuracy, this invention proposes a method for all-weather monitoring and disaster assessment of crops based on cross-modal self-distillation.

[0007] This invention employs a deep neural network based on a student-teacher architecture, constructing a total loss function that includes cross-modal alignment and intra-modal alignment. To address the "modal gap" problem, cross-modal alignment forces the model to learn shared semantic content across different modalities, thereby filtering out spurious changes caused by sensor differences. To address the "insufficient feature robustness" problem, intra-modal alignment utilizes multiple views to enhance consistency and improve the model's robustness to illumination and noise. Simultaneously, this invention introduces exponential moving average (EMA) to update teacher network parameters and designs a dynamic centering mechanism to smooth teacher outputs, effectively preventing the model collapse problem common in unsupervised learning and ensuring the stability and diversity of teacher network output distribution. Furthermore, this invention abandons complex decoders or classifiers, directly utilizing the trained backbone network to extract features and calculate Euclidean distance as the change intensity, avoiding the overfitting risk associated with training a specific classifier, and achieving high-precision, interference-resistant unsupervised change detection.

[0008] The technical solution adopted in this invention is as follows: S1) Acquire first-phase and second-phase remote sensing images of the same farmland area at different times, and apply a data augmentation strategy based on a combination of geometric transformation and photometric transformation to each remote sensing image to generate multiple viewpoint sample images; S2) Construct a cross-modal self-distillation agricultural monitoring network based on a twin architecture, including parallel branches with the same structure but different parameter update methods, namely a student network and a teacher network; input sample images from multiple perspectives into the cross-modal self-distillation agricultural monitoring network for feature extraction and mapping to a low-dimensional compact space to obtain class prediction logical values; both the student network and the teacher network consist of a backbone encoder and a projection head; define an update mechanism to update the parameters of the student network, and dynamically update the parameters of the teacher network through the exponential moving average (EMA) of the student network parameters; S3) Based on the category prediction logic value, calculate the target probability distribution of the teacher network and the predicted probability of the student network respectively; S4) Train the cross-modal self-distillation agricultural monitoring network, calculate the cross-modal distillation loss and the same-modal distillation loss, and obtain the total loss function by weighted summation; minimize the total loss through backpropagation to update the student network parameters, and update the teacher network parameters and the centering vector until a steady state is reached; S5) Using the trained student network backbone encoder, extract pixel-level feature maps from the first and second time-phase remote sensing images of the farmland area to be detected; calculate the Euclidean distance of the pixel-level feature maps and perform normalization to generate a change intensity map; use the maximum inter-class variance method to perform threshold segmentation on the change intensity map to generate a crop growth distribution map or a disaster damage mask map.

[0009] Preferably, in step S1), the specific implementation process of generating multiple viewpoint sample images is as follows: Two remote sensing images of the same farmland area were collected at different times and are respectively recorded as the first time-phase remote sensing image. Second-phase remote sensing images This study utilizes a data augmentation strategy combining geometric and photometric transformations to augment remote sensing images, constructing multi-view sample maps. Two data augmentation strategies are defined as follows: and For each pair of input images The data augmentation strategy is applied twice to generate sample images from multiple viewpoints. The calculation formula is as follows:

[0010] In the formula, and This represents two different augmentation processes using two random seeds, i.e., two data augmentation strategies; Represents first-phase remote sensing imagery application First-person view sample image generated after enhancement operation. Represents first-phase remote sensing imagery application The second-view sample image generated after the enhancement operation. Indicates second-phase remote sensing imagery application First-person view sample image generated after enhancement operation. Indicates second-phase remote sensing imagery application The second-view sample image generated after the enhancement operation. This indicates the first enhanced view. Indicates the second enhanced view; Finally, the sample images from multiple perspectives are normalized and converted into tensor format, and then input into the subsequent cross-modal self-distillation agricultural monitoring network.

[0011] Preferably, in step S2), multiple viewpoint sample images are input into a cross-modal self-distillation agricultural monitoring network for feature extraction and mapping to a low-dimensional, compact space to obtain category prediction logical values, specifically including: S2.1) Backbone encoder: The backbone encoder adopts a fully convolutional structure to extract high-dimensional spatial features of the image; First, sample images from multiple viewpoints are input into the first convolutional block, and then mapped to the feature space through a convolutional layer. Next, a batch normalization layer is used to accelerate convergence and prevent overfitting. Finally, a nonlinearity is introduced through the ReLU activation function, the corresponding calculation formula of which is:

[0012] In the formula, Represents the ReLU activation function; Indicates batch normalization; This represents a convolution operation with a kernel size of 3, a stride of 1, and padding of 1. Indicates the number of input channels; Indicates the preset feature dimension; This represents the first layer of features, whose dimensions are consistent with the spatial dimensions of the view sample image; Next, The second convolutional block is input to further extract deeper abstract features, and the corresponding formula is:

[0013] In the formula, This represents the output features of the second layer; Then Input the third convolutional block to expand the number of channels to the target feature dimension, as shown in the formula:

[0014] In the formula, This represents the output features of the third layer; Finally, The fourth convolutional block is input and subjected to a nonlinear transformation to enhance the discriminative power of the features. The corresponding formula is:

[0015] In the formula, This represents the high-dimensional spatial features of the final output of the backbone encoder. The dimension represents the feature of a high-dimensional space, where, Indicates batch size, Indicates the number of feature channels. Indicates altitude, Indicates width; S2.2) Projection Head: The projection head is connected after the backbone encoder and adopts a multilayer perceptron structure to receive high-dimensional spatial features. This is then mapped to a low-dimensional, compact space to generate class prediction logistic values ​​for calculating probability distributions; the specific implementation steps are as follows: First, global average pooling is performed on the high-dimensional spatial features output by the backbone encoder to aggregate global context information. The corresponding formula is:

[0016] In the formula, This represents the feature vector after pooling. Let represent the dimension of the feature vector after pooling, where Indicates batch size, Indicates the number of feature channels; Indicates global average pooling; Subsequently, the pooled feature vectors are input into the first fully connected layer for linear expansion to enhance the dimensionality of the features, and a smooth non-linearity is introduced through the GELU activation function, as shown in the following formula:

[0017] In the formula, This represents the GELU activation function. Indicates a fully connected layer. Represents hidden layer features; Then, the hidden layer features The input to the second fully connected layer compresses the dimension to the bottleneck low dimension, as shown in the formula:

[0018] In the formula, Indicates bottleneck layer characteristics; Next, the bottleneck layer features Normalize to the L2 norm and project it onto the unit hypersphere; the corresponding formula is:

[0019] In the formula, This represents a small constant used to prevent the denominator from being zero. express L2 norm, This represents the normalized eigenvectors; Finally, a linear layer with weight normalization maps the normalized feature vector to a lower dimension, yielding the class prediction logistic value, with the corresponding formula:

[0020] In the formula, This represents the weight matrix of the layer; This represents the category prediction logic value output by the projection head; This indicates that the weight matrix is ​​normalized using the L2 norm. The specific implementation process of the S2.3 update mechanism includes: Teacher network parameters Initialize to student network parameters With consistent parameters, during the update process, the student network completes the update by participating in gradient backpropagation, while the teacher network does not participate in gradient backpropagation and is only dynamically updated through the exponential moving average (EMA) of the student network parameters. The update formula for the teacher network parameters is as follows:

[0021] In the formula, Indicates the momentum coefficient. Indicates the teacher's network parameters. This represents the student's network parameters.

[0022] Preferably, in step S3), the calculation processes for the target probability distribution of the teacher network and the predicted probability of the student network are as follows: S3.1) The calculation process of the target probability distribution of the teacher network includes: First, the center vector is updated based on the category prediction logistic values ​​output by the teacher network. The formula for the center vector is:

[0023] In the formula, Represents the center vector. Indicates the central momentum coefficient. Indicates batch size, Indicates the current batch number The class prediction logistic value output by the teacher's network projection head for each sample. ; Subsequently, the class prediction logistic values ​​output by the teacher network are processed using the center vector and temperature coefficient to generate a sharpened target probability distribution, with the corresponding formula as follows:

[0024] In the formula, Indicates the temperature coefficient. express Activation function Indicates input sample The corresponding class prediction logic value output by the teacher's network projector. Represents the target probability distribution. Indicates the number of input networks in the current batch. One sample; S3.2) The calculation process for the predicted probability of the student network includes: Based on the category prediction logistic values ​​output by the student network, the prediction probability is calculated using the following formula:

[0025] In the formula, This represents the temperature coefficient of the student network. This represents the category prediction logical value output by the student's network projector; Indicates the predicted probability; Preferably, step S4) specifically includes: S4.1) The calculation process of the total loss function includes: First, calculate the cross-modal distillation loss. This allows images of different modalities to be aligned in the semantic space, and the corresponding formula is:

[0026] In the formula, This represents the standard cross-entropy loss; In the student network, the first The first modal image Sample images from various perspectives; In the teacher network, the first Viewpoint sample images of modal images; Indicates the first The first modal image Predicted probability of a sample image from a single viewpoint; Indicates the first Target probability distribution in the viewpoint sample map of modal imagery; Next, the teacher network is established in the first augmented view. Above, the student network is built in a second enhanced view. To maintain consistency between different views under the same mode, calculate the distillation loss under the same mode. The corresponding formula is:

[0027] In the formula, Indicates the first Second-view sample images generated from each modal image after data augmentation This indicates that the student network input was obtained. The corresponding predicted probability; Indicates the first First-view sample images generated from modal images after data augmentation This indicates input into the teacher network to obtain... The corresponding target probability distribution; Finally, the weighted summation yields the total loss function. The corresponding formula is:

[0028] In the formula, and The student network is updated via backpropagation by minimizing preset weight coefficients. S4.2) The specific process of parameter update and training iteration includes: In each training batch, the gradient cache of the student network is first zeroed out, followed by forward propagation: multiple viewpoint sample maps are input into the teacher network and the student network respectively. The forward propagation of the teacher network is performed without constructing a backward computation graph, and is only used to obtain the target probability distribution. and the category prediction logical value for the current batch; Subsequently, based on the total loss function Backpropagation is performed on the student network backbone encoder and projector to calculate the gradients of each parameter; the gradients of the student network parameters are then L2 norm clipped to restrict the global norm of the gradients to a threshold value. Within this range, the corresponding formula is:

[0029] After gradient pruning, the RMSprop optimizer is used to optimize the student network parameters. Perform gradient descent updates to drive the student network to converge gradually by minimizing the total loss; After the student network parameters are updated, the teacher network parameters are simultaneously updated using an exponential moving average method, without participating in gradient calculation. The corresponding formula is:

[0030] In the formula, Indicates the momentum coefficient. This indicates the updated teacher network parameters. For the updated student network parameters; At the same time, the global centralization vector is updated synchronously using the category prediction logical values ​​of the current batch of teacher networks, so that the output of the next batch of teachers can be decentralized using the latest centralization vector.

[0031] Preferably, in step S5), generating a crop growth distribution map or a disaster damage mask map specifically includes the following steps: First, the first temporal remote sensing image to be detected... Second-phase remote sensing images The dual-temporal image input encoder extracts pixel-level feature maps, and the corresponding formula is:

[0032] In the formula, This represents the pixel-level feature map output by the encoder. Indicates encoder; Next, the Euclidean distance between the two pixel-level feature maps at each pixel location is calculated using the following formula:

[0033] In the formula, Represents pixels Euclidean distance at; and Represents pixel-level feature maps and In pixels The feature vector of the location; Represents the L2 norm; Subsequently, the Euclidean distance is normalized to its minimum and maximum values, and then mapped to a range. The formula for generating a change intensity map within the interval is:

[0034] In the formula, and Let represent the global minimum and maximum values ​​of the Euclidean distance, respectively. A graph showing the intensity of change after normalization; Finally, the Otsu's method is used to iterate through all possible candidate thresholds. Calculate the inter-class variance between the foreground (changing region) and the background (unchanging region) at each threshold, and take the threshold with the largest inter-class variance as the global optimal threshold. The intensity change map is then binarized to generate a binary crop growth distribution map or a disaster damage mask map. The corresponding formula is:

[0035] In the formula, Represents the variance between classes. Indicates the percentage of background pixels. This indicates the percentage of pixels in the foreground. This represents the average value of the background pixels. Indicates the average value of foreground pixels. This represents the candidate threshold for traversal. This represents a map showing the distribution of crop growth or a mask map showing damage caused by disasters.

[0036] A method for all-weather monitoring and disaster assessment of crops based on cross-modal self-distillation includes a computer-readable storage medium storing a computer program. The computer program, when executed by a processor, implements the aforementioned unsupervised remote sensing image change detection method based on cross-modal self-distillation learning.

[0037] A crop all-weather monitoring and disaster assessment detection device based on cross-modal self-distillation learning includes a processor and a memory, wherein the memory stores a computer program. The computer program, when executed by the processor, implements the aforementioned unsupervised remote sensing image change detection method based on cross-modal self-distillation learning.

[0038] Compared with the prior art, the present invention has the following advantages and beneficial effects: (1) Achieving continuous and accurate agricultural monitoring around the clock and across multiple sources: This equipment relies on a cross-modal self-distillation core mechanism to drive the model to deeply mine and learn the shared crop semantic prototypes in optical and SAR images. It completely eliminates the modal gap caused by the differences in imaging mechanisms of different sensors at the feature level, effectively avoiding the problem of sensor differences being misjudged as changes in agricultural conditions. Even when optical images are missing due to weather conditions such as clouds, fog, or rain, crop growth status monitoring can be seamlessly completed through SAR images, breaking through the constraints of weather conditions on agricultural remote sensing monitoring and achieving all-time, uninterrupted agricultural monitoring.

[0039] (2) Significantly reduce the deployment and promotion costs of smart agriculture monitoring systems: This equipment adopts an unsupervised training paradigm based on exponential moving average (EMA). The entire model training process does not require manual labeling of farmland plots and disaster-damaged areas. It can automatically learn the core characteristics of crop growth through massive amounts of unlabeled multi-source remote sensing images, fundamentally solving the problems of high sample labeling costs and difficulty in obtaining samples caused by fragmented agricultural plots and a wide variety of crop types. It can quickly adapt to the monitoring needs of different agricultural production areas and different crop types, and has strong practical application and promotion potential.

[0040] (3) Significantly improves the accuracy and practicality of crop growth analysis and disaster assessment: Based on a feature space metric inference framework, this device can generate a continuous intensity map that accurately reflects the degree of crop growth change and the level of disaster damage. Then, through an adaptive threshold segmentation algorithm, it achieves pixel-level precise extraction of the changed areas. The output crop growth distribution map and disaster damage mask map have clear boundaries and accurate regions. The detection results can be directly used as an objective quantitative basis for agricultural production management work such as agricultural insurance loss assessment, crop yield prediction, and enforcement of farmland non-agricultural use, effectively improving the scientific nature of agricultural management decisions and possessing significant economic and application value.

[0041] (4) The device has strong inference efficiency and environmental adaptability: The detection method executed by this device abandons the complex decoder and classification head. In the inference stage, it only calls the trained backbone encoder to complete feature extraction and change detection, which greatly reduces the amount of computation and improves the detection inference efficiency. It can quickly respond to the needs of agricultural monitoring and disaster assessment. At the same time, relying on multi-view enhancement and same-modal consistency constraint training, the model has strong robustness to complex farmland environmental interference such as light, soil moisture, and sensor noise. It can maintain stable detection accuracy in various agricultural production scenarios. Attached Figure Description

[0042] Figure 1 This is a flowchart of the method for all-weather monitoring and disaster assessment of crops based on cross-modal self-distillation, as described in this invention. Detailed Implementation

[0043] The technical solutions of the embodiments of this application will be further described clearly and completely below with reference to the accompanying drawings. It should be noted that the described embodiments are only some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0044] To make the inventive objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be further described in detail below with reference to the accompanying drawings: In order to better understand the above-mentioned objectives, features, and advantages of this invention, the advantages of this invention will be further illustrated below by comparing the embodiments with the accompanying drawings and specific implementation methods.

[0045] This invention proposes a method for all-weather monitoring and disaster assessment of crops based on cross-modal self-distillation. The flowchart of this method is as follows: Figure 1 The steps of this method are explained in detail below: S1) Acquire first-phase and second-phase remote sensing images of the same farmland area at different times, and apply a data augmentation strategy based on a combination of geometric transformation and photometric transformation to each remote sensing image to generate multiple viewpoint sample images; Specifically, the process involves acquiring first- and second-phase remote sensing images of the same farmland monitoring area at different times, sequentially performing image preprocessing and multi-view data enhancement operations, generating multi-view sample image pairs, and constructing a sample dataset for self-supervised training. The remote sensing image acquisition includes dual-temporal images (such as crop sowing and growth periods) acquired by satellites or agricultural aerial platforms equipped with different sensors, such as optical and synthetic aperture radar (SAR). In the preprocessing stage, geometric correction and spatial registration techniques are first used to eliminate spatial positional deviations caused by imaging perspective, UAV flight attitude, or farmland topographic undulations, ensuring that pixels in different temporal images are strictly aligned in the geographic coordinate system, guaranteeing spatial consistency of crop plots. Secondly, radiometric normalization is performed, mapping pixel values ​​to a unified range to eliminate differences in data dimensions between different sensors and standardize the image radiometric characteristics. Finally, large-scale images are cropped into fixed-size image patches for model training, obtaining the training input rules for the adapted model. The image blocks are preprocessed. In the multi-view construction stage, the traditional manual annotation of farmland plots is abandoned. Instead, a data augmentation strategy combining geometric transformation and photometric transformation is adopted. This abandons the traditional manual annotation method and generates multi-view sample images with consistent content but different appearances for each preprocessed image block. Specifically, two sets of composite augmentation operation sequences with different random seeds are configured for each image block. Each set of sequences integrates geometric transformation and photometric transformation operations. The geometric transformation includes random horizontal flipping and random vertical flipping to simulate the farmland observation characteristics under different aerial survey perspectives. The photometric transformation includes color jitter, random grayscale conversion, and Gaussian blur to simulate the crop texture features under complex environments such as different weather conditions and sensor noise. Through the two sets of composite augmentation operation sequences, each image block is augmented separately to generate a first-view sample image and a second-view sample image, resulting in paired multi-view samples, which provide effective sample support for the subsequent self-supervised training of the self-distillation network. The specific implementation process is as follows: Two remote sensing images of the same farmland area were collected at different times and are respectively recorded as the first time-phase remote sensing image. Second-phase remote sensing images ; and Images can be from the same source (e.g., both are optical images) or from different sources (e.g., ... For optical images, (For SAR imagery); To enable the model to learn robust crop growth invariant features, a data augmentation strategy based on a combination of geometric and photometric transformations is used to augment the remote sensing imagery, constructing a multi-view sample map. Two sets of data augmentation operations are defined, namely... and This set includes geometric and photometric transformation strategies, specifically random horizontal flipping, random vertical flipping, color dithering (brightness, contrast, saturation, and hue perturbation), random grayscale conversion, and Gaussian blurring for each input image pair. Multiple viewpoint sample images are generated by applying two independent composite enhancements, and the calculation formula is as follows:

[0046]

[0047] In the formula, and This represents two different augmentation processes with different random seeds, i.e., two sets of data augmentation operations; Represents first-phase remote sensing imagery application First-person view sample image generated after enhancement operation. Represents first-phase remote sensing imagery application The second-view sample image generated after the enhancement operation. Indicates second-phase remote sensing imagery application First-person view sample image generated after enhancement operation. Indicates second-phase remote sensing imagery application The second-view sample image generated after the enhancement operation. This indicates the first enhanced view. Indicates the second enhanced view; Through the above operations, the diversity of views under different imaging conditions was simulated, forcing the model to focus on the essential semantic content of the image rather than irrelevant noise interference in subsequent training. Finally, these multiple view sample images were normalized and converted into tensor format, and then input into the subsequent cross-modal self-distillation agricultural monitoring network.

[0048] S2) Construct a cross-modal self-distillation agricultural monitoring network based on a twin architecture, including parallel branches with the same structure but different parameter update methods, namely a student network and a teacher network; input sample images from multiple perspectives into the cross-modal self-distillation agricultural monitoring network for feature extraction and mapping to a low-dimensional compact space to obtain class prediction logical values; both the student network and the teacher network consist of a backbone encoder and a projection head; define an update mechanism to update the parameters of the student network, and dynamically update the parameters of the teacher network through the exponential moving average (EMA) of the student network parameters; The entire framework of the cross-modal self-distillation agricultural monitoring network; Specifically, addressing the technical challenges of fragmented plots, blurred boundaries, and the susceptibility of crop growth characteristics to interference from complex field environments in agricultural remote sensing monitoring, a targeted structural design is implemented for the cross-modal self-distillation agricultural monitoring network. The key design points are as follows: ① Considering that farmland plots (especially in the hilly areas of the south or under the small-scale farming model) are often fragmented, traditional pooling downsampling operations are very likely to cause the loss of features of small plots or blurring of field ridge boundaries; therefore, the backbone network encoder abandons downsampling and adopts a fully convolutional structure to retain the complete pixel-level spatial resolution and ensure accurate perception of small plots and crop edges. ② In order to achieve all-weather monitoring and adapt to the collaborative processing requirements of remote sensing images of different modalities such as optical and SAR, the network adopts a twin parallel branch architecture, which supports two branches to extract features from optical and SAR heterogeneous images (such as optical images reflecting chlorophyll and SAR images reflecting crop structure) or images of the same modality, and to map them to a unified crop growth prototype space through a projection head, thereby eliminating the modal gap caused by the differences in imaging mechanisms of different sensors and extracting the core semantic features of crop growth shared in different modal images; ③ Given the complex farmland environment, where wind blowing leaves or instantaneous changes in sunlight can introduce noise, the teacher network is updated using the exponential moving average (EMA). This is equivalent to building a time-smoothed stable monitor that can filter out instantaneous environmental noise in the farmland and provide the student network with a more robust long-term crop growth status as a learning target.

[0049] The specific implementation process of constructing a cross-modal self-distillation agricultural monitoring network includes: S2.1) Backbone Encoder Construction: The backbone encoder is used to extract high-dimensional spatial features of the image. A fully convolutional structure is employed to avoid resolution loss due to downsampling, thus preserving pixel-level details crucial for change detection. First, multiple viewpoint sample images are input into the first convolutional block, which maps them to the feature space through convolutional layers. Then, batch normalization layers accelerate convergence and prevent overfitting. Finally, a nonlinearity is introduced through the ReLU activation function, the corresponding calculation formula of which is:

[0050] In the formula, Represents the ReLU activation function; This indicates batch normalization, which accelerates training convergence and prevents overfitting. This represents a convolution operation with a kernel size of 3, a stride of 1, and padding of 1 (ensuring that the output space size is consistent with the input). Indicates the number of input channels; Indicates the preset feature dimension; This represents the first layer of features, whose dimensions are consistent with the spatial dimensions of the view sample image; Next, The second convolutional block is input to further extract deeper abstract features, and the corresponding formula is:

[0051] In the formula, This represents the output features of the second layer, which increases the receptive field of the network by stacking convolutions; Then Inputting the third convolutional block expands the number of channels to the target feature dimension, enriching the expressive power of the features. The corresponding formula is:

[0052] In the formula, This represents the output features of the third layer; Finally, The fourth convolutional block is input and subjected to a nonlinear transformation to enhance the discriminative power of the features. The corresponding formula is:

[0053] In the formula, This represents the high-dimensional spatial features of the final output of the backbone encoder. The dimension represents the feature of a high-dimensional space, where, Indicates batch size, Indicates the number of feature channels. Indicates altitude, The width is represented by this feature map, which retains complete spatial structure information, providing a basis for subsequent pixel-level distance calculations. S2.2) Projection Head Construction: The projection head is connected after the backbone encoder and adopts a multilayer perceptron structure to receive high-dimensional spatial features. This is then mapped to a low-dimensional, compact space to generate class prediction logistic values ​​for calculating probability distributions; the specific implementation steps are as follows: First, global average pooling is performed on the high-dimensional spatial features output by the backbone encoder to aggregate global context information. The corresponding formula is:

[0054] In the formula, This represents the feature vector after pooling. Let represent the dimension of the feature vector after pooling, where Indicates batch size, Indicates the number of feature channels; This represents global average pooling, which averages the spatial dimensions of the feature map and aggregates global contextual information. This operation eliminates the influence of spatial location, allowing the model to focus on the overall semantics of the image. Subsequently, the pooled feature vectors are input into the first fully connected layer for linear expansion to enhance the dimensionality of the features, and a smooth non-linearity is introduced through the GELU activation function, as shown in the following formula:

[0055] In the formula, This represents the GELU activation function. This represents a fully connected layer that expands the dimension of the feature vector. Represents hidden layer features; Then, the hidden layer features The input to the second fully connected layer compresses the dimension to the bottleneck low dimension and removes redundant information. The corresponding formula is:

[0056] In the formula, Indicates bottleneck layer characteristics; Next, the bottleneck layer features L2 norm normalization is performed, and the feature is projected onto the unit hypersphere to enhance scale invariance. The corresponding formula is:

[0057] In the formula, This represents a small constant used to prevent the denominator from being zero. express L2 norm, This represents the normalized eigenvectors; Finally, a linear layer with weight normalization maps the normalized feature vector to a lower dimension, yielding the class prediction logistic value, with the corresponding formula:

[0058] In the formula, This represents the weight matrix of the layer. Weight normalization decouples the magnitude and direction of the weights, which helps to stabilize training. This represents the category prediction logic value output by the projection head; This means performing L2 norm normalization on the weight matrix to decouple the magnitude and direction of the weights, thus stabilizing the training.

[0059] S2.3) Updating teacher and student network parameters: Based on the student network structure, a teacher network with the exact same structure is instantiated, both including a backbone encoder and a projection head, ensuring that both can process the same input and output features of the same dimension; the teacher network parameters are... Initialize to student network parameters Consistent parameters are defined during the network construction phase, with the teacher network's update mechanism defined. The student network is updated through gradient backpropagation, while the teacher network does not participate in gradient backpropagation and is not directly updated by the loss function; instead, it is dynamically updated only through the exponential moving average (EMA) of the student network parameters. After training, the student network completes its gradient update, and then the EMA is executed to update the teacher network parameters. The corresponding update formula for the teacher network parameters is:

[0060] In the formula, Indicates the momentum coefficient. Indicates the teacher's network parameters. Indicates student network parameters; This mechanism allows the teacher network to become the historical average of the student network, resulting in a more stable and less noisy feature representation.

[0061] S3) Based on the category prediction logic value, calculate the target probability distribution of the teacher network and the predicted probability of the student network respectively; Specifically, this collaborative training process aims to enable the model to automatically learn robust crop growth characteristics without human annotation. To address the issue that large-scale planting of a single crop in farmland scenarios (with uneven category distribution) can easily lead to a homogenized model output, a dynamic centering and temperature coefficient sharpening mechanism is introduced in the teacher network target probability distribution generation stage to prevent model collapse due to homogenized output during training. This mechanism maintains a moving average center vector and subtracts the mean of the teacher output to prevent the model from "collapsed" to only be able to identify the background or a single dominant crop. The temperature coefficient is used to sharpen the probability distribution and generate high-confidence soft labels for crop categories, providing a stable and reliable basis for supervised learning for the student network. The specific implementation process includes: S3.1) Calculation of the target probability distribution of the teacher network: The center vector is updated based on the class prediction logistic values ​​output by the teacher network to prevent the model from collapsing to a single class. The formula for the center vector is:

[0062] In the formula, The center vector; The center momentum coefficient controls the degree of historical information retention; the larger the value, the smoother the center update. Batch size; For the current batch The category prediction logical value output by the teacher's network projection head for each sample; Subsequently, the center vector output by the teacher network is processed using the center vector and temperature coefficient to generate a sharpened target probability distribution (centered soft label), with the corresponding formula as follows:

[0063] In the formula, This represents the temperature coefficient. A smaller temperature coefficient results in a sharper distribution, enhancing the focus on high-confidence prototypes. express Activation function Indicates input sample The corresponding class prediction logic value output by the teacher's network projector. Represents the target probability distribution. Indicates the number of input networks in the current batch. One sample; S3.2) Based on the category prediction logistic values ​​output by the student network, calculate the prediction probability, and the corresponding formula is:

[0064] In the formula, This represents the temperature coefficient of the student network. This represents the category prediction logical value output by the student's network projector; Indicates the predicted probability; S4) Train the cross-modal self-distillation agricultural monitoring network, calculate the cross-modal distillation loss and the same-modal distillation loss, and obtain the total loss function by weighted summation; minimize the total loss through backpropagation to update the student network parameters, and update the teacher network parameters and the centering vector until a steady state is reached; The calculation of the loss function incorporates dual constraints to adapt to agricultural applications: cross-modal distillation loss addresses the challenge of all-weather monitoring by forcing the student view of optical imagery (susceptible to cloud and fog interference) to predict the teacher view of SAR imagery (all-weather imaging) (and vice versa). This mechanism forces the model to ignore the differences in sensor imaging mechanisms and extract the crop biomass and structural semantics shared by optical and SAR, thereby enabling accurate crop monitoring using SAR data even in foggy weather. Intramodal distillation loss addresses interference from the field environment by forcing the enhanced views of the same temporal imagery under different lighting or viewing angles to maintain consistent features, improving the model's robustness to complex agricultural environments in the field. By minimizing this total loss, the model ultimately learns a semantic prototype that is sensitive to crop growth status but insensitive to environmental noise. The differentiated parameter update mechanism for teacher-student networks: With the total loss function as the optimization objective, all learnable parameters of the student network are iteratively updated using the gradient backpropagation algorithm, so that the predicted probability of the student network continuously approaches the high-confidence soft label generated by the teacher network. The teacher network does not participate in gradient backpropagation, but is passively and synchronously updated based on the updated parameters of the student network using an exponential moving average method. This ensures that the parameters of the teacher network are always a smooth set of the historical iteration parameters of the student network, guaranteeing the stability and robustness of the target probability distribution output by the teacher network, and continuously providing high-quality supervision signals to the student network. S4.1) Calculation of the total loss function: Construct a total loss function that includes both cross-modal distillation loss and intra-modal distillation loss. First, calculate the cross-modal distillation loss. This forces images of different modalities to align in semantic space, eliminating modal differences. The corresponding formula is:

[0065] In the formula, This represents the standard cross-entropy loss; In the student network, the first The first modal image Sample images from various perspectives; In the teacher network, the first Viewpoint sample images of modal images; Indicates the first The first modal image Predicted probability distribution of sample images from each viewpoint; Indicates the first Target probability distribution in the viewpoint sample map of modal imagery; Through this cross-supervision, regardless of which augmented view the student network sees... still The features extracted should all be able to match another modality. The semantic prototype allows the model to learn modality-invariant semantic prototypes; Next, the teacher network is established in the first augmented view. Above, the student network is built in a second enhanced view. The above forces different views under the same modality to remain consistent, improving feature robustness. The corresponding formula is:

[0066] In the formula, Indicates the first Second-view sample images generated from each modal image after data augmentation This indicates that the student network input was obtained. The corresponding predicted probability; Indicates the first First-view sample images generated from modal images after data augmentation This indicates input into the teacher network to obtain... The corresponding target probability distribution; Both the first-view sample image and the second-view sample image originate from the same modal image and are generated through different data augmentation methods. They are semantically consistent but differ in appearance representation. In the process of calculating the same-modal distillation loss, the target probability distribution generated by the teacher network under the first-view sample image is used as a supervision signal to guide the student network to match the predicted probability distribution under the second-view sample image. This constrains the consistency of feature representation of the same modal image under different view conditions and improves the robustness of the model to factors such as illumination changes, observation angle changes and sensor noise.

[0067] Finally, the weighted summation yields the total loss function. The corresponding formula is:

[0068] In the formula, and The student network is updated via backpropagation by minimizing preset weight coefficients. S4.2) The specific process of parameter update and training iteration includes: In each training batch, the gradient cache of the student network is first zeroed out, followed by forward propagation: the sample images from the four perspectives of the current batch are input into the teacher network and the student network respectively. The forward propagation of the teacher network is performed without constructing a backward computation graph, and is only used to obtain the target probability distribution. and the category prediction logical value for the current batch; Subsequently, based on the total loss function Backpropagation is performed on the student network backbone encoder and projector to calculate the gradients of each parameter. To prevent gradient explosion from causing training instability, the gradients of the student network parameters are pruned using the L2 norm, limiting the global norm of the gradients to a threshold value. Within this range, the corresponding formula is:

[0069] After gradient pruning, the RMSprop optimizer is used to optimize the student network parameters. Perform gradient descent updates with a learning rate of The weight decay coefficient is The momentum coefficient is This drives the student network to converge gradually by minimizing the total loss; After the student network parameters are updated, the teacher network parameters are simultaneously updated using an exponential moving average method, without participating in gradient calculation. The corresponding formula is:

[0070] In the formula, Indicates the momentum coefficient. This indicates the updated teacher network parameters. For the updated student network parameters; At the same time, in accordance with the method described in S3.1), the global centralization vector is updated synchronously using the category prediction logic value of the current batch of teacher networks, so that the output of the next batch of teachers can be decentralized with the latest central vector, maintaining the stability of the output distribution and preventing the model from collapsing.

[0071] S5) Using the trained student network backbone encoder, extract pixel-level feature maps from the first and second time-phase remote sensing images of the farmland area to be detected; calculate the Euclidean distance of the pixel-level feature maps and perform normalization to generate a change intensity map; use the maximum inter-class variance method to perform threshold segmentation on the change intensity map to generate a crop growth distribution map or a disaster damage mask map. Specifically, during the inference phase, the model removes the projection head used for self-supervised training and only enables the trained student network backbone encoder as the crop feature extractor. The registered bi-temporal images of the target region are then used for this purpose. Input the encoders respectively, where, and It can represent different key nodes in the entire crop life cycle (such as the sowing period and the jointing period) or the time phases before and after a sudden agricultural disaster; the encoder extracts deep feature maps that retain pixel-level spatial structure ( These feature maps implicitly contain key agricultural information such as crop biomass and canopy structure. Subsequently, the Euclidean distance (L2 norm) between two feature maps is calculated pixel-by-pixel in the feature space to generate an original distance map. To intuitively reflect crop growth activity or the severity of damage, the original distance map undergoes max-min normalization, mapping the values ​​to... The interval is used to obtain a continuous crop condition change intensity map (CMI). Finally, the Otsu method is used to adaptively calculate the global optimal threshold, which can maximize the inter-class variance between significantly changed areas (such as fast-growing areas / disaster-stricken areas) and non-changed areas (such as fallow land / undisaster-stricken areas). Based on this, the intensity map is divided into a binary mask map to accurately output the spatial distribution of crop growth areas or the outline of disaster-stricken plots. The final crop growth distribution map or disaster damage mask map is generated. The specific implementation process includes: After the model training is completed, only the backbone encoder of the student network is retained for inference; firstly, the first temporal remote sensing image to be detected is... Second-phase remote sensing images The dual-temporal image input encoder extracts pixel-level feature maps, and the corresponding formula is:

[0072] In the formula, This represents the pixel-level feature map output by the encoder. Indicates encoder; Next, the Euclidean distance between the two pixel-level feature maps at each pixel location is calculated using the following formula:

[0073] In the formula, Represents pixels Euclidean distance at; and Represents pixel-level feature maps and In pixels The feature vector of the location; Represents the L2 norm; Subsequently, the Euclidean distance is normalized to its minimum and maximum values, and then mapped to a range. The formula for generating a change intensity map within the interval is:

[0074] In the formula, and Let represent the global minimum and maximum values ​​of the Euclidean distance, respectively. A graph showing the intensity of change after normalization; Finally, the Otsu's method is used to iterate through all possible candidate thresholds. Calculate the inter-class variance between the foreground (changing region) and the background (unchanging region) at each threshold, and take the threshold with the largest inter-class variance as the global optimal threshold. and to Binarization segmentation is performed to ultimately generate a binary-transformed crop growth distribution map or disaster damage mask map. The corresponding formula is:

[0075] In the formula, Represents the variance between classes. Indicates the percentage of background pixels. This indicates the percentage of pixels in the foreground. This represents the average value of the background pixels. Indicates the average value of foreground pixels. This represents the candidate threshold for traversal. This represents a map showing the distribution of crop growth or a mask map showing damage caused by disasters.

[0076] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0077] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for all-weather monitoring and disaster assessment of crops based on cross-modal self-distillation, characterized in that, Includes the following steps: S1) Acquire first-phase and second-phase remote sensing images of the same farmland area at different times, and apply a data augmentation strategy based on a combination of geometric transformation and photometric transformation to each remote sensing image to generate multiple viewpoint sample images; S2) Construct a cross-modal self-distillation agricultural monitoring network based on a twin architecture, including parallel branches with the same structure but different parameter update methods, namely a student network and a teacher network; input sample images from multiple perspectives into the cross-modal self-distillation agricultural monitoring network for feature extraction and mapping to a low-dimensional compact space to obtain class prediction logical values; both the student network and the teacher network consist of a backbone encoder and a projection head; define an update mechanism to update the parameters of the student network, and dynamically update the parameters of the teacher network through the exponential moving average (EMA) of the student network parameters; S3) Based on the category prediction logic value, calculate the target probability distribution of the teacher network and the predicted probability of the student network respectively; S4) Train the cross-modal self-distillation agricultural monitoring network, calculate the cross-modal distillation loss and the same-modal distillation loss, and obtain the total loss function by weighted summation; minimize the total loss through backpropagation to update the student network parameters, and update the teacher network parameters and the centering vector until a steady state is reached; The calculation process of the total loss function includes: First, calculate the cross-modal distillation loss. This allows images of different modalities to be aligned in the semantic space, and the corresponding formula is: , In the formula, This represents the standard cross-entropy loss; In the student network, the first The first modal image Sample images from various perspectives; In the teacher network, the first Viewpoint sample images of modal images; Indicates the first The first modal image Predicted probability of a sample image from a single viewpoint; Indicates the first Target probability distribution in the viewpoint sample map of modal imagery; Next, the teacher network is established in the first augmented view. Above, the student network is built in a second enhanced view. To maintain consistency between different views under the same mode, calculate the distillation loss under the same mode. The corresponding formula is: , In the formula, Indicates the first Second-view sample images generated from modal images after data augmentation This indicates that the student network input was obtained. The corresponding predicted probability; Indicates the first First-view sample images generated from modal images after data augmentation This indicates input into the teacher network to obtain... The corresponding target probability distribution; Finally, the weighted summation yields the total loss function. The corresponding formula is: , In the formula, and The student network is updated via backpropagation by minimizing preset weight coefficients. S5) Using the trained student network backbone encoder, extract pixel-level feature maps from the first and second time-phase remote sensing images of the farmland area to be detected; calculate the Euclidean distance of the pixel-level feature maps and perform normalization to generate a change intensity map; use the maximum inter-class variance method to perform threshold segmentation on the change intensity map to generate a crop growth distribution map or a disaster damage mask map.

2. The method for all-weather monitoring and disaster assessment of crops based on cross-modal self-distillation according to claim 1, characterized in that, In step S1), the specific implementation process of generating multiple viewpoint sample images is as follows: Two remote sensing images of the same farmland area were collected at different times and are respectively recorded as the first time-phase remote sensing image. Second-phase remote sensing images ; This paper describes a data augmentation strategy based on a combination of geometric and photometric transformations to enhance remote sensing images and construct multi-view sample maps. Two data augmentation strategies are defined as follows: and For each pair of input images The data augmentation strategy is applied twice to generate sample images from multiple viewpoints. The calculation formula is as follows: , , In the formula, and This represents two different augmentation processes using two random seeds, i.e., two data augmentation strategies; Represents first-phase remote sensing imagery application First-person view sample image generated after enhancement operation. Represents first-phase remote sensing imagery application The second-view sample image generated after the enhancement operation. Indicates second-phase remote sensing imagery application First-person view sample image generated after enhancement operation. Indicates second-phase remote sensing imagery application The second-view sample image generated after the enhancement operation. This indicates the first enhanced view. Indicates the second enhanced view; Finally, the sample images from multiple perspectives are normalized and converted into tensor format, and then input into the subsequent cross-modal self-distillation agricultural monitoring network.

3. The method for all-weather monitoring and disaster assessment of crops based on cross-modal self-distillation according to claim 1, characterized in that, In step S2), multiple viewpoint sample images are input into a cross-modal autodistillation agricultural monitoring network for feature extraction and mapping to a low-dimensional, compact space to obtain category prediction logical values. Specifically, this includes: S2.1) Backbone encoder: The backbone encoder adopts a fully convolutional structure to extract high-dimensional spatial features of the image; First, sample images from multiple viewpoints are input into the first convolutional block, and then mapped to the feature space through a convolutional layer. Next, a batch normalization layer is used to accelerate convergence and prevent overfitting. Finally, a nonlinearity is introduced through the ReLU activation function, the corresponding calculation formula of which is: , In the formula, Represents the ReLU activation function; Indicates batch normalization; This represents a convolution operation with a kernel size of 3, a stride of 1, and padding of 1. Indicates the number of input channels; Indicates the preset feature dimension; This represents the first layer of features, whose dimensions are consistent with the spatial dimensions of the view sample image; Next, The second convolutional block is input to further extract deeper abstract features, and the corresponding formula is: , In the formula, This represents the output features of the second layer; Then, Input the third convolutional block to expand the number of channels to the target feature dimension, as shown in the formula: , In the formula, This represents the output features of the third layer; Finally, The fourth convolutional block is input and subjected to a nonlinear transformation to enhance the discriminative power of the features. The corresponding formula is: , In the formula, This represents the high-dimensional spatial features of the final output of the backbone encoder. The dimension represents the feature of a high-dimensional space, where, Indicates batch size, Indicates the number of feature channels. Indicates altitude, Indicates width; S2.2) Projection Head: The projection head is connected after the backbone encoder and adopts a multilayer perceptron structure to receive high-dimensional spatial features. This is then mapped to a low-dimensional, compact space to generate class prediction logistic values ​​for calculating the probability distribution; the specific implementation steps are as follows: First, global average pooling is performed on the high-dimensional spatial features output by the backbone encoder to aggregate global context information. The corresponding formula is: , In the formula, This represents the feature vector after pooling. Let represent the dimension of the feature vector after pooling, where Indicates batch size, Indicates the number of feature channels; Indicates global average pooling; Subsequently, the pooled feature vectors are input into the first fully connected layer for linear expansion to enhance the dimensionality of the features, and a smooth non-linearity is introduced through the GELU activation function, as shown in the following formula: , In the formula, This represents the GELU activation function. Indicates a fully connected layer. Represents hidden layer features; Then, the hidden layer features The input to the second fully connected layer compresses the dimension to the bottleneck low dimension, and the corresponding formula is: , In the formula, Indicates bottleneck layer characteristics; Next, the bottleneck layer features Normalize to the L2 norm and project it onto the unit hypersphere; the corresponding formula is: , In the formula, This represents a small constant used to prevent the denominator from being zero. express L2 norm, This represents the normalized eigenvectors; Finally, a linear layer with weight normalization maps the normalized feature vector to a lower dimension, yielding the class prediction logistic value, with the corresponding formula: , In the formula, This represents the weight matrix of the layer; This represents the category prediction logic value output by the projection head; This indicates that the weight matrix is ​​normalized using the L2 norm.

4. The method for all-weather monitoring and disaster assessment of crops based on cross-modal self-distillation according to claim 1, characterized in that, In step S2), the specific implementation process of the update mechanism includes: Teacher network parameters Initialize to student network parameters With consistent parameters, during the update process, the student network completes the update by participating in gradient backpropagation, while the teacher network does not participate in gradient backpropagation and is only dynamically updated through the exponential moving average (EMA) of the student network parameters. The update formula for the teacher network parameters is as follows: , In the formula, Indicates the momentum coefficient. Indicates the teacher's network parameters. This represents the student's network parameters.

5. The method for all-weather monitoring and disaster assessment of crops based on cross-modal self-distillation according to claim 1, characterized in that, In step S3), the calculation processes for the target probability distribution of the teacher network and the predicted probability of the student network are as follows: S3.1) The calculation process of the target probability distribution of the teacher network includes: First, the center vector is updated based on the category prediction logistic values ​​output by the teacher network. The formula for the center vector is: , In the formula, Represents the center vector. Indicates the central momentum coefficient. Indicates batch size, Indicates the current batch number The class prediction logistic value output by the teacher's network projection head for each sample. ; Subsequently, the class prediction logistic values ​​output by the teacher network are processed using the center vector and temperature coefficient to generate a sharpened target probability distribution, with the corresponding formula as follows: , In the formula, Indicates the temperature coefficient. express Activation function Indicates input sample The corresponding class prediction logic value output by the teacher's network projector. Represents the target probability distribution. Indicates the number of input networks in the current batch. One sample; S3.2) The calculation process for the predicted probability of the student network includes: Based on the category prediction logistic values ​​output by the student network, the prediction probability is calculated using the following formula: , In the formula, This represents the temperature coefficient of the student network; This represents the category prediction logical value output by the student's network projector; This represents the predicted probability.

6. The method for all-weather monitoring and disaster assessment of crops based on cross-modal self-distillation according to claim 1, characterized in that, In step S4), the process of minimizing the total loss through backpropagation to update the student network parameters and simultaneously updating the teacher network parameters and the centering vector includes: In each training batch, the gradient cache of the student network is first zeroed out, followed by forward propagation: multiple viewpoint sample maps are input into the teacher network and the student network respectively. The forward propagation of the teacher network is performed without constructing a backward computation graph, and is only used to obtain the target probability distribution. and the category prediction logical value for the current batch; Subsequently, based on the total loss function Backpropagation is performed on the student network backbone encoder and projector to calculate the gradients of each parameter; the gradients of the student network parameters are then L2 norm clipped to restrict the global norm of the gradients to a threshold value. Within this range, the corresponding formula is: , After gradient pruning, the RMSprop optimizer is used to optimize the student network parameters. Perform gradient descent updates to drive the student network to converge gradually by minimizing the total loss; After the student network parameters are updated, the teacher network parameters are simultaneously updated using an exponential moving average method, without participating in gradient calculation. The corresponding formula is: , In the formula, Indicates the momentum coefficient. This indicates the updated teacher network parameters. For the updated student network parameters; At the same time, the global centralization vector is updated synchronously using the category prediction logical values ​​of the current batch of teacher networks, so that the output of the next batch of teachers can be decentralized using the latest centralization vector.

7. The method for all-weather monitoring and disaster assessment of crops based on cross-modal self-distillation according to claim 1, characterized in that, In step S5), a crop growth distribution map or a disaster damage mask map is generated. The specific implementation process includes: First, the first temporal remote sensing image to be detected... Second-phase remote sensing images The dual-temporal image input encoder extracts pixel-level feature maps, and the corresponding formula is: , In the formula, This represents the pixel-level feature map output by the encoder. Indicates encoder; Next, the Euclidean distance between the two pixel-level feature maps at each pixel location is calculated using the following formula: , In the formula, Represents pixels Euclidean distance at; and Represents pixel-level feature maps and In pixels The feature vector of the location; Represents the L2 norm; Subsequently, the Euclidean distance is normalized to its minimum and maximum values, and then mapped to a range. The formula for generating a change intensity map within the interval is: , In the formula, and Let represent the global minimum and maximum values ​​of the Euclidean distance, respectively. A graph showing the intensity of change after normalization; Finally, the Otsu's method is used to iterate through all possible candidate thresholds. Calculate the inter-class variance between foreground and background at each threshold, and take the threshold with the largest inter-class variance as the global optimal threshold. The intensity change map is then binarized to generate a binary crop growth distribution map or a disaster damage mask map. The corresponding formula is: , , In the formula, Represents the variance between classes. Indicates the percentage of background pixels. This indicates the percentage of pixels in the foreground. This represents the average value of the background pixels. Indicates the average value of foreground pixels. This represents the candidate threshold for traversal. This represents a map showing the distribution of crop growth or a mask map showing damage caused by disasters.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.

9. A crop all-weather monitoring and disaster assessment detection device based on cross-modal self-distillation learning, comprising a processor and a memory, wherein the memory stores a computer program, characterized in that, When the computer program is executed by the processor, it implements the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Wheat waterlogging accurate identification and disaster assessment method based on multi-temporal remote sensing image

    CN121564430A