Early disease prediction method and system of wheat stripe rust remote sensing change information driven

By progressively fusing multi-source data and using a frequency domain decoupling change detection module, combined with meteorological data and prior knowledge of diffusion patterns, the problem of insufficient multi-source data fusion and inadequate utilization of temporal change information in wheat stripe rust prediction has been solved, achieving high-precision prediction and dynamic early warning for the early stages of the disease.

CN121009293BActive Publication Date: 2026-07-31UNIV OF SCI & TECH BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
UNIV OF SCI & TECH BEIJING
Filing Date
2025-07-17
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies for predicting wheat stripe rust suffer from limited multi-source data fusion methods, insufficient utilization of temporal variation information, and weak collaborative modeling capabilities between images and environmental factors. This results in low accuracy in early disease identification, making it difficult to achieve early warning and precise control.

Method used

By employing a multi-source data progressive fusion module to deeply integrate natural images and multispectral images, combined with a frequency domain decoupled change detection module and a conditional diffusion prediction model, and through time-domain-frequency domain collaborative analysis and meteorological data and prior knowledge of diffusion patterns, the spread of diseases is dynamically simulated, enabling sensitive identification and high-precision prediction of diseases in their early stages.

Benefits of technology

It improves the ability to characterize disease features, enables sensitive identification and efficient data representation of early-stage diseases, and dynamically and accurately predicts disease development, thereby enhancing the accuracy and real-time nature of predictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009293B_ABST
    Figure CN121009293B_ABST
Patent Text Reader

Abstract

This invention provides a method and system for early prediction of wheat stripe rust driven by remote sensing change information, comprising: a multi-source data progressive fusion module, which sequentially stitches together corresponding hierarchical features from natural images and multispectral images and fuses them to output deep fusion features; a frequency domain decoupling-based change detection module, which outputs change probability maps of adjacent time points; constructing multiple sets of training samples to input into a conditional diffusion prediction model for training; collecting remote sensing image data of wheat to be predicted and corresponding meteorological data to obtain a change probability map, fusing it along with the corresponding meteorological data and prior data on diffusion patterns as a conditional vector, and inputting it along with random Gaussian noise into the trained conditional diffusion prediction model to obtain denoised prediction data, which is then processed by a visual decoder to obtain a predicted change probability map; based on this, the prediction results of the severity and distribution range of wheat stripe rust on day d are obtained. This invention can predict early wheat stripe rust.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wheat stripe rust prediction technology, and in particular to an early disease prediction method and system driven by remote sensing change information of wheat stripe rust. Background Technology

[0002] Wheat production has long been threatened by pests and diseases, with stripe rust being particularly severe. This disease is caused by the obligate parasitic fungus *Streptococcus stripe*, which specifically infects the green tissues of the plant, leading to yellowing and necrosis of the photosynthetic tissues. Furthermore, wheat stripe rust exhibits climate-sensitive epidemic characteristics, and under specific environmental conditions, it can form regional outbreaks, causing significant yield losses and is listed as a globally important crop biological disaster. Therefore, timely and accurate prediction of the occurrence and spread of stripe rust is crucial for effectively controlling this disease and ensuring wheat yield and quality. Predicting stripe rust occurrence provides a scientific basis for developing reasonable control measures, reducing pesticide use, minimizing environmental pollution, and ensuring food security. However, predicting stripe rust faces significant challenges. Conventional detection methods rely on identifying yellow lesions, but by the time lesions appear, the pathogen has already spread through the vascular system, resulting in a loss of timeliness in control measures.

[0003] Traditional statistical analysis-based methods rely on the collection and analysis of large amounts of historical data, using manual feature extraction to establish the correlation between meteorological conditions and disease occurrence. However, due to the strong nonlinear characteristics of diseases, these methods have limited data mining capabilities and less than ideal predictive results when faced with complex and ever-changing environmental factors.

[0004] In contrast, deep learning prediction methods can automatically extract high-dimensional features from massive amounts of data, breaking through the bottleneck of manual feature selection in traditional methods, and exhibiting stronger adaptability and higher prediction accuracy, thus gaining widespread application.

[0005] However, existing deep learning prediction methods have the following limitations:

[0006] 1. Existing multi-source data fusion methods are simplistic and have limited feature representation capabilities. Early-stage diseases primarily manifest as subtle changes in plant internal structure and physiological parameters, making them difficult to detect using natural images alone. Although existing methods have attempted to fuse natural images with multispectral images to enhance detection capabilities, these fusion methods are typically limited to image stacking or feature splicing, lacking in-depth exploration of the correlations between features at different levels. This results in limited fusion effectiveness, feature redundancy, or information loss, impacting disease representation capabilities and model recognition accuracy.

[0007] 2. Lack of modeling capabilities for temporal variations makes it difficult to identify weak signals during the latent period of diseases. Wheat stripe rust lesions during the latent period are typically subtle, with small variations and low signal-to-noise ratios, easily overwhelmed by noise during multi-source data fusion. However, existing deep learning methods are mostly based on direct detection frameworks, relying solely on single-temporal images to extract static features for disease identification. They fail to fully utilize the variation information contained in multi-temporal remote sensing data, making it difficult to achieve sensitive perception and accurate identification of the early evolution of diseases.

[0008] 3. Insufficient ability to collaboratively model images with environmental factors (such as meteorological data and prior knowledge of diffusion patterns) makes it difficult to characterize the dynamic propagation patterns of diseases. The occurrence and development of wheat stripe rust are influenced by time, space, and environmental conditions, exhibiting significant spatiotemporal dynamics and nonlinear characteristics. However, existing methods mostly employ traditional time-series modeling strategies, failing to effectively integrate remote sensing images with environmental factors. This results in insufficient ability to simulate disease propagation trends, making it difficult to meet the application requirements for early warning and precise control. Summary of the Invention

[0009] To address the technical problems existing in the prior art, this invention provides a method and system for early disease prediction of wheat stripe rust driven by remote sensing change information, the technical solution of which is as follows:

[0010] On the one hand, a method for early disease prediction of wheat stripe rust driven by remote sensing change information is provided, the method including:

[0011] S1. Collect and preprocess historical remote sensing image data of wheat, corresponding historical meteorological data, and prior data on diffusion modes. The historical remote sensing image data includes historical natural image data and historical multispectral image data.

[0012] S2. Input the historical remote sensing image data and train the multi-source data progressive fusion module to process the preprocessed natural images. and multispectral images Multi-scale features are sequentially spliced ​​together and fused with corresponding hierarchical features. A weighted strategy is used to enhance the response in changing regions, outputting a deep fused feature at a unified scale. ;

[0013] S3, the above Input and train a frequency-domain decoupling-based change detection module, Pairing the data in chronological order, Fourier transforms are performed on each pair, followed by Gaussian high-pass and low-pass filtering to obtain high-frequency and low-frequency decoupling features. These features are then inversely transformed and fused with the time-domain feature difference results to output a probability map of changes between adjacent time points. , , ;

[0014] S4. Using a sliding window of size d in the probability map of change. , , Multiple training samples are constructed sequentially by sliding along a fixed length. The change feature obtained by encoding the last change probability map in each training sample through a visual encoder is used as noise-added data. The remaining change probability maps in each training sample, along with the corresponding historical meteorological data and prior data of diffusion mode, are fused as a conditional vector and input into the conditional diffusion prediction model for model training.

[0015] S5. Collect remote sensing image data of wheat to be predicted for day d-1 and corresponding meteorological data. Input the remote sensing image data into the trained multi-source data progressive fusion module and the frequency domain decoupling-based change detection module to obtain the change probability map. , , , will the , , The corresponding meteorological data and prior data on diffusion patterns are fused together as a conditional vector, and then input along with random Gaussian noise into the trained conditional diffusion prediction model to perform a denoising operation, resulting in denoised prediction data. Then, through a visual decoder, the predicted probability map of change on day d is obtained. ;

[0016] S6. Based on the multispectral image in the wheat remote sensing image data to be predicted on day d-1, calculate the NDVI mask, retaining only the vegetation area, and then... The NDVI mask is multiplied pixel by pixel, and a lesion mask is generated with a preset threshold. The proportion of lesion pixels to vegetation pixels is counted as the prediction result of the severity and distribution range of wheat stripe rust on day d.

[0017] On the other hand, a wheat stripe rust early disease prediction system driven by remote sensing change information is provided, the system comprising:

[0018] The acquisition and preprocessing module is used to acquire and preprocess historical remote sensing image data of wheat, corresponding historical meteorological data, and prior data on diffusion modes. The historical remote sensing image data includes historical natural image data and historical multispectral image data.

[0019] The first training module is used to input the historical remote sensing image data and train the multi-source data progressive fusion module, which then processes the pre-processed natural image data. and multispectral images Multi-scale features are sequentially spliced ​​together and fused with corresponding hierarchical features. A weighted strategy is used to enhance the response in changing regions, outputting a deep fused feature at a unified scale. ;

[0020] The second training module is used to train the... Input and train a frequency-domain decoupling-based change detection module, Pairing the data in chronological order, Fourier transforms are performed on each pair, followed by Gaussian high-pass and low-pass filtering to obtain high-frequency and low-frequency decoupling features. These features are then inversely transformed and fused with the time-domain feature difference results to output a probability map of changes between adjacent time points. , , ;

[0021] Build a module for using a sliding window of size d in the probability map of change. , , Multiple training samples are constructed sequentially by sliding along a fixed length. The change feature obtained by encoding the last change probability map in each training sample through a visual encoder is used as noise-added data. The remaining change probability maps in each training sample, along with the corresponding historical meteorological data and prior data of diffusion mode, are fused as a conditional vector and input into the conditional diffusion prediction model for model training.

[0022] The prediction module collects remote sensing image data of wheat to be predicted for d-1 days and corresponding meteorological data. It then inputs the remote sensing image data into a trained multi-source data progressive fusion module and a frequency-domain decoupling-based change detection module to obtain a change probability map. , , , will the , , The corresponding meteorological data and prior data on diffusion patterns are fused together as a conditional vector, and then input along with random Gaussian noise into the trained conditional diffusion prediction model to perform a denoising operation, resulting in denoised prediction data. Then, through a visual decoder, the predicted probability map of change on day d is obtained. ;

[0023] The calculation module is used to calculate an NDVI mask based on the multispectral image in the wheat remote sensing image data to be predicted on day d-1, retaining only the vegetation area. The NDVI mask is multiplied pixel by pixel, and a lesion mask is generated with a preset threshold. The proportion of lesion pixels to vegetation pixels is counted as the prediction result of the severity and distribution range of wheat stripe rust on day d.

[0024] On the other hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the aforementioned early disease prediction method driven by remote sensing change information of wheat stripe rust.

[0025] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, the at least one instruction being loaded and executed by a processor to implement the above-mentioned early disease prediction method driven by remote sensing change information of wheat stripe rust.

[0026] The beneficial effects of the technical solution provided by this invention include at least the following:

[0027] 1. The multimodal progressive fusion module based on an adaptive strategy designed in this invention deeply fuses natural images with multispectral images and other multi-source data, realizing complementary utilization of different spectral information and effectively improving the ability to characterize disease features.

[0028] 2. The frequency domain decoupling-based change detection module designed in this invention fully explores the subtle change signals caused by diseases in wheat during the incubation period by analyzing the difference information of dual-phase images in the time domain, high frequency and low frequency, so as to achieve sensitive identification and efficient data representation of the early stage of disease.

[0029] 3. The conditional diffusion prediction model designed in this invention models the changing trend of disease evolution over time, and introduces meteorological data and prior knowledge of diffusion mode as constraint vectors to improve the model's ability to characterize temporal features and achieve dynamic and accurate prediction of disease development. Attached Figure Description

[0030] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 This is a flowchart of an early disease prediction method for wheat stripe rust driven by remote sensing change information, provided in an embodiment of the present invention.

[0032] Figure 2 This is a preprocessing flowchart provided in an embodiment of the present invention;

[0033] Figure 3 This is a block diagram of the multi-source data progressive fusion module provided in an embodiment of the present invention;

[0034] Figure 4This is a block diagram of the feature extraction block structure provided in an embodiment of the present invention;

[0035] Figure 5 This is a schematic diagram of high-order feature enhancement provided in an embodiment of the present invention;

[0036] Figure 6 This is a block diagram of the change detection module based on frequency domain decoupling provided in an embodiment of the present invention;

[0037] Figure 7 This is a block diagram of the visual encoder structure provided in an embodiment of the present invention;

[0038] Figure 8 This is a flowchart of the training process of the conditional diffusion prediction model provided in this embodiment of the invention;

[0039] Figure 9 This is a flowchart of the denoising operation of the conditional diffusion prediction model provided in the embodiments of the present invention;

[0040] Figure 10 This is a block diagram of the visual decoder structure provided in an embodiment of the present invention;

[0041] Figure 11 This is a block diagram of the conditional diffusion prediction model provided in the embodiments of the present invention;

[0042] Figure 12 This is a block diagram of an early disease prediction system for wheat stripe rust driven by remote sensing change information, provided in an embodiment of the present invention.

[0043] Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0044] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0045] To overcome the limitations of existing technologies in wheat stripe rust prediction, such as the reliance on a single multi-source data fusion method, insufficient utilization of temporal change information, and weak collaborative modeling capabilities between images and environmental factors, this invention aims to provide an early disease prediction method for wheat stripe rust driven by remote sensing change information. This method innovatively predicts dynamic "change information" instead of traditional static "state information." Compared to directly detecting plant status, change information can more sensitively capture weak signals during the disease's incubation period, thereby improving the accuracy and real-time performance of predictions. First, this invention utilizes a multi-source progressive fusion module based on an adaptive strategy to deeply integrate multi-source data, including natural images and multispectral images, fully leveraging the complementary features between these data sources to improve data quality. Second, based on a change detection framework, a time-domain-frequency domain collaborative analysis method is employed. By analyzing the differences in time, high-frequency, and low-frequency information between dual-temporal images, the subtle change signals caused by the disease during the wheat's incubation period are fully explored, enabling sensitive identification of changes in the early stages of the disease. Finally, a conditional diffusion prediction model was designed, using meteorological data (temperature, humidity, wind speed, etc.) and prior knowledge of diffusion methods (such as disease transmission type, transmission stage, transmission direction, diffusion intensity level, etc.) as constraint vectors. Combined with the change probability diagram, the model dynamically simulates the spatiotemporal diffusion law of stripe rust through an iterative process of forward noise addition and reverse noise reduction, thereby achieving high-quality prediction of the severity and range of wheat stripe rust.

[0046] This invention provides an early disease prediction method for wheat stripe rust driven by remote sensing change information. This method can be implemented by an electronic device, which can be a terminal or a server. Figure 1 The diagram shown is a flowchart of the method. The processing flow may include the following steps:

[0047] S1. Collect and preprocess historical remote sensing image data of wheat, corresponding historical meteorological data, and prior data on diffusion modes. The historical remote sensing image data includes historical natural image data and historical multispectral image data.

[0048] This invention utilizes a drone equipped with a MicaSense Altum-PT multispectral camera to conduct low-altitude aerial photography of farmland areas, simultaneously acquiring high-resolution data in blue, green, red, red-edge, near-infrared bands, and natural imagery. Real-time dynamic differential technology is used to precisely control the flight position and aerial photography path, ensuring temporal and spatial consistency of data acquisition and covering the entire farmland area. During the acquisition process, aerial photography missions are performed at the same time each day to reduce the impact of lighting variations and environmental factors on data quality. Image stitching technology is then used to generate a complete regional coverage map. To further improve the geometric matching accuracy of the data, high-reflectivity ground calibration points are deployed in the farmland area as geometric correction reference points for the image data, ensuring the consistency of registration across multiple time periods. During the data acquisition process, the daily acquisition time is used as the reference point. For time series indexing ( The initial base date is [date], and subsequent days are [dates] in sequence. The acquired blue, green, red, red-edge, and near-infrared band image sequences and natural image sequences are respectively represented as follows: , , , , and ( In addition, environmental information of the farmland area on the same day, including precipitation, is recorded simultaneously. ,temperature and wind speed This is to supplement the sensing time series data of disease growth.

[0049] like Figure 2 As shown, the preprocessing in this embodiment of the invention includes:

[0050] First, radiometric correction is performed on both the natural and multispectral images, followed by geometric correction using high-reflectivity ground calibration points. Next, to address the resolution difference between the natural and multispectral images, a bicubic interpolation algorithm is used to resample the natural image, achieving pixel-level registration of the heterogeneous source image. For the blue light image... Green light image Red light image Red-bordered image and near-infrared images A stacking operation is performed to combine images from different bands along the channel dimension, forming a five-channel multispectral image. This operation provides standardized input data for subsequent multi-source data fusion and feature extraction, and divides it into training and validation sets in a 9:1 ratio.

[0051] S2. Input the historical remote sensing image data and train the multi-source data progressive fusion module to process the preprocessed natural images. and multispectral images Multi-scale features are sequentially spliced ​​together and fused with corresponding hierarchical features. A weighted strategy is used to enhance the response in changing regions, outputting a deep fused feature at a unified scale. ;

[0052] Optionally, such as Figure 3 As shown, the multi-source data progressive fusion module gradually fuses perceptual information from natural images and multispectral images (at the same acquisition time point) at different levels through multi-level feature extraction and high-order feature enhancement techniques. The processing includes:

[0053] natural images In the multi-level feature extraction submodule, the input first passes through a 3D convolutional layer with a kernel of 3×3×3, and then passes through three feature extraction blocks in sequence, gradually extracting the first-level features, second-level features, and third-level features of the natural image from the shallow layer to the deep layer.

[0054] Multispectral images at the same acquisition time point The same operation was performed to obtain the first-level features, second-level features, and third-level features of the multispectral image;

[0055] Next, a progressive fusion of multi-source features is performed according to an adaptive strategy, including:

[0056] First, the first-level, second-level, and third-level features of the natural image and the multispectral image are concatenated to form first-level shallow fusion features, second-level shallow fusion features, and third-level shallow fusion features;

[0057] Next, the first-level shallow fusion feature is used as a low-order feature of the shallow layer, and the second-level shallow fusion feature is used as a high-order feature of the deep layer. They are then combined to perform high-order feature enhancement to obtain hidden shallow fusion features. The second-level shallow fusion feature is used as a low-order feature of the shallow layer, and the third-level shallow fusion feature is used as a high-order feature of the deep layer. They are then combined to perform high-order feature enhancement to obtain hidden deep fusion features.

[0058] Finally, the hidden shallow-layer fusion features are used as low-order features of the shallow layer, and the hidden deep-layer fusion features are used as high-order features of the deep layer. These are then combined to perform high-order feature enhancement, ultimately yielding the deep fusion features. (t ranges from 0 to n).

[0059] Optionally, such as Figure 4As shown, the feature extraction block, based on multi-scale convolutional fusion, introduces channel attention and spatial attention mechanisms. Channel attention focuses on the dependencies between channels, enhancing the responsiveness to key channel features by weighting the importance of each channel. Spatial attention focuses on the spatial distribution of the feature map, enhancing the perception of lesion areas and abnormal physiological locations. The two work synergistically to enhance information through channel selection and spatial focusing, thereby improving the expression quality of feature extraction. The processing includes:

[0060] The input shallow layer of the previous level features (such as...) Figure 3 The training process (using natural image primary features) first involves three branches: each branch uses convolutional layers of different sizes, each convolutional layer is followed by a ReLU activation function, and then undergoes layer normalization to stabilize the training process. These three branches generate three feature maps, which are then used... , , express;

[0061] Then, the , , Adaptive fusion is performed through a channel attention mechanism to obtain fused features. By performing adaptive fusion through a spatial attention mechanism, fused features are obtained. The fusion feature and fusion features Pixel-wise multiplication is performed, and features are integrated through a 1×1 convolution operation to output the next level of deep features (such as...). Figure 3 (Natural image secondary features)

[0062] The channel attention mechanism includes:

[0063] The , , Global information of the feature map is obtained through global average pooling and global max pooling. The results of global average pooling and global max pooling are then fed into a multilayer perceptron and summed. Finally, the weight of each channel is calculated using the sigmoid activation function. , The calculated weights Corresponding Multiply to obtain the weighted features Finally, all weighted features , , Add them together to obtain the final fusion feature. The formula is as follows:

[0064]

[0065]

[0066]

[0067] in represents the Sigmoid activation function, MLP represents the multilayer perceptron, GAP represents global average pooling, and GMP represents global max pooling.

[0068] The spatial attention mechanism includes:

[0069] right Perform average pooling and max pooling along the channel dimension respectively to obtain two spatial feature maps. and This reflects the average and maximum responses at each location. These two feature maps are then concatenated along the channel dimension, and a 7×7 convolutional layer is used to extract fused features. The spatial attention weights are then obtained after passing through a sigmoid activation function. Then spatial attention weights With features Multiply to obtain the weighted features Finally, all weighted features , , Adding them together yields the final fusion feature. The formula is as follows:

[0070]

[0071]

[0072]

[0073] in This represents the Sigmoid function, where C represents the number of channels. This represents averaging the feature map along the channel dimension. This represents performing the maximum operation on the feature map along the channel dimension. This represents a two-dimensional convolution operation with a kernel size of 7×7.

[0074] Optionally, such as Figure 5 As shown, the higher-order feature enhancement achieves progressive fusion of multi-level features through spatial scale alignment and dual-path convolution extraction. Simultaneously, it introduces average pooling and covariance pooling to jointly model global and local statistical information, enhancing the sensitivity of the fused features to spatial variations and structural differences, and improving the stability and discriminative ability of feature representation. The processing includes:

[0075] Deep high-order features (such as) Figure 3 Secondary shallow features), through upsampling operations, are recovered to the same level as the shallow low-order features (such as... Figure 3 The first-level shallow features in the model have the same spatial resolution so that they can be fused at the same scale;

[0076] Next, the high-order and low-order features are respectively entered into a dual-parallel feature extraction operation: one path directly extracts spatial features through 3×3 convolution, while the other path generates weights through 3×3 convolution and the Sigmoid activation function for spatial weighted modulation.

[0077] Next, the convolutional output of each path is multiplied pixel by pixel with the corresponding weight to achieve local weighted fusion. The results of the two paths are then added together to generate fused features.

[0078] Then, the fused features are fed into a feature recalibration module. This module first adjusts the channel dimensions using a 1×1 convolution to extract a compact global feature map. Next, the global feature map is fed into average pooling and max pooling branches respectively: the average pooling branch captures global contextual information, while the max pooling branch highlights salient region or edge features. After the two types of information are added, a 3-convolutional layer is used to fuse the local spatial structure, ultimately outputting the higher-order enhanced features (such as...). Figure 3 (Hidden shallow fusion features in the middle).

[0079] S3, the above Input and train a frequency-domain decoupling-based change detection module, Pairing the data in chronological order, Fourier transforms are performed on each pair, followed by Gaussian high-pass and low-pass filtering to obtain high-frequency and low-frequency decoupling features. These features are then inversely transformed and fused with the time-domain feature difference results to output a probability map of changes between adjacent time points. , , ;

[0080] Optionally, such as Figure 6 As shown, the frequency-domain decoupling-based change detection module separates high-frequency and low-frequency change components through Fourier transform and filters. By analyzing the differences in time, high-frequency, and low-frequency information of dual-phase images, it fully explores the subtle change signals caused by diseases in wheat during the incubation period. The processing includes:

[0081] Deeply fused features paired in chronological order (e.g.) and , and

[0082] (etc.), sequentially passing through the same shared frequency domain decoupled encoder, and then through transposed convolutional decoding with a kernel of 3, to output the probability map of changes between adjacent time steps. , , The change probability map is an image of the same size as the input image, where each pixel takes a value in the range [0, 1], representing the probability that the position of each pixel changes between adjacent time points in the image;

[0083] Among them, for adjacent deep fusion features and The frequency domain decoupling encoder first maps the components to the frequency domain using Fourier transform, and then decomposes them into high-frequency decoupling features using Gaussian high-pass filtering and Gaussian low-pass filtering. , and low-frequency decoupling characteristics , ;

[0084] Then, the high-frequency characteristic changes are calculated. and low-frequency characteristic changes ;

[0085] Then, the high-frequency feature changes were analyzed. and low-frequency characteristic changes After being transformed back to the time domain by inverse Fourier transform, the high-frequency and low-frequency variation features are obtained by sequentially passing the two-dimensional convolution with a kernel of 3, the ReLU activation function, and normalization.

[0086] Deeply fused features that have not undergone frequency domain decoupling and The temporal variation features are obtained by sequentially passing the data through a 2D convolution with a kernel of 3, the ReLU activation function, and normalization, followed by element-wise subtraction.

[0087] The high-frequency variation features, low-frequency variation features, and time-domain variation features are added and fused element by element, and then decoded using a transposed convolution with a kernel of 3 to output a variation probability map. .

[0088] S4. Using a sliding window of size d in the probability map of change. , , Multiple training samples are constructed sequentially by sliding the window with a fixed length (for example, setting the sliding window size to 6, and the first training sample being a probability map). , , , , , Then, the sliding window is moved forward one step, and the second set of training samples is the probability map of change. , , , , , Repeat the above steps; the last set of samples is a probability plot. , , , , , The last change probability map in each training sample is encoded by a visual encoder to obtain the change features as noise-added data. The remaining change probability maps in each training sample, along with the corresponding historical meteorological data and prior data of diffusion mode, are fused as a condition vector and input into the conditional diffusion prediction model for model training.

[0089] Optionally, such as Figure 7 As shown, the visual encoder extracts features from the probability map of change through local convolution and residual connections, and introduces a query-key-value mechanism to achieve dynamic weighted modeling, effectively enhancing the representation ability of key regions and patterns in the probability map of change, and improving the modeling accuracy of disease temporal evolution features. The processing includes:

[0090] The probability diagram of the input change , , , First, flatten them into one-dimensional vectors. , , , Next, feature extraction is performed four times, and then the feature changes are obtained through a multilayer perceptron. , , , ;

[0091] The feature extraction operation includes:

[0092] The one-dimensional vector , , , First, it is normalized, then subjected to a one-dimensional convolution with a kernel of 3, and then connected via residuals. , , , The sum is used as the input to the second residual network;

[0093] The second residual network generates queries, keys, and values ​​from the input through a convolution operation with a kernel of 1, weights them through an attention mechanism, and then adds them to the input through a residual connection.

[0094] Optionally, such as Figure 8 As shown, the training of the conditional diffusion prediction model is divided into a forward noise addition process and a noise prediction learning process. It simulates the disease change process, makes predictions based on multi-temporal change probability map sequences, and encodes meteorological data and prior knowledge of diffusion modes into constraint vectors to improve the model prediction accuracy and quality.

[0095] The forward noise addition process includes:

[0096] First, perform the first set of training samples, and then generate a probability plot. , , , , , Forward noise addition process:

[0097] Probability of change diagram Change features are obtained by encoding using a visual encoder. , Considered as an initial feature during forward noise addition. ;

[0098] Through a Markov process with T time steps, the initial features are... Add Gaussian noise until it is converted into a pure noise distribution. This yields a series of noisy data. ,in, , , The labels used for training the noise prediction are used to calculate the loss with the noise prediction values, and the weights of the conditional diffusion prediction model are adjusted accordingly, as expressed by the formula:

[0099]

[0100] in The cumulative noise attenuation factor is defined as follows: ; The single-step noise attenuation factor is defined as follows: ; Indicates noise scheduling parameters; Indicates Gaussian noise; Represents a normal distribution; Represents the identity matrix;

[0101] For each time step t, The noise prediction learning process, through the conditional diffusion prediction model, extracts and fuses feature information from noisy data and conditional vectors, learns the mapping relationship from feature information to noise distribution, and outputs the predicted value of the noise added at the corresponding time step, including:

[0102] The corresponding meteorological data and prior knowledge of diffusion mode are processed separately with time step t. The meteorological data is processed by conditional encoding consisting of one-dimensional 3×3 convolution and fully connected layers. The prior knowledge of diffusion mode is processed by one-hot encoding and fully connected layers. The time step t is processed by sinusoidal embedding and fully connected layers. Then the three are concatenated to obtain the constraint vector.

[0103] Transformation probability diagram , , , , After concatenation along the channel dimension, the historical vector is generated sequentially through two-dimensional convolution, flattening into a one-dimensional vector, and fully connected operations;

[0104] The constraint vector and the historical vector are concatenated to form the condition vector. ;

[0105] Add noise to the data and condition vector Simultaneously, the conditional diffusion prediction model is input, and the output is the prediction noise. ;

[0106] Then, the predicted denoised data is calculated using the following denoising formula. :

[0107]

[0108] Calculate and predict denoised data With tags The mean squared error loss between them is optimized by backpropagation to improve the model weights;

[0109] After repeating the above steps to train all groups of training samples, the training is complete.

[0110] S5. Collect remote sensing image data of the wheat to be predicted for d-1 days (e.g., collect data for 5 days if d=6) and the corresponding meteorological data. Input the remote sensing image data into the trained multi-source data progressive fusion module and the frequency domain decoupling-based change detection module to obtain the change probability map. , , , will the , , The corresponding meteorological data and prior data on diffusion patterns are fused together as a conditional vector, and then input along with random Gaussian noise into the trained conditional diffusion prediction model to perform a denoising operation, resulting in denoised prediction data. Then, through a visual decoder, the probability map of change for the predicted day d (e.g., if d=6, then the predicted day 6) is obtained. ;

[0111] Optionally, such as Figure 9 As shown, the denoising operation in S5 progressively predicts the probability map of change on day d through a reverse denoising mechanism, performs predictions based on historical multi-temporal probability maps, and encodes meteorological data and prior knowledge of diffusion patterns into constraint vectors to improve the model's prediction accuracy and quality. Specifically, this includes:

[0112] The change probability diagram , , After being stitched along the channels, historical vectors are generated sequentially through two-dimensional convolution, flattening, and fully connected operations;

[0113] The corresponding meteorological data, diffusion mode prior data, and time steps are processed separately and then concatenated to obtain the constraint vector;

[0114] The historical vector and the constraint vector are concatenated to form the condition vector.

[0115] The conditional vector and random Gaussian noise are input together into the trained conditional diffusion prediction model, using random Gaussian noise as the initial... Perform the inverse denoising process to obtain the predicted noise. And calculated using the denoising formula ,Will The predicted denoised data is obtained by repeating the process T times as input for the next iteration. ;

[0116] The result is then processed by a visual decoder to obtain the predicted probability map of change on day d. ,like Figure 10 As shown, it specifically includes:

[0117] The input prediction denoising data After being reshaped into a two-dimensional vector, the data is decoded through six stacked feature recovery operations. By introducing a query-key-value structure and residual connections, the predicted denoised data is gradually recovered. The spatial structure information effectively enhances the ability to reconstruct fine-grained disease changes, enabling more accurate prediction of change probabilities. Then, after a two-dimensional convolution with a kernel of 3, the predicted change probability map for day d is obtained. ;

[0118] The feature recovery operation includes:

[0119] The two-dimensional vector is first subjected to a transpose convolution with a kernel of 4, and then the query, key, and value are generated by convolution operations with a kernel of 1×1. The values ​​are then weighted by an attention mechanism and finally added to the two-dimensional vector through a residual connection.

[0120] Optionally, such as Figure 11 As shown, the conditional diffusion prediction model, by integrating multiple time-series statistical features, comprehensively captures the trends, critical points, and abnormal fluctuations in the disease evolution process, and introduces conditional vectors to participate in attention calculation, enabling climate and prior knowledge factors to dynamically adjust the time-series feature recovery process. The processing includes:

[0121] The input noisy data and the constraint vector Encoding operations are performed using a stacked encoding structure, including:

[0122] First, a one-dimensional convolution with a kernel of 3 is applied, followed by average pooling, minimum pooling, and maximum pooling, and then concatenated. Next, key and value are generated through convolution operations, and the constraint vector is then... The query is generated by a one-dimensional convolution operation with a kernel of 1, weighted by an attention mechanism, and finally downsampled.

[0123] Then, a decoding operation is performed using a four-stacked decoding structure to output the prediction noise. Then, the predicted denoised data is calculated using the denoising formula. ;

[0124] The decoding operation includes:

[0125] The output of the encoded structure is first subjected to a one-dimensional convolution with a kernel of 3, then concatenated after average pooling, minimum pooling, and maximum pooling, and finally generated as keys and values ​​through convolution operations with a kernel of 1. The constraint vector is then... The query is generated by a one-dimensional convolution operation with a kernel of 1, weighted by an attention mechanism, and then upsampled.

[0126] S6. Based on the multispectral image in the wheat remote sensing image data to be predicted on day d-1 (e.g., day 5 if d=6), calculate the NDVI mask, retaining only the vegetation area, and then... The NDVI mask is multiplied pixel by pixel, and a lesion mask is generated with a preset threshold. The proportion of lesion pixels to vegetation pixels is counted as the prediction result of the severity and distribution range of wheat stripe rust on day d.

[0127] The NDVI mask calculation formula of this invention is as follows:

[0128]

[0129] NIR represents near-infrared reflectance, and RED represents red reflectance; both can be obtained from multispectral images. Extract the corresponding values. The NDVI range is between -1 and 1. An NDVI > 0.2 indicates that the pixel is vegetation. The NDVI mask value is 1 in vegetation areas and 0 in non-vegetation areas.

[0130] Next, the NDVI mask and the probability of change map are compared. Pixel-by-pixel multiplication is performed, retaining only changes in vegetation areas, followed by statistical analysis to predict wheat stripe rust disease, with a preset threshold. Set to 0.3, and compare all pixel values ​​with the threshold. Comparison, when above the threshold If a pixel is considered a lesion, the value of a suspected lesion pixel is 1, and the value of other areas is 0. A lesion mask map is generated. At the same time, the proportion of lesions is calculated, that is, the total number of lesion pixels is divided by the total number of pixels in the vegetation area to predict the severity and extent of the disease.

[0131] like Figure 12 As shown, this embodiment of the invention also provides an early disease prediction system for wheat stripe rust driven by remote sensing change information, the system comprising:

[0132] The acquisition and preprocessing module 1210 is used to acquire and preprocess historical remote sensing image data of wheat, corresponding historical meteorological data, and prior data of diffusion modes. The historical remote sensing image data includes historical natural image data and historical multispectral image data.

[0133] The first training module 1220 is used to input the historical remote sensing image data and train the multi-source data progressive fusion module to process the preprocessed natural images. and multispectral images Multi-scale features are sequentially spliced ​​together and fused with corresponding hierarchical features. A weighted strategy is used to enhance the response in changing regions, outputting a deep fused feature at a unified scale. ;

[0134] The second training module 1230 is used to train the... Input and train a frequency-domain decoupling-based change detection module, Pairing the data in chronological order, Fourier transforms are performed on each pair, followed by Gaussian high-pass and low-pass filtering to obtain high-frequency and low-frequency decoupling features. These features are then inversely transformed and fused with the time-domain feature difference results to output a probability map of changes between adjacent time points. , , ;

[0135] Module 1240 is used to construct the probability map of change using a sliding window of size d. , , Multiple training samples are constructed sequentially by sliding along a fixed length. The change feature obtained by encoding the last change probability map in each training sample through a visual encoder is used as noise-added data. The remaining change probability maps in each training sample, along with the corresponding historical meteorological data and prior data of diffusion mode, are fused as a conditional vector and input into the conditional diffusion prediction model for model training.

[0136] Prediction module 1250 is used to collect remote sensing image data of wheat to be predicted for day d-1 and corresponding meteorological data. The remote sensing image data is input into the trained multi-source data progressive fusion module and the frequency domain decoupling-based change detection module to obtain the change probability map. , , , will the , , The corresponding meteorological data and prior data on diffusion patterns are fused together as a conditional vector, and then input along with random Gaussian noise into the trained conditional diffusion prediction model to perform a denoising operation, resulting in denoised prediction data. Then, through a visual decoder, the predicted probability map of change on day d is obtained. ;

[0137] Calculation module 1260 is used to calculate an NDVI mask based on the multispectral image in the wheat remote sensing image data to be predicted on day d-1, retaining only the vegetation area. The NDVI mask is multiplied pixel by pixel, and a lesion mask is generated with a preset threshold. The proportion of lesion pixels to vegetation pixels is counted as the prediction result of the severity and distribution range of wheat stripe rust on day d.

[0138] The wheat stripe rust early disease prediction system driven by remote sensing change information provided in this embodiment of the invention has a functional structure that corresponds to the wheat stripe rust early disease prediction method driven by remote sensing change information provided in this embodiment of the invention, and will not be described again here.

[0139] Figure 13This is a schematic diagram of the structure of an electronic device 1300 provided in an embodiment of the present invention. The electronic device 1300 may vary considerably due to different configurations or performance. It may include one or more central processing units (CPUs) 1301 and one or more memories 1302. The memory 1302 stores at least one instruction, which is loaded and executed by the processor 1301 to implement the steps of the above-mentioned early disease prediction method driven by remote sensing change information of wheat stripe rust.

[0140] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions that can be executed by a processor in a terminal to complete the aforementioned wheat stripe rust remote sensing change information-driven early disease prediction method. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device.

[0141] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0142] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for early disease prediction of wheat stripe rust by remote sensing change information driving, characterized in that, The method includes: S1. Collect and preprocess historical remote sensing image data of wheat, corresponding historical meteorological data, and prior knowledge of diffusion patterns. The historical remote sensing image data includes historical natural image data and historical multispectral image data. The prior knowledge of diffusion patterns includes disease transmission type, transmission stage, transmission direction, and diffusion intensity level. S2, input the historical remote sensing image data and train the multi-source data progressive fusion module, input the multi-scale features of the preprocessed natural image and multispectral image , sequentially splice the corresponding hierarchical features and fuse, adopt a weighted strategy to enhance the response of the change area, and output the deep fusion features of the unified scale ; S3, the above Input and train a frequency-domain decoupling-based change detection module, Pairing the data in chronological order, Fourier transforms are performed on each pair, followed by Gaussian high-pass and low-pass filtering to obtain high-frequency and low-frequency decoupling features. These features are then inversely transformed and fused with the time-domain feature difference results to output a probability map of changes between adjacent time points. , , ; S4. Using a sliding window of size d in the probability map of change. , , Multiple training samples are constructed sequentially by sliding along a fixed length. The change feature obtained by encoding the last change probability map in each training sample through a visual encoder is used as noise-added data. The remaining change probability maps in each training sample, along with the corresponding historical meteorological data and prior data of diffusion mode, are fused as a conditional vector and input into the conditional diffusion prediction model for model training. S5. Collect remote sensing image data of wheat to be predicted for day d-1 and corresponding meteorological data. Input the remote sensing image data into the trained multi-source data progressive fusion module and the frequency domain decoupling-based change detection module to obtain the change probability map. , , , will the , , The corresponding meteorological data and prior data on diffusion patterns are fused together as a conditional vector, and then input along with random Gaussian noise into the trained conditional diffusion prediction model to perform a denoising operation, resulting in denoised prediction data. Then, through a visual decoder, the predicted probability map of change on day d is obtained. ; S6. Based on the multispectral image of the wheat remote sensing image data to be predicted on day d-1, calculate the NDVI mask, retaining only the vegetation area, and then... The NDVI mask is multiplied pixel by pixel, and a lesion mask is generated with a preset threshold. The proportion of lesion pixels to vegetation pixels is counted as the prediction result of the severity and distribution range of wheat stripe rust on day d.

2. The method according to claim 1, characterized in that, The multi-source data progressive fusion module gradually fuses perceptual information from natural images and multispectral images at different levels through multi-level feature extraction and high-order feature enhancement techniques. The processing includes: natural images In the multi-level feature extraction submodule, the input first passes through a 3D convolutional layer with a kernel of 3×3×3, and then passes through three feature extraction blocks in sequence, gradually extracting the first-level features, second-level features, and third-level features of the natural image from the shallow layer to the deep layer. Multispectral images at the same acquisition time point The same operation was performed to obtain the first-level features, second-level features, and third-level features of the multispectral image; Next, a progressive fusion of multi-source features is performed according to an adaptive strategy, including: First, the first-level, second-level, and third-level features of the natural image and the multispectral image are concatenated to form first-level shallow fusion features, second-level shallow fusion features, and third-level shallow fusion features; Next, the first-level shallow fusion feature is used as a low-order feature of the shallow layer, and the second-level shallow fusion feature is used as a high-order feature of the deep layer. They are then combined to perform high-order feature enhancement to obtain hidden shallow fusion features. The second-level shallow fusion feature is used as a low-order feature of the shallow layer, and the third-level shallow fusion feature is used as a high-order feature of the deep layer. They are then combined to perform high-order feature enhancement to obtain hidden deep fusion features. Finally, the hidden shallow fusion features are used as low-order features of the shallow layer, and the hidden deep fusion features are used as high-order features of the deep layer. These are then combined to perform high-order feature enhancement, ultimately yielding the deep fusion features. .

3. The method according to claim 2, characterized in that, The feature extraction block, based on multi-scale convolutional fusion, introduces channel attention and spatial attention mechanisms. Channel attention focuses on the dependencies between channels, enhancing the responsiveness to key channel features by weighting the importance of each channel. Spatial attention focuses on the spatial distribution of the feature map, enhancing the perception of lesion areas and abnormal physiological locations. The synergistic effect of these two mechanisms enhances the information of channel selection and spatial focusing, improving the expression quality of feature extraction. The processing includes: The input shallow layer's previous-level features are first processed through three branches: each branch uses convolutional layers of different sizes, each convolutional layer is followed by a ReLU activation function, and then layer normalization is applied to stabilize the training process. These three branches generate three feature maps, which are then processed using... , , express; Then, the , , Adaptive fusion is performed through a channel attention mechanism to obtain fused features. Adaptive fusion is performed through spatial attention mechanism to obtain fused features. The fusion feature and fusion features Pixel-wise multiplication is performed, and features are integrated through a 1×1 convolution operation to output the next level of deep features; The channel attention mechanism includes: The , , Global information of the feature map is obtained through global average pooling and global max pooling. The results of global average pooling and global max pooling are then fed into a multilayer perceptron and summed. Finally, the weight of each channel is calculated using the sigmoid activation function. , The calculated weights Corresponding Multiply to obtain the weighted features Finally, all weighted features , , Add them together to obtain the final fusion feature. The formula is as follows: ; ; ; in represents the Sigmoid activation function, MLP represents the multilayer perceptron, GAP represents global average pooling, and GMP represents global max pooling. The spatial attention mechanism includes: right Perform average pooling and max pooling along the channel dimension respectively to obtain two spatial feature maps. and This reflects the average and maximum responses at each location. These two feature maps are then concatenated along the channel dimension, and a 7×7 convolutional layer is used to extract fused features. The spatial attention weights are then obtained after passing through a sigmoid activation function. Then spatial attention weights With features Multiply to obtain the weighted features Finally, all weighted features , , Adding them together yields the final fusion feature. The formula is as follows: ; ; ; ; in This represents the Sigmoid function, where C represents the number of channels. This represents averaging the feature map along the channel dimension. This represents performing the maximum operation on the feature map along the channel dimension. This represents a two-dimensional convolution operation with a kernel size of 7×7.

4. The method according to claim 2, characterized in that, The higher-order feature enhancement achieves progressive fusion of multi-level features through spatial scale alignment and dual-path convolution extraction. Simultaneously, it introduces average pooling and covariance pooling to jointly model global and local statistical information, enhancing the sensitivity of the fused features to spatial variations and structural differences, and improving the stability and discriminative ability of feature representation. The processing includes: The high-order features of the deep layers are restored to the same spatial resolution as the low-order features of the shallow layers through upsampling operations, so that they can be fused at the same scale; Next, the high-order and low-order features are respectively entered into a dual-parallel feature extraction operation: one path directly extracts spatial features through 3×3 convolution, while the other path generates weights through 3×3 convolution and the Sigmoid activation function for spatial weighted modulation. Next, the convolutional output of each path is multiplied pixel by pixel with the corresponding weight to achieve local weighted fusion. The results of the two paths are then added together to generate fused features. Then, the fused features are fed into the feature recalibration module. The feature recalibration module first adjusts the channel dimension through a 1×1 convolution to extract a compact global feature map. Next, the global feature map is fed into the average pooling and max pooling branches respectively: the average pooling branch captures global context information, while the max pooling branch highlights salient regional feature information or edge feature information. After the two types of information are added together, they are then fused through a 3×3 convolutional layer to fuse the local spatial structure, and finally output the high-order enhanced features.

5. The method according to claim 1, characterized in that, The frequency-domain decoupling-based change detection module separates high-frequency and low-frequency change components through Fourier transform and filters. By analyzing the differences in time, high-frequency, and low-frequency information of the dual-temporal images, it fully explores the subtle change signals caused by diseases in wheat during the latent period. The processing includes: The deep fusion features, paired in time sequence, are sequentially passed through the same shared frequency domain decoupled encoder, and then through a transposed convolutional decoder with a kernel of 3, outputting a probability map of changes between adjacent time steps. , , The change probability map is an image of the same size as the input image, where each pixel takes a value in the range [0, 1], representing the probability that the position of each pixel changes between adjacent time points in the image; Among them, for adjacent deep fusion features and The frequency domain decoupling encoder first maps the components to the frequency domain using Fourier transform, and then decomposes them into high-frequency decoupling features using Gaussian high-pass filtering and Gaussian low-pass filtering. , and low-frequency decoupling characteristics , ; Then, the high-frequency characteristic changes are calculated. and low-frequency characteristic changes ; Then, the high-frequency feature changes were analyzed. and low-frequency characteristic changes After being transformed back to the time domain by inverse Fourier transform, the high-frequency and low-frequency variation features are obtained by sequentially passing the two-dimensional convolution with a kernel of 3, the ReLU activation function, and normalization. Deeply fused features that have not undergone frequency domain decoupling and The temporal variation features are obtained by sequentially passing the data through a 2D convolution with a kernel of 3, the ReLU activation function, and normalization, followed by element-wise subtraction. The high-frequency variation features, low-frequency variation features, and time-domain variation features are added and fused element by element, and then decoded using a transposed convolution with a kernel of 3 to output a variation probability map. .

6. The method according to claim 1, characterized in that, The visual encoder extracts features from the probability map of change through local convolution and residual connections, and introduces a query-key-value mechanism to achieve dynamic weighted modeling. This effectively enhances the representation ability of key regions and patterns in the probability map of change, and improves the modeling accuracy of disease temporal evolution features. The processing includes: The probability diagram of the input change , , , First, flatten them into one-dimensional vectors. , , , Next, feature extraction is performed four times, and then the feature changes are obtained through a multilayer perceptron. , , , ; The feature extraction operation includes: The one-dimensional vector , , , First, it is normalized, then subjected to a one-dimensional convolution with a kernel of 3, and then connected via residuals. , , , The sum is used as the input to the second residual network; The second residual network generates queries, keys, and values ​​from the input through a convolution operation with a kernel of 1, weights them through an attention mechanism, and then adds them to the input through a residual connection.

7. The method according to claim 1, characterized in that, The training of the conditional diffusion prediction model is divided into a forward noise addition process and a noise prediction learning process. It simulates the disease change process, makes predictions based on multi-temporal change probability map sequences, and encodes meteorological data and prior knowledge of diffusion modes into constraint vectors to improve the model's prediction accuracy and quality. The forward noise addition process includes: First, perform the first set of training samples, and then generate a probability plot. , , , , , Forward noise addition process: Probability of change diagram Change features are obtained by encoding using a visual encoder. , Considered as an initial feature during forward noise addition. ; Through a Markov process with T time steps, the initial features are... Add Gaussian noise until it is converted into a pure noise distribution. This yields a series of noisy data. ,in, , , The labels used for training the noise prediction are used to calculate the loss with the noise prediction values, and the weights of the conditional diffusion prediction model are adjusted accordingly, as expressed by the formula: ; in The cumulative noise attenuation factor is defined as follows: ; The single-step noise attenuation factor is defined as follows: ; Indicates noise scheduling parameters; Indicates Gaussian noise; Represents a normal distribution; Represents the identity matrix; For each time step t, T, the noise prediction learning process, extracts and fuses feature information from the noisy data and conditional vectors using the conditional diffusion prediction model, learns the mapping relationship from feature information to noise distribution, and outputs the predicted value of the noise added at the corresponding time step, including: The corresponding meteorological data and prior knowledge of diffusion mode are processed separately with time step t. The meteorological data is processed by conditional encoding consisting of one-dimensional 3×3 convolution and fully connected layers. The prior knowledge of diffusion mode is processed by one-hot encoding and fully connected layers. The time step t is processed by sinusoidal embedding and fully connected layers. Then the three are concatenated to obtain the constraint vector. Transformation probability diagram , , , , After concatenation along the channel dimension, the historical vector is generated sequentially through two-dimensional convolution, flattening into a one-dimensional vector, and fully connected operations; The constraint vector and the historical vector are concatenated to form the condition vector. ; Add noise to the data and condition vector Simultaneously, the conditional diffusion prediction model is input, and the output is the prediction noise. ; Then, the predicted denoised data is calculated using the following denoising formula. : ; Calculate and predict denoised data With tags The mean squared error loss between them is optimized by backpropagation to improve the model weights; Repeat the forward noise addition process and noise prediction learning process to train all groups of training samples, and the training is complete.

8. The method according to claim 7, characterized in that, The denoising operation in S5 involves progressively predicting the probability map of change on day d using a reverse denoising mechanism. Prediction is based on historical multi-temporal probability maps, and prior knowledge of meteorological data and diffusion patterns is encoded into constraint vectors to improve model prediction accuracy and quality. Specifically, this includes: The change probability diagram , , After being stitched along the channels, historical vectors are generated sequentially through two-dimensional convolution, flattening, and fully connected operations; The corresponding meteorological data, diffusion mode prior data, and time steps are processed separately and then concatenated to obtain the constraint vector; The historical vector and the constraint vector are concatenated to form the condition vector. The conditional vector and random Gaussian noise are input together into the trained conditional diffusion prediction model, using random Gaussian noise as the initial... Perform the inverse denoising process to obtain the predicted noise. And calculated using the denoising formula ,Will The predicted denoised data is obtained by repeating the process T times as input for the next iteration. ; The result is then processed by a visual decoder to obtain the predicted probability map of change on day d. Specifically, it includes: The input prediction denoising data After being reshaped into a two-dimensional vector, the data is decoded through six stacked feature recovery operations. By introducing a query-key-value structure and residual connections, the predicted denoised data is gradually recovered. The spatial structure information effectively enhances the ability to reconstruct fine-grained disease changes, enabling more accurate prediction of change probabilities. Then, after a two-dimensional convolution with a kernel of 3, the predicted change probability map for day d is obtained. ; The feature recovery operation includes: The two-dimensional vector is first subjected to a transpose convolution with a kernel of 4, and then the query, key, and value are generated by convolution operations with a kernel of 1×1. The values ​​are then weighted by an attention mechanism and finally added to the two-dimensional vector through a residual connection.

9. The method according to claim 7, characterized in that, The conditional diffusion prediction model, by integrating multiple time-series statistical features, comprehensively captures trends, critical points, and abnormal fluctuations in the disease evolution process. It also introduces conditional vectors to participate in attention calculations, enabling climate and prior knowledge factors to dynamically adjust the time-series feature recovery process. The processing includes: The input noisy data and the constraint vector Encoding operations are performed using a stacked encoding structure, including: First, a one-dimensional convolution with a kernel of 3 is applied, followed by average pooling, minimum pooling, and maximum pooling, and then concatenated. Next, key and value are generated through convolution operations, and the constraint vector is then... The query is generated by a one-dimensional convolution operation with a kernel of 1, weighted by an attention mechanism, and finally downsampled. Then, a decoding operation is performed using a four-stacked decoding structure to output the prediction noise. Then, the predicted denoised data is calculated using the denoising formula. ; The decoding operation includes: The output of the encoded structure is first subjected to a one-dimensional convolution with a kernel of 3, then concatenated after average pooling, minimum pooling, and maximum pooling, and finally generated as keys and values ​​through convolution operations with a kernel of 1. The constraint vector is then... The query is generated by a one-dimensional convolution operation with a kernel of 1, weighted by an attention mechanism, and then upsampled.

10. A wheat stripe rust early disease prediction system driven by remote sensing change information, characterized in that, The system includes: The acquisition and preprocessing module is used to acquire and preprocess historical remote sensing image data of wheat, corresponding historical meteorological data, and prior knowledge of diffusion patterns. The historical remote sensing image data includes historical natural image data and historical multispectral image data. The prior knowledge of diffusion patterns includes disease transmission type, transmission stage, transmission direction, and diffusion intensity level. The first training module is used to input the historical remote sensing image data and train the multi-source data progressive fusion module, which then processes the pre-processed natural image data. and multispectral images Multi-scale features are sequentially spliced ​​together and fused with corresponding hierarchical features. A weighted strategy is used to enhance the response in changing regions, outputting a deep fused feature at a unified scale. ; The second training module is used to train the... Input and train a frequency-domain decoupling-based change detection module, Pairing the data in chronological order, Fourier transforms are performed on each pair, followed by Gaussian high-pass and low-pass filtering to obtain high-frequency and low-frequency decoupling features. These features are then inversely transformed and fused with the time-domain feature difference results to output a probability map of changes between adjacent time points. , , ; Build a module for using a sliding window of size d in the probability map of change. , , Multiple training samples are constructed sequentially by sliding along a fixed length. The change feature obtained by encoding the last change probability map in each training sample through a visual encoder is used as noise-added data. The remaining change probability maps in each training sample, along with the corresponding historical meteorological data and prior data of diffusion mode, are fused as a conditional vector and input into the conditional diffusion prediction model for model training. The prediction module collects remote sensing image data of wheat to be predicted for d-1 days and corresponding meteorological data. It then inputs the remote sensing image data into a trained multi-source data progressive fusion module and a frequency-domain decoupling-based change detection module to obtain a change probability map. , , , will the , , The corresponding meteorological data and prior data on diffusion patterns are fused together as a conditional vector, and then input along with random Gaussian noise into the trained conditional diffusion prediction model to perform a denoising operation, resulting in denoised prediction data. Then, through a visual decoder, the predicted probability map of change on day d is obtained. ; The calculation module is used to calculate an NDVI mask based on the multispectral image in the wheat remote sensing image data to be predicted on day d-1, retaining only the vegetation area. The NDVI mask is multiplied pixel by pixel, and a lesion mask is generated with a preset threshold. The proportion of lesion pixels to vegetation pixels is counted as the prediction result of the severity and distribution range of wheat stripe rust on day d.