Wheat waterlogging accurate identification and disaster assessment method based on multi-temporal remote sensing image
By constructing a feature matrix and combining it with a convolutional neural network to process multi-temporal remote sensing images, the problems of image misalignment and feature distortion under flood disasters were solved, enabling accurate identification and disaster assessment of wheat floods and improving the accuracy and efficiency of the assessment.
Patent Information
- Application Number
- CN202511824126.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-02-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing multi-temporal remote sensing images are easily affected by disasters in flood disaster scenarios, resulting in image misalignment and distortion of feature information, making it difficult to accurately locate the affected area, and consequently making it difficult to quantify and assess the scope and extent of the disaster's impact.
By periodically acquiring remote sensing images, extracting stable feature points to construct a feature matrix, and combining feature map sampling data from before, during, and after the disaster, the feature matrix is accurately matched to pinpoint the target area and assess the degree of disaster. Convolutional neural networks are used for feature extraction and disaster assessment.
It enables precise location of disaster-stricken areas and scientific assessment of disaster severity in multi-temporal remote sensing images affected by flooding, improving the objectivity, reliability, and timeliness of disaster assessment.
Smart Images

Figure CN121564430A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, specifically to a method for accurate identification and disaster assessment of wheat flooding based on multi-temporal remote sensing images. Background Technology
[0002] With the evolution of high-resolution satellites coupled with machine learning, crop disaster diagnosis based on multi-temporal remote sensing images has become a hot topic in precision agriculture research. Early optical index thresholding methods were greatly affected by cloud and rain interference. Later, the introduction of SAR multi-polarization water body identification and phenological feature reconstruction effectively improved the accuracy of wheat waterlogging area extraction, but the quantitative analysis of disaster damage and the expression of spatial heterogeneity are still insufficient.
[0003] Publication No. CN117372893A discloses a flood disaster assessment method based on an improved remote sensing image feature matching algorithm. The method involves collecting images of the same area at different times of disaster using a drone. Images taken before the flood are designated as template images, while images taken during and after the disaster are designated as transformed images. SURF feature point detection is performed on the two types of remote sensing images. The detected feature points are described using BRIEF feature descriptors, and the robustness of the descriptors to scale and illumination is assessed after noise is added to the disaster images due to natural environmental influences. For the template and transformed images, an improved SURF algorithm is selected for feature point matching. By comparing the number of overlapping feature points, the inundation area is automatically extracted from the remote sensing images, accurately identifying changes in flooded area. Based on the matching results, flood change information is extracted to quantify the severity of the flood disaster. However, in this technique, because feature objects in the images are submerged by the flood, the number of features is reduced, making it difficult to align multi-temporal remote sensing images and thus hindering feature comparison.
[0004] Existing multi-temporal remote sensing images are susceptible to disaster interference in flood disaster scenarios, resulting in image misalignment and feature distortion, making it difficult to accurately locate the affected area and making subsequent disaster quantitative assessment difficult. Summary of the Invention
[0005] The purpose of this invention is to solve the problems mentioned in the background art above: 1. Existing multi-temporal remote sensing images are easily interfered with by the disaster itself in flood disaster scenarios, resulting in problems such as spatial misalignment and distortion of feature information in the images, making it difficult to accurately locate the disaster area; 2. Due to the inaccurate location of the disaster-stricken area, it is impossible to effectively quantify and assess key indicators such as the scope of the disaster's impact and the degree of damage. Therefore, a method for accurate identification and disaster assessment of wheat flooding based on multi-temporal remote sensing images is proposed.
[0006] A first aspect of this invention provides a method for accurate identification and disaster assessment of wheat flooding based on multi-temporal remote sensing imagery, the method comprising: A remote sensing image set is obtained by periodically acquiring remote sensing images within a region, and multiple feature points are obtained by extracting features from the remote sensing images in the remote sensing image set. A feature matrix is constructed based on each feature point, and the feature matrix is added as a label to the corresponding remote sensing image to obtain a localization image set. The first feature map, the second feature map, and the third feature map are obtained by sampling the positioning images in the positioning image set at a preset time. The target region is obtained by aligning the feature matrices in the first feature map, the second feature map, and the third feature map, and the degree of disaster is obtained by evaluating the target region.
[0007] Optionally, feature extraction is performed on the remote sensing images in the remote sensing image set to obtain multiple feature points, including: The remote sensing image is mapped to a 64-dimensional feature space to obtain a 64-dimensional feature image. The low-level features are convolved on the 64-dimensional feature image, the middle-level features are convolved on the low-level features, the high-level features are convolved on the middle-level features, and the high-level features are input into a 3×3 convolution module to obtain the original feature tensor. After performing a global average pooling operation on the original feature tensor, channel attention weights are activated. The channel attention weights are then multiplied channel by channel with the original feature tensor to obtain a channel-weighted feature map. Max pooling is performed on the channel-weighted feature map to obtain a first weighted feature map. Average pooling is performed on the channel-weighted feature map to obtain a second weighted feature map. The first weighted feature map and the second weighted feature map are concatenated and activated to obtain spatial attention weights. The spatial attention weights are multiplied element-wise with the channel-weighted feature map to obtain an enhanced feature map. Each channel of the enhanced feature map is convolved according to the spatial dimension to obtain multiple local features. Each local feature is input into a 1×1 convolution module and fused to obtain an enhanced feature tensor.
[0008] Optionally, after inputting each local feature into a 1×1 convolution module to obtain an enhanced feature tensor, feature extraction is performed on the enhanced feature tensor to obtain feature points, including: The enhanced feature tensor is input into the decoder to obtain a 65-channel feature tensor. The probability distribution is calculated for each channel in the 65-channel feature tensor. The 65-channel feature tensor consists of 64 channel pixel positions and 1 invalid channel. The probability distribution consists of multiple valid probabilities and one invalid probability, where the invalid probability is the probability of an invalid channel. Candidate probability distributions are obtained by filtering the probability distributions according to the probability threshold, and valid probability distributions are obtained by filtering the candidate probability distributions according to the invalid probabilities. The candidate points corresponding to each valid probability in the valid probability distribution are used as feature points.
[0009] Optionally, a feature matrix can be constructed based on each feature point, including: The feature information and coordinate information of each feature point are determined. A descriptive feature matrix is constructed based on the feature information, and a spatial coordinate matrix is constructed based on the coordinate information. The descriptive feature matrix and the spatial coordinate matrix are then fused to obtain the feature matrix.
[0010] Optionally, the preset time includes a first time, a second time, and a third time; the first time is less than the second time, and the second time is less than the third time.
[0011] Optionally, the target region includes a first region, a second region, and a third region; the first feature map corresponds one-to-one with the first region, the second feature map corresponds one-to-one with the second region, and the third feature map corresponds one-to-one with the third region.
[0012] Optionally, the disaster severity is assessed by evaluating the target area, and the wheat health coefficient is obtained by identifying the third area using the target model, including: The wheat color coefficient and wheat lodging coefficient are obtained by extracting features from the third region using the target model, and the wheat health coefficient is calculated based on the wheat color coefficient and wheat lodging coefficient. The target model is an improvement based on the YOLOv8 model, including: The target model is obtained by replacing the Conv modules in the backbone network with depthwise separable convolutional modules and replacing the neck structure with an improved neck structure; the YOLOv8 module includes a backbone network and a neck structure.
[0013] Optionally, the target area is assessed to determine the degree of disaster, including: Water features are extracted from the second region and the third region to obtain the second disaster area and the third disaster area. The area of the intersection region is calculated based on the second disaster area and the third disaster area. A first area is determined based on the first region. A flood assessment coefficient is calculated based on the first area, the second affected area, the third affected area, and the area of the intersection region. An assessment coefficient is calculated based on the wheat health coefficient and the flood assessment coefficient. If the evaluation coefficient falls within the first interval, the target area is determined to be severely damaged. If the evaluation coefficient falls within the second interval, the target area is determined to be moderately damaged. If the evaluation coefficient falls within the third interval, the target area is determined to be slightly damaged.
[0014] Optionally, the workflow for improving the neck structure includes: The output of the C2f module at layer 4 in the backbone network is used as the first input, the output of the C2f module at layer 6 in the backbone network is used as the second input, and the output of the SPPF module at layer 9 in the backbone network is used as the third input. The third input is upsampled and then concatenated with the second input to obtain the first tensor. The first tensor is then input into the improved C2f module to obtain the second tensor. The second tensor is upsampled and then concatenated with the first input to obtain the third tensor. The third tensor is sequentially input into the improved C2f module and the Conv module to obtain the fourth tensor. The fourth tensor is concatenated with the third tensor to obtain the fifth tensor. The fifth tensor is sequentially input into the improved C2f module and the Conv module to obtain the sixth tensor. The 6th tensor is concatenated with the third input to obtain the 7th tensor. The 3rd tensor, the 6th tensor, and the 7th tensor are used as the output for improving the neck structure.
[0015] Optionally, the working principle of the improved C2f module includes: Obtain the original feature map, input the original feature map into a 7×7 depthwise separable convolution module to obtain the first feature map, input the first feature map into a 1×1 convolution module to obtain the second feature map, and activate the first feature map after inputting it into the 1×1 convolution module to obtain the third feature map; The second feature map is multiplied element-wise with the third feature map to obtain the fourth feature map. The fourth feature map is then input sequentially into a 1×1 convolution module and a 7×7 depthwise separable convolution module to obtain the fifth feature map. The fifth feature map is then used as the output of the improved C2f module.
[0016] The beneficial effects of this invention are: This invention proposes a method for accurate identification and disaster assessment of wheat flooding based on multi-temporal remote sensing imagery. By periodically acquiring remote sensing images and extracting stable feature points to construct a feature matrix, and combining the feature map sampling data before, during, and after the disaster, a precise match is achieved with the feature matrix. This allows for accurate identification of the target area and scientific assessment of the disaster severity. This method overcomes the problem of difficulty in accurately locating the disaster area due to flooding interference in multi-temporal remote sensing images, and improves the objectivity, reliability, and timeliness of disaster assessment. Attached Figure Description
[0017] Figure 1A flowchart of a method for accurate identification and disaster assessment of wheat flooding based on multi-temporal remote sensing images provided in an embodiment of the present invention; Figure 2 The flowchart illustrates the improved neck structure workflow of the wheat flooding accurate identification and disaster assessment method based on multi-temporal remote sensing images provided in this embodiment of the invention. Detailed Implementation
[0018] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0019] This invention provides a method for accurate identification and disaster assessment of wheat flooding based on multi-temporal remote sensing imagery. See also... Figure 1 , Figure 1 A flowchart illustrating a method for accurate identification and disaster assessment of wheat flooding based on multi-temporal remote sensing imagery, provided in an embodiment of the present invention. The method includes the following steps: S101, periodically acquire remote sensing images within the region to obtain a remote sensing image set, and extract features from the remote sensing images in the remote sensing image set to obtain multiple feature points; S102, construct a feature matrix based on each feature point, and add the feature matrix as a label to the corresponding remote sensing image to obtain a localization image set; S103, sampling the positioning images in the positioning image set at a preset time to obtain the first feature map, the second feature map and the third feature map; S104. Align the target area based on the feature matrices in the first feature map, the second feature map, and the third feature map, and assess the disaster level of the target area.
[0020] The method for accurate identification and disaster assessment of wheat flooding based on multi-temporal remote sensing images provided in this invention periodically collects remote sensing images and extracts stable feature points to construct a feature matrix. By combining feature map sampling before, during and after the disaster with the feature matrix for accurate alignment, the method achieves accurate locking of the target area and scientific assessment of the disaster severity. This overcomes the problem of difficulty in accurately locating the disaster-stricken area caused by flood interference from multi-temporal remote sensing, and improves the objectivity, reliability and timeliness of disaster assessment.
[0021] In one implementation, remote sensing images are acquired periodically, which means acquiring a sequence of images of the same area at different times. For example, within a month, four remote sensing images are taken of the same area each week.
[0022] In one implementation, feature extraction is performed on remote sensing images in a remote sensing image set to obtain multiple feature points, such as: artificial feature points, topographic feature points, and stable feature points in crop areas. Among them, artificial feature points are: corners or intersections of artificial structures such as utility poles, well houses, road corners, ditch corners, and dam edges around farmland; topographic feature points are: inflection points of plot boundaries, intersections of field ridges, and abrupt changes in terrain undulation (such as the edges of terraced fields, the junction of slopes and flat land); and stable feature points in crop areas are: non-flood-sensitive structural points in wheat fields (such as abrupt changes in texture in areas with uniform growth, and intersections of fixed ridges in wheat fields).
[0023] In one implementation, the feature matrix is composed of a descriptive feature matrix and a spatial coordinate feature matrix: the sub-feature matrix stores the grayscale or texture mode information of each feature point in an "N×D" dimension (N is the number of stable feature points after filtering, and D is the feature descriptor dimension); the spatial coordinate feature matrix records the pixel coordinates of the corresponding feature points in an "N×2" dimension, providing a spatial position reference for image registration, and the row vectors of the two matrices correspond one-to-one.
[0024] In one implementation, a first feature map, a second feature map, and a third feature map are obtained by sampling the positioning images in the positioning image set at preset times. For example, one image is taken before the flood disaster (first time), one image is taken when the flood disaster occurs (second time), and one image is taken after the flood disaster occurs (third time). The target area is obtained by aligning the feature matrices in the first feature map, the second feature map, and the third feature map. The resolution of the three feature maps is scaled and the spatial position is aligned to obtain a unified area (target area). The disaster assessment of the target area is then performed.
[0025] In one implementation, existing multi-temporal remote sensing images are difficult to accurately locate disaster-stricken areas due to flooding. By extracting spatially stable and repeatable feature points across time phases to construct a structured feature matrix, a unified positioning benchmark is provided for remote sensing images at different times, enabling accurate location of disaster-stricken areas and avoiding the impact of image misalignment or feature distortion caused by disasters on the accuracy of area positioning.
[0026] In one implementation, by accurately aligning multi-temporal positioning images and focusing on the target area, scattered remote sensing image data is transformed into disaster assessment data with spatiotemporal correlation. Through collaborative analysis of feature matrices at different disaster stages, the disaster impact process and state changes in the target area can be comprehensively captured, thereby improving the objectivity and reliability of disaster severity assessment.
[0027] In one embodiment, feature extraction is performed on remote sensing images in a remote sensing image set to obtain multiple feature points, including: The remote sensing image is mapped to a 64-dimensional feature space to obtain a 64-dimensional feature image. The 64-dimensional feature image is then convolved with low-level features, and the low-level features are further convolved with mid-level features. The mid-level features are then convolved with high-level features, and the high-level features are then input into a 3×3 convolution module to obtain the original feature tensor. After performing a global average pooling operation on the original feature tensor, channel attention weights are activated. The channel attention weights are then multiplied channel by channel with the original feature tensor to obtain a channel-weighted feature map. Max pooling is performed on the channel-weighted feature map to obtain the first weighted feature map. Average pooling is performed on the channel-weighted feature map to obtain the second weighted feature map. The first weighted feature map and the second weighted feature map are concatenated and activated to obtain the spatial attention weights. The spatial attention weights are multiplied element-wise with the channel-weighted feature map to obtain the enhanced feature map. Convolution is performed on each channel of the enhanced feature map according to the spatial dimension to obtain multiple local features. The local features are input into a 1×1 convolution module and fused to obtain the enhanced feature tensor.
[0028] In one implementation, the remote sensing image is first mapped to a 64-dimensional feature space. Then, low-level, mid-level, and high-level features are extracted through convolution in stages. Finally, the original feature tensor is obtained through 3×3 convolution. Staged convolution can capture different levels of information in the remote sensing image layer by layer. Low-level convolution retains basic details such as edges and textures, mid-level convolution integrates local semantic features, and high-level convolution extracts global abstract features. This avoids the problem that a single convolution cannot take into account multi-scale information and provides a more comprehensive feature foundation for subsequent feature processing. The high-level features are further processed through 3×3 convolution, which not only achieves reasonable compression of feature dimensions, reducing the number of channels from 256 to 128, but also retains key information. The generated original feature tensor space has a dimension of 1 / 8 of the original image. While reducing the subsequent computational load, it ensures the information density of the features and lays an efficient and high-quality input foundation for feature enhancement of multi-sensory layers.
[0029] In one implementation, at the attention mechanism level, channel attention generates weights through global average pooling and activation, weighting the original feature tensor channel by channel. This highlights key channel features for image matching (such as contours of ground features and channels corresponding to stable textures) and suppresses interference from noise or redundant channels. Spatial attention generates weights through max pooling and average pooling concatenation, weighting the channel-weighted feature map element by element. This focuses on spatially discriminative regions in the image (such as building edges and vegetation boundaries), reducing the impact of background regions on feature extraction. At the feature fusion level, local features are obtained by independently convolving each channel of the enhanced feature map according to the spatial dimension. Then, cross-channel features are fused through 1×1 convolution. This preserves the local details of each channel (such as texture features of small-scale ground features) and achieves complementarity and integration of cross-channel features. The resulting enhanced feature tensor has stronger discriminative power and highlights key information, improving the accuracy and effectiveness of subsequent feature point detection. It is particularly suitable for feature instability caused by changes in ground features and radiation differences in multi-temporal remote sensing images.
[0030] In one embodiment, after inputting each local feature into a 1×1 convolution module to obtain an enhanced feature tensor, feature extraction is performed on the enhanced feature tensor to obtain feature points, including: The enhanced feature tensor is input into the decoder to obtain a 65-channel feature tensor. The probability distribution is calculated for each channel in the 65-channel feature tensor. The 65-channel feature tensor consists of 64 channel pixel positions and 1 invalid channel. The probability distribution consists of multiple valid probabilities and one invalid probability, where the invalid probability is the probability of the invalid channel. Candidate probability distributions are obtained by filtering the probability distributions based on probability thresholds, and valid probability distributions are obtained by filtering the candidate probability distributions based on invalid probabilities. Candidate points corresponding to each valid probability in the valid probability distribution are used as feature points.
[0031] In one implementation, the enhanced feature tensor is input into the decoder to obtain a 65-channel feature tensor. The spatial dimension is 1 / 8 of the original image as the input to the decoder, and it is transformed into a 65-channel feature tensor through operations such as convolution and full connection.
[0032] In one implementation, the enhanced feature tensor consists of 64 blocks (64 channels) of 8×8 pixel size. In the 65-channel feature tensor corresponding to each spatial location, the first 64 channels correspond to the 64 sub-pixel positions within the 8×8 pixel block, which means that the 8×8 pixel block is subdivided into a finer sub-pixel grid. The value of each channel represents the potential score of the sub-pixel position to become a valid feature point. The probability distribution of each position calculated by the Softmax function is essentially a normalization of the scores of the 65 channels (64 subdivided sub-pixel positions + 1 invalid channel) corresponding to each spatial location, to obtain the probability value of each subdivided sub-pixel position and the invalid channel within the 8×8 block.
[0033] In one implementation, a candidate probability distribution is obtained by filtering the probability distribution based on a probability threshold. This refers to the channel with the highest probability among the 64 sub-pixel position channels in an 8×8 pixel block (corresponding to a sub-pixel position). Its normalized probability value must be greater than 0.01. False candidate points with extremely low probabilities, possibly generated by noise or irrelevant textures, are filtered out to ensure the candidate points have a certain basis for validity. An effective probability distribution is then obtained by filtering the candidate probability distribution based on invalid probabilities. The 65th channel out of 65 channels is an invalid channel, and its probability value represents the possibility that there are no effective feature points in the 8×8 pixel block. The probability of the sub-pixel position channel with the highest probability needs to be calculated, and it must be greater than the probability of the invalid channel. This determines whether there are truly effective feature points in the block: if the highest probability of the sub-pixel position is lower than the probability of the invalid channel, it means there are no effective feature points in the block, and the sub-pixel position will be discarded even if its probability exceeds 0.01; only when the sub-pixel position probability simultaneously satisfies both 0.01 (probability threshold) and is higher than the probability of the invalid channel will the sub-pixel position be retained as a feature point.
[0034] In one embodiment, a feature matrix is constructed based on each feature point, including: The feature information and coordinate information of each feature point are determined. A descriptive feature matrix is constructed based on the feature information, and a spatial coordinate matrix is constructed based on the coordinate information. The descriptive feature matrix and the spatial coordinate matrix are then fused to obtain the feature matrix.
[0035] In one implementation, feature information includes, for example, information about utility poles, field ridges, etc.; a feature description matrix stores grayscale or texture mode description information for each feature point, used to find the point corresponding to the same physical point in another image; coordinate information includes pixel coordinate information; and a spatial coordinate matrix stores pixel coordinate information for each feature point, used to solve the spatial transformation matrix of the two images through the coordinate relationship of corresponding points to achieve positioning and alignment.
[0036] In one implementation, the descriptive feature matrix and the spatial coordinate matrix are fused into a unified structured feature matrix (feature matrix) through a one-to-one correspondence of row vectors. This means that the grayscale / texture description information of each feature point is bound and stored with the pixel coordinate information. This approach can simultaneously take into account the discriminative power of feature matching and the accuracy of spatial positioning, avoid the problem of feature point and coordinate misalignment caused by the separation of the two types of matrices, simplify the calculation process of subsequent image alignment and target area locking, and improve the efficiency of multi-temporal remote sensing image registration and the positioning accuracy of disaster-stricken areas.
[0037] In one embodiment, the preset time includes a first time, a second time, and a third time; the first time is less than the second time, and the second time is less than the third time.
[0038] In one implementation, there are three time periods in the time series: a first time period, a second time period, and a third time period. Numerically, the first time period is less than the second time period, and the second time period is less than the third time period. For example, if the first time period is March 11, the second time period is March 15, and the third time period is March 20, then the first time period is greater than the second time period, and the second time period is greater than the third time period.
[0039] In one embodiment, the target region includes a first region, a second region, and a third region; the first feature map corresponds one-to-one with the first region, the second feature map corresponds one-to-one with the second region, and the third feature map corresponds one-to-one with the third region.
[0040] In one implementation, the target area obtained after alignment is divided into a first region, a second region, and a third region according to the disaster time sequence stage. The three regions correspond one-to-one with the first feature map (before the disaster), the second feature map (during the disaster), and the third feature map (after the disaster) collected at preset times, respectively. This achieves a precise association between each time sequence feature map and the target area, providing a clear spatiotemporal matching basis for subsequent quantitative analysis of the disaster occurrence, development, and recovery process by comparing information such as vegetation growth and water accumulation in different stages of the region, and thus scientifically assessing the degree of disaster.
[0041] In one embodiment, the degree of disaster is assessed by evaluating the target area, and the wheat health coefficient is obtained by identifying a third area through the target model, including: The wheat color coefficient and wheat lodging coefficient are obtained by extracting features from the third region using the target model, and the wheat health coefficient is calculated based on the wheat color coefficient and wheat lodging coefficient. The target model is an improvement upon the YOLOv8 model, including: The target model is obtained by replacing the Conv modules in the backbone network with depthwise separable convolutional modules and replacing the neck structure with an improved neck structure; the YOLOv8 module includes the backbone network and the neck structure.
[0042] In one implementation, RGB or HSV color space data of the wheat planting area are extracted from the aligned target area (first / second / third area), and non-wheat field background (such as soil, waterlogged areas) is removed. Based on the proportion of green component in HSV space (number of pixels in the 50-70° range of H channel / total number of wheat pixels, denoted as G_ratio) and the mean color saturation (mean of S channel, denoted as S_mean), the higher the proportion of green component and the closer the saturation is to 0.6 (typical value of healthy wheat), the better the health status. Color health normalization: a healthy wheat color benchmark is set (G_ratio0=0.8, S_mean0=0.6), and the wheat color coefficient C is calculated: C=[0.6×(G_ratio / G_ratio0)+0.4×(1-|S_mean-S_mean0| / S_mean0)], C∈[0,1], C is the wheat color coefficient, the closer the value is to 1, the healthier the color dimension, and 0.6 and 0.4 are weights.
[0043] In one implementation, the wheat planting area image is converted to grayscale, and texture features are extracted using local binary mode. The focus is on analyzing texture uniformity (denoted as T_unif) and texture contrast (denoted as T_cont). Texture features of lodged wheat: after lodging, the plants are laid flat, with high texture uniformity and low contrast; healthy upright wheat has high texture non-uniformity and high contrast. Lodging judgment criteria are set (healthy wheat T_unif0=0.3, T_cont0=0.7); Texture health normalization: the wheat lodging coefficient T (i.e., the probability of not lodging) is calculated: T=[0.5×(1-T_unif / T_unif0)+0.5×(T_cont / T_cont0)], T∈[0,1], the closer the value is to 1, the lighter the lodging degree, and 0.5 is the weight.
[0044] In one implementation, the wheat health coefficient is calculated based on the wheat color coefficient and the wheat lodging coefficient. The color health reflects the overall growth (weight 0.7), and the texture health reflects the lodging effect (weight 0.3). The weighted average is used to obtain the health coefficient H = 0.7 × C + 0.3 × T, where H ∈ [0, 1]. The closer the value is to 1, the better the wheat health.
[0045] In one implementation, the depthwise separable convolution module is a lightweight convolution module in deep learning. It first performs spatial convolution on each input channel using the corresponding convolution kernel, and then fuses cross-channel information through 1×1 pointwise convolution, which can significantly reduce the number of model parameters and computation.
[0046] In one embodiment, assessing the disaster severity of a target area includes: Water features are extracted from the second and third regions to obtain the second and third disaster-affected areas, and the area of the intersection region is calculated based on the second and third disaster-affected areas. The first area is determined based on the first region. The flood assessment coefficient is calculated based on the first area, the second affected area, the third affected area, and the area of the intersection. The assessment coefficient is calculated based on the wheat health coefficient and the flood assessment coefficient. If the evaluation coefficient falls within the first interval, the target area is determined to be severely damaged. If the assessment coefficient falls within the second interval, the target area is determined to be moderately damaged. If the evaluation coefficient falls within the third interval, the target area is determined to be slightly damaged.
[0047] In one implementation, a preliminary waterlogged area mask is extracted and segmented from the second region corresponding to the second feature map (during the disaster). Waterlogged areas other than wheat fields (such as rivers and ponds) are removed to obtain the initial waterlogged area S1 (second disaster-affected area) within the wheat field. Combined with the third region corresponding to the third feature map (after the disaster), the waterlogged area is extracted and segmented to obtain the post-disaster waterlogged area S2 (third disaster-affected area) of the wheat field. The area of wheat field continuously covered by waterlogging is calculated by the intersection of S1 and S2. J =S1∩S2.
[0048] In one implementation, the total area S of wheat fields in the first region corresponding to the first feature map (before the disaster) is calculated. Z (That is, the first area, determined by the wheat planting distribution mask), P=(S1+S2-S J ) / S Z *100%, P is the flood assessment coefficient, S1 is the second affected area, S2 is the third affected area, S J S represents the area of the intersection region. Z This is the first area.
[0049] In one implementation, the evaluation coefficient is calculated based on the wheat health coefficient and the flood assessment coefficient, and the evaluation coefficient = (weight 1 × flood assessment coefficient) / (weight 2 × wheat health coefficient); different treatments are carried out according to different damage conditions, such as artificial drainage in severely damaged areas, etc.
[0050] In one embodiment, improving the workflow of the neck structure includes: The output of the C2f module at layer 4 in the backbone network is used as the first input, the output of the C2f module at layer 6 in the backbone network is used as the second input, and the output of the SPPF module at layer 9 in the backbone network is used as the third input.
[0051] After upsampling the third input, it is concatenated with the second input to obtain the first tensor. The first tensor is then input into the improved C2f module to obtain the second tensor. After upsampling the second tensor, it is concatenated with the first input to obtain the third tensor. The third tensor is input into the improved C2f module and the Conv module in sequence to obtain the fourth tensor. The fourth tensor is concatenated with the third tensor to obtain the fifth tensor. The fifth tensor is input into the improved C2f module and the Conv module in sequence to obtain the sixth tensor. The 6th tensor is concatenated with the 3rd input to obtain the 7th tensor. The 3rd, 6th, and 7th tensors are used as the output for improving the neck structure.
[0052] In one implementation, see [link to implementation details]. Figure 2 , Figure 2 This is a flowchart illustrating the improved neck structure of the wheat flood disaster accurate identification and disaster assessment method based on multi-temporal remote sensing images provided in this invention embodiment. By fusing feature inputs from different levels of the backbone network, combined with multi-round upsampling, feature stitching, and enhanced processing of the improved C2f module, the method achieves full integration of shallow detail features and deep semantic features, improving the ability to represent waterlogged areas, abnormal growth areas, and lodging features in wheat flood disaster scenarios, and providing richer feature support for accurate identification and quantitative assessment.
[0053] One implementation method employs a structural design that combines feature stitching and multi-module progressive optimization. This approach retains key feature information while enhancing the coherence of feature propagation, reducing feature loss, and improving the accuracy for disaster-affected targets at different scales (such as small-scale water accumulation or localized collapsed areas).
[0054] In one embodiment, the improved operation of the C2f module includes: Obtain the original feature map, input the original feature map into a 7×7 depthwise separable convolutional module to obtain the first feature map, input the first feature map into a 1×1 convolutional module to obtain the second feature map, and input the first feature map into a 1×1 convolutional module and activate it to obtain the third feature map; The second feature map is multiplied element-wise with the third feature map to obtain the fourth feature map. The fourth feature map is then fed into a 1×1 convolutional module and a 7×7 depthwise separable convolutional module to obtain the fifth feature map. The fifth feature map is then used as the output of the improved C2f module.
[0055] In one implementation, the module uses a combination of 7×7 depthwise separable convolution followed by 1×1 convolution twice. The large 7×7 convolution kernel ensures that the module can obtain a wide receptive field during both the initial and subsequent feature extraction, which is beneficial for capturing a wider range of contextual information in the image. The depthwise separable convolution decomposes the standard convolution into depthwise convolution and pointwise convolution, reducing the number of computational parameters and costs. The 1×1 convolution is responsible for integrating and transforming cross-channel features, which improves the module's processing efficiency while ensuring feature extraction capabilities.
[0056] In one implementation, the module constructs a branch path containing non-linear activation and multiplies it element-wise with the output of another shortcut path. The feature map of one branch can be regarded as a set of weights, and the features of the other branch are recalibrated element-wise, thereby amplifying important features and suppressing irrelevant or noisy information. This enhances the expressive power of key features, enabling the module to more effectively filter and highlight discriminative information, thereby improving the quality of the output features.
[0057] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention should still fall within the scope of the claims of the present invention.
Claims
1. A method for accurate identification and disaster assessment of wheat flooding based on multi-temporal remote sensing imagery, characterized in that, The method includes: A remote sensing image set is obtained by periodically acquiring remote sensing images within a region, and multiple feature points are obtained by extracting features from the remote sensing images in the remote sensing image set. A feature matrix is constructed based on each feature point, and the feature matrix is added as a label to the corresponding remote sensing image to obtain a localization image set. The first feature map, the second feature map, and the third feature map are obtained by sampling the positioning images in the positioning image set at a preset time. The target region is obtained by aligning the feature matrices in the first feature map, the second feature map, and the third feature map, and the degree of disaster is obtained by evaluating the target region.
2. The method for accurate identification and disaster assessment of wheat flooding based on multi-temporal remote sensing imagery according to claim 1, characterized in that, Multiple feature points are obtained by feature extraction from the remote sensing images in the remote sensing image set, including: The remote sensing image is mapped to a 64-dimensional feature space to obtain a 64-dimensional feature image. The low-level features are convolved on the 64-dimensional feature image, the middle-level features are convolved on the low-level features, the high-level features are convolved on the middle-level features, and the high-level features are input into a 3×3 convolution module to obtain the original feature tensor. After performing a global average pooling operation on the original feature tensor, channel attention weights are activated. The channel attention weights are then multiplied channel by channel with the original feature tensor to obtain a channel-weighted feature map. Max pooling is performed on the channel-weighted feature map to obtain a first weighted feature map. Average pooling is performed on the channel-weighted feature map to obtain a second weighted feature map. The first weighted feature map and the second weighted feature map are concatenated and activated to obtain spatial attention weights. The spatial attention weights are multiplied element-wise with the channel-weighted feature map to obtain an enhanced feature map. Each channel of the enhanced feature map is convolved according to the spatial dimension to obtain multiple local features. Each local feature is input into a 1×1 convolution module and fused to obtain an enhanced feature tensor.
3. The method for accurate identification and disaster assessment of wheat flooding based on multi-temporal remote sensing imagery according to claim 2, characterized in that, After inputting each local feature into a 1×1 convolution module to obtain an enhanced feature tensor, feature extraction is performed on the enhanced feature tensor to obtain feature points, including: The enhanced feature tensor is input into the decoder to obtain a 65-channel feature tensor. The probability distribution is calculated for each channel in the 65-channel feature tensor. The 65-channel feature tensor consists of 64 channel pixel positions and 1 invalid channel. The probability distribution consists of multiple valid probabilities and one invalid probability, where the invalid probability is the probability of an invalid channel. Candidate probability distributions are obtained by filtering the probability distributions according to the probability threshold, and valid probability distributions are obtained by filtering the candidate probability distributions according to the invalid probabilities. The candidate points corresponding to each valid probability in the valid probability distribution are used as feature points.
4. The method for accurate identification and disaster assessment of wheat flooding based on multi-temporal remote sensing imagery according to claim 1, characterized in that, The feature matrix is constructed based on each feature point, including: The feature information and coordinate information of each feature point are determined. A descriptive feature matrix is constructed based on the feature information, and a spatial coordinate matrix is constructed based on the coordinate information. The descriptive feature matrix and the spatial coordinate matrix are then fused to obtain the feature matrix.
5. The method for accurate identification and disaster assessment of wheat flooding based on multi-temporal remote sensing imagery according to claim 1, characterized in that, The preset time includes a first time, a second time, and a third time; the first time is less than the second time, and the second time is less than the third time.
6. The method for accurate identification and disaster assessment of wheat flooding based on multi-temporal remote sensing imagery according to claim 1, characterized in that, The target region includes a first region, a second region, and a third region; the first feature map corresponds one-to-one with the first region, the second feature map corresponds one-to-one with the second region, and the third feature map corresponds one-to-one with the third region.
7. The method for accurate identification and disaster assessment of wheat flooding based on multi-temporal remote sensing imagery according to claim 6, characterized in that, The disaster severity is assessed by evaluating the target area, and the wheat health coefficient is obtained by identifying the third area using the target model, including: The wheat color coefficient and wheat lodging coefficient are obtained by extracting features from the third region using the target model, and the wheat health coefficient is calculated based on the wheat color coefficient and wheat lodging coefficient. The target model is an improvement based on the YOLOv8 model, including: The target model is obtained by replacing the Conv modules in the backbone network with depthwise separable convolutional modules and replacing the neck structure with an improved neck structure; the YOLOv8 module includes a backbone network and a neck structure.
8. The method for accurate identification and disaster assessment of wheat flooding based on multi-temporal remote sensing imagery according to claim 7, characterized in that, The disaster severity is assessed by evaluating the target area, including: Water features are extracted from the second region and the third region to obtain the second disaster area and the third disaster area. The area of the intersection region is calculated based on the second disaster area and the third disaster area. A first area is determined based on the first region. A flood assessment coefficient is calculated based on the first area, the second affected area, the third affected area, and the area of the intersection region. An assessment coefficient is calculated based on the wheat health coefficient and the flood assessment coefficient. If the evaluation coefficient falls within the first interval, the target area is determined to be severely damaged. If the evaluation coefficient falls within the second interval, the target area is determined to be moderately damaged. If the evaluation coefficient falls within the third interval, the target area is determined to be slightly damaged.
9. The method for accurate identification and disaster assessment of wheat flooding based on multi-temporal remote sensing imagery according to claim 7, characterized in that, The workflow for improving the neck structure includes: The output of the C2f module at layer 4 in the backbone network is used as the first input, the output of the C2f module at layer 6 in the backbone network is used as the second input, and the output of the SPPF module at layer 9 in the backbone network is used as the third input. The third input is upsampled and then concatenated with the second input to obtain the first tensor. The first tensor is then input into the improved C2f module to obtain the second tensor. The second tensor is upsampled and then concatenated with the first input to obtain the third tensor. The third tensor is sequentially input into the improved C2f module and the Conv module to obtain the fourth tensor. The fourth tensor is concatenated with the third tensor to obtain the fifth tensor. The fifth tensor is sequentially input into the improved C2f module and the Conv module to obtain the sixth tensor. The 6th tensor is concatenated with the third input to obtain the 7th tensor. The 3rd tensor, the 6th tensor, and the 7th tensor are used as the output for improving the neck structure.
10. The method for accurate identification and disaster assessment of wheat flooding based on multi-temporal remote sensing imagery according to claim 9, characterized in that, The working principle of the improved C2f module includes: Obtain the original feature map, input the original feature map into a 7×7 depthwise separable convolution module to obtain the first feature map, input the first feature map into a 1×1 convolution module to obtain the second feature map, and activate the first feature map after inputting it into the 1×1 convolution module to obtain the third feature map; The second feature map is multiplied element-wise with the third feature map to obtain the fourth feature map. The fourth feature map is then input sequentially into a 1×1 convolution module and a 7×7 depthwise separable convolution module to obtain the fifth feature map. The fifth feature map is then used as the output of the improved C2f module.
Citation Information
Patent Citations
Flood disaster assessment method based on improved remote sensing image feature matching algorithm
CN117372893A
Cited By
Crop all-weather monitoring and disaster assessment method based on cross-modal self-distillation
CN122090284A