Multi-temporal remote sensing image classification method and system under space-time attention mechanism
Through the multi-time phase remote sensing image classification method under the space-time attention mechanism, the ResNet18 network and sliding window block processing are used to solve the problem of efficient extraction of time and space information in remote sensing images, the accurate identification and classification of crop growth types is realized, and the support capacity of agricultural production is improved.
Patent Information
- Application Number
- CN202510393668.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the remote sensing data of crops with many times are complicated, and the time and space information in remote sensing images cannot be extracted efficiently, and important crop growth moments, such as key information in flowering and maturity periods, are ignored.
The space-time attention mechanism is adopted to extract multi-scale feature sequences through the ResNet18 network, combine sliding window chunking processing, generate memory key values and query key values, calculate the output of the time and space attention mechanism, and fuse multi-time sequence multi-spectral remote sensing images to generate enhanced features and restore image size.
It improves the accuracy and efficiency of crop classification of remote sensing images, can more reliably identify crop growth types, provide support for agricultural production, and realizes the automation and precise classification of remote sensing images.
Smart Images

Figure CN120259770A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of geographic information / agricultural survey, and relates to a multi-temporal remote sensing image classification method and system under a spatio-temporal attention mechanism. Background Art
[0002] The crop futures market is a crucial part of the global agricultural industrial chain, playing an indispensable role in price discovery, risk management, and resource allocation. For the effective operation of the crop futures market, it is crucial to accurately predict and grasp the changes in crop yields. Crop remote sensing image classification technology provides important information support for market participants. With the continuous progress of remote sensing technology, a large number of high-precision and high-resolution crop remote sensing images can be obtained. In particular, Sentinel-1 and Sentinel-2 satellite images contain a large amount of crop spectral information, which can reflect key data such as crop types, growth status, and yield expectations. However, how to accurately and efficiently extract this key information from a large amount of remote sensing images and apply it to the analysis and prediction of the crop futures market has always been a challenge faced by the industry.
[0003] In current research, convolutional neural networks (CNNs) have made significant progress in the field of remote sensing image classification. Their powerful feature extraction and classification capabilities have shown great potential in crop remote sensing image classification. However, most current CNN-based crop remote sensing image classification methods mainly focus on the extraction of global spatial features, not only ignoring the temporal information of multi-temporal remote sensing images but also lacking efficient extraction of spatial information, and not fully considering the characteristic change laws generated over time and the interval nature of spatial information. In addition, for important moments during the crop growth process (such as the flowering period, maturity period, etc.), traditional CNN methods have not been prominently processed, resulting in these key information being possibly ignored during the classification process. Summary of the Invention
[0004] The purpose of the present invention is to solve the technical problems in the prior art of the complexity of crop multi-temporal remote sensing data information and the inability to efficiently extract temporal and spatial information from remote sensing images, and to provide a multi-temporal remote sensing image classification method and system under a spatio-temporal attention mechanism.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] In a first aspect, the present invention discloses a multi-temporal remote sensing image classification method under a spatio-temporal attention mechanism, including:
[0007] Obtain multiple remotely sensed images of crops arranged in time series and input them into independent ResNet18 networks respectively, and extract multi-scale feature sequences containing several different levels at each moment;
[0008] Merge the last two levels in the multi-scale feature sequence along the channel direction to generate a fused feature, integrate the fused features of all previous moments into a memory module, and use the fused feature of the current latest moment alone as a query module;
[0009] Based on a sliding window, divide the memory module and the query module into blocks, and generate memory key values, query key values, and query content matrices through convolutional layers respectively; obtain the output of the temporal attention mechanism and the output of the spatial attention mechanism based on the memory key values and the query key values;
[0010] Add the output of the temporal attention mechanism and the output of the spatial attention mechanism after restoring the block structure, and perform channel merging with the query content matrix to generate an enhanced feature;
[0011] Recombine the blocks of the enhanced feature according to the original spatial positions, restore them to the input image size through convolutional layers and linear interpolation, and output the crop growth type labels to complete the classification of multi-temporal remotely sensed images of crops.
[0012] A further improvement lies in:
[0013] The obtaining of multiple remotely sensed images of crops arranged in time series and inputting them into independent ResNet18 networks respectively, and extracting multi-scale feature sequences containing several different levels at each moment includes:
[0014] Read the remotely sensed images of crops at n different moments as Input in Tensor format i (i = 1, 2,..., n), and form a List of length n in chronological order = [Input1, Input2,..., Input n ; the dimension of the single-moment remotely sensed image data is:
[0015] Input i ∈(C input , H input , W input )
[0016] where C input represents the number of channels Channels of the image, H input represents the height Heights of the image, and W input represents the width Widths of the image;
[0017] Pass the n Inputs in the List into independent Resnet18 networks respectively to obtain n multi-scale feature sequences Z i= [Z i,1 , Z i,2 , Z i,3 , Z i,4 , i = 1, 2, …, n.
[0018] The step of merging the last two levels in the multi-scale feature sequence along the channel direction to generate a fused feature, integrating the fused features of all previous moments into a memory module, and taking the fused feature of the current latest moment as a query module alone includes:
[0019] Merging Z i,3 and Z i,4 in the n multi-scale feature sequences in the channel direction to obtain a fused feature Feat i = Z i, 3cat Z i,4 , i = 1, 2, …, n; after obtaining the fused feature, reconstruct the fused features Feat i of the first n - 1 moments into a memory module Memory, and the fused feature of the nth moment is called a query module Query; the dimensions of the memory module Memory and the query module Query are:
[0020]
[0021] where T = n - 1, C feat represents the channel of the features of the memory module Memory and the query module Query, H1 represents the height of the features of the memory module Memory and the query module Query, and W1 represents the width of the features of the memory module Memory and the query module Query.
[0022] The step of dividing the memory module and the query module into blocks based on a sliding window and generating a memory key-value, a query key-value, and a query content matrix through convolutional layers respectively includes:
[0023] Dividing the memory module Memory and the query module Query into blocks based on a sliding window, where the height and width of each block (patch) are both 4, as shown in the following formula:
[0024]
[0025] where is the number of blocks;
[0026] Calculating a memory key-value Memory_key and a memory content Memory_value of the memory module Memory through an independent convolutional layer, and calculating a query key-value Query_key and a query content Query_value of the query module Query through a convolutional layer;
[0027] The characteristic dimensions of the memory key value Memory_key and the memory content Memory_value are as follows:
[0028]
[0029] The characteristic dimensions of the query key value Query_key and the query content Query_value are as follows:
[0030]
[0031] The specific processes of obtaining the output of the temporal attention mechanism and the output of the spatial attention mechanism based on the memory key value and the query key value include:
[0032] After performing channel dimension expansion and transpose operations on the query key value and the memory key value, matrix multiplication is carried out, and the temporal similarity weights are generated through Softmax. The memory values are weighted and aggregated to obtain the output of the temporal attention mechanism. Specifically:
[0033] The query key value Query_key and the memory key value Memory_key are subjected to characteristic dimension transformation as shown in the following formula:
[0034]
[0035] Among them, R represents the Reshape function, and P represents the Permute function;
[0036] The query key value Query_key and the memory key value Memory_key are multiplied matrix-wise and passed through softmax to obtain the similarity matrix S time , as shown in the following formula:
[0037]
[0038] Among them, represents matrix multiplication;
[0039] The memory content Memory_value is subjected to characteristic dimension transformation as shown in the following formula:
[0040]
[0041] The similarity matrix S time is multiplied matrix-wise with Memory_value to obtain the output Output time of the temporal attention mechanism, as shown in the following formula:
[0042]
[0043] Perform matrix multiplication after performing spatial dimension expansion and transposition operations on the query key value and the memory key value, and generate spatial similarity weights through Softmax, and weighted aggregate the memory values to obtain the output of the spatial attention mechanism; specifically:
[0044] Perform feature dimension transformation on the query key value Query_key and the memory key value Memory_key, as shown in the following formula:
[0045]
[0046] Among them, R represents the Reshape function, and P represents the Permute function;
[0047] Perform matrix multiplication on the query key value Query_key and the memory key value Memory_key and obtain the similarity matrix S through softmax space , as shown in the following formula:
[0048]
[0049] Among them, represents matrix multiplication;
[0050] Perform feature dimension transformation on the memory content Memory_value, as shown in the following formula:
[0051]
[0052] Multiply the similarity matrix S time with Memory_value to obtain the output Output of the spatial attention mechanism space , as shown in the following formula:
[0053]
[0054] The above-mentioned adding the output of the temporal attention mechanism and the output of the spatial attention mechanism after restoring the block structure, and performing channel merging with the query content matrix to generate enhanced features includes:
[0055] Perform feature dimension transformation on the output Output of the temporal attention mechanism time and the output Output of the spatial attention mechanism space , as shown in the following formula:
[0056]
[0057] Perform feature dimension transformation on the output Output of the temporal attention mechanism time and the output Output of the spatial attention mechanism spaceAdd them together and merge with the query content Query_value in the channel dimension to obtain the final output output, as shown in the following formula:
[0058] Output = (Output time + Output space ) cat Query_value ∈ (N, 2×C value , 4, 4).
[0059] Recombining the blocks of the enhanced features according to the original spatial positions, restoring them to the input image size through a convolutional layer and linear interpolation, and outputting the crop growth status labels to complete the multi-temporal remote sensing image classification of crops includes:
[0060] Stitching back the N blocks of the final output output according to the original positions, passing through a convolutional layer and obtaining the output label Label through a linear interpolation algorithm, as shown in the following formula:
[0061]
[0062] Among them, Conv represents the convolutional layer, and Interpolate represents the linear interpolation operation.
[0063] In a second aspect, the present invention discloses a multi-temporal remote sensing image classification system under a spatio-temporal attention mechanism, including:
[0064] A multi-scale feature sequence extraction module, configured to obtain multiple temporally arranged crop remote sensing images and input them into independent ResNet18 networks respectively, and extract multi-scale feature sequences including several different levels at each moment;
[0065] A fused feature generation module, configured to merge the last two levels in the multi-scale feature sequence along the channel direction to generate a fused feature, integrate the fused features of all previous moments into a memory module, and use the fused feature of the current latest moment alone as a query module;
[0066] A spatio-temporal attention calculation module, configured to block the memory module and the query module based on a sliding window, generate a memory key-value, a query key-value, and a query content matrix respectively through a convolutional layer; obtain a time attention mechanism output and a space attention mechanism output based on the memory key-value and the query key-value;
[0067] A feature fusion output module, configured to add the time attention mechanism output and the space attention mechanism output after restoring the block structure, and merge them with the query content matrix in the channel dimension to generate an enhanced feature;
[0068] The feature reconstruction decoding module is used to reorganize the enhanced feature blocks according to the original spatial position, restore them to the input image size through convolutional layers and linear interpolation, output crop growth status labels, and complete the classification of multi-temporal remote sensing images of crops.
[0069] In a third aspect, the present invention discloses an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the multi-temporal remote sensing image classification method under the above-mentioned spatiotemporal attention mechanism when executing the computer program.
[0070] In a fourth aspect, the present invention discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the multi-temporal remote sensing image classification method under the spatiotemporal attention mechanism described in any one of claims 1 to 7 is implemented.
[0071] Compared with the prior art, the present invention has the following beneficial effects:
[0072] The invention discloses a multi-temporal remote sensing image classification method under a spatiotemporal attention mechanism. Aiming at the problem of complicated information of multi-temporal remote sensing data of crops, the invention can efficiently and accurately extract key high-frequency spatiotemporal features by fusing multi-series multispectral remote sensing images and using a dual attention mechanism of time and space. This mechanism enhances the model's ability to capture complex spatiotemporal information, making the classification process more accurate. The invention adopts a carefully designed sliding window strategy, that is, sliding window block processing to optimize the feature extraction process. This strategy ensures full utilization of time and space information, avoids omission or redundancy of information, and thus improves the effectiveness and efficiency of feature extraction. By fusing multi-scale features, using a spatiotemporal attention mechanism and optimizing the feature extraction process, the invention significantly improves the accuracy of crop classification of remote sensing images. This enables the method to more reliably identify the growth type of crops in practical applications, provide strong support for agricultural production, and provide a reference for market sales pricing of agricultural products. The invention realizes the automation and precision classification of remote sensing images, greatly improving the efficiency and effect of remote sensing image processing. This method can quickly process a large amount of remote sensing data, providing possibilities for real-time monitoring and decision support in the field of agricultural remote sensing. The present invention can complete the prediction task of planting types in advance based on historical remote sensing images, and provide a basis for crop area estimation and yield calculation. This is of great significance for agricultural management, international grain trade decision-making, etc., and brings practical application value to the field of agricultural remote sensing.
[0073] The present invention discloses a multi-temporal remote sensing image classification system under a spatio-temporal attention mechanism. The system adopts a modular design, including a multi-scale feature sequence extraction module, a fused feature generation module, a spatio-temporal attention calculation module, a feature fusion output module, and a feature reconstruction and decoding module. Each module undertakes a specific function, making the system structure clear and easy to understand and maintain. The multi-scale feature sequence extraction module uses an independent ResNet18 network to extract features from multiple temporally arranged crop remote sensing images, obtaining a multi-scale feature sequence containing several different levels. This multi-scale feature extraction method can capture information at different scales and levels in the image, providing a rich feature basis for subsequent spatio-temporal attention calculation. The fused feature generation module merges the last two levels in the multi-scale feature sequence along the channel direction to generate a fused feature, integrates the fused features of all previous time steps into a memory module, and uses the fused feature of the current latest time step as a query module. The spatio-temporal attention calculation module performs block processing on the memory module and the query module based on a sliding window, generates memory key-value, query key-value, and query content matrices through convolutional layers, and calculates the output of the temporal attention mechanism and the output of the spatial attention mechanism. This way of fusing spatio-temporal information can more comprehensively consider the temporal and spatial relationships in the image, improving the accuracy of classification. The feature fusion output module adds the output of the temporal attention mechanism and the output of the spatial attention mechanism after restoring the block structure, and performs channel merging with the query content matrix to generate an enhanced feature. The feature reconstruction and decoding module reorganizes the blocks of the enhanced feature according to the original spatial positions, restores them to the input image size through convolutional layers and linear interpolation, and outputs the crop growth status labels. This way of feature fusion and reconstruction can retain the spatial information of the image while enhancing the expression ability of the features, making the classification result more accurate and reliable. The multi-temporal remote sensing image classification system under the spatio-temporal attention mechanism proposed by the present invention can be applied to multiple fields such as crop growth type monitoring, planting area estimation, and yield prediction, providing strong support for agricultural production and management. The system has the characteristics of automation, high efficiency, and accuracy, can greatly improve the efficiency and effect of remote sensing image processing, and has broad application prospects and market value. Description of the Drawings
[0074] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0075] Figure 1 It is a flowchart of a multi-temporal remote sensing image classification method under a spatio-temporal attention mechanism in the present invention;
[0076] Figure 2 It is a calculation step diagram of a multi-temporal remote sensing image classification method under a spatio-temporal attention mechanism in the present invention;
[0077] Figure 3 It is a network structure diagram of a multi-temporal remote sensing image classification method under a spatio-temporal attention mechanism in the present invention;
[0078] Figure 4 It is a schematic diagram of the Shiftwindow (sliding window) strategy in a multi-temporal remote sensing image classification method under a spatio-temporal attention mechanism in the present invention;
[0079] Figure 5 It is a module diagram of a multi-temporal remote sensing image classification system under a spatio-temporal attention mechanism in the present invention. Detailed implementation manners
[0080] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0081] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0082] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0083] The present invention will be further described in detail below with reference to the accompanying drawings:
[0084] See Figure 1 , the embodiments of the present invention disclose a multi-temporal remote sensing image classification method under a spatio-temporal attention mechanism, including:
[0085] S1, obtaining multiple remotely sensed images of crops arranged in time series and respectively inputting them into independent ResNet18 networks to extract multi-scale feature sequences including several different levels at each moment;
[0086] S2. Merge the last two levels in the multi-scale feature sequence along the channel direction to generate fused features, integrate the fused features at all previous moments into a memory module, and use the fused feature at the current latest moment as a query module alone;
[0087] S3. Based on a sliding window, divide the memory module and the query module into blocks, and generate a memory key-value, a query key-value, and a query content matrix through convolutional layers respectively; obtain the output of the temporal attention mechanism and the output of the spatial attention mechanism based on the memory key-value and the query key-value;
[0088] S4. Add the output of the temporal attention mechanism and the output of the spatial attention mechanism after restoring the block structure, and perform channel merging with the query content matrix to generate enhanced features;
[0089] S5. Recombine the blocks of the enhanced features according to the original spatial positions, restore them to the input image size through convolutional layers and bilinear interpolation, and output the crop growth type labels to complete the classification of multi-temporal remote sensing images of crops.
[0090] The present invention discloses a method for classifying multi-temporal remote sensing images under a spatio-temporal attention mechanism. Aiming at the problem of complex information in multi-temporal remote sensing data of crops, the present invention can efficiently and accurately extract key high-frequency spatio-temporal features by fusing multi-temporal multi-spectral remote sensing images and with the help of dual spatio-temporal attention mechanisms. This mechanism enhances the model's ability to capture complex spatio-temporal information, making the classification process more accurate. The present invention adopts a carefully designed sliding window strategy, that is, sliding window block processing to optimize the feature extraction process. This strategy ensures the full utilization of temporal and spatial information, avoids information omission or redundancy, and thus improves the effectiveness and efficiency of feature extraction. By fusing multi-scale features, using spatio-temporal attention mechanisms, and optimizing the feature extraction process, the present invention significantly improves the accuracy of crop classification in remote sensing images. This enables the method to more reliably identify the growth status of crops in practical applications and provides strong support for agricultural production. The present invention realizes the automatic and precise classification of remote sensing images, greatly improving the efficiency and effect of remote sensing image processing. This method can quickly process a large amount of remote sensing data, providing the possibility for real-time monitoring and decision support in the field of agricultural remote sensing. The present invention can complete the prediction task of planting types in advance based on historical remote sensing images, providing a basis for crop area estimation and yield calculation. This is of great significance for agricultural management, international food trading decisions, etc., bringing practical application value to the field of agricultural remote sensing.
[0091] Participate Figure 2 and Figure 3 , and the content of the present invention will be described in detail below in conjunction with specific embodiments:
[0092] Taking Sentinel-2 data as the data source and the specific implementation object, first, the multi-temporal data of Sentinel-2 in South Dakota (including 30 moments) is cut, and the cutting standard is that the coordinate regions of the data at different moments are the same under the same coordinate system.
[0093] A multi-temporal remote sensing image classification method under a spatio-temporal attention mechanism in the present invention includes the following steps:
[0094] Step 1: Read the multi-temporal remote sensing images and input them into an independent Resnet18 module for spatial feature extraction; read the remote sensing images of crops at n different moments as Input in Tensor format i (i = 1, 2, …, n), and form a List of length n in chronological order = [Input1, Input2, …, Input n ; The dimension of the single-moment remote sensing image data is:
[0095] Input i ∈(C input , H input , W input )
[0096] where C input represents the number of channels of the image Channels, H input represents the height of the image Heights, and W input represents the width of the image Widths;
[0097] Input the n Inputs in the List into independent Resnet18 networks respectively to obtain n multi-scale feature sequences Z i = [Z i,1 , Z i,2 , Z i,3 , Z i,4 , i = 1, 2, …, n.
[0098] That is, read the remote sensing images at 30 different moments as inputs in Tensor format and form a List of length 30 in chronological order = [Input1, Input2, …, Input 30 . The data dimension of Input i is (13, 48, 48). Enter the Resnet18 network to obtain the multi-scale feature sequence. That is: Input the 30 Inputs in the List into 30 independent Resnet18 networks respectively to obtain 30 multi-scale feature sequences Z i = [Z i,1 , Z i,2 , Z i,3 , Z i,4, i = 1, 2, …, 30.
[0099] Step 2: Merge Z i,3 and Z i,4 in the channel direction to obtain the fused feature Feat i = Z i,3 cat Z i,4 , i = 1, 2, …, n; After obtaining the fused feature, reconstruct the fused features Feat i of the first n - 1 moments into the memory module Memory, and the fused feature at the nth moment is called the query module Query; The dimensions of the memory module Memory and the query module Query are:
[0100]
[0101] where T = n - 1, C feat represents the channel of the features of the memory module Memory and the query module Query, H1 represents the height of the features of the memory module Memory and the query module Query, and W1 represents the width of the features of the memory module Memory and the query module Query.
[0102] That is, after obtaining the fused feature, reconstruct the fused features Feat i of the first 29 moments into a memory module Memory in Tensor format, whose data dimension is (29, 768, 24, 24), and the fused feature at the 30th moment is called the query module Query, whose data dimension is (768, 24, 24).
[0103] Step 3: As Figure 4 shown, it is the schematic diagram of the Shift window strategy, where Height, Width, and Channel represent the height, width, and number of channels of a single remote sensing image respectively, n represents n moments, and the Shift window strategy will Figure 3 the Memory_key, Memory_value, Query_key, and Query_value in
[0104] be divided into N patches of equal size by a sliding window with a height and width of 4, and after division, they are connected to the subsequent self-attention module to extract temporal and spatial information.
[0105] Use the Shift window to block Memory and Query.
[0106]
[0107] Among them, is the number of blocks;
[0108] That is, the data is sliced into 36 patches with a height and width of 4 respectively. At this time, the dimension of the Memory module is (36, 29, 768, 4, 4), and the dimension of the Query module is (36, 768, 4, 4).
[0109] Calculate the memory key Memory_key and memory content Memory_value of the memory module Memory through an independent convolutional layer, and calculate the query key Query_key and query content Query_value of the query module Query through a convolutional layer;
[0110] The feature dimensions of the memory key Memory_key and the memory content Memory_value are:
[0111]
[0112] The feature dimensions of the query key Query_key and the query content Query_value are:
[0113]
[0114] Where N = 36, T = 29, C key = 256, C value = 512.
[0115] Step 4: After performing channel dimension expansion and transpose operations on the query key and the memory key, perform matrix multiplication, generate a temporal similarity weight through Softmax, and weighted aggregate the memory values to obtain the output of the temporal attention mechanism; specifically:
[0116] Perform feature dimension transformation on the query key Query_key and the memory key Memory_key, as shown in the following formula:
[0117]
[0118] Among them, R represents the Reshape function, and P represents the Permute function;
[0119] Perform matrix multiplication on the query key Query_key and the memory key Memory_key and obtain a similarity matrix S time , as shown in the following formula:
[0120]
[0121] Among them, Represents matrix multiplication.
[0122] Perform a feature dimension transformation on the memory content Memory_value as shown in the following formula:
[0123]
[0124] Multiply the similarity matrix S time with Memory_value through matrix multiplication to obtain the output Output of the temporal attention mechanism time as shown in the following formula:
[0125]
[0126] Step Five: After performing spatial dimension expansion and transpose operations on the query key value and the memory key value, perform matrix multiplication, generate spatial similarity weights through Softmax, and aggregate the memory values with weights to obtain the output of the spatial attention mechanism; specifically:
[0127] Perform a feature dimension transformation on the query key value Query_key and the memory key value Memory_key as shown in the following formula:
[0128]
[0129] where R represents the Reshape function and P represents the Permute function;
[0130] Multiply the query key value Query_key with the memory key value Memory_key through matrix multiplication and obtain the similarity matrix S through softmax space as shown in the following formula:
[0131]
[0132] where represents matrix multiplication;
[0133] Perform a feature dimension transformation on the memory content Memory_value as shown in the following formula:
[0134]
[0135] Multiply the similarity matrix S time with Memory_value through matrix multiplication to obtain the output Output of the spatial attention mechanism space as shown in the following formula:
[0136]
[0137] Step Six: Take the output Output of the temporal attention mechanismtime Perform a feature dimension transformation on the output of the spatial attention mechanism space as shown in the following formula:
[0138]
[0139] Add the output of the temporal attention mechanism time to the output of the spatial attention mechanism space and merge them with the query content Query_value in the channel dimension to obtain the final output output, as shown in the following formula:
[0140] Output = (Output time + Output space ) cat Query_value ∈ (N, 2×C value , 4, 4);
[0141] At this time, the data dimension of output is (36, 1024, 4, 4).
[0142] Step 7: Stitch the N blocks of the final output output back in their original positions, pass through a convolutional layer, and obtain the output label Label through the bilinear interpolation algorithm, as shown in the following formula:
[0143]
[0144] where Conv represents the convolutional layer and Interpolate represents the bilinear interpolation operation.
[0145] That is, stitch the 36 patches of output back in their original positions to obtain an output with a data dimension of (1024, 24, 24), pass through a convolutional layer, and obtain the output label Label through the bilinear interpolation algorithm, as shown in Equation (17). Finally, the dimension of the output Label is (4, 48, 48).
[0146] After completing the above steps on part of the data in South Dakota (1354 training images and 339 validation images), the classification accuracy of crop remote sensing images is shown in Table 1:
[0147] Table 1 Classification accuracy table of crop remote sensing images
[0148] P R F1 IOU Background 1 0.98 0.99 0.98 Corn 0.88 0.86 0.87 0.77 Soybean 0.90 0.85 0.87 0.78 Others 0.98 0.99 0.98 0.97 Average 0.940 0.919 0.929 0.874
[0149] As can be seen from the above table, using the method in the present invention to train and validate on part of the data in South Dakota, the average IOU of the classification of the main crops (soybeans and corn) reaches 77.5%, and the average IOU of the classification of the entire image reaches 87.4%, achieving a relatively excellent result in the crop classification algorithm, which proves the effectiveness of the present method.
[0150] See Figure 5 , the present invention discloses a multi-temporal remote sensing image classification system under a spatio-temporal attention mechanism, including:
[0151] A multi-scale feature sequence extraction module, configured to obtain multiple remotely sensed images of crops arranged in time series and respectively input them into independent ResNet18 networks, and extract multi-scale feature sequences including several different levels at each moment;
[0152] A fused feature generation module, configured to merge the last two levels in the multi-scale feature sequence along the channel direction to generate a fused feature, integrate the fused features of all previous moments into a memory module, and use the fused feature of the current latest moment as a query module alone;
[0153] A spatio-temporal attention calculation module, configured to block the memory module and the query module based on a sliding window, and respectively generate a memory key-value, a query key-value, and a query content matrix through a convolutional layer; obtain a time attention mechanism output and a spatial attention mechanism output based on the memory key-value and the query key-value;
[0154] A feature fusion output module, configured to add the time attention mechanism output and the spatial attention mechanism output after restoring the block structure, and perform channel merging with the query content matrix to generate an enhanced feature;
[0155] A feature reconstruction decoding module, configured to reorganize the blocks of the enhanced feature according to the original spatial positions, restore them to the input image size through a convolutional layer and linear interpolation, and output crop growth type labels to complete the classification of multi-temporal remote sensing images of crops.
[0156] The present invention discloses a multi-temporal remote sensing image classification system under a spatio-temporal attention mechanism. The system adopts a modular design and includes a multi-scale feature sequence extraction module, a fused feature generation module, a spatio-temporal attention calculation module, a feature fusion output module, and a feature reconstruction and decoding module. Each module undertakes a specific function, making the system structure clear and easy to understand and maintain. The multi-scale feature sequence extraction module uses an independent ResNet18 network to extract features from multiple temporally arranged crop remote sensing images, obtaining a multi-scale feature sequence containing several different levels. This multi-scale feature extraction method can capture information at different scales and levels in the image, providing a rich feature basis for subsequent spatio-temporal attention calculation. The fused feature generation module merges the last two levels in the multi-scale feature sequence along the channel direction to generate a fused feature, integrates the fused features of all previous time steps into a memory module, and uses the fused feature of the current latest time step as a query module. The spatio-temporal attention calculation module performs block processing on the memory module and the query module based on a sliding window, generates a memory key-value, a query key-value, and a query content matrix through a convolutional layer, and calculates the output of the temporal attention mechanism and the output of the spatial attention mechanism. This way of fusing spatio-temporal information can more comprehensively consider the temporal and spatial relationships in the image, improving the accuracy of classification. The feature fusion output module adds the output of the temporal attention mechanism and the output of the spatial attention mechanism after restoring the block structure, and performs channel merging with the query content matrix to generate an enhanced feature. The feature reconstruction and decoding module reorganizes the blocks of the enhanced feature according to the original spatial positions, restores them to the input image size through a convolutional layer and linear interpolation, and outputs the crop growth status label. This way of feature fusion and reconstruction can retain the spatial information of the image while enhancing the expressive ability of the features, making the classification result more accurate and reliable. The multi-temporal remote sensing image classification system under the spatio-temporal attention mechanism proposed by the present invention can be applied to multiple fields such as crop growth status monitoring, planting area estimation, and yield prediction, providing strong support for agricultural production and management. The system has the characteristics of automation, high efficiency, and accuracy, can greatly improve the efficiency and effect of remote sensing image processing, and has broad application prospects and market value.
[0157] The third object of the present invention is to provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the multi-temporal remote sensing image classification method under the spatio-temporal attention mechanism.
[0158] The multi-temporal remote sensing image classification method under the spatio-temporal attention mechanism includes the following steps:
[0159] Obtain multiple temporally arranged crop remote sensing images and respectively input them into an independent ResNet18 network to extract multi-scale feature sequences containing several different levels at each time step;
[0160] Merge the last two levels in the multi-scale feature sequence along the channel direction to generate a fused feature, integrate the fused features of all previous moments into a memory module, and use the fused feature of the current latest moment as a query module alone;
[0161] Based on a sliding window, divide the memory module and the query module into blocks, and generate a memory key-value, a query key-value, and a query content matrix through convolutional layers respectively; obtain the output of the temporal attention mechanism and the output of the spatial attention mechanism based on the memory key-value and the query key-value;
[0162] Add the output of the temporal attention mechanism and the output of the spatial attention mechanism after restoring the block structure, and perform channel merging with the query content matrix to generate an enhanced feature;
[0163] Reorganize the blocks of the enhanced feature according to the original spatial positions, restore them to the input image size through convolutional layers and linear interpolation, and output the crop growth status label to complete the classification of multi-temporal remote sensing images of crops.
[0164] The fourth object of the present invention is to provide a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the multi-temporal remote sensing image classification method under the spatio-temporal attention mechanism is implemented.
[0165] The multi-temporal remote sensing image classification method under the spatio-temporal attention mechanism includes the following steps:
[0166] Obtain multiple remotely sensed images of crops arranged in time series and input them into independent ResNet18 networks respectively, and extract multi-scale feature sequences containing several different levels at each moment;
[0167] Merge the last two levels in the multi-scale feature sequence along the channel direction to generate a fused feature, integrate the fused features of all previous moments into a memory module, and use the fused feature of the current latest moment as a query module alone;
[0168] Based on a sliding window, divide the memory module and the query module into blocks, and generate a memory key-value, a query key-value, and a query content matrix through convolutional layers respectively; obtain the output of the temporal attention mechanism and the output of the spatial attention mechanism based on the memory key-value and the query key-value;
[0169] Add the output of the temporal attention mechanism and the output of the spatial attention mechanism after restoring the block structure, and perform channel merging with the query content matrix to generate an enhanced feature;
[0170] Reorganize the blocks of the enhanced feature according to the original spatial positions, restore them to the input image size through convolutional layers and linear interpolation, and output the crop growth type label to complete the classification of multi-temporal remote sensing images of crops.
[0171] Those skilled in the art will understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0172] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0173] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0174] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0175] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A multi-temporal remote sensing image classification method under a spatio-temporal attention mechanism, characterized in that, Including: Obtain multiple remotely sensed images of crops arranged in time series and input them into independent ResNet18 networks respectively, and extract multi-scale feature sequences including several different levels at each moment; Merge the last two levels in the multi-scale feature sequence along the channel direction to generate a fused feature, integrate the fused features of all previous moments into a memory module, and use the fused feature of the current latest moment alone as a query module; Based on a sliding window, divide the memory module and the query module into blocks, and generate a memory key-value, a query key-value, and a query content matrix through convolutional layers respectively; Obtain the output of the temporal attention mechanism and the output of the spatial attention mechanism based on the memory key-value and the query key-value; Add the output of the temporal attention mechanism and the output of the spatial attention mechanism after restoring the block structure, and perform channel merging with the query content matrix to generate an enhanced feature; Recombine the blocks of the enhanced feature according to the original spatial position, restore it to the input image size through convolutional layers and linear interpolation, and output the crop growth type label to complete the classification of multi-temporal remotely sensed images of crops.
2. The multi-temporal remote sensing image classification method under the spatio-temporal attention mechanism according to claim 1, wherein The step of obtaining multiple remotely sensed images of crops arranged in time series and inputting them into independent ResNet18 networks respectively, and extracting multi-scale feature sequences including several different levels at each moment includes: Read the remote sensing images of crops at n different times as Input in Tensor format i (i = 1, 2, …, n), and form a list of length n in chronological order: List = [Input1, Input2, …, Input n ; The dimension of the remote sensing image data at a single time is: Input i ∈(C input ,H input ,W input ) Among them, C input represents the Channels of the image, H input represents the Heights of the image, W input represents the Widths of the image; The n Inputs in the List are respectively input into independent Resnet18 networks to obtain n multi-scale feature sequences Z i = [Z i,1 , Z i,2 , Z i,3 , Z i,4 , where i = 1, 2, …, n.
3. The multi-temporal remote sensing image classification method under the spatio-temporal attention mechanism according to claim 2, wherein, The step of merging the last two levels in the multi-scale feature sequence along the channel direction to generate a fused feature, integrating the fused features of all previous moments into a memory module, and using the fused feature of the current latest moment alone as a query module includes: Merge Z in the Z of n multi-scale feature sequences i,3 and Z i,4 in the channel direction to obtain the fused feature Feat i = Z i,3 catZ i,4 , i = 1, 2, …, n; after obtaining the fused feature, reconstruct the fused features Feat i at the previous n - 1 time steps into the memory module Memory, and the fused feature at the nth time step is called the query module Query; the dimensions of the memory module Memory and the query module Query are: where T = n - 1, C feat represents the channels of the features of the memory module Memory and the query module Query, H1 represents the height of the features of the memory module Memory and the query module Query, and W1 represents the width of the features of the memory module Memory and the query module Query.
4. The multi-temporal remote sensing image classification method under the spatio-temporal attention mechanism according to claim 3, wherein The step of dividing the memory module and the query module into blocks based on a sliding window and generating a memory key-value, a query key-value, and a query content matrix through convolutional layers respectively includes: Divide the memory module Memory and the query module Query into blocks based on a sliding window, and the height and width of each block (patch) are 4, as shown in the following formula: Among them, is the number of blocks; Calculate the memory key-value Memory_key and the memory content Memory_value of the memory module Memory through an independent convolutional layer, and calculate the query key-value Query_key and the query content Query_value of the query module Query through a convolutional layer; The feature dimensions of the memory key-value Memory_key and the memory content Memory_value are: The feature dimensions of the query key-value Query_key and the query content Query_value are:
5. The multi-temporal remote sensing image classification method under the spatio-temporal attention mechanism according to claim 4, wherein Specifically, obtaining the output of the temporal attention mechanism and the output of the spatial attention mechanism based on the memory key-value and the query key-value includes: Perform a matrix multiplication after performing channel dimension expansion and transpose operations on the query key-value and the memory key-value, and generate a temporal similarity weight through Softmax, and weighted aggregation of the memory values to obtain the output of the temporal attention mechanism; specifically: Perform feature dimension transformation on the query key-value Query_key and the memory key-value, as shown in the following formula: Among them, R represents the Reshape function, and P represents the Permute function; The query key Query_key and the memory key Memory_key are multiplied in matrix and passed through softmax to obtain the similarity matrix S time , as shown in the following formula: Among them, represents matrix multiplication; Perform feature dimension transformation on the memory content Memory_value, as shown in the following formula: Multiply the similarity matrix S time with Memory_value to obtain the output Output of the temporal attention mechanism time , as shown in the following formula: Perform matrix multiplication after performing spatial dimension expansion and transposition operations on the query key value and the memory key value, generate spatial similarity weights through Softmax, and weighted aggregate the memory values to obtain the output of the spatial attention mechanism. Specifically: Perform feature dimension transformation on the query key value Query_key and the memory key value Memory_key as shown in the following formula: where, R represents the Reshape function, and P represents the Permute function; The query key value Query_key is matrix-multiplied with the memory key value Memory_key and passed through softmax to obtain the similarity matrix S space , as shown in the following formula: Among them, represents matrix multiplication; Perform feature dimension transformation on the memory content Memory_value as shown in the following formula: Multiply the similarity matrix S time by Memory_value to obtain the output Output of the spatial attention mechanism space , as shown in the following equation:
6. The multi-temporal remote sensing image classification method under the spatio-temporal attention mechanism according to claim 5, wherein The step of adding the outputs of the temporal attention mechanism and the spatial attention mechanism after restoring the block structure, and performing channel merging with the query content matrix to generate enhanced features includes: Output of the temporal attention mechanism time and the output of the spatial attention mechanism space are subjected to a feature dimension transformation as shown in the following formula: Output of the temporal attention mechanism time is added to the output of the spatial attention mechanism space and merged with the query content Query_value in the channel dimension to obtain the final output output, as shown in the following formula: Output=(Output time +Output space )cat Query_value∈(N,2×C value ,4,4).
7. The multi-temporal remote sensing image classification method under the spatio-temporal attention mechanism according to claim 6, characterized in that, The step of reorganizing the blocks of the enhanced features according to the original spatial positions, restoring them to the input image size through a convolutional layer and linear interpolation, and outputting the crop growth status labels to complete the multi-temporal remote sensing image classification of crops includes: Stitch back the N blocks of the final output output according to the original positions, and obtain the output label Label through a convolutional layer and the linear interpolation algorithm as shown in the following formula: where, Conv represents the convolutional layer, and Interpolate represents the linear interpolation operation.
8. The multi-temporal remote sensing image classification system under the spatio-temporal attention mechanism according to claim 1, characterized in that Including: A multi-scale feature sequence extraction module, which is used to obtain multiple temporally arranged crop remote sensing images and input them into independent ResNet18 networks respectively, and extract multi-scale feature sequences including several different levels at each moment; A fused feature generation module, which is used to merge the last two levels in the multi-scale feature sequence along the channel direction to generate a fused feature, integrate the fused features of all previous moments into a memory module, and use the fused feature of the current latest moment as a query module alone; A spatio-temporal attention calculation module, which is used to divide the memory module and the query module into blocks based on a sliding window, and generate memory key values, query key values, and query content matrices through convolutional layers respectively; Obtain the output of the temporal attention mechanism and the output of the spatial attention mechanism based on the memory key value and the query key value; A feature fusion output module, which is used to add the outputs of the temporal attention mechanism and the spatial attention mechanism after restoring the block structure, and perform channel merging with the query content matrix to generate enhanced features; A feature reconstruction decoding module, which is used to reorganize the blocks of the enhanced features according to the original spatial positions, restore them to the input image size through a convolutional layer and linear interpolation, and output the crop growth status labels to complete the multi-temporal remote sensing image classification of crops.
9. An electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the multi-temporal remote sensing image classification method under the spatio-temporal attention mechanism according to any one of claims 1-7.
10. A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the multi-temporal remote sensing image classification method under the spatio-temporal attention mechanism according to any one of claims 1-7.
Citation Information
Cited By
Large-scale MIMO channel estimation method fusing double attention mechanism and TCN-BiLSTM network
CN121000559A
Large-scale MIMO channel estimation method fusing dual attention mechanism and TCN-BiLSTM network
CN121000559B