Remote sensing image multi-source heterogeneous data fusion processing method and system
By combining convolutional neural networks and cross-temporal attention mechanisms, the problems of semantic inconsistency and insufficient detail preservation in remote sensing image fusion were solved, achieving efficient and accurate fusion of multi-source remote sensing data and improving the spatial resolution and semantic consistency of remote sensing images.
Patent Information
- Application Number
- CN202510767510.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Existing remote sensing image fusion methods struggle to effectively capture the complex semantic relationships between data acquired by different sensors, lack a modeling mechanism for dynamic evolution over time, and fail to retain sufficient local detail information, resulting in inconsistent semantic expression, poor stability, and blurred edge structures in the fusion results.
A convolutional neural network is used to extract global information. Combined with a lightweight multi-scale pyramid model and a cross-temporal attention mechanism, an interactive network model with a progressive fusion strategy and hybrid contrastive learning is used to achieve multi-level structural perception and semantic consistency restoration of remote sensing images.
It significantly improves the spatial resolution and semantic clarity of remote sensing images, enhances the complementarity of information in the temporal and spatial dimensions, improves the accuracy and controllability of the fusion results, takes into account computational efficiency, and is suitable for efficient and high-precision fusion processing of multi-source remote sensing data.
Smart Images

Figure CN120808082A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of remote sensing image data fusion, and particularly relates to a remote sensing image multi-source heterogeneous data fusion processing method and system. BACKGROUND
[0002] With the rapid development of remote sensing technology, multi-source remote sensing data presents diversified characteristics in spatial resolution, temporal resolution and spectral resolution, and these data have the characteristics of strong complementarity and rich information, and play an important role in the fields of city planning, environmental monitoring, disaster warning and the like. In order to fully utilize the information potential of multi-source remote sensing data and improve the remote sensing image interpretation accuracy and application efficiency, the multi-source heterogeneous data fusion technology has become one of the key research directions.
[0003] However, the existing remote sensing image fusion method still has many deficiencies, mainly including the following aspects:
[0004] (1) The traditional method has weak modeling ability for the semantic consistency of multi-source heterogeneous data, and most of the current fusion methods rely on hand-designed features or simple superposition strategies based on pixel level, which is difficult to effectively capture the complex semantic relationship between data obtained by different sensors, resulting in inconsistent semantic expression of the fusion result, affecting the accuracy of data monitoring.
[0005] (2) Lack of modeling mechanism for dynamic evolution of time dimension, the existing technology often processes each phase remote sensing image independently, and does not fully mine the trend information of ground object target changing with time, resulting in poor stability of the fusion result in dynamic scene, and it is difficult to meet the needs of long time sequence remote sensing monitoring.
[0006] (3) Insufficient preservation of local detail information, fuzzy edge structure, in the process of multi-source data fusion, many algorithms pay more attention to the overall feature expression, ignoring the effective extraction and preservation of local details of the image, resulting in blurred boundaries and unclear ground object contours of the fused remote sensing image, affecting the identification and interpretation of key ground objects. SUMMARY
[0007] The purpose of the present application is to provide a remote sensing image multi-source heterogeneous data fusion processing method and system, to solve the technical problems that the existing technology is difficult to effectively capture the complex semantic relationship between data obtained by different sensors, resulting in inconsistent semantic expression of the fusion result, lack of modeling mechanism for dynamic evolution of time dimension, and insufficient preservation of local detail information, fuzzy edge structure and the like.
[0008] To solve the above technical problems, the present application specifically provides the following technical scheme:
[0009] The first aspect of the present application provides a remote sensing image multi-source heterogeneous data fusion processing method, comprising the following steps:
[0010] The time-continuous remote sensing image data adopts a convolutional neural network as an encoder to extract global information, extracts image local features through cropping operations, and adopts a lightweight multi-scale pyramid model combined with a feature semantic extraction module to capture local feature information in the remote sensing image data.
[0011] The local feature information is extracted through a cross-time domain attention mechanism constructed by a cyclic matrix to extract mutual information between different modal variables, and filter highly correlated cross-time domain key information.
[0012] The cross-time domain key information is established by a cross-time domain perception hierarchical aggregation module, the local feature information is aggregated on the corresponding global information according to the time domain, and the detail information and edge information of the remote sensing image are obtained.
[0013] The detail information and edge information of the remote sensing image are fused according to the perception hierarchical level by using an asymptotic fusion strategy, global fusion data is obtained, a loss function is introduced to the global fusion data to narrow the semantic gap, and semantic fusion information is obtained.
[0014] A hybrid contrast learning interaction network model is used to repair missing information in the semantic fusion information, and complete real-time remote sensing image multi-source heterogeneous fusion data is obtained.
[0015] As a preferred scheme of the present application, the time-continuous remote sensing image data adopts a convolutional neural network as an encoder to extract global information, comprising:
[0016] The multi-source remote sensing image data is sequentially input into a deep convolutional neural network in time sequence, a convolutional layer is used as the main part, a multi-level feature extraction module is constructed, and the spatial feature expression of the image is gradually extracted through stacking convolutional layers and pooling layers.
[0017] The spatial feature expression adopts an attention mechanism to capture the semantic information of the remote sensing image, maps each remote sensing image into a high-dimensional feature vector, retains its key spatial distribution rule and time evolution trend, and outputs a high-level feature vector.
[0018] The global feature vector of the high-level feature vector is extracted, and the global feature vector is used as the global information in the multi-source heterogeneous data fusion stage.
[0019] As a preferred scheme of the present application, the image local features are extracted through cropping operations, and a multi-scale hierarchical feature aggregation module is used to obtain multi-scale hierarchical aggregated features.
[0020] The global information is segmented by a multi-scale hierarchical feature aggregation module of the convolution layer, the multi-scale hierarchical feature aggregation module is provided with K channels, each channel is provided with a corresponding dilatation rate r n ;
[0021] The number of scales of the multi-scale hierarchical feature aggregation module is set as M, and the size of each channel is set as M / 6;
[0022] The size of the convolution layer is constrained by a pyramid model, three 4*4 convolution blocks are connected to aggregate local features of the image;
[0023] The feature information of each feature pyramid channel of the pyramid model is added successively by hierarchical feature aggregation to obtain multi-scale hierarchical aggregation features.
[0024] As a preferred scheme of the present application, the multi-scale hierarchical aggregation features are combined with a feature semantic extraction module to capture local feature information in remote sensing image data, including:
[0025] The multi-scale hierarchical aggregation features are acquired by the feature semantic extraction module to obtain local feature information, the feature semantic extraction module is composed of three different convolutions, namely pointwise convolution, ordinary convolution and dilated convolution, the pointwise convolution, ordinary convolution and dilated convolution are respectively convolved with the multi-scale hierarchical aggregation features based on time domain;
[0026] Each feature after convolution is output by a normalization layer to obtain aggregated cascaded features with the same target dimension, and the aggregated cascaded features are decoded by a multi-scale attention mechanism to output local feature information, and the expression is as follows:
[0027]
[0028] Wherein, F represents the encoded feature, respectively represent the Relu and Sigmoid activation function weights of the convolution layer, F forcos represents the aggregated cascaded feature, F s represents the local feature information output after decoding by the multi-scale attention mechanism, BN{PointwiseConv(F), BN{OrdinaryConv(F), BN{DilationConv(F) respectively correspond to the convolution operation of the pointwise convolution, ordinary convolution and dilated convolution, and Concat() represents a feature aggregation function.
[0029] As a preferred scheme of the present application, the cross-time domain attention mechanism is constructed by a cyclic matrix to extract mutual information between different modal variables of the local feature information, and highly correlated cross-time domain key information is screened, including:
[0030] The local feature information extracted at multiple time points is organized in time sequence as a feature sequence D = {d1, d2, …, dT}, where T represents the time length of the feature sequence. t}, d t represents the local feature map at the t-th time step.
[0031] A cyclic matrix M is constructed according to the feature sequence D to build the mutual relationship between features at different time points, and the expression of the cyclic matrix M is:
[0032] M = similarity (d i , d j )
[0033] Wherein, the similarity measure of the cyclic matrix M is obtained by cosine similarity calculation, d i , d j represent the local feature maps at the i-th and j-th time steps, respectively.
[0034] An attention weight matrix is constructed through the cyclic matrix to quantify the feature distribution between time points, and the expression of the attention weight matrix is:
[0035] A i = softmax (M i )
[0036] Wherein, M i represents the cyclic matrix between the current time and the i-th time, and A i represents the attention weight between the current time and the i-th time.
[0037] The feature sequence D is weighted and summed according to the attention weight matrix to obtain an enhanced cross-time domain feature representation,
[0038]
[0039] Wherein, f j represents the time domain feature at the j-th time, A i,j represents the attention weight matrix at the i-th and j-th times, T represents the time length of the feature sequence, and f i represents the enhanced cross-time domain feature representation at the i-th time.
[0040] According to the attention weight matrix, the attention weights are sorted, and the time points with higher weights and the corresponding cross-time domain feature representations are screened out to extract highly correlated cross-time domain key information.
[0041] As a preferred scheme of the present application, a cross-time domain perception hierarchical aggregation module is established for the cross-time domain key information, the local feature information is aggregated on the corresponding global information according to the time domain, the detail information and edge information of the remote sensing image are obtained, including:
[0042] The cross-time domain feature representation at the i moment is taken as the input of the local feature, and the global feature at the current time point is input at the same time.
[0043] A cross-time domain perception module based on an attention mechanism is constructed, and the attention weight between the cross-time domain feature representation at the current moment and the global feature at different moments is calculated.
[0044] The global information at the historical moment is introduced into the local feature expression of the current frame through a weighted fusion function, and an enhanced local feature representation with time sequence memory capability is obtained.
[0045] The enhanced local feature representation is input into the convolution layer of the feature semantic extraction module, a U-Net hierarchical aggregation architecture is adopted, the enhanced local feature representation is sampled layer by layer, and the global feature at the corresponding level is spliced.
[0046] An attention gate module is embedded in the convolution layer of the feature semantic extraction module, information sensitive to details and edges is screened and strengthened, and corresponding detail information and edge information are obtained.
[0047] As a preferred scheme of the present application, the detail information and edge information of the remote sensing image are fused according to the perception hierarchical level by using an asymptotic fusion strategy, global fusion data is obtained, including:
[0048] A multi-level perception feature map is constructed based on a deep convolutional neural network according to the detail information and edge information, multi-level feature maps are extracted, and a perception hierarchical level from shallow to deep is formed.
[0049] An attention weighted fusion module is introduced on each layer of the perception hierarchical level, the detail information and edge information extracted at the current level are fused with the global feature at the corresponding level, and the output after each layer fusion is taken as the input of the next layer fusion, and layer-by-layer enhanced fusion data is obtained.
[0050] An edge-guided fusion strategy is adopted for the fusion data, the edge and texture information of the image are fused, and the local details and edge features of the remote sensing image are obtained.
[0051] The fusion proportion is adjusted in real time on the perception hierarchical level, and the global fusion data is obtained after multi-level gradual fusion.
[0052] As a preferred scheme of the present application, a loss function is introduced for the global fusion data to narrow the semantic gap, semantic fusion information is obtained, including:
[0053] The cross-entropy loss function is used to calculate the correlation of the corresponding semantics on each layer of the perception hierarchical level for the global fusion data, and the expression is:
[0054]
[0055] Wherein, β1 and β2 represent the weight coefficients on the main perception hierarchical level and the secondary perception hierarchical level respectively, The weight parameter value is represented by β1 and β2, The true value of the fusion data mapped to the main perception hierarchical level and the secondary perception hierarchical level is represented by L CE The correlation between images contained in the attention weight matrix is represented by β1 and β2.
[0056] As a preferred scheme of the present application, the interactive network model using mixed contrast learning is used to repair the missing information in the semantic fusion information, and complete real-time remote sensing image multi-source heterogeneous fusion data is obtained, including:
[0057] An interactive neural network structure containing a main branch and an auxiliary branch is constructed, wherein the main branch is used to extract the deep representation of the current semantic fusion information, and the auxiliary branch is used to extract the global features contained in the cross-time domain key information;
[0058] A mixed contrast learning strategy is introduced between the main branch and the auxiliary branch, and the remote sensing image data at the same time is subjected to contrast learning, and the sample similarity is quantified by maximizing the loss function;
[0059] The cross-attention mechanism is used to extract the potential information of each layer of the perception hierarchical level at the current time which can be used to repair the missing area;
[0060] Based on the contrast learning, the semantic consistency representation is obtained, and the content reconstruction is performed on the missing or low-confidence area in the semantic fusion graph combined with the context information;
[0061] The repaired semantic fusion feature map is outputted, and is mapped back to the pixel space to form complete remote sensing image fusion data.
[0062] In a second aspect of the present application, a remote sensing image multi-source heterogeneous data fusion processing system is provided, comprising:
[0063] A data acquisition module is used to acquire multi-source remote sensing image data from different sensor types, different time sequences and different spatial resolutions;
[0064] A feature extraction module is connected with the data acquisition module, and uses a convolutional neural network as an encoder to process the time-continuous remote sensing image data and extract the global feature information of the image;
[0065] The local feature enhancement module is connected with the feature extraction module, extracts a region of interest through an image cropping operation, and captures local details and edge information in the remote sensing image by using a lightweight multi-scale pyramid model combined with a feature semantic extraction module;
[0066] The cross-time-domain modeling module is connected with the local feature enhancement module, constructs a cross-time-domain attention mechanism based on a circulant matrix, extracts the mutual relationship between different time phases and different modal variables, and screens highly correlated cross-time-domain key information;
[0067] The hierarchical aggregation module is connected with the cross-time-domain modeling module, establishes a cross-time-domain perception hierarchical aggregation structure, and fuses local feature information to corresponding global features layer by layer according to time domain information, so as to enhance the detail expression capability of the remote sensing image.
[0068] The asymptotic fusion module is connected with the hierarchical aggregation module, adopts an asymptotic fusion strategy, and performs feature fusion at each perception hierarchical level to generate global fusion data with high spatial resolution and semantic consistency.
[0069] The semantic repair module is connected with the asymptotic fusion module, introduces an interactive network model of mixed contrast learning, and intelligently repairs missing or abnormal areas in semantic fusion information.
[0070] The loss optimization module is connected with the semantic repair module, optimizes and trains the global fusion data by designing a multi-task loss function, and obtains efficient fusion data.
[0071] Compared with the prior art, the present application has the following beneficial effects:
[0072] The present application extracts high-level semantic global features of remote sensing images by the feature extraction module combined with a convolutional neural network and an attention mechanism, captures local details and edge information by the local feature enhancement module through image cropping operation and a lightweight multi-scale pyramid model, realizes effective perception of the multi-level structure of remote sensing images, significantly improves the spatial resolution and semantic clarity of the fusion result, establishes a cross-time-domain perception hierarchical aggregation structure by the hierarchical aggregation module, fuses local features to corresponding global features layer by layer according to time sequence information, further strengthens the information complementarity of remote sensing images in time and spatial dimensions, thereby more accurately restores the details and edge information of the image, avoids the information redundancy or conflict problem that may occur in the traditional method, improves the controllability of the fusion process and the accuracy of the result, and at the same time, takes into account the calculation efficiency, builds a "global-local-time-semantic" collaborative deep fusion framework, and realizes efficient and high-precision fusion processing of multi-source remote sensing data. BRIEF DESCRIPTION OF DRAWINGS
[0073] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings needed to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those skilled in the art, other drawings can also be obtained from the provided drawings without creative labor.
[0074] Figure 1 A remote sensing image multi-source heterogeneous data fusion processing method flow chart is provided for the embodiments of the present application.
[0075] Figure 2 A remote sensing image multi-source heterogeneous data fusion processing system block diagram is provided for the embodiments of the present application. DETAILED DESCRIPTION
[0076] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0077] As shown in Figure 1 and Figure 2 , the present application provides a remote sensing image multi-source heterogeneous data fusion processing method, comprising the following steps:
[0078] The convolutional neural network is used as an encoder to extract global information of the time-continuous remote sensing image data, the image local features are extracted through the cropping operation, and the lightweight multi-scale pyramid model is combined with the feature semantic extraction module to capture the local feature information in the remote sensing image data.
[0079] In the present embodiment, the convolutional neural network is used as an encoder to extract the global information of the remote sensing image, and the local area features are obtained by combining the image cropping operation, and then the lightweight multi-scale pyramid model is combined with the feature semantic extraction module to capture the local feature information, which realizes the multi-level perception of the detail structure, edge information and semantic content in the remote sensing image, and significantly enhances the data expression ability and fusion precision.
[0080] The cross-time domain attention mechanism is constructed by a cyclic matrix to extract the mutual information between different modal variables, and the highly correlated cross-time domain key information is screened.
[0081] In this embodiment, a cross-time domain attention mechanism is constructed using a circulant matrix to effectively model the dynamic change relationship between different time points, extract mutual information between different modal variables, and filter out highly relevant cross-time domain key information, thereby improving the modeling capability and stability of the model for time series remote sensing data, and being suitable for processing multi-source heterogeneous data with time continuity.
[0082] A cross-time domain perception hierarchical aggregation module is established for the cross-time domain key information, and the local feature information is aggregated on the corresponding global information according to the time domain to obtain the detail information and edge information of the remote sensing image.
[0083] In this embodiment, by establishing a cross-time domain perception hierarchical aggregation module, the local features are aggregated layer by layer to the corresponding global features according to the time sequence information, realizing information fusion and complementarity between time dimension and space dimension, and further improving the detail restoration capability and edge clarity of the remote sensing image in the time evolution process.
[0084] The detail information and edge information of the remote sensing image are fused according to the perception hierarchical level by using an asymptotic fusion strategy, global fusion data is obtained, a loss function is introduced for the global fusion data to narrow the semantic gap, and semantic fusion information is obtained.
[0085] In this embodiment, an asymptotic fusion strategy based on perception level is used to sequentially fuse the detail and edge information according to the hierarchical structure, avoiding the information conflict or redundancy problem that may occur in traditional fusion methods, improving the controllability and robustness of the fusion result, and reducing the computational complexity and improving the algorithm efficiency.
[0086] In this embodiment, a loss function is used to optimize the fusion process, effectively narrowing the semantic gap between different modal data, and enhancing the consistency and explainability of the fusion result at the high-level semantic level.
[0087] A hybrid contrast learning interactive network model is used to repair the missing information in the semantic fusion information, and complete real-time remote sensing image multi-source heterogeneous fusion data is obtained.
[0088] In this embodiment, a hybrid contrast learning interactive network model is used to intelligently repair the missing part in the semantic fusion information, which not only improves the integrity and accuracy of the fusion data, but also enhances the adaptive ability and real-time response performance of the system, and is suitable for remote sensing monitoring and decision support in complex environments.
[0089] Convolutional neural networks are used as encoders to extract global information from time-continuous remote sensing image data, including:
[0090] The multi-source remote sensing image data is sequentially input into a deep convolutional neural network in time sequence, a multi-level feature extraction module is constructed with a convolutional layer as the backbone, and the spatial feature expression of the image is gradually extracted through stacking of the convolutional layer and the pooling layer;
[0091] In this embodiment, by sequentially inputting the multi-source remote sensing image data into the deep convolutional neural network in time sequence, and combining the stacking structure of the convolutional layer and the pooling layer, the spatial feature expression of the image can be extracted layer by layer while retaining the time evolution trend, thereby providing high-dimensional semantic representation containing dynamic change information for subsequent fusion processing.
[0092] The attention mechanism is used to capture the semantic information of the remote sensing image based on the spatial feature expression, each remote sensing image is mapped into a high-dimensional feature vector, the key spatial distribution rule and the time evolution trend are retained, and a high-level feature vector is output;
[0093] In this embodiment, the attention mechanism is introduced based on the spatial feature extraction, which can effectively filter out the key areas with significant semantic meaning in the remote sensing image, improve the recognition ability of the model to complex ground objects, and finally map each image into a high-quality high-dimensional feature vector, which has stronger discriminability and fusibility.
[0094] The global feature vector of the high-level feature vector is extracted, and the global feature vector is used as global information in the multi-source heterogeneous data fusion stage.
[0095] In this embodiment, the high-level feature vector not only contains the key spatial distribution rule of the original image, but also fuses the semantic consistency and change information between multiple time phases to form a unified and compact global feature representation.
[0096] In this embodiment, the global feature vector has good generalization ability and modal adaptability, and can establish semantic association between data of different sensors, different resolutions or different imaging modes, thereby improving the consistency and compatibility of multi-source heterogeneous remote sensing images in the fusion process.
[0097] The image local feature is extracted through the cropping operation, and a multi-scale hierarchical feature aggregation module is used to obtain a multi-scale hierarchical aggregation feature, including:
[0098] The multi-scale hierarchical feature aggregation module is used to segment the local information based on the global information of the convolutional layer, the multi-scale hierarchical feature aggregation module is provided with K channels, each channel is provided with a corresponding inflation rate r n ;
[0099] In this embodiment, by cropping the original image or the global feature map, focusing on the key area, effectively extracting the local features of the image, enhancing the model's perception of small targets, edge structures and complex ground object types in remote sensing images, and improving the spatial resolution and semantic clarity of the fusion result.
[0100] The number of scales of the multi-scale hierarchical feature aggregation module is set to M, and the size of each channel is set to M / 6;
[0101] The size of the convolution layer is constrained by the pyramid model, and three 4*4 convolution blocks are connected to aggregate local features of the image;
[0102] In this embodiment, a pyramid model is used to construct a multi-scale convolution block, three convolution modules of different scales are connected in cascade to form a feature expression structure from coarse to fine, realizing information complementation and fusion between different levels of features, and significantly enhancing the model's recognition and expression ability for multi-scale targets in remote sensing images.
[0103] The feature information of each feature pyramid channel of the pyramid model is added successively to obtain multi-scale hierarchical aggregation features.
[0104] In this embodiment, by successively adding and fusing the feature information output by each level of the pyramid, the key structural information at different scales is retained, the information loss problem caused by single-scale feature extraction is avoided, and the consistency and anti-interference ability of feature representation are improved.
[0105] The multi-scale hierarchical aggregation features are combined with the feature semantic extraction module to capture local feature information in remote sensing image data, including:
[0106] The multi-scale hierarchical aggregation features are obtained by the feature semantic extraction module, and the feature semantic extraction module is composed of three different convolutions of point-by-point convolution, ordinary convolution and dilated convolution. The point-by-point convolution, ordinary convolution and dilated convolution are respectively convolved with the multi-scale hierarchical aggregation features as the time domain reference.
[0107] In this embodiment, the feature semantic extraction module uses three different structures of point-by-point convolution, ordinary convolution and dilated convolution to process multi-scale hierarchical aggregation features in parallel, which can capture multi-level semantic information such as channel compression, spatial structure and large-scale context, and improve the recognition accuracy and generalization ability of the model for complex ground object targets in remote sensing images.
[0108] Each convolutional feature is output through a normalization layer to obtain an aggregated cascaded feature with the same target dimension. The aggregated cascaded feature is decoded by a multi-scale attention mechanism to output local feature information, and its expression is:
[0109]
[0110] wherein F represents an encoding feature, respectively represent the Relu and Sigmoid activation function weights of the convolution layer, F forcos represents an aggregated cascade feature, F s represents local feature information output after decoding by a multi-scale attention mechanism, BN{PointwiseConv(F), BN{OrdinaryConv(F), BN{DilationConv(F) respectively correspond to the convolution operations of pointwise convolution, ordinary convolution and dilated convolution, and Concat() represents a feature aggregation function.
[0111] In this embodiment, by normalizing the outputs after the three convolution operations and concatenating them into an aggregated cascade feature of uniform dimension, the consistency and stability of feature representation are enhanced, high-quality input is provided for subsequent attention decoding, and the dimension mismatch problem in the multi-scale feature fusion process is effectively avoided.
[0112] The local feature information is extracted through a cross-time domain attention mechanism constructed by a circulant matrix to extract mutual information between different modal variables, and to filter highly correlated cross-time domain key information, including:
[0113] The local feature information extracted at multiple time points is organized in time sequence as a feature sequence D={d1, d2, …, d t}, d t represents the local feature map at the t-th time step;
[0114] A circulant matrix M is constructed according to the feature sequence D to construct the mutual relationship between features at different times, and the expression of the circulant matrix M is:
[0115] M=similarity(d i ,d j )
[0116] Wherein the similarity measure of the circulant matrix M is obtained by cosine similarity calculation, d i , d j respectively represent the local feature maps at the i-th and j-th time steps;
[0117] In this embodiment, the local feature maps at multiple time points are organized in sequence as a feature sequence, and a circulant matrix based on cosine similarity is constructed, which can effectively model the evolution law of remote sensing images in the time dimension and reveal the potential semantic association between different time points.
[0118] The attention weight matrix is constructed by the cyclic matrix, and the feature distribution between each time point is quantified, and the expression of the attention weight matrix is:
[0119] A i = softmax(M i )
[0120] Wherein, M i represents the cyclic matrix between the current time and the i-th time, A i represents the attention weight between the current time and the i-th time;
[0121] In this embodiment, the attention weight matrix is constructed based on the cyclic matrix, the feature distribution difference between each time point is quantified, and the enhanced cross-time domain feature representation is further generated by weighted summation, which not only preserves the continuity of the time sequence information, but also highlights the contribution of the high correlation time, and significantly improves the time perception ability of the model.
[0122] According to the attention weight matrix, the feature sequence D is weighted and summed to obtain an enhanced cross-time domain feature representation,
[0123]
[0124] Wherein, f j represents the time domain feature at j time, A i,j represents the attention weight matrix at i, j time, T represents the time length of the feature sequence, f i represents the enhanced cross-time domain feature representation at i time;
[0125] According to the attention weight matrix, the attention weight is sorted, and the time point with high weight and the corresponding cross-time domain feature representation are screened out, and the highly correlated cross-time domain key information is extracted.
[0126] In this embodiment, by sorting the attention weight, the time point with high weight and the corresponding cross-time domain feature representation are screened out, which can quickly locate the multi-modal variable information highly related to the current time, reduce redundant calculation, and improve the response speed and accuracy of the fusion system.
[0127] The cross-time domain perception hierarchical aggregation module is established for the cross-time domain key information, and the local feature information is aggregated on the corresponding global information according to the time domain, so as to obtain the detail information and edge information of the remote sensing image, including:
[0128] The cross-time domain feature representation at i time is input as the input of the local feature, and the global feature at the current time point is also input;
[0129] a cross-time domain perception module based on an attention mechanism is constructed to calculate attention weights between the cross-time domain feature representation at the current moment and global features at different moments;
[0130] In this embodiment, the attention weights between the current moment and global features at different moments are calculated by constructing a cross-time domain perception module based on an attention mechanism, and the global information at the historical moment is introduced into the local feature expression of the current frame by using a weighted fusion function, thereby realizing the memory function of the time sequence data. This method effectively captures the dynamic change law in the time dimension and improves the understanding ability of the model for complex space-time scenes.
[0131] The global information at the historical moment is introduced into the local feature expression of the current frame by using a weighted fusion function, and an enhanced local feature representation with time sequence memory ability is obtained;
[0132] The enhanced local feature representation is input into the convolution layer of the feature semantic extraction module, a U-Net hierarchical aggregation architecture is adopted to sample the enhanced local feature representation layer by layer and splice it with the global features at the corresponding level;
[0133] In this embodiment, under the U-Net hierarchical aggregation architecture, the enhanced local feature representation is sampled layer by layer and spliced with the global features at the corresponding level, realizing the effective combination of local details and global context information. This hierarchical feature fusion method not only retains rich spatial structure information, but also enhances the understanding and adaptability of the model to large-scale environment background.
[0134] An attention gate module is embedded in the convolution layer of the feature semantic extraction module to filter and strengthen information sensitive to details and edges; and corresponding detail information and edge information is obtained.
[0135] In this embodiment, the attention gate module embedded in the convolution layer of the feature semantic extraction module can intelligently filter and strengthen information sensitive to details and edges, which enables the model to significantly improve the recognition accuracy of subtle structures in remote sensing images while maintaining the consistency of overall semantics.
[0136] The detail information and edge information of the remote sensing image are fused according to the perception hierarchical levels by using an asymptotic fusion strategy to obtain global fusion data, including:
[0137] A multi-level perception feature map is constructed based on a deep convolutional neural network according to the detail information and edge information to extract multi-level feature maps and form perception hierarchical levels from shallow to deep;
[0138] In this embodiment, by extracting multi-level feature maps based on a deep convolutional neural network and forming a perception hierarchical level from shallow to deep, the complete information chain from low-level edge texture to high-level semantic objects in remote sensing images can be effectively captured.
[0139] An attention weighted fusion module is introduced on each layer of the perception hierarchical level, the detail information and edge information extracted at the current level are fused with the global features of the corresponding level, and the output after fusion at each layer is taken as the input for fusion of the next layer to obtain fusion data enhanced layer by layer;
[0140] In this embodiment, an attention weighted fusion module is introduced on each layer of the perception hierarchical level, the detail and edge information of the current level are adaptively fused with the global features of the corresponding level, which not only enhances the information retention degree of the key region, but also effectively suppresses the interference of redundant or conflicting information, significantly improving the accuracy and consistency of the fusion result.
[0141] An edge guided fusion strategy is adopted for the fusion data to fuse the edge and texture information of the image to obtain the local details and edge features of the remote sensing image.
[0142] The fusion ratio is adjusted in real time on the perception hierarchical level, and after multi-level progressive fusion, global fusion data is obtained.
[0143] In this embodiment, the output after fusion at each layer is taken as the input for fusion of the next layer, realizing the step-by-step optimization and enhancement of information between different levels, forming a good feature transmission and accumulation effect, and making the final fusion result have stronger structural integrity and semantic coherence.
[0144] A loss function is introduced for the global fusion data to narrow the semantic gap and obtain semantic fusion information, including:
[0145] A cross-entropy loss function is used to calculate the relevance of the corresponding semantics on each layer of the perception hierarchical level for the global fusion data, and its expression is:
[0146]
[0147] wherein β1 and β2 represent the weight coefficients on the primary perception hierarchical level and the secondary perception hierarchical level, respectively, represents the weight parameter value, respectively represent the true values of the fusion data mapped to the primary perception hierarchical level and the secondary perception hierarchical level, and L CE represents the correlation between images contained in the attention weight matrix.
[0148] In this embodiment, by introducing a cross-entropy loss function, the semantic correlation between the primary level and the secondary level is calculated on each layer of the perception hierarchical level, which can effectively quantify the semantic deviation between different modalities or remote sensing data at different time points, thereby guiding the model to learn more consistent high-level semantic representation, significantly improving the fusion quality of multi-source heterogeneous data.
[0149] The interactive network model using hybrid contrast learning repairs missing information in the semantic fusion information to obtain complete real-time remote sensing image multi-source heterogeneous fusion data, including:
[0150] An interactive neural network structure containing a main branch and an auxiliary branch is constructed, wherein the main branch is used to extract deep representations of the current semantic fusion information, and the auxiliary branch is used to extract global features contained in the cross-time domain key information;
[0151] In this embodiment, by designing an interactive neural network architecture containing a main branch and an auxiliary branch, deep representations of the current semantic fusion information and global features of the cross-time domain key information are extracted respectively, joint modeling of spatial semantics and temporal evolution information is realized, and the understanding ability of dynamic changes and static structures in remote sensing images is enhanced.
[0152] A hybrid contrast learning strategy is introduced between the main branch and the auxiliary branch, and remote sensing image data at the same time is subjected to contrast learning, and the similarity of the quantized samples is maximized by a loss function;
[0153] In this embodiment, a hybrid contrast learning mechanism is introduced between the main branch and the auxiliary branch, and the similarity between samples of remote sensing images at the same time and from different perspectives or modalities is maximized, which can effectively enhance the semantic consistency and discriminability of feature representation, and significantly improve the robustness of the model to noise, occlusion and other interference factors.
[0154] Cross-attention mechanism is used to extract potential information on each layer of the perception hierarchical level at the current time that can be used to repair missing areas;
[0155] In this embodiment, cross-attention mechanism is used to extract context potential information on the perception hierarchical level that can be used to repair missing areas, accurately locate features related to the semantics of missing areas, and thus realize high-quality content reconstruction and improve the detail integrity and visual quality of remote sensing images.
[0156] Based on contrast learning, a semantic consistent representation is obtained, and a missing or low-confidence area in the semantic fusion graph is reconstructed based on context information;
[0157] The repaired semantic fusion feature map is output and mapped back to the pixel space to form complete remote sensing image fusion data.
[0158] In this embodiment, the repaired semantic fusion feature map has higher integrity and accuracy, and can be further mapped back to the pixel space to generate complete remote sensing image fusion data, so that the remote sensing image fusion data achieves a high level in terms of structural integrity, semantic consistency and visual quality.
[0159] A second embodiment: a remote sensing image multi-source heterogeneous data fusion processing system, comprising:
[0160] A data acquisition module is configured to acquire multi-source remote sensing image data from different sensor types, different time sequences and different spatial resolutions.
[0161] A feature extraction module is connected to the data acquisition module and uses a convolutional neural network as an encoder to process time-continuous remote sensing image data and extract global feature information of the image.
[0162] A local feature enhancement module is connected to the feature extraction module, extracts the region of interest through image cropping operation, and captures the local details and edge information in the remote sensing image by using a lightweight multi-scale pyramid model combined with a feature semantic extraction module.
[0163] A cross-time-domain modeling module is connected to the local feature enhancement module, constructs a cross-time-domain attention mechanism based on a circulant matrix, extracts the mutual relationship between different time phases and different modal variables, and filters the highly correlated cross-time-domain key information.
[0164] A hierarchical aggregation module is connected to the cross-time-domain modeling module, establishes a cross-time-domain perception hierarchical aggregation structure, and fuses the local feature information into the corresponding global feature layer by layer according to the time domain information to enhance the detail expression ability of the remote sensing image.
[0165] An asymptotic fusion module is connected to the hierarchical aggregation module, adopts an asymptotic fusion strategy, and performs feature fusion at each perception hierarchical level to generate global fusion data with high spatial resolution and semantic consistency.
[0166] A semantic repair module is connected to the asymptotic fusion module, introduces an interactive network model of hybrid contrast learning, and intelligently repairs the missing or abnormal areas in the semantic fusion information.
[0167] A loss optimization module is connected to the semantic repair module, optimizes and trains the global fusion data by designing a multi-task loss function, and obtains efficient fusion data.
[0168] In the embodiment, the whole system has the ability to comprehensively support multi-source heterogeneous data input and collaborative modeling, significantly improves the information expression ability through the global-local collaborative feature extraction mechanism, combines the cross-time domain attention modeling and the hierarchical aggregation structure to strengthen the spatiotemporal information complementarity, adopts the asymptotic fusion strategy to improve the fusion accuracy and efficiency, avoids the information redundancy or conflict problem that may occur in the traditional method, and improves the controllability of the fusion process and the accuracy of the result.
[0169] The application extracts the high-level semantic global features of the remote sensing image through the feature extraction module combined with the convolutional neural network and the attention mechanism, the local feature enhancement module captures the local details and edge information through the image clipping operation and the lightweight multi-scale pyramid model, realizes the effective perception of the multi-level structure of the remote sensing image, significantly improves the spatial resolution and semantic clarity of the fusion result, adopts the hierarchical aggregation module to establish the cross-time domain perception hierarchical aggregation structure, and further strengthens the information complementarity of the remote sensing image in the time and space dimensions, so that the details and edge information of the image are more accurately restored, the information redundancy or conflict problem that may occur in the traditional method is avoided, the controllability of the fusion process and the accuracy of the result are improved, and the calculation efficiency is considered at the same time. Through the construction of the "global-local-time-semantic" collaborative deep fusion framework, the efficient and high-precision fusion processing of the multi-source remote sensing data is realized.
[0170] The above embodiments are only exemplary embodiments of the application and are not used to limit the application, and the protection scope of the application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements to the application within the spirit and protection scope of the application, and such modifications or equivalent replacements are also regarded as falling within the protection scope of the application.
Claims
1. A remote sensing image multi-source heterogeneous data fusion processing method, characterized in that: The following steps are involved: A convolutional neural network is used as an encoder to extract global information from time-continuous remote sensing image data, and local image features are extracted through cropping operations. A lightweight multi-scale pyramid model is used in combination with a feature semantic extraction module to capture local feature information in remote sensing image data. A cross-temporal attention mechanism is constructed for the local feature information through a circulant matrix to extract mutual information between different modal variables and screen highly correlated cross-temporal key information; Establishing a cross-temporal perception hierarchical aggregation module for the cross-temporal key information, aggregating the local feature information on the corresponding global information according to the time domain, and obtaining the detail information and edge information of the remote sensing image; The detail information and edge information of the remote sensing image are fused respectively according to the perception hierarchical level using a gradual fusion strategy to obtain global fusion data, and a loss function is introduced into the global fusion data to narrow the semantic gap and obtain semantic fusion information; An interactive network model of hybrid contrastive learning is used to repair missing information in the semantic fusion information and obtain complete real-time remote sensing image multi-source heterogeneous fusion data.
2. The method for fusion processing of multi-source heterogeneous remote sensing image data according to claim 1, characterized in that: A convolutional neural network is used as an encoder to extract global information from time-continuous remote sensing image data, including: Multi-source remote sensing image data are input into the deep convolutional neural network in chronological order. With the convolution layer as the backbone, a multi-level feature extraction module is constructed. By stacking convolution layers and pooling layers, the spatial feature expression of the image is gradually extracted. The attention mechanism is used to capture the semantic information of the remote sensing image for the spatial feature expression, and each remote sensing image is mapped into a high-dimensional feature vector, retaining its key spatial distribution law and temporal evolution trend, and outputting a high-level feature vector; A global feature vector of the high-level feature vector is extracted, and the global feature vector is used as global information in a multi-source heterogeneous data fusion stage.
3. A remote sensing image multi-source heterogeneous data fusion processing method according to claim 2, characterized in that: Extracting local features of the image through a cropping operation, and applying a multi-scale hierarchical feature aggregation module to the local features of the image to obtain multi-scale hierarchical aggregation features, including: In the convolution layer, the global information is segmented by a multi-scale hierarchical feature aggregation module for local information. The multi-scale hierarchical feature aggregation module is provided with K channels, and each channel is provided with a corresponding expansion rate r n ; The number of scales of the multi-scale hierarchical feature aggregation module is set to M, and the size of each channel is set to 1M / 6; A pyramid model is used to constrain the size of the convolutional layer, and three 4*4 convolutional blocks are connected to aggregate local image features; The feature information of each feature pyramid channel of the pyramid model is successively added through hierarchical feature aggregation to obtain multi-scale hierarchical aggregation features.
4. The method for fusion processing of multi-source heterogeneous remote sensing image data according to claim 3, characterized in that: The multi-scale hierarchical aggregation features are combined with a feature semantic extraction module to capture local feature information in remote sensing image data, including: The multi-scale hierarchical aggregation features are passed through a feature semantic extraction module to obtain local feature information. The feature semantic extraction module is composed of three different convolutions: point-by-point convolution, ordinary convolution, and dilated convolution. The point-by-point convolution, ordinary convolution, and dilated convolution are used as the reference in the time domain to convolve the multi-scale hierarchical aggregation features respectively; Each convolution feature is passed through the normalization layer to output the aggregated cascade feature with the same target dimension. The aggregated cascade feature is decoded through the multi-scale attention mechanism to output the local feature information, which is expressed as: Among them, F represents the encoding feature, Represents the Relu and Sigmoid activation function weights of the convolutional layer, F forcos represents the aggregated cascade feature, F s It represents the local feature information output after decoding by the multi-scale attention mechanism. BN{PointwiseConv(F), BN{OrdinaryConv(F), BN{DilationConv(F) correspond to the convolution operations of pointwise convolution, ordinary convolution and dilated convolution respectively. Concat() represents the feature aggregation function.
5. The method for fusion processing of multi-source heterogeneous remote sensing image data according to claim 4, characterized in that: The local feature information is constructed through a circulant matrix to construct a cross-temporal attention mechanism to extract the mutual information between different modal variables and screen highly correlated cross-temporal key information, including: The local feature information extracted at multiple time points is organized into a feature sequence D = {d1, d2, ..., d t }, d t Represents the local feature map of the t-th time step; The circulant matrix M is constructed based on the feature sequence D to build the relationship between features at different times. The expression of the circulant matrix M is: M=similarity(d i ,d j ) Among them, the similarity measure of the circulant matrix M is calculated using cosine similarity, d i d j Represent the local feature maps of the i-th and j-th time steps respectively; The attention weight matrix is constructed by the circulant matrix to quantify the feature distribution between each time point. The expression of the attention weight matrix is: A i =softmax(M i ) Among them, M i Represents the circulant matrix between the current moment and the i-th moment, A i Represents the attention weight between the current moment and the i-th moment; The feature sequence D is weighted and summed according to the attention weight matrix to obtain an enhanced cross-temporal feature representation. Among them, f j represents the time domain characteristics at time j, A i,j represents the attention weight matrix at time i and j, T represents the time length of the feature sequence, and f i represents the cross-temporal feature representation at time i after enhancement; The attention weights are sorted according to the attention weight matrix, and the time points with higher weights and the corresponding cross-time domain feature representations are screened out to extract highly correlated cross-time domain key information.
6. A remote sensing image multi-source heterogeneous data fusion processing method according to claim 5, characterized in that: A cross-temporal perception hierarchical aggregation module is established for the cross-temporal key information, and the local feature information is aggregated on the corresponding global information according to the time domain to obtain the detail information and edge information of the remote sensing image, including: The cross-temporal feature representation at time i is used as the input of the local feature, and the global feature at the current time point is input at the same time; Construct a cross-temporal perception module based on the attention mechanism to calculate the attention weight between the cross-temporal feature representation at the current moment and the global features at different moments; The global information of historical moments is introduced into the local feature expression of the current frame through a weighted fusion function to obtain an enhanced local feature representation with temporal memory capability. Input the enhanced local feature representation into the convolution layer of the feature semantic extraction module, adopt the U-Net layered aggregation architecture to sample the enhanced local feature representation layer by layer, and splice it with the global features of the corresponding layer; An attention gating module is embedded in the convolutional layer of the feature semantic extraction module to filter and enhance information that is sensitive to details and edges; and corresponding detail information and edge information are obtained.
7. The method for fusion processing of multi-source heterogeneous remote sensing image data according to claim 6, characterized in that: The detail information and edge information of the remote sensing image are fused separately according to the perception hierarchical level using a gradual fusion strategy to obtain global fusion data, including: Constructing a multi-level perceptual feature map based on a deep convolutional neural network according to the detail information and edge information, extracting multi-level feature maps, and forming a perceptual hierarchical hierarchy from shallow to deep; An attention-weighted fusion module is introduced at each level of the perception hierarchy to fuse the detail information and edge information extracted at the current level with the global features of the corresponding level. The fused output of each level is used as the input of the next level to obtain layer-by-layer enhanced fusion data. An edge-guided fusion strategy is adopted for the fused data to fuse edge and texture information of the image and obtain local details and edge features of the remote sensing image; The fusion ratio is adjusted in real time at the perception hierarchical level, and global fusion data is obtained through multi-level progressive fusion.
8. The method for fusion processing of multi-source heterogeneous remote sensing image data according to claim 7, characterized in that: Introducing a loss function into the global fusion data to narrow the semantic gap and obtain semantic fusion information includes: The cross entropy loss function is used to calculate the relevance of the corresponding semantics at each layer of the perception hierarchy for the global fusion data, and its expression is: Among them, β1 and β2 represent the weight coefficients at the primary perception level and the secondary perception level, respectively. represents the weight parameter value, They represent the true values of the fusion data mapped to the primary perception level and the secondary perception level, respectively. CE Represents the correlation between images contained in the attention weight matrix.
9. The method for fusion processing of multi-source heterogeneous remote sensing image data according to claim 8, characterized in that: The hybrid contrastive learning interactive network model is used to repair the missing information in the semantic fusion information and obtain complete real-time remote sensing image multi-source heterogeneous fusion data, including: Constructing an interactive neural network structure comprising two main and auxiliary branches, wherein the main branch is used to extract the deep representation of the current semantic fusion information, and the auxiliary branch is used to extract the global features contained in the cross-temporal key information; A hybrid contrast learning strategy is introduced between the main and auxiliary branches to perform contrast learning on the remote sensing image data at the same moment, and to quantify the sample similarity by maximizing the loss function; A cross-attention mechanism is used to extract potential information that can be used to repair the missing areas at each level of the perception hierarchy at the current moment; It obtains semantic consistency representation based on contrastive learning and reconstructs the content of missing or low-confidence areas in the semantic fusion graph by combining contextual information. Output the restored semantic fusion feature map and map it back to the pixel space to form complete remote sensing image fusion data.
10. A remote sensing image multi-source heterogeneous data fusion processing system, characterized in that: include: Data acquisition module, used to acquire multi-source remote sensing image data from different sensor types, different time series and different spatial resolutions; A feature extraction module is connected to the data acquisition module and uses a convolutional neural network as an encoder to process the time-continuous remote sensing image data and extract the global feature information of the image; A local feature enhancement module is connected to the feature extraction module to extract the region of interest through image cropping and to capture local details and edge information in remote sensing images using a lightweight multi-scale pyramid model combined with a feature semantic extraction module; A cross-temporal modeling module is connected to the local feature enhancement module, and constructs a cross-temporal attention mechanism based on a circulant matrix to extract the relationship between variables of different phases and modalities, and screen highly correlated cross-temporal key information; A hierarchical aggregation module is connected to the cross-temporal modeling module to establish a cross-temporal perception hierarchical aggregation structure, which integrates local feature information into corresponding global features layer by layer based on temporal information to enhance the detail expression capability of remote sensing images; An asymptotic fusion module, connected to the hierarchical aggregation module, adopts an asymptotic fusion strategy to fuse features at each perception level to generate global fusion data with high spatial resolution and semantic consistency; A semantic restoration module is connected to the asymptotic fusion module and introduces an interactive network model of hybrid contrastive learning to intelligently repair missing or abnormal areas in the semantic fusion information; The loss optimization module is connected to the semantic restoration module and optimizes the training of the global fusion data by designing a multi-task loss function to obtain efficient fusion data.
Citation Information
Patent Citations
Multi-scale fusion remote sensing image semantic segmentation method and system
CN115512103A
Remote sensing image segmentation method based on channel enhancement and cross-level multi-input features
CN119380018A
Brain cognitive state layered recognition method based on space-time attention
CN119513566A
Multi-scale feature remote sensing image semantic segmentation method and system
CN120032126A
Landslide recognition method based on laplacian pyramid remote sensing image fusion
US11521377B1
Cited By
Remote sensing image production method and system
CN121305318A
Model optimization method, image processing method and system for remote sensing dynamic monitoring
CN122048700A
Confrontation environment global image reconstruction method for single-machine perception information
CN122289450A
An anti-environment global image reconstruction method for single-machine perception information
CN122289450B