High-resolution satellite data semantic analysis and intelligent identification system and method
By generating multi-scale adaptive grids and spatiotemporal correlation matrices, the instability problem of construction target identification in high-resolution satellite images was solved, and refined, spatiotemporally consistent semantic parsing of temporary construction areas was achieved, improving identification accuracy and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUNAN CHUANGXIN WEILI TECH CO LTD
- Filing Date
- 2026-01-19
- Publication Date
- 2026-05-08
AI Technical Summary
In temporary construction scenarios on urban roads, targets such as construction fences, material stacks, and small machinery in high-resolution satellite images have weak textures and unstable outlines, making them easy to confuse with shadows and road damage. They also have obvious stage-specific characteristics, making it difficult for existing methods to reliably identify the real construction area, resulting in high rates of misjudgment and missed judgment.
By analyzing the texture complexity, edge density, and saliency response of high-resolution satellite imagery, multi-scale adaptive grid cells are generated and spatial indexes are established. Multi-temporal images are loaded for multi-scale convolutional coding and gradient sequence extraction to construct a spatiotemporal correlation matrix, predict semantic change trends, and combine them with a cross-spatiotemporal stable reference frame for fusion and correction to generate a comprehensive confidence matrix to output the corrected semantic recognition results.
It achieves fine segmentation and spatiotemporally consistent semantic parsing of high-resolution satellite imagery, improves recognition accuracy, reduces the impact of noise and structural residuals, and obtains more accurate and stable semantic recognition results.
Smart Images

Figure CN121999495A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of high-resolution remote sensing image analysis based on artificial intelligence, specifically to a high-resolution satellite data semantic parsing and intelligent recognition system and method. Background Technology
[0002] In temporary construction scenarios on urban roads, targets such as construction barriers, material stacks, and small machinery typically occupy only a few pixels in high-resolution satellite images. Their textures are weak, their outlines unstable, and they are easily confused with shadows, reflective areas, and road damage, making it difficult for models to extract reliable features. Furthermore, temporary construction has distinct phases, with significant differences in the appearance of the same location across different time slices, leading to feature drift. Existing methods largely rely on single-phase images, failing to utilize temporal patterns, lacking the ability to analyze differences in neighborhood structures, and lacking mechanisms to handle inconsistencies in multi-phase features. This results in difficulty in consistently identifying genuine temporary construction areas, leading to high rates of false positives and false negatives. Therefore, designing a high-resolution satellite data semantic parsing and intelligent recognition system and method to improve recognition accuracy is essential. Summary of the Invention
[0003] To address the shortcomings of existing technologies, this invention provides a high-resolution satellite data semantic parsing and intelligent recognition system and method, which has the advantage of improving recognition accuracy and solves the problems mentioned in the background technology.
[0004] To achieve the aforementioned goal of improving recognition accuracy, this invention provides the following technical solution: a high-resolution satellite data semantic parsing and intelligent recognition method, comprising the following steps: Texture complexity, edge density, and saliency response analysis are performed on high-resolution satellite imagery. Dynamic scaling is performed based on regional differences to generate multi-scale adaptive grid cells and establish a spatial index. Based on multi-scale adaptive grids and spatial indexing, the corresponding multi-temporal images are loaded, and multi-scale convolutional coding, structural description, gradient sequence extraction and brightness difference calculation are performed on each grid to generate multi-scale spatiotemporal feature sequences. Based on the multi-scale spatiotemporal feature sequence and spatial index, the neighborhood grid set is determined, the spatiotemporal coupling relationship between grids is constructed, the spatiotemporal correlation matrix is generated and the semantic change trend of the main grid is predicted, the scale alignment and local perturbation suppression of multi-temporal features are performed, and a cross-spatiotemporal stable reference frame is constructed. By fusing the cross-temporal and spatial stable reference frame, semantic change trend, spatiotemporal correlation matrix, and the preliminary semantic judgment of the current frame, the confidence scores of structural consistency, temporal consistency, and neighborhood correlation are calculated to form a comprehensive confidence matrix. Based on the difference comparison between the comprehensive confidence matrix and historical recognition records, the weights of key features and matching parameters are adjusted, and the corrected grid semantic recognition result is output.
[0005] Preferably, the process of generating multi-scale adaptive mesh cells and establishing spatial indexes is as follows: The image is initially divided into blocks according to spatial location and resolution, and the texture entropy, orientation gradient statistics and edge sharpness index are calculated for each block. The importance of the segmented regions is quantitatively assessed by combining the saliency response map. Based on the spatial detail density, self-similarity and structural complexity within the regions, the images are divided into different scale levels. An adaptive grid generation strategy is adopted for each scale level to generate multi-scale adaptive grid cells, and the grid number, geographical range, level identifier and neighborhood relationship are recorded to establish a spatial index.
[0006] Preferably, the process of loading corresponding multi-temporal images based on multi-scale adaptive grids and spatial indexes is as follows: Based on the grid number and geographical extent recorded in the spatial index, image blocks corresponding to grid units of different scales are extracted from the multi-temporal image dataset; Perform geometric correction, brightness normalization, and cloud / fog removal on the image; Image blocks are organized and cached according to the scale hierarchy of a multi-scale adaptive grid.
[0007] Preferably, the process of generating multi-scale spatiotemporal feature sequences is as follows: Based on multi-temporal image blocks corresponding to grids at each scale, convolutional coding, structural texture description, directional gradient sequence extraction, and local brightness difference analysis are performed on the image blocks within each grid cell to obtain the primary feature set for each time node. The primary feature sets of the same grid cell at different time nodes are sorted by timestamp, and time synchronization, cross-scale alignment, outlier feature marking and feature sparsification are performed. Multi-scale spatiotemporal feature sequences are constructed based on feature stability, integrity, and temporal density.
[0008] Preferably, the process of generating the spatiotemporal correlation matrix and predicting the semantic change trend of the main grid is as follows: The spatial neighborhood set of the main grid is determined based on the spatial index, and the corresponding temporal neighborhood set is constructed based on the timestamp; Joint analysis is performed on the multi-scale spatiotemporal feature sequences of the main grid and its neighborhood to calculate feature similarity, directional gradient consistency and local variation coupling degree. A spatiotemporal correlation matrix is generated based on the coupling degree, and the semantic change trend of the main grid is predicted accordingly.
[0009] Preferably, the process of constructing a stable reference frame across time and space is as follows: Based on the spatiotemporal correlation matrix and the predicted semantic change trend of the main grid, scale alignment, local noise perturbation suppression, structural residual filtering and neighborhood consistency enhancement are performed on the multi-temporal feature sequences of each grid. By utilizing the spatial adjacency, temporal continuity, and cross-scale dependence represented by the spatiotemporal correlation matrix, the processed spatiotemporal features are jointly analyzed to extract structurally stable regions, continuous change patterns, and key change responses, and anomalous disturbance features are weighted and suppressed. The extracted stable structures and change patterns are integrated according to grid number, spatial location and scale level to generate a cross-temporal and spatiotemporal stable reference frame.
[0010] Preferably, the process of fusing the cross-temporal stable reference frame, semantic change trend, spatiotemporal correlation matrix, and preliminary semantic determination of the current frame is as follows: The initial semantic determination of the current frame is compared with the structural features of the cross-temporal stable reference frame on a grid-by-grid basis, and the spatial consistency, local texture matching degree and cross-scale response difference of each grid are calculated. Using the neighborhood coupling information provided by the spatiotemporal correlation matrix, neighborhood compensation and cross-temporal weighted correction are performed on meshes with structural deviations or local anomalies; Based on the predicted semantic change trend and compared with historical and neighborhood information, time series corrections are made for the initially identified possible drift areas to generate preliminary fusion estimates.
[0011] Preferably, the process of forming the comprehensive confidence matrix is as follows: For the fusion semantic estimation results of each grid, the structural consistency confidence is calculated, including the stability of features within the grid, texture matching degree, and cross-scale feature residuals. Based on the comparison of historical multi-temporal features and the prediction of semantic change trends in time series, the temporal consistency confidence is calculated, including feature drift, trend matching degree and abnormal change markers. Using the spatiotemporal correlation matrix and neighborhood grid feature information, the confidence of neighborhood correlation is calculated, including neighborhood similarity, local coupling degree and neighborhood deviation correction coefficient; According to the preset weighting strategy, the confidence scores of structural consistency, temporal consistency, and neighborhood correlation are fused to generate a comprehensive confidence matrix.
[0012] Preferably, the process of outputting the corrected grid semantic recognition result is as follows: The comprehensive confidence matrix is compared with the historical identification records grid by grid, and the confidence deviation of each grid in terms of structure, time and neighborhood dimensions is calculated. For grids with deviations exceeding a preset threshold, the weights of key features, including texture features, gradient direction, and brightness response, are adjusted based on confidence contribution and historical weights. Based on the adjusted feature weights and grid matching parameters, the grid semantic classification probability is recalculated, and neighborhood interpolation or weighted fusion is performed on uncertain grids to form the final output of the corrected grid semantic judgment result.
[0013] A high-resolution satellite data semantic parsing and intelligent recognition system includes: Adaptive Mesh Module: Performs texture, edge, and saliency analysis on satellite imagery and performs dynamic scaling to generate multi-scale adaptive meshes and spatial indexes; Spatiotemporal feature module: Loads multi-temporal images and extracts multi-scale spatiotemporal features such as convolutional coding, structural description, gradient sequence and brightness difference on each grid; The association modeling module determines the neighborhood grid based on spatiotemporal features and spatial indexes, constructs the spatiotemporal coupling relationship between grids and predicts semantic change trends, while generating a stable reference frame across spatiotemporal regions. Confidence fusion module: It fuses the stable reference frame, change trend, correlation matrix and current frame determination to calculate the comprehensive confidence of structural consistency, temporal consistency and neighborhood correlation; The adaptive correction module compares the overall confidence level with historical recognition records, adjusts the feature weights and matching parameters, and generates the corrected semantic recognition results.
[0014] Compared with existing technologies, the present invention provides a high-resolution satellite data semantic parsing and intelligent recognition system and method, which has the following beneficial effects: This invention achieves refined segmentation and systematic management of high-resolution satellite imagery by constructing a multi-scale adaptive grid and establishing a spatial index. This enables the accurate extraction and organization of features from different spatial regions and scales. By generating multi-scale spatiotemporal feature sequences from multi-temporal images, it can capture changes in ground features and dynamic patterns over time, improving the continuity and consistency of semantic analysis. Through neighborhood grid set and spatiotemporal correlation matrix analysis, it can predict the semantic change trend of the main grid. Furthermore, by combining a cross-spatiotemporal stable reference frame to perform scale alignment and perturbation suppression on multi-temporal features, it helps to reduce the impact of local noise and structural residuals. In addition, by fusing preliminary judgment results, spatiotemporal correlation information, and predicted trends, it calculates a comprehensive confidence matrix and compares it with historical data, dynamically correcting the weights of key features and matching parameters, thereby obtaining more accurate and stable grid semantic recognition results. This achieves refined, spatiotemporally consistent, and continuously reliable semantic analysis capabilities for high-resolution remote sensing images. Attached Figure Description
[0015] Figure 1This is a schematic diagram of the method of the present invention; Figure 2 This is a schematic diagram of the structure of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] Example 1: Please refer to Figure 1 As shown in the embodiment of the present invention, a high-resolution satellite data semantic parsing and intelligent recognition method includes the following steps: S1: Perform texture complexity, edge density, and saliency response analysis on high-resolution satellite imagery, perform dynamic scaling based on regional differences, generate multi-scale adaptive grid cells, and establish a spatial index.
[0018] The process of generating multi-scale adaptive mesh elements and establishing spatial indexes in S1 is as follows: The image is initially divided into blocks according to spatial location and resolution, and the texture entropy, orientation gradient statistics and edge sharpness index are calculated for each block. The input high-resolution satellite imagery is divided into initial blocks according to a preset spatial grid. Each block corresponds to a certain geographic coordinate range. The block size is automatically determined based on the image resolution. After grayscale or multi-band fusion processing is performed on each block, the local texture entropy is calculated using a sliding window. The gradient intensity in each direction is statistically analyzed and a histogram of directional gradients is generated. At the same time, the Canny operator is applied to calculate the edge intensity, and the sharpness index is calculated on the edge response map. Finally, each block will obtain a set of numerical indicators, including texture entropy, gradient statistics in each direction, and edge sharpness.
[0019] The importance of the segmented regions is quantitatively assessed by combining the saliency response map. Based on the spatial detail density, self-similarity and structural complexity within the regions, the images are divided into different scale levels. Texture, gradient, and edge metrics are mapped to saliency response maps. An importance score for each block is generated using a weighted overlay or normalization method. The spatial detail density within each block is calculated. The variance and gradient changes of different pixel values are statistically analyzed through local windows to quantify structural complexity. Simultaneously, the repeatability of local texture patterns is evaluated using a self-similarity method. Based on the scores and structural metrics, the image is divided into several scale levels, each corresponding to a specific block size and spatial resolution.
[0020] An adaptive grid generation strategy is adopted for each scale level to generate multi-scale adaptive grid cells, and the grid number, geographical range, level identifier and neighborhood relationship are recorded to establish a spatial index; Within each scale level, the grid density is dynamically adjusted based on the importance and spatial complexity of the blocks. High-importance areas are divided into smaller grids to increase detail resolution, while low-complexity areas are divided into larger grids to reduce computation. Each grid is assigned a unique number and its corresponding geographic coordinates are recorded, indicating its scale level. At the same time, the neighboring grid relationships are determined according to the up, down, left, right and diagonal directions.
[0021] S2: Based on multi-scale adaptive grids and spatial indexing, load the corresponding multi-temporal images, perform multi-scale convolutional coding, structural description, gradient sequence extraction and brightness difference calculation on each grid, and generate multi-scale spatiotemporal feature sequences.
[0022] The process of loading multi-temporal images based on multi-scale adaptive grids and spatial indexes in S2 is as follows: Based on the grid number and geographical extent recorded in the spatial index, image blocks corresponding to grid units of different scales are extracted from the multi-temporal image dataset; The number and corresponding geographic coordinate range of each grid cell are obtained through spatial indexing. The coordinates are matched with the image georeference information to extract the corresponding area image block from the stored multi-temporal images. The image data of each time node is geographically mapped, the coordinate range is converted into a pixel index, and the rectangular image block corresponding to the grid is obtained by cropping in the image matrix through row and column indexing. The image block is then identified and stored according to the grid number and time label.
[0023] Perform geometric correction, brightness normalization, and cloud / fog removal on the image; The extracted image blocks are first geometrically corrected to align images acquired at different times to a unified geographic coordinate system. Affine transformation or projection transformation methods are used to adjust the rotation, translation, and scaling deviations of the images. Subsequently, the image blocks are brightness standardized to normalize the pixel grayscale values of multi-temporal images to a unified range or perform color correction based on reference bands. For cloud and fog areas, the coverage area is identified and a mask is generated by using threshold-based brightness and color analysis methods or multi-band index calculation.
[0024] Image blocks are organized and cached according to the scale hierarchy of a multi-scale adaptive grid; After correction and removal, the image blocks are classified and managed according to their grid numbers and corresponding scale levels. A data structure is created for each scale level, and an index mapping table is built in memory or on disk to associate each image block with its grid number, timestamp, and spatial location, ensuring that data at any grid and time node can be accessed quickly.
[0025] The process of generating multi-scale spatiotemporal feature sequences in S2 is as follows: Based on multi-temporal image blocks corresponding to grids at each scale, convolutional coding, structural texture description, directional gradient sequence extraction, and local brightness difference analysis are performed on the image blocks within each grid cell to obtain the primary feature set for each time node. For each grid cell at each scale level, the image block at the corresponding time node is read and input into the convolution calculation module. The local feature response of the image pixels is calculated using a predefined convolution kernel to obtain the convolutional coding feature map. The convolutional coding result is further described by structure and texture, and the local pixel intensity distribution, edge direction and gray-level co-occurrence matrix are statistically analyzed to form texture pattern features. Then, the directional gradient sequence is calculated within the image block. The gradient direction and amplitude changes are calculated according to the pixel neighborhood, and the gradient changes of continuous pixels are serialized to form features. Finally, the local brightness difference is calculated for the pixels of the image block, and the brightness difference between the center pixel and the surrounding pixels is recorded to form brightness change features.
[0026] The primary feature sets of the same grid cell at different time nodes are sorted by timestamp, and time synchronization, cross-scale alignment, outlier feature marking and feature sparsification are performed. For each grid cell, the primary feature set of that cell at different time nodes is extracted, sorted according to the timestamp of the acquisition time to form a time series, and the feature series is time-synchronized to ensure that the features at different time nodes have a uniform time interval. The time step can be adjusted by interpolation or time alignment algorithms. Cross-scale alignment processing aligns the features corresponding to grids at different scales through spatial interpolation or feature mapping relationships, so that multi-scale features at the same location can be analyzed accordingly. Outliers in the feature series are marked by statistical thresholds or change amplitude detection methods and ignored or weighted down in subsequent processing. The feature vector is sparsified by selecting features with high contribution values or performing dimensionality reduction encoding to reduce redundant information while retaining key structural and brightness response information.
[0027] Multi-scale spatiotemporal feature sequences are constructed based on feature stability, integrity, and temporal density. For the feature time series of the same grid cell, calculate the stability index of each feature, including the consistency of fluctuation range and direction at continuous time nodes, calculate the integrity index, count the proportion of effective features to total features, and combine the time node density information to determine the credibility of features at each time step in the sequence. Based on stability, integrity and time density, select feature points, and combine the key features in the continuous time series in chronological order to generate a multi-scale spatiotemporal feature sequence.
[0028] S3: Determine the neighborhood grid set based on multi-scale spatiotemporal feature sequences and spatial indices, construct the spatiotemporal coupling relationship between grids, generate a spatiotemporal correlation matrix and predict the semantic change trend of the main grid, perform scale alignment and local perturbation suppression on multi-temporal features, and construct a cross-spatiotemporal stable reference frame.
[0029] The process of generating the spatiotemporal correlation matrix and predicting the semantic change trend of the main grid in S3 is as follows: The spatial neighborhood set of the main grid is determined based on the spatial index, and the corresponding temporal neighborhood set is constructed based on the timestamp; Using the spatial index information of each main grid, the grid cells adjacent to the main grid in geographic coordinates are found. The adjacency relationship is calculated by grid number and geographic range to form a spatial neighborhood set. In the time dimension, based on the acquisition timestamps of multi-temporal images, the same grids and their neighborhood features at adjacent time nodes are aligned to form a temporal neighborhood set for each main grid.
[0030] Joint analysis is performed on the multi-scale spatiotemporal feature sequences of the main grid and its neighborhood to calculate feature similarity, directional gradient consistency and local variation coupling degree. Multi-scale spatiotemporal feature sequences of the main grid and its neighboring grids at various scale levels and time nodes are extracted. Feature vectors at the same time step are compared to calculate feature similarity. The matching degree of structural features, texture features, and brightness difference features is evaluated by calculating Euclidean distance or cosine similarity. At the same time, the angle difference of the directional gradient sequence is calculated and continuity analysis is performed. The consistency of gradient direction within the neighborhood is statistically analyzed to form a gradient consistency index. The coupling degree of local feature change amplitude is calculated. The feature changes of the main grid are compared with the feature changes of the neighboring grids. The coupling degree value of local change is obtained by weighted averaging or convolution operation.
[0031] A spatiotemporal correlation matrix is generated based on the coupling degree, and the semantic change trend of the main grid is predicted accordingly. Step 3: Generate a spatiotemporal correlation matrix based on the coupling degree, and predict the semantic change trend of the main grid accordingly. Spatial neighborhood coupling degree, temporal neighborhood coupling degree, and cross-scale coupling degree are arranged according to grid number and time series to form a three-dimensional matrix representing the coupling relationship between the main grid and neighboring grids in space, time, and scale. Each dimension of the matrix is stored as a spatial correlation matrix, a temporal correlation matrix, and a cross-scale correlation matrix, respectively. Based on this, the multi-scale spatiotemporal feature sequence of the main grid is predicted in combination with the spatiotemporal correlation matrix. The feature change vector of future time nodes is generated by using linear combination, weighted summation, or trend calculation method based on coupling degree, and the semantic state change trend of the main grid is determined accordingly.
[0032] The process of constructing a stable reference frame across time and space in S3 is as follows: Based on the spatiotemporal correlation matrix and the predicted semantic change trend of the main grid, scale alignment, local noise perturbation suppression, structural residual filtering and neighborhood consistency enhancement are performed on the multi-temporal feature sequences of each grid. For each grid cell, based on its spatial neighborhood, temporal neighborhood, and cross-scale correlation information recorded in the spatiotemporal correlation matrix, the multi-temporal feature sequences are aligned according to the scale hierarchy to ensure that features at each scale can directly correspond at the same spatial location and time node. For local outliers or noise disturbances in the feature sequences, the local variance or brightness gradient change is calculated, and the influence of noise on the features is suppressed by sliding window filtering or local weighted averaging. For structural residuals, i.e., non-smooth changes in features between adjacent times or scales, threshold filtering or smooth interpolation is performed to adjust abrupt or discontinuous features to the range of neighborhood consistency. The neighborhood feature weighted fusion method is used to enhance the similar features of the grid and its neighborhood.
[0033] By utilizing the spatial adjacency, temporal continuity, and cross-scale dependence represented by the spatiotemporal correlation matrix, the processed spatiotemporal features are jointly analyzed to extract structurally stable regions, continuous change patterns, and key change responses, and anomalous disturbance features are weighted and suppressed. The feature matrices of each grid and its neighborhood at various time nodes and scales are arranged by spatial index and timestamp to form a three-dimensional joint feature matrix. Based on spatial adjacency, the features of adjacent grids are averaged or weighted and fused to identify structural feature regions that remain stable in multiple neighborhoods. For the time dimension, the continuity and consistency of features over time are calculated, and feature vectors that exhibit stable or continuously changing patterns in continuous time series are extracted. For cross-scale dependencies, the feature consistency of the same grids in different scale levels is analyzed, and key feature responses that recur at multiple scales are retained. At the same time, unstable or anomalous perturbation features are weighted and suppressed according to coupling weights to reduce their impact on the stable reference frame.
[0034] The extracted stable structures and change patterns are integrated according to grid number, spatial location and scale level to generate a cross-temporal and spatiotemporal stable reference frame; The structurally stable regions, continuous change patterns, and key change responses are indexed according to grid number, arranged in a geographic coordinate system according to spatial location, and the feature sets at different scale levels are uniformly identified. Through multidimensional arrays or data table structures, the spatial characteristics, temporal variation characteristics, and scale characteristics of each grid unit are uniformly recorded to form a searchable and indexable spatiotemporal stable reference frame.
[0035] S4: Integrate the cross-temporal stable reference frame, semantic change trend, spatiotemporal correlation matrix with the preliminary semantic judgment of the current frame, calculate the confidence of structural consistency, temporal consistency and neighborhood correlation, and form a comprehensive confidence matrix.
[0036] The process of fusing the cross-spatiotemporal stable reference frame, semantic change trend, spatiotemporal correlation matrix, and preliminary semantic determination of the current frame in S4 is as follows: The initial semantic determination of the current frame is compared with the structural features of the cross-temporal stable reference frame on a grid-by-grid basis, and the spatial consistency, local texture matching degree and cross-scale response difference of each grid are calculated. For each grid cell in the current frame, based on its number and geographical location in the spatial index, the structural features and change pattern features of the corresponding grid are retrieved from the spatiotemporal stable reference frame. The preliminary semantic judgment results of the current frame grid are compared with the stable features in the reference frame to calculate the spatial consistency value, including the feature position overlap and boundary similarity. At the same time, the local texture matching degree is calculated, and the texture similarity is obtained by convolution kernel or feature vector dot product. In the cross-scale dimension, the feature response values corresponding to each scale level are extracted, and the difference matrix is calculated.
[0037] Using the neighborhood coupling information provided by the spatiotemporal correlation matrix, neighborhood compensation and cross-temporal weighted correction are performed on meshes with structural deviations or local anomalies; Based on the spatial adjacency, temporal continuity, and cross-scale dependence recorded in the spatiotemporal correlation matrix, neighborhood compensation processing is performed on grids that deviate from the stable reference frame in the current frame. Feature values of neighboring grids in the spatial and temporal dimensions are extracted, and the features of the main grid are corrected by weighted averaging or local fusion algorithms to make them consistent with the features of the neighborhood. For grids with structural deviations, cross-temporal feature weighted correction is used to integrate the features of adjacent time nodes by linear or nonlinear weighting and adjust the semantic values of the grids.
[0038] Based on the predicted semantic change trend and compared with historical and neighborhood information, the time series correction is performed on the initially identified possible drift areas to generate preliminary fusion estimates. For each grid cell, the predicted semantic change trend is extracted and compared with historical frame data to calculate its rate of change and direction. If the preliminary judgment result shows abnormal drift or abrupt change, the correction factor is calculated based on the time series characteristics and smoothed by combining neighborhood and cross-scale information. The corrected feature value is updated to the preliminary fusion estimate through the time series weighted fusion algorithm.
[0039] The process of forming the comprehensive confidence matrix in S4 is as follows: For the fusion semantic estimation results of each grid, the structural consistency confidence is calculated, including the stability of features within the grid, texture matching degree, and cross-scale feature residuals. For each grid cell, its internal multi-scale feature vectors, including texture features, gradient direction features, and brightness difference features, are extracted from the fused semantic estimation results. The stability of the internal features of the grid is calculated, and the feature fluctuation is quantified by statistically analyzing the variance and standard deviation of each scale feature in the time series. The texture matching degree is calculated using convolution kernels or feature vector similarity to measure the overlap between the current grid features and the corresponding grid features in the cross-temporal stable reference frame. In the cross-scale dimension, the residuals of the features at each scale level are calculated, and the various indicators are combined to form the structural consistency confidence of the grid.
[0040] Based on the comparison of historical multi-temporal features and the prediction of semantic change trends in time series, the temporal consistency confidence is calculated, including feature drift, trend matching degree and abnormal change markers. For each grid cell, historical multi-temporal feature sequences are extracted and combined with the predicted semantic change trend. The current fused semantic estimate is compared with the time series to calculate the feature drift. The rate of change is obtained by the difference between the current feature and the historical feature and the time interval. The trend matching degree is calculated. The current change is quantified by the time series correlation coefficient or the least squares fitting method to determine whether it conforms to the predicted trend. At the same time, abnormal change points are marked to form an abnormal change index. The drift, trend matching degree and abnormal marking are combined to calculate the temporal consistency confidence of each grid cell.
[0041] Using the spatiotemporal correlation matrix and neighborhood grid feature information, the confidence of neighborhood correlation is calculated, including neighborhood similarity, local coupling degree and neighborhood deviation correction coefficient; The spatiotemporal correlation matrix is used to obtain the set of neighboring grids for each grid in space, time and across scales. The feature vectors of the neighboring grids are extracted and similarity calculation is performed with the features of the main grid. The neighborhood similarity is calculated, and the matching degree between the neighbors is measured by feature vector dot product or cosine similarity. The local coupling degree is calculated, and the synergy between the main grid and the neighboring grids is quantified by the variance and difference distribution of the neighborhood features. For grids with deviations, the neighborhood deviation correction coefficient is calculated, and the neighborhood correlation confidence is generated in a comprehensive manner.
[0042] According to the preset weighting strategy, the confidence scores of structural consistency, temporal consistency, and neighborhood correlation are fused to generate a comprehensive confidence matrix; The structural consistency confidence, temporal consistency confidence, and neighborhood correlation confidence of each grid are weighted and fused according to a preset weighting strategy. The weighting method can be linear weighting, which multiplies each confidence by a weight coefficient and then sums them, or it can be normalized before merging. Finally, the fusion results are organized into a matrix according to grid number, spatial location, and scale level to form a complete comprehensive confidence matrix.
[0043] S5: Based on the comprehensive confidence matrix and historical recognition records, a difference comparison is performed, the weights of key features and matching parameters are adjusted, and the corrected grid semantic recognition result is output.
[0044] The process of outputting the corrected grid semantic recognition result in S5 is as follows: The comprehensive confidence matrix is compared with the historical identification records grid by grid, and the confidence deviation of each grid in terms of structure, time and neighborhood dimensions is calculated. For each grid cell, three types of confidence values—structural consistency, temporal consistency, and neighborhood correlation—are extracted from the comprehensive confidence matrix. At the same time, the historical confidence and feature information of the corresponding grid are extracted from the historical identification records. The current confidence is compared with the historical confidence in three dimensions on a grid-by-grid basis, and the structural deviation, temporal deviation, and neighborhood deviation are calculated respectively to form a grid-level confidence deviation vector.
[0045] For grids with deviations exceeding a preset threshold, the weights of key features, including texture features, gradient direction, and brightness response, are adjusted based on confidence contribution and historical weights. For grid cells with confidence deviations exceeding a preset threshold, key features are extracted from the fused semantic estimation results, including texture feature vectors, directional gradient sequences, and brightness response values. Based on the confidence contribution and historical weight coefficients, the weights of each key feature are adjusted by increasing or decreasing. The adjustment strategy adopts a weighted update method, which proportionally corrects the weights of different features within the grid to ensure that the weight updates are consistent with historical information and neighborhood features. At the same time, the updated feature weights are recorded.
[0046] Based on the adjusted feature weights and grid matching parameters, the grid semantic classification probability is recalculated, and neighborhood interpolation or weighted fusion is performed on uncertain grids to form the final output of the corrected grid semantic judgment result. Using the adjusted key feature weights, the multi-scale feature vectors of each grid are weighted and summed or feature fused to generate a new semantic feature description of the grid. Based on the feature description and classification model or probability calculation formula, the semantic classification probability distribution of the grid is recalculated. For grids with low or uncertain classification probabilities, the features and classification probabilities of their spatial neighboring grids are extracted. The semantic judgment values of uncertain grids are corrected through neighborhood interpolation methods or weighted fusion strategies to generate stable grid semantic classification results. The corrected semantic classification results of each grid unit are organized and recorded according to grid number, spatial location and scale level to generate complete grid semantic recognition results.
[0047] Example 2: As Figure 2 As shown, a high-resolution satellite data semantic parsing and intelligent recognition system includes: Adaptive Mesh Module: Performs texture, edge, and saliency analysis on satellite imagery and performs dynamic scaling to generate multi-scale adaptive meshes and spatial indexes; Spatiotemporal feature module: Loads multi-temporal images and extracts scale spatiotemporal features of convolutional coding, structural description, gradient sequence and brightness difference on each grid; The association modeling module determines the neighborhood grid based on spatiotemporal features and spatial indexes, constructs the spatiotemporal coupling relationship between grids and predicts semantic change trends, while generating a stable reference frame across spatiotemporal regions. Confidence fusion module: It fuses the stable reference frame, change trend, correlation matrix and current frame determination to calculate the comprehensive confidence of structural consistency, temporal consistency and neighborhood correlation; The adaptive correction module compares the overall confidence level with historical recognition records, adjusts the feature weights and matching parameters, and generates the corrected semantic recognition results.
[0048] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0049] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for semantic parsing and intelligent recognition of high-resolution satellite data, characterized in that, Includes the following steps: Texture complexity, edge density, and saliency response analysis are performed on high-resolution satellite imagery. Dynamic scaling is performed based on regional differences to generate multi-scale adaptive grid cells and establish a spatial index. Based on multi-scale adaptive grids and spatial index loading of corresponding multi-temporal images, multi-scale convolutional coding, structural description, gradient sequence extraction and brightness difference calculation are performed on each grid to generate multi-scale spatiotemporal feature sequences. Based on the multi-scale spatiotemporal feature sequence and spatial index, the neighborhood grid set is determined, the spatiotemporal coupling relationship between grids is constructed, the spatiotemporal correlation matrix is generated and the semantic change trend of the main grid is predicted, the scale alignment and local perturbation suppression of multi-temporal features are performed, and a cross-spatiotemporal stable reference frame is constructed. By fusing the cross-temporal and spatial stable reference frame, semantic change trend, spatiotemporal correlation matrix, and the preliminary semantic judgment of the current frame, the confidence scores of structural consistency, temporal consistency, and neighborhood correlation are calculated to form a comprehensive confidence matrix. Based on the difference comparison between the comprehensive confidence matrix and historical recognition records, the weights of key features and matching parameters are adjusted, and the corrected grid semantic recognition result is output.
2. The high-resolution satellite data semantic parsing and intelligent recognition method according to claim 1, characterized in that, The process of generating multi-scale adaptive mesh elements and establishing spatial indexes is as follows: The image is initially divided into blocks according to spatial location and resolution, and the texture entropy, orientation gradient statistics and edge sharpness index are calculated for each block. The importance of the segmented regions is quantitatively assessed by combining the saliency response map. Based on the spatial detail density, self-similarity and structural complexity within the regions, the images are divided into different scale levels. An adaptive grid generation strategy is adopted for each scale level to generate multi-scale adaptive grid cells, and the grid number, geographical range, level identifier and neighborhood relationship are recorded to establish a spatial index.
3. The high-resolution satellite data semantic parsing and intelligent recognition method according to claim 2, characterized in that, The process of loading multi-temporal images based on multi-scale adaptive grids and spatial indexing is as follows: Based on the grid number and geographical extent recorded in the spatial index, image blocks corresponding to grid units of different scales are extracted from the multi-temporal image dataset; Perform geometric correction, brightness normalization, and cloud / fog removal on the image; Image blocks are organized and cached according to the scale hierarchy of a multi-scale adaptive grid.
4. The high-resolution satellite data semantic parsing and intelligent recognition method according to claim 3, characterized in that, The process of generating multi-scale spatiotemporal feature sequences is as follows: Based on multi-temporal image blocks corresponding to grids at each scale, convolutional coding, structural texture description, directional gradient sequence extraction, and local brightness difference analysis are performed on the image blocks within each grid cell to obtain the primary feature set for each time node. The primary feature sets of the same grid cell at different time nodes are sorted by timestamp, and time synchronization, cross-scale alignment, outlier feature marking and feature sparsification are performed. Multi-scale spatiotemporal feature sequences are constructed based on feature stability, integrity, and temporal density.
5. The high-resolution satellite data semantic parsing and intelligent recognition method according to claim 4, characterized in that, The process of generating the spatiotemporal correlation matrix and predicting the semantic change trend of the main grid is as follows: The spatial neighborhood set of the main grid is determined based on the spatial index, and the corresponding temporal neighborhood set is constructed based on the timestamp; Joint analysis is performed on the multi-scale spatiotemporal feature sequences of the main grid and its neighborhood to calculate feature similarity, directional gradient consistency and local variation coupling degree. A spatiotemporal correlation matrix is generated based on the coupling degree, and the semantic change trend of the main grid is predicted accordingly.
6. The high-resolution satellite data semantic parsing and intelligent recognition method according to claim 5, characterized in that, The process of constructing a stable reference frame across time and space is as follows: Based on the spatiotemporal correlation matrix and the predicted semantic change trend of the main grid, scale alignment, local noise perturbation suppression, structural residual filtering and neighborhood consistency enhancement are performed on the multi-temporal feature sequences of each grid. By utilizing the spatial adjacency, temporal continuity, and cross-scale dependence represented by the spatiotemporal correlation matrix, the processed spatiotemporal features are jointly analyzed to extract structurally stable regions, continuous change patterns, and key change responses, and anomalous disturbance features are weighted and suppressed. The extracted stable structures and change patterns are integrated according to grid number, spatial location and scale level to generate a cross-temporal and spatiotemporal stable reference frame.
7. The high-resolution satellite data semantic parsing and intelligent recognition method according to claim 6, characterized in that, The process of fusing the cross-temporal stable reference frame, semantic change trend, spatiotemporal correlation matrix, and preliminary semantic determination of the current frame is as follows: The initial semantic determination of the current frame is compared with the structural features of the cross-temporal stable reference frame on a grid-by-grid basis, and the spatial consistency, local texture matching degree and cross-scale response difference of each grid are calculated. Using the neighborhood coupling information provided by the spatiotemporal correlation matrix, neighborhood compensation and cross-temporal weighted correction are performed on meshes with structural deviations or local anomalies; Based on the predicted semantic change trend and compared with historical and neighborhood information, time series corrections are made for the initially identified possible drift areas to generate preliminary fusion estimates.
8. The high-resolution satellite data semantic parsing and intelligent recognition method according to claim 7, characterized in that, The process of forming the comprehensive confidence matrix is as follows: For the fusion semantic estimation results of each grid, the structural consistency confidence is calculated, including the stability of features within the grid, texture matching degree, and cross-scale feature residuals. Based on the comparison of historical multi-temporal features and the prediction of semantic change trends in time series, the temporal consistency confidence is calculated, including feature drift, trend matching degree and abnormal change markers. Using the spatiotemporal correlation matrix and neighborhood grid feature information, the confidence of neighborhood correlation is calculated, including neighborhood similarity, local coupling degree and neighborhood deviation correction coefficient; According to the preset weighting strategy, the confidence scores of structural consistency, temporal consistency, and neighborhood correlation are fused to generate a comprehensive confidence matrix.
9. The high-resolution satellite data semantic parsing and intelligent recognition method according to claim 8, characterized in that, The process of outputting the corrected grid semantic recognition result is as follows: The comprehensive confidence matrix is compared with the historical identification records grid by grid, and the confidence deviation of each grid in terms of structure, time and neighborhood dimensions is calculated. For grids with deviations exceeding a preset threshold, the weights of key features, including texture features, gradient direction, and brightness response, are adjusted based on confidence contribution and historical weights. Based on the adjusted feature weights and grid matching parameters, the grid semantic classification probability is recalculated, and neighborhood interpolation or weighted fusion is performed on uncertain grids to form the final output of the corrected grid semantic judgment result.
10. A high-resolution satellite data semantic parsing and intelligent recognition system, applied to the method described in any one of claims 1-9, characterized in that, include: Adaptive Mesh Module: Performs texture, edge, and saliency analysis on satellite imagery and performs dynamic scaling to generate multi-scale adaptive meshes and spatial indexes; Spatiotemporal feature module: Loads multi-temporal images and extracts multi-scale spatiotemporal features such as convolutional coding, structural description, gradient sequence and brightness difference on each grid; The association modeling module determines the neighborhood grid based on spatiotemporal features and spatial indexes, constructs the spatiotemporal coupling relationship between grids and predicts semantic change trends, while generating a stable reference frame across spatiotemporal regions. Confidence fusion module: It fuses the stable reference frame, change trend, correlation matrix and current frame determination to calculate the comprehensive confidence of structural consistency, temporal consistency and neighborhood correlation; The adaptive correction module compares the overall confidence level with historical recognition records, adjusts the feature weights and matching parameters, and generates the corrected semantic recognition results.