A 3D Gaussian-sprayed multi-temporal satellite remote sensing image ground change detection method

CN122821279APending Publication Date: 2026-09-25JINGGANGSHAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610969552.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]部分方法尝试引入数字表面模型对影像进行正射校正以消除视角差异影响,但高精度数字表面模型获取成本高,时效性也难以保证

Benefits of technology

第一,该方法将变化检测从依赖几何配准的像素比较转化为基于语义特征向量的跨时相检索。语义特征向量学习表达的是地物物理材质属性而非特定视角下的外观,使得同一地物在前后时相影像中的语义特征具有相似性。检测过程不执行时相间像素级几何精配准操作,宽松空间权重仅根据卫星成像几何参数提供空间引导,不要求高斯基元在后时相影像上的精确投影位置。通过在后时相影像中检索语义匹配区域,未变化地物的存在性得分较高,已消失地物因缺乏语义匹配而得分较低,从而减少因建筑物高度和视角差异引起的伪变化。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821279A_ABST
    Figure CN122821279A_ABST
Patent Text Reader

Abstract

The application discloses a 3D Gaussian sputtering multi-temporal satellite remote sensing image ground surface change detection method and relates to the technical field of remote sensing image processing. The method comprises the following steps: acquiring multi-temporal high-resolution satellite images of a target region; training a three-dimensional Gaussian sputtering model by using multi-view images of a previous time phase, wherein each Gaussian cell has geometric parameters, color parameters and a semantic feature vector; inputting a later time phase image into a pre-trained semantic coding network to obtain a two-dimensional semantic feature map; calculating the similarity between the semantic feature vector of each Gaussian cell and the semantic features at each position in the later time phase semantic feature map, taking the maximum value after being weighted by a loose spatial weight as an existence score; and determining the Gaussian cell with a score lower than a threshold value as a change and outputting a three-dimensional detection result. The application realizes change detection through cross-time phase semantic feature retrieval, avoids pixel-level geometric fine registration between time phases and solves the problem of false changes caused by building height and imaging angle difference in urban areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of remote sensing image processing technology, and more specifically, to a method for detecting surface changes in multi-temporal satellite remote sensing images using 3D Gaussian sputtering. Background Technology

[0002] Change detection using multi-temporal high-resolution satellite remote sensing imagery is an important technical means in fields such as urban monitoring and illegal construction identification. By analyzing satellite images of the same geographical area acquired at different times, changes such as the addition, demolition, or alteration of objects on the ground can be identified.

[0003] In high-resolution satellite imagery of urban areas, buildings generally have a certain height. When satellites image at different times, the viewing angle often differs, causing the side facades and shadows of the same building to be projected onto different ground pixel locations. This projection shift results in noticeable pixel differences between images from different time periods for buildings that have remained unchanged.

[0004] Current change detection methods typically rely on pixel-level geometrical registration of consecutive temporal images, followed by pixel differencing or feature comparison after spatial alignment. However, two-dimensional registration models struggle to eliminate projection distortions caused by differences in building height and viewing angle. Even if the overall registration accuracy meets requirements, pixel differences in building facade areas may still exceed the registration error tolerance and be misjudged as surface changes.

[0005] Some methods attempt to introduce digital surface models to orthorectify images to eliminate the influence of viewpoint differences; however, acquiring high-precision digital surface models is costly and timeliness is difficult to guarantee. When the digital surface model is inconsistent with the actual ground features, the correction process may introduce new errors. Furthermore, methods that perform 3D reconstruction of consecutive temporal images and then compare them will see the errors from the two independent reconstructions accumulate and propagate to the change detection stage, affecting the detection results. Therefore, a method for detecting surface changes in multi-temporal satellite remote sensing images using 3D Gaussian sputtering is proposed to address the above problems. Summary of the Invention

[0006] To overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a method for detecting surface changes in multi-temporal satellite remote sensing images using 3D Gaussian sputtering. This method aims to address the problem that in the detection of changes in multi-temporal satellite remote sensing images, the method relies on pixel-level geometrical registration between time phases, while two-dimensional registration is difficult to eliminate projection distortion caused by differences in building height and imaging viewing angle, leading to misjudgment of viewing angle differences as surface changes.

[0007] To achieve the above objectives, the present invention provides the following technical solution: A method for detecting surface changes in multi-temporal satellite remote sensing images using 3D Gaussian sputtering includes the following steps: S1: Acquire at least two high-resolution satellite images of the target area from different perspectives in the preceding time phase and one high-resolution satellite image in the following time phase. The preceding and following time phase images constitute a multi-temporal image pair. The multi-view images in the preceding time phase provide multi-view geometric constraints for the subsequent training of the 3D Gaussian sputtering model, while the single image in the following time phase serves as the data to be detected for change detection.

[0008] S2: A three-dimensional Gaussian sputtering model is trained using multi-view satellite imagery from previous time periods. The model consists of multiple Gaussian elements, each with geometric parameters, color parameters, and semantic feature vectors.

[0009] During training, the parameters of the Gaussian primitives are jointly optimized using color reconstruction loss and semantic consistency loss. The color reconstruction loss calculates the difference between the rendered colors and the colors of the real image, constraining the Gaussian primitives to accurately fit the scene's geometry and appearance. The semantic consistency loss calculates the difference between the rendered semantic feature map and the semantic feature maps extracted from images from previous time-series perspectives.

[0010] Since multi-view images observe the same surface object from different angles, the same material may appear in different colors under different views, but its physical material properties remain unchanged. By minimizing the semantic consistency loss of multi-view images, the semantic feature vector is constrained to learn to express the physical material properties of the surface at the corresponding spatial location, rather than recording the instantaneous color or lighting conditions under a specific view.

[0011] The semantic feature vectors obtained in this way are stable under changes in illumination and viewing angle, providing a reliable foundation for subsequent cross-temporal semantic retrieval.

[0012] The parameters of the Gaussian unit are jointly optimized using the following loss function: ; in, For the total loss, For color reconstruction loss, For semantic consistency loss, These are preset weight coefficients. A smaller value is chosen during the initial training phase. The value is initially set to prioritize fitting geometric structure and color information to the model, and then gradually increased. The value is adjusted to strengthen the optimization of semantic feature vectors by semantic consistency constraints. The value of can ensure the accuracy of geometric and appearance fitting while obtaining a stable expression of the material properties of ground features.

[0013] The process of obtaining semantic feature vectors is as follows: Semantic feature maps of previous temporal images were extracted using a pre-trained semantic coding network. By differentiable rendering, the semantic feature vectors of each Gaussian unit are projected onto each viewpoint to obtain the rendered semantic feature map. The difference between the two is used as the semantic consistency loss for optimization. Under the multi-view constraint, the semantic feature vector of each Gaussian unit converges into a compact vector expressing the material properties of that spatial location.

[0014] The parameters of the pre-trained semantic encoding network remain fixed throughout the training process and subsequent feature extraction, and the same network is used for both phases. This ensures that the semantic features of both phases are in the same feature space and are comparable, which is a prerequisite for effective cross-phase semantic similarity calculation.

[0015] The geometric parameters of the Gaussian element include its 3D center position and covariance matrix, which describes the shape and orientation of the Gaussian element in 3D space. Color parameters are represented by spherical harmonic coefficients, which, as coefficients of a set of orthogonal basis functions, can express the continuous color variation of the Gaussian element under different viewpoints. The semantic feature vector has a lower dimension than the color parameters, representing the material properties of land features in a compact form while reducing the computational cost of subsequent cross-temporal retrieval.

[0016] S3: Input the satellite imagery from the later time phase into the same pre-trained semantic coding network as in S2 to obtain a two-dimensional semantic feature map. Since the same network is used for both time phases and the network parameters remain unchanged, the weights and biases of the convolutional kernels in each layer of the network remain constant. The semantic features of the same type of land surface material in the images from both time phases are in the same feature space and have high similarity. This provides a basis for feature space consistency in the cross-time phase semantic similarity calculation in S4.

[0017] S4: For each Gaussian element in the model obtained in S2, calculate the similarity between its semantic feature vector and the semantic features of each pixel position in the two-dimensional semantic feature map, and use the relaxed spatial weights determined by the satellite imaging geometric parameters to weight the similarity, and take the global maximum value after weighting as the existence score of the Gaussian element in the later time phase.

[0018] In high-resolution satellite imagery of urban areas, the height of buildings and the difference in viewing angles during different time phases can cause shifts in the projected position of building tops on the image. If precise projection positions are used for semantic matching, this shift may prevent the semantic features of unchanged features from finding a match at the corresponding location, thus leading to misjudgment as changed features.

[0019] The relaxed spatial weights are constructed based on satellite imaging geometric parameters and provide spatial guidance for the possible imaging positions of Gaussian elements in later temporal images, reflecting only spatial plausibility without requiring precise projection positions. Even if the projection positions of ground features on the image shift due to changes in viewing angle, a high score can still be obtained as long as their semantic features match within a reasonable spatial range, thus avoiding semantic matching failures caused by projection position deviations.

[0020] The loose space weights are constructed as follows: Using the rational polynomial coefficient parameters of the preceding and following temporal images, the three-dimensional center position of the Gaussian element is projected onto the following temporal image plane to obtain a rough projection position. A two-dimensional Gaussian weighted map is constructed with this location as the center. The weights are at their maximum value at the center and gradually decrease towards the edges, with the rate of decrease controlled by the standard deviation.

[0021] The standard deviation is determined by the range of differences between the maximum height of buildings and the satellite view within the target area, so that the coverage of the weighted map includes the maximum projection offset caused by changes in building height and view, providing a loose spatial guide for the possible imaging positions of Gaussian elements.

[0022] Existence score is calculated using the following formula: ; in, For existence score, For Gaussian elements, the semantic feature vector is... Location in a two-dimensional semantic feature map semantic features at the location For similarity function, For the loose space weight in position The value at that location, For all spatial locations Take the maximum value.

[0023] The calculation searches for the region in the entire post-temporal image space that best matches the semantics of Gaussian meta-analysis and has a reasonable spatial location, and combines semantic similarity and spatial reasonableness into a single score.

[0024] If the ground features have not changed, and there are semantically highly similar regions within a reasonable spatial range in the later time-phase image, the spatial weight and semantic similarity of this location are both high, and the weighted score is close to the maximum value.

[0025] If the ground features have disappeared, there will be a lack of semantic matching regions in the later time-phase images. Even if some locations show a certain degree of similarity due to coincidence in materials, the weighted score will be at a low level because the spatial weight of that location is low or the similarity itself is not high.

[0026] Therefore, the existence score can effectively distinguish between unchanged features and features that have disappeared.

[0027] Similarity function Cosine similarity is used, calculated by dividing the inner product of two vectors by the product of their respective magnitudes. Cosine similarity is insensitive to vector magnitudes and only reflects the directional consistency of two semantic feature vectors. While differences in illumination intensity between temporal images may cause an overall scaling of the response amplitudes of each component of the semantic feature vectors, the orientation of the same ground feature material remains similar. Cosine similarity can provide a high similarity value, ensuring that the existence score is unaffected by the overall scaling of the semantic feature magnitudes.

[0028] S5: Goughbred elements with an existence score below a preset threshold are classified as variable Goughbred elements.

[0029] The preset threshold is determined based on the statistical distribution characteristics of the similarity function on unchanged land features. Known unchanged surface areas within the target region are selected as samples, and the existence score of Gaussian elements in these samples is calculated. The preset quantile of this score distribution is then used as the threshold. This method adapts the threshold to the image quality, land feature type, and data characteristics of the target region, reducing the impact of differences in data characteristics across different regions on the consistency of detection results.

[0030] S6: Output all changes in Gaussian elements as 3D detection results of surface changes.

[0031] The 3D detection result of land surface changes is a 3D point cloud carrying change labels. The 3D center positions and covariance matrices of the Gaussian elements labeled as changed represent the 3D geometric contours of vanished land objects in the previous time phase. The 3D center positions determine the spatial coordinates of the land features, while the scaling and orientation of the covariance matrix determine the extent and spatial orientation of the land features in 3D space. This output directly provides 3D spatial information of changed land features, which can be further used for determining the spatial extent and type identification of changed targets.

[0032] The technical effects and advantages of this invention are as follows: First, this method transforms change detection from pixel-by-pixel comparison relying on geometric registration to cross-temporal retrieval based on semantic feature vectors. Semantic feature vectors learn to represent the physical material properties of ground features rather than their appearance from a specific viewpoint, ensuring similar semantic features for the same ground feature across different temporal images. The detection process does not perform pixel-level precise geometric registration between temporal images; the relaxed spatial weights provide spatial guidance solely based on satellite imaging geometric parameters, without requiring precise projection positions of Gaussian elements on later temporal images. By retrieving semantically matching regions in later temporal images, unchanged ground features receive higher presence scores, while vanished ground features receive lower scores due to a lack of semantic matching, thus reducing spurious changes caused by differences in building height and viewpoint.

[0033] Second, this method trains a 3D Gaussian sputtering model under pre-temporal multi-view constraints, reconstructing the 3D geometry of the scene. The output is a 3D point cloud with variation labels. The 3D center positions and covariance matrices of the Gaussian elements labeled as variations can directly reflect the 3D spatial contours of the vanished surface objects, providing 3D structural information for subsequent analysis. Attached Figure Description

[0034] Figure 1 This is a schematic diagram of the implementation process of the method of the present invention; Figure 2 This is a schematic diagram of the training process of the three-dimensional Gaussian sputtering model of the present invention; Figure 3 This is a schematic diagram of the output structure of the three-dimensional change detection results of the present invention. Detailed Implementation

[0035] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0036] Example 1 As attached Figures 1 to 3 The present invention discloses a method for detecting surface changes in multi-temporal satellite remote sensing images using 3D Gaussian sputtering. The technical solution of the present invention will be described in detail below according to the step sequence.

[0037] S1: Acquire multi-temporal high-resolution satellite images of the target area.

[0038] The target area is the city's core area, which contains multiple buildings with a height exceeding 20 meters.

[0039] The first phase imagery consists of three-view stereo images acquired during a single pass of the high-resolution stereo mapping satellite, including forward-looking, emergent, and retrograde views. The second phase imagery consists of a single emergent image acquired during another pass of the same satellite. The acquisition time interval between the two phase images is one year.

[0040] Both temporal images are accompanied by rational polynomial coefficient parameters, which describe the geometric model during satellite imaging. These rational polynomial coefficient parameters record the mapping relationship between sensor position, attitude, and ground coordinates. They can project three-dimensional ground coordinates onto two-dimensional image coordinates, or inversely calculate ground coordinates from image coordinates. The preceding and subsequent temporal images constitute a multi-temporal image pair.

[0041] In another implementation, the preceding and following time-phase images are from different satellite sensors. The preceding time-phase images use satellite data with multi-view imaging capabilities, while the following time-phase images use satellite data that only provides single-view imaging. This approach is suitable when the same satellite sensor cannot acquire data at the required time phase due to mission schedule or weather conditions.

[0042] S2: Train a 3D Gaussian sputtering model using multi-view satellite imagery from previous time phases.

[0043] The three-dimensional Gaussian sputtering model consists of multiple Gaussian primitives, each carrying three types of parameters: geometric parameters, color parameters, and semantic feature vectors.

[0044] The geometric parameters include the three-dimensional center position and the covariance matrix. The covariance matrix describes the shape and orientation of the Gaussian element in three-dimensional space. It is a 3×3 positive definite matrix that can be decomposed into a rotation matrix and a scaling matrix, which control the orientation of the Gaussian element and the extent of its extension along each axis, respectively.

[0045] Color parameters are represented by spherical harmonic coefficients, which are the coefficients of a set of orthogonal basis functions used to express the color changes presented by Gaussian elements under different viewpoints. The semantic feature vector is a low-dimensional vector, with a dimension lower than that of the spherical harmonic coefficients, used to express the surface physical material properties of the spatial location corresponding to the Gaussian element, such as concrete, glass curtain walls, or vegetation.

[0046] The method of model initialization is explained below.

[0047] As one implementation method, a sparse point cloud is generated from previous temporal multi-view images using a motion reconstruction structure method. The motion reconstruction structure method receives previous temporal three-view stereo images as input, detects feature points in each image and performs multi-view matching, recovers the camera pose and 3D point coordinates using epipolar geometry constraints, and outputs the position and color information of each 3D point in the sparse point cloud.

[0048] The positions of each 3D point in the sparse point cloud are used as the initial values ​​for the 3D center positions of Gaussian elements, and the color mean of the corresponding pixel of the 3D point is used as the initial value for the color parameters. The covariance matrix of each Gaussian element is initialized to an isotropic Gaussian distribution, that is, a matrix with equal diagonal elements and zero off-diagonal elements.

[0049] A semantic feature vector is appended to each Gaussian element, using random initialization. This approach is suitable for scenarios with rich textures in previous temporal images and reliable matching of corresponding points across different viewpoints.

[0050] As another implementation, when a reliable sparse point cloud cannot be obtained through motion reconstruction methods, such as when feature point matching fails due to large areas of water, cloud cover, or areas with weak texture in previous temporal images, the 3D Gaussian sputtering model directly initializes a preset number of Gaussian elements randomly within the 3D spatial range of the target region. The preset number is set according to the area of ​​the target region and the scene complexity.

[0051] The three-dimensional center position of each Gaussian element is uniformly and randomly sampled in the three-dimensional space corresponding to the target region. The covariance matrix is ​​initialized to an isotropic Gaussian distribution of a preset size. The color parameters and semantic feature vectors are also randomly initialized.

[0052] During training, in addition to optimizing the parameters of existing Gaussian elements, adaptive density control is also applied to them. Specifically, every preset number of iterations, the gradient norm and transparency value of each Gaussian element are calculated. Gaussian elements whose gradients exceed a preset gradient threshold are split or cloned to increase the density of Gaussian elements in regions rich in geometric details.

[0053] The splitting operation divides a Gaussian primitive (GRP) into two along the direction of maximum variance in its covariance matrix, while the cloning operation copies the GRP and fine-tunes its position. GRPs whose transparency consistently falls below a preset transparency threshold are deleted to remove redundant GRPs that contribute little to scene rendering. Through these density controls, the model gradually adjusts from its initial state to a distribution that matches the geometric complexity of the scene.

[0054] The following section explains pre-trained semantic coding networks and how they are used.

[0055] The pre-trained semantic encoding network employs a residual convolutional neural network pre-trained on a remote sensing scene classification dataset. Pre-trained weights are loaded before training, and the parameters remain fixed throughout the training process, not participating in gradient updates. The network receives a single satellite image as input, extracts hierarchical features through multiple convolutional and pooling operations, and outputs a multi-channel two-dimensional semantic feature map.

[0056] Because the network was pre-trained on remote sensing images containing various land cover types, the responses of different channels correspond to different land surface materials and texture patterns, and the same type of material has similar activation patterns in the feature map.

[0057] In another implementation, when computational resources are limited or storage space is strictly required, the pre-trained semantic coding network employs a lightweight convolutional neural network pre-trained on a remote sensing semantic segmentation dataset, resulting in a smaller number of output channels. For example, the standard network outputs 128 channels, while the lightweight network outputs 64 channels.

[0058] At this point, the dimension of the semantic feature vector is also reduced accordingly, for example, from 16 dimensions to 8 dimensions. When calculating the semantic consistency loss, since the number of channels in the rendered semantic feature map is inconsistent with the number of channels in the extracted semantic feature map, a learnable linear projection layer is used to map the low-dimensional rendered semantic feature map to the same number of channels as the extracted semantic feature map, and then the difference between the two is calculated. The parameters of the linear projection layer are a channel number transformation matrix, which is optimized during model training and is not retained after training; it only serves the purpose of dimension matching during the training phase.

[0059] The training process is carried out through iterative optimization. Each iteration performs the following operations: select a viewpoint image from the previous time phase, and project the semantic feature vectors of each Gaussian unit in that viewpoint onto the image plane through differentiable rendering to obtain a rendered semantic feature map.

[0060] Differentiable rendering calculates the projection position and shape of each Gaussian primitive on the image plane based on its 3D center position, covariance matrix, and camera parameters from the current viewpoint. Then, it sorts these primitives by depth from farthest to nearest, and performs a blended projection accumulation based on their transparency. Primitives closer to the camera with higher transparency occlude those farther away. Simultaneously, the image from the same viewpoint is input into a pre-trained semantic encoding network to obtain the extracted semantic feature map output by the network.

[0061] Training uses a joint loss function. The color reconstruction loss is calculated. This refers to the difference between the rendered color and the actual image color, calculated pixel-by-pixel using the L1 norm and then averaged. Semantic consistency loss is then calculated. This refers to the difference between the rendered semantic feature map and the extracted semantic feature map. Similarly, the absolute value of the difference is calculated position-by-position and channel-by-channel using the L1 norm, and then averaged. Total loss according to Calculation, where These are preset weighting coefficients.

[0062] In the initial training phase Choose a small value, such as 0.01, to prioritize fitting geometric structure and color information to the model. Gradually increase this value as iterations progress. Values, for example, every 1000 iterations will be Increase the value by 0.01 until a preset upper limit is reached to strengthen the optimization of semantic consistency constraints on semantic feature vectors. Calculate the gradient of the total loss with respect to the parameters of each Gaussian element through backpropagation, and update the geometric parameters, color parameters, and semantic feature vectors of each Gaussian element based on the gradient.

[0063] Through joint optimization with multi-view constraints, the semantic feature vector gradually converges into a vector representing the physical material properties of the surface at the corresponding spatial location. For the same concrete roof, the colors rendered by the spherical harmonic coefficient differ from the frontal, oblique, and rearward views due to variations in lighting angle and observation direction. However, under the semantic consistency constraint across multiple views, the semantic feature maps rendered from all three views must be consistent with the extracted semantic features from the corresponding views. Figure 1 Therefore, the semantic feature vector converges to a vector expressing the material property of "concrete," rather than the color feature under a specific lighting condition. This semantic feature vector is stable to changes in lighting and viewing angle.

[0064] When dot product similarity is used as the similarity function in subsequent steps, an additional L2 norm regularization constraint on the semantic feature vector is added to the training loss function. The regularization term is the square root of the sum of squares of each component of the semantic feature vector multiplied by a preset regularization coefficient. This term makes the magnitude of the semantic feature vector approach 1 during training.

[0065] Since the result of dot product similarity is affected by the vector magnitude, when the magnitudes of both vectors are 1, the dot product similarity degenerates into cosine similarity. By using regularization constraints to make the magnitudes approach 1, the stability of the magnitudes in cross-temporal similarity calculations can be maintained, avoiding similarity distortion caused by differences in magnitudes.

[0066] S3: Semantic feature map of the extracted temporal image.

[0067] A single frontal view image from a later time phase is input into the same pre-trained semantic encoding network as in S2. The network parameters remain fixed, and it outputs a two-dimensional semantic feature map.

[0068] The spatial resolution of this two-dimensional semantic feature map is the same as that of the input image. The semantic feature at each spatial location is a vector with the same dimension as the semantic feature vector in S2, and both are located in the same feature space.

[0069] Since the same semantic coding network is used across different time phases and the network parameters remain unchanged, the weights and biases of the convolutional kernels in each layer of the network remain constant. Therefore, the semantic features of the same type of surface material exhibit high similarity in the images from different time phases. For example, the semantic feature vector directions of corresponding areas of the same concrete building should be highly consistent across different time phases. This provides a foundation for feature space consistency in subsequent cross-temporal semantic retrieval.

[0070] S4: For each Gaussian element in the three-dimensional Gaussian sputtering model obtained in S2, calculate its existence score.

[0071] This step processes each Gaussian element in the model obtained from S2 one by one, performing the construction of the loose space weights and the calculation of the existence score for each Gaussian element in turn.

[0072] First, a loose spatial weight is constructed. Using the rational polynomial coefficient parameters of the preceding and following temporal images, the three-dimensional center position of the Gaussian element is orthogonally projected onto the following temporal image plane through the rational polynomial coefficients to obtain a two-dimensional pixel coordinate, which serves as the coarse projection position.

[0073] Since the geometric model described by the coefficient parameters of a rational polynomial is a polynomial ratio function, its accuracy is limited by the order of the polynomial and the calibration accuracy of the coefficients. In addition, since buildings have a certain height, the projected position of the top of the building on different time phase images will be significantly shifted when the satellite imaging angle changes. This shift is positively correlated with the building height and the angle of difference in the viewing angle. Therefore, there is a deviation between the rough projected position and the actual imaging position of the ground object on the later time phase image.

[0074] A two-dimensional Gaussian weighted map is constructed with this rough projection location as the center. This weighted map covers the entire spatial extent of the subsequent temporal image, in terms of spatial location. value at The weight is determined by the Euclidean distance from the location to the rough projection location and the standard deviation of the Gaussian distribution, with the center having the largest weight and gradually decreasing towards the periphery.

[0075] The standard deviation is determined by taking the maximum height of buildings within the target area. The maximum height is the maximum height value among buildings exceeding 20m as described in S1. The absolute value of the difference between the imaging angle of the preceding and following temporal frontal images is taken as the angle difference angle. The angle of difference can be calculated using the ephemeris parameters of the two temporal images and the sensor pointing angle.

[0076] Standard deviation according to Sure; in Preset scaling factor, The theoretical value of the maximum projected offset of the building's top due to viewing angle differences, multiplied by a scaling factor. The offset is then converted to the standard deviation of a Gaussian distribution, extending the effective coverage of the weighted map from the center to a region approximately 2 to 3 times the standard deviation, sufficient to encompass the maximum projection offset. This weighted map provides lenient spatial guidance for the possible imaging positions of Gaussian elements on later temporal images, reflecting spatial rationality through weight attenuation, without requiring the position to precisely coincide with the coarse projection position.

[0077] In another implementation, when the difference in building height and the difference in viewing angle within the target area are small, the standard deviation is directly adopted as a preset fixed value, without the need to calculate it separately for each Gaussian element, thereby reducing computational overhead.

[0078] Then, the existence score of the Gaussian unit is calculated. The semantic feature vector of the Gaussian unit is then... With each spatial location in the two-dimensional semantic feature map semantic features at the location Calculate similarity for each value. Then, compare the obtained similarity values ​​with the relaxed space weighted graph. The value at the corresponding position in By multiplying each position, a weighted similarity map is obtained.

[0079] The maximum value across all spatial locations in the weighted similarity graph is taken as the existence score of the Gaussian unit. ,Right now ,in Indicates all spatial locations Take the maximum value above. This is the similarity function.

[0080] The existence score means searching for the region in the later temporal image that best matches the semantics of the Gaussian meta-analysis and has a reasonable spatial location.

[0081] If the surface object corresponding to the Gaussian element remains unchanged between the preceding and following time phases, then there should be a location in the later time phase image near the coarse projection position that is semantically highly similar, and the direction of the semantic feature vector at this location should be similar to... The similarity is consistent, with a cosine similarity close to 1, and the spatial weight at this position is relatively high, resulting in a high weighted similarity value and an existence score close to 1.

[0082] If the surface object has been demolished, then in the subsequent temporal image, there are areas lacking semantic matching within a reasonable location range, and the semantic feature vectors at each location are all... There are significant differences, and the similarity value is low.

[0083] Even if a certain similarity occurs at a certain location due to a coincidence in materials, the spatial weight of that location is lower because it is far from the rough projection location, or the similarity itself is not high. The combination of these two factors results in a low maximum value after weighting, and the existence score is significantly lower than 1.

[0084] Similarity function There are multiple implementation methods: In one implementation, Cosine similarity is used, calculated by dividing the inner product of two vectors by the product of their respective magnitudes. The value of cosine similarity ranges from -1 to 1, with values ​​closer to 1 indicating more consistent orientations. Cosine similarity is insensitive to vector magnitudes, measuring only directional consistency. It is suitable for scenarios where changes in illumination between images from different time periods cause an overall scaling of semantic feature magnitudes, but the orientation remains consistent. For example, when the same building is imaged on sunny and cloudy days, the feature response amplitudes of each channel may scale due to different light intensities, but the orientation of the feature vectors remains similar; in this case, cosine similarity can still provide a high similarity value.

[0085] In another implementation, Dot product similarity is used to directly calculate the inner product of two vectors. This method eliminates the need for magnitude calculation and division, resulting in lower computational cost. It is suitable for scenarios where the lighting conditions between preceding and following images differ little and the semantic feature magnitudes are relatively stable. This method requires the use of L2 norm regularization constraints during S2 training to ensure that the magnitudes of the semantic feature vectors approach 1. When both magnitudes are 1, dot product similarity is equivalent to cosine similarity, as both reflect directional consistency.

[0086] When the target region is large and the number of Gaussian elements is large, the search range for the maximum value can be limited. The weight value is taken with the coarse projection position as the center. The region exceeding the preset cutoff threshold is considered the valid search range. The cutoff threshold can be set to 0.1, corresponding to a position where the weight decays to approximately one-tenth of the center value.

[0087] Weighted similarity is calculated only at spatial locations within the effective search range, and the maximum value is taken. Regions outside the effective search range are not included in the calculation. Since the two-dimensional Gaussian distribution decays rapidly far from the center, and the effective search range only covers a small part of the weight map, limiting the search range reduces the number of spatial locations where similarity needs to be calculated for each Gaussian element, thereby reducing computational overhead.

[0088] S5: Determine the change in Gaussian elements.

[0089] The existence score calculated by S4 With preset threshold Compare. If If the Gaussian element is determined to be a changed Gaussian element, it means that the surface object corresponding to the Gaussian element has undergone a real change between the previous and next time phases. For example, the demolition of a building causes the disappearance of the feature at that location, or the construction of a new building on an open space causes a change in the original surface cover.

[0090] Preset threshold There are multiple ways to determine this: In one implementation, known unchanging surface areas within the target region are selected as invariant samples, such as city squares, main roads, or rooftop areas of large public buildings that have remained unchanged for many consecutive years. Existence scores of each Gaussian unit within the invariant sample are extracted, and the mean of these scores is calculated. and standard deviation .Will Set as This method allows the threshold to be adaptively adjusted based on the image quality, land cover type, and data characteristics of different target areas.

[0091] In another implementation, when reliable, unchanging ground feature samples are lacking within the target area, A pre-calibrated fixed value is used on an independent validation dataset. This validation dataset contains multiple sets of multi-temporal satellite image pairs with known change labels, and is similar to the target area in terms of sensor type, spatial resolution, and land cover type. Change detection is performed on the validation dataset with different threshold values. The detection precision and recall are calculated at each threshold, and the F1 score is calculated using the precision and recall. The threshold that maximizes the F1 score is selected as the pre-calibrated fixed value. The F1 score is the harmonic mean of precision and recall, which comprehensively reflects the change detection performance. When the land cover type and image characteristics of the target area differ significantly from those of the validation dataset, the former adaptive thresholding method is preferable.

[0092] S6: Output the three-dimensional detection results of surface changes.

[0093] Output all Gaussian elements in S5 that are determined to be changing.

[0094] In one implementation, all variable Gaussian primitives are extracted from the original 3D Gaussian sputtering model to form independent 3D point clouds. The extraction operation copies or removes all parameters of the variable Gaussian primitives from the original model's data structure and stores them separately as new point cloud data. Each variable Gaussian primitive carries its 3D center position, covariance matrix, and variation label. The 3D center position and covariance matrix of the Gaussian primitives labeled as variable together represent the 3D geometric contours of vanished surface objects in the previous time phase. The 3D center position determines the spatial coordinates of the feature, the scaling portion of the covariance matrix determines the extent of the Gaussian primitive in various spatial directions, and its orientation portion determines the spatial orientation of the Gaussian primitive; for example, a vertically oriented Gaussian primitive represents the facade of a building.

[0095] Furthermore, spatial density clustering can be performed on the set of changing Gaussian elements. A density-based clustering algorithm, such as DBSCAN, is used to determine whether elements belong to the same cluster based on the Euclidean distance between their three-dimensional center locations. Changing Gaussian elements with similar spatial locations are clustered together to distinguish different changing building targets and determine the spatial extent of each changing target, avoiding the confusion of adjacent but different building elements as the same target.

[0096] In another implementation, the complete structure of the original 3D Gaussian sputtering model is preserved without extraction or separation. Instead, a variation label field is added to each Gaussian primitive, with the value indicating either change or no change. The output is a complete set of Gaussian primitives, each carrying its geometric parameters, color parameters, semantic feature vector, and variation label.

[0097] By querying the change label field, changed hyperscale elements can be filtered out and displayed separately. Alternatively, changed and unchanged hyperscale elements can be rendered in different colors, presenting the before and after states in the same 3D view for comparative analysis. This facilitates the visualization of 3D changes in cities or the comparative analysis of changes before and after.

[0098] The entire detection process does not include pixel-level geometric registration between previous and subsequent temporal images. Change detection is achieved through cross-temporal retrieval of semantic feature vectors. Since semantic feature vectors learn to represent the physical material properties of ground features rather than their appearance under a specific viewpoint, they exhibit a certain degree of stability against changes in illumination and viewpoint. Relaxed spatial weights only provide spatial guidance and do not require precise matching of projection positions; therefore, the detection results are unaffected by differences in imaging viewpoints caused by building height.

[0099] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for detecting surface changes in multi-temporal satellite remote sensing images using 3D Gaussian sputtering, characterized in that, include: S1: Acquire at least two high-resolution satellite images of the target area from different perspectives in the earlier time phase and one high-resolution satellite image in the later time phase, wherein the earlier time phase image and the later time phase image constitute a multi-time phase image pair. S2: Train a three-dimensional Gaussian sputtering model using the multi-view satellite images from the previous time phase. The model consists of multiple Gaussian primitives, each of which has geometric parameters, color parameters, and semantic feature vectors. During training, the parameters of the Gaussian unit are jointly optimized using color reconstruction loss and semantic consistency loss, where the semantic consistency loss is the difference between the rendered semantic feature map and the semantic feature map extracted from the previous temporal images. S3: Input the satellite images from the later time phase into a pre-trained semantic coding network to obtain a two-dimensional semantic feature map; S4: For each Gaussian element in the model obtained in S2, calculate the similarity between its semantic feature vector and the semantic features of each pixel position in the two-dimensional semantic feature map; The similarity is weighted using a relaxed spatial weight determined by satellite imaging geometric parameters, and the global maximum value after weighting is taken as the existence score of the Gaussian element in the later temporal phase. S5: The Gaussian elements whose existence score is lower than a preset threshold are determined to be variable Gaussian elements; S6: Output all changes in Gaussian elements as 3D detection results of surface changes.

2. The method for detecting surface changes in multi-temporal satellite remote sensing images using 3D Gaussian sputtering according to claim 1, characterized in that, The parameters of the Gaussian element described in S2 are jointly optimized using the following loss function: ; in, L1 is the total loss, L2 is the color reconstruction loss, and L3 is the semantic consistency loss. These are the preset weighting coefficients.

3. The method for detecting surface changes in multi-temporal satellite remote sensing images using 3D Gaussian sputtering according to claim 1, characterized in that, The semantic feature vector described in S2 is obtained as follows: The semantic feature maps of the previous temporal images are extracted using a pre-trained semantic coding network. The semantic feature vectors are then projected onto each viewpoint using differentiable rendering to obtain rendered semantic feature maps. The difference between the two is used as the semantic consistency loss for optimization. The parameters of the pre-trained semantic coding network remain fixed during the S2 training process and the S3 feature extraction process, and it is the same network.

4. The method for detecting surface changes in multi-temporal satellite remote sensing images using 3D Gaussian sputtering according to claim 1, characterized in that, The loose space weights described in S4 are constructed as follows: Using the rational polynomial coefficient parameters of the preceding and following temporal images, the three-dimensional center position of the Gaussian element is projected onto the following temporal image plane to obtain a coarse projection position. A two-dimensional Gaussian weighted map is constructed with the rough projection location as the center. The standard deviation of the weighted map is determined by the maximum height of buildings in the target area and the range of differences in satellite viewpoints.

5. The method for detecting surface changes in multi-temporal satellite remote sensing images using 3D Gaussian sputtering according to claim 1, characterized in that, The existence score described in S4 is calculated using the following formula: ; in, For existence score, For semantic feature vectors, For position semantic features at the location For similarity function, The loose space weights at position The value at that location, For all spatial locations Take the maximum value.

6. The method for detecting surface changes in multi-temporal satellite remote sensing images using 3D Gaussian sputtering according to claim 5, characterized in that, The similarity function Cosine similarity is calculated by dividing the inner product of the two vectors by the product of their respective magnitudes.

7. The method for detecting surface changes in multi-temporal satellite remote sensing images using 3D Gaussian sputtering according to claim 1, characterized in that, The geometric parameters of the Gaussian elements in S2 include the three-dimensional center position and the covariance matrix, the color parameters are represented by spherical harmonic coefficients, and the dimension of the semantic feature vector is lower than the dimension of the color parameters.

8. The method for detecting surface changes in multi-temporal satellite remote sensing images using 3D Gaussian sputtering according to claim 1, characterized in that, The preceding and following time-phase images are high-resolution satellite images of urban areas, and the target area includes buildings with a height exceeding 20m. The detection process does not include pixel-level geometric registration between the preceding and following temporal images.

9. The method for detecting surface changes in multi-temporal satellite remote sensing images using 3D Gaussian sputtering according to claim 1, characterized in that, The three-dimensional detection result of surface changes described in S6 is a three-dimensional point cloud carrying change labels. The three-dimensional center position and covariance matrix of the Gaussian elements marked as changing represent the three-dimensional geometric contours of the disappeared surface objects in the previous time phase.

10. The method for detecting surface changes in multi-temporal satellite remote sensing images using 3D Gaussian sputtering according to claim 1 or 5, characterized in that, The preset threshold mentioned in S5 is determined based on the statistical distribution characteristics of the similarity function on unchanged ground features.