A method, equipment, medium, and product for detecting dynamic changes in aerospace three-dimensional scenes.

CN122200415BActive Publication Date: 2026-08-14HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-15
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

然而,二维方法难以克服成像角度差异导致的伪变化(如树木阴影移动、建筑物视角遮挡)

Benefits of technology

神经辐射场(Neural Radiance Fields,NeRF)作为一种隐式场景表示,能实现高质量的新视角合成,但是相关NeRF变化检测方法多通过渲染图像差异进行检测,缺乏显式的三维几何约束,导致变化定位精度不足。本申请提供了一种空天立体场景动态变化检测方法、设备、介质及产品,通过包括体渲染损失函数和一致性损失函数的总损失函数对时空神经占据场模型进行训练,能够提升变化定位精度;通过计算同一空间位置在不同时相占据栅格确定候选变化体素集合,并结合两个时相的渲染影像的光度差异确定高置信度变化区域,实现了三维空间内的高精度动态变化检测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122200415B_ABST
    Figure CN122200415B_ABST
Patent Text Reader

Abstract

This application discloses a method, device, medium, and product for detecting dynamic changes in aerospace stereoscopic scenes, relating to the fields of remote sensing image processing, computer vision, and 3D reconstruction. The method includes: acquiring aerospace stereoscopic image data of a target area at two different time phases; constructing a spatiotemporal neural occupancy field model for each time phase; training the spatiotemporal neural occupancy field model using a total loss function; extracting occupancy grids for each time phase based on the explicit occupancy probabilities of the two time phases to determine a candidate set of change voxels; generating rendered images of the two time phases from the same virtual viewpoint, calculating the luminosity difference between the two rendered images, and determining high-confidence change regions based on the candidate set of change voxels; and determining the dynamic change detection result based on the high-confidence change regions, thus achieving high-precision dynamic change detection in 3D space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of remote sensing image processing, computer vision and 3D reconstruction technology, and in particular to a method, device, medium and product for detecting dynamic changes in aerospace 3D scenes. Background Technology

[0002] Currently, change detection in aerospace remote sensing images mainly focuses on two-dimensional pixel-level analysis. However, two-dimensional methods struggle to overcome pseudo-changes caused by differences in imaging angles (such as shifting tree shadows or building viewpoint occlusion). With the development of 3D reconstruction technology, change detection methods based on point clouds or meshes have gradually emerged, but traditional multi-view stereo matching (MVS) is prone to generating holes in areas with weak texture and is difficult to generate images from new perspectives. Summary of the Invention

[0003] The purpose of this application is to provide a method, device, medium, and product for detecting dynamic changes in aerospace three-dimensional scenes, which can achieve high-precision dynamic change detection in three-dimensional space.

[0004] To achieve the above objectives, this application provides the following solution: Firstly, this application provides a method for detecting dynamic changes in aerospace three-dimensional scenes, including: Acquire aerospace stereo imagery data of the target area at two different time phases; the aerospace stereo imagery data includes several remote sensing images; For each time phase, a spatiotemporal neural occupancy field model is constructed; the spatiotemporal neural occupancy field model takes the three-dimensional spatial coordinates of the spatial location in the remote sensing image and the observation direction as input, and predicts the volume density, explicit occupancy probability and predicted color of the remote sensing image. Using the total loss function, the spatiotemporal neural occupancy model is spatiotemporally and jointly optimized and trained based on the aerospace stereo image data of each time phase to obtain the trained spatiotemporal neural occupancy model for each time phase; the total loss function includes a volume rendering loss function and a consistency loss function; the consistency loss function is used to constrain the consistency between explicit occupancy probability and implicit occupancy probability. Based on the explicit occupancy probabilities of remote sensing images from two time phases, the occupancy raster of each time phase is extracted to determine the candidate change voxel set. Using a pre-trained spatiotemporal neural occupancy field model corresponding to each time period, render images of two time periods under the same virtual view are generated. The luminance difference between the render images of the two time periods is calculated, and high-confidence change regions are determined based on the luminance difference and the candidate change voxel set. The dynamic change detection results of the target area are determined based on the high confidence change area.

[0005] Secondly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the processor executes the computer program to implement the steps of the above-described method for detecting dynamic changes in aerospace three-dimensional scenes.

[0006] Thirdly, this application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method for detecting dynamic changes in aerospace three-dimensional scenes.

[0007] Fourthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method for detecting dynamic changes in aerospace three-dimensional scenes.

[0008] According to the specific embodiments provided in this application, this application has the following technical effects: Neural Radiance Fields (NeRF), as an implicit scene representation, can achieve high-quality synthesis of new perspectives. However, most NeRF change detection methods rely on differences in rendered images, lacking explicit 3D geometric constraints, resulting in insufficient change localization accuracy. This application provides a method, device, medium, and product for dynamic change detection in aerospace 3D scenes. By training a spatiotemporal neural occupancy field model with a total loss function including a volume rendering loss function and a consistency loss function, the accuracy of change localization can be improved. By calculating the occupancy grid of the same spatial location in different temporal phases to determine the candidate set of change voxels, and combining the photometric differences between the rendered images of the two temporal phases to determine high-confidence change regions, high-precision dynamic change detection in 3D space is achieved. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly described below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort: Figure 1 This is an application environment diagram of a method for detecting dynamic changes in aerospace three-dimensional scenes according to an embodiment of this application; Figure 2 A flowchart illustrating a method for detecting dynamic changes in a three-dimensional aerospace scene according to an embodiment of this application; Figure 3 A schematic diagram illustrating the specific process of a method for detecting dynamic changes in a three-dimensional aerospace scene according to an embodiment of this application; Figure 4This is a schematic diagram of a spatiotemporal neural occupancy field model structure provided in an embodiment of this application; Figure 5 This is a schematic diagram illustrating the principle of spatiotemporal change detection according to an embodiment of this application; Figure 6 A schematic diagram of remote sensing images at two different time phases provided in an embodiment of this application; Figure 6 (a) in the middle is Temporal remote sensing images; Figure 6 (b) in the middle is Temporal remote sensing images; Figure 7 A schematic diagram of the rendering results for two time phases provided in an embodiment of this application; Figure 7 (a) in the middle is The rendering result of the time phase; Figure 7 (b) in the middle is The rendering result of the time phase; Figure 8 This is a schematic diagram of the change detection results provided in an embodiment of this application; Figure 8 In the diagram, (a) represents the difference in occupancy probability; Figure 8 (b) in the diagram represents the region of truth value variation; Figure 9 This is a schematic diagram of the training process loss curve provided in an embodiment of this application; Figure 9 (a) in the figure represents the rendering loss curve; Figure 9 (b) in the figure represents the consistency loss curve; Figure 10 This is a schematic diagram of the spatial distribution of a three-dimensional occupying grid provided in an embodiment of this application; Figure 10 (a) in the middle is The time phase occupies the grid; Figure 10 (b) in the middle is The time phase occupies the grid; Figure 10 (c) in the figure represents the difference in the number of occupancy grids between the two time phases; Figure 11 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0010] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0011] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0012] The aerospace 3D scene dynamic change detection method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be set up independently, integrated into server 104, or placed in the cloud or on other servers. Terminal 102 can send aerospace stereo image data to server 104. After receiving the aerospace stereo image data, server 104 constructs a spatiotemporal neural occupancy field model for each phase. Using a total loss function, it trains the spatiotemporal neural occupancy field model based on aerospace stereo image data from two different phases. Based on the explicit occupancy probabilities of the two phases, it extracts the occupancy grid for each phase, determines the candidate change voxel set, generates rendered images of the two phases from the same virtual viewpoint, calculates the luminosity difference between the two rendered images, and determines high-confidence change regions based on the candidate change voxel set. Based on the high-confidence change regions, it determines the dynamic change detection result. Server 104 can feed back the obtained dynamic change detection result to terminal 102. In addition, in some embodiments, the method for detecting dynamic changes in aerospace stereoscopic scenes can also be implemented by the server 104 or the terminal 102 separately. For example, the terminal 102 can directly perform dynamic changes detection of aerospace stereoscopic scenes on two different time phases of aerospace stereoscopic image data, or the server 104 can obtain two different time phases of aerospace stereoscopic image data from the data storage system and perform dynamic changes detection of aerospace stereoscopic scenes on the two different time phases of aerospace stereoscopic image data.

[0013] The terminal 102 can be, but is not limited to, various desktop computers and laptops. The server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers, or it can be a cloud server.

[0014] In one exemplary embodiment, such as Figure 2 As shown, a method for detecting dynamic changes in a three-dimensional aerospace scene is provided. This method is executed by a computer device, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, the method is applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps 201 to 206.

[0015] Step 201: Acquire aerospace stereo imagery data of the target area at two different time phases; the aerospace stereo imagery data includes several remote sensing images.

[0016] Step 202: For each time phase, construct a spatiotemporal neural occupancy field model; the spatiotemporal neural occupancy field model takes the three-dimensional spatial coordinates of the spatial location in the remote sensing image and the observation direction as input, and predicts the volume density, explicit occupancy probability and predicted color of the remote sensing image.

[0017] Step 203: Using the total loss function, the spatiotemporal neural occupancy model is spatiotemporally and jointly optimized and trained based on the aerospace stereo image data of each time phase to obtain the trained spatiotemporal neural occupancy model for each time phase; the total loss function includes a volume rendering loss function and a consistency loss function; the consistency loss function is used to constrain the consistency between explicit occupancy probability and implicit occupancy probability.

[0018] Step 204: Based on the explicit occupancy probabilities of the remote sensing images of the two time phases, extract the occupancy raster of the two time phases respectively to determine the candidate change voxel set.

[0019] Step 205: Use the pre-trained spatiotemporal neural occupancy field model corresponding to each time to generate rendered images of two time phases under the same virtual viewpoint, calculate the luminance difference between the rendered images of the two time phases, and determine the high confidence change region based on the luminance difference and the candidate change voxel set.

[0020] Step 206: Determine the dynamic change detection results of the target area based on the high-confidence change area.

[0021] By implementing steps 201 to 206 above, the implicit continuous representation of NeRF is combined with the explicit discrete representation of the spatial occupancy grid to construct a unified spatiotemporal field. The spatiotemporal neural occupancy field model is trained by a total loss function including a volume rendering loss function and a consistency loss function, which can improve the accuracy of change localization. By calculating the change in occupancy probability of the same spatial location under different temporal occupancy grids, the candidate set of change voxels is determined, and the high-confidence change region is determined by combining the photometric difference between the rendered images of the two temporal phases, thus realizing high-precision dynamic change detection in three-dimensional space.

[0022] like Figure 3 As shown, the specific process of the above method includes: data acquisition and preprocessing; construction of a three-dimensional NeRF-Occupancy unified field (i.e., construction of a spatiotemporal neural occupancy field model with two phases), corresponding to step 202 above; spatiotemporal joint optimization, corresponding to step 203 above; dynamic change detection, corresponding to steps 204-205 above; output and visualization.

[0023] This embodiment uses satellite stereo imagery of illegal construction monitoring in a certain urban area in 2019 and 2023 as an example, taking this area as the target area; dual-view stereo imagery of the target area is selected. 2019 is used as... The timeframe is based on 2023. Time phase: Acquire aerospace stereo imagery data at two different time phases. , for The first phase of time Zhang remote sensing images, The total number of remote sensing images represents the total number of images in the two-temporal-phase aerospace stereo imagery data, which includes image sets from both phases. Define 2019. The collection of images of the time phase is , They are respectively The first and second remote sensing images in the temporal image set; 2023 The collection of images of the time phase is , They are respectively The first and second remote sensing images in the temporal image set. The corresponding camera extrinsic parameter set is: , including rotation matrix Translation vector Points in the world coordinate system Mapping to the camera coordinate system, the mapping formula is as follows: , These are coordinates in the camera coordinate system. Radiometric correction and geometric registration are performed on all remote sensing images to eliminate atmospheric and system geometric errors and ensure consistency in color space across multiple temporal images. Correction can be achieved using an RPC (Rational Function Correction) model, calculated... Phase and The set of camera extrinsic parameters corresponding to the set of images in time phase and ,in, They are respectively Corresponding camera extrinsic parameters; They are respectively The corresponding camera extrinsic parameters. Each remote sensing image includes the three-dimensional spatial coordinates and observation direction of several spatial locations within the image.

[0024] Constructing a three-dimensional NeRF-Occupancy unified field: targeting the first Phase, construct the scene function corresponding to that phase (i.e., spatiotemporal neural occupancy field model). This function converts three-dimensional spatial coordinates. and two-dimensional observation direction Mapped to color c and volume density .

[0025] (1); in, A three-dimensional point in world coordinates. It is the set of real numbers; It is a unit direction vector; The predicted RGB color (i.e., the predicted color); For volume density, It is the set of positive real numbers. For the first Model parameters of the spatiotemporal neural occupancy field model of the time phase.

[0026] The spatiotemporal neural occupancy field model includes a feature extraction layer and a color output network. The color output network includes a density prediction head, an occupancy prediction head, and a color prediction head. The feature extraction layer extracts features from the three-dimensional spatial coordinates of several spatial locations in the remote sensing image, obtaining intermediate layer feature vectors. The color output network includes a density prediction head, an occupancy prediction head, and a color prediction head. The density prediction head and the occupancy prediction head predict the volume density and explicit occupancy probability of the remote sensing image based on the intermediate layer feature vectors, respectively. The color prediction head predicts the predicted color of the remote sensing image based on the intermediate layer feature vectors and the observation direction. Figure 4 As shown, the feature extraction layer (pre-layer) and the color prediction head (post-layer) constitute the MLP model. The feature extraction layer consists of the first N layers of the Multilayer Perceptron (MLP) model, and the color prediction head consists of the last KN layers of the MLP model. In this embodiment, the MLP model consists of 9 sequentially connected fully connected layers, and N can be 6. Therefore, the feature extraction layer consists of the first 6 fully connected layers of the MLP model, and the color prediction head consists of the last 3 fully connected layers of the MLP model. The density prediction head includes sequentially connected fully connected layers and ReLU activation function layers. The occupancy prediction head includes sequentially connected fully connected layers and Sigmoid activation function layers. This spatiotemporal neural occupancy field model first constructs a consistent neural radiation field from multiple perspectives using aerospace stereo imagery and derives an explicit spatial occupancy raster branch in the feature space. The implicit field is constrained by the volume rendering equation, and the explicit field is constrained by the occupancy probability, with both being jointly optimized. Based on this, dynamic change detection in three-dimensional space is achieved by calculating the difference in occupancy probability and the difference in rendering luminance between time sequences.

[0027] The model parameters of the feature extraction layer are This part maps the input spatial coordinates to intermediate layer feature vectors. , For the feature extraction layer, the first The weight matrix of the layer, For the feature extraction layer, the first The bias vector of the layer.

[0028] The model parameters of the density prediction head and the occupancy prediction head are expressed as follows: This part will input the intermediate layer feature vector Mapped to volume density respectively and explicit occupancy probability ; These are the weight matrix and bias vector of the fully connected layer in the density prediction head, respectively; These are the weight matrix and bias vector of the ReLU activation function layer in the density prediction head, respectively; These are the weight matrix and bias vector of the fully connected layer in the prediction head, respectively; These are the weight matrix and bias vector of the Sigmoid activation function layer in the prediction head, respectively.

[0029] The model parameters of the color prediction head are: This part will use the intermediate layer feature vectors The vector concatenated with the observation direction is mapped to the final predicted color. These are the weight matrix and bias vector of the color prediction head, respectively.

[0030] In summary, the model parameters of the spatiotemporal neural occupancy field model are the union of these three sets of parameters: ,in It indicates and.

[0031] The forward propagation process of the spatiotemporal neural occupancy field model includes the following steps 1 to 3.

[0032] Step 1, Feature Extraction: Convert coordinates Input the first 6 layers of the MLP model and extract the feature vectors of the intermediate layers. ; Step 2, Density and Occupation Prediction: Based on the intermediate layer feature vectors, the volume density is output through the density prediction head and the occupancy prediction head, respectively. and explicit occupancy probability .

[0033] (2); (3); In the formula, To occupy the grid The explicit occupancy probability value of the corresponding voxel represents the probability that the spatial location is occupied by an object; an explicit three-dimensional occupancy grid is embedded in the implicit field of the spatiotemporal neural occupancy field model. Let the voxel mesh size be... , These represent the length, width, and height of the grid cell, respectively; for any voxel center Its explicit occupancy probability It is predicted from the feature vectors of the intermediate layer; and These represent the fully connected layers of the density prediction head and the occupancy prediction head, respectively.

[0034] Step 3, Color Prediction: Concatenate the intermediate layer feature vector with the observation direction, input it into the color prediction head, and output the predicted color.

[0035] Spatiotemporal joint optimization: Define the total loss function while constraining the rendering quality of the spatiotemporal neural occupancy field model and the geometric consistency of the occupancy raster. Set the batch size to B=1024, and randomly sample the ray set R on each remote sensing image.

[0036] 1) Volume rendering loss function: for remote sensing images The rendered color of a certain pixel ray r in the image. Approximate calculation is performed using integration. Let the ray equation be r(s) = o + sd, where o is the optical center and s is the ray step size parameter. At the near-section... and far section The formula for calculating the rendered color is as follows: Several points are sampled between points. (4); in, The number of sampling points between the near section and the far section can be 64; For the first The cumulative transmittance at the sampling point (cumulative transmittance represents the optical fiber's velocity from the camera through the sampling point) is the same as the cumulative transmittance at the sampling point. Before each sampling point (The proportion of remaining light energy after each sampling point). ; For the first Volume density at each sampling point; For the first The sampling point and the first The distance between each sampling point; For the first Opacity of each sampling point; For the first The sampling point and the first The distance between each sampling point For the first The distance from each sampling point to the center of the camera For the first The distance from each sampling point to the center of the camera; For the first The predicted color for each sampling point.

[0037] Rendering loss function Defined as: (5); in, and Pixel rays in remote sensing images The rendered colors and the actual pixel colors; For light batches; This represents the 2-norm.

[0038] 2) Consistency loss: To correlate implicit density In addition to explicit occupancy, a consistency constraint is introduced. Theoretically, the volume density and occupancy probability of a point in space should satisfy a physical correlation. The formula for calculating the implicit occupancy probability is as follows: (6); in, It is the implicit occupancy probability derived from the volume density predicted by the model; This represents the distance between adjacent sampling points, where a sampling point is a point between the near-section and the far-section. voxels Volume density; voxels Volume density; It is a numerically stable term and is a very small positive number.

[0039] Consistency loss function It is expressed as follows: (7); in, voxels The corresponding explicit occupancy probability; For voxels The implicit occupancy probability is derived from the volume density integral. To occupy a grid.

[0040] 3) Solid geometric constraints: Utilizing conjugate point constraints of solid image pairs. For ray pairs with the same name... and The depth distribution at the spatial intersection should be consistent. A depth smoothing loss is introduced as a solid geometric constraint, and the solid geometric constraint function is expressed as follows: (8); in, For pixel light The expected depth is calculated using the following formula (9); For pixel rays The gradient of the expected depth; It represents the 1-norm.

[0041] (9); in, For the first The distance from each sampling point to the center of the camera.

[0042] The total loss function is expressed as follows: (10); in, This is the total loss function; , and These are the weight parameters for the volume rendering loss function, the consistency loss function, and the solid geometry constraint function, respectively.

[0043] The model parameters of the spatiotemporal neural occupancy field model for the two time phases are iteratively updated by the Adam optimizer until the loss converges (i.e. the value of the total loss function no longer decreases significantly), thus obtaining the trained spatiotemporal neural occupancy field model for each time phase.

[0044] In step 204 above, the occupancy grids of the two time phases are extracted based on the explicit occupancy probabilities of the remote sensing images of the two time phases to determine the candidate change voxel set. Specifically, this includes: extracting the occupancy grids of the two time phases based on the explicit occupancy probabilities of the remote sensing images of the two time phases; calculating the change in occupancy probability of the same spatial location under the occupancy grids of different time phases based on the occupancy grids of the two time phases; when the change in occupancy probability of the same spatial location under the occupancy grids of different time phases is greater than a set occupancy probability change threshold, the voxel corresponding to the spatial location is determined as a candidate change voxel; all candidate change voxels constitute the candidate change voxel set.

[0045] The principle of spatiotemporal change detection is as follows: Figure 5 As shown. Dynamic change detection: for two time phases and The trained model parameters were obtained respectively. , and the corresponding occupied grid , .

[0046] 3D Occupation Difference Detection: Calculating the change in occupancy probability at the same spatial location. : (11); in, , They are respectively in Phase and voxels (grids) at the same spatial location at the same time The explicit occupancy probability; if If the voxel is determined to be a candidate variable voxel, then a set of candidate variable voxels is obtained. , , To set a threshold for the change in occupancy probability, in a specific example, A value of 0.6 can be taken. In this embodiment, if the change in the occupancy probability of a certain area is found to be significantly greater than the set threshold for the change in occupancy probability, it indicates that the area is a newly added building.

[0047] Photometric consistency verification: To eliminate spurious changes caused by seasonal variations or lighting, a trained spatiotemporal neural occupancy field model is used to generate rendered images from the same virtual viewpoint. The virtual camera pose (i.e., the virtual top-down viewpoint) is then set. Render the images of the two phases separately.

[0048] The image rendering process is as follows: 1) Setting up the virtual camera: Determine the camera pose (intrinsic and extrinsic parameters) for the virtual viewpoint; 2) Generating rays: For each pixel of the virtual viewpoint image, generate a ray r that originates from the optical center and passes through the pixel based on the camera parameters; 3) Ray sampling: Sample N points along the ray r between the near and far planes to obtain the 3D spatial coordinates x and the observation direction d of these sampling points; 4) Model inference: Input the 3D spatial coordinates and observation direction d into the trained spatiotemporal neural occupancy field model, and the model outputs the predicted color c and volume density for each sampling point. 5) Volume rendering integral: Using the volume rendering equation, combined with predicted color and volume density, calculate the cumulative transmittance and weight of the ray, and finally sum the weighted values ​​to obtain the rendered color value of the pixel. 6) Image generation: Repeat the above process for all pixels, and the resulting color value matrix is ​​the rendered image under the virtual viewpoint.

[0049] Calculate the luminosity difference between the rendered images from two different time phases: (12); in, The difference in luminosity between the rendered images of the two time phases; for Rendered images of time phases; for The rendered image of the time phase.

[0050] Define the set of pixels that show significant changes on the image plane. for: (13); in, To set a difference threshold.

[0051] The luminosity difference between the rendered images from two different time phases is back-projected into 3D space and compared with the candidate set of variable voxels. Find the intersection to obtain the high-confidence variation region. This involves projecting the candidate set of variable voxels onto the image plane, retaining only the voxels whose projection points fall on the image plane. Voxels of the region yield high-confidence variation regions: (14); in, voxels The projection coordinates of the center on the virtual mirror plane.

[0052] In step 205 above, the high-confidence change region is determined based on the photometric difference and the candidate change voxel set. Specifically, this includes: back-projecting the photometric difference of the rendered images of the two time phases into three-dimensional space, retaining voxels in the candidate change voxel set whose photometric difference is greater than a set difference threshold, and determining the regions corresponding to all voxels whose photometric difference is greater than the set difference threshold as high-confidence change regions.

[0053] In step 206 above, determining the dynamic change detection result of the target area based on the high confidence change area specifically includes: converting the high confidence change area into a three-dimensional point cloud or mesh model with geographic coordinates, and determining the dynamic change detection result of the target area.

[0054] Output and Visualization: Convert high-confidence variation areas into 3D point clouds or mesh models with geographic coordinates. Converted to OBJ format and overlaid onto a GIS map, the area was confirmed as an illegal construction zone. Simultaneously, the user can specify any flight path, and the system utilizes two corresponding trained spatiotemporal neural occupancy field models. and Real-time generation of dual-time comparison videos.

[0055] To verify the effectiveness of the method in this application, this embodiment constructs a simulation experiment system based on the PyTorch framework to simulate a scenario of detecting changes in aerospace stereo image data.

[0056] Simulation Environment Configuration: The simulation experiment uses a Python 3.10 environment, with main dependencies including PyTorch 2.0, NumPy, and Matplotlib. The hardware platform is a standard CPU, fully verifying the universality of the method. Network parameters are set as follows: MLP layer count D=6, number of channels per layer W=256; number of light sampling points... =32; Batch size B=1024; Learning rate Number of training iterations =800.

[0057] Simulation data design: Two phases of satellite imagery data were simulated and generated, with an image size of 64×64 pixels. Phase (2019): Includes a dark green ground background with RGB values ​​of approximately (0.15, 0.35, 0.15). Timeframe (2023): Three new building areas were added at the same location: 1. A large red rectangular building (20×20), with RGB values ​​approximately (0.9, 0.15, 0.15); 2. A medium blue rectangular building (20×18), with RGB values ​​approximately (0.15, 0.15, 0.9); 3. A small yellow rectangular building (13×13), with RGB values ​​approximately (0.9, 0.85, 0.15). Remote sensing imagery from both timeframes is shown below. Figure 6 As shown.

[0058] The model successfully reconstructed the scene's appearance information using volume rendering equations. The rendered image was highly consistent with the input image, validating the network's learning ability. The rendering results for the two time phases are as follows: Figure 7 As shown. The change detection results are as follows. Figure 8 As shown, the occupancy probability difference is displayed in the form of a heatmap: it shows the difference in occupancy grid values ​​between two time phases. The highlighted areas in the heatmap represent high-probability change locations, which highly match the actual building locations. The true value change area shows the change area (white area) after thresholding, clearly marking the locations of the three newly added buildings.

[0059] from Figure 9 It can be observed that the rendering loss decreases rapidly with the number of iterations and tends to converge, indicating that the model has successfully learned the texture information of the scene. The consistency loss also shows a decreasing trend, verifying the effectiveness of the physical consistency constraint between the explicit occupancy raster and the implicit density field. The models in both time phases reach convergence, and the loss values ​​are similar, indicating that the model has good adaptability to data from different time phases.

[0060] Depend on Figure 10 It can be clearly seen The timing phase showed a significant increase in the probability of occupancy at the building location, among which, Figure 10 (c) in the figure shows the difference in the grid occupancy between the two time phases, with the highlighted area consistent with the actual change position.

[0061] Simulation experiments show that the proposed aerospace 3D scene dynamic change detection method has the following advantages: 1) It can accurately detect changes in newly added buildings in 3D space. Compared with traditional 2D change detection, this method directly outputs the change results in 3D space, overcoming projection distortion and achieving high positioning accuracy; 2) The fusion of explicit occupancy grid and implicit NeRF field effectively improves geometric constraints; strong anti-interference ability: through the multi-view synthesis capability of NeRF, the influence of non-structural changes such as shadows and illumination can be eliminated; 3) The loss function is reasonably designed and the model converges stably; through the consistency constraint of explicit occupancy grid and implicit density field, the accuracy of 3D geometry is enhanced; 4) The method has good adaptability to simulated remote sensing image data; 5) High integrity: by utilizing the implicit interpolation capability of NeRF, the holes in the weak texture area of ​​traditional MVS point cloud are filled, making the change detection more complete.

[0062] This application also provides an application scenario in which the above-mentioned aerospace stereoscopic scene dynamic change detection method is applied. Specifically, the aerospace stereoscopic scene dynamic change detection method provided in this embodiment can be applied to an illegal building identification scenario. The illegal building identification scenario includes an image acquisition stage, an aerospace stereoscopic scene dynamic change detection link, and an illegal building identification stage; aerospace stereoscopic image data enters the aerospace stereoscopic scene dynamic change detection link from the image acquisition stage, obtains the corresponding dynamic change detection results through human-machine collaboration, and then enters the downstream illegal building identification stage. The aerospace stereoscopic scene dynamic change detection method provided in this embodiment belongs to the aerospace stereoscopic scene dynamic change detection link. Specifically, in the process of detecting dynamic changes in aerospace stereo scenes using aerospace stereo image data, a spatiotemporal neural occupancy field model can be constructed for each time phase. The spatiotemporal neural occupancy field model is trained using a total loss function based on aerospace stereo image data from two different time phases. Based on the explicit occupancy probabilities of the two time phases, the occupancy grid of each time phase is extracted, a candidate set of change voxels is determined, and rendered images of the two time phases are generated under the same virtual viewpoint. The luminosity difference between the rendered images of the two time phases is calculated, and a high-confidence change region is determined based on the candidate set of change voxels. The dynamic change detection result is then determined based on the high-confidence change region.

[0063] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 11As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs in the non-volatile storage media to run. The database stores dynamic change detection data for aerospace stereoscopic scenes. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for detecting dynamic changes in aerospace stereoscopic scenes.

[0064] Those skilled in the art will understand that Figure 11 The structures shown are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0065] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0066] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0067] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0068] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0069] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0070] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for detecting dynamic changes in a three-dimensional aerospace scene, characterized in that, The method includes: Acquire aerospace stereo imagery data of the target area at two different time phases; the aerospace stereo imagery data includes several remote sensing images; For each time phase, a spatiotemporal neural occupancy field model is constructed; the spatiotemporal neural occupancy field model takes the three-dimensional spatial coordinates of the spatial location in the remote sensing image and the observation direction as input, and predicts the volume density, explicit occupancy probability and predicted color of the remote sensing image. Using the total loss function, the spatiotemporal neural occupancy model is spatiotemporally jointly optimized and trained based on the aerospace stereo image data of each time phase, resulting in a trained spatiotemporal neural occupancy model for each time phase. The total loss function includes a volume rendering loss function and a consistency loss function. The consistency loss function is used to constrain the consistency between explicit occupancy probabilities and implicit occupancy probabilities. The formula for calculating the implicit occupancy probability is as follows: ; in, Represents the natural base; This represents the distance between adjacent sampling points, where a sampling point is a point between the near-section and the far-section. voxels Volume density; ; in, For consistency loss function; The explicit occupancy probability corresponding to the voxel; For voxels The implicit occupancy probability is derived from the volume density integral. To occupy a grid; Based on the explicit occupancy probabilities of remote sensing images from two time phases, the occupancy raster of each time phase is extracted to determine the candidate change voxel set. Using a pre-trained spatiotemporal neural occupancy field model corresponding to each time period, render images of two time phases under the same virtual viewpoint are generated. The luminance difference between the render images of the two time phases is calculated, and high-confidence change regions are determined based on the luminance difference and the candidate set of change voxels. Specifically, the determination of high-confidence change regions based on the luminance difference and the candidate set of change voxels includes: projecting the candidate set of change voxels onto the image plane, retaining voxels in the candidate set of change voxels whose luminance difference is greater than a set difference threshold, and determining the regions corresponding to all voxels whose luminance difference is greater than the set difference threshold as high-confidence change regions. The dynamic change detection results of the target area are determined based on the high confidence change area.

2. The method for detecting dynamic changes in aerospace three-dimensional scenes according to claim 1, characterized in that, The spatiotemporal neural occupancy model includes a feature extraction layer and a color output network; the color output network includes a density prediction head, an occupancy prediction head, and a color prediction head; the feature extraction layer and the color prediction head constitute an MLP model, the feature extraction layer consists of the first N layers of the MLP model, and the color prediction head consists of the last KN layers of the MLP model; the density prediction head includes a fully connected layer and a ReLU activation function layer connected in sequence; the occupancy prediction head includes a fully connected layer and a Sigmoid activation function layer connected in sequence.

3. The method for detecting dynamic changes in aerospace three-dimensional scenes according to claim 1, characterized in that, The total loss function is expressed as follows: ; ; ; in, This is the total loss function; , and Volume rendering loss function Consistency loss function and solid geometry constraint functions Weight parameters; and Pixel rays in remote sensing images The rendered color and the pixel's true color; For light batches; Represents the 2-norm; Pixel Ray Expected depth; For the ray along the pixel The gradient of the expected depth; It represents the 1-norm.

4. The method for detecting dynamic changes in aerospace three-dimensional scenes according to claim 2, characterized in that, The formula for calculating rendered colors is as follows: ; in, This represents the number of sampling points between the near-section and the far-section. For the first Cumulative transmittance at each sampling point; Opacity; For the first The predicted color for each sampling point.

5. The method for detecting dynamic changes in aerospace three-dimensional scenes according to claim 1, characterized in that, Based on the explicit occupancy probabilities of remote sensing images from two different time phases, occupancy gratings for each phase are extracted to determine the candidate change voxel set, specifically including: Based on the explicit occupancy probabilities of remote sensing images from two different time phases, the occupancy grids for the two time phases are extracted respectively. Calculate the change in occupancy probability of the same spatial location under different occupancy grids in two time phases; When the change in the occupancy probability of the same spatial location under different temporal phases is greater than the set threshold for the change in occupancy probability, the voxel corresponding to the spatial location is determined as a candidate variable voxel; all candidate variable voxels constitute a candidate variable voxel set.

6. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that the processor executes the computer program to implement the method for detecting dynamic changes in aerospace three-dimensional scenes according to any one of claims 1-5.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the method for detecting dynamic changes in aerospace three-dimensional scenes as described in any one of claims 1-5.

8. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the method for detecting dynamic changes in aerospace three-dimensional scenes as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Method and device for three-dimensional reconstruction of remote sensing image

    CN117765172A

  • Change detection method based on nerve radiation field

    CN119559410A