Local dynamic updating method and system for three-dimensional geographic scene
By combining a multi-head self-attention mechanism and a neural radiation field diffusion model, the problem of insufficient accuracy and coherence in traditional 3D geographic scene update methods in complex environments is solved, and efficient and accurate local dynamic updates of urban geographic scenes are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-03-31
AI Technical Summary
Traditional 3D geographic scene update methods are weak in updating local scenes under complex spatial interference factors, and ignore the projection relationship between different viewpoints during texture restoration, resulting in insufficient overall coherence and accuracy of the updated 3D geographic scene.
A multi-head self-attention mechanism is used to capture long-range contextual dependencies in urban geographic scene information datasets. A local dynamic update model for the scene is constructed by combining neural radiation field and diffusion model. The model is pre-trained using multi-view image data and laser point cloud data to generate a lightweight scene update package for local dynamic updates.
It improves the accuracy and multi-view consistency of 3D geographic scene updates, ensures the rationality and coherence of generated data, and achieves efficient dynamic updates of urban geographic scenes.
Smart Images

Figure CN121767577A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of 3D scene construction technology, and more specifically, to a method and system for local dynamic updating of 3D geographic scenes. Background Technology
[0002] With the increasing demands for timeliness and accuracy of spatial data in fields such as digital twin cities and smart city planning, 3D geographic scenes, as an important carrier for intuitively presenting urban forms and supporting decision analysis, have gradually become an important indicator for measuring the level of smart city construction due to their dynamic updating capability. Because urban features are constantly changing, 3D geographic scenes need to promptly perceive these changes and update them dynamically. However, traditional 3D geographic scene updating methods still have room for improvement in terms of detail accuracy, structural rationality, and multi-view coherence.
[0003] On the one hand, traditional methods mostly employ single-scale feature extraction, such as methods based on fixed-resolution 2D feature map construction. Single-scale feature extraction methods often struggle to accurately extract detailed features of small-scale geographic entities and fail to fully represent the overall structure of large-scale geographic objects, often resulting in insufficient accuracy of the extracted data. On the other hand, when using motion recovery structures for 3D geographic scene reconstruction, traditional solutions are prone to affecting the integrity and geometric accuracy of the reconstructed model in scenarios such as building occlusion and tree cover. Furthermore, when introducing generative diffusion models for texture restoration, they often neglect the projection relationship between different viewpoints, causing color inconsistencies and detail misalignments in the generated textures when viewed across different viewpoints, thus resulting in certain defects in the overall coherence of the updated 3D geographic scene. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a method and system for local dynamic updating of three-dimensional geographic scenes. This addresses the technical issues in existing technologies, such as the weak anti-interference ability of local scene updates under complex spatial interference factors and the tendency to neglect the projection relationship between different viewpoints when performing texture restoration.
[0005] The purpose and effectiveness of the present invention, which provides a method and system for local dynamic updating of a three-dimensional geographic scene, are achieved through the following specific technical means:
[0006] A method for local dynamic updating of a 3D geographic scene includes:
[0007] Data is collected from the three-dimensional geospatial scene of the city to obtain real-time urban geographic scene information data and construct the corresponding feature pyramid to generate an urban geographic scene information dataset.
[0008] Based on the urban geographic scene update recognition mechanism, scene update recognition is performed on the urban geographic scene information dataset to obtain scene update recognition results;
[0009] A scene local dynamic update model is constructed based on the neural radiation field and diffusion model, and the model is pre-trained.
[0010] The scene update recognition results are input into the scene local dynamic update model to update the urban geographic scene and obtain urban geographic scene update data.
[0011] The city geographic scene update data is packaged into a lightweight scene update package and sent to the client for local dynamic scene updates.
[0012] As a further aspect of the present invention, scene update recognition is performed on an urban geographic scene information dataset based on an urban geographic scene update recognition mechanism to obtain scene update recognition results, including:
[0013] By using a multi-head self-attention mechanism to capture long-distance contextual dependencies in urban geographic scene information datasets, semantic relationships between spatial features in urban geographic scene information datasets can be obtained.
[0014] A historical urban geographic scene information dataset is acquired and used as a benchmark for scene update identification. By analyzing the cross-correlation between the scene update identification benchmark and the urban geographic scene information dataset, a scene update attention weight map is generated. The scene update attention weight map is used to label regional information where the urban geographic scene may change.
[0015] Based on the scene update attention weight map, the urban geographic scene information dataset and the scene update recognition benchmark are weighted and fused to generate a binary change mask, a multi-class semantic segmentation map, change confidence and change type data, which are then output as the scene update recognition result.
[0016] Among them, the binary change mask is used to label whether the urban geographic scene has changed, the multi-class semantic segmentation map is used to label different urban geographic scene categories, the change confidence is used to quantify the reliability of changes in the urban geographic scene, and the change type data is used to describe the types of changes in the urban geographic scene.
[0017] As a further aspect of the present invention, a scene update attention weight map is generated by analyzing the cross-correlation between the scene update recognition benchmark and the urban geographic scene information dataset, including:
[0018] Feature extraction is performed on the scene update recognition benchmark and the urban geographic scene information dataset to generate updated feature maps and historical feature maps;
[0019] Based on a preset projection mechanism, the updated feature map and the historical feature map are mapped to the geographic scene feature space and then L2 normalized.
[0020] Centered on each pixel of the projected updated feature map and the historical feature map, local 3D feature hypercubes of two preset sizes, namely the first scale and the second scale, are extracted to construct the first scale local 3D feature hypercube group and the second scale local 3D feature hypercube group.
[0021] Each local 3D feature hypercube group contains a local 3D feature hypercube of the same size extracted from the updated feature map and the historical feature map, respectively.
[0022] Geographic scene discrimination information is extracted from each local 3D feature hypercube group using a micro multilayer perceptron. The geographic scene discrimination information is multiplied element-wise to obtain a structural similarity tensor. The structural similarity tensor is then subjected to global average pooling in both the channel dimension and the spatial dimension to obtain similarity response values.
[0023] Adaptively weighted fusion is performed on the similarity response values corresponding to the first-scale local 3D feature hypercube group and the second-scale local 3D feature hypercube group of the same pixel to generate a fused similarity response value.
[0024] Based on the fusion similarity response value corresponding to each pixel, an initial scene update attention weight map is constructed, and the initial scene update attention weight map is processed by an activation function to generate a scene update attention weight map.
[0025] As a further aspect of the present invention, a scene local dynamic update model is constructed based on the neural radiation field and diffusion model, and the model is pre-trained, including:
[0026] The scene local dynamic update model includes a neural radiation field submodule and a multi-view diffusion network submodule.
[0027] A first training dataset consisting of multi-view image data of urban buildings and corresponding laser point cloud data is obtained. The multi-view image data of urban buildings includes at least five perspectives of urban buildings: front, back, left, right and top. The neural radiation field submodule is pre-trained using a scene structure training strategy.
[0028] A second training dataset consisting of building texture images under different lighting conditions, weather conditions, and building structure occlusion conditions was obtained, and a building texture training strategy was used to pre-train the multi-view diffusion network submodule.
[0029] A third training dataset consisting of real urban geographic scene update cases was obtained, and the neural radiation field submodule and the multi-view diffusion network submodule were jointly pre-trained by combining scene structure loss function, building texture loss function and cross-modal consistency loss function.
[0030] As a further aspect of the present invention, the method further includes:
[0031] The scene structure training strategy is expressed as follows: dynamically adjusting the view acquisition method from the first training dataset and the number of views input to the neural radiation field submodule, and weighting and fusing the neural radiation field loss function with the BIM-based geometric constraint loss function to construct a scene structure loss function, and pre-training based on the scene structure loss function;
[0032] The architectural texture training strategy is to construct an architectural texture loss function by weighted fusion of the L2 loss function and the consistency loss function, perform pre-training based on the architectural texture loss function, and gradually increase the weight of the consistency loss function from the initial set value to the predetermined consistency weight threshold as the pre-training process progresses. The consistency loss function is composed of a weighted fusion of the L1 loss function and the epipolar constraint loss function.
[0033] As a further aspect of the present invention, the scene update recognition result is input into the scene local dynamic update model to perform urban geographic scene update, thereby obtaining urban geographic scene update data, including:
[0034] The neural radiation field submodule of the scene local dynamic update model is used to analyze the urban scene structure of the scene update recognition results and obtain urban scene structure data.
[0035] The multi-view diffusion network submodule of the scene local dynamic update model extracts texture features from the scene update recognition results to obtain building texture feature data.
[0036] Urban scene structure data and building texture feature data are fused to generate updated urban geographic scene data.
[0037] As a further aspect of the present invention, urban geographic scene update data is encapsulated into a lightweight scene update package and sent to the client for local dynamic scene updates, including:
[0038] The urban scene structure data in the urban geographic scene update data is subjected to grid compression using the Draco algorithm, and its vertex coordinates are quantized using an octree.
[0039] Texture encoding is performed on the architectural texture feature data in the updated urban geographic scene data.
[0040] Based on urban geographic scene update data, historical urban geographic scenes are updated to generate updated urban geographic scenes.
[0041] The changes in vertex coordinates and patch topology of the corresponding regions in the updated urban geographic scene and the historical urban geographic scene are calculated, patch difference data is obtained, and the update operation type is recorded by operation code. The update operation type includes at least add, delete and modify operations.
[0042] The processed vertex coordinates, face difference data, and opcodes are packaged into a lightweight scene update package and sent to the client.
[0043] As a further aspect of the present invention, obtaining patch difference data and recording the update operation type via operation codes includes:
[0044] The lightweight scene update package contains forward update instructions from the updated city geo scene to the historical city geo scene, as well as reverse rollback instructions from the historical city geo scene to the updated city geo scene.
[0045] The farthest point sampling algorithm is used to sample multiple vertices from the updated urban geographic scene and the historical urban geographic scene respectively. The transformation matrix between the updated urban geographic scene and the historical urban geographic scene is obtained by the iterative nearest point algorithm. Based on the transformation matrix, vertex matching is performed on the updated urban geographic scene and the historical urban geographic scene. The difference transformation of the region corresponding to the vertex coordinates is recorded synchronously to obtain the patch difference data. The corresponding operation code is generated according to the preset rules to generate a positive update instruction.
[0046] The forward update instruction is reverse-mapped according to the preset reverse mapping rules to generate a reverse rollback instruction. The forward update instruction and the reverse rollback instruction are bound by a preset attribute field. The forward update instruction and the reverse rollback instruction can be switched by changing the value of the preset attribute field.
[0047] As a further aspect of the present invention, data is collected from the three-dimensional geospatial scene of the city to obtain real-time urban geographic scene information data and construct a corresponding feature pyramid to generate an urban geographic scene information dataset, including:
[0048] Multi-source data collection is performed on urban areas to obtain real-time urban geographic scene information data, which includes at least RGB optical image data, multispectral image data, and laser point cloud data.
[0049] A combination of the dark target method and the atmospheric scattering model is used to perform radiometric correction on RGB optical image data, and the corresponding color and texture features are extracted to generate an RGB optical image feature pyramid, which is then output as the RGB optical image feature set.
[0050] The multispectral image data is atmospherically corrected based on the preset atmospheric correction module, generating a multispectral image feature pyramid, which is then output as a multispectral image feature set.
[0051] The laser point cloud data is denoised by combining statistical filtering and radius filtering, and the corresponding digital surface model, normal vector map and intensity map are obtained. Multiple laser point cloud feature pyramids are generated and output as laser point cloud feature sets.
[0052] A feature pyramid of urban geographic scene information is constructed based on RGB optical image feature set, multispectral image feature set, and laser point cloud feature set.
[0053] A local dynamic update system for a three-dimensional geographic scene includes:
[0054] The data acquisition module is used to acquire real-time urban geographic scene information data and construct a corresponding feature pyramid to generate an urban geographic scene information dataset.
[0055] An update recognition module is used to perform scene update recognition on an urban geographic scene information dataset and obtain scene update recognition results.
[0056] The model building module is used to build a scene local dynamic update model based on the neural radiation field and diffusion model, and to perform model pre-training;
[0057] The scene update module is used to input the scene update recognition result into the scene local dynamic update model to update the urban geographic scene and obtain urban geographic scene update data.
[0058] A 3D construction module is used to encapsulate urban geographic scene update data into a lightweight scene update package and send it to the client for local dynamic scene updates.
[0059] Based on the above, the embodiments of this application realize the data collection of the three-dimensional geospatial scene of the city, obtain real-time urban geographic scene information data and construct the corresponding feature pyramid, generate urban geographic scene information dataset, and improve the accuracy of feature extraction of the collected data by constructing the feature pyramid.
[0060] Based on the urban geographic scene update recognition mechanism, scene update recognition is performed on the urban geographic scene information dataset to obtain scene update recognition results. In the process of scene update recognition, the three-dimensional structural features of the urban geographic scene information dataset are extracted in the form of local three-dimensional feature hypercube, which improves the accuracy of scene update recognition and provides a comprehensive and accurate data foundation for subsequent data analysis.
[0061] A scene local dynamic update model is constructed based on the neural radiation field and diffusion model, and the model is pre-trained. The scene update recognition results are input into the scene local dynamic update model to update the urban geographic scene and obtain urban geographic scene update data. The neural radiation field sub-module and the multi-view diffusion network sub-module are constructed through the neural radiation field and diffusion model respectively. By pre-training the two sub-modules respectively, the model's ability to repair under complex occlusion and maintain the consistency of the generated data from multiple perspectives is trained. Finally, the conflict between the two sub-modules is eliminated by joint training, thereby ensuring the rationality of the generated urban geographic scene update data.
[0062] Urban geographic scene update data is packaged into a lightweight scene update package and sent to the client, thereby enabling local dynamic updates of the 3D geographic scene. Attached Figure Description
[0063] Figure 1 This is a schematic diagram of the execution flow of a method for local dynamic updating of a three-dimensional geographic scene provided in an embodiment of the present invention;
[0064] Figure 2 This is a schematic diagram of a local dynamic update system for a three-dimensional geographic scene provided in an embodiment of the present invention. Detailed Implementation
[0065] The embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are used to illustrate the technical solutions of the present invention, but should not be used to limit the scope of protection of the present invention.
[0066] Example:
[0067] As attached Figure 1 , Figure 2 As shown:
[0068] This invention provides a method for local dynamic updating of a three-dimensional geographic scene, applicable to three-dimensional scene construction, comprising the following steps:
[0069] Step S1: Collect data on the three-dimensional geospatial scene of the city, obtain real-time urban geographic scene information data, construct the corresponding feature pyramid, and generate an urban geographic scene information dataset.
[0070] Specifically, multi-source data is collected from urban areas to obtain real-time urban geographic scene information data. The real-time urban geographic scene information data includes at least RGB optical image data, multispectral image data, and laser point cloud data, and historical data corresponding to the real-time urban geographic scene information data is also obtained.
[0071] A combination of the dark target method and the atmospheric scattering model was used to perform radiometric correction on RGB optical image data, and the corresponding color and texture features were extracted to generate an RGB optical image feature set.
[0072] Understandably, the dark target method is used to extract dark target regions, such as shadow areas and vegetation-covered areas, from RGB optical image data. Assuming that the spectral reflectance of the pixels corresponding to the dark target regions is close to zero, aerosol optical thickness inversion is performed to obtain initial aerosol optical thickness data. The initial aerosol optical thickness data is then input into an atmospheric scattering model, and by solving the radiative transfer equation, atmospheric path radiation such as atmospheric molecular scattering and aerosol scattering is separated, thereby performing radiative correction.
[0073] Using ResNet-50 as the backbone network, multiple feature maps with resolutions ranging from small to large are extracted sequentially from RGB optical image data. Based on these feature maps, an RGB optical image feature pyramid with a feature pyramid structure is formed. The RGB optical image feature pyramid is output as an RGB optical image feature set. For example, RGB optical image feature maps with resolutions of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original RGB optical image resolution are collected respectively. These five feature maps form a feature pyramid structure with resolutions ranging from low to high, thereby enabling the capture of urban geographic scene information at different scales.
[0074] Atmospheric correction is performed on multispectral image data based on a preset atmospheric correction module to generate a multispectral image feature set.
[0075] In one possible embodiment, atmospheric correction is performed on the multispectral image data using the FLAASH module of ENVI software. This module, based on radiative transfer theory, accurately removes the influence of atmospheric components such as water vapor and ozone on spectral absorption, obtains the NDVI and NDWI corresponding to the atmospherically corrected multispectral image data, and marks multiple feature regions in the multispectral image data, such as vegetation areas and water bodies, using a threshold segmentation method. The multispectral image data after marking feature regions is used as the feature map of the lowest level of the multispectral image feature pyramid. The resolution of this layer's feature map is reduced to 1 / 2 of the original resolution using bilinear interpolation, and feature extraction is performed using a convolutional neural network to obtain a new feature map, which is then used as the feature map of the next higher level. The above bilinear interpolation and feature extraction operations are repeated until the multispectral image feature pyramid is constructed. The multispectral image feature pyramid is then output as a multispectral image feature set. For example, to construct a multispectral image feature pyramid with 5 levels, the bilinear interpolation and feature extraction operations need to be repeated 4 times.
[0076] The laser point cloud data is denoised by combining statistical filtering and radius filtering, and the corresponding digital surface model, normal vector map and intensity map are obtained to generate a laser point cloud feature set.
[0077] In one possible embodiment, a statistical filtering algorithm is used to remove outliers from the laser point cloud data. For example, the number of neighboring points is set to 15 and the standard deviation multiple is 2.5. The 3σ principle is used to statistically test the Euclidean distance between each spatial point and its neighboring points in the laser point cloud data, identify and remove abnormal data points that deviate from the normal distribution characteristics, such as non-target dynamic interference points generated by birds or clouds, thereby reducing data noise.
[0078] Radius filtering is defined as traversing the laser point cloud dataset using a KD-Tree to retrieve the number of neighboring points of each point within a specified radius. Only core points with a neighboring point count not less than a preset minimum neighboring point count threshold are retained as valid data. Radius filtering can effectively extract point cloud data of key geographic elements such as building facades and road paving surfaces, thereby improving the geometric representation accuracy of 3D scenes.
[0079] Based on the filtered laser point cloud data, a digital surface model, normal vector map, and intensity map are generated. The corresponding laser point cloud feature pyramids are generated according to the above bilinear interpolation and feature extraction operations, and the laser point cloud feature pyramids are output as the laser point cloud feature set.
[0080] A feature pyramid of urban geographic scene information is constructed based on RGB optical image feature set, multispectral image feature set, and laser point cloud feature set.
[0081] Specifically, data standardization is achieved by unifying the spatial resolution, data format, and storage structure of RGB optical image feature sets, multispectral image feature sets, and laser point cloud feature sets, thereby generating urban geographic scene information datasets.
[0082] Step S2: Based on the urban geographic scene update recognition mechanism, perform scene update recognition on the urban geographic scene information dataset and obtain the scene update recognition results.
[0083] In this embodiment, step S2 includes:
[0084] Step S21: Obtain the semantic relationships between spatial features in the urban geographic scene information dataset.
[0085] Specifically, a multi-head self-attention mechanism is used to capture long-distance contextual dependencies in urban geographic scene information datasets, thereby obtaining semantic associations between spatial features in the urban geographic scene information datasets.
[0086] In one possible embodiment, a multi-head self-attention mechanism is used to map each feature vector in the urban geographic scene information dataset to multiple query vector spaces, key vector spaces, and value vector spaces to generate feature vectors in the corresponding vector spaces. The dot product similarity between the query feature vector and all key feature vectors is obtained and normalized to generate corresponding weight coefficients. The weight coefficients are used to quantify the degree of correlation between each feature vector in the urban geographic scene information dataset. For example, when analyzing the spatial correlation between "building windows" and "walls", the higher the corresponding weight coefficients, the stronger the spatial dependence between the two.
[0087] The weight coefficients are weighted and summed with the corresponding value feature vectors to aggregate the feature vectors with semantic relationships in the urban geographic scene information dataset, generating multiple aggregated feature vector sets. The cosine similarity between feature vectors in the aggregated feature vector set is obtained. Only when the cosine similarity is greater than a preset semantic similarity threshold is it determined that there is a semantic relationship between feature vectors in the aggregated feature vector set.
[0088] Step S22: Generate a scene update attention weight map.
[0089] Specifically, a historical urban geographic scene information dataset is acquired and used as a benchmark for scene update identification. By analyzing the cross-correlation between the scene update identification benchmark and the urban geographic scene information dataset, a scene update attention weight map is generated. The scene update attention weight map is used to label regional information where the urban geographic scene may change.
[0090] In one possible embodiment, features are extracted from the scene update identification benchmark and the urban geographic scene information dataset to generate updated feature maps and historical feature maps.
[0091] Based on a preset projection mechanism, the updated feature map and the historical feature map are mapped to the geographic scene feature space and then L2 normalized.
[0092] Centered on each pixel of the projected updated feature map and the historical feature map, local 3D feature hypercubes of two preset sizes, namely the first scale and the second scale, are extracted to construct the first scale local 3D feature hypercube group and the second scale local 3D feature hypercube group.
[0093] Each local 3D feature hypercube group contains a local 3D feature hypercube of the same size extracted from the updated feature map and the historical feature map, respectively.
[0094] Geographic scene discrimination information is extracted from each local 3D feature hypercube group using a miniature multilayer perceptron. The geographic scene discrimination information is multiplied element-wise to obtain a structural similarity tensor. The structural similarity tensor is then subjected to global average pooling in both the channel dimension and the spatial dimension to obtain similarity response values.
[0095] Adaptively weighted fusion is performed on the similarity response values corresponding to the first-scale local 3D feature hypercube group and the second-scale local 3D feature hypercube group of the same pixel to generate a fused similarity response value.
[0096] Based on the fusion similarity response value corresponding to each pixel, an initial scene update attention weight map is constructed, and the initial scene update attention weight map is processed by an activation function to generate a scene update attention weight map.
[0097] Step S23: Generate scene update recognition results based on scene update attention weight map.
[0098] Specifically, based on the scene update attention weight map, the urban geographic scene information dataset is weighted and fused with the scene update recognition benchmark to generate a binary change mask, a multi-class semantic segmentation map, change confidence and change type data, which are then output as the scene update recognition result.
[0099] Understandably, binary change masks are used to label whether urban geographic scenes have changed, multi-class semantic segmentation maps are used to label different categories of urban geographic scenes, change confidence is used to quantify the reliability of changes in urban geographic scenes, and change type data is used to describe the types of changes in urban geographic scenes.
[0100] In one possible implementation, the attention weight map is updated according to the scene. Construct new and old feature weight matrices respectively and To ensure that areas undergoing change in the urban geographic scene have a higher weight, the following settings are configured: By setting Ensure that the weight of unchanged areas in the urban geographic scene is not too low.
[0101] Will , The feature sets are obtained by element-wise multiplying the corresponding feature data in the urban geographic scene information dataset and the scene update recognition benchmark, respectively. Based on deep learning, the real-time feature set and the historical feature set are adaptively fused to obtain the fused feature set. The data in the fused feature set is mapped to the interval [0, 1] by the Sigmoid activation function. The mapped data value is used as the change confidence. For example, if a feature data in the fused feature set is converted to a value of 0.7 after being mapped by the Sigmoid activation function, then the change confidence of the corresponding feature data is 0.7. The mapped data is then binary divided by a preset threshold. For example, data with a value greater than 0.5 is forced to be 1, and data with a value less than 0.5 is forced to be 0. The fused feature set after binary division is used as a binary mask for output. For example, a value of 1 in the binary mask indicates that a change in the region has occurred, and a value of 0 indicates that a change in the region has occurred.
[0102] Semantic segmentation is performed on the fused feature set using the Softmax activation function. The Softmax activation function outputs a 512×512 matrix, with elements taking values of {0, 1, 2, 3, 4}, corresponding to five change types: "unchanged", "newly built", "demolished", "new vegetation", and "road reconstruction". The determination is only effective when the probability of a certain change type is not lower than a preset threshold. For example, the probability of "newly built" is 0.1, and the probability of "road reconstruction" is 0.7. Assuming the preset threshold is 0.6, the change type of the current area is determined to be "road reconstruction".
[0103] By using preset encoding rules, the semantically segmented regional change types are encoded into vector form, and this vector is output as change type data. For example, suppose common changes in a city's 3D geographic scene are summarized into eight basic change types: new construction, demolition, increase in building height, decrease in building height, start of construction, completion of construction, change of land feature attributes, and other types of changes. An eight-dimensional vector is generated using digital encoding, i.e., [1, 0, 0, 0, 0, 0, 0, 0] can be used to represent that the change type of the current area is new construction, and [0, 0, 5, 0, 0, 1, 0, 0] can be used to represent that the change type of the current area is that the building height has increased by 5 meters after construction is completed.
[0104] For regions where the change confidence level is lower than a preset confidence threshold, periodic secondary verification is performed. For example, data from regions where the change confidence level is lower than the preset confidence threshold is randomly selected periodically to generate a verification task data package. The verification task data package contains the corresponding image data of the region, a binary change mask, change confidence level, and change type data. The verification task data package is then pushed to a web-based expert platform. Experts verify the task and mark the data through the platform. The marking options include three options: "Confirm Change", "Reject Change", and "Pending Observation". If "Confirm Change" is selected, the binary change mask is set to 1; if "Reject Change" is selected, the binary change mask is set to 0; if "Pending Observation" is selected, the detection is repeated after 7 days. If the number of "Reject Change" markings exceeds a preset stability threshold during the periodic secondary verification, the process returns to step S21 and repeats all the contents of step S2. After repeating step S2, a secondary verification is performed again until the number of "Reject Change" markings is lower than the preset stability threshold.
[0105] Step S3: Construct a local dynamic update model for the scene based on the neural radiation field and diffusion model, and perform model pre-training.
[0106] Specifically, the scene local dynamic update model includes a neural radiation field submodule and a multi-view diffusion network submodule. First, the scene structure extraction capability and building texture extraction capability of the scene local dynamic update model are trained by pre-training the neural radiation field submodule and the multi-view diffusion network submodule respectively. Then, the neural radiation field submodule and the multi-view diffusion network submodule are jointly pre-trained to coordinate the feature extraction capabilities between the two modules.
[0107] In this embodiment, step S3 includes:
[0108] Step S31: Pre-train the neural radiation field submodule.
[0109] Specifically, a first training dataset is obtained, consisting of multi-view image data of urban buildings and corresponding laser point cloud data. The multi-view image data of urban buildings includes at least five perspectives: front, back, left, right, and top. The neural radiation field submodule is pre-trained using a scene structure training strategy.
[0110] In one possible embodiment, the scene structure training strategy is expressed as dynamically adjusting the method of view acquisition from the first training dataset and the number of views input to the neural radiation field submodule, and weightedly fusing the neural radiation field loss function with the BIM-based geometric constraint loss function to construct a scene structure loss function, and pre-training based on the scene structure loss function.
[0111] Specifically, the scene structure loss function can be expressed as: ,in, For scene structure loss, It is expressed as mean squared error loss, used to measure the pixel-level difference between the rendered color and the real image, thereby ensuring the visual realism of the 3D scene; This is represented as the geometric constraint loss based on Building Information Modeling (BIM). By incorporating pre-acquired BIM industry standard parameters into the training, the model is forced to generate a 3D geometric structure that conforms to real-world engineering specifications. The geometric constraints include at least roof slope constraints and wall verticality constraints. The roof slope constraints are set based on pre-acquired data from a city-level BIM component library, such as a residential roof slope constraint of [20°, 25°], and the corresponding roof slope constraint loss is obtained using the L1 loss function. The wall verticality constraints are set based on pre-acquired data from a city-level BIM component library, and the corresponding wall verticality constraint loss is obtained using the L2 loss function, such as the allowable error for the verticality constraint loss of office buildings not exceeding 0.5°.
[0112] Understandably, to improve the model's adaptability to complex urban environments, a data sampling strategy based on spatial distance and probability weighting is adopted. In the early stages of pre-training, multiple nearest source views of the local area to be updated are selected as training data. The nearest source view refers to the view data in the 3D geographic scene that has the highest similarity to the current local area to be updated in terms of spatial location, viewpoint, texture features, etc. For example, in the first 1 / 10 of the total training steps, the 10 nearest source views that are spatially closest to the target area are selected as training data. As training progresses, the number of input views is gradually reduced until the number of input views reaches a predetermined value. For example, one nearest source view is reduced simultaneously for every 10% of the training stage, thereby gradually focusing on the core features of the local area to be updated, and finally retaining two representative views.
[0113] Probability weighting means randomly selecting training views from the first training dataset with an 80% probability to enhance the model's ability to generalize data from different perspectives, and fixedly selecting the nearest source view with a 20% probability to maintain the accuracy of capturing local details, thereby avoiding the occurrence of overfitting problems.
[0114] Step S32: Pre-train the multi-view diffusion network submodule.
[0115] Specifically, a second training dataset is obtained, consisting of building texture images under different lighting conditions, weather conditions, and building structure occlusion conditions. A building texture training strategy is then used to pre-train the multi-view diffusion network submodule.
[0116] In one possible embodiment, the architectural texture training strategy is to construct an architectural texture loss function by weighted fusion of the L2 loss function and the consistency loss function, perform pre-training based on the architectural texture loss function, and gradually increase the weight of the consistency loss function from the initial set value to a predetermined consistency weight threshold as the pre-training process progresses, such as gradually increasing from the initial weight of 0.1 to 0.3. The consistency loss function is composed of a weighted fusion of the L1 loss function and the epipolar constraint loss function.
[0117] Understandably, while using the architectural texture training strategy for pre-training, image denoising is performed on the second training dataset simultaneously. By obtaining the mean square error between the denoised image data and the original image data without denoising, the error loss between the denoised image data and the original image data without denoising is quantified.
[0118] Specifically, the architectural texture loss function can be expressed as follows: ,in This is represented as architectural texture loss. This represents the mean square error between the denoised image data and the original data. This is represented as the color difference loss of the same pixel under different viewpoints, obtained according to the L1 loss function. Let be the epipolar constraint loss for the second training dataset, and and Together they constitute the consistency loss function.
[0119] Step S33: Perform joint pre-training on the neural radiation field submodule and the multi-view diffusion network submodule.
[0120] Specifically, a third training dataset consisting of real urban geographic scene update cases is obtained, and the loss results obtained from the scene structure loss function, building texture loss function, and cross-modal consistency loss function are weighted and fused according to a preset weight allocation ratio to generate a joint loss function, such as scene structure loss: building texture loss: cross-modal consistency loss = 3:3:4. The joint loss result is obtained according to the joint loss function, and the neural radiation field submodule and the multi-view diffusion network submodule are jointly pre-trained based on the joint loss result.
[0121] Understandably, the cross-modal consistency loss function can be expressed in the following form: ,in This is represented as cross-modal consistency loss. This represents the total number of data collection points. Represented as the first The geometric normal vector output by the neural radiation field submodule at each acquisition point Local direction vectors of textures generated by the multi-view diffusion network submodule The cosine of the included angle is used to quantify the angular deviation between the two vectors; Represented as the first The geometric surface color values output by the neural radiation field submodule at each acquisition point Color values of textures generated by the multi-view diffusion network submodule The mean square error is used to quantify the color difference in the output data of the two sub-modules; , These are weighting coefficients, with a default value of 0.5.
[0122] Step S4: Input the scene update recognition result into the scene local dynamic update model to update the urban geographic scene and obtain urban geographic scene update data.
[0123] Specifically, the neural radiation field submodule of the scene local dynamic update model performs urban scene structure analysis on the scene update recognition results to obtain urban scene structure data. The multi-view diffusion network submodule of the scene local dynamic update model extracts texture features from the scene update recognition results to obtain building texture feature data. The urban scene structure data and building texture feature data are then fused to generate urban geographic scene update data.
[0124] In one possible embodiment, the urban geographic scene is divided into multiple dynamic update units based on the binary change mask contained in the scene update recognition result. The neural radiation field submodule and the multi-view diffusion network submodule extract the semantic labels of the multi-category semantic segmentation map corresponding to each dynamic update unit according to the pre-established city-level BIM component library, obtain the BIM parameters of the land features corresponding to the area to be updated, and generate urban scene structure data and building texture feature data according to the scene structure loss function and building texture loss function contained in the module. The urban scene structure data and building texture feature data are jointly regulated and feature fused according to the joint loss function contained in the scene local dynamic update model to generate urban geographic scene update data.
[0125] Step S5: Package the city geographic scene update data into a lightweight scene update package and send it to the client for local dynamic scene updates.
[0126] Specifically, the Draco algorithm is used to compress the urban scene structure data in the urban geographic scene update data into a grid, and the vertex coordinates are quantized using an octree.
[0127] Texture encoding is performed on the architectural texture feature data in the updated urban geographic scene data.
[0128] Based on urban geographic scene update data, historical urban geographic scenes are updated to generate updated urban geographic scenes.
[0129] The changes in vertex coordinates and patch topology of the corresponding regions in the updated urban geographic scene and the historical urban geographic scene are calculated to obtain patch difference data. The update operation type is recorded through operation codes, and the update operation type includes at least add, delete and modify operations.
[0130] The processed vertex coordinates, face difference data, and opcodes are packaged into a lightweight scene update package and sent to the client.
[0131] Understandably, the lightweight scene update package contains forward update instructions from the updated city geographic scene to the historical city geographic scene, as well as reverse rollback instructions from the historical city geographic scene to the updated city geographic scene.
[0132] The farthest point sampling algorithm is used to sample multiple vertices from the updated urban geographic scene and the historical urban geographic scene respectively. The transformation matrix between the updated urban geographic scene and the historical urban geographic scene is obtained by the iterative nearest point algorithm. Based on the transformation matrix, vertex matching is performed on the updated urban geographic scene and the historical urban geographic scene. The difference transformation of the region corresponding to the vertex coordinates is recorded synchronously. The corresponding operation code is generated according to the preset rules to generate a positive update instruction.
[0133] The forward update instruction is reverse-mapped according to the preset reverse mapping rules to generate a reverse rollback instruction. The forward update instruction and the reverse rollback instruction are bound by a preset attribute field. The forward update instruction and the reverse rollback instruction can be switched by changing the value of the preset attribute field.
[0134] This invention provides a local dynamic update system for three-dimensional geographic scenes, applicable to three-dimensional scene construction, including:
[0135] The data acquisition module is used to acquire real-time urban geographic scene information data and construct an urban geographic scene information feature pyramid based on the real-time urban geographic scene information data.
[0136] An update recognition module is used to perform scene update recognition on the urban geographic scene information feature pyramid and obtain scene update recognition results.
[0137] The model building module is used to build a scene local dynamic update model based on the neural radiation field and diffusion model, and to perform model pre-training;
[0138] The scene update module is used to input the scene update recognition result into the scene local dynamic update model to update the urban geographic scene and obtain urban geographic scene update data.
[0139] A 3D construction module is used to encapsulate urban geographic scene update data into a lightweight scene update package and send it to the client for local dynamic scene updates.
[0140] The specific usage and function of this embodiment are as follows:
[0141] First, data is collected from the three-dimensional geospatial scene of the city to obtain real-time urban geographic scene information data and construct a corresponding feature pyramid to generate an urban geographic scene information dataset. By constructing a feature pyramid, the accuracy of feature extraction from real-time urban geographic scene information data is improved, making the generated urban geographic scene information dataset more accurate and providing a comprehensive and accurate data foundation for subsequent data analysis.
[0142] Next, based on the urban geographic scene update recognition mechanism, scene update recognition is performed on the urban geographic scene information dataset to obtain scene update recognition results. In the process of scene update recognition, the three-dimensional structural features of the urban geographic scene information dataset are extracted in the form of local three-dimensional feature hypercube, thereby improving the accuracy of scene update recognition.
[0143] Secondly, a scene local dynamic update model is constructed based on the neural radiation field and diffusion model, and the model is pre-trained. The scene update recognition results are input into the scene local dynamic update model to update the urban geographic scene and obtain urban geographic scene update data. The neural radiation field sub-module and the multi-view diffusion network sub-module are constructed through the neural radiation field and diffusion model respectively. By pre-training the two sub-modules respectively, the model's ability to repair under complex occlusion and maintain the consistency of the generated data from multiple perspectives is trained. Finally, the conflict between the two sub-modules is resolved by joint training, thereby ensuring the rationality of the generated urban geographic scene update data.
[0144] Finally, the urban geographic scene update data is packaged into a lightweight scene update package and sent to the client, thereby enabling local dynamic updates of the 3D geographic scene.
[0145] Furthermore, embodiments of the present invention also provide an electronic device, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method in Embodiment 1 described above.
[0146] The following is a detailed introduction to the various components of the electronic device:
[0147] In this context, the processor is the control center of the electronic device. It can be a single processor or a collective term for multiple processing elements. For example, a processor can be one or more central processing units (CPUs), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement Embodiment 1 of this invention, such as one or more digital signal processors (DSPs) or one or more field-programmable gate arrays (FPGAs).
[0148] The processor can perform various functions of an electronic device by running or executing software programs stored in memory and by calling data stored in memory.
[0149] The memory is used to store the software program that executes the solution of the present invention, and the execution is controlled by the processor. For specific implementation methods, please refer to the above method embodiments, which will not be repeated here.
[0150] The memory can be a real-only memory (ROM) or other type of static storage device capable of storing static information and instructions, a random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only (CD-ROM), or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. The memory can be integrated with the processor or exist independently and coupled to the processor through an interface circuit of an electronic device; this embodiment of the invention does not specifically limit this.
[0151] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via limited means (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0152] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0153] It should be understood that, in the embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0154] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for local dynamic updating of a three-dimensional geographic scene, characterized in that, The method includes: Data is collected from the three-dimensional geospatial scene of the city to obtain real-time urban geographic scene information data and construct the corresponding feature pyramid to generate an urban geographic scene information dataset. Based on the urban geographic scene update recognition mechanism, scene update recognition is performed on the urban geographic scene information dataset to obtain scene update recognition results; A scene local dynamic update model is constructed based on the neural radiation field and diffusion model, and the model is pre-trained. The scene update recognition results are input into the scene local dynamic update model to update the urban geographic scene and obtain urban geographic scene update data. The city geographic scene update data is packaged into a lightweight scene update package and sent to the client for local dynamic scene updates.
2. The method for local dynamic updating of a three-dimensional geographic scene according to claim 1, characterized in that, Based on the urban geographic scene update recognition mechanism, scene update recognition is performed on the urban geographic scene information dataset to obtain scene update recognition results, including: By using a multi-head self-attention mechanism to capture long-distance contextual dependencies in urban geographic scene information datasets, semantic relationships between spatial features in urban geographic scene information datasets can be obtained. A historical urban geographic scene information dataset is acquired and used as a benchmark for scene update identification. By analyzing the cross-correlation between the scene update identification benchmark and the urban geographic scene information dataset, a scene update attention weight map is generated. The scene update attention weight map is used to label regional information where the urban geographic scene may change. Based on the scene update attention weight map, the urban geographic scene information dataset and the scene update recognition benchmark are weighted and fused to generate a binary change mask, a multi-class semantic segmentation map, change confidence and change type data, which are then output as the scene update recognition result. Among them, the binary change mask is used to label whether the urban geographic scene has changed, the multi-class semantic segmentation map is used to label different urban geographic scene categories, the change confidence is used to quantify the reliability of changes in the urban geographic scene, and the change type data is used to describe the types of changes in the urban geographic scene.
3. The method for local dynamic updating of a three-dimensional geographic scene according to claim 2, characterized in that, By analyzing the cross-correlation between the scene update recognition benchmark and the urban geographic scene information dataset, a scene update attention weight map is generated, including: Feature extraction is performed on the scene update recognition benchmark and the urban geographic scene information dataset to generate updated feature maps and historical feature maps; Based on a preset projection mechanism, the updated feature map and the historical feature map are mapped to the geographic scene feature space and then L2 normalized. Centered on each pixel of the projected updated feature map and the historical feature map, local 3D feature hypercubes of two preset sizes, namely the first scale and the second scale, are extracted to construct the first scale local 3D feature hypercube group and the second scale local 3D feature hypercube group. Each local 3D feature hypercube group contains a local 3D feature hypercube of the same size extracted from the updated feature map and the historical feature map, respectively. Geographic scene discrimination information is extracted from each local 3D feature hypercube group using a micro multilayer perceptron. The geographic scene discrimination information is multiplied element-wise to obtain a structural similarity tensor. The structural similarity tensor is then subjected to global average pooling in both the channel dimension and the spatial dimension to obtain similarity response values. Adaptively weighted fusion is performed on the similarity response values corresponding to the first-scale local 3D feature hypercube group and the second-scale local 3D feature hypercube group of the same pixel to generate a fused similarity response value. Based on the fusion similarity response value corresponding to each pixel, an initial scene update attention weight map is constructed, and the initial scene update attention weight map is processed by an activation function to generate a scene update attention weight map.
4. The method for local dynamic updating of a three-dimensional geographic scene according to claim 1, characterized in that, A scene local dynamic update model is constructed based on the neural radiation field and diffusion model, and the model is pre-trained, including: The scene local dynamic update model includes a neural radiation field submodule and a multi-view diffusion network submodule. A first training dataset consisting of multi-view image data of urban buildings and corresponding laser point cloud data is obtained. The multi-view image data of urban buildings includes at least five perspectives of urban buildings: front, back, left, right and top. The neural radiation field submodule is pre-trained using a scene structure training strategy. A second training dataset consisting of building texture images under different lighting conditions, weather conditions, and building structure occlusion conditions was obtained, and a building texture training strategy was used to pre-train the multi-view diffusion network submodule. A third training dataset consisting of real urban geographic scene update cases was obtained, and the neural radiation field submodule and the multi-view diffusion network submodule were jointly pre-trained by combining scene structure loss function, building texture loss function and cross-modal consistency loss function.
5. The method for local dynamic updating of a three-dimensional geographic scene according to claim 4, characterized in that, The method further includes: The scene structure training strategy is expressed as follows: dynamically adjusting the view acquisition method from the first training dataset and the number of views input to the neural radiation field submodule, and weighting and fusing the neural radiation field loss function with the BIM-based geometric constraint loss function to construct a scene structure loss function, and pre-training based on the scene structure loss function; The architectural texture training strategy is to construct an architectural texture loss function by weighted fusion of the L2 loss function and the consistency loss function, perform pre-training based on the architectural texture loss function, and gradually increase the weight of the consistency loss function from the initial set value to the predetermined consistency weight threshold as the pre-training process progresses. The consistency loss function is composed of a weighted fusion of the L1 loss function and the epipolar constraint loss function.
6. The method for local dynamic updating of a three-dimensional geographic scene according to claim 1, characterized in that, The scene update recognition results are input into the scene local dynamic update model to update the urban geographic scene, obtaining urban geographic scene update data, including: The neural radiation field submodule of the local dynamic update model of the scene is used to analyze the urban scene structure of the scene update recognition results and obtain urban scene structure data. The multi-view diffusion network submodule of the scene local dynamic update model extracts texture features from the scene update recognition results to obtain building texture feature data. Urban scene structure data and building texture feature data are fused to generate updated urban geographic scene data.
7. The method for local dynamic updating of a three-dimensional geographic scene according to claim 1, characterized in that, The city geographic scene update data is packaged into a lightweight scene update package and sent to the client for local dynamic scene updates, including: The urban scene structure data in the urban geographic scene update data is subjected to grid compression using the Draco algorithm, and its vertex coordinates are quantized using an octree. Texture encoding is performed on the architectural texture feature data in the updated urban geographic scene data. Based on urban geographic scene update data, historical urban geographic scenes are updated to generate updated urban geographic scenes. The changes in vertex coordinates and patch topology of the corresponding regions in the updated urban geographic scene and the historical urban geographic scene are calculated, patch difference data is obtained, and the update operation type is recorded by operation code. The update operation type includes at least add, delete and modify operations. The processed vertex coordinates, face difference data, and opcodes are packaged into a lightweight scene update package and sent to the client.
8. The method for local dynamic updating of a three-dimensional geographic scene according to claim 7, characterized in that, Acquire patch difference data and record the update operation type using the operation code, including: The lightweight scene update package contains forward update instructions from the updated city geo scene to the historical city geo scene, as well as reverse rollback instructions from the historical city geo scene to the updated city geo scene. The farthest point sampling algorithm is used to sample multiple vertices from the updated urban geographic scene and the historical urban geographic scene respectively. The transformation matrix between the updated urban geographic scene and the historical urban geographic scene is obtained by the iterative nearest point algorithm. Based on the transformation matrix, vertex matching is performed on the updated urban geographic scene and the historical urban geographic scene. The difference transformation of the region corresponding to the vertex coordinates is recorded synchronously to obtain the patch difference data. The corresponding operation code is generated according to the preset rules to generate a positive update instruction. The forward update instruction is reverse-mapped according to the preset reverse mapping rules to generate a reverse rollback instruction. The forward update instruction and the reverse rollback instruction are bound by a preset attribute field. The forward update instruction and the reverse rollback instruction can be switched by changing the value of the preset attribute field.
9. The method for local dynamic updating of a three-dimensional geographic scene according to claim 1, characterized in that, Data is collected from the three-dimensional geospatial scene of the city to obtain real-time urban geographic scene information data and construct a corresponding feature pyramid, generating an urban geographic scene information dataset, including: Multi-source data collection is performed on urban areas to obtain real-time urban geographic scene information data, which includes at least RGB optical image data, multispectral image data, and laser point cloud data. A combination of the dark target method and the atmospheric scattering model is used to perform radiometric correction on RGB optical image data, and the corresponding color and texture features are extracted to generate an RGB optical image feature pyramid, which is then output as the RGB optical image feature set. The multispectral image data is atmospherically corrected based on the preset atmospheric correction module, generating a multispectral image feature pyramid, which is then output as a multispectral image feature set. The laser point cloud data is denoised by combining statistical filtering and radius filtering, and the corresponding digital surface model, normal vector map and intensity map are obtained. Multiple laser point cloud feature pyramids are generated and output as laser point cloud feature sets. A feature pyramid of urban geographic scene information is constructed based on RGB optical image feature set, multispectral image feature set, and laser point cloud feature set.
10. A local dynamic update system for a three-dimensional geographic scene, characterized in that, include: The data acquisition module is used to acquire real-time urban geographic scene information data and construct a corresponding feature pyramid to generate an urban geographic scene information dataset. An update recognition module is used to perform scene update recognition on an urban geographic scene information dataset and obtain scene update recognition results. The model building module is used to build a scene local dynamic update model based on the neural radiation field and diffusion model, and to perform model pre-training; The scene update module is used to input the scene update recognition result into the scene local dynamic update model to update the urban geographic scene and obtain urban geographic scene update data. A 3D construction module is used to encapsulate urban geographic scene update data into a lightweight scene update package and send it to the client for local dynamic scene updates.