An artificial intelligence-based digital scene automatic generation method and system

CN122473390BActive Publication Date: 2026-09-29CHANGCHUN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610965904.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-29
Estimated Expiration
2046-07-01

AI Technical Summary

Technical Problem

[0005]本申请提供一种基于人工智能的数字场景自动生成方法和系统,用以解决现有技术中在应对冰雪文旅活动期间目标历史场景多重遮挡叠加场景时,所存在的临时实体误判和光影投影纹理混淆的技术问题,其中,目标历史场景包括历史街区或历史建筑

Benefits of technology

[0005]本申请提供一种基于人工智能的数字场景自动生成方法和系统,用以解决现有技术中在应对冰雪文旅活动期间目标历史场景多重遮挡叠加场景时,所存在的临时实体误判和光影投影纹理混淆的技术问题,其中,目标历史场景包括历史街区或历史建筑。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122473390B_ABST
    Figure CN122473390B_ABST
Patent Text Reader

Abstract

The application provides an artificial intelligence-based digital scene automatic generation method and system, and relates to the technical field of digital cultural heritage, wherein the method comprises the following steps: acquiring an ice-snow activity pre-day scene, an ice-snow activity during-day scene and an ice-snow activity during-night scene of the same historical block; forming a historical building ontology reference scene according to the ice-snow activity pre-day scene; forming an ice-snow activity layer according to the historical building ontology reference scene and the ice-snow activity during-day scene; forming a light and shadow display layer according to the historical building ontology reference scene, the ice-snow activity layer and the ice-snow activity during-night scene; combining a city heritage information model to complete a historical ontology layer; and generating a digital scene by combining each scene layer according to a target display state. The application improves the accuracy of historical building digitalization restoration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of digital cultural heritage protection technology, and in particular to a method and system for automatically generating digital scenes based on artificial intelligence. Background Technology

[0002] With the development of digital preservation and cultural tourism integration of historical districts, multi-view 3D reconstruction and scene rendering based on multi-source data (such as daytime images, laser point clouds, etc.) have become important means for online display, immersive experience and protection and renewal assessment.

[0003] However, capturing the ice and snow cultural tourism scenes in cold-region cities presents unique challenges: temporary entities such as snow and ice sculptures can create real geometric occlusion on buildings; while dynamic light and shadow and projection animations during nighttime activities can visually alter the inherent colors and textures of building surfaces, resulting in extremely varied appearances of the same building at different times.

[0004] Existing 3D reconstruction and image restoration methods struggle to handle such complex scenes, easily misrepresenting temporary entities as permanent building components or misinterpreting dynamic lighting as inherent building textures. This results in generated digital scenes that can only reflect a mixed state at a specific moment, unable to flexibly switch between different modes such as "historical ontology restoration," "ice and snow activity display," or "nighttime light and shadow," severely limiting the application flexibility of digital scenes. Summary of the Invention

[0005] This application provides an artificial intelligence-based method and system for automatically generating digital scenes, which solves the technical problems of temporary entity misjudgment and confusion of light and shadow projection textures in the prior art when dealing with multiple overlapping scenes of target historical scenes during ice and snow cultural tourism activities. The target historical scenes include historical blocks or historical buildings.

[0006] To address the aforementioned technical problems, in a first aspect, this application provides a method for automatically generating digital scenes based on artificial intelligence, comprising: Acquire daytime scenes before ice and snow activities, daytime scenes during ice and snow activities, and nighttime scenes during ice and snow activities for the same target historical scene. The target historical scene includes the historical building itself, temporary ice and snow entities, and dynamic light and shadow projections. Multi-view 3D reconstruction and semantic segmentation of building components were performed on daytime scenes before ice and snow activities to form a benchmark scene of historical building ontology with component-level semantic annotations; The ice and snow material semantic segmentation is performed on the daytime scene during the ice and snow activities, and the ice and snow candidate areas are subjected to three-dimensional change detection on the continuous surface of the historical building body in the benchmark scene to form an ice and snow activity layer that is added relative to the historical building body. Based on the baseline scene of the historical building and the snow and ice activity layer, the nighttime image sequence in the nighttime scene during the snow and ice activity is surface-unfolded, snow and ice area masked, and projection pattern boundary tracked to form a light and shadow display layer attached to the surface of the historical building. Based on the shading area represented by the snow and ice activity layer and the coverage area represented by the light and shadow display layer, a historical ontology layer is supplemented by combining the urban heritage information model. By combining the historical ontology layer, the ice and snow activity layer, and the light and shadow display layer according to the target display state, a digital scene of the target historical scene is generated.

[0007] Furthermore, before supplementing the historical entity layer based on the shading area represented by the snow and ice activity layer and the coverage area represented by the light and shadow display layer, and combining the urban heritage information model, the process also includes: overlaying the dynamic light and shadow projections in at least two projection periods onto the surface of the historical building to form a projection coverage trajectory; extracting the inherent appearance content of the historical building from the surface segments not continuously covered by the projection pattern in the projection coverage trajectory; and using the inherent appearance content as the appearance basis for supplementing the historical entity layer.

[0008] Furthermore, after classifying the snow-covered, ice-sculpted, and snow-sculpted portions into the ice and snow activity layer, the method also includes: along the contact edge between the snow-covered, ice-sculpted, or snow-sculpted portions and the main body of the historical building, following the extension direction of the roof ridge, eaves, door and window frames, step edges, and street paving seams, generating a main body continuity line interrupted by the temporary ice and snow entity; defining the orientation and boundary of the obscured historical building components through the main body continuity line.

[0009] Furthermore, before supplementing the historical ontology layer based on the shading area represented by the ice and snow activity layer and the coverage area represented by the light and shadow display layer, and combining the urban heritage information model, the process also includes: obtaining exhibition data corresponding to ice and snow activities, including ice sculpture locations, snow sculpture locations, projected building surfaces, and projection time periods; mapping the ice sculpture locations to the ice sculpture parts in the ice and snow activity layer, and mapping the snow sculpture locations to the snow sculpture parts in the ice and snow activity layer; mapping the projected building surfaces to the dynamic light and shadow projections in the light and shadow display layer spatially, and mapping the projection time periods to the dynamic light and shadow projections in the light and shadow display layer temporally.

[0010] Secondly, this application provides an artificial intelligence-based automatic digital scene generation system, comprising: The acquisition module is used to acquire daytime scenes before ice and snow activities, daytime scenes during ice and snow activities, and nighttime scenes during ice and snow activities for the same target historical scene. The target historical scene includes historical blocks or historical buildings, and the target historical scene includes the historical building itself, temporary ice and snow entities, and dynamic light and shadow projections. The first module is used to perform multi-view 3D reconstruction and semantic segmentation of architectural components for daytime scenes before ice and snow activities, forming a historical building ontology benchmark scene with component-level semantic annotations. The second forming module is used to perform semantic segmentation of ice and snow materials in daytime scenes during ice and snow activities, and to perform three-dimensional change detection on the continuous surface of the ice and snow candidate area and the historical building body in the benchmark scene to form an ice and snow activity layer that is added relative to the historical building body. The third module is used to perform surface unfolding, ice and snow area masking, and projection pattern boundary tracking on the nighttime image sequence in the nighttime scene during the ice and snow activities, based on the historical building's base scene and the ice and snow activity layer, to form a light and shadow display layer attached to the surface of the historical building. The completion module is used to complete the historical ontology layer based on the shading area represented by the snow and ice activity layer and the coverage area represented by the light and shadow display layer, combined with the urban heritage information model. The generation module is used to combine the historical ontology layer, the ice and snow activity layer, and the light and shadow display layer according to the target display state to generate a digital scene of the target historical scene.

[0011] In this application, by acquiring scene data from multiple time periods before and after the ice and snow activities, multi-view 3D reconstruction and semantic segmentation of building components are first performed on the daytime scene before the ice and snow activities to obtain the baseline scene of the historical building body. Then, semantic segmentation of ice and snow materials and 3D change detection are performed on the daytime scene during the ice and snow activities to obtain the ice and snow activity layer. Surface unfolding, ice and snow area masking, and projection pattern boundary tracking are performed on the nighttime scene during the ice and snow activities to obtain the light and shadow display layer. After forming the light and shadow display layer, a projection coverage trajectory is formed by superimposing at least two projection time periods, and the inherent appearance content is extracted from surface fragments not continuously covered by the projection pattern. After forming the ice and snow activity layer, a body continuation line is generated along the contact edge between the temporary ice and snow entity and the historical building body to limit the direction and boundary of the obscured historical building components. Before supplementing the historical body layer, spatial or temporal attribute correspondences can be made between the ice sculpture locations, snow sculpture locations, projected building surfaces, and projection time periods in conjunction with the exhibition data. This allows for the differentiation of historical buildings, temporary ice and snow entities, and dynamic light and shadow projections within the same generation process. By using an urban heritage information model to supplement the areas of historical buildings that are obscured or covered, and by flexibly combining different scene layers according to different target display states, a digital scene of the target historical scene that meets various display needs can be generated. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is an overall flowchart of an artificial intelligence-based automatic digital scene generation method disclosed in this application.

[0014] Figure 2 This is a schematic diagram of a semantic segmentation network structure for building components disclosed in this application.

[0015] Figure 3 This application discloses a flowchart of the formation process of an active ice and snow layer.

[0016] Figure 4 This is a schematic diagram of projection coverage trajectory analysis disclosed in this application. Detailed Implementation

[0017] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. The described embodiments are only some embodiments of this application, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0018] See Figure 1 As shown in the embodiment of this application, an automatic digital scene generation method based on artificial intelligence is disclosed, which is applied to the digital scene generation of target historical scenes during ice and snow cultural tourism activities. The target historical scenes include historical blocks or historical buildings. The method includes steps S1 to S6.

[0019] S1: Obtain the daytime scene before the ice and snow event, the daytime scene during the ice and snow event, and the nighttime scene during the ice and snow event for the same target historical scene.

[0020] Among them, the target historical scene refers to the historical block or historical building that needs to be generated as a digital scene for ice and snow cultural tourism activities; when the target historical scene is a historical block, the target historical scene includes multiple historical buildings and the street and alley spaces connecting the historical buildings; when the target historical scene is a historical building, the target historical scene includes a single historical building and its adjacent collectable space.

[0021] The target historical scene comprises the historical building itself, temporary ice and snow structures, and dynamic light and shadow projections. The historical building itself refers to the permanent architectural components and enclosing structure of the historical building, including roofs, walls, eaves, doors, windows, steps, and street paving. Temporary ice and snow structures are three-dimensional physical structures made of ice or snow temporarily placed in the target historical scene during ice and snow cultural tourism activities, including snow covering the surface of the historical building and ice and snow sculptures placed independently in passageways. Dynamic light and shadow projections refer to the light shows, projection animations, and dynamic color changes projected onto the surface of the historical building during the nighttime phase of ice and snow cultural tourism activities; the patterns and colors of the dynamic light and shadow projections change periodically over time.

[0022] Since the impact of ice and snow cultural tourism activities on historical buildings includes two types: spatial obstruction and apparent coverage, spatial obstruction comes from the accumulation and placement of temporary ice and snow entities, while apparent coverage comes from the projection of dynamic light and shadow at night, and the two influencing factors occur at different times and under different conditions, it is necessary to collect scene data at three different time periods in order to peel off various influencing factors layer by layer in subsequent steps.

[0023] Specifically, the daytime scenes before ice and snow activities, the daytime scenes during ice and snow activities, and the nighttime scenes during ice and snow activities were collected as follows: The daytime scene before ice and snow activities refers to the scene data obtained by image collection of historical blocks under the conditions that ice and snow activities have not yet been set up, the surface of the historical buildings is not covered by snow, and there are no ice sculptures or snow sculptures placed in the block space. The daytime scene before ice and snow activities contains the original appearance information of the historical buildings when they are not covered by temporary ice and snow entities and are not disturbed by dynamic light and shadow projection. The collection method can be set according to actual needs. For example, in this application, UAV oblique photogrammetry technology is used to collect images of the historical blocks from five directions. The forward overlap is set to no less than 80%, the lateral overlap is set to no less than 60%, and the ground resolution of the image is set to no more than 2 cm.

[0024] Daytime scenes during ice and snow activities refer to scene data acquired by capturing images of historical blocks under daytime conditions when ice and snow activities have been set up, the surfaces of historical buildings may be covered with snow, ice and snow sculptures have been placed in the street space, but the nighttime dynamic light and shadow projection has not yet been turned on. Because there is no interference from dynamic light and shadow projection in the daytime acquisition environment, daytime scenes during ice and snow activities can clearly reflect the spatial relationship between temporary ice and snow entities and historical buildings.

[0025] Nighttime scenes during ice and snow activities refer to scene data acquired by capturing images of historical blocks during the nighttime performance phase of ice and snow activities, under the condition that the surface of the historical building is covered by dynamic light and shadow projection. Since the content of the dynamic light and shadow projection changes periodically over time, the acquisition of nighttime scenes during ice and snow activities needs to continuously acquire image sequences at a rate of no less than 15 frames per second within multiple projection periods to cover the complete change cycle of the dynamic light and shadow projection content. The acquisition duration can be determined according to the cycle of the dynamic light and shadow projection. For example, in this application, the acquisition duration is set to be no less than one complete cycle of the dynamic light and shadow projection.

[0026] In practical engineering applications, assuming a historical district includes three traditional wooden buildings and two bluestone-paved streets, the daytime scenes before the ice and snow activities were collected one week before the event setup, with 1200 images acquired. The daytime scenes during the ice and snow activities were collected on the third day of the event, and it can be observed that the snow cover on the roofs was about 15 to 30 centimeters thick, and four ice sculptures and two snow sculptures were arranged in the streets. The nighttime scenes during the ice and snow activities were collected from 19:00 to 21:00 on the same day, with a dynamic light and shadow projection cycle of 8 minutes, and a total of three complete image sequences were collected.

[0027] This application established a complete data chain by collecting scene data of historical blocks at three different time periods, including the original state of the historical building itself, the state of the temporary snow and ice superimposed, and the state of the dynamic light and shadow projection superimposed, providing original data support for the layer-by-layer stripping of various influencing factors in subsequent steps.

[0028] Before processing the scene data for the three time periods in subsequent steps, it is necessary to register the daytime scene before the ice and snow event, the daytime scene during the ice and snow event, and the nighttime scene during the ice and snow event to the same scene coordinate system. In this application, the 3D reconstruction coordinate system of the daytime scene before the ice and snow event is used as the unified scene coordinate system. The 3D reconstruction results of the daytime scene during the ice and snow event and the nighttime scene during the ice and snow event are rigidly registered with the 3D reconstruction results of the daytime scene before the ice and snow event through the iterative nearest point registration algorithm. After registration, the root mean square distance residual between the scene data of each time period is not greater than the root mean square error of the 3D reconstruction accuracy, thereby ensuring that the spatial comparison and change detection across time periods in subsequent steps are carried out under a unified spatial reference.

[0029] S2: Perform multi-view 3D reconstruction and semantic segmentation of building components for daytime scenes before ice and snow activities to form a historical building ontology benchmark scene with component-level semantic annotations.

[0030] In step S1, the daytime scene before the snow and ice event was acquired. Since this scene reflects the original state of the historical building before any temporary snow or ice covering or dynamic light and shadow projection interference, it can serve as a reference for subsequent scene layer separation. However, the daytime scene before the snow and ice event is only raw image data and lacks 3D spatial representation and component-level semantic annotation. Therefore, this step does not directly obtain the reference scene from the image. Instead, it first restores the spatial surface of the historical building through multi-view 3D reconstruction, and then assigns component categories such as roof, walls, eaves, doors and windows, steps, and street paving to the spatial surface through semantic segmentation of building components. Finally, a historical building reference scene with component-level semantic annotation is formed. The historical building reference scene provides a quantifiable spatial reference for subsequent 3D change detection, projection surface positioning, and component completion.

[0031] Specifically, step S2 includes at least two consecutive processing levels: the first level is multi-view 3D reconstruction, which restores the spatial surface of the historical building from multi-view images of the daytime scene before the snow and ice activities; the second level is semantic segmentation of building components, which identifies component categories such as roofs, walls, eaves, doors and windows, steps, and street paving on the spatial surface, and forms a continuous surface of the entity with component affiliation. Through these two processing levels, subsequent steps can use the continuous surface of the entity as a common object for cross-time period comparison, projection positioning, and component completion.

[0032] The specific implementation process of this step includes steps S21 to S23.

[0033] S21: Perform multi-view 3D reconstruction of the daytime scene before ice and snow activities to form the spatial surface of the historical building.

[0034] Among them, multi-view 3D reconstruction refers to the technical method of restoring the three-dimensional geometric shape of the subject by using multiple images obtained from different shooting angles and processing steps such as feature matching, camera pose estimation and dense point cloud generation. Spatial surface refers to a continuous three-dimensional mesh surface composed of triangular facets. Each triangular facet on the spatial surface has three attributes: three-dimensional coordinates, normal vector and texture color value.

[0035] The specific 3D reconstruction algorithm can be selected according to actual needs. For example, in this application, the method based on motion recovery structure is first used to extract feature points and perform multi-view matching on 1200 images to estimate the camera pose parameters of each image and generate a sparse point cloud. Then, the sparse point cloud is densified by the multi-view stereo matching algorithm to generate a dense point cloud. Finally, the Poisson surface reconstruction algorithm is used to transform the dense point cloud into a continuous triangular mesh spatial surface. The Poisson surface reconstruction algorithm takes the 3D coordinates and normal vector of each point in the dense point cloud as input, obtains the implicit surface function by solving the Poisson equation, and then extracts the isosurface from the implicit surface function to generate a triangular mesh spatial surface. The reconstruction accuracy is measured by the root mean square error. In this application, the root mean square error of the reconstruction accuracy is set to be no higher than 0.03 meters.

[0036] S22: Perform semantic segmentation of building components on the spatial surface to obtain the corresponding continuous surfaces of the roof, walls, eaves, doors and windows, steps and street paving.

[0037] Among them, semantic segmentation of building components refers to the process of classifying and labeling each triangular facet on a three-dimensional spatial surface according to the type of building component it belongs to. A continuous surface refers to the three-dimensional surface region corresponding to a set of triangular facets that belong to the same component category and are spatially connected. The criteria for determining spatial connectivity are: two triangular facets are considered adjacent when they share at least one edge. The largest connected region formed by starting from a certain triangular facet and gradually expanding to all adjacent triangular facets of the same category is a continuous surface.

[0038] See Figure 2As shown, this application employs a point cloud-based deep learning semantic segmentation network to achieve component-level classification. This point cloud-based deep learning semantic segmentation network is a specific application of deep learning technology in artificial intelligence. It achieves intelligent recognition of building components through layer-by-layer nonlinear feature mapping of a multilayer perceptron and hierarchical spatial abstraction using farthest point sampling. The structure of the semantic segmentation network is as follows: The triangular mesh spatial surface obtained from the 3D reconstruction in step S21 is converted into a point cloud with a point density of no less than 2500 sampling points per square meter according to a uniform sampling strategy. A nine-dimensional feature vector is formed by the three-dimensional coordinates (x, y, z), normal vectors (nx, ny, nz), and texture color values ​​(R, G, B) of each sampling point in the point cloud, serving as the network input. The encoder 101 of the semantic segmentation network consists of four feature extraction modules. Each feature extraction module maps the input features to a high-dimensional feature space through a multilayer perceptron with shared weights, and then... The far-point sampling algorithm downsamples the point cloud layer by layer to expand the receptive field. The output feature dimensions of the four feature extraction modules are 64, 128, 256 and 512 respectively. The decoder 102 of the semantic segmentation network restores the features of each layer to the original resolution through feature interpolation and skip connections 103. The skip connections 103 of the decoder 102 concatenate the features of the corresponding layer in the encoder 101 with the features of the current layer in the decoder 102. The final output layer 104 is a six-channel point-by-point classifier. The six channels correspond to six component categories: roof, wall, eaves, doors and windows, steps and street paving.

[0039] The correspondence between the six component categories of the semantic segmentation network and the components of historical buildings is shown in Table 1.

[0040] Table 1. Semantic Segmentation Category Mapping Table for Building Components

[0041] As shown in Table 1, the six component categories of the semantic segmentation network cover the common architectural component types in historical districts. The range of architectural components and typical materials corresponding to each category serve as the annotation basis for training the semantic segmentation network.

[0042] The training process of the semantic segmentation network is as follows: The training dataset contains manually labeled point cloud samples of at least 50 historical buildings. The training set and validation set are divided in an 8:2 ratio. The validation set contains labeled point clouds of at least 10 historical buildings. Each sampling point is labeled with a corresponding component category number. The number of labeled samples for each component category is kept balanced through a random oversampling strategy. During training, the cross-entropy loss function is used. The cross-entropy loss function is calculated by calculating the negative logarithm of the probability value predicted by the semantic segmentation network for the i-th sampling point in the point cloud to belong to the true category, and then averaging the negative logarithmic values ​​of all sampling points as the total loss. The training optimizer uses the Adam optimizer, with an initial learning rate set to 0.001. The learning rate decay strategy uses cosine annealing decay. Every 50 training rounds, the learning rate is decayed from the current value to the initial value according to the cosine function. The learning rate is reduced to one percent, then restored to the initial learning rate to restart the decay cycle; the batch size for each training batch is set to 16 point cloud blocks, and each point cloud block contains 4096 sampling points; the output feature dimensions of the multilayer perceptron in the four-layer feature extraction module of the encoder are 64, 128, 256 and 512 dimensions respectively, and each multilayer perceptron consists of two fully connected layers, with batch normalization and ReLU activation function used between the two fully connected layers; the sampling ratio of the farthest point sampling algorithm in the four layers is 4:1, 4:1, 4:1 and 2:1 respectively, that is, after downsampling the original input of 4096 sampling points through four layers, 32 sampling points are retained; when the labeled sample size of each component category is unbalanced, a random oversampling strategy is used to make the sample size of the minority category consistent with the sample size of the majority category, and the oversampling factor is the ratio of the sample size of the majority category to the sample size of the minority category rounded up; The training rounds are set to be no less than 200 rounds. The model evaluation metrics are point-by-point classification accuracy and average intersection-union ratio (IU). Point-by-point classification accuracy is the ratio of the number of correctly classified sampling points in the validation set to the total number of sampling points in the validation set. The average IU is the arithmetic mean of the IU of the six component categories. The IU of each category is the ratio of the number of correctly predicted sampling points of that category to the union of the number of predicted sampling points and the number of actual sampling points of that category. Training stops when the point-by-point classification accuracy on the validation set is no less than 92% and the average IU is no less than 0.75. In the implementation scenario of this application, the IU of each category is as follows: roof 0.88, wall 0.85, eaves 0.72, doors and windows 0.74, steps 0.71, and street paving 0.82.

[0043] The reason why the semantic segmentation network can perform building component classification is that the encoder 101 expands the local receptive field of each sampling point through layer-by-layer downsampling, enabling the semantic segmentation network to capture the local geometric features and global structural features of building components, such as the large-area gentle slope of the roof, the linear overhang of the eaves, and the regular rectangular recess of the doors and windows; the decoder 102 retains the detailed information of each layer through skip connections 103, enabling the semantic segmentation network to accurately locate the boundaries between components when restoring the original resolution.

[0044] S23: Write the continuous surface of the body into the scene space according to the component affiliation of the historical building body to form the benchmark scene of the historical building body.

[0045] Among them, component attribution refers to the combination of the specific building unit number and component type number to which each continuous surface of the building belongs. For example, the roof component number of building B1 is B1-RF-01, and the component number of the first wall of building B1 is B1-WL-01. Through the component attribution identifier, the spatial location of each continuous surface of the building in the entire historical district and the building to which it belongs can be accurately located in subsequent steps.

[0046] In practical engineering applications, the daytime scene before ice and snow activities in a historical district is processed. Step S21 generates a spatial surface containing approximately 5.8 million triangular facets from 1200 images through multi-view 3D reconstruction. In step S22, the semantic segmentation network outputs the component category point by point for the sampling points corresponding to the 5.8 million triangular facets, resulting in 12 continuous faces for the roof, 24 continuous faces for the walls, 18 continuous faces for the eaves, 36 continuous faces for the doors and windows, 6 continuous faces for the steps, and 4 continuous faces for the street paving, totaling 100 continuous faces. These are then written into the scene space in step S23, forming the baseline scene of the historical building.

[0047] This application transforms the historical building ontology in the daytime scene before ice and snow activities into a three-dimensional spatial representation with component-level semantic annotations through multi-view three-dimensional reconstruction and semantic segmentation of building components. This enables the historical building ontology benchmark scene to include not only three-dimensional geometric location, but also component type, component affiliation, and ontology continuous surface information, thereby providing a unified spatial and semantic benchmark for subsequent three-dimensional change detection after semantic segmentation of ice and snow materials, surface unfolding of nighttime image sequences, and historical building ontology completion.

[0048] S3: Perform semantic segmentation of ice and snow materials in daytime scenes during ice and snow activities, and perform 3D change detection on the continuous surface of the ice and snow candidate area and the historical building body in the baseline scene to form an additional ice and snow activity layer relative to the historical building body.

[0049] In step S2, a baseline scene of the historical building has been formed. Since the baseline scene reflects the original spatial form of the historical building before the ice and snow activities, and the daytime scene during the ice and snow activities includes temporary ice and snow entities such as snow accumulation, ice sculptures, and snow sculptures added on top of the original spatial form, this step first performs semantic segmentation of ice and snow materials in the daytime scene during the ice and snow activities to obtain candidate ice and snow regions that may be composed of ice or snow. Then, the candidate ice and snow regions are subjected to three-dimensional change detection with the continuous surface of the building in the baseline scene to confirm the activity-period space occupied by the candidate ice and snow regions located outside the continuous surface of the building and relative to the historical building. Finally, based on the contact relationship and reflection characteristics between the activity-period space occupied surface and the historical building, it is divided into snow accumulation, ice sculpture, and snow sculpture parts and classified into the ice and snow activity layer. The ice and snow activity layer refers to the independent scene data layer formed after separating all the temporary ice and snow entities added during the ice and snow activities from the historical building.

[0050] Specifically, step S3 does not directly treat the entire area of ​​difference between the two time periods as the ice and snow activity layer. Instead, it first excludes non-ice and snow material areas through semantic segmentation of ice and snow materials. Then, it determines whether the candidate ice and snow areas are located outside the continuous surface of the building body through three-dimensional change detection methods such as the signed distance from point to surface. Finally, it distinguishes between snow accumulation, ice sculpture, and snow sculpture by rules such as the contact area ratio and the proportion of specular reflection components. The resulting ice and snow activity layer has both material category information and spatial occupancy information relative to the historical building body.

[0051] See Figure 3 As shown, the specific implementation process of this step includes steps S31 to S34.

[0052] S31: Perform semantic segmentation of ice and snow materials in daytime scenes during ice and snow activities to obtain candidate ice and snow regions.

[0053] Among them, ice and snow material semantic segmentation refers to the process of distinguishing the surface area composed of ice or snow from the brick, stone and wood surface area of ​​the historical building body in the 3D reconstruction results of daytime scenes during ice and snow activities based on material spectral features and surface texture features. Ice and snow candidate region refers to the 3D surface area covered by the set of triangular facets marked as ice or snow material after ice and snow material semantic segmentation.

[0054] First, the daytime scene during the ice and snow activities needs to be reconstructed in three dimensions using the same multi-view three-dimensional reconstruction method as in step S21 to obtain the triangular mesh space surface of the daytime scene during the ice and snow activities. Then, the semantic segmentation of ice and snow materials is performed on the triangular mesh space surface. In this application, the characteristics of high reflectivity and low texture complexity of ice and snow surfaces in the visible light band are utilized, and the surface reflectivity and the mean texture gradient are used as two indicators for joint screening.

[0055] For each triangular facet in the triangular mesh space of a daytime scene during snow and ice activities, the ratio of the average gray value of that facet across all visible viewpoints to the incident light intensity is calculated as the surface reflectivity. Surface reflectivity is a dimensionless quantity, ranging from 0 to 1. The incident light intensity is estimated as follows: a grayscale standard calibration plate is placed in the scene during each image acquisition. The known reflectivity of the grayscale standard calibration plate is 0.18. The estimated incident light intensity under the current acquisition conditions is obtained by reading the grayscale value of the grayscale standard calibration plate in the image and dividing it by its known reflectivity of 0.18. The unit is grayscale value; when it is not possible to continuously place the grayscale standard calibration plate during the acquisition process, a reference surface with known reflectivity in the scene (such as a clean plaster wall, whose typical reflectivity is about 0.4 to 0.5) can be selected for back calculation; calculate the arithmetic mean of the texture gradient of the triangular facet in all visible view images as the texture gradient mean. The texture gradient is calculated by the Sobel operator, and the unit of the texture gradient mean is grayscale value per pixel; when the surface reflectivity is greater than 0.7 and the texture gradient mean is less than the preset texture gradient threshold, the triangular facet is marked as part of the ice and snow candidate area.

[0056] The preset texture gradient threshold is determined as follows: In the training sample set containing both ice and snow surfaces and building material surfaces, the mean distribution of texture gradients for ice and snow surfaces and building material surfaces are statistically analyzed respectively. The 95th percentile value of the mean texture gradient distribution for ice and snow surfaces is taken as the preset texture gradient threshold. In the implementation scenario of this application, the preset texture gradient threshold is set to 15 gray values ​​per pixel. The above threshold values ​​are preferred values ​​in the implementation scenario of this application. In scenarios with different lighting conditions, different camera exposure parameters, or different material combinations, each threshold can be dynamically adjusted according to the calibration samples: The surface reflectance threshold and the preset texture gradient threshold can be obtained by collecting no less than 200 ice and snow surface samples and no less than 200 building material surface samples in the target scene, and statistically analyzing the distribution of the two types of samples on the two indices of surface reflectance and mean texture gradient. The index value corresponding to the intersection of the two distributions is used as the adaptive threshold. When the two distributions overlap and the intersection point is not unique, the index value that maximizes the segmentation accuracy of ice and snow materials on the validation set is selected as the threshold.

[0057] S32: Perform three-dimensional change detection on the candidate ice and snow region and the continuous surface of the body to form the active period space occupancy surface located outside the continuous surface of the body.

[0058] Among them, three-dimensional change detection refers to spatially registering the three-dimensional reconstruction results of the daytime scene during the ice and snow activities with the historical building's reference scene, calculating the point-by-point distance between the three-dimensional surfaces of the two time periods, and determining the area with a distance greater than the preset change detection distance threshold as the area where spatial change has occurred. The activity period spatial occupancy surface refers to the three-dimensional surface area covered by the set of triangular facets that have been determined to have undergone spatial change after three-dimensional change detection and are located outside the continuous surface of the building.

[0059] Before performing 3D change detection, it is necessary to first unify the 3D reconstruction results of the daytime scene during the ice and snow activities with the historical building reference scene into the same coordinate system. In this application, the iterative nearest point registration algorithm is used for spatial registration. The iterative nearest point registration algorithm is a registration method that finds the correspondence between two sets of point clouds through iteration and solves the optimal rigid body transformation parameters. The registration accuracy requirement is that the root mean square distance between the two sets of point clouds after registration is not greater than the root mean square error of the 3D reconstruction accuracy in step S21.

[0060] After spatial registration, for each sampling point q in the ice and snow candidate region, calculate the point-to-surface signed distance d between sampling point q and the nearest continuous surface of the historical building reference scene. The formula for calculating the point-to-surface signed distance d is d=(qp)·n; where q is the 3D coordinate vector of the sampling point to be detected in the ice and snow candidate region, in meters; p is the 3D coordinate vector of the projection point on the nearest continuous surface of the historical building reference scene, in meters; n is the unit normal vector of the continuous surface of the projection point p at the projection point p, which is a dimensionless vector; qp is the displacement vector between sampling point q and projection point p, in meters; (qp)·n represents the dot product operation of the displacement vector and the unit normal vector. Since the unit of the displacement vector is meters and the unit normal vector is dimensionless, the result of the dot product operation is in meters, therefore the dimension of d is meters.

[0061] When the signed distance d from the point to the surface is greater than the preset change detection distance threshold, it is determined that a spatial change has occurred at the sampling point q, and a positive value of d indicates that the sampling point q is located outside the continuous surface of the body. The preset change detection distance threshold is determined as follows: take three times the root mean square error of the three-dimensional reconstruction accuracy in step S21 as the preset change detection distance threshold to eliminate the interference of reconstruction noise. For example, when the root mean square error of the three-dimensional reconstruction accuracy is 0.03 meters, the preset change detection distance threshold is set to 0.09 meters.

[0062] In practical engineering applications, assuming the three-dimensional coordinates of a sampling point q in the ice and snow candidate area are [10.50, 5.20, 3.80] meters, and the three-dimensional coordinates of the projection point p on the nearest continuous surface of the body are [10.50, 5.20, 3.55] meters, and the unit normal vector n at the projection point p is [0, 0, 1], then the displacement vector qp = [0, 0, 0.25] meters, d = [0, 0, 0.25]·[0, 0, 1] = 0.25 meters. Since 0.25 meters is greater than the preset change detection distance threshold of 0.09 meters and d is a positive value, it is determined that a spatial change has occurred at the sampling point q and it is located outside the continuous surface of the body.

[0063] S33: Divide the space occupied during the activity period into the snow-covered part attached to the surface of the historical building, the ice sculpture part occupying the passage space of the street, and the snow sculpture part occupying the passage space of the street.

[0064] Each connected region in the space occupied during the activity period is classified step by step according to two criteria: the first criterion is the contact area ratio. The second criterion is the proportion of specular reflection component. Among them, the snow-covered part refers to the temporary ice and snow entities attached to the surface of the historical building, the ice sculpture part refers to the temporary ice and snow entities made of ice material and placed independently in the street space, and the snow sculpture part refers to the temporary ice and snow entities made of snow material and placed independently in the street space.

[0065] Contact area ratio Used to distinguish between snow-covered and non-snow-covered areas, with a contact area ratio of... The calculation formula (1) is as follows:

[0066] In the above formula, The area covered by the set of boundary patches in a connected region of the activity space occupancy surface where the Euclidean distance between the boundary patch and the continuous surface of the historical building in the reference scene is less than a preset contact distance threshold is the area in square meters. The preset contact distance threshold is set to twice the root mean square error of the 3D reconstruction accuracy in step S21, which is 0.06 meters. This represents the total surface area of ​​the connected region, in square meters. This is the ratio of contact areas. Since both the numerator and denominator are in square meters, dividing them will result in... It is a dimensionless quantity, and its value ranges from 0 to 1.

[0067] When the contact area ratio When the contact area ratio exceeds the preset threshold of 0.3, the connected area is determined to be a snow-covered area. The preset contact area ratio threshold of 0.3 is based on the following: snow is usually attached to the surface of the historical building in a thin layer. Therefore, the ratio of the contact area between the boundary surface of the snow-covered area and the continuous surface of the building to the total area is usually greater than 0.3. On the other hand, ice sculptures and snow sculptures are usually placed in the street space as independent three-dimensional solid forms, with only a small amount of contact between the bottom and the ground or steps. Therefore, the contact area ratio between ice sculptures and snow sculptures is usually less than 0.3.

[0068] When the contact area ratio When the value is not greater than 0.3, further calculations are made based on the proportion of the specular reflection component. Distinguishing between ice sculptures and snow sculptures, the proportion of specular reflection components. The calculation formula (2) is as follows:

[0069] In the above formula (2), This represents the average intensity of specular reflection light from the connected region at multiple viewing angles, expressed in grayscale values. This represents the average total reflected light intensity of the connected region across multiple viewing angles, expressed in grayscale values. It is a dimensionless quantity, and its value ranges from 0 to 1.

[0070] Specular reflection component The separation method is as follows: This application adopts a multi-view reflection component separation method based on the Phong reflection model. The Phong reflection model determines the total reflected light intensity of the surface. Decomposed into diffuse reflection components With specular reflection component The sum of the diffuse reflection components It remains constant from all viewing angles, while the specular reflection component... The specular reflection exhibits characteristics that change with the viewing angle. Therefore, for a connected region of the space occupied during the activity period, the intensity of reflected light in that region should be collected from at least five different viewing angles. The five viewing angles should be evenly distributed around the surface normal vector of the connected region to ensure that at least one viewing angle does not contain specular reflection specular highlights. The minimum value of the reflected light intensity observed from all viewing angles is taken as the diffuse reflection component. The estimated value is obtained by subtracting the diffuse reflection component from the observed reflected light intensity at each viewpoint. The estimated values ​​are then averaged to obtain the specular reflection component. The estimated value.

[0071] When the proportion of specular reflection component If the specular reflection threshold is greater than 0.4, the connected region is determined to be an ice sculpture; otherwise, it is determined to be a snow sculpture. The basis for setting the specular reflection threshold to 0.4 is that the surface of ice is translucent and has strong specular reflection characteristics in daytime environments, and the proportion of specular reflection is usually greater than 0.4; the surface of snow is opaque white and has strong diffuse reflection characteristics, and the proportion of specular reflection is usually less than 0.4.

[0072] The classification rules for temporary ice and snow entities are summarized in Table 2.

[0073] Table 2. Rules for Classifying Temporary Entities of Ice and Snow

[0074] As shown in Table 2, the classification process is a step-by-step process: first, based on the contact area ratio... To distinguish between snow-covered and non-snow-covered areas, the non-snow-covered areas were then evaluated based on the proportion of specular reflection components. To further distinguish between ice sculptures and snow sculptures, the contact area ratio threshold of 0.3 and the specular reflection threshold of 0.4 are preferred values ​​for the implementation scenario of this application. In practical applications, they can be adaptively determined based on the verification samples: the contact area ratio threshold can be determined by collecting no less than 30 known categories of temporary ice and snow entity samples in the target scene, statistically analyzing the contact area ratio distribution of snow-covered samples and non-snow-covered samples respectively, and taking the contact area ratio corresponding to the intersection of the two distributions as the adaptive threshold; the specular reflection threshold can be determined by collecting no less than 20 ice sculpture samples and 20 snow sculpture samples in the target scene, statistically analyzing the specular reflection component proportion distribution of the two types of samples respectively, and taking the specular reflection component proportion value that maximizes the classification accuracy of the verification samples as the adaptive threshold.

[0075] S34: The snow-covered area, ice sculpture area and snow sculpture area are classified into the snow and ice activity layer.

[0076] After classifying the snow-covered, ice-sculpted, and snow-sculpted sections into the snow and ice activity layer, it is also necessary to generate the body continuity line. The body continuity line refers to the geometric outline of the historical building components that still exists in the area covered by the temporary snow and ice entity. The physical meaning of the body continuity line is that the outline of the historical building components was continuous before being covered by the temporary snow and ice entity. The coverage of the temporary snow and ice entity only covers a local segment of the component outline, and the component outline still continues in the covered area.

[0077] The method for generating the body continuation line is as follows: along the contact edge between the snow-covered part, ice sculpture part, or snow sculpture part and the historical building body, extract the endpoints of the visible component outline segments on both sides of the contact edge, and take the direction vector within a certain length range of the ends of the visible component outline segments on both sides. The direction vector is calculated by linearly fitting the last five sampling points at the end of the visible component outline segment to obtain the straight line direction. Then, take the arithmetic mean of the direction vectors on both sides as the extension direction of the body continuation line. Connect the two endpoints with straight lines along the arithmetic mean direction vector to generate the body continuation line.

[0078] The rules for generating the body continuation line are shown in Table 3.

[0079] Table 3. Rules for Generating Body Continuation Lines

[0080] As shown in Table 3, since the ridge, eaves, door and window frames, step edges and paving seams of historical buildings are usually straight in a local area, it is physically reasonable to use a straight connection method to generate the body continuation line. For the outline of the component with a curved direction, an intermediate control point can be added between the two endpoints and a segmented straight connection method can be used to approximate it.

[0081] This application uses semantic segmentation of ice and snow materials, 3D change detection, and hierarchical classification to separate and classify temporary ice and snow entities added during ice and snow activities from the historical building body and categorize them into the ice and snow activity layer. Since the 3D change detection uses the aforementioned historical building body reference scene with component-level semantic annotations as a reference, the ice and snow activity layer can simultaneously characterize the material type, spatial occupancy, and occlusion relationship of the temporary ice and snow entities relative to the historical building body, while generating body continuation lines to provide constraints for the subsequent geometric completion of the historical building body.

[0082] S4: Based on the baseline scene of the historical building and the snow and ice activity layer, the nighttime image sequence in the nighttime scene during the snow and ice activity is surface unfolded, the snow and ice area is masked and the projection pattern boundary is tracked to form a light and shadow display layer attached to the surface of the historical building.

[0083] In step S3, an ice and snow activity layer has been formed. Since the ice and snow activity layer has identified the positions of the snow accumulation, ice sculptures, and snow sculptures in three-dimensional space, when processing the night scene during the ice and snow activity, it is necessary to first register the night image sequence with the historical building's reference scene and project it onto the two-dimensional parametric plane of the historical building's walls, eaves, and paved surfaces. Then, based on the ice and snow activity layer, the ice and snow areas in the unfolded night image sequence are masked to exclude the spatial areas occupied by temporary ice and snow entities. In addition to temporary ice and snow entities, the night scene during the ice and snow activity also includes dynamic light and shadow projections projected onto the surface of the historical building. If the visual content of the historical building's surface in the night scene during the ice and snow activity is directly used as the inherent texture of the historical building, visual information unrelated to the historical building will be introduced into the digital scene. Therefore, it is also necessary to perform projection pattern boundary detection and optical flow tracking on the retained night image content of the historical building's surface to separate the dynamic light and shadow projections from the surface of the historical building and establish a spatial correspondence between the dynamic light and shadow projection content and the surface of the historical building.

[0084] Specifically, step S4 includes at least five processing stages: surface unrolling, snow and ice area masking, projection pattern boundary detection, optical flow tracing, and correspondence writing. Surface unrolling is used to convert the nighttime image content on the three-dimensional surface into a two-dimensional parametric planar image; snow and ice area masking is used to eliminate the interference of snow, ice sculptures, and snow sculptures on projection recognition; projection pattern boundary detection is used to extract the coverage contour of the dynamic light and shadow projection in the effective area; optical flow tracing is used to obtain the trajectory of the projection pattern boundary moving with the projection time period; and correspondence writing is used to write the dynamic light and shadow projection and the surface of the historical building to which it is attached into the light and shadow display layer.

[0085] Among them, the light and shadow display layer refers to the independent scene data layer formed by separating the content of the dynamic light and shadow projection and its correspondence with the surface of the historical building from the night scene during the ice and snow activities through surface unfolding, ice and snow area masking, projection pattern boundary detection, optical flow tracking and three-dimensional back projection.

[0086] The specific implementation process of this step includes steps S41 to S47.

[0087] S41: Register and project the nighttime image sequence from the nighttime scene during the ice and snow activities onto a two-dimensional parametric plane of the walls, eaves and paved surfaces of the historical building to form an unfolded nighttime image sequence.

[0088] Surface unwrapping refers to the process of spatially registering a sequence of nighttime images from a nighttime scene during a snow and ice event with a reference scene of a historical building, and then projecting the registered nighttime image content onto a two-dimensional parametric plane of the corresponding continuous surface of the historical building reference scene. The purpose of surface unwrapping is to convert nighttime texture information on a three-dimensional surface into a two-dimensional planar image, enabling subsequent operations such as snow and ice area masking, projection pattern boundary detection, and optical flow tracing to be performed on the unwrapped two-dimensional parametric planar image.

[0089] This application employs a piecewise minimum distortion unfolding method to map a three-dimensional spatial surface to a two-dimensional parameterized plane. The specific implementation process is as follows: First, based on the semantic segmentation results of the building components in step S22, each continuous surface of the body is treated as an independent unfolding segment. For continuous surfaces of the body that are approximately planar (such as walls and street paving), the three-dimensional coordinates are directly projected onto the orthogonal plane of the normal vector of the plane using orthogonal projection. Orthogonal projection does not produce area distortion or angle distortion. For continuous surfaces of the body that have curvature (such as eaves arc surfaces and bracket arch surfaces), a parameterization method based on least squares conformal mapping is adopted. This method solves for the optimal parameterized coordinates by minimizing the quadratic energy function of the angle change of the triangular facets before and after unfolding, so as to keep the local angle relationship unchanged. The area distortion rate after conformal mapping does not exceed five percent. For the seam areas between each unfolding segment, an overlap band with a width of not less than 10 pixels is set to ensure the continuity across segments.

[0090] S42: Mask the snow and ice areas in the unfolded night image sequence. Specifically, mask the snow, ice sculptures and snow sculptures corresponding to the snow and ice activity layer in the unfolded night image sequence, while preserving the night image content of the surface of the historical building.

[0091] Since the boundary of the temporary ice and snow entity in three-dimensional space has been determined in step S3, the three-dimensional spatial boundary of the temporary ice and snow entity can be projected onto the two-dimensional parametric plane after being unfolded in step S41 to generate a binary mask. The binary mask is a matrix of the same size as the two-dimensional parametric plane image, composed of zero and one values. The area with a value of zero in the binary mask corresponds to the invalid area covered by the temporary ice and snow entity, and the area with a value of one corresponds to the valid area of ​​the surface of the historical building. Only the nighttime image content of the valid area is retained for subsequent processing.

[0092] S43: Perform projection pattern boundary detection and optical flow tracking on the nighttime image content of the historical building's surface to obtain the projection pattern boundary that moves on the surface of the historical building as the projection time period changes.

[0093] The projection pattern boundary refers to the outline of the projection pattern of the dynamic light and shadow projection on the surface of the historical building. Since the content of the dynamic light and shadow projection changes over time, the position of the projection pattern boundary on the surface of the historical building also moves over time. Therefore, this step first detects the feature points of the projection pattern boundary on the unfolded two-dimensional parametric planar image, and then performs optical flow tracking on the feature points of the projection pattern boundary to obtain the motion trajectory of the projection pattern boundary as the projection period changes.

[0094] First, feature points of the projection pattern boundary need to be detected on the unfolded two-dimensional parametric plane image of each frame. The specific method for feature point detection is as follows: the Canny edge detection operator is used to perform edge detection on the unfolded two-dimensional parametric plane image. The high threshold of the Canny edge detection operator is set to 100 gray values ​​and the low threshold is set to 50 gray values. Among the detected edge pixels, edge pixels located in the effective area and whose edge gradient magnitude is greater than the preset edge gradient threshold are selected as feature points of the projection pattern boundary. The preset edge gradient threshold is set to twice the preset texture gradient threshold in step S31, i.e., 30 gray values ​​per pixel, to ensure that the detected feature points correspond to the boundary of the dynamic light and shadow projection pattern rather than the texture boundary of the surface of the historical building itself.

[0095] Furthermore, to further eliminate the possibility of misidentifying the inherent texture boundaries of the historical building's surface as projection pattern boundaries, this application introduces a time-varying criterion as a supplementary screening condition: for each edge feature point detected in consecutive frames, the average displacement and average color change of the feature point in five adjacent frames are calculated. When the average displacement is less than 1 pixel and the average color change is less than 5 grayscale values, the feature point is determined to be the inherent texture boundary of the historical building's surface rather than the projection pattern boundary, and it is removed from the set of projection pattern boundary feature points. The basis for the above determination is that the pattern content of dynamic light and shadow projection changes periodically with time, and the projection pattern boundary will inevitably produce observable displacement or color change in adjacent frames, while the inherent texture boundaries of the historical building's surface, such as brick joints, wooden painted decorations, and window frames, remain unchanged in position and color in all frames.

[0096] Then, optical flow tracking is performed on the detected projection pattern boundary feature points. Optical flow tracking refers to the method of calculating the motion vector of pixels between consecutive image frames to track the trajectory of the target contour changing over time. In this application, the Lucas-Kanade sparse optical flow algorithm based on pyramid layering is used for tracking. The Lucas-Kanade sparse optical flow algorithm is based on the assumption of grayscale constancy. The least squares solution of the optical flow constraint equation is performed within a local window of size w multiplied by w to obtain the two-dimensional motion vector of the feature points. The optical flow constraint equation is shown in equation (3) as follows:

[0097] In the above formula (3), This represents the grayscale gradient of the image in the x-direction, expressed in grayscale values ​​per pixel. This represents the grayscale gradient of the image in the y-direction, expressed in grayscale values ​​per pixel. The x-component of the motion vector, in pixels per frame; The component of the motion vector in the y-direction, in pixels per frame; The gray-level temporal gradient of an image between adjacent frames, assuming the time interval between adjacent frames is normalized to one frame. The unit is grayscale value per frame; The unit of measurement is grayscale value per pixel multiplied by pixels per frame, which equals grayscale value per frame. The dimension of this value is also grayscale. The dimension of is grayscale value per frame. Therefore, the dimension of the three terms in the above formula is grayscale value per frame. Since the dimensions are consistent, the equation is valid.

[0098] In this application, the number of pyramid layers is set to 4, and the local window size w is set to 21 pixels. The motion trajectory of the projection pattern boundary on the surface of the historical building is obtained by the motion vector of feature points between consecutive frames.

[0099] S44: Based on the projection pattern boundary and the historical building's reference scene, determine the surface of the historical building to which the dynamic light and shadow projection is attached.

[0100] The position coordinates of the projection pattern boundary obtained in step S43 on the unfolded two-dimensional parametric plane are back-projected to three-dimensional space through parametric inverse transformation to determine the three-dimensional spatial surface area covered by the dynamic light and shadow projection. This three-dimensional spatial surface area is the surface of the historical building to which the dynamic light and shadow projection is attached. The specific implementation of parametric inverse transformation is as follows: During the unfolding process in step S41, the mapping relationship between the three-dimensional vertex coordinates of each triangular facet and the corresponding two-dimensional parametric coordinates is recorded. During parametric inverse transformation, the triangular facet containing the target coordinate point on the two-dimensional parametric plane is first determined. Then, the centroid coordinates of the target coordinate point are calculated using the coordinates of the three vertices of the triangular facet on the two-dimensional parametric plane. Finally, the three-dimensional coordinates of the three vertices of the triangular facet are weighted and interpolated using the same centroid coordinates to obtain the coordinates of the target coordinate point in three-dimensional space. Since the centroid coordinate interpolation is unique and continuous within the triangular facet, the accuracy of parametric inverse transformation depends on the size of the triangular facet. In this application, the average side length of the triangular facet is about 0.02 meters, so the positioning accuracy of the inverse transformation is better than 0.02 meters.

[0101] S45: Write the correspondence to the light and shadow display layer. Specifically, write the correspondence between the dynamic light and shadow projection attached to the surface of the historical building and the surface of the historical building to which the dynamic light and shadow projection is attached into the light and shadow display layer.

[0102] S46: Overlay dynamic light and shadow projections from at least two projection periods onto the surface of the historical building to form a projection coverage trajectory. See also Figure 4 As shown, after completing steps S41 to S45 to form the light and shadow display layer, projection coverage trajectory analysis is performed. The projection coverage trajectory refers to the cumulative time trajectory of the area covered by dynamic light and shadow projection on the same historical building surface during multiple projection periods. The projection coverage trajectory is used to record the spatial location and duration of the same historical building surface covered by dynamic light and shadow projection during different projection periods.

[0103] The specific process is as follows: Dynamic light and shadow projections in at least two projection periods are superimposed on the surface of the historical building to which they are attached. A projection period refers to a time segment divided into equal time intervals within a complete dynamic light and shadow projection cycle. The duration of each projection period is set as the dynamic light and shadow projection cycle divided by the number of segments. For example, when the cycle is 8 minutes and divided into 16 projection periods, the duration of each projection period is 30 seconds. For each pixel position on the surface of each historical building, the ratio of the number of time periods covered by dynamic light and shadow projections in all projection periods to the total number of projection periods is calculated. This ratio is used as the continuous coverage ratio of the pixel position. The continuous coverage ratios of each pixel position are arranged according to the corresponding surface of the historical building to obtain the projection coverage trajectory.

[0104] S47: Extract the inherent appearance of the historical building from the surface segments not continuously covered by the projection pattern in the projection coverage trajectory, and use the inherent appearance as the basis for supplementing the historical building layer.

[0105] Specifically, pixel locations with a continuous coverage ratio of less than 1 are identified as surface segments not continuously covered by the projected pattern. The color and texture information collected during the time period when the pixel location is not covered by the dynamic light and shadow projection is the inherent appearance content of the pixel location. When color information is collected for a pixel location in multiple uncovered time periods, the median of the color values ​​in each uncovered time period is taken as the inherent appearance content of the pixel location to eliminate the interference of random noise. When a pixel location is covered by the dynamic light and shadow projection in all projection time periods, i.e., the continuous coverage ratio is equal to 1, the inherent appearance content of the pixel location cannot be directly extracted. In this case, the inherent appearance content of the pixel locations with a continuous coverage ratio of less than 1 in the four neighboring areas of the pixel location is used as the estimated value by inverse distance weighted interpolation. If there are no usable pixel locations in the four neighboring areas, the search range is gradually expanded until usable inherent appearance content is found. Since the color information collected at night is affected by the ambient light color temperature and deviates from the inherent color during the day, color correction is required. The color correction method is as follows: taking the daytime color value of the unobstructed area on the same building surface in the daytime scene before the ice and snow activities in step S1 as a reference, calculate the channel mean difference between the color value of the nighttime uncovered period and the daytime color value in the same area, and use this difference to perform channel-by-channel translation correction on all inherent appearance content extracted at night.

[0106] In practical engineering applications, the light and shadow display layer was extracted from 24 continuous surfaces of the walls of three buildings in a historical district. In three complete dynamic light and shadow projection cycles, step S43 tracked approximately 8,600 feature points of the projection pattern boundary. Through projection coverage trajectory analysis, it was found that approximately 22% of the area on the wall surface of the three buildings was not covered by dynamic light and shadow projection during all projection periods. From these uncovered areas, a total of 38 samples of inherent appearance content of three types of historical materials, namely brick walls, wooden pillars, and plaster, were extracted.

[0107] This application separates dynamic light and shadow projection from the inherent texture of the surface of the historical building through surface unfolding, snow and ice area masking, projection pattern boundary detection, optical flow tracing, projection coverage trajectory analysis, and inherent appearance content extraction. It establishes a spatial correspondence between the dynamic light and shadow projection content and the surface of the historical building, and extracts inherent appearance content samples for subsequent appearance replacement.

[0108] S5: Based on the shading area represented by the snow and ice activity layer and the coverage area represented by the light and shadow display layer, the historical ontology layer is supplemented by combining the urban heritage information model.

[0109] In steps S3 and S4, an ice and snow activity layer and a light and shadow display layer are formed, respectively. Since the ice and snow activity layer represents the geometrically occluded areas of the historical building caused by snow accumulation, ice sculptures, and snow sculptures, and the light and shadow display layer represents the visually covered areas of the historical building's surface texture caused by dynamic light and shadow projection, the occluded or covered areas in the historical building's baseline scene need to be supplemented with external knowledge to restore the complete appearance of the historical building before the impact of ice and snow activities. However, geometric inference alone cannot determine the type of building components and historical material information of the occluded areas; therefore, it is necessary to introduce historical archive information recorded in the urban heritage information model as a basis for supplementation. The historical building layer refers to the historical building data layer that has had its complete geometric shape and inherent appearance restored after supplementation processing.

[0110] Therefore, the input in step S5 is not the abstract data layer name itself, but rather the occlusion area, contact edge, and body continuation line recorded in the snow and ice activity layer, and the coverage area, projection coverage trajectory, and inherent appearance content recorded in the light and shadow display layer. Based on the above regional information and combined with the urban heritage information model, the completion module performs component morphology completion and appearance content replacement on the geometrically missing areas and the appearance pollution areas, respectively.

[0111] A city heritage information model (CIM) is a digital information model based on Building Information Modeling (BIM) technology, integrating multi-dimensional heritage information such as conservation archives, survey records, material samples, and landscape planning for historical buildings. The data source for the CIM is archival materials generated by urban planning and cultural heritage protection authorities during the identification and conservation planning process of historical buildings. The construction process of the CIM is as follows: A professional institution with historical building surveying qualifications, based on the requirements of the conservation planning and using total station measurement data and existing survey drawings, establishes a three-dimensional component model of each historical building on the BIM platform. The spatial reference coordinate system of the model adopts the standard 2000 geodetic coordinate system or a local independent engineering coordinate system. The modeling accuracy is such that the component positioning error does not exceed 0.1 meters. The coordinates of the center point of the geometric bounding box of each component are used as the reference coordinates of that component and written into the model attribute table. Based on the geometric model, the building component type, historical material name, conservation level label, and streetscape information from the conservation archives are entered into the corresponding component attribute fields. The data structure of the urban heritage information model is as follows: each historical building is an independent record unit. Each record includes the building number, building name, list of building components, type number of each component, reference coordinates of each component, historical material name of each component, protection level label of each component, and the appearance information of the street where the building is located. An example of the data structure of the urban heritage information model is shown in Table 4.

[0112] Table 4. Example of data structure for urban heritage information model

[0113] As shown in Table 4, each record in the urban heritage information model uniquely identifies a building component through its building number and component type number, and records the historical material name, protection level, and streetscape information of the component.

[0114] S50: Before supplementing the historical ontology layer by combining the urban heritage information model with the shading area represented by the ice and snow activity layer and the coverage area represented by the light and shadow display layer, exhibition data corresponding to ice and snow activities can also be obtained.

[0115] The exhibition data includes ice sculpture locations, snow sculpture locations, projected building surfaces, and projection time periods. This data is generated by the organizers of the ice and snow activities during the planning phase. Ice and snow sculpture locations are the marked coordinates of the ice and snow sculptures on the target historical scene plan. Projected building surfaces are the building facade or surface number corresponding to each projection device. Projection time periods are the start and end times of each projection device or time segments within the projection cycle. The coordinates of the ice sculpture locations in the exhibition data are transformed to the coordinate system of the historical building's reference scene using the same iterative nearest-neighbor registration algorithm as in step S32. The transformed ice sculpture location coordinates are then matched with the centroid coordinates of the ice sculptures in the ice and snow activity layer using nearest-neighbor matching to achieve positional correspondence. The matching distance threshold is set to 3 meters. The positional correspondence between the snow sculpture locations and the snow sculptures in the ice and snow activity layer is the same as that between the ice sculpture locations. The projected building surfaces are spatially correlated with the historical building surfaces to which the dynamic light and shadow projections in the light and shadow display layer are attached, and the projection time periods are temporally correlated with the dynamic light and shadow projections in the light and shadow display layer. By matching the spatial and temporal attributes of the exhibition materials, we can verify the spatial location and projection sequence of the recognition results in the ice and snow activity layer and the light and shadow display layer, and provide auxiliary basis for determining the occluded and covered areas when supplementing the historical ontology layer.

[0116] The specific implementation process of this step includes steps S51 to S55.

[0117] S51: The body missing region is formed based on the shading area represented by the snow and ice active layer.

[0118] The missing part of the entity refers to the portion of the historical building entity that is spatially occupied by temporary ice and snow entities in the historical building entity baseline scene. The geometric boundary of the missing part of the entity is determined by the contact edge between each temporary ice and snow entity in the ice and snow activity layer and the historical building entity.

[0119] S52: The apparent contamination area of ​​the body is formed based on the coverage area represented by the light and shadow display layer.

[0120] The apparent contamination area refers to the part of the historical building that is covered by dynamic light and shadow projection in the historical building's baseline scene. The apparent contamination area is geometrically intact, but the visual information of the historical building's surface in this area is replaced by the content of the dynamic light and shadow projection.

[0121] S53: Establish a correspondence between the continuous surface of the historical building in the benchmark scene and the information on the type of building components, historical materials, protection level and street appearance recorded in the urban heritage information model to obtain the correspondence of heritage attributes of the continuous surface.

[0122] The method for establishing the correspondence between heritage attributes of the continuous surface of the ontology is as follows: First, the coordinate reference system of the urban heritage information model is unified to the same engineering coordinate system as the 3D reconstruction result in step S21 through control point coordinate transformation. The coordinate transformation method is a seven-parameter Helmert transformation based on no less than four control points with the same name. The control points are the clearly identifiable building corners in the historical district, and the coordinate transformation residual is no more than 0.1 meters. Then, the spatial centroid coordinates of each continuous surface of the ontology are used as matching features and nearest neighbor matching is performed with the reference coordinates of the building components recorded in the coordinate-transformed urban heritage information model. The reference coordinates of the building components in the urban heritage information model refer to the coordinates of the center point of the geometric bounding box of the component in the building information model. The matching distance threshold is set to 2 meters. The basis for determining this threshold is that the distance between the centroids of different components of the same historical building is usually greater than 2 meters, while the coordinate deviation of the same component in the two data sources is usually less than 2 meters after coordinate transformation. Therefore, the 2-meter threshold can ensure the success rate of correct matching while avoiding mismatching across components. The successfully matched continuous surface of the ontology inherits all attribute records of the corresponding building component in the urban heritage information model.

[0123] When the nearest neighbor matching distance exceeds the 2-meter threshold but is less than 4 meters, or when there are two or more candidate components within the 2-meter threshold, a secondary verification is initiated: First, it is verified whether the component category of the continuous surface of the subject to be matched is consistent with the component type recorded in the urban heritage information model of the candidate component; second, the angle between the average normal vector of the continuous surface of the subject to be matched and the reference normal vector of the candidate component is calculated, and the candidate component is excluded when the angle is greater than 45 degrees; third, the ratio between the bounding box size of the continuous surface of the subject to be matched and the bounding box size of the candidate component recorded in the urban heritage information model is compared, and when the ratio is less than 0.5 or less... Candidate components with a score greater than 2.0 are excluded. Finally, the floor and facade numbers recorded in the urban heritage information model are used to verify whether the floor and facade orientation of the continuous surface of the subject to be matched matches the floor and facade orientation of the candidate component. If multiple candidate components still exist after the above two verifications, they are sorted according to the weighted scores of three indicators: matching distance, normal vector angle, and bounding box size ratio, with weights of 0.5, 0.3, and 0.2, respectively. The candidate component with the highest comprehensive score is selected as the final matching result. When all candidate components fail the two verifications, the continuous surface of the subject is marked as pending manual confirmation.

[0124] S54: Based on the information on the types of building components, historical materials, protection levels, and streetscapes recorded in the urban heritage information model, as well as the correspondence between the heritage attributes of the continuous surface of the main body, determine the types of building components, historical materials, protection levels, and streetscapes of the missing areas of the main body. Using the main body continuity line, the continuity relationship between adjacent building components, and the arrangement relationship of similar components as constraints, supplement the component forms of the missing areas of the main body.

[0125] The specific method for component morphology completion is as follows: The continuity relationship of adjacent building components refers to the topological adjacency relationship established when two different types of continuous surfaces share at least one boundary line segment in the historical building body reference scene. For example, the continuous surface of the roof body and the continuous surface of the wall body are adjacent through the eaves line segment. The continuity relationship of adjacent building components is recorded in the form of an adjacency matrix, which records the coordinates of the shared boundary line segment of each pair of adjacent continuous surfaces. The arrangement relationship of similar components refers to the regularity in geometric scale and spatial spacing between multiple continuous surfaces of the same component type in the same building or the same block. For example, multiple windows of the same building are arranged at equal intervals on the wall and have the same size. The arrangement relationship of similar components is characterized by two parameters: the average size of the geometric bounding box of each similar continuous surface and the average adjacent spacing.

[0126] The specific algorithm for component morphology completion is as follows: First, on the boundary of the missing body region, using the body continuation line generated in step S34 as the direction constraint of the occluded component outline, the cross-sectional outlines of the visible component outlines at both ends of the body continuation line are linearly interpolated along the direction of the body continuation line to generate the intermediate cross-sectional sequence of the component outlines in the occluded region. Adjacent cross-sections are connected by triangular facets to form a component surface mesh. Second, based on the shared boundary line segments recorded in the continuation relationship of adjacent building components, the newly generated component surface mesh in the completion region is vertex-stitched with the boundary of the adjacent visible body continuous surface. The stitching method is to project the vertices on the boundary of the completion mesh onto the boundary line segments of the adjacent visible body continuous surface and merge them into the same vertex to maintain the continuity of the topological connection. Third, using the average geometric bounding box size and the average adjacent spacing recorded in the arrangement relationship of similar components as constraints, the completed component surface mesh is scaled and adjusted so that its geometric scale is consistent with the continuous surface of the same type of visible body. The scaling factor is the ratio of the average geometric bounding box size of the continuous surface of the same type of body to the current geometric bounding box size of the completion mesh.

[0127] For example, when the wall of a building B1 is obscured by an ice sculpture, the visible outline sections of the walls on both sides of the ice sculpture's contact edge are linearly interpolated along the body's continuity line to generate a triangular mesh of the obscured wall. Then, the upper boundary of this mesh is stitched to the lower boundary of the roof body's continuous surface, and the lower boundary is stitched to the upper boundary of the step body's continuous surface. Finally, the mesh is scaled and adjusted according to the average scale of the other visible walls of building B1.

[0128] S55: Based on the historical material attribution and inherent appearance of the building components belonging to the pollution area, replace the appearance formed by dynamic light and shadow projection to form a historical entity layer.

[0129] The specific method of appearance replacement is as follows: Based on the historical material type recorded in the urban heritage information model of the continuous surface to which the appearance pollution area belongs, select the inherent appearance content sample that matches the name of the historical material type from the inherent appearance content samples extracted in step S4; when there are multiple inherent appearance content samples of the same historical material type, the selection rule is: calculate the Euclidean distance between each candidate sample and the spatial position of the appearance pollution area in the historical building's reference scene, and select the candidate sample with the closest spatial distance as the texture source tile, based on the fact that similar material surfaces with similar spatial positions have high similarity in weathering degree, color change and texture details; extend the inherent appearance content sample to the entire appearance pollution area through a texture synthesis method based on tile splicing.

[0130] The specific process of the texture synthesis method based on tile splicing is as follows: An inherent appearance content sample is used as the texture source tile. The size of the texture source tile can be set according to actual needs. For example, in this application, the size of the texture source tile is set to 64 pixels by 64 pixels. Texture source tiles are placed row by row and column by column within the surface contamination area. An overlapping area is set between adjacent tiles. The width of the overlapping area is set to one-sixth of the width of the texture source tile. That is, under the condition that the size of the texture source tile is 64 pixels, the width of the overlapping area is 10 pixels. The basis for setting the width of the overlapping area to one-sixth of the tile width is that: if the width of the overlapping area is too small, the transition at the seam will not be smooth enough; if the width of the overlapping area is too large, it will reduce the utilization rate of the effective texture. In this application, after experimental comparison of 5 different overlapping ratios, it was found that the one-sixth ratio achieves a better balance between splicing quality and texture utilization rate.

[0131] In each overlapping region, a minimum seam path is searched using a dynamic programming algorithm. The search process for the minimum seam path is as follows: First, the pixel-by-pixel seam energy at each pixel position within the overlapping region is defined. Pixel-by-pixel seam energy The calculation formula (4) is as follows:

[0132] In the above formula, , , These are the red, green, and blue channel values ​​of the first patch at the i-th pixel position in the overlapping area, respectively, and are integers ranging from 0 to 255. , , These are the red, green, and blue channel values ​​of the second patch at the i-th pixel position in the overlapping area, respectively, and their values ​​are integers ranging from 0 to 255. The square of each channel difference is the square of the gray value, and the sum of the squares of the three channels is also the square of the gray value.

[0133] Then, starting from the top row of the overlapping region, for each pixel position in each row, calculate the minimum cumulative seam energy from the top row to that pixel position. The cumulative seam energy at each pixel position is equal to the pixel-wise seam energy at that pixel position. Add the minimum value of the cumulative seam energy among the three pixels adjacent to the current pixel position in the previous row, calculate row by row until the bottom row, and then backtrack from the pixel position with the minimum cumulative seam energy in the bottom row to the top row to obtain the minimum seam path. Then, crop and stitch the two adjacent tiles along the minimum seam path.

[0134] In practical engineering applications, assuming the size of a certain surface contamination area is 320 pixels by 240 pixels, the size of the texture source tile is 64 pixels by 64 pixels, and the width of the overlapping area is 10 pixels, then approximately 6 tiles (320 / (64-10)=5.9) are needed in the horizontal direction, and approximately 5 tiles (240 / (64-10)=4.4) are needed in the vertical direction. A total of about 30 tile placements are required. Each time a tile is placed, a minimum seam path is searched in the overlapping area in both the horizontal and vertical directions. Since one-sixth of 64 pixels is 10.67 pixels, rounded down to 10 pixels, there may be a remaining area on the right or bottom side of the last tile in each row or column that is less than the width of a complete tile step. For the remaining area, a cropped fragment of the texture source tile is used to fill it. A 10-pixel overlap is still set between the cropped fragment and the previous tile, and the same minimum seam path search method is used to stitch them together.

[0135] After steps S51 to S55 are completed, a total of 18 missing areas and 32 polluted areas are identified in a certain historical block. Step S53 obtains 168 component type information and 23 historical material records of three buildings from the urban heritage information model. Step S54 completes the component forms for the 18 missing areas. Step S55 replaces the appearance content generated by dynamic light and shadow projection with the inherent appearance content samples for the 32 polluted areas, and finally forms a complete historical ontology layer.

[0136] This application uses historical archive information from the urban heritage information model as a basis for supplementation, and restores the component form of the area covered by temporary ice and snow entities by constraints such as the ontological continuity line. It replaces the surface texture of the area contaminated by dynamic light and shadow projection with samples of inherent appearance content, thus forming a complete historical ontological layer.

[0137] S6: Combine the historical ontology layer, the ice and snow activity layer, and the light and shadow display layer according to the target display state to generate a digital scene of the target historical scene.

[0138] In steps S2 to S5, the historical building body base scene, the ice and snow activity layer, the light and shadow display layer, and the historical body layer have been formed respectively. Since different application scenarios have different requirements for the display content of the target historical scene digital scene, this step achieves flexible combination of different scene layers by selecting the target display state.

[0139] Among them, the target display state refers to the digital scene display mode selected by the user based on specific application needs.

[0140] This step defines three target display states, and the combination relationship between the three target display states and the scene layer is shown in Table 5: Table 5 Relationship between Target Display Status and Scene Layer Combination

[0141] As shown in Table 5, the historical entity restoration state corresponds to only combining the historical entity layer, that is, only displaying the complete appearance of the historical building entity after the supplementation process in step S5; the ice and snow activity display state corresponds to combining the historical entity layer and the ice and snow activity layer, that is, displaying the snow accumulation part, ice sculpture part and snow sculpture part on the basis of the complete appearance of the historical building entity; the night tour light and shadow display state corresponds to combining the historical entity layer, the ice and snow activity layer and the light and shadow display layer, that is, further displaying the effect of dynamic light and shadow projection on the basis of the ice and snow activity display state.

[0142] After generating the digital scene of the target historical scene, the historical ontology layer, the ice and snow activity layer, and the light and shadow display layer need to be saved as independently callable scene layers. Independent callable means that each scene layer is stored as an independent data file. When loading the digital scene, one or more scene layer files can be selectively loaded according to the target display state. Different loading combinations produce different versions of the target historical scene digital scene.

[0143] In practical engineering applications, the digital scene generated for a certain target historical scene contains three independently callable scene layer files. When the target display state is set to the historical ontology restoration state, only the historical ontology layer file is loaded. When the target display state is set to the ice and snow activity display state, both the historical ontology layer file and the ice and snow activity layer file are loaded. When the target display state is set to the night tour light and shadow display state, all three scene layer files are loaded.

[0144] This application distinguishes and forms independent scene layers for historical buildings, temporary ice and snow entities, and dynamic light and shadow projections in the target historical scene within the same generation process. The scene layers are flexibly combined according to different target display states, thereby generating a digital scene of the target historical scene that meets multiple needs such as historical building restoration, ice and snow activity display, and night tour light and shadow display in a single data acquisition and processing process. This improves the accuracy of digital restoration of historical buildings in ice and snow cultural tourism scenes, enhances the adaptability of digital scenes in different display modes, and reduces the workload of repeated acquisition and modeling of the same target historical scene.

[0145] This application provides an artificial intelligence-based automatic digital scene generation system, the system comprising: The acquisition module is used to acquire daytime scenes before ice and snow activities, daytime scenes during ice and snow activities, and nighttime scenes during ice and snow activities for the same target historical scene. The target historical scene includes historical blocks or historical buildings, and includes the historical building itself, temporary ice and snow entities, and dynamic light and shadow projections.

[0146] The first module is used to perform multi-view 3D reconstruction and semantic segmentation of architectural components for daytime scenes before ice and snow activities, forming a historical building ontology benchmark scene with component-level semantic annotations.

[0147] The second forming module is used to perform semantic segmentation of ice and snow materials in daytime scenes during ice and snow activities, and to perform three-dimensional change detection on the continuous surface of the ice and snow candidate area and the historical building body in the benchmark scene, forming an additional ice and snow activity layer relative to the historical building body.

[0148] The third forming module is used to perform surface unfolding, ice and snow area masking, and projection pattern boundary tracking on the nighttime image sequence in the nighttime scene during the ice and snow activities, based on the historical building's base scene and the ice and snow activity layer, to form a light and shadow display layer attached to the surface of the historical building.

[0149] The completion module is used to complete the historical ontology layer based on the shading area represented by the snow and ice activity layer and the coverage area represented by the light and shadow display layer, combined with the urban heritage information model.

[0150] The generation module is used to combine the historical ontology layer, the ice and snow activity layer, and the light and shadow display layer according to the target display state to generate a digital scene of the target historical scene.

[0151] The AI-based automatic digital scene generation system of this application is used to implement the aforementioned AI-based automatic digital scene generation method. To ensure that the system embodiments and method steps correspond and support each other, the data flow and collaborative relationships between the modules are described below.

[0152] Specifically, the three types of scene data output by the acquisition module serve as the input basis for the first formation module, the second formation module, and the third formation module; the first formation module outputs the historical building body baseline scene and sends it to the second formation module, the third formation module, and the supplementation module respectively; the second formation module outputs the ice and snow activity layer and sends it to the third formation module, the supplementation module, and the generation module; the third formation module outputs the light and shadow display layer and sends it to the supplementation module and the generation module; the supplementation module outputs the historical body layer and sends it to the generation module.

[0153] Furthermore, the third forming module may include an unfolding unit, a masking unit, a tracking unit, and a writing unit. The unfolding unit is used to register and project the nighttime image sequence from the nighttime scene during the ice and snow activities onto a two-dimensional parametric plane of the walls, eaves, and paved surfaces of the historical building, forming an unfolded nighttime image sequence; the masking unit is used to mask the snow, ice sculpture, and snow sculpture parts corresponding to the ice and snow activity layer in the unfolded nighttime image sequence, preserving the nighttime image content of the historical building's surface; the tracking unit is used to perform projection pattern boundary detection and optical flow tracking on the nighttime image content of the historical building's surface, obtaining the projection pattern boundary that moves on the surface of the historical building during the projection period; the writing unit is used to determine the surface of the historical building to which the dynamic light and shadow projection is attached based on the projection pattern boundary and the historical building's reference scene, and write the correspondence between the dynamic light and shadow projection and the attached historical building surface into the light and shadow display layer.

[0154] The system may also include a trajectory analysis unit, a continuation line generation unit, and an exhibition data correspondence unit. The trajectory analysis unit is used to overlay the dynamic light and shadow projections in at least two projection periods onto the surface of the historical building to form a projection coverage trajectory, and extract the inherent appearance content of the surface segments that are not continuously covered by the projection pattern; the continuation line generation unit is used to generate a continuation line along the contact edge between the snow-covered part, ice sculpture part, or snow sculpture part and the historical building, and to define the direction and boundary of the obscured historical building components through the continuation line; the exhibition data correspondence unit is used to correspond the ice sculpture points with the ice sculpture parts, the snow sculpture points with the snow sculpture parts, the projected building surface with the dynamic light and shadow projection in space, and the projection period with the dynamic light and shadow projection in terms of time attributes.

[0155] The system may also include a storage module, which is used to save the historical ontology layer, the ice and snow activity layer and the light and shadow display layer as scene layer files that can be called independently, so that the generation module can call one or more scene layer files according to the target display state to form the corresponding target historical scene digital scene version.

[0156] The above provides a detailed description of the method and system for automatically generating digital scenes based on artificial intelligence, as provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A method for automatically generating digital scenes based on artificial intelligence, characterized in that, include: Acquire daytime scenes before ice and snow activities, daytime scenes during ice and snow activities, and nighttime scenes during ice and snow activities for the same target historical scene. The target historical scene includes the historical building itself, temporary ice and snow entities, and dynamic light and shadow projections. Multi-view 3D reconstruction and semantic segmentation of building components were performed on the daytime scene before the ice and snow activities to form a historical building ontology benchmark scene with component-level semantic annotations. The ice and snow material semantic segmentation is performed on the daytime scene during the ice and snow activities, and the ice and snow candidate regions are subjected to three-dimensional change detection with the continuous surface of the historical building body in the reference scene to form an ice and snow activity layer added relative to the historical building body. Based on the historical building's reference scene and the ice and snow activity layer, the nighttime image sequence in the nighttime scene during the ice and snow activity is surface-unfolded, ice and snow area masked, and projection pattern boundary tracked to form a light and shadow display layer attached to the surface of the historical building. Based on the shading area represented by the snow and ice activity layer and the coverage area represented by the light and shadow display layer, a historical ontology layer is supplemented by combining the urban heritage information model. The historical ontology layer, the ice and snow activity layer, and the light and shadow display layer are combined according to the target display state to generate a digital scene of the target historical scene.

2. The method according to claim 1, characterized in that, Based on the historical building's baseline scene and the snow and ice activity layer, the nighttime image sequence in the nighttime scene during the snow and ice activity is surface-unfolded, snow and ice area masked, and projection pattern boundary tracked to form a light and shadow display layer attached to the surface of the historical building, including: The nighttime image sequence from the nighttime scene during the ice and snow activities is registered and projected onto a two-dimensional parametric plane of the walls, eaves and paved surfaces of the historical building to form an unfolded nighttime image sequence. In the unfolded nighttime image sequence, the snow-covered, ice-sculpted, and snow-sculpted portions corresponding to the ice and snow activity layer are masked to preserve the nighttime image content of the surface of the historical building. Projection pattern boundary detection and optical flow tracking are performed on the nighttime image content of the historical building's surface to obtain the projection pattern boundary that moves on the surface of the historical building as the projection time period changes. Based on the boundary of the projection pattern and the reference scene of the historical building, the surface of the historical building to which the dynamic light and shadow projection is attached is determined; The correspondence between the dynamic light and shadow projections attached to the surface of the historical building and the surface of the historical building to which the dynamic light and shadow projections are attached is written into the light and shadow display layer.

3. The method according to claim 2, characterized in that, Before supplementing the historical ontology layer based on the shading area represented by the snow and ice activity layer and the coverage area represented by the light and shadow display layer, and combining it with the urban heritage information model, the following is also included: The dynamic light and shadow projections in at least two projection periods are superimposed on the surface of the historical building to form a projection coverage trajectory. Extract the inherent appearance of the historical building from the surface segments that are not continuously covered by the projected pattern in the projection coverage trajectory; The inherent appearance content is used as the appearance basis for supplementing the historical ontology layer.

4. The method according to claim 3, characterized in that, The process of constructing a historical ontology layer based on the shading area represented by the snow and ice activity layer and the coverage area represented by the light and shadow display layer, combined with the urban heritage information model, includes: Based on the shielding area represented by the ice and snow activity layer, a body missing area is formed, and the body missing area corresponds to the part of the historical building body occupied by the ice and snow temporary entity. The apparent pollution area of ​​the building body is formed based on the coverage area represented by the light and shadow display layer, and the apparent pollution area of ​​the building body corresponds to the part of the historical building body covered by dynamic light and shadow projection. Establish a correspondence between the continuous surface of the historical building in the benchmark scene and the information on the type of building components, historical materials, protection level and street appearance recorded in the urban heritage information model to obtain the correspondence of heritage attributes of the continuous surface. Based on the architectural component types, historical materials, protection levels, and streetscape information recorded in the urban heritage information model, and the correspondence of heritage attributes of the continuous surface of the main body, the architectural component types, historical materials, protection levels, and streetscape information of the missing area of ​​the main body are determined. The component form of the missing area of ​​the main body is supplemented by constraints such as the main body continuation line, the continuation relationship of adjacent architectural components, and the arrangement relationship of similar components. The main body continuation line refers to the geometric outline of the historical architectural components that still exist in the area where they are temporarily covered by ice and snow. Based on the historical material attribution of the building components to which the apparent pollution area belongs and the inherent apparent content, the apparent content formed by the dynamic light and shadow projection is replaced to form a historical ontological layer.

5. The method according to claim 1, characterized in that, The process involves performing multi-view 3D reconstruction and semantic segmentation of architectural components on the daytime scene before the ice and snow activities to form a historical building ontology benchmark scene with component-level semantic annotations, including: Multi-view 3D reconstruction of the daytime scene before the ice and snow activities was carried out to form the spatial surface of the historical building body; The spatial surface is semantically segmented to obtain the corresponding continuous surfaces of the roof, walls, eaves, doors and windows, steps and street paving. The continuous surface of the body is written into the scene space according to the component affiliation of the historical building body to form the benchmark scene of the historical building body.

6. The method according to claim 5, characterized in that, The step of performing semantic segmentation of ice and snow materials in the daytime scene during the ice and snow activities, and performing 3D change detection on the ice and snow candidate regions and the continuous surfaces of the historical building in the baseline scene to form a new ice and snow activity layer relative to the historical building, includes: The daytime scenes during the ice and snow activities are subjected to semantic segmentation of ice and snow materials to obtain ice and snow candidate regions; The candidate ice and snow region and the continuous surface of the body are subjected to three-dimensional change detection to form an active period space occupancy surface located outside the continuous surface of the body. The activity period space occupancy area is divided into the snow part attached to the surface of the historical building, the ice sculpture part occupying the passage space of the target historical scene, and the snow sculpture part occupying the passage space of the target historical scene. The snow-covered portion, the ice sculpture portion, and the snow sculpture portion are classified into the snow and ice activity layer.

7. The method according to claim 6, characterized in that, After classifying the snow-covered portion, the ice sculpture portion, and the snow sculpture portion into the snow and ice activity layer, the following is also included: Along the contact edge between the snow-covered part, the ice sculpture part, or the snow sculpture part and the main body of the historical building, following the extension direction of the roof ridge, eaves, door and window frames, step edges, and street paving seams, a main body continuity line interrupted by the temporary ice and snow entity is generated. The orientation and boundaries of the obscured historical building components are defined by the body continuation line.

8. The method according to claim 1, characterized in that, include: The target display state is one of the following: historical ontology restoration state, ice and snow activity display state, and night tour light and shadow display state; The restored state of the historical ontology corresponds to the combination of the historical ontology layer; The ice and snow activity display state corresponds to the combination of the historical ontology layer and the ice and snow activity layer; The nighttime light and shadow display corresponds to the combination of the historical entity layer, the ice and snow activity layer, and the light and shadow display layer.

9. The method according to claim 2, characterized in that, Before supplementing the historical ontology layer based on the shading area represented by the snow and ice activity layer and the coverage area represented by the light and shadow display layer, and combining it with the urban heritage information model, the following is also included: Obtain the exhibition information corresponding to the ice and snow activities, including ice sculpture locations, snow sculpture locations, projected building surfaces, and projection time periods; The positions of the ice sculpture points are matched with the positions of the ice sculpture parts in the ice and snow activity layer, and the positions of the snow sculpture points are matched with the positions of the snow sculpture parts in the ice and snow activity layer. Spatially correspond the projected building surface to the dynamic light and shadow projection in the light and shadow display layer, and correspond the projection time period to the dynamic light and shadow projection in the light and shadow display layer in terms of time attributes.

10. The method according to claim 1, characterized in that, After generating the digital scene of the target historical scene by combining the historical ontology layer, the ice and snow activity layer, and the light and shadow display layer according to the target display state, the process further includes: The historical ontology layer, the ice and snow activity layer, and the light and shadow display layer are each saved as scene layers that can be called independently; At least one of the scene layers is invoked according to the target display state to form a digital scene version of the target historical scene.

Citation Information

Patent Citations

  • Twin power station collision visual simulation method based on physical simulation

    CN121211870A

  • Unmanned aerial vehicle refined surveying and mapping method for digital elevation on natural ice and snow surface

    CN122015767A