Multi-source perception three-dimensional security scene reconstruction method for digital twin park

By collecting multi-source sensing data in the digital twin park and generating unified timestamp and quality marker records, and combining the geometrically consistent scene skeleton results to perform dynamic target removal and boundary distortion suppression, the problems of boundary ghosting and scale drift in multi-source data fusion are solved, and high-precision and stable 3D security scene reconstruction is achieved.

CN121962470BActive Publication Date: 2026-07-24ASCEND IT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ASCEND IT CO LTD
Filing Date
2026-03-30
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In digital twin parks, existing technologies lack unified timestamps and pose registration for the fusion of multi-source sensing data, leading to boundary ghosting, scale drift, and incorrect writing of noise points, making it difficult to meet the accuracy, stability, and accountability requirements of security scenarios.

Method used

By collecting multi-source sensing data and writing it into the data source registration record, generating a unified timestamp, constructing a multi-source data sequence and quality marker record, and combining the geometrically consistent scene skeleton results to perform dynamic target removal and boundary distortion suppression, multi-modal deep fusion and point cloud densification are achieved to generate a three-dimensional security scene model.

Benefits of technology

It improves the accuracy, stability, and adaptability to severe weather in the 3D representation of park security, reduces reliance on manual intervention, and ensures the traceability and verifiability of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962470B_ABST
    Figure CN121962470B_ABST
Patent Text Reader

Abstract

This invention discloses a method for reconstructing multi-source sensing 3D security scenes in digital twin parks, belonging to the field of multi-source sensing 3D security scene reconstruction technology in digital twin parks. It addresses the problems of inconsistent spatiotemporal calibers of multi-source data, difficulty in constraining geometric benchmarks, and the tendency for dynamic targets to be fixed, leading to difficulties in accurate reconstruction and traceable fixation of 3D security scenes. The method registers the data sources and generates a unified timestamp using four timestamps, completing cross-sensor extrinsic and intrinsic parameter calibration. Based on oblique image multi-view geometric modeling and laser point cloud registration, a dense color point cloud is generated through multimodal depth fusion of sparse depth maps from video frames. After dynamic target removal and distortion suppression, block-based feature registration and stitching are performed to solidify and output a 3D security scene model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-source sensing 3D security scene reconstruction technology for digital twin parks, and more specifically, to a method for multi-source sensing 3D security scene reconstruction for digital twin parks. Background Technology

[0002] Digital twin parks require the integration of oblique imagery, laser point clouds, millimeter-wave point clouds, and security videos into a queryable and traceable 3D security scene for intrusion detection, patrol review, and situational awareness simulation. Existing methods often model or stitch data offline based on individual sensors, lacking unified registration and verification of data sources, poses, and timestamps. Cross-source correlation often relies on manual point selection or experience-based alignment. When extrinsic parameter drift, unstable timing, or occlusion causes sparse depth breaks, the fusion results are prone to boundary ghosting, thin surfaces, and scale drift, with dynamic targets such as personnel and vehicles being fixed into static models. Furthermore, the large area of ​​the park, numerous data collection batches, and continuous updates mean that the lack of quality marking and geometric consistency verification allows local errors to accumulate during stitching. Additionally, under conditions such as rain and fog, millimeter-wave and laser echoes differ significantly; without geometric constraints and dynamic removal mechanisms, noise points can be mistakenly written in, failing to meet the accuracy, stability, and traceability requirements of security scenarios.

[0003] To address the above problems, this invention proposes a solution. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a method for reconstructing multi-source perception three-dimensional security scenes for digital twin parks, so as to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: In a preferred embodiment, it includes: Collect multi-source sensing data and write it into the data source registration record, generate a unified timestamp, and construct a multi-source data sequence and quality label record; Construct a 3D model of the tilted image and align it with the laser point cloud reference to obtain a geometrically consistent scene skeleton result, and write it into the quality mark record; A set of colored point cloud frames is generated using the geometrically consistent scene skeleton result as a geometric constraint. A temporal correspondence of corresponding points is established between adjacent colored point cloud frames, and the displacement velocity is calculated. The velocity threshold is dynamically adjusted based on the alignment instability coefficient, coverage compression, and branch conflict index. The 3D points are classified into dynamic target points and solidified suppression points and processed in combination with pixel motion and quality mark records. When the laser point cloud has gaps or the availability of quality mark records decreases under rain, fog, strong reflection, or occlusion conditions, the millimeter-wave point cloud is transformed to the unified spatial reference system of the park based on the extrinsic rotation matrix and extrinsic translation vector and participates in merging. Only millimeter-wave points that are determined to be usable by the quality mark records and whose nearest neighbor distance from the millimeter-wave point to the geometrically consistent scene skeleton result is not greater than a preset distance threshold are allowed to participate in merging. When the branch conflict index is not less than a preset threshold, the millimeter-wave point is further restricted from falling into the neighborhood of solidified suppression points. A dense static colored point cloud set and quality mark records are output. Register, stitch, and merge dense static color point cloud sets to generate triangular meshes and project security video frame textures to output a 3D security scene model.

[0006] In a preferred embodiment, for oblique images, laser point clouds, millimeter-wave point clouds, and security video frames, a data source registration record is first established for each data object and the source credentials and original carrier index are written. Then, multi-source sensing data is collected synchronously and the corresponding pose and lens intrinsic parameter matrix are recorded. Then, a unified timestamp is generated by performing four-timestamp time synchronization interaction with each acquisition device using a time synchronization reference node.

[0007] In a preferred embodiment, by constructing the correspondence constraints between point clouds and images and solving for the extrinsic rotation matrix and extrinsic translation vector with the minimum reprojection error as the criterion, a calibration parameter record is formed together with the lens intrinsic parameter matrix; for various types of data, the acquisition integrity, timestamp validity, pose validity, and measurement saturation or missing quality marker fields are calculated and written, and written into the quality marker record, resulting in a multi-source data sequence that can be expressed under the same spatiotemporal reference and can be used according to the constraints of the quality marker record.

[0008] In a preferred embodiment, based on the oblique image sequence with unified timestamp and pose metadata, image feature points that can be repeatedly located are first extracted in each oblique image and a matching relationship is established between overlapping viewpoints. The exterior orientation parameters of each oblique image are solved and the image sparse point cloud is obtained by back calculation. Then, global joint optimization is performed on all observations with the minimum pixel reprojection error as the criterion. Subsequently, dense point clouds of the image are generated using multi-view parallax constraints, and triangular meshes and texture mappings are constructed based on these to obtain a 3D model of the tilted image. At the same time, the laser point cloud sequence is transformed to a unified spatial reference system of the park according to pose metadata and multiple frames are merged to form a laser point cloud reference. Rigid body alignment transformation is solved by iterative nearest point registration to eliminate residual translation and rotation deviations, so that the 3D model of the tilted image is aligned with the laser point cloud reference and a geometrically consistent 3D scene skeleton result is output. Finally, the nearest neighbor distance statistics between the point set of the 3D model of the tilted image and the point set of the laser point cloud reference are calculated to verify geometric consistency and written into the quality mark record. When the alignment residual exceeds a preset threshold, a check mark is written to constrain the output accuracy.

[0009] In a preferred embodiment, the system takes a sequence of security video frames, laser point clouds, and millimeter-wave point clouds with uniform timestamps, recorded with calibration parameters, as input, and a geometrically consistent scene skeleton as geometric constraint. First, security video frames and corresponding laser point cloud frames are selected according to a uniform timestamp or a short time window. The laser point cloud frames are transformed to the coordinate system of the security camera device through an extrinsic rotation matrix and an extrinsic translation vector, and then projected onto the pixel plane according to the lens intrinsic parameter matrix to generate a sparse depth map. Subsequently, the texture edge information of the security video frames and the distance observation information of the sparse depth map are used as constraints to perform multimodal depth fusion inference to obtain a dense depth map. Then, the dense depth map is back-projected to generate color point cloud frames with bound colors, and uniformly transformed to a uniform spatial reference system of the park. Multi-view color point cloud frames within the same uniform timestamp or short time window are merged to form a dense static color point cloud set.

[0010] In a preferred embodiment, the alignment instability coefficient is constructed from the alignment jump variable; The coverage compression is constructed from the effective coverage rate of the boundary neighborhood; The branch conflict index is constructed by the degree of deviation between the color branch prediction and the depth branch prediction in the boundary neighborhood; The corresponding time sequence of the same point is as follows: under the unified spatial reference system of the park, the three-dimensional point at time t is used as the query point, the nearest neighbor point is searched in the color point cloud frame at time t+Δt, and the corresponding relationship is established when the nearest neighbor distance is not greater than the preset spatial threshold and the quality mark records of the two frames are determined to be available. Where Δt is the time interval between two adjacent frames corresponding to the same timestamp.

[0011] In a preferred embodiment, points that meet the conditions of displacement velocity, pixel motion, and quality mark recording are marked as dynamic target points and removed from the color point cloud frame set; points whose displacement velocity exceeds the threshold but are in a state of alignment instability, coverage breakage, or high risk of branch conflict, or whose pixel motion is insufficient for mutual verification, are marked as solidification suppression points. They are temporarily not included in the merging of the dense static color point cloud set within the scrolling window to suppress the solidification of ghosting surfaces and thin surfaces. When the above dynamic quantities continuously and stably meet the threshold conditions, the solidification suppression mark is removed and they are allowed to participate in the merging of the dense static color point cloud set. If the dynamic target point determination conditions are met during the observation period, they are removed. The neighborhood of the solidification inhibition point is a spatial neighborhood constructed with the coordinates of the solidification inhibition point in the unified spatial reference system of the park as the center and with a preset spatial radius.

[0012] In a preferred embodiment, the dense static color point cloud set is spatially divided into point cloud blocks with overlapping regions, and unusable points are removed based on quality marker records and necessary downsampling is performed; then, point cloud features are extracted from the blocks, and cross-block corresponding point pairs are obtained by filtering within the overlapping regions; the rigid body stitching transformation between blocks is solved based on the cross-block corresponding point pairs, and fine registration is completed by combining iterative nearest point registration; all blocks are repeatedly stitched, and overlapping covered points are merged for consistency based on quality marker records and residual distance indicators; A triangular mesh is generated from a globally dense static color point cloud set. The security video frame texture is projected onto the mesh surface using the lens intrinsic parameter matrix, extrinsic parameter rotation matrix, and extrinsic parameter translation vector, and then solidified to output a 3D security scene model.

[0013] The technical effects and advantages of this invention for multi-source sensing 3D security scene reconstruction in digital twin parks: This invention ensures that each type of data is traceable and constrained by using source credentials and quality markers throughout the acquisition, calibration, modeling, fusion, and stitching processes; it uses laser point cloud benchmarks to align tilted modeling results, suppressing scale drift and providing a verifiable geometric skeleton; it drives multimodal fusion with sparse depth observations and dynamically adjusts thresholds based on alignment jumps, coverage compression, and branch conflicts to achieve dynamic target removal and boundary distortion suppression; and it solidifies and outputs a continuous textured 3D scene after block-level feature registration and fine registration, improving the accuracy, stability, and adaptability to severe weather in the 3D representation of park security, facilitating continuous updates, and reducing reliance on manual intervention. Attached Figure Description

[0014] Figure 1 This is a flowchart illustrating the implementation of the multi-source perception 3D security scene reconstruction method for digital twin parks according to the present invention.

[0015] Figure 2 This is a timing diagram of the multi-source perception 3D security scene reconstruction method for digital twin parks according to the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] Example This invention discloses a method for reconstructing multi-source sensing 3D security scenes for digital twin parks, such as... Figure 1 As shown, it includes: Step 1: Acquisition and spatiotemporal calibration of multi-source sensing data; In this step, the acquisition of multi-source sensing data, source registration, unified timestamp generation, and cross-sensor spatial calibration and alignment are performed to obtain tilted images, laser point clouds, millimeter-wave point clouds, and security video frames that can be expressed under the same spatiotemporal reference, and to form corresponding calibration parameters and quality marks.

[0018] First, a data source registration record is established for each data object entering this method, and a source credential corresponding to each data object is written at the time of acquisition. The data source registration record includes at least: a data source type marker, an acquisition device identifier, an acquisition time recording method marker, an acquisition pose recording method marker, a sampling frequency or frame rate, a resolution or point cloud parameter, and an original carrier index; wherein, the original carrier index points to an immutable path or hash identifier of the oblique image file, point cloud file, or video frame sequence, so that any data object has a verifiable source credential.

[0019] The source credential is a verifiable credential marker used to verify the data object collection process and the collection subject. It points to the collection log number, device signature value, collection session number or equivalent verifiable credential, so that the data source registration record can be traced and verified.

[0020] Subsequently, multi-source sensing data was collected simultaneously within the park: a drone equipped with an oblique camera was used to acquire oblique image sequences, and the drone pose information corresponding to each oblique image was recorded; a LiDAR was deployed on the ground to collect laser point cloud sequences and record the radar pose; when it was necessary to enhance perception in adverse weather conditions and at long distances, millimeter-wave point cloud sequences were collected and the radar pose was recorded; and the video stream output by the security camera was retrieved, the security video frame sequence was extracted, and the pose and lens intrinsic parameters of the security camera were recorded.

[0021] The aforementioned pose can be obtained by an inertial measurement unit, odometer, positioning receiver, or equivalent pose measurement method, and written into the data source registration record along with the data object.

[0022] Subsequently, to eliminate local clock differences between different acquisition devices, this step generates a unified timestamp. Specifically, using the time synchronization reference node as the time base, a four-timestamp time synchronization interaction is performed with any acquisition device; The time synchronization reference node is a time synchronization device or time server deployed in the campus network. Its time reference can be provided by satellite time synchronization, a high-stability clock source or an equivalent time reference, and it provides time synchronization interaction services to the outside world as the only time reference for the generation of unified timestamps.

[0023] like Figure 2 As shown, the time when the acquisition device sends the request is The timing reference node receives the time as follows: The timing reference node sends the response at the following time. The data acquisition device receives the response at the following time. Under the condition of approximately symmetrical communication delay, this step calculates the estimated clock offset of the acquisition device relative to the timing reference node. for: ; This will be used to record the local timestamp written by the acquisition device for any data object. Corrected to a uniform timestamp : ; in, This is a clock offset estimate; , , , Four timestamps for time synchronization interaction; The data is collected using the device's local timestamp. To standardize timestamps, this step will... Metadata fields of oblique images, laser point clouds, millimeter-wave point clouds, and security video frames are written to form a multi-source data sequence under a unified timeline.

[0024] Next, this step performs cross-sensor spatial calibration, enabling point clouds and images to be mapped to each other under the same spatial reference. Taking LiDAR and security camera devices as an example, let the three-dimensional points in the LiDAR coordinate system be denoted as... The three-dimensional points in the coordinate system of the security camera device are Then the two are rotated through the extrinsic rotation matrix. Translation vector with extrinsic parameters satisfy: ; And Projected onto image pixel coordinates At that time, the lens intrinsic parameter matrix was used. The pinhole imaging model is derived from the projection normalization operator. Get pixel coordinates: ; To solve the extrinsic parameters , The value of is determined in this step, which involves constructing a point cloud during the acquisition phase. Image correspondence constraints are obtained by extracting several correspondences through target calibration, natural features, or artificial features. The extrinsic parameters are solved using the criterion of minimizing reprojection error, so that the following objective function is minimized: ; in, This is the projection normalization operator from 3D to 2D; Let be the pixel coordinates of the i-th corresponding point in the image; Let be the three-dimensional coordinates of the i-th corresponding point in the lidar coordinate system. For the spatial relationships between the millimeter-wave radar and the security camera device, and between the millimeter-wave radar and the lidar, this step can use a similar extrinsic parameter solution method to obtain the corresponding extrinsic parameters, and then write these extrinsic parameters along with the lens intrinsic parameters into the calibration parameter record.

[0025] Finally, this step writes quality tags to the acquired data and updates the quality tag records: for oblique images, laser point clouds, millimeter-wave point clouds and security video frames, quality tags such as acquisition integrity, timestamp validity, pose validity and measurement saturation or missing are calculated and written respectively; when there are missing timestamps, unusable poses, large-area holes in the point cloud or severely blurred images, the corresponding unusable tag or tag to be verified is written, so that the validity of multi-source data at the same time can be constrained according to the quality tags.

[0026] Step 2: Align and fuse oblique image 3D modeling with laser point cloud reference; In this step, this method constructs a 3D model of the park based on the oblique image sequence with unified timestamps and pose metadata obtained in Step 1; and uses the laser point cloud sequence obtained in Step 1 to form a laser point cloud reference, aligning the 3D model of the oblique image to the unified spatial reference system of the park, to obtain a geometrically consistent 3D scene skeleton result. Here, the 3D model of the oblique image refers to the 3D point cloud or triangular mesh and its texture representation obtained from the oblique image sequence through multi-view geometry; the laser point cloud reference refers to a high-precision 3D point set obtained by lidar scanning based on actual measured distances, used as a geometric reference constraint.

[0027] First, this method performs multi-view geometric solving on the oblique image sequence to generate an initial 3D structure. Specifically, image feature points with stable geometric discriminative properties are extracted within each oblique image. These image feature points refer to pixel positions that can be repeatedly located in the image plane and their local descriptive information. Subsequently, feature point matching relationships are established between oblique images of adjacent or overlapping viewpoints. The matching relationship refers to the corresponding pixel observation pairs of the same spatial entity on different oblique images.

[0028] Based on the above matching relationship, this method solves for the exterior orientation parameters of each tilted image and simultaneously estimates the three-dimensional coordinates of the observed spatial points to obtain the image sparse point cloud. The exterior orientation parameters are represented by a rotation matrix and a translation vector, used to describe the pose of the tilted camera coordinate system relative to the unified spatial reference system of the park. The image sparse point cloud is a set of three-dimensional points calculated from the matching relationship, used to describe the initial geometric skeleton of the park scene.

[0029] To ensure that the exterior orientation parameters and the sparse point cloud of the image converge globally, this method performs global joint optimization on all oblique image observations, with the minimum pixel reprojection error as the criterion. Let the pixel observation coordinates of the j-th spatial point in the i-th oblique image be... The three-dimensional coordinates of this spatial point in the unified spatial reference system of the park are: The exterior orientation parameters of the tilting camera device corresponding to the i-th tilted image are the rotation matrix. With translation vector The lens intrinsic parameter matrix is Then the projected pixel coordinates of the spatial point on the i-th tilted image are transformed from three-dimensional to two-dimensional by the projection operator. The following is given: ; in, These are the projected pixel coordinates predicted from the current parameters. This is a projection normalization operator that normalizes homogeneous coordinates to pixel coordinates. The pixel reprojection error is defined as follows: and The method aims to minimize the sum of squared global reprojection errors by finding the Euclidean distance. ; in, This is the set of indices for which there are valid pixel observations; The norm is 2. Through the above solution, a convergent set of exterior orientation parameters and a sparse point cloud of the image are obtained.

[0030] After obtaining the sparse point cloud from the image, this method further generates a dense point cloud from the image. The dense point cloud referred to here is a set of dense 3D points obtained by performing pixel-by-pixel or region-by-region depth estimation on the scene surface based on the sparse point cloud from the image, utilizing the parallax constraints of the oblique image sequence under multiple viewpoints. After generating the dense point cloud from the image, this method constructs an image triangular mesh from the dense point cloud using a triangular partitioning method. The image triangular mesh refers to a continuous surface model formed by connecting multiple triangular facets. Based on the projection relationship of each triangular facet onto the oblique image, the color texture of the oblique image is mapped onto the surface of the image triangular mesh, thereby obtaining a 3D model of the oblique image that includes geometry and texture.

[0031] Subsequently, this method constructs a laser point cloud reference and completes the reference alignment of the oblique image 3D model. Specifically, the laser point cloud sequence acquired in step one is transformed to a unified spatial reference system of the park according to its pose metadata, and multiple frames of laser point clouds are spatially merged to obtain a reference point set covering the main static structures of the park, which serves as the laser point cloud reference. To eliminate residual translation and rotation deviations between the oblique image 3D model and the laser point cloud reference, this method uses an iterative nearest-point registration algorithm to solve the rigid body alignment transformation. Let the 3D point set extracted from the oblique image 3D model be { The three-dimensional point set in the laser point cloud reference is The iterative nearest-point registration algorithm performs a process for each point in each iteration. Select its nearest neighbor in the laser point cloud reference. As corresponding points, solve for the rotation matrix. With translation vector Minimize the following objective function: ; in, It is a three-dimensional rotation matrix. Let be a three-dimensional translation vector, and c(k) be the nearest neighbor index function. After completing one solution, { }by , The nearest neighbor correspondence is updated and re-established until the objective function decreases below a preset threshold or the number of iterations reaches the upper limit. Through the above iterative nearest point registration algorithm, the 3D model of the tilted image is geometrically aligned to the laser point cloud reference, thus forming a geometrically consistent scene skeleton model under the unified spatial reference system of the park.

[0032] After completing the benchmark alignment, this method performs a geometric consistency check on the oblique image 3D model and writes the check result into a quality tag record. The geometric consistency check includes at least: calculating the nearest neighbor distance statistics from the oblique image 3D model point set to the laser point cloud benchmark point set under the unified spatial reference frame of the park, and determining whether the alignment residual meets the preset threshold based on this; when the alignment residual exceeds the threshold, the corresponding data object or corresponding spatial region is written into the check mark to ensure that the output scene skeleton model is verifiable and constrainable in terms of geometric accuracy.

[0033] Step 3: Parallel dynamic target removal through multimodal deep fusion and point cloud densification; In this step, under the unified spatial reference frame of the park, the method takes as input the calibration parameter records obtained in step one, which include the lens intrinsic parameter matrix and the extrinsic rotation matrix and extrinsic translation vector between each sensor, as well as the security video frame sequence, laser point cloud sequence and millimeter wave point cloud sequence with unified timestamps obtained in step one, and the geometrically consistent scene skeleton result obtained in step two as geometric constraints. It performs multimodal deep fusion of images and point clouds and point cloud densification, while removing the three-dimensional points corresponding to dynamic targets such as personnel and vehicles. The output is a dense static color point cloud set and its quality mark record that can be used to express three-dimensional security scenes.

[0034] First, this method uses a unified timestamp as an index to select security video frames and corresponding laser point cloud frames within the same unified timestamp or the same short time window. The laser point cloud frames are then transformed to the coordinate system of the security camera device corresponding to the security video frame using the extrinsic rotation matrix and extrinsic translation vector obtained in step one. Finally, the laser point cloud frames are projected onto the pixel plane of the security video frame according to the lens intrinsic parameter matrix to generate a sparse depth map.

[0035] It should be noted that the sparse depth map referred to here is a depth image on the pixel grid of a security video frame, in which depth values ​​are recorded only at the pixel positions hit by the laser point cloud projection, and the remaining pixel positions are empty; this sparse depth map serves as the geometric observation input for deep learning inference, rather than generating depth out of thin air.

[0036] Subsequently, this method performs multimodal depth fusion inference on the security video frame and the sparse depth map to obtain a dense depth map. The multimodal depth fusion inference referred to here is a learning-based mapping with a convolutional neural network as its core, which fuses the texture and edge information provided by the security video frame and the distance observation information provided by the sparse depth map at the feature layer, thereby filling in the missing depth on the pixel grid and enhancing the consistency of the depth boundary.

[0037] Furthermore, this step employs a two-branch structure to achieve this inference: one is the color branch, which takes security video frames as input and outputs a predicted depth map. ,in The first branch represents pixel coordinates; the second branch is the depth branch, which takes a sparse depth map as input and outputs a predicted depth map. Simultaneously, the network outputs two confidence weight maps corresponding to the pixels. and The confidence weight map is used to characterize the depth confidence of the corresponding branch at that pixel location. Based on the above output, this step generates a dense depth map using a weighted fusion method. : ; in, For dense depth maps in pixels The depth value at that location; Depth prediction for the color branch output; The depth prediction output for the depth branch; and The color branch and the depth branch are respectively located at the pixel level. Confidence weight at the point; To prevent extremely small positive numbers with a denominator of zero, this fusion method ensures that the numerical source of the dense depth map is always constrained by both image texture cues and sparse distance observations, rather than inferred solely from a single image source.

[0038] To ensure the controllability of the dense depth map in terms of geometric scale and boundary morphology, this step uses a supervised depth map constructed from the sparse depth map when employing supervised training or online self-calibration for multimodal depth fusion inference. As a verifiable distance constraint, the supervised depth is valid only at non-empty pixels in the sparse depth map.

[0039] This step constructs an effective pixel mask using the non-empty pixel locations of the sparse depth map, and only participates in the masking process at pixels where the effective pixel mask is true. The value of is summed with the loss function.

[0040] Next, this step uses the training objective function L, which is composed of both the depth consistency term and the depth gradient consistency term: ; in, For monitoring the depth-effective set of pixels; It is a norm; For the gradient operator of the pixel grid; and The non-negative weighting coefficients are used to balance numerical accuracy and boundary consistency. Through this objective function, the dense depth map is consistent with the supervised depth at pixels with real measurement data, and its gradient changes in boundary morphology are consistent with those of the supervised depth, thereby suppressing depth smearing and edge drift.

[0041] After obtaining the dense depth map, this step back-projects the dense depth map to generate a color point cloud frame, and then uniformly transforms it to the park's unified spatial reference frame. The color point cloud frame referred to here means that each 3D point contains both spatial coordinates and color information bound to the security video frame. Specifically, for any pixel coordinate... Construct homogeneous pixel vectors Using the lens intrinsic parameter matrix With dense depth value Calculate its three-dimensional points in the coordinate system of the security camera device. : ; The three-dimensional point is then transformed into the unified spatial reference system of the park through the pose transformation of the security camera device. At the same time, the security video frames are divided into pixels. The color value at that location is bound to This process forms color point cloud frames. By repeating the above process on multiple security video frames within the same unified timestamp or the same short time window, a multi-view color point cloud frame set can be formed and merged into a dense point set under the unified spatial reference system of the park.

[0042] It should be noted that, under the conditions that the clock offset estimate has been obtained based on the four-timestamp timing interaction and a unified timestamp has been generated, and the laser point cloud has been projected onto the security video frame to form a sparse depth map based on the extrinsic rotation matrix and extrinsic translation vector, and then a dense depth map has been obtained by fusing the color branch and the depth branch, an unconventional state of evidence splitting and caliber reversal may occur near the same static structure: the alignment relationship of the unified timestamp changes discontinuously between adjacent frames, causing instantaneous mismatch between the security video frame and the laser point cloud frame associated under the same unified timestamp; at the same time, the sparse depth map forms a break zone with alternating missing effective depth pixels near the structure boundary, causing the distance observation constraint of the depth branch to exhibit an unstable distribution at the boundary.

[0043] Consequently, the depth predictions from the color branch and the depth branch deviate significantly from each other in the neighborhood of the fracture zone. These deviations migrate along with the alignment and fracture zone location, and the dominant relationship flips within a short window. This results in ghosting and flaking surfaces appearing near the boundaries of what should be a single continuous surface in the fused dense depth map. Ghosting refers to two similar but non-overlapping depth layers of the same structure in the dense depth map, while flaking refers to floating depth layers with approximately zero thickness near the real surface. This distortion further manifests as spurious structural bands moving along structural boundaries in the back-projected color point cloud frame set, exhibiting positional drift and morphological reversal, leading to inconsistent geometric representations of the same static structure in continuous acquisitions.

[0044] Therefore, in this embodiment, after forming the set of colored point cloud frames, this step performs culling on the 3D points corresponding to dynamic targets and controlled suppression on apparent drift points caused by discontinuous changes in the unified timestamp alignment relationship and sparse depth map breaks, so that the output point set focuses on the static scene and avoids solidifying ghosting surfaces and thin sheets near the static structural boundaries. The dynamic targets referred to here are objects whose spatial positions change significantly over time within the acquisition time window, including people and vehicles; the apparent drift points referred to here are non-real motion representation points caused by instantaneous mismatch of the unified timestamp alignment relationship, alternating loss of effective pixels in the sparse depth map, or deviation migration of depth prediction between the color branch and the depth branch. These points may appear as false structural bands moving along structural boundaries in 3D space, but their cause is not the movement of real targets.

[0045] To distinguish between the two types of points mentioned above, this step involves comparing adjacent unified timestamps t and... Establish temporal correspondences between corresponding points in the color point cloud frames and calculate displacement velocity.

[0046] The temporal correspondence of corresponding points referred to in this step refers to the three-dimensional point at time t within the unified spatial reference system of the park. Let t+ be the query point. Searching for its nearest neighbor in a colored point cloud frame + And under the condition that the nearest neighbor distance is not greater than a preset spatial threshold and both frames of data quality are determined to be usable, + As The corresponding point; if the condition is not met, it is determined that there is no reliable correspondence between the point and the velocity quantity at that time step, and it will not participate in the velocity quantity calculation.

[0047] Let the coordinates of a three-dimensional point corresponding to the same spatial location at time t be... At any moment The coordinates are Then the displacement velocity quantity v is defined as follows: ; in, It is the time interval between two adjacent frames corresponding to the same timestamp, which is obtained by differentiating the same timestamp field.

[0048] Next, under the condition that dynamic targets are susceptible to discontinuous changes in alignment relationships based solely on a fixed velocity threshold, this step introduces three dynamic quantities directly related to scene distortion, and dynamically adjusts the velocity threshold accordingly. The first dynamic quantity is the alignment jump variable. The clock offset estimate is obtained by the time stamp synchronization calculation in step one four. The construct is defined as the absolute value of the offset difference between adjacent sampling periods: ; and with threshold Normalizing it yields the alignment instability coefficient. : ; in, The first is the alignment jump threshold, used to define the criteria for determining discontinuous changes in alignment relationships. The second dynamic quantity is the coverage compression quantity. It is constructed by the effective pixel ratio of the sparse depth map generated by projection within the structural boundary neighborhood. The structural boundary neighborhood refers to the dense depth map. Let the set of pixels centered at the location of a significant change in the depth gradient be denoted as . The sparse depth map at the pixel level When a valid depth observation exists at a given location, that pixel is called a valid pixel, and the set of valid pixels is denoted as . Based on this, the effective coverage rate of the boundary neighborhood is defined. With covering pressure : ; ; in, This indicates the number of elements in the set. The third dynamic variable is the branch conflict index. Its depth prediction is output by the color branch. Depth prediction with depth branch output The degree of deviation within the neighborhood of the structural boundary is defined as... Robust aggregation on: ; in, For median operators; To prevent extremely small positive numbers with a denominator of zero.

[0049] Based on the above three dynamic quantities, this step expands the original fixed velocity threshold into a velocity threshold that dynamically adjusts over time. This is to reduce the misleading influence of alignment instability, coverage breaks, and branch conflicts on dynamic target identification. Let the reference velocity threshold be... If the weighting coefficients are α, β, γ, then the following definition is made: ; Wherein, α, β, γ are non-negative coefficients used to characterize the adjustment strength of alignment instability, cover breakage, and branch conflict on the discrimination threshold.

[0050] Furthermore, while employing a dynamic velocity threshold, this step retains the pixel motion at the image level as a mutual verification constraint. The pixel displacement field between adjacent security video frames is defined as optical flow, which refers to the pixel displacement vector field of the same physical point in two image frames; for three-dimensional points... Pixel position corresponding to projection The optical flow amplitude at that pixel location is recorded as... And set the pixel motion threshold as .when and Furthermore, if the data quality marker corresponding to the point is determined to be usable, the point is marked as a dynamic target point; the set of points marked as dynamic target points is removed from the set of colored point cloud frames and does not participate in the merging of static point sets.

[0051] when However, if any of the following conditions occurs: Alignment instability coefficient Not less than the preset threshold or cover the amount of pressure Not less than the preset threshold or branch conflict index Not less than the preset threshold or pixel motion No more than If a point is found to be a solidified suppression point, it is marked as such. A solidified suppression point, as referred to here, is a point that does not participate in the static point set merging within the current short window, but is also not directly deleted as a dynamic target. It is used to suppress ghosting surfaces and thin surfaces caused by alignment jumps, break zones, and branch conflicts from being written into the static scene. For solidified suppression points, this step continuously observes their corresponding points within the scrolling window. , , To ensure stability, when the aforementioned dynamic quantities are all below their respective thresholds for at least M consecutive sampling periods, the solidification suppression flag is removed, and the point is allowed to participate in the static point set merging under the condition that its quality flag is deemed usable; if the point meets the dynamic target determination conditions during the observation period, it is converted into a dynamic target point and is removed.

[0052] in, To align the instability coefficient threshold, To cover the pressure threshold, The threshold value for the branch conflict index is M, where M is the number of consecutive stable sampling periods used to remove the solidification inhibition marker. The nearest neighbor distance threshold is set for millimeter-wave point cloud points to geometrically consistent scene skeleton results; all of the above thresholds are pre-set or calibrated parameters, and are written into the parameter configuration record for reproduction.

[0053] When gaps appear in the laser point cloud or the usability of the quality marker decreases under conditions such as rain, fog, strong reflection, or occlusion, this step transforms the millimeter-wave point cloud to the unified spatial reference frame of the park using the extrinsic rotation matrix and extrinsic translation vector from step one. After completing dynamic target removal and solidification suppression, it is merged with the static color point cloud set. Among these steps, the millimeter-wave point cloud is only used when its quality marker is determined to be usable and the nearest neighbor distance from its point to the geometrically consistent scene skeleton result from step two is not greater than a preset distance threshold. Under certain conditions, it participates in merging to suppress the formation of false structural bands at the static structural boundary by unstable echoes; simultaneously, when the corresponding branch conflict index of this region... Not less than At that time, only millimeter-wave points that meet the nearest neighbor distance constraint and do not fall into the neighborhood of the solidified suppression point are allowed to participate in the merging, so as to avoid solidifying noise points into thin sheets in the state of evidence splitting.

[0054] The solidification suppression point neighborhood is a spatial neighborhood constructed with the coordinates of the solidification suppression point in the unified spatial reference system of the park as the center and a preset spatial radius rp. When the Euclidean distance from the candidate millimeter-wave point to any solidification suppression point is less than rp, it is determined that the millimeter-wave point falls into the solidification suppression point neighborhood.

[0055] It should be noted that the tilted image described in step two is acquired by the tilted camera device, and the security video frame described in step three is output by the security camera device. The two are different acquisition subjects, and their poses and intrinsic parameters are written into the calibration parameter record and referenced separately in subsequent processing.

[0056] Step 4: Feature-guided point cloud registration and stitching, and solidification output of 3D security scene; First, this method performs spatial partitioning and overlap organization on the dense static color point cloud set. Specifically, the dense static color point cloud set under the unified spatial reference frame of the park is divided into several point cloud blocks according to spatial extent, and overlapping areas are retained between adjacent point cloud blocks. Here, a point cloud block refers to a subset of points with spatial boundaries, and the overlapping area refers to the area jointly covered by adjacent point cloud blocks in space, used to establish cross-block registration constraints. For each point cloud block, this method removes points marked as unusable based on quality labels, and performs necessary downsampling while preserving the geometric structure, so that the computational cost of subsequent feature calculation and matching is controlled and does not change the main shape of the scene structure.

[0057] Subsequently, this method constructs point cloud features for each point cloud block and uses these features to drive the extraction of cross-block correspondences. The point cloud features referred to here are fixed-length vectors obtained by encoding each 3D point and its neighborhood geometric relationships, used to characterize the local geometric structure of that point. To obtain discriminative point cloud features, this step employs a graph residual neural network to extract features from the point cloud blocks: the graph residual neural network is a neural network that uses a point cloud neighborhood graph as its computational carrier and residual connections as a stable training method; the neighborhood graph consists of an edge set formed by each point within the point cloud block and its nearest neighbors, thus representing the local geometric relationships of the point cloud as a graph structure. The graph residual neural network outputs a point cloud feature vector for each point. ,in The coordinates of this point in the unified spatial reference system of the park are: This is the point cloud feature vector for that point. Through this feature extraction, point cloud blocks at the same structural location within overlapping regions will exhibit similar point cloud feature vectors.

[0058] After obtaining the point cloud features, this method extracts the point correspondence between any pair of point cloud blocks with overlapping regions. To avoid the high computational cost caused by full pairwise comparisons, this step employs a sequential similarity detection algorithm for similarity screening: the sequential similarity detection algorithm compares the point cloud feature vectors sequentially, narrowing down the candidate set step by step, and outputs the candidate correspondence set with the highest similarity. Specifically, let the point cloud feature vector of the candidate points in block A be... The point cloud feature vector of the candidate points in block B is In this step, the similarity between the two is defined as cosine similarity. : ; in, For vector dot product, It is a norm 2. The sequential similarity detection algorithm outputs m pairs of corresponding points within overlapping regions based on similarity. ,in The coordinates of a 3D point in block A. The coordinates of the three-dimensional points in block B are given, and each pair of corresponding points satisfies the condition that the similarity of their point cloud feature vectors ranks high in the candidate set and is selected through consistency constraints.

[0059] In one embodiment, the sequential similarity detection algorithm takes the set of point cloud feature vectors in the overlapping region as input, firstly retrieves the initial candidate corresponding set by approximate nearest neighbor search for each candidate point, then sequentially filters from high to low cosine similarity, and introduces geometric consistency constraints in each round of filtering to eliminate inconsistent candidates, until the size of the candidate set converges to a preset upper limit or the similarity is lower than a preset threshold, thereby outputting a set of corresponding point pairs for solving rigid body splicing transformation.

[0060] After obtaining the set of corresponding point pairs, this step solves for the rigid body splicing transformation from block A to block B. The rigid body splicing transformation referred to here is a three-dimensional rigid body transformation composed of rotation matrices and translation vectors, used to eliminate residual rigid body deviations between the two blocks. Let the rigid body splicing transformation consist of a rotation matrix... With translation vector This means that this step solves the problem by minimizing the squared error of corresponding point pairs. , : ; in, It is a three-dimensional rotation matrix. It is a three-dimensional translation vector. and Let be the three-dimensional coordinates of the k-th pair of corresponding points. This is obtained through the solution. , This serves as the initial alignment result for aligning block A to block B.

[0061] After obtaining the initial alignment results, this step uses the iterative nearest-point registration algorithm to perform fine registration on the rigid body stitching transformation. The iterative nearest-point registration algorithm in the current... , Under the given conditions, for each point in block A, select the nearest neighbor in block B as the new corresponding point, and repeatedly solve the following objective function until convergence: ; in, For a point in block A, For block B and The nearest corresponding point, c(i), is the nearest neighbor correspondence index function. After fine registration converges, the residual distance statistics of block A and block B in the overlapping area meet the preset threshold constraint, and the residual distance statistics are written into the quality mark record as the geometric accuracy index of this stitching.

[0062] This step involves extracting features from overlapping point cloud blocks, performing sequential similarity detection, solving rigid body stitching transformations, and applying an iterative nearest-point registration algorithm for precise registration. Based on the stitching results, the point clouds of each block are unified under a single spatial reference frame within the same area, resulting in a continuous and consistent globally dense static color point cloud. For points repeatedly covered in multiple stitching operations, this step performs consistency merging based on their quality marker records and local residual distance indices, ensuring the output point set is geometrically continuous and color-stable.

[0063] Subsequently, this step solidifies the globally dense static color point cloud into a 3D security scene model. Specifically, this step estimates the normal vector of each point based on the globally dense static color point cloud. The normal vector refers to the unit vector representing the orientation of the local surface. After obtaining the normal vector, a surface reconstruction algorithm is used to generate a triangular mesh from the point cloud. The triangular mesh is a continuous surface representation formed by connecting multiple triangular facets. After the triangular mesh is generated, this step uses the lens intrinsic parameter matrix, extrinsic parameter rotation matrix, and extrinsic parameter translation vector from step one to project the visible texture of the security video frame in the unified spatial reference frame of the park onto the surface of the triangular mesh, forming a texture map. This allows the 3D security scene model to simultaneously possess geometric continuity and texture realism. For areas with occlusion, missing viewpoints, or inconsistent colors in the texture map, this step uses quality marker records as constraints, prioritizing security video frames with usable quality markers and small reprojection residuals for texture assignment to obtain a stable and consistent texture representation.

[0064] In one embodiment, the surface reconstruction algorithm employs either the Poisson surface reconstruction algorithm or the spherical pivot surface reconstruction algorithm to generate a triangular mesh output from a point cloud input constrained by normal vectors.

[0065] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0066] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.

[0067] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and inventive constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0068] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0069] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0070] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for reconstructing multi-source sensing 3D security scenes in digital twin parks, characterized in that, include: Collect multi-source sensing data and write it into the data source registration record, generate a unified timestamp, and construct a multi-source data sequence and quality label record; Construct a 3D model of the tilted image and align it with the laser point cloud reference to obtain a geometrically consistent scene skeleton result, and write it into the quality mark record; A set of colored point cloud frames is generated using the geometrically consistent scene skeleton result as a geometric constraint. A temporal correspondence of corresponding points is established between adjacent colored point cloud frames, and the displacement velocity is calculated. The velocity threshold is dynamically adjusted based on the alignment instability coefficient, coverage compression, and branch conflict index. The 3D points are classified into dynamic target points and solidified suppression points and processed in combination with pixel motion and quality mark records. When the laser point cloud has gaps or the availability of quality mark records decreases under rain, fog, strong reflection, or occlusion conditions, the millimeter-wave point cloud is transformed to the unified spatial reference system of the park based on the extrinsic rotation matrix and extrinsic translation vector and participates in merging. Only millimeter-wave points that are determined to be usable by the quality mark records and whose nearest neighbor distance from the millimeter-wave point to the geometrically consistent scene skeleton result is not greater than a preset distance threshold are allowed to participate in merging. When the branch conflict index is not less than a preset threshold, the millimeter-wave point is further restricted from falling into the neighborhood of solidified suppression points. A dense static colored point cloud set and quality mark records are output. Register, stitch, and merge dense static color point cloud sets to generate triangular meshes and project security video frame textures to output a 3D security scene model.

2. The method for reconstructing multi-source sensing three-dimensional security scenes for digital twin parks according to claim 1, characterized in that: For oblique images, laser point clouds, millimeter-wave point clouds, and security video frames, a data source registration record is first established for each data object and the source credentials and original carrier index are written. Then, multi-source sensing data are collected synchronously and the corresponding pose and lens intrinsic parameter matrix are recorded. Finally, a unified timestamp is generated by performing four-timestamp time synchronization interaction with each acquisition device using the time synchronization reference node.

3. The method for reconstructing multi-source sensing three-dimensional security scenes for digital twin parks according to claim 2, characterized in that: By constructing the correspondence constraints between point clouds and images and solving for the extrinsic rotation matrix and extrinsic translation vector with the minimum reprojection error as the criterion, a calibration parameter record is formed together with the lens intrinsic parameter matrix. For various types of data, the acquisition integrity, timestamp validity, pose validity, and measurement saturation or missing quality marker fields are calculated and written, and written into the quality marker record, resulting in a multi-source data sequence that can be expressed under the same spatiotemporal reference and can be used according to the constraints of the quality marker record.

4. The method for reconstructing multi-source sensing three-dimensional security scenes for digital twin parks according to claim 1, characterized in that: Based on the oblique image sequence with unified timestamp and pose metadata, we first extract repeatable image feature points in each oblique image and establish matching relationships between overlapping viewpoints. We then solve the exterior orientation parameters of each oblique image and back-calculate the sparse point cloud of the image. Finally, we perform global joint optimization on all observations with the minimum pixel reprojection error as the criterion. Subsequently, dense point clouds of the image are generated using multi-view parallax constraints, and triangular meshes and texture mappings are constructed based on these to obtain a 3D model of the tilted image. At the same time, the laser point cloud sequence is transformed to a unified spatial reference system of the park according to pose metadata and multiple frames are merged to form a laser point cloud reference. Rigid body alignment transformation is solved by iterative nearest point registration to eliminate residual translation and rotation deviations, so that the 3D model of the tilted image is aligned with the laser point cloud reference and a geometrically consistent 3D scene skeleton result is output. Finally, the nearest neighbor distance statistics between the point set of the 3D model of the tilted image and the point set of the laser point cloud reference are calculated to verify geometric consistency and written into the quality mark record. When the alignment residual exceeds a preset threshold, a check mark is written to constrain the output accuracy.

5. The method for reconstructing a multi-source sensing 3D security scene for a digital twin park according to claim 1, characterized in that: Using calibration parameters and security video frame sequences, laser point cloud sequences, and millimeter-wave point cloud sequences with unified timestamps as input, and geometrically consistent scene skeleton results as geometric constraints, security video frames and corresponding laser point cloud frames are first selected according to unified timestamps or short time windows. The laser point cloud frames are transformed to the coordinate system of the security camera device through extrinsic rotation matrix and extrinsic translation vector, and projected onto the pixel plane according to the lens intrinsic parameter matrix to generate a sparse depth map. Then, the texture edge information of the security video frames and the distance observation information of the sparse depth map are used as constraints to perform multimodal depth fusion inference to obtain a dense depth map. The dense depth map is then back-projected to generate color point cloud frames with bound colors, and uniformly transformed to a unified spatial reference system of the park. Multi-view color point cloud frames within the same unified timestamp or short time window are merged to form a dense static color point cloud set.

6. The method for reconstructing a multi-source sensing 3D security scene for a digital twin park according to claim 1, characterized in that: The alignment instability coefficient is constructed from the alignment jump variable; The coverage compression is constructed from the effective coverage rate of the boundary neighborhood; The branch conflict index is constructed by the degree of deviation between the color branch prediction and the depth branch prediction in the boundary neighborhood; The corresponding time sequence of the same point is as follows: under the unified spatial reference system of the park, the three-dimensional point at time t is used as the query point, the nearest neighbor point is searched in the color point cloud frame at time t+Δt, and the corresponding relationship is established when the nearest neighbor distance is not greater than the preset spatial threshold and the quality mark records of the two frames are determined to be available. Where Δt is the time interval between two adjacent frames corresponding to the same timestamp.

7. The method for reconstructing a multi-source sensing 3D security scene for a digital twin park according to claim 6, characterized in that: Points that meet the conditions of displacement velocity, pixel motion, and usable quality mark records are marked as dynamic target points and removed from the color point cloud frame set. Points whose displacement velocity exceeds the threshold but are in a state of alignment instability, coverage breakage, or high risk of branch conflict, or whose pixel motion is insufficient for mutual verification, are marked as solidification suppression points. They are temporarily not included in the merging of dense static color point cloud sets within the scrolling window to suppress the solidification of ghosting surfaces and thin surfaces. When the alignment instability coefficient, coverage compression amount, and branch conflict index are lower than the corresponding preset thresholds for no less than M consecutive sampling periods, the solidification suppression mark is removed, and they are allowed to participate in the merging of dense static color point cloud sets if the quality mark record is determined to be usable. If they meet the dynamic target point determination conditions during the observation period, they are removed. Where M is the preset number of continuous stable sampling periods, and the corresponding preset thresholds include the alignment instability coefficient threshold, the coverage compression threshold, and the branch conflict index threshold; The neighborhood of the solidification inhibition point is a spatial neighborhood constructed with the coordinates of the solidification inhibition point in the unified spatial reference system of the park as the center and with a preset spatial radius.

8. The method for reconstructing a multi-source sensing 3D security scene for a digital twin park according to claim 1, characterized in that: The dense static color point cloud set is spatially divided into point cloud blocks with overlapping regions. Unusable points are removed based on quality marker records, and necessary downsampling is performed. Then, point cloud features are extracted from the blocks, and cross-block corresponding point pairs are obtained by filtering within the overlapping regions. Rigid body stitching transformation between blocks is solved based on cross-block corresponding point pairs, and fine registration is completed by combining iterative nearest point registration. All blocks are stitched repeatedly, and overlapping covered points are merged for consistency based on quality marker records and residual distance indicators. A triangular mesh is generated from a globally dense static color point cloud set. The security video frame texture is projected onto the mesh surface using the lens intrinsic parameter matrix, extrinsic parameter rotation matrix, and extrinsic parameter translation vector, and then solidified to output a 3D security scene model.