Video point cloud data fusion and surface reconstruction method based on digital twin scene

By constructing a time-continuous trajectory model and fusing multi-dimensional feature mapping data, the spatiotemporal coupling problem of video point cloud data under unsteady motion was solved, achieving high-precision digital twin scene reconstruction, eliminating motion distortion and texture misalignment, and improving reconstruction quality.

CN121482333BActive Publication Date: 2026-04-07BEIJING ZHIHUI YUNZHOU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-09
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In non-steady motion conditions, the spatiotemporal coupling problem between the continuous data stream characteristics of lidar and the rolling shutter effect of camera leads to motion distortion after video point cloud data fusion, resulting in problems such as geometric layering, texture mapping misalignment, and surface topology tearing in the reconstructed digital twin model.

Method used

By synchronously acquiring data from cameras, lidar, and inertial measurement units, a time-continuous trajectory model is constructed using B-spline basis functions. The pose transformation matrix of video images and laser point cloud data is solved. Combined with perspective projection and voxelized truncated symbolic distance fields, multi-dimensional feature mapping data is generated and weighted fusion is performed to generate a digital twin 3D surface.

Benefits of technology

It accurately inverts the instantaneous pose of the mobile acquisition terminal, decouples the spatiotemporal coupling problem, eliminates motion distortion, improves the geometric accuracy and visual realism of digital twin scene reconstruction, avoids geometric layering and texture misalignment, and improves reconstruction quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482333B_ABST
    Figure CN121482333B_ABST
Patent Text Reader

Abstract

The application provides a video point cloud data fusion and curved surface reconstruction method based on a digital twinborn scene, and belongs to the technical field of video processing. The application synchronously collects video, laser point cloud and inertial data. Then, the inertial data is mapped to a Lie group manifold space, and a continuous trajectory model is constructed by using B-spline fitting. Based on the model, a first pose transformation matrix of each row of pixels at an exposure instant and a second pose transformation matrix of each laser point are accurately solved. Then, the point cloud is projected to a global coordinate system by using the second matrix, and combined with the first matrix, an internal parameter and a perspective principle, multi-dimensional feature mapping data of space-time registration is generated. Finally, the data is input into a voxelized truncated signed distance field, distance values and color weights are synchronously updated, and a high-fidelity digital twinborn three-dimensional curved surface is generated by extracting a zero level set isosurface. The application can improve modeling accuracy and visual effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of video processing technology, and in particular relates to a method for video point cloud data fusion and surface reconstruction based on digital twin scenes. Background Technology

[0002] With the rapid development of digital twin technology, the real-time fusion of texture information from video images with geometric information from laser point clouds to reconstruct 3D surfaces has become a key technology for building high-fidelity digital twin scenes. This technology can simultaneously restore the precise spatial structure and realistic surface texture of the physical world, and has broad application prospects in fields such as smart city construction, autonomous driving simulation testing, and remote inspection of industrial facilities.

[0003] In existing technologies, visual or laser odometry based on discrete keyframes is generally used for pose estimation. This assumes that the sensor is stationary or moving at a constant speed during the single frame acquisition time. The video frame is projected onto the corresponding point cloud frame through external calibration parameters to achieve data fusion. Then, the fused data is spatially stitched and surface generated through a meshing algorithm based on discrete point clouds to construct a three-dimensional model.

[0004] However, when the data acquisition device is in unsteady motion conditions such as handheld or drone-borne operation, existing technologies struggle to address the spatiotemporal coupling between the continuous data stream characteristics of LiDAR and the rolling shutter effect of the camera. This results in motion distortion in the fused data that is difficult to eliminate, leading to geometric layering, texture mapping misalignment, and surface topological tearing in the reconstructed digital twin model. Therefore, existing technologies suffer from poor modeling accuracy and visual effects. Summary of the Invention

[0005] The purpose of this application is to provide a method, system, electronic device and storage medium for video point cloud data fusion and surface reconstruction based on digital twin scenes, so as to solve the problems of poor modeling accuracy and visual effect in the prior art.

[0006] To address the aforementioned technical problems, in a first aspect, this application provides a method for video point cloud data fusion and surface reconstruction based on a digital twin scene, comprising:

[0007] Simultaneously acquire video image sequences and laser point cloud data of the target scene collected by a mobile acquisition terminal integrating a camera, lidar and inertial measurement unit during continuous movement, as well as inertial measurement data of the mobile acquisition terminal;

[0008] Inertial measurement data is mapped to the Lie group manifold space, and the motion trajectory of the mobile acquisition terminal is fitted with a high-order continuous model using B-spline basis functions to construct a time-continuous continuous trajectory model.

[0009] The first pose transformation matrix at the exposure time of each row of pixels in the video image sequence and the second pose transformation matrix at the emission time of each laser point in the laser point cloud data are calculated based on the continuous trajectory model.

[0010] The laser point cloud data is projected from the radar coordinate system to the digital twin global coordinate system based on the second pose transformation matrix. Based on the first pose transformation matrix and the camera's intrinsic parameters, the projection mapping relationship between the video image sequence and the laser point cloud data after transformation to the digital twin global coordinate system is established using the perspective projection principle. This generates multidimensional feature mapping data, which includes spatiotemporal registration geometric coordinate information and texture attribute information.

[0011] Multidimensional feature mapping data is used as the update source input to the voxelized truncated symbolic distance field. The symbolic distance values ​​in the voxel grid are updated using geometric coordinate information and the color weight values ​​in the voxel grid are updated using texture attribute information. The zero level set isosurface is extracted from the updated truncated symbolic distance field to generate a digital twin 3D surface.

[0012] In one feasible implementation, the method further includes:

[0013] The inertial measurement data is processed by time-domain differentiation to calculate the time derivative magnitudes of the acceleration and angular velocity vectors, and a scalar sequence of motion intensity corresponding to different time points is generated based on the time derivative magnitudes.

[0014] In the process of mapping inertial measurement data to the Lie group manifold space and using B-spline basis functions to perform high-order continuous fitting on the motion trajectory of the mobile acquisition terminal to construct a time-continuous continuous trajectory model, the time node interval of the B-spline basis functions is set to the first preset interval during the time period when the value of the motion intensity scalar sequence is greater than the preset dynamic threshold.

[0015] During the time period when the value of the scalar sequence of motion intensity is less than or equal to the dynamic threshold, the time node interval of the B-spline basis function is set to the second preset interval, wherein the value of the first preset interval is less than the value of the second preset interval.

[0016] In one feasible implementation, inertial measurement data is mapped to a Lie group manifold space, and the motion trajectory of the mobile acquisition terminal is fitted with a high-order continuous model using B-spline basis functions to construct a time-continuous trajectory model, including:

[0017] The angular velocity and acceleration components in the inertial measurement data are processed by discrete integration in the time dimension and converted into pose increment constraint information representing rotation and translation attributes in the Lie group manifold space.

[0018] Multiple trajectory control points, including spatial attitude parameters and translation parameters, are set in the Lie group manifold space. The corresponding time series nodes are determined according to the sampling period of the mobile acquisition terminal. Multiple basis functions are used to perform weight allocation and weighted accumulation processing on the trajectory control points in the time axis direction to obtain the motion fitting curve representing the initial motion state.

[0019] Deviation convergence calculation is performed on the pose increment constraint information and motion fitting curve to determine the final pose parameters of the trajectory control points. The trajectory control points, including the final pose parameters, and basis functions are then parameterized to obtain a continuous trajectory model.

[0020] In one feasible implementation, the calculation of the first pose transformation matrix at the exposure time of each row of pixels in the video image sequence, and the second pose transformation matrix at the emission time of each laser point in the laser point cloud data, based on a continuous trajectory model, includes:

[0021] The instantaneous exposure time of each row of pixels in the video image sequence is determined based on the line scan rate of the camera, and the instantaneous pulse time of each laser point in the laser point cloud data is determined based on the pulse frequency of the lidar.

[0022] The instantaneous exposure time of the row and the instantaneous pulse time of the point are imported into the continuous trajectory model to determine the multiple trajectory control points adjacent to each row pixel and each laser point in the time axis direction. The contribution weight distribution of the multiple adjacent trajectory control points in the instantaneous exposure time of the row and the instantaneous pulse time is calculated using the basis function.

[0023] By using the contribution weight distribution to smooth and superimpose multiple trajectory control points including the final pose parameters, the first pose transformation matrix representing the spatial position and orientation of each row of pixels at the exposure time, and the second pose transformation matrix representing the spatial position and orientation of each laser point at the emission time are calculated.

[0024] In one feasible implementation, the laser point cloud data is projected from the radar coordinate system to the digital twin global coordinate system according to the second pose transformation matrix. Based on the first pose transformation matrix and the camera's intrinsic parameters, a projection mapping relationship is established between the video image sequence and the laser point cloud data transformed to the digital twin global coordinate system using the principle of perspective projection. This generates multi-dimensional feature mapping data, which includes spatiotemporal registration geometric coordinate information and texture attribute information, including:

[0025] The second pose transformation matrix is ​​used to perform spatial coordinate transformation on the laser point cloud data, mapping the local three-dimensional coordinate position of each laser point in the radar coordinate system to spatial positioning information in the digital twin global coordinate system.

[0026] Based on the first pose transformation matrix, the spatial positioning information is processed by inverse viewpoint transformation to calculate the relative spatial position parameters of each laser point in the camera's observation coordinate system.

[0027] By using the camera's intrinsic parameters to perform perspective projection calculations on the relative spatial point parameters, the three-dimensional points in space are projected onto a two-dimensional imaging plane to obtain the pixel coordinate index of each laser point in the video image sequence.

[0028] Color component data is extracted from video image sequences based on pixel coordinate indexes, and then aggregated with spatial positioning information as texture attribute information to generate multidimensional feature mapping data.

[0029] In one feasible implementation, the method further includes, before extracting color component data from the video image sequence based on pixel coordinate indices:

[0030] Extract the numerical components of the relative spatial point position parameters in the direction perpendicular to the camera imaging plane as the point cloud measurement depth;

[0031] Detect whether the pixel coordinate index exceeds the preset pixel boundary range of the video image sequence;

[0032] If the point cloud measurement depth is within the effective detection range of the camera and the pixel coordinate index does not exceed the preset pixel boundary range, the mapping relationship between the corresponding spatial positioning information and the pixel coordinate index is retained; otherwise, the corresponding spatial positioning information is marked as a non-visible point to determine the effective mapping point for participating in color component data extraction.

[0033] In one feasible implementation, multidimensional feature mapping data is used as an update source input to a voxelized truncated symbolic distance field. The symbolic distance values ​​within the voxel mesh are updated using geometric coordinate information, and the color weight values ​​within the voxel mesh are updated using texture attribute information. Finally, a zero-level set isosurface is extracted from the updated truncated symbolic distance field to generate a digital twin 3D surface, including:

[0034] The voxel position of the geometric coordinate information in space is determined based on the global coordinate system of the digital twin, and the target voxel mesh within the preset cutoff distance range of the voxel position is determined.

[0035] Calculate the projected distance from the center point of the target voxel mesh to the geometric coordinate information, and perform a weighted fusion process with the projected distance and the pre-stored distance value in the target voxel mesh to obtain the updated symbolic distance value;

[0036] Perform a spatiotemporal co-mapping between the color values ​​in the texture attribute information and the updated symbolic distance values. Accumulate and calculate the color components within the target voxel mesh according to the weight ratio used when updating the symbolic distance values ​​to obtain the updated color weight values.

[0037] In a truncated symbolic distance field composed of multiple target voxel grids, the zero-crossing spatial locations where the symbolic distance values ​​between two adjacent target voxel grids alternate between positive and negative are retrieved, and the zero-level set isosurface is extracted by performing triangular patch fitting on the zero-crossing spatial locations.

[0038] Secondly, this application provides a video point cloud data fusion and surface reconstruction system based on a digital twin scene, including:

[0039] The acquisition module is used to simultaneously acquire video image sequences and laser point cloud data of the target scene collected by the mobile acquisition terminal, which integrates a camera, lidar and inertial measurement unit, during continuous movement, as well as the inertial measurement data of the mobile acquisition terminal.

[0040] The module is used to map inertial measurement data to the Lie group manifold space and use B-spline basis functions to perform high-order continuous fitting on the motion trajectory of the mobile acquisition terminal to construct a time-continuous continuous trajectory model.

[0041] The calculation module is used to calculate the first pose transformation matrix at the exposure time of each row of pixels in the video image sequence based on the continuous trajectory model, and the second pose transformation matrix at the emission time of each laser point in the laser point cloud data.

[0042] The generation module is used to project the laser point cloud data from the radar coordinate system to the digital twin global coordinate system according to the second pose transformation matrix, and to establish the projection mapping relationship between the video image sequence and the laser point cloud data after transformation to the digital twin global coordinate system based on the first pose transformation matrix and the camera's intrinsic parameters, using the perspective projection principle, and to generate multi-dimensional feature mapping data, which includes spatiotemporal registration geometric coordinate information and texture attribute information.

[0043] The generation module is also used to input multidimensional feature mapping data as an update source into the voxelized truncated symbolic distance field, update the symbolic distance value in the voxel grid using geometric coordinate information and update the color weight value in the voxel grid using texture attribute information, and extract the zero level set isosurface from the updated truncated symbolic distance field to generate a digital twin 3D surface.

[0044] Thirdly, this application provides an electronic device, comprising:

[0045] Memory, used to store computer programs;

[0046] A processor is used to execute computer programs to implement the steps of the video point cloud data fusion and surface reconstruction method based on digital twin scenes as described in the first aspect above.

[0047] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps of the video point cloud data fusion and surface reconstruction method based on a digital twin scene as described in the first aspect above.

[0048] The video point cloud data fusion and surface reconstruction method based on digital twin scenes provided in this application maps inertial measurement data to a Lie group manifold space and uses B-spline basis functions to construct a time-continuous trajectory model. This overcomes the temporal resolution limitations of existing technologies based on the discrete frame assumption, accurately inverting the instantaneous pose of each row of pixels exposure and each laser point emission moment during the non-steady-state motion of the mobile acquisition terminal. This decouples the spatiotemporal coupling problem between the continuous data stream characteristics of the lidar and the rolling shutter effect of the camera, eliminating motion distortion caused by high-frequency jitter or rapid movement of the equipment at the source. Furthermore, combined with the voxelization update of the truncated symbolic distance field and the zero level set extraction mechanism, the method uses spatiotemporally precisely registered multidimensional feature mapping data for weighted fusion. This smooths geometric noise while ensuring surface continuity, avoiding geometric layering, texture misalignment, and surface topological tearing in the reconstructed model, thus improving the geometric accuracy and visual realism of digital twin scene reconstruction under mobile non-steady-state acquisition conditions.

[0049] Furthermore, by performing time-domain differential processing on the inertial measurement data, the dynamic abrupt changes of the mobile acquisition terminal under unsteady motion conditions can be captured in real time. During periods of violent shaking or variable speed movement of the equipment, the time node interval of the B-spline basis function is adaptively reduced to improve the fitting degree of freedom and instantaneous expression accuracy of the continuous trajectory model for high-frequency motion details; while during the stable motion phase, the node interval is increased accordingly to effectively suppress computational redundancy while ensuring the global smoothness of the trajectory. This non-uniform node distribution strategy based on changes in motion intensity enables the constructed trajectory model to accurately reproduce the microscopic motion trajectory under complex working conditions, fundamentally eliminating the spatiotemporal coupling deviation between line-by-line exposure of the rolling shutter and point-by-point emission of laser pulses. Therefore, this application solves the problem of motion distortion and geometric tearing that is difficult to eliminate in the process of mobile modeling, and improves the texture mapping accuracy and high-fidelity reconstruction quality of digital twin 3D surfaces. Attached Figure Description

[0050] To more clearly illustrate the technical solutions of the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 A flowchart illustrating a method for video point cloud data fusion and surface reconstruction based on a digital twin scene, provided in an embodiment of this application;

[0052] Figure 2 A schematic diagram illustrating the generation of a pose transformation matrix provided in an embodiment of this application;

[0053] Figure 3 A schematic diagram illustrating a truncated symbol distance field update and surface reconstruction provided for an embodiment of this application;

[0054] Figure 4 A schematic diagram of the structure of a video point cloud data fusion and surface reconstruction system based on a digital twin scene provided in this application embodiment;

[0055] Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in one embodiment of this application. Detailed Implementation

[0056] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0057] To address the problems of existing technologies, embodiments of this application provide a method, apparatus, device, computer storage medium, and computer program product for video point cloud data fusion and surface reconstruction based on a digital twin scene. The method for video point cloud data fusion and surface reconstruction based on a digital twin scene provided in this application embodiment will be described below first.

[0058] Figure 1 The illustration shows a flowchart of a video point cloud data fusion and surface reconstruction method based on a digital twin scene provided in one embodiment of this application.

[0059] like Figure 1 As shown, the video point cloud data fusion and surface reconstruction method based on digital twin scenarios includes:

[0060] S101: Simultaneously acquire video image sequences and laser point cloud data of the target scene collected by a mobile acquisition terminal integrating a camera, lidar, and inertial measurement unit during continuous movement, as well as the inertial measurement data of the mobile acquisition terminal.

[0061] A mobile data acquisition terminal is a hardware device that integrates a visual capture unit, a laser detection unit, and an inertial sensing unit, and supports flexible movement within physical space. The target scenario refers to spatial environments such as industrial computer rooms, production workshops, or power facilities that require 3D digital reconstruction.

[0062] A video image sequence refers to a set of color pixel frames captured by a camera along a timeline, represented as... It is used to record the texture and color of environmental surfaces. Laser point cloud data refers to a large-scale set of spatial geometric points formed by active detection by lidar. Each laser point is represented as a vector including its three-dimensional position and reflection intensity. Inertial measurement data refers to the sequence of high-frequency motion parameters acquired by an inertial measurement unit (IMU), which can be represented as a set of six-dimensional motion vectors including acceleration and angular velocity components. .

[0063] In practical implementation, the mobile data acquisition terminal moves continuously along irregular paths within the target scene, activating its internal hardware clock synchronization mechanism to ensure the synchronous acquisition of various data types. Taking the digital twin modeling process of a power room as an example, the camera... Frequency of acquiring video image sequences Each frame The resolution is Color images are captured to show the colors of the transformer casing. Simultaneously, the lidar acquires laser point cloud data describing the device's shape. Assuming each laser point Represented as a vector This represents the location and reflection intensity of the device's surface in space. Meanwhile, the inertial measurement unit (IMU) uses... High-frequency sampling records the jitter and speed changes during movement, generating inertial measurement data. Assuming that the instantaneous motion vector is... Manifestation The triaxial acceleration and rotation rate were recorded. The initial exposure time of each frame was then extracted. Emission bias time of each laser point and the sampling time of inertial data Using a unified time reference to sequence video images Laser point cloud data and inertial measurement data Execution time alignment association.

[0064] S102: Map the inertial measurement data to the Lie group manifold space, and use B-spline basis functions to perform high-order continuous fitting on the motion trajectory of the mobile acquisition terminal to construct a time-continuous continuous trajectory model.

[0065] A Lie group manifold space is a nonlinear mathematical space used to describe the rotational and translational states of a mobile data acquisition terminal in three-dimensional space. It can be a special Euclidean group. Its advantage lies in ensuring the orthogonality of the rotation matrix and avoiding the singularity problem caused by Euler angle representation. B-spline basis functions refer to a set of local support basis functions defined on the time node vector, which can provide a smooth weight distribution according to the set curve order.

[0066] In the detailed implementation of the plan, the first step is to extract the acquired inertial measurement data. This includes triaxial acceleration vectors With the three-axis angular velocity vector Using the discrete integral algorithm to... and Perform pre-integration processing to obtain the initial pose sequence corresponding to the inertial sampling time. Each of them These are all transformation matrices in the manifold space of a Lie group. Next, a node sequence is set on the time axis of the mobile acquisition terminal's motion. And initialize a set of trajectory control points in the Lie group space. Higher-order continuous fitting refers to determining the optimal control point parameters through nonlinear optimization techniques, so that the generated motion trajectory satisfies the continuity of second-order or higher-order derivatives in terms of position, velocity, and acceleration, thereby eliminating motion jumps caused by discrete sampling.

[0067] Then, the B-spline basis functions are used to determine the trajectory control points. A time-based weighted accumulation process is performed to construct the initial trajectory. To improve fitting accuracy, a nonlinear optimization algorithm is used to perform iterative bias calculation, such as the Levenberg-Marquardt method, based on the initial pose sequence. For the observation target, the trajectory control points The pose parameters are iteratively corrected until the generated trajectory pose matches the observed pose. The Lie group manifold residuals converge to a preset range. Finally, the optimized control point parameters and basis functions are parameterized to obtain a time-continuous trajectory model. A continuous trajectory model refers to a trajectory model whose time is... The parameterized pose function with independent variables can output the global pose of the mobile acquisition terminal at any instant within the acquisition period.

[0068] S103: Based on the continuous trajectory model, calculate the first pose transformation matrix at the exposure time of each row of pixels in the video image sequence, and the second pose transformation matrix at the emission time of each laser point in the laser point cloud data.

[0069] The first pose transformation matrix is ​​a homogeneous matrix describing the instantaneous spatial position and rotation angle of the camera's optical center in the global coordinate system over time, denoted as: The second pose transformation matrix includes the translation and attitude attributes of the lidar scanning center at the pulse emission moment, denoted as: .

[0070] First, the line scan period of a single row of pixels is obtained through the camera's hardware parameters. Then, combined with the start trigger time of the video frame, the line exposure time of each row of pixels is calculated. ,in The frame start time, For row index, The scanning interval is defined as the row scan interval. Simultaneously, the point emission time of each point in the laser point cloud data is determined using the point clock frequency of the lidar. Subsequently, and These are then incorporated into the time-continuous trajectory model established in the previous steps. Within the Lie group manifold space, this model is used to retrieve the basis function support intervals where the target time falls, and to extract multiple trajectory control points that function within these intervals. The weighting coefficients are calculated at the current time using the basis functions. For these trajectory control points Perform weighted interpolation to calculate the camera's position. Moment And lidar in Moment .

[0071] S104: Project the laser point cloud data from the radar coordinate system to the digital twin global coordinate system according to the second pose transformation matrix. Based on the first pose transformation matrix and the camera's intrinsic parameters, establish the projection mapping relationship between the video image sequence and the laser point cloud data after transformation to the digital twin global coordinate system using the perspective projection principle. Generate multidimensional feature mapping data, which includes spatiotemporal registration geometric coordinate information and texture attribute information.

[0072] First, retrieve the laser point cloud data. laser point in The second pose transformation matrix corresponding to the laser point at the emission time is used. , laser point The transformation is from the radar local coordinate system to the digital twin global coordinate system. The digital twin global coordinate system refers to the world reference coordinate system established in the digital twin space to unify all sensor data and virtual entities. Through this matrix operation, global spatial points can be obtained as geometric coordinate information for spatiotemporal registration. The geometric coordinate information for spatiotemporal registration refers to the three-dimensional spatial positions determined after motion distortion elimination processing and under a unified time reference.

[0073] Next, in order to establish global spatial points With video image sequence The connection between them requires establishing a projection mapping relationship. The projection mapping relationship refers to the mathematical transformation logic between spatial three-dimensional points and video pixels, established based on the principles of geometric optics. At this point, the first pose transformation matrix at the corresponding moment is obtained. and utilize The inverse matrix will point to the global space. The coordinates are transformed to the camera's observation coordinate system to obtain the relative spatial point parameters. Then, the principles of perspective projection are used in conjunction with the camera's intrinsic parameters. The relative spatial position parameters are projected onto the camera's imaging plane. Camera intrinsic parameters, including focal length, optical center coordinates, and lens distortion coefficients, define the geometric projection characteristics of the camera's imaging plane. Through this projection calculation, the pixel coordinate index corresponding to each laser point in the video image frame is determined. .

[0074] Finally, based on pixel coordinate indexing From video image sequences Color component data is extracted from the corresponding pixel rows. The color component data is defined as texture attribute information, i.e., the color values ​​from the video image frame corresponding to spatial points. The texture attribute information is then compared with the global spatial points. The geometric coordinate information is paired and aggregated at the data field level to generate multidimensional feature mapping data. Multidimensional feature mapping data refers to a composite dataset that associates and encapsulates three-dimensional geometric position components and color and texture components at the field level.

[0075] S105: The multidimensional feature mapping data is used as the update source input to the voxelized truncated symbolic distance field. The symbolic distance value in the voxel grid is updated using geometric coordinate information and the color weight value in the voxel grid is updated using texture attribute information. The zero level set isosurface is extracted from the updated truncated symbolic distance field to generate a digital twin 3D surface.

[0076] First, the generated multidimensional feature map data is used as the update source and input into the voxelized truncated symbolic distance field (TSDF). A truncated symbolic distance field (TSDF) is a numerical field that divides three-dimensional space into a regular cubic grid. Each grid cell, or voxel, stores the weighted average distance to the nearest object surface, and this distance value is constrained within a preset truncation interval. Using the geometric coordinate information in the multidimensional feature map data, the specific voxel position in space is located, and the target voxel grid within the preset truncation distance range of that position is determined. The projected distance from the center point of the target voxel grid to the geometric coordinate information is calculated to obtain the instantaneous symbolic distance value. The symbolic distance value refers to the positive and negative distances from the voxel center to the reconstructed surface; its magnitude represents the distance, and the positive or negative sign distinguishes whether the voxel is located outside or inside the object.

[0077] Next, a weighted fusion process is performed between this instantaneous value and the pre-stored distance values ​​within the voxel grid to obtain the updated symbolic distance value. Simultaneously, the color weight values ​​within the voxel grid are updated using texture attribute information from the multidimensional feature map data. Texture attribute information refers to the color values ​​from video image frames corresponding to spatial points, typically represented as three-dimensional color vectors. Color weight values ​​are numerical values ​​used to represent the reliability of voxel colors or the cumulative number of observations. A spatiotemporal co-mapping is performed between the color values ​​in the texture attribute information and the updated symbolic distance values. The color components within the target voxel mesh are then cumulatively calculated according to the weight ratio used when updating the symbolic distance values.

[0078] After completing the cumulative update of all data, the zero-level set isosurface is extracted from the updated truncated symbolic distance field. The zero-level set isosurface is a continuous spatial set consisting of all points in the symbolic distance field with a symbolic distance value of zero, which directly corresponds to the real geometric surface of the observed object in a physical sense. By performing triangular patch fitting on the zero-intersection spatial locations where the symbolic distance values ​​alternate between positive and negative, a digital twin 3D surface is finally generated.

[0079] This embodiment maps inertial measurement data to a Lie group manifold space and uses B-spline basis functions to construct a time-continuous trajectory model, overcoming the temporal resolution limitations of existing technologies based on the discrete frame assumption. It accurately inverts the instantaneous pose of each row of pixels exposure and each laser point emission moment during the non-steady-state motion of the mobile acquisition terminal, thereby decoupling the spatiotemporal coupling problem between the continuous data stream characteristics of the lidar and the rolling shutter effect of the camera, eliminating motion distortion caused by high-frequency jitter or rapid movement of the equipment from the source. Furthermore, it combines the voxelization update of the truncated symbolic distance field and the zero level set extraction mechanism, and uses the spatiotemporally precisely registered multidimensional feature mapping data for weighted fusion, which smooths geometric noise while ensuring surface continuity, avoiding geometric layering, texture misalignment, and surface topological tearing in the reconstructed model, and improving the geometric accuracy and visual realism of digital twin scene reconstruction under mobile non-steady-state acquisition conditions.

[0080] In one feasible implementation, the method further includes:

[0081] The inertial measurement data is processed by time-domain differentiation to calculate the time derivative magnitudes of the acceleration and angular velocity vectors, and a scalar sequence of motion intensity corresponding to different time points is generated based on the time derivative magnitudes.

[0082] A scalar sequence of motion intensity refers to a set of values ​​that evolves over time, represented as: Each element Both represent the dynamic disturbance intensity of the device at a specific sampling instant.

[0083] During the implementation of the plan, inertial measurement data was acquired. The three-axis acceleration vector With the three-axis angular velocity vector The acceleration derivative vector is obtained by calculating the observation increment between adjacent sampling times and dividing it by the sampling time interval. With the derivative vector of angular velocity Next, the sum of the Euclidean norms of the two vectors is calculated to obtain the scalar value of the motion intensity at the current instant. .

[0084] In a specific embodiment, during the digital twin modeling task, the mobile data acquisition terminal is held by the operator and used while moving between shelves. When the operator experiences occasional shaking due to uneven ground, the system can detect the movement at any given time. At the time Between these points, the acceleration vector abruptly changes from a steady-state value to a value with significant fluctuations. Through the aforementioned difference and magnitude summation operations, the resulting motion intensity scalar value jumps from a near-zero level to a higher dynamic value. By performing the above operations on all sampling points, a motion intensity scalar sequence reflecting the jitter characteristics at each point in the acquisition path is finally generated. .

[0085] In the process of mapping inertial measurement data to the Lie group manifold space and using B-spline basis functions to perform high-order continuous fitting on the motion trajectory of the mobile acquisition terminal to construct a time-continuous continuous trajectory model, the time node interval of the B-spline basis functions is set to the first preset interval during the time period when the value of the motion intensity scalar sequence is greater than the preset dynamic threshold.

[0086] During the time period when the value of the scalar sequence of motion intensity is less than or equal to the dynamic threshold, the time node interval of the B-spline basis function is set to the second preset interval, wherein the value of the first preset interval is less than the value of the second preset interval.

[0087] The dynamic threshold refers to a preset boundary value used to distinguish between steady motion and violent motion, denoted as: The first preset interval refers to the fine-grained time step size used for the high-frequency fluctuation segment during trajectory fitting. The second preset interval refers to the coarse-grained time step size used for the low-frequency stable segment, and its value is greater than the first preset interval.

[0088] By introducing this non-uniform time node distribution mechanism based on motion state, B-spline curves can possess more local shape control points in areas of complex motion, while maintaining the simplicity of mathematical expression and global smoothness in stable areas. During the implementation of the scheme, the generated motion intensity scalar sequence... Each value in the kinetic threshold Real-time comparison is performed. During the process of fitting the motion trajectory of the mobile acquisition terminal using B-spline basis functions, the distribution density of time nodes is dynamically switched based on the comparison results.

[0089] In a specific embodiment, a dynamic threshold can be set. When the mobile data acquisition terminal performs rapid turns or obstacle avoidance maneuvers during the inspection process, it detects the motion intensity scalar during that time period. Exceeded the threshold This is determined to be unsteady motion. Subsequently, the time interval of the B-spline basis function is set to a smaller first preset interval, such as 0.01 seconds, within this time period to enhance the trajectory model's fitting accuracy for steering details. Conversely, when the equipment moves at a constant speed along the straight path in the warehouse, the motion intensity scalar... Drop below the threshold At this point, the time interval is set to a larger second preset interval, for example... Seconds. Through this adaptive node configuration method, the constructed time-continuous trajectory model can reproduce complex microscopic motion trajectories with extremely high accuracy, while also effectively suppressing computational redundancy.

[0090] This embodiment, by performing time-domain differential processing on inertial measurement data, can capture the dynamic abrupt changes of a mobile acquisition terminal under unsteady motion conditions in real time. During periods of violent shaking or variable speed movement, the time node interval of the B-spline basis function is adaptively reduced to improve the fitting freedom and instantaneous expression accuracy of the continuous trajectory model for high-frequency motion details; while during the stable motion phase, the node interval is increased accordingly to effectively suppress computational redundancy while ensuring global trajectory smoothness. This non-uniform node distribution strategy based on changes in motion intensity enables the constructed trajectory model to accurately reproduce the microscopic motion trajectory under complex conditions, fundamentally eliminating the spatiotemporal coupling deviation between line-by-line exposure of the rolling shutter and point-by-point emission of laser pulses. Therefore, this application solves the problem of motion distortion and geometric tearing that is difficult to eliminate in the process of mobile modeling, and improves the texture mapping accuracy and high-fidelity reconstruction quality of digital twin 3D surfaces.

[0091] In one feasible implementation, S102: Mapping the inertial measurement data to the Lie group manifold space, and using B-spline basis functions to perform high-order continuous fitting on the motion trajectory of the mobile acquisition terminal to construct a time-continuous continuous trajectory model, including:

[0092] S1021: The angular velocity and acceleration components in the inertial measurement data are processed by discrete integration in the time dimension and converted into pose increment constraint information representing rotation and translation attributes in the Lie group manifold space.

[0093] The generation of pose increment constraint information refers to the mathematical expression of the rotation angle change and translation distance increment generated by moving the acquisition terminal between adjacent time sampling points. In the Lie group manifold space, it is reflected as an algebraic constraint describing the local motion trend.

[0094] During the implementation of the plan, the angular velocity vector in the inertial measurement data was used as a basis. and acceleration vector By employing a pre-integration algorithm to perform numerical accumulation within each extremely short sampling interval, the measured values ​​in the sensor coordinate system are mapped to the Lie group manifold. Space. In digital twin modeling scenarios, for the acquired motion observation sequences By performing an exponential mapping on the angular velocities of adjacent time points and a quadratic integral on the acceleration, an incremental matrix representing the local pose change can be obtained. .

[0095] S1022: In the Lie group manifold space, multiple trajectory control points including spatial attitude parameters and translation parameters are set, and the corresponding time series nodes are determined according to the sampling period of the mobile acquisition terminal. Multiple basis functions are used to perform weight allocation and weighted accumulation processing on the trajectory control points in the time axis direction to obtain the motion fitting curve representing the initial motion state.

[0096] Trajectory control points refer to a set of core parameter points set in the Lie group manifold space to determine the geometry and dynamic trajectory of the B-spline curve. Each point includes a three-dimensional rotation matrix and a three-dimensional translation vector. Time series nodes refer to a series of discrete time points distributed according to a specific density within the time period covered by the motion trajectory, used to define the support range of the basis functions.

[0097] First, the time series node sequence is determined based on the sampling period and total duration of the mobile acquisition terminal. This is then performed in the Lie group manifold space. Multiple trajectory control point sets are distributed in the middle. Next, a fourth-order B-spline basis function is selected to calculate the time weight allocation of each control point at the current fitting time. A set of control point vectors including the initial position and attitude guess can be defined, and these points are weighted and accumulated through the basis function. Its mathematical expression is shown in formula (1).

[0098] (1)

[0099] in, represent The first-order basis function is Time for the first The contribution weights of each control point are distributed to obtain a motion fitting curve that initially reflects the movement trend of the equipment and represents its initial motion state. The motion fitting curve refers to the curve formed by interpolating or approximating the discrete control points using basis functions, relating to the time variable. Continuous pose function.

[0100] S1023: Perform deviation convergence calculation on the pose increment constraint information and motion fitting curve to determine the final pose parameters of the trajectory control points, and parametrically express the trajectory control points including the final pose parameters and basis functions to obtain a continuous trajectory model.

[0101] The final pose parameters refer to the state of the trajectory control points determined through optimization in the Lie group manifold space, which includes the corrected rotation matrix and translation vector.

[0102] First, a nonlinear least squares-based objective function is constructed, treating the pose increment constraint information obtained in S1021 as the true observation value, and the motion fitting curve generated in S1022 as the variable to be optimized. Through bias convergence calculation, the Lie group distance between the output pose and the actual observed pose at each sampling time is minimized. Specifically, the error objective function... An example of the formula can be shown in formula (2).

[0103] (2)

[0104] in, For at any time The observation transformation matrix obtained from inertial measurement data, For the B-spline fitted curve at time 10 ... The predicted transformation matrix, This means mapping the error in the Lie group space to its corresponding Lie algebra tangent space to calculate the vector deviation under the Euclidean norm.

[0105] In a specific embodiment, the objective function can be optimized by calling the Levonburg Marquardt optimizer. Iterative solutions are performed. Taking a 2-second trajectory of a mobile data acquisition terminal in a warehouse as an example, the trajectory control points are continuously adjusted. The parameters are configured to ensure that the curve trajectory determined by each set of control points closely matches the pose sequence obtained from IMU integration. When the residual change during the iteration process is less than a preset minimum threshold, the computation is considered to have converged, and the values ​​of the control points are locked as the final pose parameters. Finally, these parameters are encapsulated with fourth-order B-spline basis functions to obtain the final continuous trajectory model.

[0106] In one feasible implementation, S103: Calculating the first pose transformation matrix at the exposure time of each row of pixels in the video image sequence, and the second pose transformation matrix at the emission time of each laser point in the laser point cloud data, based on a continuous trajectory model, includes:

[0107] S1031: Determine the instantaneous exposure time of each row of pixels in the video image sequence based on the line scan rate of the camera, and determine the instantaneous pulse time of each laser point in the laser point cloud data based on the pulse frequency of the lidar.

[0108] Line exposure instantaneous time refers to the physical moment when a specific row of pixels captures light information under a rolling shutter mechanism. Point pulse instantaneous time refers to the microscopic timestamp of a single laser pulse emitted by the laser detection unit and the recorded echo, which is usually determined by a high-precision clock counter inside the lidar.

[0109] Based on the obtained line scan rate of the camera and the pulse frequency of the lidar For any image frame in a video image sequence, based on its frame start time... and the row number of the target pixel row The instantaneous exposure time of the row pixels is calculated by accumulating the row intervals. Simultaneously, for the laser point cloud data, based on the emission sequence index of the laser points... The instantaneous time of the corresponding point pulse is calculated. In a specific embodiment, in the digital twin modeling scenario, two key time datasets are generated: row exposure time series. and point pulse time series .

[0110] S1032: Import the instantaneous exposure time of the row and the instantaneous pulse time of the dot pulse into the continuous trajectory model, determine the multiple trajectory control points adjacent to each row pixel and each laser point in the time axis direction, and use the basis function to calculate the contribution weight distribution of the multiple adjacent trajectory control points in the instantaneous exposure time of the row and the instantaneous pulse time.

[0111] The contribution weight distribution refers to a series of normalized coefficient values ​​calculated using B-spline basis functions under a specific time variable, which reflects the influence weight of adjacent control points on the current fitted pose.

[0112] The calculated instantaneous time series of line exposure can be used to solve the problem. and point pulse time series Import a pre-built continuous trajectory model. The model is based on each time parameter. Within the support interval of the node vector, retrieve a set of trajectory control points associated with it. ,in Let be the order of the B-spline. Then, the contribution weight distribution corresponding to that moment is calculated using a basis function iterative algorithm. In a specific embodiment, this is applied to the exposure instant of a certain row in a warehouse scene. The continuous trajectory model determines the sequence of neighboring control points along the time axis and calculates the weight vector. Each component in the vector It strictly corresponds to the contribution ratio of each neighboring control point to the instantaneous pose.

[0113] Figure 2 A schematic diagram of generating a pose transformation matrix is ​​shown in one embodiment of this application.

[0114] S1033: The contribution weight distribution is used to smooth and superimpose multiple trajectory control points, including the final pose parameters, to calculate the first pose transformation matrix representing the spatial position and orientation of each row of pixels at the exposure time, and the second pose transformation matrix representing the spatial position and orientation of each laser point at the emission time.

[0115] The first pose transformation matrix refers to the instantaneous position of the camera optical center in the global coordinate system. Transformation matrix The second pose transformation matrix refers to the global pose matrix of the lidar scanning center at the detection moment. .

[0116] The calculated contribution weight distribution can be used to perform smooth superposition processing of Lie group space on multiple neighboring trajectory control points. For example... Figure 2The diagram shown is a schematic of the pose transformation matrix solution. The pose is solved by Lie group interpolation as shown in formula (3).

[0117] (3)

[0118] in, These are the final pose parameters of the trajectory control points. To contribute weight components in the weight distribution, and These represent the exponential and logarithmic mappings between Lie groups and Lie algebras, respectively.

[0119] In the modeling process, for each row of pixels in the video sequence, formula (3) and the corresponding... Solve for the first pose transformation matrix Simultaneously, for each laser point in the laser point cloud, its emission time is used... Solve for the corresponding second pose transformation matrix Because the mobile acquisition terminal has slight displacements during movement, the calculated pose sequence shows a trend of microscopic changes over time. This ensures that each row of pixels and each laser point can obtain motion compensation that best matches physical reality, solving the common problems of texture mapping misalignment and geometric topology tearing in digital twin models.

[0120] In one feasible implementation, S104: The laser point cloud data is projected from the radar coordinate system to the digital twin global coordinate system according to the second pose transformation matrix. Based on the first pose transformation matrix and the camera's intrinsic parameters, a projection mapping relationship is established between the video image sequence and the laser point cloud data transformed to the digital twin global coordinate system using the perspective projection principle. This generates multi-dimensional feature mapping data, which includes spatiotemporal registration geometric coordinate information and texture attribute information, including:

[0121] S1041: The second pose transformation matrix is ​​used to perform spatial coordinate transformation on the laser point cloud data, mapping the local three-dimensional coordinate position of each laser point in the radar coordinate system to spatial positioning information in the digital twin global coordinate system.

[0122] The global spatial alignment radar coordinate system of a laser point cloud refers to a local right-handed coordinate system established with the geometric center of the laser radar as the origin. The local three-dimensional coordinate point refers to the original spatial coordinates, including the distance and angle information of the target object, measured by the laser radar at a specific pulse moment, and can be represented as... .

[0123] During the implementation of the solution, spatial positioning information refers to the absolute spatial coordinates of the laser point in the digital twin world coordinate system after coordinate system transformation. This can be achieved by extracting the local three-dimensional coordinates of the laser point from the laser point cloud data. Retrieve the pre-calculated second pose transformation matrix corresponding to the laser point at the emission time. The second pose transformation matrix is ​​used to perform spatial coordinate transformation on the local 3D coordinate point. By performing linear algebraic operations of matrix-vector multiplication, the spatial positioning information of the point in the digital twin global coordinate system is calculated. .

[0124] For example, in a modeling task of an industrial warehouse, for a laser point on the edge of a shelf collected at a certain moment, the local vector of this point in the radar coordinate system is... Multiply by the pose matrix at the current instant This process maps the data to the global map coordinate system of the warehouse. It eliminates laser point offset caused by continuous movement of the mobile acquisition terminal, ensuring the geometric accuracy of the point cloud data in global space.

[0125] S1042: Perform inverse perspective transformation processing on the spatial positioning information based on the first pose transformation matrix to calculate the relative spatial position parameters of each laser point in the camera's observation coordinate system.

[0126] The observation coordinate system refers to a system with the camera's optical center as the origin, and The camera's local reference space is where the axial direction coincides with the optical axis. Inverse viewpoint transformation refers to the process of converting a point's position from a unified global coordinate system back to the camera's observation viewpoint at a specific exposure instant.

[0127] During the implementation of the scheme, the relative spatial position parameter refers to the three-dimensional coordinates describing the laser point's position relative to the current instantaneous camera attitude in the observation coordinate system. The first pose transformation matrix at the current laser point's line exposure time is obtained. Calculate the inverse of the first pose transformation matrix. And use this inverse matrix to transform the spatial positioning information into global space. Perform reverse perspective transformation processing.

[0128] In a specific embodiment, for points that have already been aligned to the global coordinate system... By calculating its projected position from the camera's current viewpoint, using Solve for relative spatial position parameters When the handheld mobile acquisition terminal rotates or moves rapidly, even a tiny time increment can be compensated for in real time through this inverse transformation, so that the global spatial point can be accurately restored to the camera's field of view when the current row of pixels is exposed, eliminating the misalignment caused by the rolling shutter effect.

[0129] S1043: Using the camera's intrinsic parameters, perform perspective projection calculations on the relative spatial point parameters to project the three-dimensional points in space onto the two-dimensional imaging plane, obtaining the pixel coordinate index of each laser point in the video image sequence.

[0130] Perspective projection calculation refers to the mathematical process of mapping three-dimensional spatial points to a two-dimensional image coordinate system based on the geometric model of pinhole imaging. During implementation, the pixel coordinate index refers to the row and column position number of the laser point after projection into the video image matrix, typically represented as... By retrieving the camera's intrinsic parameters Using the camera's intrinsic parameters to determine the relative spatial position parameters. Perspective projection calculations are performed to establish a geometric mapping from three-dimensional spatial points to two-dimensional pixels. The calculation process uses the following projection formula (4).

[0131] (4)

[0132] in, This represents the depth factor from the laser point to the camera's imaging plane. In a specific embodiment, in a modeled industrial warehouse scenario, by mapping the relative spatial position of a point on the equipment surface in the observation coordinate system onto a video frame image, the pixel coordinate index corresponding to that point at the current moment can be calculated. .

[0133] S1044: Extract color component data from the video image sequence based on pixel coordinate index, and aggregate the color component data as texture attribute information and spatial positioning information to generate multidimensional feature mapping data.

[0134] The aggregation of multidimensional feature mapping data to generate color component data refers to the red, green, and blue color intensity values ​​stored at specific pixel coordinates in a video image, which can be represented as a color vector. Data field aggregation refers to the process of merging interrelated geometric coordinate attributes and texture color attributes into a single composite data structure.

[0135] During the implementation of the solution, the pixel coordinate index obtained from the solution can be used as a reference. Color component data is retrieved and extracted from the pixel matrix of the corresponding video frame image. The extracted color component data is then used as texture attribute information, along with the previously calculated spatial positioning information. Perform data field aggregation.

[0136] For example, in a warehouse modeling task, the global coordinates of laser points located on the wall surface are... Color vectors sampled from the video The data is encapsulated to generate multidimensional feature mapping data including six-dimensional attributes. The final feature dataset not only includes the real physical space coordinates, but also accurately attaches the corresponding instantaneous texture information.

[0137] In one feasible implementation, before S1044: extracting color component data from the video image sequence based on pixel coordinate indices, the method further includes:

[0138] The numerical components of the relative spatial point position parameters in the direction perpendicular to the camera imaging plane are extracted as the point cloud measurement depth.

[0139] Point cloud depth extraction refers to the projected distance of a spatial sampling point relative to the camera's optical center along the principal optical axis in the camera's observation coordinate system. During implementation, the relative spatial point position parameters obtained previously are used... The numerical components in the direction perpendicular to the camera's imaging plane are extracted. The relative spatial position parameters are represented as a three-dimensional coordinate vector. ,in The axis usually coincides with the optical axis, so it can be directly extracted. Axis components as depth measurement of point cloud Specifically, in the digital twin modeling task of an industrial warehouse, each set of data is acquired in real time. The depth components are stored as a depth dataset distributed over time. .

[0140] Detect whether the pixel coordinate index exceeds the preset pixel boundary range of the video image sequence.

[0141] The preset pixel boundary range refers to the effective imaging area determined by the physical resolution of the camera's image sensor, and is composed of a set of coordinate constraints in the horizontal and vertical directions. Detecting whether the pixel coordinate indices exceed this range is to exclude spatial points that are projected outside the imaging plane and cannot be used to obtain effective texture and color.

[0142] During the implementation of the solution, the image width of the video image sequence is obtained. With height The pixel coordinate index calculated by perspective projection operation and Logical determinations are performed against the corresponding boundary intervals. In a specific embodiment, for the projected coordinates of any laser sampling point within the industrial warehouse, the determination logic is as follows: Detection Does it meet the requirements? ,and Does it meet the requirements? If the result is true, it means that the point falls within the currently captured video frame; otherwise, it is determined that the coordinates are out of bounds.

[0143] If the point cloud measurement depth is within the effective detection range of the camera and the pixel coordinate index does not exceed the preset pixel boundary range, the mapping relationship between the corresponding spatial positioning information and the pixel coordinate index is retained; otherwise, the corresponding spatial positioning information is marked as a non-visible point to determine the effective mapping point for participating in color component data extraction.

[0144] An effective mapping point refers to a laser sampling point that possesses high-fidelity texture mapping conditions and is within the current camera's field of view and the sensor's reliable observation range. A non-visible point refers to a spatial geometric location that cannot be reliably correlated with video color because it is outside the lens's angle of view or exceeds the effective measurement range.

[0145] During the implementation of the plan, the effective detection range of the pre-set camera is determined. Depth measurement in point clouds If the pixel coordinate index is within the effective detection range and does not exceed the preset pixel boundary range, the mapping relationship between the corresponding spatial positioning information and the pixel coordinate index is retained, and it is determined as a valid mapping point. If any of the above judgment conditions are false, the corresponding spatial positioning information is marked as a non-visible point.

[0146] For example, in warehouse modeling, when the operator turns their handheld terminal to a corner, causing some shelf points to extend beyond the field of view, or the depth of the points to exceed the threshold for clear texture imaging, the situation may occur. At that time, these points are automatically filtered out, and only high-quality mapping relationships are retained. Through this filtering mechanism, erroneous mappings caused by occlusion or outside the field of view are eliminated, so that the points that participate in the extraction of color components in S1044 are all high-quality valid points, thereby ensuring the accuracy of the final digital twin model texture mounting.

[0147] Figure 3 A schematic diagram of truncated symbol distance field update and surface reconstruction provided in one embodiment of this application is shown.

[0148] In one feasible implementation, S105: The multidimensional feature mapping data is input as an update source to the voxelized truncated symbolic distance field; the symbolic distance values ​​within the voxel mesh are updated using geometric coordinate information; the color weight values ​​within the voxel mesh are updated using texture attribute information; and the zero-level set isosurface is extracted from the updated truncated symbolic distance field to generate a digital twin 3D surface, including:

[0149] S1051: Determine the voxel position in space based on the global coordinate system of the digital twin, and determine the target voxel mesh within the preset cutoff distance range of the voxel position.

[0150] Voxel position maps geometric coordinate information to discretized spatial coordinates after a 3D spatial mesh index. Preset cutoff distance range. It refers to the distance threshold limit used to define the scope of geometric updates in a truncated symbolic distance field.

[0151] During the implementation of the plan, such as Figure 3 As shown, the target voxel mesh This refers to a mesh in three-dimensional space divided into numerous tiny cubic units, subject to specific geometric sampling points. A set of local processing units for observing the impact. In practice, a three-dimensional voxel mesh space is established based on the digital twin's global coordinate system. The voxel position of the geometric coordinate information in space is determined by dividing the geometric coordinate information by a preset voxel resolution and performing a floor operation. Subsequently, within a preset cutoff distance range centered on this voxel position... Search for all target voxel meshes that meet the distance constraints.

[0152] For example, in the process of modeling the internal facilities of an industrial warehouse, a geometric sampling point that includes global coordinate information... Mapped to 3D mesh index and with a step size of Within a preset cutoff distance range, a set of target voxel meshes adjacent to it is identified. .

[0153] S1052: Calculate the projected distance from the center point of the target voxel mesh to the geometric coordinate information, and perform a weighted fusion process with the projected distance and the pre-stored distance value in the target voxel mesh to obtain the updated symbolic distance value.

[0154] Projected distance refers to the signed metric deviation from the geometric center point of the target voxel mesh to the location of the geometric coordinate information in the sensor's viewing direction. The signed distance value is a scalar value used to represent the distance between the center point of the voxel mesh and the actual physical surface; its positive or negative sign is used to distinguish whether the voxel is located inside or outside the object.

[0155] During the implementation of the scheme, the spatial displacement vector of the center point of each target voxel grid and the current geometric coordinate information is calculated. The projection distance is obtained by calculating the projection modulus of this vector in the sensor's observation direction. Subsequently, a weighted fusion process is performed between the projected distance and the pre-stored distance values ​​within the target voxel mesh to obtain the updated symbolic distance value. The specific weighted fusion formula can be shown in formula (5).

[0156] (5)

[0157] in, The updated symbolic distance value, The original cumulative weight, The current projection distance, This serves as the updated weight for the current observation. For example, for a specific shelf support sampling point, this process fuses the geometric features collected during multiple continuous movements into the voxel field, effectively eliminating discrete noise generated by single-frame sensor measurements.

[0158] S1053: Perform a spatiotemporal co-mapping between the color values ​​in the texture attribute information and the updated symbolic distance values. Accumulate and calculate the color components within the target voxel mesh according to the weight ratio used when updating the symbolic distance values ​​to obtain the updated color weight values.

[0159] Color weight values ​​refer to numerical vectors stored within a voxel grid that reflect the average color attribute and the reliability of color observations at that spatial location. During the implementation of the scheme, such as... Figure 3 As shown in the mid-voxel update processing section, the texture attribute (RGB) color values ​​in the multidimensional feature map data are obtained. The color components within the target voxel grid are cumulatively calculated according to the weight ratio used when updating the symbolic distance value, achieving synchronous updates of color and geometry in the spatiotemporal dimension, thus obtaining the updated color weight values.

[0160] For example, updating weights geometrically. This is also applied to the weighted average calculation of color components. In the specific scenario of warehouse modeling, for the identified valid mapping points, their color vectors are... Co-mapping is performed with the color data pre-stored within the voxels. By calculating the weighted cumulative values ​​of the color components, each target voxel mesh not only records the positional features of the surface but also incorporates high-fidelity texture features from the video image sequence.

[0161] S1054: In a truncated symbolic distance field composed of multiple target voxel grids, search for zero-crossing spatial locations where the symbolic distance values ​​between two adjacent target voxel grids alternate between positive and negative, and extract the zero-level set isosurface by performing triangular patch fitting on the zero-crossing spatial locations.

[0162] A zero-crossing spatial location refers to the boundary position in a truncated sign distance field where the sign distance value changes from positive (+) to negative (-) or from negative to positive; it represents the physical surface of an object. A zero-level set isosurface is a continuous geometric surface formed by connecting all spatial points with a sign distance value of zero.

[0163] During the implementation of the plan, such as Figure 3 As shown, the triangular patch fitting process constructs a triangular mesh set to approximate the real surface of the 3D object by connecting the vertices generated from the zero-crossing spatial locations. This process performs a mesh traversal search in the truncated signed distance field, detecting whether the sign of the signed distance value between two adjacent target voxel meshes has changed, thereby retrieving the zero-crossing spatial locations. Subsequently, a linear interpolation algorithm is used to accurately locate the zero-crossing points on the mesh edges, and the zero-level set isosurface is extracted through fitting.

[0164] For example, the moving cube algorithm is used to generate isosurface vertices on the edges of each voxel cell and connect them to form triangles. In a surface reconstruction example of a complex industrial equipment, a digital twin 3D surface was successfully generated by performing retrieval and fitting on the full voxel mesh. The final generated surface has a smooth topological structure and realistic surface texture, effectively overcoming the geometric layering and topological tearing problems in traditional modeling methods.

[0165] Figure 4 This application provides a schematic diagram of a specific implementation of a video point cloud data fusion and surface reconstruction system based on a digital twin scene, referring to... Figure 4 The system may include:

[0166] The acquisition module 410 is used to simultaneously acquire video image sequences and laser point cloud data of the target scene collected by the mobile acquisition terminal, which integrates a camera, lidar and inertial measurement unit, during continuous movement, as well as the inertial measurement data of the mobile acquisition terminal.

[0167] Module 420 is used to map inertial measurement data to the Lie group manifold space and use B-spline basis functions to perform high-order continuous fitting on the motion trajectory of the mobile acquisition terminal to construct a time-continuous continuous trajectory model.

[0168] The calculation module 430 is used to calculate the first pose transformation matrix at the exposure time of each row of pixels in the video image sequence and the second pose transformation matrix at the emission time of each laser point in the laser point cloud data based on the continuous trajectory model.

[0169] The generation module 440 is used to project the laser point cloud data from the radar coordinate system to the digital twin global coordinate system according to the second pose transformation matrix, and to establish the projection mapping relationship between the video image sequence and the laser point cloud data after transformation to the digital twin global coordinate system based on the first pose transformation matrix and the camera's intrinsic parameters, using the perspective projection principle, and to generate multi-dimensional feature mapping data, which includes spatiotemporal registration geometric coordinate information and texture attribute information.

[0170] The generation module 440 is also used to input multidimensional feature mapping data as an update source into the voxelized truncated symbolic distance field, update the symbolic distance value in the voxel grid using geometric coordinate information and update the color weight value in the voxel grid using texture attribute information, and extract the zero level set isosurface from the updated truncated symbolic distance field to generate a digital twin three-dimensional surface.

[0171] The video point cloud data fusion and surface reconstruction system based on digital twin scenes in this application embodiment is used to implement the aforementioned video point cloud data fusion and surface reconstruction method based on digital twin scenes. Therefore, the specific implementation of the video point cloud data fusion and surface reconstruction system based on digital twin scenes can be found in the embodiment section of the video point cloud data fusion and surface reconstruction method based on digital twin scenes above. The specific implementation can be referred to the description of the corresponding embodiments, which will not be repeated here.

[0172] Figure 5 A schematic diagram of the hardware structure of an electronic device provided in one embodiment of this application is shown.

[0173] The electronic device may include a processor 510 and a memory 520 storing computer program instructions.

[0174] Specifically, the processor 510 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0175] Memory 520 may include mass storage for data or instructions. For example, and not limitingly, memory 520 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 520 may include removable or non-removable (or fixed) media. Where appropriate, memory 520 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 520 is non-volatile solid-state memory.

[0176] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect of this disclosure.

[0177] The processor 510 reads and executes computer program instructions stored in the memory 520 to implement any of the video point cloud data fusion and surface reconstruction methods based on digital twin scenes in the above embodiments.

[0178] In one example, the electronic device may also include a communication interface 530 and a bus 540. Wherein, such as Figure 5 As shown, the processor 510, memory 520, and communication interface 530 are connected through bus 540 and complete communication with each other.

[0179] The communication interface 530 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0180] Bus 540 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 540 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.

[0181] The electronic device can execute the video point cloud data fusion and surface reconstruction method based on digital twin scene in the embodiments of this application, thereby realizing the video point cloud data fusion and surface reconstruction method based on digital twin scene described in conjunction with the accompanying drawings.

[0182] Furthermore, in conjunction with the video point cloud data fusion and surface reconstruction method based on digital twin scenes in the above embodiments, this application embodiment can provide a computer-readable storage medium for implementation. This computer-readable storage medium stores computer program instructions; when executed by a processor, these computer program instructions implement any of the video point cloud data fusion and surface reconstruction methods based on digital twin scenes in the above embodiments.

[0183] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0184] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0185] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0186] The aspects of this application have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0187] The foregoing has provided a detailed description of the video point cloud data fusion and surface reconstruction method, system, electronic device, and storage medium based on a digital twin scene provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A method for video point cloud data fusion and surface reconstruction based on digital twin scenes, characterized in that, include: The system simultaneously acquires video image sequences and laser point cloud data of the target scene collected by a mobile acquisition terminal integrating a camera, lidar, and inertial measurement unit during continuous movement, as well as the inertial measurement data of the mobile acquisition terminal. The inertial measurement data is mapped to the Lie group manifold space, and the motion trajectory of the mobile acquisition terminal is fitted with a high-order continuous model using B-spline basis functions to construct a time-continuous continuous trajectory model. Based on the continuous trajectory model, the first pose transformation matrix at the exposure time of each row of pixels in the video image sequence and the second pose transformation matrix at the emission time of each laser point in the laser point cloud data are calculated. The laser point cloud data is projected from the radar coordinate system to the digital twin global coordinate system according to the second pose transformation matrix. Based on the first pose transformation matrix and the intrinsic parameters of the camera, the projection mapping relationship between the video image sequence and the laser point cloud data after transformation to the digital twin global coordinate system is established using the perspective projection principle, thereby generating multidimensional feature mapping data. The multidimensional feature mapping data includes spatiotemporal registration geometric coordinate information and texture attribute information. The multidimensional feature mapping data is used as an update source input to the voxelized truncated symbolic distance field. The symbolic distance values ​​in the voxel grid are updated using the geometric coordinate information and the color weight values ​​in the voxel grid are updated using the texture attribute information. The zero level set isosurface is extracted from the updated truncated symbolic distance field to generate a digital twin 3D surface.

2. The method according to claim 1, characterized in that, The method further includes: The inertial measurement data is processed by time-domain differentiation to calculate the time derivative magnitudes of the acceleration vector and angular velocity vector, and a scalar sequence of motion intensity corresponding to different time moments is generated based on the time derivative magnitudes. In the process of mapping the inertial measurement data to the Lie group manifold space and using B-spline basis functions to perform high-order continuous fitting on the motion trajectory of the mobile acquisition terminal to construct a time-continuous continuous trajectory model, during the time period when the value of the motion intensity scalar sequence is greater than the preset dynamic threshold, the time node interval of the B-spline basis functions is set to the first preset interval. During the time period when the value of the motion intensity scalar sequence is less than or equal to the dynamic threshold, the time node interval of the B-spline basis function is set to a second preset interval, wherein the value of the first preset interval is less than the value of the second preset interval.

3. The method according to claim 1, characterized in that, The step of mapping the inertial measurement data to a Lie group manifold space and using B-spline basis functions to perform high-order continuous fitting on the motion trajectory of the mobile acquisition terminal to construct a time-continuous continuous trajectory model includes: The angular velocity and acceleration components in the inertial measurement data are processed by discrete integration in the time dimension and converted into pose increment constraint information representing rotation and translation attributes in the Lie group manifold space. Multiple trajectory control points, including spatial attitude parameters and translation parameters, are set in the Lie group manifold space. The corresponding time series nodes are determined according to the sampling period of the mobile acquisition terminal. Multiple basis functions are used to perform weight allocation and weighted accumulation processing on the trajectory control points in the time axis direction to obtain the motion fitting curve representing the initial motion state. The deviation convergence calculation is performed on the pose increment constraint information and the motion fitting curve to determine the final pose parameters of the trajectory control points. The trajectory control points including the final pose parameters and the basis functions are then parameterized to obtain the continuous trajectory model.

4. The method according to claim 3, characterized in that, The step of calculating the first pose transformation matrix at the exposure time of each row of pixels in the video image sequence based on the continuous trajectory model, and the second pose transformation matrix at the emission time of each laser point in the laser point cloud data, includes: The instantaneous exposure time of each row of pixels in the video image sequence is determined based on the line scan rate of the camera, and the instantaneous pulse time of each laser point in the laser point cloud data is determined based on the pulse frequency of the lidar. The instantaneous exposure time of the row and the instantaneous pulse time of the point are imported into the continuous trajectory model to determine the multiple trajectory control points adjacent to each row pixel and each laser point in the time axis direction. The contribution weight distribution of the multiple adjacent trajectory control points in the instantaneous exposure time of the row and the instantaneous pulse time is calculated using the basis function. The contribution weight distribution is used to smooth and superimpose multiple trajectory control points, including the final pose parameters, to calculate the first pose transformation matrix representing the spatial position and orientation of each row of pixels at the exposure time, and the second pose transformation matrix representing the spatial position and orientation of each laser point at the emission time.

5. The method according to claim 1, characterized in that, The process involves projecting the laser point cloud data from the radar coordinate system to the digital twin global coordinate system based on the second pose transformation matrix, and establishing a projection mapping relationship between the video image sequence and the laser point cloud data transformed to the digital twin global coordinate system using the perspective projection principle based on the first pose transformation matrix and the intrinsic parameters of the camera, thereby generating multidimensional feature mapping data. This multidimensional feature mapping data includes spatiotemporal registration geometric coordinate information and texture attribute information, including: The second pose transformation matrix is ​​used to perform spatial coordinate transformation processing on the laser point cloud data, mapping the local three-dimensional coordinate position of each laser point in the radar coordinate system to spatial positioning information in the digital twin global coordinate system. Based on the first pose transformation matrix, the spatial positioning information is subjected to inverse viewpoint transformation processing to calculate the relative spatial position parameters of each laser point in the camera's observation coordinate system. Using the intrinsic parameters of the camera, perspective projection calculation is performed on the relative spatial point parameters to project the three-dimensional points in space onto the two-dimensional imaging plane, thereby obtaining the pixel coordinate index of each laser point in the video image sequence. Color component data is extracted from the video image sequence based on the pixel coordinate index, and the color component data is aggregated with the spatial positioning information as the texture attribute information to generate the multidimensional feature mapping data.

6. The method according to claim 5, characterized in that, Before extracting color component data from the video image sequence based on the pixel coordinate index, the method further includes: The numerical components of the relative spatial point position parameters in the direction perpendicular to the camera imaging plane are extracted as the point cloud measurement depth; Detect whether the pixel coordinate index exceeds the preset pixel boundary range of the video image sequence; If the point cloud measurement depth is within the effective detection range of the camera and the pixel coordinate index does not exceed the preset pixel boundary range, the mapping relationship between the corresponding spatial positioning information and the pixel coordinate index is retained; otherwise, the corresponding spatial positioning information is marked as a non-visible point to determine the effective mapping point for participating in the color component data extraction.

7. The method according to claim 1, characterized in that, The process of inputting the multidimensional feature mapping data as an update source into a voxelized truncated symbolic distance field, updating the symbolic distance values ​​within the voxel mesh using the geometric coordinate information, updating the color weight values ​​within the voxel mesh using the texture attribute information, and extracting the zero-level set isosurface from the updated truncated symbolic distance field to generate a digital twin 3D surface includes: The voxel position of the geometric coordinate information in space is determined according to the digital twin global coordinate system, and the target voxel mesh within the preset cutoff distance range of the voxel position is determined. Calculate the projection distance from the center point of the target voxel mesh to the geometric coordinate information, and perform a weighted fusion process on the projection distance and the pre-stored distance value in the target voxel mesh to obtain the updated symbolic distance value; Perform a spatiotemporal co-mapping between the color values ​​in the texture attribute information and the updated symbolic distance value, and accumulate the color components in the target voxel grid according to the weight ratio used when updating the symbolic distance value to obtain the updated color weight value. In the truncated symbolic distance field composed of multiple target voxel grids, the zero-crossing spatial locations where the symbolic distance values ​​between two adjacent target voxel grids alternate between positive and negative are retrieved, and the zero-level set isosurface is extracted by performing triangular patch fitting on the zero-crossing spatial locations.

8. A video point cloud data fusion and surface reconstruction system based on digital twin scenes, characterized in that, include: The acquisition module is used to simultaneously acquire video image sequences and laser point cloud data of the target scene collected by a mobile acquisition terminal integrating a camera, lidar and inertial measurement unit during continuous movement, as well as the inertial measurement data of the mobile acquisition terminal. The module is used to map the inertial measurement data to the Lie group manifold space and use B-spline basis functions to perform high-order continuous fitting on the motion trajectory of the mobile acquisition terminal to construct a time-continuous continuous trajectory model. The calculation module is used to calculate the first pose transformation matrix at the exposure time of each row of pixels in the video image sequence and the second pose transformation matrix at the emission time of each laser point in the laser point cloud data based on the continuous trajectory model. The generation module is used to project the laser point cloud data from the radar coordinate system to the digital twin global coordinate system according to the second pose transformation matrix, and to establish the projection mapping relationship between the video image sequence and the laser point cloud data after transformation to the digital twin global coordinate system according to the first pose transformation matrix and the intrinsic parameters of the camera, using the perspective projection principle, to generate multi-dimensional feature mapping data, which includes spatiotemporal registration geometric coordinate information and texture attribute information; The generation module is also used to input the multidimensional feature mapping data as an update source into the voxelized truncated symbolic distance field, update the symbolic distance value in the voxel grid using the geometric coordinate information and update the color weight value in the voxel grid using the texture attribute information, and extract the zero level set isosurface from the updated truncated symbolic distance field to generate a digital twin three-dimensional surface.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor is configured to execute the computer program to implement the steps of the video point cloud data fusion and surface reconstruction method based on a digital twin scene as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, enables the video point cloud data fusion and surface reconstruction method based on a digital twin scene as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • High-fidelity three-dimensional reconstruction method fusing attitude prior and geometric constraint

    CN120931839A

  • Updating a 3D map of an environment

    WO2023219682A1