Digital twinborn live-action construction system and method based on multi-source video image fusion
By using multi-source video image fusion technology, the problems of low efficiency in multi-source data collaborative processing and model update are solved, realizing high-precision, real-time digital twin reality construction, providing quantitative evaluation methods, and ensuring the dynamic synchronization and reliability of the model and reality.
Patent Information
- Application Number
- CN202511518186.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2026-02-03
AI Technical Summary
Existing technologies lack the ability to collaboratively process multi-source data, have limited synchronization and preprocessing accuracy, insufficient reliability of fusion reconstruction and model update efficiency, are prone to geometric conflicts, have incomplete model verification and quality assessment systems, and lack accurate measurement of the fit with the real scene.
By constructing a data acquisition unit, a multi-source synchronization unit, a preprocessing and calibration unit, a fusion and reconstruction unit, a real-time update and consistency maintenance unit, and a visualization and verification unit, and by employing technologies such as unified timestamps, interface protocols, collaborative control, time-series calibration, noise suppression, distortion correction, multi-resolution representation, and spatial indexing, a multi-source data stream with consistent calibration and clear confidence is generated, enabling efficient construction and real-time updating of the 3D model.
It improves the geometric and photometric consistency of multi-source data, reduces fusion and reconstruction errors, achieves detailed integrity and real-time updates of digital twin models, provides quantitative evaluation methods, and ensures dynamic synchronization and reliability between the model and the real scene.
Smart Images

Figure CN121459239A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent video analysis, in particular to a digital twin real scene construction system and method based on multi-source video image fusion. BACKGROUND
[0002] The core of the digital twin real scene construction technology is to construct a three-dimensional model that is consistent with the physical scene in real time by integrating multi-source perception data, providing support for scene monitoring and management in the fields of intelligent manufacturing and smart city. The key of this technology lies in solving the problems of ordered collection, accurate synchronization, quality optimization, efficient fusion and real-time model updating of multi-source data. Only by overcoming the difficulties of data time sequence asynchronization, format heterogeneity and quality fluctuation through a systematic technical solution can the geometric accuracy and time consistency of the digital twin real scene be guaranteed, meeting the needs of model precision and real-time performance in actual applications.
[0003] In the prior art, related patents have explored in the field of multi-source data processing and digital twin model construction. For example, Chinese patent CN202410932282.2 discloses a three-dimensional video fusion method and system based on digital twin technology, including data collection, data processing, model establishment, data correlation, video fusion, analysis results and data storage steps. Its advantage lies in constructing a digital twin model by collecting entity data of the region, importing and fusing the processed video data with the model, and simultaneously detecting and optimizing the fusion effect and quality to present a more realistic scene, improve visual appeal and immersion, and help users understand the equipment operation status in the region. Chinese patent CN202510475311.1 discloses an artificial intelligence video processing method based on digital twin. This method solves the time inconsistency and missing frame problems of multi-source video data through a timestamp alignment mechanism and time interpolation method, ensuring data integrity and synchronization. It uses multi-source integration technology and three-dimensional convolution network to fuse multi-view video data and construct a high-precision virtual model, optimizes dynamic scene mapping through real-time parameter updating and Kalman filtering, improves model precision and real-time performance, predicts target trajectory and classifies behavior patterns through long short-term memory network and attention mechanism, provides support for intelligent decision-making, improves data processing efficiency and accuracy, and enhances the adaptability and intelligence level of the system in dynamic scenes.
[0004] Although the above technical solutions have corresponding design advantages, the above technical solutions still have the following technical defects: firstly, the multi-source data collaborative processing capability is insufficient, and the synchronization and preprocessing accuracy is limited: for CN202410932282.2, it only focuses on the fusion of video data and digital twin model, does not involve the collaborative collection design of multi-source sensing data such as depth sensor, inertial sensor and position information, and does not develop a scheme for key preprocessing links such as data noise classification suppression and lens distortion correction, so that the input data quality is difficult to support high-precision digital twin model construction; for CN202510475311.1, although the time stamp alignment and interpolation method is used to improve the time sequence problem, the event-level data matching and clock drift estimation algorithm is not introduced, the synchronization deviation caused by the clock offset in the long-term operation of the sensor cannot be solved, and the spatial matching accuracy of multi-source data is not optimized, which hides the error risk for subsequent fusion reconstruction. Secondly, the fusion reconstruction reliability and model updating efficiency are insufficient, and geometric conflicts are prone to occur: CN202410932282.2 does not construct a data confidence labeling system, the weight distribution in the fusion process lacks accurate basis, and does not design an incremental updating mechanism, which needs to use full-amount reconstruction mode in the face of scene changes, resulting in low model updating efficiency; CN202510475311.1 realizes data fusion through a three-dimensional convolution network, but does not consider view consistency detection and occlusion perception, which is easy to cause geometric distortion of the reconstructed model due to conflicts between different view data, and does not use a multi-resolution storage strategy, resulting in high resource consumption of model storage and processing. Thirdly, the model verification and quality evaluation system is incomplete, and the precise measurement of the real scene consistency is lacking: the above two patents do not establish a quantitative verification system combined with data confidence, only indirectly reflect the technical value through subjective effect optimization or intelligent decision support, which cannot quantitatively evaluate the consistency degree of the digital twin model and the physical real scene from the core dimensions of geometric accuracy and texture consistency, and it is difficult to guarantee the reliability of the model in actual scene application. In view of this, we propose a digital twin real scene construction system based on multi-source video image fusion. SUMMARY
[0005] The present application aims to provide a digital twin real scene construction system based on multi-source video image fusion and a method thereof, to solve the problems of insufficient multi-source data collaborative processing capability, limited synchronization and preprocessing accuracy, insufficient fusion reconstruction reliability and model updating efficiency, prone to geometric conflicts, and incomplete model verification and quality evaluation system, and lack of precise measurement of real scene consistency in the background art.
[0006] To solve the above technical problems, one of the purposes of the present application is to provide a digital twin real scene construction system based on multi-source video image fusion, comprising: A data acquisition unit is configured to orderly acquire and standardize multi-source perception data, integrate unified timestamp standards, interface protocols, preliminary compression and packaging mechanisms, realize cooperative acquisition of cameras, depth sensors, inertial sensors and position information, and output correlatable data streams to downstream; A multi-source synchronization unit is configured to maintain the timing consistency of multi-channel asynchronous data streams, combine timing calibration, event-level matching and clock drift estimation algorithms, complete time alignment and sensor identification alignment of multi-source data, output synchronized multi-modal frame sequences, and ensure the timing accuracy of subsequent data processing; A preprocessing and calibration unit is configured to optimize the quality of multi-modal data and guarantee the geometric and photometric consistency, adopt noise suppression, distortion correction, initial estimation of internal and external parameters, and spatial matching technology based on the improved SIFT+FLANN algorithm to generate input data with consistent calibration and confidence annotation; A fusion and reconstruction unit is configured to construct a digital twin three-dimensional scene model based on multi-source data, generate three-dimensional models in the form of meshes and voxels through hierarchical confidence weighted multi-resolution representation, view consistency detection, occlusion perception fusion and incremental geometry construction process, and output versions and confidence indicators to meet the geometric restoration requirements of digital twins; A real-time update and consistency maintenance unit is configured to maintain the time consistency of digital twin models and real scenes and realize incremental updates, based on spatial index-based incremental change detection, model difference merging and version control strategies, to integrate new data into existing models and guarantee global and local consistency, and output change logs to realize scene dynamic synchronization; A visualization and verification unit is configured to verify the reproducibility of digital twin models and evaluate the quantitative effects, combined with multi-view rendering, consistency index calculation and comparison test functions with calibrated scenes, to output visualization results, technical performance index reports and quantitative data.
[0007] As a further improvement of the technical solution, the data acquisition unit includes a time synchronization module, an interface adaptation module, a data preprocessing module, a cooperative control module and a data stream output module, wherein: The time synchronization module integrates high-precision clock sources and network time synchronization protocols to generate format-unified timestamps for the acquisition data of cameras, depth sensors, inertial sensors and position information, and realizes time reference alignment of multi-source perception data through periodic clock calibration mechanisms; The interface adaptation module configures a layered communication protocol stack, supports a data transmission protocol of multiple types of sensors and control instructions, respectively adapts data transmission requirements of a depth sensor, a camera, position information and a sensor control instruction, and realizes protocol conversion and data access of heterogeneous sensors; The data preprocessing module adopts a lossless compression algorithm to compress image type raw data, removes outliers in inertial sensor data through statistical filtering criteria, and standardizes and packages the processed data according to a preset format. The cooperative control module synchronously controls a camera exposure time, a depth sensor sampling period, an inertial sensor data reading frequency and a position information refreshing time through a cooperative mechanism of a hardware trigger signal and a software scheduling instruction, and guarantees frame-level acquisition synchronization of multi-source perception data. The data stream output module constructs a ring buffer structure based on shared memory, integrates time stamps generated by the time synchronization module, raw data accessed by the interface adaptation module and data packets packaged by the data preprocessing module according to a “frame number-sensor ID-time stamp” association rule, forms an associated data stream and transmits it to the multi-source synchronization unit, and adopts a data integrity checking mechanism in the transmission process.
[0008] As a further improvement of the technical solution, the multi-source synchronization unit includes a timing calibration and drift estimation module, an event matching and identification alignment module and a synchronous frame sequence output module, wherein: The timing calibration and drift estimation module performs timing calibration through a sliding window and a least square method, and estimates clock drift through a chi-square test, outputs calibrated time stamps and clock drift calibration factors, and realizes time reference consistency of multi-source data. The event matching and identification alignment module identifies and matches data of the same event in different sensors based on time stamps and event identifications, generates event-level association information, and aligns data of different sensors through comparison of sensor identification and time stamps, to ensure consistency of multi-source data at the event level and sensor identification, and provide statistical information of alignment error. The synchronous frame sequence output module integrates multi-source data that has undergone timing calibration, event matching and identification alignment processing according to the “frame number-sensor ID-time stamp” association rule, forms a synchronous multi-modal frame sequence, and outputs it to the preprocessing and calibration unit, and adopts a data integrity checking mechanism in the transmission process to ensure data accuracy and integrity.
[0009] As a further improvement of the technical solution, the preprocessing and calibration unit includes a noise suppression module and a distortion correction module, wherein: The noise suppression module is used for noise suppression of multi-modal data, extracts dynamic noise features through a Kalman filtering algorithm , quantize the random noise based on the time-series fluctuation law of the "prediction - measurement" residual; extract the trend noise features through sliding window adaptive filtering , adjust the filtering intensity according to the smoothness of the data within the window (enhance filtering when the variance within the window > threshold); extract the abnormal noise features through multi-source data joint verification , identify the isolated noise points deviating from the multi-source data consistency; The distortion correction module performs correction based on the lens distortion model, and the lens distortion model quantizes the radial deformation law of the lens through the radial distortion parameters , and quantizes the tangential offset caused by the lens assembly deviation through the tangential distortion parameters ; the radial distortion parameters and tangential distortion parameters are obtained by collecting multi-view images through a checkerboard calibration plate, performing corner detection and least squares fitting; after obtaining and , combined with the pixel coordinates and the distance from the optical center , calculate the corrected coordinates through the distortion model to achieve image geometric correction and ensure the photometric consistency of multi-source data.
[0010] As a further improvement of this technical solution, the preprocessing and calibration unit further includes an internal and external parameter estimation module and a spatial matching module, where: The internal and external parameter estimation module uses feature point matching combined with the least squares method to achieve parameter estimation: through SIFT-based feature point matching and the RANSAC algorithm, complete the calibration of the internal parameter matrix ; through minimizing the reprojection error, iteratively optimize the external parameters, namely the rotation matrix , the translation vector , and output the calibrated and consistent data; The spatial matching module generates the input data with confidence annotations based on the improved SIFT+FLANN algorithm, specifically including: Extract scale-invariant feature points through the SIFT algorithm with adaptive parameter adjustment, and dynamically adjust the Gaussian kernel scale and the number of difference layers of the Gaussian difference pyramid according to the illumination intensity and texture complexity of the input image to enhance the responsiveness and distinguishability of feature points in low-light and weak-texture scenarios; Calculate the Euclidean distance of the feature point descriptors using the FLANN algorithm , and introduce a dynamic threshold based on the local descriptor similarity distribution to screen the valid matching pairs; For the screened matching pairs, further remove the outliers using the weighted RANSAC algorithm; combined with the internal parameter matrix , rotation matrix and translation vector , mapping feature points of multi-source data to a unified coordinate system, based on the number of matching point pairs , distance variance and feature point stability weight , fusing the matching number proportion, distance variance and average weight to calculate the matching confidence , completing the spatial matching of multi-source data and outputting the result with confidence annotation.
[0011] As a further improvement of the technical solution, the fusion and reconstruction unit includes a hierarchical confidence weighting module, a multi-resolution representation module and a geometric incremental expansion module, wherein: The hierarchical confidence weighting module assigns dynamic weights (weights are negatively correlated with the number of matching point pairs and distance variance, and positively correlated with sensor accuracy) to different sensor data based on the matching confidence output by the preprocessing and calibration unit, performs weighted fusion on multi-source observations of the same spatial position, and generates fusion data with confidence labels; The multi-resolution representation module divides resolution levels according to scene complexity (high resolution for close-range areas and low resolution for long-range areas), organizes fusion data through octree structure, retains millimeter-level geometric details in high-resolution levels, and compresses redundant information in low-resolution levels, achieving efficient storage and representation of three-dimensional models; The geometric incremental expansion module dynamically expands the octree structure based on time-series new data: for newly collected scene areas, new octree nodes are added at the corresponding resolution level and filled with fusion data; for update data of already reconstructed areas, only the geometric information of the corresponding nodes is updated, and through node hash index, the fast splicing of historical models and incremental data is realized, avoiding the redundant calculation of full reconstruction.
[0012] As a further improvement of the technical solution, the fusion and reconstruction unit further includes a view consistency detection module, an occlusion-aware fusion module and a local geometric update module, wherein: The view consistency detection module performs reprojection error calculation on the same name feature points of multi-source data, and marks and removes contradictory data when the error exceeds the threshold, and retains valid data with consistent view angles; The occlusion-aware fusion module constructs a scene depth map based on depth sensor data, identifies occlusion areas by comparing depth values under different views, uses weighted fusion for non-occlusion areas, and prioritizes high-confidence sensor data for occlusion areas to generate conflict-free fusion results; The local geometry update module locates the model update area for the fused data: for the newly added scene area, a new mesh surface or a filled voxel is generated through triangulation; for the changed part of the already reconstructed area, the original geometry information is deleted and replaced with the geometry structure corresponding to the new fused data, and the model version identifier is updated synchronously to ensure the timeliness and accuracy of the three-dimensional model.
[0013] As a further improvement of the technical solution, the real-time updating and consistency maintaining unit comprises a spatial index construction module, an incremental change detection module, and a model difference merging and version control module, wherein: The spatial index construction module receives the mesh or voxel form three-dimensional model output by the fusion and reconstruction unit, simultaneously receives the updated area spatial range, new geometry features, and updated model version identifier output by the local geometry update module, constructs an R-tree index according to the scene space division basic unit, and synchronously updates the index item of the corresponding spatial unit for the updated area marked by the local geometry update module; The incremental change detection module receives the new data with confidence annotation output by the preprocessing and calibration unit, locates the corresponding model spatial unit through the R-tree index, and compares the feature hash values of the new data and the existing model; when the difference degree exceeds the threshold value (the threshold value is related to the matching confidence of the preprocessing and calibration unit), the changed area is marked and the geometry difference is extracted; The model difference merging and version control module merges the changed area: non-conflicting changes directly replace the model information, and conflicting changes preferentially retain the data according to the "collection time + matching confidence"; after merging, the R-tree index version number and hash value are updated, and a change log containing the coordinates, type, timestamp, and version number of the changed area is generated.
[0014] As a further improvement of the technical solution, the visualization and verification unit comprises a multi-view rendering module, a consistency index calculation module, a calibration scene comparison test module, and a result output module, wherein: The multi-view rendering module provides multi-angle view rendering based on the mesh or voxel form three-dimensional model output by the fusion and reconstruction unit; The consistency index calculation module calculates two types of quantitative indexes, namely geometry precision and texture consistency, according to the mesh or voxel form three-dimensional model output by the fusion and reconstruction unit, and the index calculation process is associated with the confidence index output by the fusion and reconstruction unit; The calibration scene comparison test module compares the mesh or voxel form three-dimensional model with the preset calibration scene, and calculates the deviation between the model and the calibration scene; The result output module is used to output the visualization result, technical performance index report, and quantitative data.
[0015] The second object of the present application is to provide a digital twin real scene construction method based on multi-source video image fusion. S100, multi-source perception data acquisition and standardization: integrate unified time stamp, interface protocol and preliminary compression packet mechanism, generate unified time stamp through high-precision clock synchronization, adapt to the transmission needs of multiple sensors, compress image data and remove abnormal values of inertial sensors, cooperatively control the acquisition time through hardware triggering and software scheduling, integrate data according to "frame number-sensor ID-time stamp" to form associated data stream, and check the integrity during transmission; S200, multi-source data stream timing synchronization: the data stream of S100 is calibrated by using sliding window and least square method, the clock drift is estimated by chi-square test, the data is matched based on time stamp and event identification and aligned based on sensor identification, and after integration, a synchronous multi-modal frame sequence is formed, and the integrity is checked during transmission; S300, multi-modal data preprocessing and calibration: noise suppression, distortion correction, internal and external parameter estimation and spatial matching are performed on the frame sequence of S200, and calibration data with confidence annotation are generated; S400, multi-source data fusion and three-dimensional reconstruction: based on the confidence of S300, the sensor data is weighted and fused, and the multi-resolution data is organized by octree according to the scene complexity; contradictory data with projection error exceeding the threshold are removed, occluded areas are identified by depth map and high-confidence data are preferentially retained; the octree is dynamically expanded, and three-dimensional models in the form of grids and voxels, versions and confidence indicators are generated and updated; S500, model real-time updating and consistency maintenance: based on the three-dimensional model and update information of S400, an R-tree index is constructed; the model unit is located by R-tree, the feature hash values of new data and existing model are compared, and if the difference exceeds the threshold, the change is marked; the differences are combined according to "acquisition time + confidence", the R-tree and version are updated, and the change log is generated; S600, model visualization and verification: multi-view rendering is performed on the three-dimensional model of S400, and the confidence is combined to calculate the geometric accuracy and texture consistency indicators; the statistical deviation is compared with the preset calibration scene, and the visualization result, performance report and quantitative data are output.
[0016] Compared with the prior art, the present application has the following advantages: 1. The present application constructs a data acquisition unit containing unified time stamp standard, interface protocol and cooperative control mechanism, and cooperates with a multi-source synchronization unit based on sliding window, least square method and clock drift estimation algorithm, effectively solving the problem of asynchronous multi-source perception data acquisition and heterogeneous format, and being able to output multi-modal data stream with consistent timing reference and association, laying a stable data foundation for subsequent digital twin model construction; 2. The present application relies on classified noise suppression, refined lens distortion correction, internal and external parameter accurate estimation and spatial matching technology with confidence, which improves the quality fluctuation and geometric deviation of multi-modal data, generates input data with consistent calibration and clear confidence, improves the geometric and photometric consistency of multi-source data, and reduces the error risk of subsequent fusion reconstruction; 3. The present application solves the problems of multi-source data fusion conflict, model geometric distortion and low storage efficiency by using hierarchical confidence weighted multi-source data fusion method, combining with view consistency detection, occlusion perception fusion and multi-resolution representation fusion reconstruction unit, which can generate complete detail and geometric conflict-free grid / voxel three-dimensional model to meet the geometric restoration requirements of digital twin scene; 4. The present application avoids the redundant calculation of traditional full reconstruction through real-time updating unit based on spatial index construction, incremental change detection and model difference merging and version control, effectively solves the problems of low update efficiency and difficult global consistency of digital twin model, and can quickly integrate new data into existing model and generate change log to realize dynamic synchronization of model and real scene; 5. The present application makes up for the defect of lack of quantitative evaluation of digital twin model through visualization and verification unit containing multi-view rendering, geometric precision and texture consistency index calculation and calibration scene comparison test, which can quantify the degree of coincidence between model and physical scene from core dimension and provide reproducible verification basis for model reliability. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 The system framework of the present application is shown in the figure; The meanings of various labels in the figure are as follows: 100, data acquisition unit; 110, time synchronization module; 120, interface adaptation module; 130, data preprocessing module; 140, cooperative control module; 150, data stream output module; 200, multi-source synchronization unit; 210, time sequence calibration and drift estimation module; 220, event matching and identification alignment module; 230, synchronized frame sequence output module; 300, preprocessing and calibration unit; 310, noise suppression module; 320, distortion correction module; 330, internal and external parameter estimation module; 340, spatial matching module; 400, fusion and reconstruction unit; 410, hierarchical confidence weighting module; 420, multi-resolution representation module; 430, geometric incremental expansion module; 440, view consistency detection module; 450, occlusion perception fusion module; 460, local geometric update module; 500, real-time updating and consistency maintenance unit; 510, spatial index construction module; 520, incremental change detection module; 530, model difference merging and version control module; 600. Visualization and Verification Unit; 610. Multi-view Rendering Module; 620. Consistency Index Calculation Module; 630. Calibration Scene Comparison Test Module; 640. Result Output Module. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0019] like Figure 1 As shown, this embodiment provides a digital twin reality construction system based on multi-source video image fusion, including: The data acquisition unit 100 is used to systematically acquire and standardize the output of multi-source sensing data. It integrates a unified timestamp standard, interface protocol, preliminary compression and packetization mechanism to achieve collaborative acquisition of camera, depth sensor, inertial sensor and position information, and outputs a correlated data stream downstream. In this embodiment, the data acquisition unit 100 includes a time synchronization module 110, an interface adaptation module 120, a data preprocessing module 130, a collaborative control module 140, and a data stream output module 150, wherein: The time synchronization module 110 integrates a high-precision clock source and a network time synchronization protocol to generate a unified timestamp for the data collected by the camera, depth sensor, inertial sensor and location information, and achieves time reference alignment of multi-source sensing data through a periodic clock calibration mechanism. Specifically, the high-precision clock source for the time synchronization module 110 can be a temperature-compensated crystal oscillator, and the network time synchronization protocol preferentially adopts IEEE 1588PTPv2, achieving time reference alignment of multi-source data through hardware timestamps. The initial calibration period of the periodic clock calibration mechanism can be set according to scenario requirements. During the calibration process, the time deviation of multi-source data is continuously monitored: if the deviation of multiple consecutive calibrations is within a preset reasonable range, the calibration period can be appropriately extended to reduce resource consumption; if the deviation exceeds the preset range, the calibration period is shortened to improve synchronization stability. A timestamp interpolation algorithm is used during the calibration process to ensure the continuity of data flow within the calibration interval and avoid data gaps caused by synchronization adjustments.
[0020] The interface adapter module 120 is configured with a layered communication protocol stack, which supports data transmission protocols for multiple types of sensors and control commands. It adapts to the data transmission requirements of depth sensors, cameras, location information and sensor control commands, and realizes protocol conversion and data access for heterogeneous sensors. Specifically, the layered communication protocol stack is configured in layers according to functions: the physical layer supports common interface types such as USB3.0, GigEVision, RS485, etc., wherein camera data is preferentially transmitted through the GigEVision interface (adapted to the demand for high-bandwidth image data), the depth sensor is accessed through the USB3.0 interface (balance between transmission rate and hardware compatibility), and the position information module is transmitted through the RS485 interface (adapted to long-distance and low-power consumption scenarios); the data link layer adopts a CRC check mechanism to preliminarily check the integrity of the transmitted data; and the application layer defines special data frame structures for different sensors, for example, a camera data frame includes device parameter fields such as exposure time and gain, a depth sensor data frame includes a data validity mask field, and a sensor control instruction frame encapsulates start-stop control and parameter adjustment operation instructions in JSON format, ensuring the readability and compatibility of the instruction transmission.
[0021] The data preprocessing module 130 adopts a lossless compression algorithm to compress image-type raw data, removes outliers in inertial sensor data through statistical filtering criteria, and standardizes the processed data in a preset format; Specifically, the lossless compression algorithm is designed differently according to data types: image-type raw data adopts a PNG compression algorithm based on predictive coding to reduce storage occupancy while preserving the original pixel distribution characteristics; depth data adopts run-length encoding to compress spatial redundancy information; and inertial sensor data has a small data volume, so the storage structure is optimized by rearranging the data format to reduce storage fragmentation. During the outlier removal process of inertial sensor data, the data is first subjected to sliding window mean calculation, then data deviating from the window mean by more than a preset range is marked as an outlier, and the outlier is replaced through interpolation of the previous and subsequent frames, while the time node of the occurrence of the outlier and the replacement method are recorded, providing a basis for subsequent data checking and problem tracing.
[0022] The cooperative control module 140 synchronously controls the camera exposure time, depth sensor sampling period, inertial sensor data reading frequency, and position information refresh timing through a cooperative mechanism of hardware trigger signals and software scheduling instructions, ensuring the frame-level acquisition synchronization of multi-source perception data; Specifically, the hardware trigger signal is generated by the FPGA and outputs a fixed-width 3.3V TTL level pulse, which is connected to the camera external trigger interface, the depth sensor synchronization pin, respectively, to realize the acquisition time synchronization at the hardware level; the software scheduling instruction is based on the priority scheduling mechanism of the real-time operating system, wherein the camera exposure control task is set to the highest priority to ensure that the camera exposure time strictly corresponds to the hardware trigger pulse, and the inertial sensor data reading task adopts periodic scheduling, and the scheduling period and the hardware trigger frequency maintain an integer multiple relationship, avoiding conflicts between the sampling operations of different sensors, and ensuring the frame-level acquisition synchronization of multi-source data.
[0023] The data stream output module 150 constructs a shared memory-based ring buffer structure, integrates the time stamp generated by the time synchronization module 110, the raw data accessed by the interface adaptation module 120, and the data packet encapsulated by the data preprocessing module 130 according to the "frame number-sensor ID-time stamp" association rule, forms an associated data stream, and transmits the associated data stream to the multi-source synchronization unit 200. A data integrity checking mechanism is used in the transmission process.
[0024] Specifically, the size of the ring buffer structure is configured according to the maximum data amount of a single frame and the peak processing demand of the system, and sufficient space is reserved to cope with sudden data amount growth and avoid buffer overflow. During data integration, the "frame number" is a 64-bit self-incrementing integer, which is initialized to 0 after the system is powered on and automatically incremented by 1 after each frame of data encapsulation is completed. The "sensor ID" is an 8-bit enumeration value, which assigns a unique identifier to different types of sensors (such as different enumeration values for cameras, depth sensors, and inertial sensors). The "time stamp" uses the UTC time combined with a microsecond-level offset format to ensure the uniqueness and readability of the time information. The data integrity checking mechanism adds a 32-bit CRC check value at the end of each data packet. The receiving end (multi-source synchronization unit 200) performs CRC checking after receiving the data. If the checking fails, a retransmission request is sent through feedback instructions. If the retransmission fails after reaching the preset upper limit, the frame data is marked as invalid and skipped, avoiding invalid data from blocking the transmission of subsequent normal data streams.
[0025] The multi-source synchronization unit 200 is used to maintain the timing consistency of multiple asynchronous data streams, and combines timing calibration, event-level matching, and clock drift estimation algorithms to complete the time alignment and sensor identification alignment of multi-source data, output a synchronized multi-modal frame sequence, and ensure the timing accuracy of subsequent data processing. In this embodiment, the multi-source synchronization unit 200 includes a timing calibration and drift estimation module 210, an event matching and identification alignment module 220, and a synchronized frame sequence output module 230. The timing calibration and drift estimation module 210 performs timing calibration through a sliding window and least squares method, and estimates clock drift through chi-square test, outputs the calibrated time stamp and clock drift calibration factor, and realizes the consistency of the time reference of multi-source data. Specifically, the size of the sliding window is determined according to the sampling period of the multi-source sensor: if the sampling period difference of the sensors is small, the window size is set to 5-10 frames to ensure that the window contains enough data points to reflect the timing trend; if the sampling period difference is large, the window is set according to the sampling period of the high-frequency sensor, and the low-frequency data is interpolated to complete the window, ensuring that the data amount of each sensor in the window is matched.
[0026] Specifically, the specific implementation logic of the least squares timing calibration is: taking the timestamp of a high-frequency sensor (such as an inertial sensor) as a reference, calculating the deviation value of the timestamp of other sensors in the window from the reference timestamp, linearly fitting the deviation value to obtain a linear equation of “sensor type-time deviation”, and correcting the timestamp of each sensor based on the equation to realize the timing alignment of multi-source data in the window.
[0027] Specifically, in the clock drift estimation process of the chi-square test, the calibration deviation values of the continuous three sliding windows are first counted, the chi-square statistic of the theoretical distribution (assuming normal distribution) and the actual distribution of the deviation values is calculated: if the statistic is less than the preset threshold, it is determined that there is no significant clock drift, and the current calibration period is maintained; if the statistic exceeds the threshold, it is determined that there is a drift, the rate of change of the deviation value with time is fitted by linear fitting, a clock drift calibration factor (the factor is linearly related to the drift rate) is generated, and the calibration factor is fed back to the time synchronization module 110 in real time, which is used for time reference correction in the subsequent data acquisition stage to avoid drift accumulation.
[0028] The event matching and identification alignment module 220 identifies and matches the data of the same event in different sensors based on the timestamp and event identification, generates event-level association information, and aligns the data of different sensors by comparing the sensor identity and the timestamp, to ensure the consistency of multi-source data at the event level and the sensor identification, and to provide statistical information of the alignment error; Specifically, the event identification is defined based on physical scene features or sensor data features: for example, in an industrial production line scene, “device start-stop event” corresponds to the light-dark change of the device outline in the camera image, the vibration mutation of the inertial sensor, and the stillness / movement switch of the position information; in the smart city scene, “vehicle passing event” corresponds to the appearance of the vehicle outline in the camera image and the distance mutation of the depth sensor. The event identification is triggered by a preset feature threshold, such as when the gray scale change of the camera image exceeds 50% or the vibration acceleration of the inertial sensor exceeds 2 m / s², the event trigger point is marked.
[0029] Specifically, in the timestamp matching process, a certain reference sensor (such as a camera, because the image data event features are the most intuitive) corresponding to the event trigger point is taken as the center, a time matching window (the window size is determined according to the event duration of the scene, such as 100-500 ms for the device start-stop event) is set, and the timestamp data of other sensors in the window is retrieved. The data falling within the window is determined as the associated data of the same event; if a sensor has no data in the window, the supplementary data corresponding to the event is generated by interpolating the front and rear frame data to ensure the integrity of the event data.
[0030] Specifically, the sensor identification alignment step adopts a unified 8-bit enumerated sensor ID (consistent with the sensor ID definition of the data acquisition unit 100, such as 0x01 for a camera and 0x02 for a depth sensor), and a "identification check bit" is added to the data packet header field. The check bit is calculated by splicing the sensor ID and the last four bits of the device serial number. The receiving end confirms the uniqueness and correctness of the sensor identification by comparing the check bit with the preset sensor information table. The statistical information of the alignment error includes the "event matching success rate" (the number of successfully matched events / the total number of events), the "identification check failure times", and the "timestamp deviation range of each sensor". The statistical results are stored at a minute level for subsequent fault troubleshooting and parameter optimization reference.
[0031] The synchronous frame sequence output module 230 integrates the multi-source data after time sequence calibration, event matching, and identification alignment according to the "frame number-sensor ID-timestamp" association rule, forms a synchronous multi-modal frame sequence, and outputs it to the preprocessing and calibration unit 300. A data integrity check mechanism is used in the transmission process to ensure the accuracy and integrity of the data.
[0032] Specifically, the data integration logic of the synchronous frame sequence output module 230 needs to be compatible with the data acquisition unit 100. The "frame number" inherits the 64-bit incremental integer sequence of the data acquisition unit 100 to ensure that the frame numbers of the same original data frame are consistent in the acquisition and synchronization steps. The "sensor ID" follows the 8-bit enumeration value to avoid identification confusion. The "timestamp" uses the calibrated UTC time plus the microsecond offset format, which is completely synchronized with the calibrated timestamp output by the time sequence calibration module.
[0033] Further, when encapsulating data, the synchronous frame is organized in the structure of "frame header-sensor data block-check field". The frame header includes the frame number, the sensor ID list, and the total data length. The sensor data block is arranged in ascending order of sensor ID, and each data block includes the calibrated timestamp, the original data (or interpolated supplementary data), and the event identification of the sensor. The check field uses 32-bit CRC check (consistent with the check algorithm of the data acquisition unit 100), covering the entire content of the frame header and the sensor data block.
[0034] Specifically, after receiving the synchronous frame, the receiving end (preprocessing and calibration unit 300) first checks the CRC value. If the check fails, it immediately sends a retransmission request to the synchronous frame sequence output module 230. The retransmission request includes the frame number and sensor ID of the invalid frame. The sending end only retransmits the corresponding invalid data block to avoid resource waste caused by full-frame retransmission. If the same frame data is retransmitted more than three times and still fails to check, the frame is marked as "invalid frame" and skipped, and the invalid frame information (frame number, invalid reason, and occurrence time) is recorded. When the system is idle, the time sequence calibration parameters are re-optimized to reduce the probability of subsequent invalid frame generation.
[0035] The preprocessing and calibration unit 300 is used for quality optimization and geometric photometric consistency guarantee of multi-modal data, adopts noise suppression, distortion correction, initial estimation of internal and external parameters, and spatial matching technology based on improved SIFT+FLANN algorithm to generate input data with consistent calibration and confidence annotation; In the embodiment, the preprocessing and calibration unit 300 includes a noise suppression module 310 and a distortion correction module 320, wherein: The noise suppression module 310 is used for noise suppression of multi-modal data, extracts dynamic noise features by Kalman filtering algorithm , quantifies random noise based on the time series fluctuation law of the "predicted-actual" residual; extracts trend noise features by sliding window adaptive filtering , adjusts the filtering strength according to the smoothness of the data in the window (enhances filtering when the variance in the window is greater than the threshold); extracts abnormal noise features by multi-source data joint verification , identifies isolated noise points that deviate from the consistency of multi-source data; Specifically, the noise suppression module 310 respectively adopts Kalman filtering, sliding window adaptive filtering, and multi-source data joint verification processing for dynamic noise, trend noise, and abnormal noise in multi-modal data, and the specific process is as follows: Dynamic noise processing: first, initialize the state parameters of the sensor (such as the initial speed and acceleration of the inertial sensor) and the process noise covariance matrix (preset according to the sensor type, such as the shutter shake characteristics of the camera value reference), and the observation noise covariance matrix (determined according to the sensor factory accuracy); then process each frame of data in the loop according to the process of "state prediction (estimate the current state according to the state at the last time, output ) → covariance prediction (update the uncertainty of state prediction, output ) → Kalman gain calculation (balance the weight of prediction and observation, output ) → state update (correct the predicted state with the measured value , output the optimal state ) → covariance update (update the state uncertainty, output )"; finally, statistics the "predicted-actual" residual of continuous m frames (usually take 30 frames, ensure to cover the time series fluctuation), quantifies the random noise intensity through the time series fluctuation law of the residual, generates dynamic noise feature value , which is used for subsequent judgment of the severity of dynamic interference in the data; Trend noise processing: first set the sliding window size according to the sensor sampling rate (such as 10Hz sensor set window for 10 frames, corresponding to 1 second data, to ensure that the data trend change can be captured); when the window is sliding frame by frame, the mean value (reflecting the overall trend of the data in the window) and the variance (reflecting the degree of data fluctuation) of the data in the window are calculated in real time; if the variance in the window is greater than the preset threshold (set according to the scene, such as position data , set to a value that can identify significant drift), indicating that the data fluctuation is large, the filter strength coefficient is adjusted to 0.8 (enhanced filtering to suppress trend drift), otherwise is adjusted to 0.3 (weak filtering to retain data details); then use to calculate the weighted data of the original data of the current frame and the window mean value, and obtain the filtered data ; finally, compare all original data in the window with the corresponding filtered data, and take the maximum difference as the trend noise feature value , which quantifies the influence of trend bias in the data; Abnormal noise processing: first classify the same type of sensor data collected at the same time (such as the image gray data of all cameras, the acceleration data of all inertial sensors); calculate the mean value (reflecting the consistent trend of multi-source data) and standard deviation (reflecting the normal dispersion range of multi-source data) of the same type of data at the same time; then calculate the deviation of a single sensor data from the mean value (to determine whether the data deviates from the consistent trend of multi-source data); set the abnormal threshold to 3 times the standard deviation (to avoid misjudging normal dispersion data as abnormal); if the deviation of a single data exceeds the threshold, mark the data as an abnormal noise point (corresponding to ), otherwise mark it as normal (corresponding to ); the subsequent data processing link will automatically exclude the abnormal points marked as 1, to avoid the influence of isolated noise points on the consistency of multi-modal data.
[0036] The distortion correction module 320 performs correction based on a lens distortion model, which quantifies the radial deformation of the lens through radial distortion parameters , and quantifies the tangential deviation caused by the tangential deviation of the lens assembly through tangential distortion parameters ; the radial distortion parameters and the tangential distortion parameters are obtained by collecting multi-view images through a chessboard calibration plate, and are obtained through corner point detection and least squares fitting; after obtaining and , combined with the pixel coordinates and the distance from the optical center , the corrected coordinates are calculated through the distortion model to realize image geometric correction and ensure the photometric consistency of multi-source data.
[0037] Specifically, the distortion correction processing specifically includes lens distortion parameter fitting and pixel coordinate correction, and specifically includes: Lens distortion parameter fitting (checkerboard calibration): first, collect multi-view images through a checkerboard calibration board, and then obtain radial distortion parameters and tangential distortion parameters through corner point detection and least square fitting The core formula is as follows: Corner point observation residual error ; Least square optimization target ; Wherein, , is the actual observation coordinate of the corner point (detected from the collected image), , is the theoretical calibration coordinate of the corner point (calculated according to the actual size of the checkerboard); the calculation logic is to collect 10-20 checkerboard images at different angles, detect the corner point coordinates, and then fit the distortion parameters with the target of “minimum residual square sum” to ensure that the parameters can accurately reflect the actual distortion characteristics of the lens.
[0038] Pixel coordinate correction: after obtaining the distortion parameters, the pixel coordinates and the distance from the optical center are combined to calculate the corrected coordinates through a distortion model, and the core formula is as follows: Distance from the optical center ; Radial distortion correction ; ; Tangential distortion correction ; ; Wherein, , is the camera optical center coordinate (usually half of the image resolution), , is the coordinate after radial distortion correction only, , is the final coordinate after radial and tangential distortion correction; the calculation logic is to calculate the distance from the optical center for each pixel of each image, and then complete the radial and tangential correction in sequence to generate a geometric correction image, thereby ensuring the photometric consistency of multi-source data.
[0039] In the embodiment, the preprocessing and calibration unit 300 further includes an intrinsic and extrinsic parameter estimation module 330 and a spatial matching module 340, wherein: The intrinsic and extrinsic parameter estimation module 330 estimates the parameters by matching feature points and using the least square method: through SIFT-based feature point matching and RANSAC algorithm, the intrinsic matrix calibration; external parameters (rotation matrix , translation vector ) are iteratively optimized by minimizing re-projection error Specifically, the parameter estimation process specifically includes intrinsic matrix calibration and extrinsic (rotation matrix , translation vector ) optimization, which is realized by feature point matching combined with least squares method, and specifically includes: intrinsic matrix calibration: calibration is completed by SIFT-based feature point matching and RANSAC algorithm, and the core formula is intrinsic matrix form , and re-projection error ; wherein the homogeneous coordinate re-projection formula is ; in the formula, is a 3x3 camera intrinsic matrix, , is , axial direction focal length (converted from lens focal length and pixel size), is a pixel distortion coefficient (usually set to 0), , is the camera principal point coordinate (consistent with the optical center coordinate), is the re-projection error, , is the actual image coordinate of the feature point, , is the re-projected image coordinate, is a 3x3 rotation matrix (without unit), is a 3x1 translation vector, , , is the world coordinate of the feature point; the calculation logic is to extract SIFT feature points from multi-view images, to remove outliers after RANSAC, and to fit the intrinsic matrix with the goal of "minimum re-projection error" , and to verify whether the error is within an acceptable range after fitting by a new image; extrinsic optimization: iteratively optimized by minimizing re-projection error, and the core formula is the extrinsic optimization target , wherein: ; ; ; in the formula, is the number of feature points used for optimization (integer), is the Reprojection error of each feature point , For the first The actual image coordinates of each feature point , For the first Reprojected coordinates of feature points (unit: pixels). , , Equal to rotation matrix The element (unitless). , , Translation vector elements, , , For the first The world coordinates of each feature point; the calculation logic is a fixed intrinsic parameter matrix. ,initialization For identity matrix, The zero vector is iteratively adjusted using gradient descent. and The process continues until the sum of the errors from three consecutive iterations is less than a preset value, at which point calibrated data is output.
[0040] The spatial matching module 340 generates confidence-labeled input data based on an improved SIFT+FLANN algorithm, specifically including: Scale-invariant feature points are extracted using the SIFT algorithm with adaptive parameter adjustment, based on the illumination intensity of the input image. Texture complexity Dynamically adjust the Gaussian kernel scale of the difference of Gaussians pyramid Difference layer number Enhance the responsiveness and discriminativeness of feature points in low-light and weak-texture scenes; Calculate the Euclidean distance of feature point descriptors using the FLANN algorithm. It also introduces a dynamic threshold based on the distribution of local descriptor similarity. Filter valid matching pairs; For the selected matching pairs, the weighted RANSAC algorithm is used to further eliminate outliers; combined with the intrinsic parameter matrix Rotation matrix and translation vector This maps feature points from multi-source data to a unified coordinate system, based on the number of matching point pairs. Distance variance and feature point stability weights (The matching confidence is calculated by fusing feature point repetition detection rate, response intensity, and descriptor discriminability), and the matching confidence is calculated by combining the proportion of fused matches, distance variance, and average weight. It completes spatial matching of multi-source data and outputs results with confidence labels.
[0041] Specifically, the spatial matching process includes three parts: adaptive SIFT feature point extraction, FLANN matching and outlier removal, and matching confidence calculation. It generates input data with confidence labels based on an improved SIFT+FLANN algorithm, specifically including: Adaptive SIFT Feature Point Extraction: This method extracts scale-invariant feature points using the SIFT algorithm with adaptively adjusted parameters. The core formula is: Gaussian kernel scale ; Difference layer number ; Feature point response value ; in, The pyramid group number (integer). The number of layers within the group (integer). The interval between layers within a group (an integer, fixed at 3). For the image width and height, The scaling factor is the adjacent Gaussian kernel; the calculation logic is based on... and Dynamic adjustment and (Increases under low light / weak texture) ,Increase By using nonmaximum suppression to filter feature points, the responsiveness and discriminativeness of feature points in low-light and weak-texture scenes are enhanced.
[0042] FLANN matching and outlier removal: The FLANN algorithm is used to calculate the Euclidean distance of feature point descriptors and filter matching pairs. Then, the weighted RANSAC algorithm is used to remove outliers. The core formula is: Euclidean distance ; Dynamic threshold ; Weighted RANSAC weights ; in, The Euclidean distance for a 128-dimensional SIFT descriptor. , For feature points , The Dimensional descriptor (0-1 normalized). To match the mean distance, is the distance standard deviation; the calculation logic is to calculate the distance by using the FLANN algorithm, and then calculate the distance standard deviation according to the formula Filter the effective matching pairs, and then remove the outliers by the weighted re-projection error (based on the intrinsic matrix , rotation matrix and translation vector ) and keep the inliers for subsequent matching.
[0043] Matching confidence calculation: map the multi-source data feature points to the unified coordinate system by combining the intrinsic matrix , rotation matrix and translation vector , and calculate the matching confidence, the core formula is: Matching number ratio ; Distance variance ; Average weight ; Matching confidence ; Wherein, is the number of effective matching pairs (integer), is the initial number of matching pairs (integer), is the distance of the i-th effective matching pair, is the average distance of the effective matching pairs, is the average weight (unitless), is the matching confidence (0-1); the calculation logic is to calculate the matching confidence by fusing the three factors according to the weight after counting the effective matching pair parameters. , complete the spatial matching of multi-source data and output the results with confidence annotation.
[0044] The fusion and reconstruction unit 400 is used to construct a digital twin three-dimensional scene model based on multi-source data, generate a three-dimensional model in the form of a grid or a voxel through a hierarchical confidence weighted multi-resolution representation, a view consistency detection, a occlusion perception fusion and an incremental geometry construction process, and output a version and a confidence index to meet the geometric restoration requirements of the digital twin; In this embodiment, the fusion and reconstruction unit 400 includes a hierarchical confidence weighting module 410, a multi-resolution representation module 420 and a geometric incremental expansion module 430, wherein: The hierarchical confidence weighting module 410 assigns dynamic weights (the weight is negatively correlated with the number of matching pairs and the distance variance, and is positively correlated with the sensor accuracy) to different sensor data based on the matching confidence output by the preprocessing and calibration unit 300, performs weighted fusion on the multi-source observation values of the same space position, and generates fusion data with a confidence label; Specifically, the hierarchical confidence weighting module 410 first obtains the matching confidence score output by the preprocessing unit. Number of matching point pairs Distance variance The sensor accuracy level is used to allocate weights based on a "base weight + dynamic adjustment" method: the base weight is determined by the sensor accuracy level (the higher the accuracy, the larger the initial value), and then... Adjustment( (If the weight exceeds the threshold, increase the weight; if the weight falls below the threshold, decrease the weight). Fine-tuning ( (Weight reduction for values exceeding the threshold). Perform weighted fusion on multi-source observations at the same spatial location. ), and append the fused data from each sensor The weighted confidence labels ensure that the results are traceable.
[0045] The multi-resolution representation module 420 divides the resolution levels according to the scene complexity (high resolution is used for the near-field area and low resolution is used for the far-field area), and organizes and merges the data through an octree structure. The high-resolution level retains millimeter-level geometric details, while the low-resolution level compresses redundant information, thereby achieving efficient storage and representation of the 3D model. Specifically, the multi-resolution representation module 420 divides the resolution levels according to "camera-target distance + feature point density": high resolution is used for near-field areas (close distance, high feature point density), and low resolution is used for far-field areas (far distance, low feature point density). The octree organization and fusion process is as follows: with the scene bounding box as the root node of the octree, the near-field area is recursively divided into millimeter-level child nodes (preserving geometric details), while the far-field area is divided only 1-2 times (the child node size is larger, and similar data is merged to compress redundancy). By organizing and fusing data through the octree, the geometric restoration and storage efficiency are balanced.
[0046] The geometric incremental expansion module 430 dynamically expands the octree structure based on the newly added time-series data: for newly acquired scene areas, octree nodes are added at the corresponding resolution level and the fused data is filled; for updated data of reconstructed areas, only the geometric information of the corresponding nodes is updated, and the historical model and incremental data are quickly spliced together through the node hash index, avoiding redundant calculations of full reconstruction.
[0047] Specifically, the geometric incremental expansion module 430 first assigns a unique hash index to each node of the octree (generated by combining "node level + parent node index + child node sequence number" to ensure uniqueness), and establishes a mapping table between the index and the node's spatial range (recording the world coordinate system coordinate intervals covered by the node). After receiving new time-series data, it determines the region type through "coordinate comparison": the world coordinates of the new data (based on the intrinsic parameter matrix of the preprocessing stage) are... External reference / The mapping result is compared with the mapping table. If no node is matched, it is a "new scene area", and if a node is matched, it is "reconstructed area update data". For the new area: determine the node division times according to the multi-resolution level, create a new sub-node and assign an index, obtain the corresponding area fusion data from the hierarchical confidence weighting module, and segment and fill according to the node space range; for the reconstructed area: only load the existing data of the matched node, replace the old data at the corresponding position in the node with the new fusion data (such as updating the depth value of a part of the node, without replacing the entire node), update the original index, only mark the "update timestamp", rely on the index to quickly associate the historical model, and avoid the redundant calculation of full reconstruction.
[0048] In the embodiment, the fusion and reconstruction unit 400 further includes a view consistency detection module 440, an occlusion-aware fusion module 450, and a local geometry update module 460, wherein: The view consistency detection module 440 performs reprojection error calculation on the same name feature points of the multi-source data, and marks and removes the contradictory data when the error exceeds the threshold value, and retains the effective data with consistent views; Specifically, the same name feature point pairs (such as matching feature points of the camera and the depth sensor) of the multi-source data, the intrinsic matrix , the rotation matrix and the translation vector are obtained from the preprocessing and calibration unit 300. The reprojection error calculation process is as follows: taking the feature point world coordinates of a certain sensor (such as the main camera) as the reference, the world coordinates are reprojected to the image coordinate system of another sensor through the , and the Euclidean distance between the reprojected coordinates and the actual observed feature point coordinates of another sensor is calculated, that is, the reprojection error; Specifically, the error threshold is set according to the average matching confidence in the preprocessing stage: if the average (matching reliability is high), the threshold is set to 3-5 pixels; if (matching reliability is low), the threshold is set to 1-2 pixels. The feature point pairs with error exceeding the threshold are marked as "contradictory data", and the corresponding fusion data (such as the spatial fusion value calculated based on the point pair) is deleted synchronously, and the effective data with error meeting the standard is retained to avoid geometric distortion of the three-dimensional model caused by view conflict. At the same time, the setting of 0.7 is based on the industry consensus of feature matching reliability in the field of computer vision.
[0049] The occlusion-aware fusion module 450 constructs a scene depth map based on the depth sensor data, identifies the occlusion area by comparing the depth values under different views, adopts weighted fusion for the non-occlusion area, and prioritizes high-confidence sensor data for the occlusion area to generate a conflict-free fusion result; Specifically, the occlusion perception fusion module 450 constructs a scene depth map based on the preprocessed depth data, and identifies an occlusion area through multi-view depth value comparison (a small depth value is a foreground occlusion area, and a large depth value is a background non-occlusion area). The non-occlusion area adopts hierarchical confidence weighted fusion, and the occlusion area is preferentially retained The highest sensor data (close to the average fusion of the confidence), generates a conflict-free fusion result and adds an occlusion label.
[0050] The local geometry update module 460 locates the model update area for the fused data: for the newly added scene area, a new mesh surface or a filled voxel is generated through triangulation; for the changed part of the already reconstructed area, the original geometry information is deleted and replaced with the geometry structure corresponding to the new fused data, and the model version identifier is updated synchronously to ensure the timeliness and accuracy of the three-dimensional model.
[0051] Specifically, the local geometry update module 460 locates the update area through comparison of new and old data features (the proportion of newly added feature points exceeds a threshold or new feature points appear). The newly added area is triangulated to generate a mesh or fill a voxel; the already reconstructed area only replaces the old geometry data of the corresponding node, and the model version identifier (including the basic version and the update timestamp) is updated synchronously to ensure the timeliness of the model while preserving the continuity of the geometry splicing (such as a mesh without cracks).
[0052] The real-time update and consistency maintenance unit 500 is used for maintaining the time consistency of the digital twin model and the real scene and realizing incremental update, based on the incremental change detection, model difference merging and version control strategy of the spatial index, the new data is integrated into the existing model and the global and local consistency is guaranteed, and the change log is output to realize scene dynamic synchronization; In the embodiment, the real-time update and consistency maintenance unit 500 includes a spatial index construction module 510, an incremental change detection module 520, and a model difference merging and version control module 530, wherein: The spatial index construction module 510 is based on the mesh or voxel form three-dimensional model output by the fusion and reconstruction unit 400, simultaneously receives the update area space range, new geometry features and updated model version identifier output by the local geometry update module 460, divides the basic unit according to the scene space to construct the R-tree index, and synchronously updates the index item of the corresponding spatial unit for the update area marked by the local geometry update module 460; Specifically, the R-tree index construction and update area index synchronous processing specifically includes: The module first determines the basic unit partitioning rules for the R-tree index—combining the resolution of the 3D model output by the fusion and reconstruction unit 400 (such as millimeter level for close-up and centimeter level for distant view) with the actual spatial range of the scene, the scene is divided into uniform spatial basic units. The size of the basic unit is set to "10-20 times the highest resolution of the model" to ensure that a single basic unit can contain sufficient model details while avoiding excessively large units that would reduce indexing efficiency.
[0053] Furthermore, the construction process of the R-tree index is as follows: the overall spatial range of the scene is taken as the root node of the R-tree. The root node contains four core fields: "spatial range, list of child node pointers, model data reference (pointing to the mesh / voxel data of the corresponding basic unit), and current version identifier". The divided basic units are taken as leaf nodes of the R-tree. The spatial range of each leaf node corresponds to a basic unit. The model data reference is directly associated with the 3D model data of the unit output by the fusion and reconstruction unit 400. According to the principle of "minimum spatial range overlap", the leaf nodes are aggregated into non-leaf nodes layer by layer, and finally a complete R-tree index structure is formed. The index file is stored locally or in a distributed database, supporting fast read and write.
[0054] Furthermore, for the updated region (including spatial extent, new geometric features, and updated version identifier) output by the local geometry update module 460, the index synchronization process is as follows: using the "spatial extent search" function of the R-tree index, locate all R-tree nodes (including leaf nodes and non-leaf nodes) that spatially overlap with the updated region; for leaf nodes, directly update their "model data reference" to the model data corresponding to the new geometric features, and synchronously modify the "version identifier" to the version number output by the local geometry update module 460; for non-leaf nodes, recalculate the boundaries of the spatial extents of all their child nodes, update their own "spatial extent" field, and ensure that the spatial extent of the parent node can completely cover the child nodes; after the index update is completed, generate an "index update record" (including the updated node ID, old version number, new version number, and update time) for subsequent version backtracking.
[0055] The incremental change detection module 520 receives new data with confidence labels output by the preprocessing and calibration unit 300, locates the corresponding model spatial unit through R-tree indexing, and compares the feature hash value of the new data with that of the existing model; when the difference exceeds the threshold (the threshold is associated with the matching confidence of the preprocessing and calibration unit 300), the changed area is marked and the geometric difference is extracted. Specifically, the incremental change detection processing of the incremental change detection module 520 includes: R-tree localization: Receives the output of the preprocessing unit with confidence level. For new data, extract its world coordinate range and locate the corresponding leaf node (i.e., the model spatial unit affected by the new data) through R-tree "spatial range matching" to avoid traversing the entire model.
[0056] Feature Hashing: If the unit is a mesh model, extract the vertex coordinates (keep 3 decimal places), triangle index, and normal vector of the existing model and new data, and calculate the hash value using SHA-256. If it is a voxel model, extract the voxel value and confidence, and calculate the hash value using MD5 (adapt to the amount of voxel data).
[0057] Difference judgment and change marking: Calculate the ratio of the Hamming distance (binary difference bit) of the hash value to the total number of bits to get the difference degree; the threshold value is associated When the threshold value is 0.1-0.2, When the threshold value is 0.2-0.3, When the threshold value is 0.3-0.4. If the difference degree exceeds the threshold value, mark it as a changed area, and extract the geometric difference (addition / deletion of vertices / triangles of mesh, numerical change position of voxel).
[0058] Model difference merging and version control module 530 merges the changed area: non-conflicting changes directly replace the model information, and conflicting changes are prioritized to retain data according to "collection time + matching confidence"; after merging, update the R-tree index version number and hash value, and generate a change log containing the changed area coordinates, type, timestamp, and version number.
[0059] Specifically, the difference merging and version control processing of the model difference merging and version control module 530 includes: conflict judgment and processing: non-conflicting changes are directly replaced by new data; Conflicting changes have overlaps, and are first compared by collection time (new data is prioritized), and if the time difference is less than 5 seconds, they are compared (high confidence is prioritized), and the conflict processing result is recorded.
[0060] Index and version update: after merging, update the "model data reference" and version identifier (format "main version number. Secondary version number", increment the secondary version number by 1 each time) of the corresponding R-tree leaf node, and recalculate the node model hash value to overwrite the old value.
[0061] Change log generation: the log contains "log ID, changed area coordinate range, change type (addition / modification / deletion), new data timestamp, old / new version number, conflict processing record, log timestamp", stored in JSON, supporting model backtracking and tracing.
[0062] Visualization and verification unit 600 is used for reproducibility verification and quantitative effect evaluation of digital twin model, combined with multi-view rendering, consistency index calculation and comparison test function with calibration scene, output visualization result, technical performance index report and quantitative data.
[0063] In the present embodiment, the visualization and verification unit 600 comprises a multi-view rendering module 610, a consistency index calculation module 620, a calibration scene contrast test module 630 and a result output module 640, wherein: The multi-view rendering module 610 provides multi-angle view rendering based on the grid or voxel form three-dimensional model output by the fusion and reconstruction unit 400; Specifically, the multi-view rendering module 610 first receives the grid or voxel three-dimensional model output by the fusion and reconstruction unit 400, and in the first step, the model core data is parsed: for the grid model, the vertex coordinates, triangular face index and texture mapping relationship are extracted, and for the voxel model, the voxel space coordinates, color value and transparency parameter are extracted, and then these data are imported into the rendering buffer (including vertex buffer, index buffer and texture map buffer) to provide data support for subsequent rendering.
[0064] The consistency index calculation module 620 calculates two types of quantitative indexes of geometric accuracy and texture consistency according to the grid or voxel form three-dimensional model output by the fusion and reconstruction unit 400, and the index calculation process is associated with the confidence index output by the fusion and reconstruction unit 400; Geometric accuracy index calculation: For the grid model: first, the vertices with a confidence of ≥0.5 are selected (low confidence data is filtered to avoid error interference), the reference true value is taken from the feature points or high-precision laser scanning data preset in the calibration scene, the three-dimensional Euclidean distance between each selected vertex and the corresponding reference point (single vertex error) is calculated, and then the overall average geometric error is obtained by weighting and summing the "vertex error x vertex confidence" and dividing by the total confidence; For the voxel model: the effective voxels with a confidence of ≥0.5 are also selected, the number of voxels with inconsistent voxel values (such as occupancy state, density) and reference true values is counted, and the voxel deviation rate is obtained by "inconsistent voxel number ÷ total effective voxel number", reflecting the geometric restoration accuracy at the voxel level.
[0065] Texture consistency index calculation: Color deviation: the pixel area corresponding to the reference image (preset high-definition real photo) of the model texture map is extracted, the RGB three-color difference value of the corresponding pixels is calculated and the absolute value average (single area color deviation) is taken, and then the overall average deviation is weighted by the texture confidence, and the texture area with a confidence of <0.5 is filtered; Texture integrity: the area of the region covered with effective texture (confidence ≥0.5) on the model surface is counted, and the texture integrity proportion is obtained by dividing the total area of the model surface, reflecting the overall degree of texture coverage.
[0066] The calibration scene contrast test module 630 compares the grid or voxel form three-dimensional model with the preset calibration scene, and counts the deviation between the model and the calibration scene; Specifically, the calibration scene comparison test module 630 first calls preset calibration scene data, which includes a standard geometric body (such as a cube with a known side length, a sphere with a known radius, or a plane calibration plate with a known grid spacing) with known accurate parameters and a texture calibration object (such as a checkerboard or a two-dimensional code pattern, with the size and color parameters recorded in advance), and then completes the comparison between the model and the calibration scene through three steps: Feature matching and spatial alignment: key features (such as cube vertices, calibration plate corner points, and texture pattern feature points) corresponding to the calibration scene are extracted from the three-dimensional model, the ICP (Iterative Closest Point) algorithm is used to match the model feature points with the calibration scene reference feature points, and the spatial transformation relationship (rotation matrix and translation vector) between the model and the calibration scene is calculated to ensure that they are compared in the same coordinate system; Deviation statistics and analysis: Geometric position deviation: the three-dimensional position deviation (Euclidean distance between the coordinates of the model feature points and the coordinates of the corresponding points in the calibration scene) of the matched feature points is calculated, and the average deviation, maximum deviation, and deviation distribution range (such as the interval of 90% deviation values) are statistically analyzed; Size deviation: for standard geometric bodies (such as cube side length and sphere diameter) in the calibration scene, the corresponding size (model measurement value) is calculated through the model feature points, and the relative size deviation is obtained by " | model measurement value - scene actual value | ÷ scene actual value ", which reflects the model size restoration accuracy; Consistency judgment and visualization: compare the statistical deviation with the preset application scene allowed threshold (such as the industrial detection scene threshold is stricter than the urban market scene), mark the areas with deviation exceeding the threshold (such as the size deviation of a certain edge of the model is too large), and generate a deviation heat map - use color gradient (such as blue to red representing deviation from small to large) to visually display the deviation degree of each area of the model, which facilitates quick positioning of problem parts.
[0067] The result output module 640 is used to output the visualization results, technical performance index report, and quantitative data.
[0068] The embodiment also provides a digital twin real scene construction method based on multi-source video image fusion, and a digital twin real scene construction system based on multi-source video image fusion. S200, multi-source data stream timing synchronization: adopt sliding window + least square method to calibrate the data stream of S100, estimate clock drift through chi-square test, match data based on timestamp and event identification and align sensor identification, integrate to form a synchronized multi-modal frame sequence, and check integrity during transmission; S300, multi-modal data preprocessing and calibration: perform noise suppression, distortion correction, internal and external parameter estimation and spatial matching on the frame sequence of S200, and generate calibration data with confidence annotation; S400, multi-source data fusion and three-dimensional reconstruction: assign weights to sensor data based on the confidence of S300 and perform weighted fusion, organize multi-resolution data in octree according to scene complexity; eliminate contradictory data with projection error exceeding threshold, identify occlusion area through depth map and preferentially retain high-confidence data; dynamically expand octree to generate, update and obtain version and confidence index of three-dimensional model in grid and voxel form; S500, model real-time updating and consistency maintenance: construct R-tree index based on the three-dimensional model and update information of S400; locate model units through R-tree, compare feature hash values of new data of S300 and existing model, and mark changes if the difference exceeds threshold; merge differences according to "acquisition time + confidence", update R-tree and version, and generate change log; S600, model visualization and verification: perform multi-view rendering on the three-dimensional model of S400, calculate geometric precision and texture consistency index in combination with confidence; compare and statistically analyze deviations with preset calibration scene, and output visualization results, performance report and quantitative data.
[0069] Those of ordinary skill in the art can understand that the processes for implementing all or part of the steps of the above embodiments can be completed by hardware, or by programs instructing relevant hardware to complete.
[0070] The basic principles, main features and advantages of the present application are shown and described above. Those skilled in the art should understand that the present application is not limited by the above embodiments, and the above embodiments and descriptions in the specification are only preferred examples of the present application and are not intended to limit the present application. Without departing from the spirit and scope of the present application, various changes and improvements can be made to the present application, and these changes and improvements all fall within the scope of the claimed present application. The scope of protection of the present application is defined by the appended claims and their equivalents.
Claims
1. A digital twin reality construction system based on multi-source video image fusion, characterized in that, include: The data acquisition unit (100) is used to collect and standardize multi-source sensing data in an orderly manner, integrate a unified timestamp standard, interface protocol, preliminary compression and packetization mechanism, realize the collaborative acquisition of camera, depth sensor, inertial sensor and position information, and output a correlated data stream downstream; Multi-source synchronization unit (200) is used to maintain the timing consistency of multiple asynchronous data streams. It combines timing calibration, event-level matching and clock drift estimation algorithms to complete the timing alignment of multi-source data and sensor identification alignment, and outputs a synchronized multimodal frame sequence to ensure the timing accuracy of subsequent data processing. The preprocessing and calibration unit (300) is used to optimize the quality of multimodal data and ensure geometric photometric consistency. It uses noise suppression, distortion correction, initial estimation of intrinsic and extrinsic parameters, and spatial matching technology based on the improved SIFT+FLANN algorithm to generate input data with consistent calibration and confidence labeling. The fusion and reconstruction unit (400) is used to construct a digital twin three-dimensional scene model based on multi-source data. Through hierarchical confidence weighted multi-resolution representation, view consistency detection, occlusion perception fusion and incremental geometry construction process, it generates a three-dimensional model in the form of mesh and voxels, and outputs version and confidence index to meet the geometric restoration requirements of digital twin. Real-time update and consistency maintenance unit (500) is used to maintain the time consistency between the digital twin model and the real scene and to realize incremental updates. Based on the incremental change detection, model difference merging and version control strategy of spatial index, new data is integrated into the existing model and global and local consistency is ensured. Change logs are output to realize dynamic synchronization of the scene. The visualization and verification unit (600) is used to perform reproducibility verification and quantitative effect evaluation of the digital twin model. It combines multi-view rendering, consistency index calculation and comparison test with the calibration scene to output visualization results, technical performance index reports and quantitative data.
2. The digital twin reality construction system based on multi-source video image fusion according to claim 1, characterized in that, The data acquisition unit (100) includes a time synchronization module (110), an interface adaptation module (120), a data preprocessing module (130), a collaborative control module (140), and a data stream output module (150), wherein: The time synchronization module (110) integrates a high-precision clock source and a network time synchronization protocol to generate a unified timestamp for the data collected by the camera, depth sensor, inertial sensor and location information, and achieves time reference alignment of multi-source sensing data through a periodic clock calibration mechanism. The interface adaptation module (120) is configured with a layered communication protocol stack, which supports data transmission protocols for multiple types of sensors and control commands, respectively adapts to the data transmission requirements of depth sensors, cameras, location information and sensor control commands, and realizes protocol conversion and data access for heterogeneous sensors. The data preprocessing module (130) uses a lossless compression algorithm to compress the original image data, removes outliers from the inertial sensor data using statistical filtering criteria, and standardizes and packages the processed data according to a preset format. The collaborative control module (140) synchronously controls the camera exposure time, depth sensor sampling period, inertial sensor data reading frequency, and position information refresh timing through a collaborative mechanism of hardware trigger signals and software scheduling instructions. The data stream output module (150) constructs a ring buffer structure based on shared memory, which integrates the timestamp generated by the time synchronization module (110), the raw data accessed by the interface adaptation module (120), and the data packets encapsulated by the data preprocessing module (130) according to the association rule of "frame number-sensor ID-timestamp" to form an associative data stream and transmit it to the multi-source synchronization unit (200). A data integrity verification mechanism is adopted during the transmission process.
3. The digital twin reality construction system based on multi-source video image fusion according to claim 2, characterized in that, The multi-source synchronization unit (200) includes a timing calibration and drift estimation module (210), an event matching and identifier alignment module (220), and a synchronization frame sequence output module (230), wherein: The timing calibration and drift estimation module (210) performs timing calibration through a sliding window and the least squares method, and estimates clock drift through the chi-square test, outputting the calibrated timestamp and clock drift calibration factor to achieve consistency of time reference for multi-source data. The event matching and identification alignment module (220) identifies and matches data of the same event in different sensors based on timestamps and event identifiers, generates event-level association information, and performs identification alignment on data from different sensors by comparing sensor identity identifiers and timestamps, ensuring consistency of multi-source data in terms of event level and sensor identifiers, and providing statistical information on alignment error; The synchronous frame sequence output module (230) integrates the multi-source data after time-series calibration, event matching and identifier alignment according to the association rule of "frame number-sensor ID-time stamp" to form a synchronous multimodal frame sequence, and outputs it to the preprocessing and calibration unit (300). A data integrity verification mechanism is adopted during the transmission process.
4. The digital twin reality construction system based on multi-source video image fusion according to claim 3, characterized in that, The preprocessing and calibration unit (300) includes a noise suppression module (310) and a distortion correction module (320), wherein: The noise suppression module (310) is used to suppress noise in multimodal data and extract dynamic noise features using a Kalman filter algorithm. Random noise is quantified based on the temporal fluctuation patterns of the "prediction-measurement" residuals; trend noise features are extracted through sliding window adaptive filtering. The filtering intensity is adjusted based on the smoothness of the data within the window; abnormal noise features are extracted through joint verification of multi-source data. Identify isolated noise points that deviate from the consistency of multi-source data; The distortion correction module (320) performs correction based on a lens distortion model, and the lens distortion model quantifies the radial deformation law of the lens through radial distortion parameters while quantifying the tangential offset caused by the assembly deviation of the lens through tangential distortion parameters ; the radial distortion parameters and tangential distortion parameters are obtained by collecting multi-view images through a checkerboard calibration plate, followed by corner detection and least squares fitting; after obtaining and , combined with the pixel coordinates and the distance from the optical center , the corrected coordinates are calculated through the distortion model to achieve image geometric correction.
5. The digital twin reality construction system based on multi-source video image fusion according to claim 4, characterized in that, The preprocessing and calibration unit (300) further includes an intrinsic and extrinsic parameter estimation module (330) and a spatial matching module (340), wherein: The intrinsic and extrinsic parameter estimation module (330) uses feature point matching combined with the least squares method to achieve parameter estimation: through SIFT-based feature point matching and the RANSAC algorithm, the intrinsic parameter matrix is completed. The calibration is performed by iteratively optimizing the external parameters, i.e., the rotation matrix, by minimizing the reprojection error. Translation vector Output data that is consistent with the calibration; The spatial matching module (340) generates confidence-labeled input data based on the improved SIFT+FLANN algorithm, specifically including: Scale-invariant feature points are extracted using the SIFT algorithm with adaptive parameter adjustment, based on the illumination intensity of the input image. Texture complexity Dynamically adjust the Gaussian kernel scale of the difference of Gaussians pyramid Difference layer number Enhance the responsiveness and discriminativeness of feature points in low-light and weak-texture scenes; Calculate the Euclidean distance of feature point descriptors using the FLANN algorithm. It also introduces a dynamic threshold based on the distribution of local descriptor similarity. Filter valid matching pairs; For the selected matching pairs, the weighted RANSAC algorithm is used to further eliminate outliers; combined with the intrinsic parameter matrix Rotation matrix and translation vector This maps feature points from multi-source data to a unified coordinate system, based on the number of matching point pairs. Distance variance and feature point stability weights The matching confidence score is calculated by combining the proportion of matching numbers, distance variance, and average weight. It completes spatial matching of multi-source data and outputs results with confidence labels.
6. The digital twin reality construction system based on multi-source video image fusion according to claim 5, characterized in that, The fusion and reconstruction unit (400) includes a hierarchical confidence weighting module (410), a multi-resolution representation module (420), and a geometric increment expansion module (430), wherein: The hierarchical confidence weighting module (410) assigns dynamic weights to different sensor data based on the matching confidence level output by the preprocessing and calibration unit (300), performs weighted fusion on multi-source observations at the same spatial location, and generates fused data with confidence labels. The multi-resolution representation module (420) divides the resolution levels according to the scene complexity, organizes and merges data through an octree structure, retains millimeter-level geometric details in the high-resolution level, and compresses redundant information in the low-resolution level, thereby achieving efficient storage and representation of the three-dimensional model. The geometric incremental expansion module (430) dynamically expands the octree structure based on the newly added time-series data: for newly collected scene areas, octree nodes are added at the corresponding resolution level and fused data is filled; for updated data of reconstructed areas, only the geometric information of the corresponding nodes is updated, and the historical model and incremental data are quickly spliced together through the node hash index.
7. The digital twin reality construction system based on multi-source video image fusion according to claim 6, characterized in that, The fusion and reconstruction unit (400) further includes a view consistency detection module (440), an occlusion perception fusion module (450), and a local geometry update module (460), wherein: The viewpoint consistency detection module (440) performs reprojection error calculation on the same feature points of multi-source data. When the error exceeds the threshold, it is marked as contradictory data and removed, while retaining valid data with consistent viewpoints. The occlusion perception fusion module (450) constructs a scene depth map based on depth sensor data, identifies occluded areas by comparing depth values from different perspectives, performs weighted fusion on non-occluded areas, and prioritizes the retention of high-confidence sensor data for occluded areas to generate conflict-free fusion results. The local geometry update module (460) locates the model update region for the fused data: for newly added scene regions, it generates new mesh surfaces or fill voxels through triangulation; for the changed parts of the reconstructed regions, it deletes the original geometric information and replaces it with the geometric structure corresponding to the new fused data, and updates the model version identifier synchronously.
8. The digital twin reality construction system based on multi-source video image fusion according to claim 7, characterized in that, The real-time update and consistency maintenance unit (500) includes a spatial index construction module (510), an incremental change detection module (520), and a model difference merging and version control module (530), wherein: The spatial index construction module (510) is based on the mesh or voxel form three-dimensional model output by the fusion and reconstruction unit (400), and simultaneously receives the updated region spatial range, new geometric features and updated model version identifier output by the local geometry update module (460). It constructs an R-tree index according to the scene space division of basic units, and synchronously updates the index items of the corresponding spatial units for the updated region marked by the local geometry update module (460). The incremental change detection module (520) receives new data with confidence labels output by the preprocessing and calibration unit (300), locates the corresponding model spatial unit through R-tree indexing, and compares the feature hash value of the new data with that of the existing model; when the difference exceeds the threshold, the changed area is marked and the geometric difference is extracted. The model difference merging and version control module (530) merges the changed areas: non-conflicting changes directly replace the model information, and conflicting changes retain data with priority according to "collection time + matching confidence level"; after merging, the R tree index version number and hash value are updated, and a change log containing the coordinates, type, timestamp and version number of the changed area is generated.
9. The digital twin reality construction system based on multi-source video image fusion according to claim 8, characterized in that, The visualization and verification unit (600) includes a multi-view rendering module (610), a consistency index calculation module (620), a calibration scene comparison test module (630), and a result output module (640), wherein: The multi-view rendering module (610) provides multi-angle view rendering based on the mesh or voxel form of the 3D model output by the fusion and reconstruction unit (400). The consistency index calculation module (620) calculates two types of quantitative indicators, namely geometric accuracy and texture consistency, based on the mesh or voxel form three-dimensional model output by the fusion and reconstruction unit (400), and the index calculation process is associated with the confidence index output by the fusion and reconstruction unit (400). The calibration scene comparison test module (630) compares the three-dimensional model in mesh or voxel form with the preset calibration scene and counts the deviation between the model and the calibration scene; The result output module (640) is used to output visualization results, technical performance index reports and quantitative data.
10. A method for constructing a digital twin reality scene based on multi-source video image fusion, based on the digital twin reality scene construction system based on multi-source video image fusion as described in any one of claims 1-9, characterized in that, Includes the following steps: S100, Multi-source sensing data acquisition and standardization: Integrates unified timestamps, interface protocols and preliminary compression packet mechanisms, generates unified timestamps through high-precision clock synchronization, adapts to the transmission requirements of multiple types of sensors, compresses image data and removes abnormal values from inertial sensors, controls the acquisition timing through hardware triggering and software scheduling, integrates data into a correlated data stream according to "frame number-sensor ID-timestamp", and verifies integrity during transmission; S200, multi-source data stream timing synchronization: The data stream of S100 is calibrated by sliding window + least squares method, clock drift is estimated by chi-square test, data is matched based on timestamp and event identifier and aligned with sensor identifier, and integrated to form a synchronous multimodal frame sequence, and integrity is checked during transmission; S300, Multimodal Data Preprocessing and Calibration: Perform noise suppression, distortion correction, intrinsic and extrinsic parameter estimation and spatial matching on the frame sequence of S200, and generate calibration data with confidence labels; S400, multi-source data fusion and 3D reconstruction: Based on the confidence level of S300, sensor data are weighted and fused, and multi-resolution data are organized using an octree according to scene complexity; contradictory data with reprojection errors exceeding the threshold are removed, occluded areas are identified through depth maps and high-confidence data is retained first; the octree is dynamically expanded to generate and update 3D models and versions in the form of meshes and voxels, as well as confidence indices; S500, Real-time Model Updates and Consistency Maintenance: Construct an R-tree index based on the S400 3D model and update information; Locate model units through the R-tree, compare the feature hash values of new S300 data with existing models, and mark changes if the differences exceed the threshold; Merge differences according to "collection time + confidence level", update the R-tree and version, and generate change logs; S600 Model Visualization and Verification: Perform multi-view rendering on the S400 3D model, calculate geometric accuracy and texture consistency indices based on confidence level; compare the statistical deviation with the preset calibration scene, and output visualization results, performance reports and quantitative data.
Citation Information
Patent Citations
Three-dimensional video fusion method and system based on digital twin technology
CN118864723A
Artificial intelligence video processing method based on digital twinning
CN120388319B