Scene inspection video and three-dimensional model automatic roaming system and method

By deploying a hardware-level time synchronization module and dual-channel transmission technology in the camera equipment, the problem of insufficient spatiotemporal alignment accuracy in automatic roaming was solved, realizing efficient joint roaming of video stream and 3D model, and improving positioning accuracy and system efficiency.

CN121842355APending Publication Date: 2026-04-10GUANGZHOU AEBELL ELECTRICAL TECH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing automatic roaming methods suffer from insufficient spatiotemporal alignment accuracy and significant time synchronization errors between the video stream and the 3D model, leading to phenomena such as asynchrony between the video view and the model's perspective and target positioning deviation during roaming.

Method used

A hardware-level time synchronization module is deployed in the camera equipment to perform lens distortion calibration and sensor parameter calibration. The video stream and equipment status data are transmitted through dual transmission channels. Combined with UDP and TCP protocols, timestamp dynamic compensation is performed, and real-time feature extraction and compression are performed at the device end or edge end. The backend system completes accurate matching and mapping based on the hardware timestamp.

Benefits of technology

It improves the spatiotemporal alignment accuracy, solves the problem of asynchronous video and model perspectives, reduces computational pressure and transmission latency, and enables efficient joint animation between video and 3D model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121842355A_ABST
    Figure CN121842355A_ABST
Patent Text Reader

Abstract

The invention discloses a scene inspection video and three-dimensional model automatic roaming system and method, and belongs to the technical field of inspection automatic roaming, and the automatic roaming method comprises the following steps: deploying a plurality of camera devices in an inspection scene, constructing a virtual camera model corresponding to the camera devices in a three-dimensional model, parameter matching of the virtual camera and the physical camera shooting equipment is completed; synchronously acquiring video stream data and equipment state data of an inspection scene through camera equipment; respectively transmitting the video stream data and the equipment state data by adopting double transmission channels; performing real-time feature extraction and compression on the video stream data at the equipment end or the edge end, and fusing the equipment state data, the calibration parameters and the extracted video feature data; the virtual camera in the three-dimensional model is driven to automatically roam according to the video feature track and the preset path, and the automatic roaming method integrates a multi-device collaborative acquisition mechanism and an edge computing gateway, and meets the requirements of large-scale and complex inspection scenes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of automatic roaming inspection, in particular to a scene inspection video and three-dimensional model automatic roaming system and method. BACKGROUND

[0002] In the scene inspection work of industrial plants, power stations and environmental protection parks, the scene inspection video and three-dimensional model automatic roaming technology can realize the linkage roaming of video real scene and three-dimensional model, improve the inspection efficiency and fault positioning accuracy, reduce the workload of manual inspection, and has wide coverage and fast speed.

[0003] Patent with publication number CN114187414B discloses a roadway three-dimensional roaming inspection method and system, which includes creating a three-dimensional real scene model consistent with the actual underground roadway and equipment distribution; the real-time parameters of the equipment are actively transmitted to the three-dimensional real scene model through specific configuration; the model is displayed and roaming inspection is performed through a display platform, which includes global roaming inspection and in-roadway roaming inspection, and in the global roaming inspection mode, it can be switched to in-roadway roaming inspection of any device or roadway through the operation event of the display platform; when monitoring the actual underground situation, automatic inspection of the underground can be completed through the inspection system, which is simple to operate and safe and reliable.

[0004] Patent with publication number CN112366821A discloses a three-dimensional video intelligent inspection system and method, which includes an AR inspection system including an AR device and an inspection robot; a three-dimensional imaging system including a holographic scanning device, a holographic image device and a data acquisition module; an interactive center system including a control system; using AR interaction technology, it can comprehensively display on-site device pictures, real-time information and alarm information, and can display real-time video tracking of mobile devices such as inspection robots and inspection recorders, and can control mobile video devices for accurate positioning inspection; it can directly operate the three-dimensional model, pull the video automatic tracking device object, realize full-station omnidirectional video roaming, automatically focus on key video inspection points according to the set inspection path and time interval, and realize full-omnidirectional high-quality video inspection of the substation.

[0005] The patent with publication number CN118301291A discloses a video inspection system based on a three-dimensional model, a parameter synchronous debugging method and a control method thereof, which comprises a physical camera, a virtual camera, a three-dimensional display module, a video display module, a three-dimensional model, and a video inspection software. The virtual camera with the same characteristics as the physical camera is set in the three-dimensional model for each camera through the inspection software, forming a mirror camera group. The virtual camera is operated through the inspection software, and the command is sent to the mirror physical camera. The same perspective three-dimensional picture and video picture can be presented in the three-dimensional display module and the video display module, realizing the three-dimensional model traction video linkage function. Through the combination operation of multiple cameras, the three-dimensional model traction video roaming of the whole scene is realized. The operation process of the mirror camera is recorded, the inspection path is automatically planned, the cameras on the path are combined in order, and the three-dimensional model automatically traction video roaming of the planned path is realized according to the record. The existing part of the automatic roaming method has the problem of insufficient space-time alignment accuracy. The time synchronization error of the video stream and the three-dimensional model is large, and the space mapping deviation is obvious, resulting in the phenomenon that the video picture and the model view are out of synchronization in the roaming process, the target positioning deviates, etc.

[0006] In view of the above problems, it is urgent to make innovative design on the basis of the original automatic roaming method. SUMMARY

[0007] The purpose of the present application is to provide a scene inspection video and three-dimensional model automatic roaming system and method to solve the problem of insufficient space-time alignment accuracy of the part of the automatic roaming method in the background art. The time synchronization error of the video stream and the three-dimensional model is large, and the space mapping deviation is obvious, resulting in the phenomenon that the video picture and the model view are out of synchronization in the roaming process, the target positioning deviates, etc.

[0008] In the first aspect, the present application provides a scene inspection video and three-dimensional model automatic roaming method. The automatic roaming method comprises the following steps: Deploy multiple camera equipment with hardware-level time synchronization module in the inspection scene, complete lens distortion calibration and sensor parameter calibration of the camera equipment, solidify the calibration parameters to the device firmware, and at the same time, construct a virtual camera model corresponding to the camera equipment in the three-dimensional model, complete the parameter matching of the virtual camera and the physical camera equipment.

[0009] Synchronously collect video stream data and equipment state data of the inspection scene through the camera equipment. The video stream data is collected at a preset frame rate and embedded with a nanosecond-level hardware timestamp. The equipment state data is collected at a preset frequency and associated with the hardware timestamp of the latest video stream.

[0010] The video stream data and the device state data are transmitted through two transmission channels respectively, the high-bandwidth channel transmits the video stream data using the UDP protocol, and the low-bandwidth channel transmits the device state data and the calibration parameters using the TCP protocol, and the transmission delay is monitored in real time during the transmission process and dynamic timestamp compensation is performed.

[0011] The video stream data is extracted and compressed in real time at the device end or the edge end, the device state data, the calibration parameters and the extracted video feature data are fused, and preliminary space-time alignment is completed.

[0012] The preprocessed fused data is received by the backend system, accurate matching of the video stream data and the device state data is completed based on the hardware timestamp, the video features are mapped to the three-dimensional model coordinate system by combining the point cloud registration algorithm, and space-time alignment of the video and the three-dimensional model is realized.

[0013] Based on the space-time alignment result, the virtual camera in the three-dimensional model automatically roams according to the video feature track and the preset path, synchronously associates the corresponding video stream picture and realizes picture-in-picture superposition or split-screen display, supports one-key switching of the user from the three-dimensional model roaming interface to the corresponding real scene video picture, when the video feature recognizes an abnormal target in inspection, automatically triggers the virtual camera to focus on the abnormal area and marks, synchronously calls the historical inspection video clip of the area for comparison and display, realizes the linkage automatic roaming of the scene inspection video and the three-dimensional model, and helps the staff to quickly locate the abnormal position and trace the abnormal evolution process.

[0014] Preferably, the lens distortion calibration adopts Zhang Zhengyou calibration method, the calibration parameters include an intrinsic matrix, a distortion coefficient and an extrinsic matrix, and the calibration parameters are associated with a timestamp generation module of a hardware-level time synchronization module to ensure accurate matching of the video frames after distortion correction and the timestamp information; the hardware-level time synchronization module includes a PTP slave clock chip and a GPS Beidou dual-mode time service module; the PTP slave clock chip establishes a master-slave synchronization link with a PTP master clock device deployed in the inspection scene, and realizes sub-microsecond-level time synchronization; the GPS Beidou dual-mode time service module is applied to part of the mobile camera devices in the inspection scene without network coverage, and realizes alignment of the time reference with the PTP master clock through satellite time service signals.

[0015] Preferably, a main camera device in the plurality of camera devices integrates a PTP master clock gateway, the main camera device manages the time reference of each slave camera device through a preset master-slave communication protocol; when multiple devices are cooperatively collected, the main camera device sends a time synchronization calibration instruction to the slave camera devices first, and after the time reference of the whole system is unified, a hardware trigger signal is sent to control all camera devices to start data collection at the same time; the main camera device receives timestamp feedback information of each slave camera device in real time, triggers secondary calibration for the devices with time deviation, and ensures that the video stream data and the device state data collected by multiple devices are based on the same time reference.

[0016] Preferably, two independent acquisition threads are arranged in the camera equipment, respectively for video stream data acquisition and device state data acquisition; the video stream data acquisition thread acquires at a frame rate of 25fps-30fps, the device state data acquisition thread acquires at a frequency of 1Hz-10Hz, and the latest hardware timestamp of the video acquisition thread is read through a synchronous triggering mechanism.

[0017] Preferably, the device state data includes running state parameters and spatial pose parameters of the camera equipment, the running state parameters include temperature, shutter speed, anti-shake state and exposure parameters, and the spatial pose parameters include positioning coordinates and attitude angles; the nanosecond-level hardware timestamp adopts a three-tuple structure of UTC time, device unique identifier and data type identifier, wherein the data type identifier is used to distinguish video stream data and device state data, and to ensure time reference tracing and accurate correlation of multi-source data.

[0018] Preferably, the dual transmission channel logic layer adopts dual ports and dual protocols in the same network, channel one is used for transmitting video stream data and adopts UDP protocol, and channel two is used for transmitting device state data and adopts TCP protocol; the timestamp dynamic compensation method is that the camera equipment records the data sending timestamp, the backend records the receiving timestamp, the transmission delay duration is calculated, and the real data acquisition time is restored by subtracting the delay duration from the receiving timestamp.

[0019] Preferably, the video stream data is encoded and transmitted in layers using the SVC scalable coding technology to form layered code streams of base layer and enhancement layer; the base layer carries low-resolution and low-code-rate core video information and is transmitted preferentially using the UDP protocol to ensure real-time performance; the enhancement layer carries high-definition detail information and is dynamically transmitted on demand according to the bandwidth feedback and image quality requirements of the backend system, realizing bandwidth adaptive video stream transmission optimization, and taking into account real-time performance and image quality requirements; the video encoding adopts a frame-level data hard binding mechanism to encapsulate I-frames and device state data and camera equipment external parameter at the corresponding time, to generate time and space aligned data packets with timestamp index, to maintain the highest transmission priority of the data packets through a transmission priority scheduling mechanism, to transmit P-frames and B-frames and video differential data and relative time offset based on I-frames, and to record the relative time offset with nanosecond-level precision to ensure the time correlation of non-key frames and I-frames.

[0020] Preferably, the heterogeneous edge computing chip built into the camera device uses a lightweight deep learning feature extraction network to complete real-time feature extraction of video stream data. The extracted features include semantic features, geometric features, and region of interest coordinates of the inspection target. The original video frames are encoded with a high compression ratio and then stored and archived locally. The extracted feature tensor data, device status data, and calibration parameters are encapsulated in a preset data format and transmitted synchronously, realizing the synergistic optimization of front-end data dimensionality reduction and back-end computational burden reduction. Multiple camera devices construct a distributed data aggregation node through an edge computing gateway, which uses a time-sensitive network protocol to receive multi-source data from each camera device. The edge computing gateway has a built-in lightweight spatiotemporal fusion algorithm, which completes the time synchronization calibration and spatial coordinate unification of multi-device data based on hardware timestamps. After completing the initial spatiotemporal alignment, a standardized fused data frame is generated and then pushed to the back-end system through a bandwidth adaptive transmission strategy, realizing edge-cloud collaborative computing power distribution and data preprocessing optimization.

[0021] Preferably, the precise matching of the video stream data and device status data is specifically achieved by the backend system extracting the triplet structure hardware timestamp of each data packet in the fused data, establishing a one-to-one correspondence between the video stream data and the device status data through timestamp indexing, and eliminating abnormal data with timestamp deviations exceeding the threshold; the point cloud registration algorithm adopts the ICP algorithm and the NDT algorithm, first performing coarse registration between the point cloud data corresponding to the video features and the point cloud data of the 3D model, and then optimizing the spatial mapping deviation through fine registration, finally accurately mapping the video features to the global coordinate system of the 3D model, realizing the dual precise alignment of the video and the 3D model in the time dimension and the spatial dimension.

[0022] Secondly, this application provides a scene inspection video and 3D model automatic roaming system for executing a scene inspection video and 3D model automatic roaming method. The automatic roaming system includes: The equipment deployment initialization module is used for the deployment configuration and parameter initialization of camera equipment in inspection scenarios. It completes lens distortion calibration and sensor parameter calibration, solidifies the calibration parameters into the equipment firmware, and builds a matching 3D virtual camera model.

[0023] The multi-source data synchronous acquisition module is used to realize the synchronous acquisition and timestamp embedding of video stream data and equipment status data in the inspection scenario.

[0024] The data transmission module employs dual transmission channels to separately transmit video streams and device status data, enabling efficient and low-latency transmission of multi-source data and ensuring time synchronization accuracy during data transmission.

[0025] The front-end data processing and fusion module performs data dimensionality reduction, feature extraction, and preliminary spatiotemporal alignment at both the device and edge levels to reduce the computational burden on the back-end.

[0026] A backend space-time alignment module establishes a one-to-one correspondence between the video stream and the device state data based on the triple hardware timestamp, eliminates abnormal data, and is used to realize accurate matching of fused data and mapping and alignment of video features to the three-dimensional model coordinate system.

[0027] An automatic roaming driving interaction module drives the three-dimensional model virtual camera to roam, and is used to realize deep linkage of video and three-dimensional model and user interaction.

[0028] Compared with the prior art, the beneficial effects of the present application are: a multi-device cooperative collection mechanism and an edge computing gateway convergence processing scheme adapt to the needs of large-scale and complex inspection scenes.

[0029] By integrating a hardware-level time synchronization module in the camera equipment and embedding a nanosecond-level hardware timestamp, the time synchronization accuracy is guaranteed from the data collection source, the time drift problem caused by the soft clock is avoided, the time error is greatly shortened, and the spatial distortion error is reduced from the source through lens distortion calibration and sensor parameter calibration.

[0030] Dual transmission channel separation is adopted to transmit video stream data and device state data, the advantages of UDP and TCP protocols are combined, the state data is prevented from being blocked by high-bandwidth video stream, and through a timestamp dynamic compensation mechanism, the transmission delay error caused by network fluctuations is offset, and the real-time and reliability of data transmission are improved.

[0031] Video feature extraction, data fusion and preliminary space-time alignment are completed at the device end or the edge end, the backend computing pressure is greatly reduced, the data transmission amount is reduced, the space-time alignment efficiency and accuracy are further improved, and the problems of virtual-real disconnection and roaming lag in the prior art are effectively solved. BRIEF DESCRIPTION OF DRAWINGS

[0032] Figure 1 The flowchart of the present application.

[0033] Figure 2 The module diagram of the present application. DETAILED DESCRIPTION

[0034] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0035] This application provides a method for automatic roaming of scene inspection video and 3D model. The core of this method is to deploy multiple camera devices with hardware-level time synchronization modules in the inspection scene, complete lens distortion calibration and sensor parameter calibration of the camera devices, and solidify the calibration parameters into the device firmware. Simultaneously, a virtual camera model corresponding to the camera devices is constructed in the 3D model to achieve parameter matching between the virtual camera and the physical camera devices. Video stream data and device status data of the inspection scene are synchronously collected by the camera devices. The video stream data is collected at a preset frame rate and embedded with nanosecond-level hardware timestamps, while the device status data is collected at a preset frequency and associated with the hardware timestamp of the most recent frame of the video stream. Dual transmission channels are used to transmit the video stream data and device status data separately. The high-bandwidth channel uses the UDP protocol to transmit the video stream data, while the low-bandwidth channel uses the TCP protocol to transmit the device status data and calibration parameters. During transmission, transmission delay is monitored in real time and timestamps are dynamically compensated. The method also involves monitoring the transmission delay at the device end or edge end. The video stream data undergoes real-time feature extraction and compression. Device status data, calibration parameters, and extracted video feature data are fused to achieve initial spatiotemporal alignment. The backend system receives the preprocessed fused data, accurately matches the video stream data with the device status data based on hardware timestamps, and maps video features to the 3D model coordinate system using point cloud registration algorithms, achieving spatiotemporal alignment between the video and the 3D model. Based on the spatiotemporal alignment results, the virtual camera in the 3D model automatically roams along the video feature trajectory and preset path, synchronously associating with the corresponding video stream images and achieving picture-in-picture overlay or split-screen display. Users can switch to the corresponding real-scene video image with one click on the 3D model roaming interface. When video features identify an abnormal target during inspection, the virtual camera automatically focuses on and marks the abnormal area, simultaneously retrieving historical inspection video clips of that area for comparison and display. This achieves automatic linkage between the scene inspection video and the 3D model, helping staff quickly locate abnormal locations and trace the evolution of abnormalities.

[0036] Example 1: To better understand the above technical solution, the following will provide a detailed description of the technical solution in conjunction with the accompanying drawings and specific implementation methods. (Refer to...) Figure 1 As shown, this figure is an exemplary flowchart of the scene inspection video and 3D model automatic roaming method according to this embodiment of the application. The automatic roaming method includes the following steps: S1. Inspection equipment deployment and initialization: Deploy multiple camera devices with hardware-level time synchronization modules in the inspection scenario, complete the lens distortion calibration and sensor parameter calibration of the camera devices, solidify the calibration parameters into the device firmware, and at the same time build a virtual camera model corresponding to the camera devices in the 3D model to complete the parameter matching between the virtual camera and the physical camera devices.

[0037] The existing part of the camera equipment mostly uses soft clock to embed time stamp, which is easy to cause millisecond level time drift due to system load fluctuation. When video stream and device state data are mixed and transmitted, state data is easy to be blocked by high bandwidth video stream, resulting in transmission delay. Feature extraction and distortion correction of video frames are completed in the backend, which not only increases the computing pressure of the backend, but also further amplifies the time error, and cannot meet the needs of high-precision inspection scenes.

[0038] In the embodiment, the lens distortion calibration adopts Zhang Zhengyou calibration method, the calibration parameters include intrinsic matrix, distortion coefficient and extrinsic matrix, and the calibration parameters are associated with the time stamp generation module of the hardware level time synchronization module to ensure accurate matching of the video frames after distortion correction and the time stamp information. The hardware level time synchronization module includes a PTP slave clock chip and a GPS Beidou dual-mode time service module. The PTP slave clock chip establishes a master-slave synchronization link with the PTP master clock device deployed in the inspection scene to realize sub-microsecond level time synchronization. The GPS Beidou dual-mode time service module is applied to part of the mobile camera equipment in the inspection scene without network coverage, and realizes time reference alignment with the PTP master clock through satellite time service signal.

[0039] In the embodiment, the main camera equipment in the plurality of camera equipment integrates a PTP master clock gateway, and the main camera equipment manages the time reference of each slave camera equipment through a preset master-slave communication protocol. When multiple devices cooperate to collect, the main camera equipment first sends a time synchronization calibration instruction to the slave camera equipment, and then sends a hardware trigger signal after completing the unification of the system time reference, to control all camera equipment to start data collection at the same time. The main camera equipment receives the time stamp feedback information of each slave camera equipment in real time, triggers secondary calibration for the equipment with time deviation, and ensures that the video stream data and device state data collected by multiple devices are based on the same time reference.

[0040] It should be noted that when the device collects video frames, a nanosecond level high-precision time stamp is directly embedded in the original data frame header of each frame of image, which contains UTC time and device unique identifier. At the same time, the time stamp of the same source as the video frame is embedded in the data packet of the device state data, such as temperature, shutter speed, anti-shake state and positioning coordinates. Avoiding the time drift of soft clock due to system load fluctuation, ensuring that the time reference of video frames and device state data is completely consistent, the time stamp error can be controlled within 100 ns, which is much better than the millisecond level error of traditional soft clock. For mobile camera equipment such as unmanned aerial vehicle inspection and mobile inspection robot, GPS Beidou time service can solve the problem of time synchronization in the environment without network. For fixed camera (such as factory monitoring), PTP hardware synchronization can realize clock unification among multiple devices.

[0041] In a specific implementation, before the device is shipped or deployed on site, the camera lens distortion is calibrated and the sensor parameters are calibrated. The calibration parameters (internal parameter matrix, distortion coefficient and external parameter matrix) are fixed in the device firmware and uploaded to the backend system in synchronization with the video stream and status data. The effects of lens radial distortion and tangential distortion on image feature extraction are eliminated from the source, improving the accuracy of subsequent point cloud registration and feature matching, indirectly reducing the error of spatial mapping. For example, in power inspection, the edge features of insulators can be more accurately aligned with the insulator model in the three-dimensional model after distortion correction.

[0042] S2, multi-source data synchronous acquisition, video stream data and device status data of the inspection scene are synchronously acquired by the camera equipment. The video stream data is acquired at a preset frame rate and embedded with a nanosecond-level hardware timestamp. The device status data is acquired at a preset frequency and associated with the hardware timestamp of the latest video stream.

[0043] In this embodiment, two independent acquisition threads are provided in the camera equipment, which are respectively used for video stream data acquisition and device status data acquisition. The video stream data acquisition thread is acquired at a frame rate of 25fps-30fps, and the device status data acquisition thread is acquired at a frequency of 1Hz-10Hz. The latest hardware timestamp of the video acquisition thread is read through a synchronous triggering mechanism.

[0044] In this embodiment, the device status data includes running state parameters and spatial pose parameters of the camera equipment. The running state parameters include temperature, shutter speed, anti-shake state and exposure parameters. The spatial pose parameters include positioning coordinates and attitude angles. The nanosecond-level hardware timestamp adopts a three-tuple structure of UTC time, device unique identifier and data type identifier. The data type identifier is used to distinguish between video stream data and device status data, ensuring the time reference traceability and accurate association of multi-source data.

[0045] It should be noted that the synchronous triggering mechanism is provided in the device firmware, and the video stream acquisition and device status data acquisition are divided into two independent but synchronous acquisition threads. The video thread acquires images at a fixed frame rate, and each frame is embedded with a hardware timestamp. The status thread acquires device status data at a required frequency, and reads the latest timestamp of the video thread each time to ensure that the timestamps of the status data and the latest video frame are strictly aligned, avoiding the timestamp offset of the video frame caused by the lag of the status data acquisition, solving the problem of timestamp misalignment between the video frame and the status data, and being especially suitable for scenes where the device status data includes position, attitude and other key spatial information.

[0046] In a specific implementation, for the multi-camera networking inspection scene, a hardware trigger interface is added to the camera equipment, a synchronous trigger signal is sent by the master device to control all slave devices to start video collection and state data collection at the same time, the problem of inconsistent collection start time among multiple devices is solved, and video splicing misplacement and three-dimensional model perspective deviation caused by the difference in device start time are avoided, which is especially suitable for panoramic splicing inspection and large-scale scene modeling requirements.

[0047] S3, separate data transmission, using double transmission channels to transmit video stream data and device state data respectively, high bandwidth channel transmits video stream data using UDP protocol, low bandwidth channel transmits device state data and calibration parameters using TCP protocol, and real-time monitoring of transmission delay and time stamp dynamic compensation are performed during transmission.

[0048] In this embodiment, the dual transmission channel logic layer uses dual ports and dual protocols in the same network, channel one is used for transmitting video stream data using UDP protocol, and channel two is used for transmitting device state data using TCP protocol; the time stamp dynamic compensation method is that the camera equipment records the data sending time stamp, the back end records the receiving time stamp, the transmission delay duration is calculated, and the real data collection time is restored by subtracting the delay duration from the receiving time stamp.

[0049] The dual transmission channel can avoid blocking of state data by large volume video frames, ensure that state data arrives at the back end in time, and the back end system can quickly match video frames with corresponding device state data through time stamp association, eliminating time misalignment caused by transmission delay.

[0050] It should be noted that a transmission delay monitoring field is added in the transmission protocol, the camera equipment records the sending time stamp when sending the data packet, the back end records the receiving time stamp when receiving, the delay duration of single transmission is calculated, and the back end system dynamically compensates the time stamp of the video frame and the state data according to the delay duration, such as 50ms delay, the real data collection time is restored by subtracting 50ms from the receiving time stamp, offsetting the transmission delay error caused by network fluctuations, which is especially suitable for long-distance data transmission between edge inspection equipment and cloud systems.

[0051] In this embodiment, the video stream data is hierarchically encoded and transmitted using the SVC scalable coding technology to form a hierarchical code stream of the base layer and the enhancement layer; the base layer carries low-resolution and low-code-rate core video information and is transmitted preferentially using the UDP protocol to ensure real-time performance; the enhancement layer carries high-definition detail information and is dynamically transmitted on demand according to the bandwidth feedback of the back-end system and the quality requirement, to realize bandwidth-adaptive video stream transmission optimization, taking into account the real-time performance and the quality requirement; the video coding uses a frame-level data hard binding mechanism to encapsulate the I frame and the device state data and the camera device extrinsic parameter at the corresponding time, to generate a time and space aligned data packet with a timestamp index, and to maintain the highest transmission priority of the data packet through a transmission priority scheduling mechanism; the P frame and the B frame transmit video differential data and a relative time offset based on the I frame, and the relative time offset is recorded with nanosecond-level precision to ensure the time correlation of the non-key frame and the I frame.

[0052] The I frame is the core frame for video feature extraction, and after being bound with the state data, the spatial mapping precision of the key frame can be greatly improved; at the same time, the additional data transmission amount of the non-key frame is reduced, and the bandwidth pressure is reduced.

[0053] S4, front-end preprocessing, real-time feature extraction and compression of video stream data at the device end or edge end, fusion of device state data, calibration parameters and extracted video feature data, and completion of preliminary time and space alignment.

[0054] In this embodiment, the heterogeneous edge computing chip built-in in the camera device uses a lightweight deep learning feature extraction network to complete real-time feature extraction of the video stream data, and the extracted features include semantic features, geometric features and region of interest coordinates of the inspection target; the original video frame is encoded with a high compression ratio and stored locally for archiving; the extracted feature tensor data, device state data and calibration parameters are encapsulated in a preset data format and transmitted synchronously, to realize collaborative optimization of front-end data dimension reduction and back-end computing load reduction; multiple camera devices construct a distributed data aggregation node through an edge computing gateway, receive multi-source data of each camera device using a time-sensitive network protocol, and the edge computing gateway is built-in with a lightweight time and space fusion algorithm to complete time synchronization calibration and spatial coordinate unification of multi-device data based on a hardware timestamp, to generate a standardized fusion data frame after preliminary time and space alignment, and then push the data frame to the back-end system through a bandwidth adaptive transmission strategy, to realize power distribution and data preprocessing optimization of edge and cloud collaboration.

[0055] It should be noted that the edge computing chip built in the camera equipment is used to extract features of the video frame in real time at the equipment end, such as the corner points, edges and texture features of the inspection target, and the feature data is transmitted in combination with the timestamp and equipment state data, the original video frame can be compressed or stored locally as needed, the backend does not need to extract features from the full amount of video frame, and the feature data transmitted by the equipment end is directly matched with the three-dimensional model, which greatly reduces the computing delay of the backend; at the same time, the amount of video data transmission is reduced, and the time error caused by transmission delay is reduced.

[0056] In a specific implementation, an edge computing gateway is deployed at the inspection site, video streams, state data and laser radar point cloud data of multiple camera equipment in the same area are converged to the gateway, the gateway completes preliminary space-time alignment locally, and then transmits the fused aligned data to the backend system, a large amount of computing tasks are sunk from the cloud to the edge, the amount of data transmitted and the delay of long-distance transmission are reduced; at the same time, the edge can quickly respond to local inspection requirements, real-time abnormality early warning, and further improve the real-time performance of the system.

[0057] S5, accurate space-time alignment at the backend, the backend system receives the preprocessed fused data, accurately matches the video stream data and the equipment state data based on the hardware timestamp, maps the video features to the three-dimensional model coordinate system by combining the point cloud registration algorithm, and realizes the space-time alignment of the video and the three-dimensional model.

[0058] In this embodiment, the accurate matching of the video stream data and the equipment state data is specifically achieved by extracting the three-tuple structure hardware timestamp of each data packet in the fused data by the backend system, establishing a one-to-one correspondence between the video stream data and the equipment state data through the timestamp index, and eliminating abnormal data with a timestamp deviation exceeding a threshold value; the point cloud registration algorithm adopts the ICP algorithm and the NDT algorithm, first performs coarse registration on the point cloud data corresponding to the video features and the point cloud data of the three-dimensional model, and then optimizes the spatial mapping deviation through fine registration, finally accurately maps the video features to the global coordinate system of the three-dimensional model, and realizes the double accurate alignment of the video and the three-dimensional model in the time dimension and the space dimension.

[0059] It should be noted that the front-end preprocessing module preliminarily cleans the collected data, eliminates invalid video frames and abnormal device state data caused by device jitter and signal interference, and standardizes different types of data into a unified format. Among them, the video stream data is encapsulated as a frame sequence with a timestamp index, the device state data is arranged as a structured dictionary with timestamp-device ID-state parameter-value, the attitude and position data are associated to the corresponding inspection device ID, forming a preliminary fusion data packet, which is pushed to the backend system through HTTP / HTTPS protocol. In view of the possible network interruption and data packet loss in scene inspection, the front-end preprocessing module will perform integrity check on the fusion data. If data loss or discontinuous timestamp is detected, a local retransmission mechanism will be triggered to ensure the continuity of the fusion data received by the backend in the time dimension, avoiding the influence of data discontinuity on the subsequent alignment accuracy.

[0060] The triple structure hardware timestamp of each data packet in the fusion data is defined as the collection timestamp, preprocessing timestamp and pushing timestamp, as follows: The collection timestamp is the original hardware timestamp of the data collected by the front-end device sensor and camera, which is generated by the high-precision crystal oscillator clock built-in the device, ensuring that the T1 of the video frames and device state data collected at the same time is completely consistent.

[0061] The preprocessing timestamp is the timestamp of the data cleaning and standardization completed by the front-end preprocessing module, which is used to mark the completion node of the preprocessing process, avoiding the time deviation caused by preprocessing time consumption.

[0062] The pushing timestamp is the timestamp of the fusion data pushed by the front-end to the back-end. After receiving the data, the back-end will calculate the difference between T3 and the back-end system time. If the difference exceeds 500ms (which can be dynamically configured according to the inspection scene), it is determined that there is an abnormal network transmission delay, triggering the timestamp calibration mechanism, taking the back-end system time as the reference to compensate and correct T1 and T2.

[0063] S6, automatic roaming driving, based on the spatio-temporal alignment result, driving the virtual camera in the three-dimensional model to automatically roam along the video feature track and the preset path, synchronously associating the corresponding video stream picture and realizing picture-in-picture superposition or split-screen display, supporting users to switch to the corresponding real scene video picture in one key in the three-dimensional model roaming interface, when the video feature recognizes the inspection abnormal target, automatically triggering the virtual camera to focus on the abnormal area and marking, synchronously calling the historical inspection video segment of the area for comparison and display, realizing the linkage automatic roaming of scene inspection video and three-dimensional model, helping the staff to quickly locate the abnormal position and trace the abnormal evolution process.

[0064] It should be noted that based on the spatio-temporal alignment result, a video feature track following and preset path adaptive dual-mode driving mechanism is constructed to realize intelligent motion control of the virtual camera, as follows: Video feature track following mode: The backend system extracts the actual motion track of the inspection device from the spatio-temporal alignment result, i.e., the physical space track corresponding to the video feature, optimizes the track continuity through a track smoothing algorithm, drives the virtual camera to move along the optimized track, and simultaneously synchronizes the attitude parameters of the inspection device, such as the heading angle, pitch angle, and roll angle, to ensure that the shooting angle and depth of field range of the virtual camera are completely consistent with those of the actual inspection camera, thereby realizing virtual roaming and replicating the real inspection process.

[0065] Pre-set path adaptation mode: The staff can preset an inspection path based on the three-dimensional model, such as a key inspection route planned according to the importance level of the equipment or a full-coverage inspection route divided according to the region. The backend system matches the pre-set path nodes with the video feature track through the position mapping relationship in the spatio-temporal alignment result, drives the virtual camera to move along the pre-set path, and automatically retrieves the video stream picture corresponding to the path node, thereby realizing roaming along the planned path and linkage with the live video.

[0066] The two modes support seamless switching, and the movement speed of the virtual camera can be dynamically adapted according to the video frame rate to avoid the problem of asynchronous roaming angle and video picture.

[0067] It should be noted that, in order to meet the needs of different inspection and monitoring scenes, two core linkage display modes, picture-in-picture overlay and split-screen display, are designed, and the user interaction experience is optimized, as follows: Picture-in-picture overlay mode: The live video picture is overlaid in the specified area of the three-dimensional model roaming main interface, and the overlay window synchronously displays the key information such as the current roaming time point and the corresponding device ID, thereby facilitating the staff to view the three-dimensional space layout while real-time checking the live details.

[0068] Split-screen display mode: Two split-screen layouts, i.e., two-split-screen and three-split-screen, are supported, and the three-split-screen mode can realize the trinity linkage of three-dimensional roaming angle positioning, live video detail observation, and device parameter real-time monitoring, and the split-screen ratio can be self-defined.

[0069] The interaction convenience is strengthened, and the user can click any device or region in the three-dimensional model roaming interface to switch to the corresponding live video picture, and automatically jump to the inspection time node of the device or region. Conversely, the user can click a feature target such as a device valve or a pipeline interface in the live video interface, and the virtual camera in the three-dimensional model will automatically position to the spatial position corresponding to the target and highlight the mark, thereby realizing bidirectional precise linkage between the three-dimensional and live scenes.

[0070] Embodiment two, the present application provides a scene inspection video and three-dimensional model automatic roaming system, referring to Figure 2As shown, the figure is a module diagram of the scene inspection video and three-dimensional model automatic roaming system according to the embodiment of the present application, and the automatic roaming system comprises: A device deployment initialization module is configured to initialize the deployment configuration and parameters of the camera device in the inspection scene, complete lens distortion calibration and sensor parameter calibration, solidify the calibration parameters to the device firmware, and build a matching three-dimensional virtual camera model.

[0071] A multi-source data synchronous acquisition module is configured to realize synchronous acquisition and timestamp embedding of the inspection scene video stream data and device state data.

[0072] A data transmission module is configured to separate the transmission of the video stream and the device state data by using a double transmission channel, so as to realize efficient and low-delay transmission of the multi-source data and guarantee the time synchronization accuracy in the data transmission process.

[0073] A front-end data processing and fusion module is configured to complete data dimension reduction, feature extraction and preliminary space-time alignment at the device end and the edge end, so as to reduce the computing pressure of the back end.

[0074] A back-end space-time alignment module is configured to establish a one-to-one correspondence between the video stream and the device state data based on a triple hardware timestamp, eliminate abnormal data, and realize accurate matching of the fusion data and mapping and alignment of the video features to the three-dimensional model coordinate system.

[0075] An automatic roaming driving interaction module is configured to drive the three-dimensional model virtual camera to roam, so as to realize deep linkage and user interaction between the video and the three-dimensional model.

[0076] A plurality of fixed camera devices integrated with IEEE 1588v2 standard PTP slave clock chips are deployed in a plurality of inspection areas of a factory, a mobile camera device integrated with a GPS Beidou dual-mode time module is carried on a mobile inspection robot, Zhang Zhengyou calibration method is used to calibrate the lens distortion of all camera devices, the intrinsic matrix, distortion coefficient and extrinsic matrix are obtained, the calibration parameters are solidified to the device firmware, in the three-dimensional BIM model of the industrial factory, virtual camera models corresponding to 10 fixed camera devices and 2 mobile camera devices are built, the parameters of each physical camera device are inputted, the parameter matching of the virtual camera and the physical camera device is completed, the main fixed camera device sends a GPIO hardware trigger signal to the other 9 fixed camera devices, so as to ensure that all fixed camera devices start to collect synchronously, and the mobile inspection robot realizes time synchronization with the fixed camera device through the GPS Beidou time module.

[0077] Two independent acquisition threads are arranged in each camera device, a video acquisition thread acquires industrial plant video stream data at a frame rate of 30 fps, and a hardware timestamp at a nanosecond level containing UTC time and a unique identifier of the device is embedded in each frame of video data; a device state acquisition thread acquires device state data at a frequency of 5 Hz, the data including device temperature, shutter speed, anti-shake state and positioning coordinates, the latest hardware timestamp of the video acquisition thread is read each time the device state data is acquired, and time correlation of the device state data and the video frame is achieved.

[0078] Each camera device is configured with dual network cards as dual transmission channels, a high-bandwidth channel 1 transmits video stream data using UDP protocol, the video stream data is encoded using SVC scalable coding of H.265, low-resolution base layer video stream is preferentially transmitted, and a low-bandwidth channel 2 transmits device state data and calibration parameters using TCP protocol, in the transmission process, the camera device records the sending timestamp of each data packet, the backend system records the receiving timestamp when receiving the data packet, the transmission delay duration is calculated, the receiving timestamp is subtracted by the delay duration to restore the real time of data acquisition, and in the video encoding process, each I frame is generated, the I frame is hard-bound with the device state data at the corresponding time, a space-time aligned data packet is generated and preferentially transmitted, and P frames and B frames only transmit video data and time offset relative to the nearest I frame.

[0079] Through the NVIDIA Jetson edge computing chip built in each camera device, real-time feature extraction is performed on the video stream data, corner, edge and texture features of the production device are extracted, the original video frame is compressed by H.265 and stored locally, only the extracted feature data, device state data and calibration parameters are synchronously transmitted to the edge computing gateway, the edge computing gateway aggregates multi-source data of 10 fixed camera devices and 2 mobile inspection robots, preliminary space-time alignment is completed based on the hardware timestamp, and fusion data is generated.

[0080] The backend system receives the fusion data transmitted by the edge computing gateway, accurately matches the video feature data and the device state data based on the nanosecond-level hardware timestamp, maps the extracted video features to the coordinate system of the three-dimensional BIM model of the industrial plant by using the ICP point cloud registration algorithm, and realizes accurate space-time alignment of the video and the three-dimensional model.

[0081] Based on the space-time alignment result, the backend system drives the virtual camera in the three-dimensional BIM model to automatically roam according to the video feature track, automatically drives the virtual camera to jump to the corresponding abnormal position when the video feature recognizes an abnormal production device, synchronously and associatively displays the real-time video stream picture of the position, realizes linkage automatic roaming of the industrial plant inspection video and the three-dimensional model, and facilitates the staff to quickly locate the fault position.

[0082] Although the present application has been described in detail with reference to the foregoing embodiments, the technical solutions recorded in the foregoing embodiments can be modified, or some of the technical features can be replaced by equivalent features, by those skilled in the art, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for automatic roaming of scene inspection video and 3D model, characterized in that, The automatic roaming method includes the following steps: In the inspection scenario, multiple camera devices with hardware-level time synchronization modules are deployed to complete the lens distortion calibration and sensor parameter calibration of the camera devices. The calibration parameters are then fixed to the device firmware. At the same time, a virtual camera model corresponding to the camera device is constructed in the 3D model to complete the parameter matching between the virtual camera and the physical camera device. The video stream data and equipment status data of the inspection scene are collected synchronously by the camera equipment. The video stream data is collected at a preset frame rate and embedded with a nanosecond-level hardware timestamp. The equipment status data is collected at a preset frequency and associated with the hardware timestamp of the most recent frame of the video stream. Dual transmission channels are used to transmit video stream data and device status data respectively. The high-bandwidth channel uses the UDP protocol to transmit video stream data, while the low-bandwidth channel uses the TCP protocol to transmit device status data and calibration parameters. At the same time, the transmission delay is monitored in real time and timestamps are dynamically compensated. Real-time feature extraction and compression of video stream data are performed at the device or edge, and device status data, calibration parameters and extracted video feature data are fused to complete preliminary spatiotemporal alignment; The backend system receives the pre-processed fused data, completes the accurate matching of video stream data and device status data based on hardware timestamps, and maps video features to the three-dimensional model coordinate system by combining point cloud registration algorithm to achieve spatiotemporal alignment between video and three-dimensional model; Based on the spatiotemporal alignment results, the virtual camera in the 3D model is driven to automatically roam along the video feature trajectory and preset path, synchronously associating with the corresponding video stream and realizing picture-in-picture overlay or split-screen display. Users can switch to the corresponding real-scene video screen with one click on the 3D model roaming interface. When the video features identify the abnormal target during inspection, the virtual camera is automatically triggered to focus on the abnormal area and mark it. The historical inspection video clips of the area are retrieved simultaneously for comparison and display, realizing the linkage and automatic roaming of scene inspection video and 3D model, helping staff to quickly locate the abnormal location and trace the abnormal evolution process.

2. The method for automatic roaming of scene inspection video and 3D model according to claim 1, characterized in that: The lens distortion calibration adopts the Zhang Zhengyou calibration method. The calibration parameters include the intrinsic parameter matrix, distortion coefficients and extrinsic parameter matrix. The calibration parameters are associated with the timestamp generation module of the hardware-level time synchronization module to ensure that the video frames after distortion correction are accurately matched with the timestamp information. The hardware-level time synchronization module includes a PTP slave clock chip and a GPS / BeiDou dual-mode time synchronization module; PTP establishes a master-slave synchronization link with the PTP master clock device deployed in the inspection scenario through the clock chip, achieving sub-microsecond time synchronization; The GPS-BeiDou dual-mode timing module is used in inspection scenarios for mobile camera equipment without network coverage, achieving time alignment with the PTP master clock through satellite timing signals.

3. The method for automatic roaming of scene inspection video and 3D model according to claim 1, characterized in that: The master camera among the multiple camera devices integrates a PTP master clock gateway, and the master camera manages the time base of each slave camera through a preset master-slave communication protocol; When multiple devices are collecting data collaboratively, the master camera device first sends a time synchronization calibration command to the slave camera devices to unify the time base of the entire system before sending a hardware trigger signal to control all camera devices to start data acquisition at the same time. The main camera receives real-time timestamp feedback from each slave camera and triggers secondary calibration for devices with time discrepancies, ensuring that the video stream data and device status data collected by multiple devices are based on the same time reference.

4. The method for automatic roaming of scene inspection video and 3D model according to claim 1, characterized in that: The camera device is equipped with two independent acquisition threads, which are used for video stream data acquisition and device status data acquisition, respectively. The video stream data acquisition thread acquires data at a frame rate of 25fps-30fps, while the device status data acquisition thread acquires data at a frequency of 1Hz-10Hz. The latest hardware timestamp of the video acquisition thread is read through a synchronous triggering mechanism.

5. The method for automatic roaming of scene inspection video and 3D model according to claim 1, characterized in that: The device status data includes the camera device's operating status parameters and spatial pose parameters. The operating status parameters include temperature, shutter speed, image stabilization status, and exposure parameters. The spatial pose parameters include positioning coordinates and attitude angles. The nanosecond-level hardware timestamp adopts a triplet structure of UTC time, device unique identifier, and data type identifier. The data type identifier is used to distinguish video stream data from device status data, ensuring the time reference traceability and accurate correlation of multi-source data.

6. The method for automatic roaming of scene inspection video and 3D model according to claim 1, characterized in that: The dual transmission channel logic layer adopts dual ports and dual protocols in the same network. Channel one is used to transmit video stream data using the UDP protocol, and channel two is used to transmit device status data using the TCP protocol. The timestamp dynamic compensation method is as follows: the camera device records the data transmission timestamp, the backend records the reception timestamp, the transmission delay duration is calculated, and the reception timestamp is subtracted from the delay duration to restore the true time of data acquisition.

7. The method for automatic roaming of scene inspection video and 3D model according to claim 1, characterized in that: The video stream data is transmitted using SVC scalable coding technology in a layered encoding process, forming a layered bitstream of a base layer and an enhancement layer. The base layer carries core video information with low resolution and low bit rate, and uses UDP protocol for priority transmission to ensure real-time performance. The enhancement layer carries high-definition detail information and dynamically transmits it on demand based on the bandwidth feedback and image quality requirements of the backend system, realizing bandwidth-adaptive video stream transmission optimization that balances real-time performance and image quality requirements. The video encoding uses a frame-level data hard binding mechanism to encapsulate I-frames with corresponding device status data and camera device external parameters, generating spatiotemporally aligned data packets with timestamp indexes. The data packets are kept at the highest transmission priority through a transmission priority scheduling mechanism. P-frames and B-frames transmit video differential data and relative time offsets based on I-frames. The relative time offsets are recorded with nanosecond precision to ensure the temporal correlation between non-critical frames and I-frames.

8. The method for automatic roaming of scene inspection video and 3D model according to claim 1, characterized in that: The camera device has a built-in heterogeneous edge computing chip that uses a lightweight deep learning feature extraction network to complete real-time feature extraction of video stream data. The extracted features include semantic features, geometric features, and region of interest coordinates of the inspection target. The original video frames are encoded with a high compression ratio and then stored and archived locally. The extracted feature tensor data, device status data, and calibration parameters are encapsulated in a preset data format and transmitted synchronously, realizing the collaborative optimization of front-end data dimensionality reduction and back-end computational burden reduction. Multiple camera devices form a distributed data aggregation node through an edge computing gateway. The gateway uses a time-sensitive network protocol to receive multi-source data from each camera device. The edge computing gateway has a built-in lightweight spatiotemporal fusion algorithm that uses hardware timestamps to complete time synchronization calibration and spatial coordinate unification of multi-device data. After initial spatiotemporal alignment, it generates standardized fused data frames, which are then pushed to the backend system through a bandwidth adaptive transmission strategy. This enables edge and cloud collaborative computing power distribution and data preprocessing optimization.

9. The method for automatic roaming of scene inspection video and 3D model according to claim 1, characterized in that: The precise matching of video stream data and device status data is achieved by the backend system extracting the triplet structure hardware timestamp of each data packet in the fused data, establishing a one-to-one correspondence between video stream data and device status data through timestamp indexing, and eliminating abnormal data with timestamp deviations exceeding the threshold. The point cloud registration algorithm uses the ICP algorithm and the NDT algorithm. First, it performs coarse registration between the point cloud data corresponding to the video features and the point cloud data of the 3D model. Then, it optimizes the spatial mapping deviation through fine registration. Finally, it accurately maps the video features to the global coordinate system of the 3D model, achieving precise alignment between the video and the 3D model in both time and space dimensions.

10. A scene inspection video and 3D model automatic roaming system, used to execute the scene inspection video and 3D model automatic roaming method as described in any one of claims 1 to 9, characterized in that, The automatic roaming system includes: The equipment deployment initialization module is used for the deployment configuration and parameter initialization of camera equipment in inspection scenarios. It completes lens distortion calibration and sensor parameter calibration, solidifies the calibration parameters into the equipment firmware, and builds a matching 3D virtual camera model. The multi-source data synchronous acquisition module is used to realize the synchronous acquisition and timestamp embedding of video stream data and equipment status data in the inspection scenario; The data transmission module uses dual transmission channels to separately transmit video streams and device status data, enabling efficient and low-latency transmission of multi-source data and ensuring time synchronization accuracy during data transmission. The front-end data processing and fusion module performs data dimensionality reduction, feature extraction, and preliminary spatiotemporal alignment at both the device and edge levels to reduce the computational burden on the back-end. The backend spatiotemporal alignment module establishes a one-to-one correspondence between video streams and device status data based on triplet hardware timestamps, and removes abnormal data. This is used to achieve accurate matching of fused data and mapping and alignment of video features to the three-dimensional model coordinate system. The automatic roaming and interactive module drives the virtual camera of the 3D model to roam, enabling deep linkage between video and 3D model and user interaction.

Citation Information

Patent Citations

  • Three-dimensional video intelligent inspection system and inspection method

    CN112366821A

  • A method and system for three-dimensional roaming inspection of tunnels

    CN114187414B

  • Video inspection system based on three-dimensional model, parameter synchronous debugging method and control method thereof

    CN118301291A