Mobile video and live-action three-dimensional rapid fusion method

By selecting key frames on mobile devices and using dynamic reuse and micro-compensation mechanisms to process non-key frames, the problem of fusing mobile device video data with real-world 3D models was solved, achieving efficient real-time matching of video and 3D models.

CN120876262AActive Publication Date: 2025-10-31TIANJIN SURVEYING & MAPPING INST CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511367078.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2025-10-31
Estimated Expiration
2045-09-24

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively integrate video data collected by mobile devices with real-world 3D models. In particular, traditional calibration methods fail when the device's position and orientation change, making effective matching difficult.

Method used

The method involves selecting key frames at predetermined frequencies, calculating rotation matrices based on the pose information of the key frames, and processing non-key frames through dynamic reuse and micro-compensation mechanisms. It also utilizes sensors such as GNSS and IMU to acquire real-time position and attitude information of the device, and dynamically adjusts the rotation matrices and positions of non-key frames.

Benefits of technology

It significantly reduces the computational complexity of the system, improves the response speed and frame rate of fusion processing, and meets the requirements of high real-time applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_55
    Figure SMS_55
  • Figure SMS_69
    Figure SMS_69
  • Figure SMS_71
    Figure SMS_71
Patent Text Reader

Abstract

The invention discloses a mobile video and live-action three-dimensional rapid fusion method, which comprises the following steps of: selecting key frames according to a preset frequency, and calculating a rotation matrix of the key frames based on pose information corresponding to the key frames; judging whether the modulus length of the angular velocity vector corresponding to the non-key frame is greater than an angular velocity threshold value or not; if not, the non-key frame directly selects the rotation matrix of the adjacent key frame, and the position is updated according to the moving speed; and if yes, calculating a rotation compensation matrix corresponding to the non-key frame according to the angular velocity vector, integrating the rotation compensation matrix with the rotation matrix of the adjacent key frame to obtain the rotation matrix of the non-key frame, and updating the position according to the moving speed. According to the method, a non-key frame rotation matrix dynamic multiplexing and micro-compensation mechanism is introduced, so that the average calculation complexity of the system is remarkably reduced, the response speed and the frame rate of integral fusion processing are greatly improved, and the requirements of remote command and other application scenes with high real-time requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video processing technology, specifically to a method for rapid fusion of mobile video and real-world 3D scenes. Background Technology

[0002] Existing real-scene 3D video fusion technologies typically rely on cameras in fixed positions (such as surveillance cameras). Their core function is to project video footage onto the surface of a 3D model using pre-calibrated coordinates. 3D visualization and virtual reality technologies enable map manipulation, venue information location queries, security monitoring, and comprehensive information display on 3D electronic maps. The system allows for very fast map browsing on B / S clients, largely meeting the practical needs of the system. Currently, the construction of 3D geographic information platforms is still in its early stages. Further development and improvement are needed in areas such as maintaining smooth browsing of massive 3D models, integrating data from existing 2D systems into 3D scenes, integrating existing 2D geographic information systems into 3D geographic information platforms, and providing corresponding analysis and decision support modules for specific business applications (such as population information queries and flight simulations).

[0003] CN106874436A discloses a multi-source image fusion imaging system for a 3D police geographic information platform. It addresses the technical problem of low realism in simulated objective worlds in existing technologies. The system includes an oblique remote sensing photogrammetry device, a near-ground mobile photogrammetry device, a real-scene 3D modeling subsystem, a building information modeling subsystem, a 2D GIS subsystem, and a police information subsystem. The oblique remote sensing photogrammetry device and the near-ground mobile photogrammetry device are connected to the real-scene 3D modeling subsystem. The real-scene 3D modeling subsystem and the building information modeling subsystem are connected to a virtual reality city information model construction system. The virtual reality city information model construction system, the 2D GIS subsystem, and the police information subsystem are connected to the multi-source fusion subsystem. Its advantages include breakthroughs in consistent global positioning, ground-air integration, indoor-outdoor integration, and interface and standardization. However, this technology faces bottlenecks for video data collected by mobile terminals such as mobile phones and mobile recorders, making fusion difficult: First, the position and orientation of mobile devices are constantly changing, rendering traditional calibration fusion methods that rely on static presets ineffective; second, the motion states of mobile videos are often complex, making effective matching difficult with traditional fusion methods. These shortcomings limit the application of mobile terminal video in existing real-scene 3D fusion systems. Summary of the Invention

[0004] The purpose of this invention is to provide a method for rapid fusion of mobile video and real-world 3D scenes to solve the problems mentioned in the background art.

[0005] To achieve the above-mentioned technical effects, the present invention adopts the following solution: A method for rapid fusion of mobile video and real-world 3D scene includes the following steps: Select keyframes at a predetermined frequency. Calculate the rotation matrix of the keyframe based on its pose information; Determine whether the magnitude of the angular velocity vector corresponding to a non-keyframe is greater than the angular velocity threshold; If not, the non-keyframe directly uses the rotation matrix of the adjacent keyframe and updates the position according to the movement speed; If so, the rotation compensation matrix corresponding to the non-key frame is calculated based on the angular velocity vector and integrated with the rotation matrices of the adjacent key frames to obtain the rotation matrix of the non-key frame. At the same time, the position is updated based on the movement speed.

[0006] The above technical solution also includes calculating a composite motion metric for each non-key frame relative to the previous key frame. The composite motion metric reflects the rate of change of position, the rate of change of attitude, and their coupling degree. If it exceeds a predetermined threshold, the non-key frame is selected as a key frame.

[0007] In the above technical solution, the video acquisition terminal continuously acquires the device's geographical location, three-dimensional pose information, and video information in real time, generating pose information corresponding to keyframes. The pose information includes position information and velocity. acceleration Pitch angle Yaw angle Roll angle and the angular velocity vector ω.

[0008] In the above technical solution, the calculation method of the composite motion measurement index is as follows: , is a composite motion metric, where a is the acceleration corresponding to a non-keyframe, ω is the angular velocity vector corresponding to a non-keyframe, and α, β, and γ are weighting coefficients.

[0009] In the above technical solution, the rotation compensation matrix for: ; , The rotation matrix for adjacent keyframes. The rotation matrix obtained by integration for the current non-keyframe, where ω is the angular velocity vector of the adjacent keyframe, ω x ω y ω z These represent the angular velocities about the x, y, and z axes, respectively. It is the difference between the current non-keyframe timestamp and the most recent keyframe timestamp.

[0010] By adopting the above technical solution, compared with the prior art, the beneficial effects of the present invention are as follows: The present invention introduces a dynamic reuse and micro-compensation mechanism for non-critical frame rotation matrices, concentrating core computing resources on processing a small amount of high-value critical frame data. For a large number of non-critical frames, projection parameters are quickly obtained through dynamic reuse and micro-compensation mechanisms. This processing strategy significantly reduces the average computational complexity of the system, greatly improves the overall fusion processing response speed and frame rate, and meets the needs of command and other application scenarios with high real-time requirements. Detailed Implementation

[0011] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0012] The method for rapid fusion of mobile video and real-world 3D scene according to the present invention includes the following steps: Keyframes are selected at a predetermined frequency, meaning that the selection of keyframes uses a fixed time interval extraction mechanism, such as setting a time window baseline value. =0.5s, based on the start time of the video. Based on this, keyframe time series are generated according to the following formula: , Extract timestamp The corresponding video frames serve as keyframes, and the extraction time interval for these keyframes can be adjusted according to actual needs. The video acquisition terminal is equipped with an integrated multi-source sensor assembly, including a built-in satellite positioning module, inertial measurement unit, orientation sensor, and video acquisition module, which continuously acquires the device's geographical location and 3D attitude information in real time. The pose information of the keyframes is generated by fusing data from GNSS, IMU, and other multi-source sensors. This pose information includes position information P(x,y,z), where P(x,y,z) represents spatial coordinates based on an independent coordinate system, such as the 3D coordinates in the Tianjin 2000 coordinate system, in meters (m), and velocity... acceleration Pitch angle Yaw angle Roll angle And the angular velocity vector ω. To achieve pose information reading of keyframes, the video acquisition device (such as a camera) and all attitude sensors (velocity, acceleration, angular velocity, etc.) are first made to operate based on the same time reference (such as a system timestamp or hardware synchronization signal) to ensure that the acquisition time of all devices can be accurately aligned. Then, the data acquisition triggered by the keyframe timestamp is solved. The acquired data can be obtained by taking the average value within a predetermined window range to obtain the pose information corresponding to the keyframe. This lays a key foundation for the accurate registration of video frames captured in motion with the 3D scene, solves the problem of real-scene video fusion during mobile video acquisition, and aligns the device clock with the central server through a time synchronization protocol, laying the foundation for subsequent spatiotemporal data fusion.

[0013] To avoid the loss of key information caused by collecting key frames at a fixed frequency when the motion amplitude is too large, the invention further includes calculating a composite motion metric index for each non-key frame relative to the previous key frame. The composite motion metric index reflects the rate of change of position, the rate of change of attitude, and their coupling degree. If it exceeds a predetermined threshold, the candidate non-key frame is selected as a key frame. Key frames obtained in this way are processed in the same way as key frames selected at fixed time intervals. This invention adopts a dynamic key frame discrimination mechanism, breaking through the traditional fixed time interval key frame selection method.

[0014] The specific rules are as follows: By default, two frames are extracted per second as candidate keyframes to ensure the lowest possible spatiotemporal resolution. Then, a composite motion metric M is constructed to reflect the rate of change of position, the rate of change of pose, and their coupling degree. When M is greater than a threshold, it indicates a significant change in motion state, and the current frame is forcibly marked as a keyframe. The calculation method for the composite motion metric is as follows: Where M is a composite motion metric. Let ω be the acceleration vector of the currently calculated non-keyframe, and ω be the angular velocity vector of the calculated non-keyframe. , , , which is a weighting coefficient used to determine the contribution of acceleration and rotational motion to the "compositeness". When M is greater than the threshold, it is determined that the motion state has changed significantly, and the frame is forcibly selected as a key frame, with its position information and attitude parameters recorded as key position parameters.

[0015] When designing specific parameters, the first step is to analyze the physical magnitude and distribution: calculate the data from a large number of samples. (usually 0~2g) (typically 0~2 rad / s) The magnitude and distribution. Discovery. In general motion, the range of variation is relatively small and the movement is stable. and The changes are dramatic and even greater in rotational and combined motions. Then, the independent components affect the test: fixed. =0, =0, adjust It was found that relying solely on acceleration is insufficient to effectively capture changes in rotation and combined motion; smaller weights are required. (Fixed) =0, =0, adjust Angular velocity is sensitive to rotation; assigning it a higher weight can better reflect attitude changes. (Fixed) =0, =0, adjust Coupling terms When acceleration and rotation occur simultaneously in non-gravitational directions (such as braking / acceleration during turning), the threshold increases significantly, which is key to identifying strong composite motions and requires high weighting. Then, iterative adjustments are made: parameters are gradually increased, and thresholds are gradually decreased, observing changes in projection accuracy to ultimately determine the parameters and thresholds. Finally, parameter adjustments are made for different scenarios: High-dynamic / intense motion scenarios: slightly increase the threshold to raise the trigger threshold and reduce keyframes generated by ordinary intense motion; Gentle motion scenarios: slightly decrease the threshold to improve sensitivity and capture more subtle motion changes; Strong rotation-dominated scenarios: decrease... ,Increase Laboratory tests showed that... =1, =5, =5, threshold =0.25, which can be appropriately reduced to improve accuracy. In specific implementation, the video acquisition module continuously outputs while the attitude sensor continuously acquires and caches attitude data. The data processing unit calculates the above composite motion metrics in real time. When the composite motion metrics corresponding to a certain ordinary frame reach or exceed the threshold, the data saving mechanism is triggered, the timestamp and pose data of the non-key frame (ordinary frame) are extracted, and the non-key frame is bound and stored as a key frame.

[0016] Then, the rotation matrix of the keyframe is calculated based on the pose information corresponding to the keyframe. Only keyframes are used to calculate the rotation matrix, which greatly reduces the amount of processing and makes it easier to implement on mobile devices; specifically, for any keyframe, the rotation matrix... The parameters that need to be obtained in the calculation include: pitch angle Yaw angle Roll angle Rotation matrix of keyframes calculate:

[0017]

[0018]

[0019]

[0020] Rotation about the z-axis Rotation matrix of angle, To bypass z The rotation angle of the shaft, i.e., the yaw angle; Rotation about the y-axis Rotation matrix of angle, The angle of rotation about the y-axis is the pitch angle. To bypass x Axis rotation Rotation matrix of angle, To bypass x The rotation angle of the shaft, i.e., the roll angle.

[0021] Then construct the projection matrix. ,

[0022] This refers to the location information corresponding to the key frame, that is, the spatial location information of the sensor at the timestamp corresponding to the key frame.

[0023] Then, the keyframes are transformed using perspective projection to map the video pixels (x, y, z) to the 3D model coordinates (X, Y, Z), thus achieving the mapping from pixel coordinates to world coordinates.

[0024] The homogeneous coordinates (4-dimensional vector) in the world coordinate system represent the position of a point in space. (X,Y,Z) are the three-dimensional coordinates of the point in the world coordinate system; the last dimension is fixed at 1 (which is the normalization of homogeneous coordinates, which facilitates translation through matrix multiplication). For the camera intrinsic parameter matrix, This refers to the column and row indexes of pixels in the image.

[0025] For non-keyframes, firstly, it is determined whether the magnitude of the angular velocity vector corresponding to the non-keyframe is greater than the angular velocity threshold. If not, the rotation matrix of the adjacent keyframe is directly used for the non-keyframe, and the position is updated according to the movement speed and time difference of the adjacent keyframe. If yes, the rotation compensation matrix of the non-keyframe is calculated based on the angular velocity vector of the adjacent keyframe and integrated with the rotation matrix of the adjacent keyframe to obtain the rotation matrix of the non-keyframe. At the same time, the position is updated according to the movement speed and time difference of the adjacent keyframe. After calculating the rotation matrix and position of the non-keyframe, its corresponding projection matrix can be calculated in the same way as the keyframe. Then, the video pixels are mapped to the coordinates of the three-dimensional model through perspective projection transformation. In this invention, adjacent keyframes refer to the most recent previous keyframe.

[0026] Specifically, for each non-keyframe: when ω ≤ threshold (tested in the lab, threshold set to 0.03 rad / s), since the keyframe selection involves small time intervals, small pose changes, and small position changes, the rotation matrix of the nearest adjacent keyframe is directly reused, and only the position is updated; the dynamic keyframe extraction mechanism ensures that the pose changes between adjacent keyframes are minimal, and the selection of keyframes also ensures... The error introduced by reuse is negligible, as it is extremely small.

[0027] Non-keyframe position information The update uses quadratic kinematic equations. For acceleration and velocity, due to the short time interval and small changes, the position information of adjacent keyframes is directly selected. acceleration and speed To reduce the amount of calculation while maintaining a certain level of accuracy, that is,

[0028] When ω > threshold, this invention employs a rotation compensation matrix, i.e., a micro-rotation compensation mechanism. The micro-rotation compensation algorithm is activated to calculate a rotation compensation matrix, i.e., a micro-rotation matrix, which is then fitted with the rotation matrix of the most recent previous keyframe. This approach improves efficiency while reducing error. for:

[0029] , To integrate the rotation matrices of the current non-keyframes, ω is the rotation matrix between adjacent keyframes, and ω is the angular velocity vector between adjacent keyframes, typically the triaxial angular velocities measured by the sensor. x ω y ω z These represent the angular velocities about the x, y, and z axes, respectively. The difference between the current non-keyframe timestamp and the nearest adjacent keyframe timestamp is used to update the position information of the non-keyframe in the same way as described above.

[0030] For non-keyframes, after calculating their respective rotation matrices and position information, the video pixel points (x,y,z) are mapped to the 3D model coordinates (X,Y,Z) through perspective projection transformation in the same way as keyframes, that is, the mapping from pixel coordinates to world coordinates is realized, which will not be elaborated here.

[0031] This invention proposes a lightweight rotation matrix dynamic reuse mechanism. This mechanism, through a hierarchical processing strategy and a dynamic micro-compensation algorithm, fully utilizes the position parameters carried by keyframes. While ensuring the projection accuracy of non-keyframes, this method effectively reduces the computational load of the projection process, significantly lowering computational complexity while maintaining the projection accuracy of non-keyframes. It effectively solves the real-time fusion computation bottleneck caused by frequent changes in device posture in mobile scenes.

[0032] This invention continuously acquires the geographical location and 3D attitude information of a mobile video device in real time through multi-source sensors (GNSS, IMU, gyroscope) integrated on the device. For videos with complex motion, keyframes are captured according to corresponding rules. The system prioritizes processing the projection matrix of the received keyframes, calculates the accurate fusion parameters under the key perspective, and accurately projects and maps them onto the observation plane corresponding to the 3D reality model through a perspective projection transformation model. A dynamic reuse and micro-compensation mechanism for non-keyframe rotation matrices is constructed. Through the dynamic reuse and micro-compensation mechanism, the computational load of fusion between mobile video and 3D reality is significantly reduced while ensuring the projection accuracy of non-keyframes.

Claims

1. A method for rapid fusion of mobile video and real-world 3D scene, characterized in that, Includes the following steps, Select keyframes at a predetermined frequency. Calculate the rotation matrix of the keyframe based on its pose information; Determine whether the magnitude of the angular velocity vector corresponding to a non-keyframe is greater than the angular velocity threshold; If not, the non-keyframe directly uses the rotation matrix of the adjacent keyframe and updates the position according to the movement speed; If so, the rotation compensation matrix corresponding to the non-key frame is calculated based on the angular velocity vector and integrated with the rotation matrices of the adjacent key frames to obtain the rotation matrix of the non-key frame. At the same time, the position is updated based on the movement speed.

2. The method for rapid fusion of mobile video and real-world 3D scene as described in claim 1, characterized in that, It also includes calculating a composite motion metric for each non-key frame relative to the previous key frame. The composite motion metric reflects the rate of change of position, the rate of change of attitude, and their coupling degree. If it exceeds a predetermined threshold, the non-key frame is selected as a key frame.

3. The method for rapid fusion of mobile video and real-world 3D scene as described in claim 1 or 2, characterized in that, The video acquisition terminal continuously acquires the device's geographic location, 3D pose information, and video information in real time, generating pose information corresponding to keyframes. This pose information includes position information and velocity. acceleration Pitch angle Yaw angle Roll angle and the angular velocity vector ω.

4. The method for rapid fusion of mobile video and real-world 3D scene as described in claim 2, characterized in that, The calculation method for the aforementioned composite motion measurement index is as follows: , M is the composite motion metric, a is the acceleration corresponding to the non-keyframe, ω is the angular velocity vector corresponding to the non-keyframe, and α, β, and γ are weighting coefficients.

5. The method for rapid fusion of mobile video and real-world 3D scene as described in claim 1, characterized in that, The rotation compensation matrix for: ; , The rotation matrix for adjacent keyframes. The rotation matrix obtained by integration for the current non-keyframe, where ω is the angular velocity vector of the adjacent keyframe, ω x ω y ω z These represent the angular velocities about the x, y, and z axes, respectively. It is the difference between the current non-keyframe timestamp and the most recent keyframe timestamp.

Citation Information

Patent Citations

  • Multi-source image fusion imaging system of three-dimensional police geographic information platform

    CN106874436A

  • Fusion method and system of mobile video and geographic scene and electronic equipment

    CN111582022A

  • Mobile robot vision SLAM key frame adaptive screening method based on neural network

    CN113076988A

  • Intelligent agent navigation map updating method, equipment and medium

    CN116539025A

  • External parameter calibration method and device based on vehicle-mounted camera

    CN119338919A