Video fusion method and device based on three-dimensional scene, and storage medium

By acquiring video frame data from motion devices and constructing mapping relationships, and by optimizing camera pose data using geographic data, precise fusion of 3D scenes and videos is achieved. This solves the problems of unnatural fusion and insufficient real-time performance in existing technologies, and improves the fusion effect and user experience.

CN121052994BActive Publication Date: 2026-02-06WSGRI SMART CITY(WUHAN) ENGINEERING TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511598953.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-02-06
Estimated Expiration
2045-11-04

AI Technical Summary

Technical Problem

In existing technologies, the methods for fusing 3D scenes and videos are time-consuming and laborious, making it difficult to achieve the desired fusion effect. Furthermore, the fusion is unnatural and lacks real-time performance in large-scale scenes.

Method used

By acquiring multiple video frames from a camera mounted on motion equipment, extracting image feature points and constructing mapping relationships, and using strong constraint optimization terms from geographic data to iteratively solve the camera pose data, the conversion parameters between the video frame images and the 3D scene model are determined, thereby achieving accurate fusion of video and 3D scene.

Benefits of technology

It improves the accuracy and robustness of camera pose data calculation, enhances the real-time nature and naturalness of video fusion, ensures spatial alignment and dynamic overlay of video content and 3D models, and improves user immersion and interaction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121052994B_ABST
    Figure CN121052994B_ABST
Patent Text Reader

Abstract

The application provides a three-dimensional scene-based video fusion method and device and a storage medium, and belongs to the technical field of video monitoring. The method comprises the following steps: acquiring video frame data collected by a camera installed on a motion device, wherein the video frame data comprises a video frame image and a corresponding timestamp; acquiring space-time information of the motion device, wherein the space-time information comprises geographic data and a corresponding timestamp; extracting image feature points of each video frame image, and constructing a mapping relationship between the image feature points and the geographic data; based on the mapping relationship, iteratively solving pose data of the camera by minimizing a re-projection error, wherein a target function corresponding to the re-projection error comprises a strong constraint optimization of the geographic data; based on the pose data, determining conversion parameters between a first coordinate system of the video frame image and a second coordinate system of a pre-constructed three-dimensional scene model, obtaining video fusion parameters, and fusing the multiple video frame data with the three-dimensional scene model, thereby improving the fusion effect of the three-dimensional scene and the video.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video monitoring, and in particular to a video fusion method and device based on a three-dimensional scene and a storage medium. BACKGROUND

[0002] With the development of three-dimensional modeling and video monitoring technology, it has become a key requirement in the fields of smart city management and industrial safety supervision to fuse real-time video pictures with a pre-constructed three-dimensional scene. The fusion of video content and a three-dimensional scene is crucial for improving user experience and enhancing reality immersion.

[0003] In related technologies, the fusion of a video and a three-dimensional scene relies on complex geometric transformations and manual adjustments, which not only consumes time and effort but also fails to achieve an ideal fusion effect. Although automatic fusion methods improve efficiency to some extent, they still have problems such as unnatural fusion and insufficient real-time performance when dealing with large-scale scenes. SUMMARY

[0004] Therefore, it is necessary to provide a video fusion method and device based on a three-dimensional scene and a storage medium to solve the technical problems of unnatural fusion and insufficient real-time performance of a video based on a three-dimensional scene in the prior art, which leads to poor fusion effect of a three-dimensional scene and a video.

[0005] To solve the above technical problems, in a first aspect, the present application provides a video fusion method based on a three-dimensional scene, comprising:

[0006] obtaining a camera installed on a motion device, collecting multiple frames of video frame data in a target area, each frame of the video frame data including a video frame image and a corresponding first timestamp, and obtaining spatiotemporal information of the motion device, the spatiotemporal information including geographic data and a corresponding second timestamp;

[0007] extracting image feature points of each frame of the video frame image, and constructing a mapping relationship between the image feature points and the geographic data based on the first timestamp and the second timestamp;

[0008] based on the mapping relationship, iteratively solving pose data of the camera by minimizing a re-projection error, the target function corresponding to the re-projection error including a strong constraint optimization term of the geographic data;

[0009] based on the pose data, determining conversion parameters between a first coordinate system of the video frame image and a second coordinate system of a pre-constructed three-dimensional scene model, obtaining video fusion parameters, and fusing the multiple frames of video frame data with the three-dimensional scene model.

[0010] In a possible implementation, before the strong constraint of the geographic data is included in the target function corresponding to the re-projection error by minimizing the re-projection error based on the mapping relationship, the target function further includes:

[0011] a penalty term corresponding to the strong constraint of the geographic data is determined;

[0012] a target function corresponding to the re-projection error is obtained based on the penalty term and an expression corresponding to the re-projection error.

[0013] In a possible implementation, the view feature corresponding to the model view is extracted, including:

[0014] the pose data of the camera is iteratively solved by minimizing the re-projection error based on the mapping relationship, including:

[0015] the pose data of the camera is calculated by using a nonlinear least squares optimization algorithm to iteratively solve the target function with the mapping relationship as observation data of an iterative solving process and the minimization of the target function as an optimization objective, wherein a sparse matrix decomposer is used in each iteration.

[0016] In a possible implementation, the assembly feature corresponding to the model assembly diagram is extracted, including:

[0017] the image feature point of each frame of the video frame image is extracted, including:

[0018] the feature point of the motion device is extracted from each frame of the key frame by using an improved ORB algorithm, and the improved ORB algorithm includes at least one of adjusting a preset corner response threshold, introducing a scale pyramid to extract a feature point, retaining a feature point by non-maximum suppression, and improving a dimension of a feature descriptor.

[0019] the motion of the feature point in adjacent video frame images is tracked by using an optical flow method to generate a feature point motion vector, and the image feature point is obtained.

[0020] In a possible implementation, the mapping relationship between the image feature point and the geographic data is constructed based on the first timestamp and the second timestamp, including:

[0021] for each image feature point, initial geographic data that matches is determined based on the first timestamp and the second timestamp, to form a plurality of initial mapping pairs;

[0022] a re-projection error is calculated based on the initial mapping pairs, and an initial mapping pair that does not meet a matching condition is removed by using a random sample consensus algorithm according to a re-projection error result;

[0023] When the number of initial mapping pairs meeting the matching condition is less than the preset number threshold, then continue to acquire image feature points corresponding to a new video frame image until the number of initial mapping pairs meeting the matching condition is greater than or equal to the preset number threshold, and the mapping relationship is obtained.

[0024] In a possible implementation, the determining, based on the pose data, of the conversion parameter between the first coordinate system of the video frame image and the second coordinate system of the pre-constructed three-dimensional scene model to obtain the video fusion parameter comprises the following steps.

[0025] Converting the first coordinate system of the video frame image into the Cartesian coordinate system of the three-dimensional scene model through UTM projection based on the pose data;

[0026] Determining a scaling factor between mapping of three-dimensional dimensions of the three-dimensional scene model to two-dimensional image pixel dimensions of the video frame image;

[0027] Constructing a pose rotation matrix corresponding to the pose data according to the Cartesian coordinate system and the scaling factor;

[0028] Based on the pose rotation matrix, constructing a projection error equation of the three-dimensional scene model and the video frame image by using an optical flow method;

[0029] According to the projection error equation, predicting an optimal pose increment by using a Kalman filtering method;

[0030] Stacking the optimal pose increment to the current pose data to update the pose data;

[0031] Based on the updated pose data, the conversion parameter is obtained by minimizing the projection error equation;

[0032] According to the conversion parameter and the scaling factor, the video fusion parameter is generated.

[0033] In a possible implementation, after the determining, based on the pose data, of the conversion parameter between the first coordinate system of the video frame image and the second coordinate system of the pre-constructed three-dimensional scene model to obtain the video fusion parameter, the method further comprises the following steps.

[0034] Constructing a first target function and a second target function, the first target function being a first expression of an absolute value of a difference between an optimized position point and a projected theoretical position point, and the second target function being an optical flow error of the video frame image and a structural similarity error of the video frame image;

[0035] Based on the first target function and the second target function, iteratively optimizing the video fusion parameter.

[0036] In a possible implementation, the iterative optimization of the video fusion parameter based on the first target function and the second target function comprises:

[0037] In each iteration, a first solving result is obtained by solving with the optimization objective of minimizing the first target function;

[0038] The video fusion parameter is iteratively optimized with the optimization objective of minimizing the second target function with the first solving result as a constraint condition.

[0039] In a second aspect, the present application further provides a video fusion device based on a three-dimensional scene, comprising:

[0040] An acquisition unit is configured to acquire a camera installed on a motion device, collect a plurality of frames of video frame data in a target area, each frame of the video frame data comprising a video frame image and a corresponding first timestamp, and acquire spatio-temporal information of the motion device, the spatio-temporal information comprising geographic data and a corresponding second timestamp;

[0041] A construction unit is configured to extract image feature points of each frame of the video frame image, and construct a mapping relationship between the image feature points and the geographic data based on the first timestamp and the second timestamp;

[0042] A calculation unit is configured to iteratively calculate pose data of the camera by minimizing a re-projection error based on the mapping relationship, the target function corresponding to the re-projection error comprising a strong constraint optimization item of the geographic data;

[0043] A fusion unit is configured to determine conversion parameters between a first coordinate system of the video frame image and a second coordinate system of a three-dimensional scene model constructed in advance based on the pose data, to obtain video fusion parameters, and to fuse the plurality of frames of video frame data with the three-dimensional scene model.

[0044] In a third aspect, the present application further provides an electronic device comprising a memory and a processor, wherein:

[0045] The memory is configured to store a program;

[0046] The processor is coupled with the memory and is configured to execute the program stored in the memory to implement steps in the video fusion method based on a three-dimensional scene in any of the above implementation manners.

[0047] In a fourth aspect, the present application further provides a computer readable storage medium for storing computer readable programs or instructions, which can implement steps in the video fusion method based on a three-dimensional scene in any of the above implementation manners when executed by a processor.

[0048] The beneficial effects of the present application are:

[0049] The video fusion method based on a three-dimensional scene provided by the present application acquires multiple frames of video frame data collected by a camera installed on a moving device in a target area, each frame of video frame data including a video frame image and a corresponding first timestamp, and acquires spatiotemporal information of the moving device, the spatiotemporal information including geographic data and a corresponding second timestamp, ensuring data collection efficiency and accuracy and providing an accurate, real-time and reliable basis for subsequent fusion of video content and a three-dimensional scene; image feature points of each frame of video frame image are extracted, and a mapping relationship between the image feature points and the geographic data is constructed based on the first timestamp and the second timestamp, ensuring accurate time alignment, which can enable the image feature points to be accurately mapped to corresponding positions in a three-dimensional scene model, and at the same time, in a dynamic scene, the mapping relationship can track changes in the image feature points over time, thereby supporting accurate rendering and fusion of dynamic objects; based on the mapping relationship, pose data of the camera is iteratively solved by minimizing a re-projection error, the target function corresponding to the re-projection error including a strongly constrained optimization term of the geographic data, which fully utilizes accurate data of vehicle position and attitude provided by GPS, reduces errors that may be caused by relying only on visual information, greatly improves the accuracy and robustness of the camera pose data solution, and is conducive to enhancing the real-time and naturalness of subsequent video fusion; based on the pose data, conversion parameters between a first coordinate system of the video frame image and a second coordinate system of the three-dimensional scene model pre-constructed are determined, video fusion parameters are obtained, and the multiple frames of video frame data are fused with the three-dimensional scene model, ensuring the consistency and coherence of the video frame image and the three-dimensional scene model in vision, realizing spatial alignment and dynamic superposition of the video content and the three-dimensional model, making the fusion of the three-dimensional scene model and the image video frame more real-time and natural, and improving the fusion effect of the three-dimensional scene and the video. BRIEF DESCRIPTION OF DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.

[0051] Figure 1 An embodiment flowchart of the video fusion method based on a three-dimensional scene provided by the present application;

[0052] Figure 2 An embodiment flowchart of the video fusion method based on a three-dimensional scene provided by the present application;

[0053] Figure 3An embodiment structure schematic diagram of a three-dimensional scene-based video fusion device provided by the present application is shown. DETAILED DESCRIPTION

[0054] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0055] In the description of the embodiments of the present application, unless otherwise specified, the meaning of “a plurality of” is two or more. The association relationship of “and / or” describing the associated objects indicates that there can be three relationships, for example: A and / or B can represent the following three cases: A exists alone, A and B exist simultaneously, and B exists alone.

[0056] The “first”, “second” and the like described in the embodiments of the present application are only for the purpose of description, and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the technical features limited by “first” and “second” can explicitly or implicitly include at least one of the features.

[0057] In this document, the reference to “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily all refer to the same embodiment, nor does it necessarily refer to a particular independent or alternative embodiment, in isolation or in combination with other embodiments. It is explicitly and implicitly understood by a person skilled in the art that the embodiments described herein can be combined with other embodiments.

[0058] The present application provides a three-dimensional scene-based video fusion method, device and storage medium, which are described below respectively.

[0059] The three-dimensional scene-based video fusion method provided by the embodiments of the present application can be applied to the scene of video monitoring, such as combining real-time monitoring video pictures with three-dimensional models for display. Through this fusion, the monitoring video can be embedded into a three-dimensional environment as a part, and displayed and analyzed in a virtual scene, which helps to provide more comprehensive and accurate visual information.

[0060] The execution subject of the video fusion method based on a three-dimensional scene provided in the embodiments of the present application can be a video fusion device based on a three-dimensional scene provided in the embodiments of the present application, or a server device, a physical host or a user equipment (UE) or other types of electronic devices integrated with the video fusion device based on a three-dimensional scene, wherein the video fusion device based on a three-dimensional scene can be implemented in a hardware or software manner, and the UE can be a terminal device such as a smart phone, a tablet computer, a notebook computer, a palm computer, a desktop computer or a personal digital assistant (PDA).

[0061] Figure 1 An embodiment flow diagram of the video fusion method based on a three-dimensional scene provided in the present application is shown in FIG. 1. Figure 1 The video fusion method based on a three-dimensional scene includes the following steps.

[0062] S101, acquiring a camera installed on a motion device, collecting a plurality of frames of video frame data in a target area, each frame of the video frame data including a video frame image and a corresponding first timestamp, and acquiring spatiotemporal information of the motion device, the spatiotemporal information including geographic data and a corresponding second timestamp.

[0063] In the embodiments, the motion device is a pre-calibrated device, which can be a vehicle. The motion device is described as a vehicle in the following.

[0064] The vehicle is installed with a camera, which is used to drive in a target area of the camera, such as a monitoring area, to ensure that the collection of the video frame image covers the full field of view. A plurality of frames of video frame data are collected by the camera, and each frame of the video frame data includes a frame of video frame image collected by the camera and a corresponding timestamp, i.e., a first timestamp.

[0065] The spatiotemporal information of the motion device includes geographic data and a corresponding timestamp, i.e., a second timestamp. The geographic data can be GPS data corresponding to the center of mass of the motion device. The longitude and latitude ( λ, φ ) and the height ( h ) of the center of mass of the motion device, such as a vehicle, can be recorded by high-precision GPS data, the attitude angle (such as the pitch angle θ , the yaw angle ψ and the roll angle γ ) of the vehicle body can be obtained by inertial measurement unit data, the video is recorded by the vehicle-mounted camera, the timestamp is strictly aligned with the GPS / IMU data, and the spatiotemporal trajectory data set D is formed. p i , v i , t i )∣i = 1, 2,..., n}, wherein p i is a geographic coordinate, v i is a video frame image, t i is a timestamp, i.e., a first timestamp or a second timestamp.

[0066] The inventor found that for the calibration of a moving device, the conventional calibration method usually adopts a static target (such as a checkerboard) arranged manually in the scene, and the parameters are solved after the target image is shot by the camera, or a total station or other equipment is used to directly measure the installation position of the camera, which is time-consuming and laborious, and is difficult to adapt to the batch calibration demand of a large range scene (such as a highway, a city road network), and the high-precision total station measurement equipment has a high cost and needs to be operated by professional personnel, which is not suitable for large-scale engineering. In view of this, the calibration method of the moving device (such as a vehicle) of the embodiment can be that a high-precision GPS (such as a positioning accuracy ≤ 10 cm), an inertial measurement unit (such as an IMU angular velocity accuracy ≤ 0.1° / h) and an odometer are carried on the calibrated moving device (such as a vehicle) to drive at a preset speed (such as 5-10 km / h) through the monitoring area of the camera, so that accurate video frame data and space-time information can be obtained in real time, and the subsequent fusion efficiency is improved.

[0067] In one specific embodiment, a certain model of vehicle can be selected as the calibration vehicle chassis, an aluminum alloy support is added to the roof of the vehicle, a GNSS receiver (such as a combination of GPS and Beidou dual mode, static positioning accuracy 1 cm, dynamic positioning accuracy 10 cm, update frequency 10 Hz) is fixed, the antenna plane is kept horizontal with the roof to ensure stable satellite signal reception; an inertial measurement unit is installed in the center console of the vehicle, connected to the vehicle power supply by a hard line, fixed by a three-point support, and the IMU coordinate system is aligned with the vehicle body coordinate system (such as a deviation ≤ 0.5°); an incremental wheel speed encoder (such as a resolution of 0.1 mm) is installed on the rear axle, which rotates synchronously with the wheel through gear engagement, and the signal line is connected to the vehicle data acquisition unit to supplement the mileage data in the GPS signal blind area (such as a tunnel). The calibration vehicle drives at a certain speed (such as 8 km / h) in the camera monitoring area, the speed is stably controlled by the vehicle ECU speed limiting module, the driving track covers two-thirds of the left, middle and right areas of the camera field of view, and the single driving distance is ≥ 100 m to ensure that the video frame image acquisition covers the full field of view.

[0068] In another specific embodiment, the Precision Time Protocol (PTP) can be used to synchronize the clocks of GPS, IMU, and camera through Gigabit Ethernet, with a timestamp accuracy of 1 ms, ensuring that the time alignment error between geographic data and video frame images is < 50 ms. The data acquisition frequency is set as follows: GPS data 10 Hz, IMU data 100 Hz, camera (Basler ace 2, 1920x1050@30fps), and video frame data is recorded synchronously through hardware triggering, with each frame accompanied by a timestamp ti (in UTC millisecond format).

[0069] geographic data p i stored in the format of wherein denotes the geographic coordinates, denote longitude, latitude, and height, respectively. calculated in real time by GPS, denotes the attitude angle, denote the pitch angle, yaw angle, and roll angle, respectively. obtained by quaternion conversion by IMU; video frame v i stored in PNG format, with a resolution of 1920x1050, and the file name of each video frame image contains a timestamp t i (e.g., 20250618_143022_001.png); the spatiotemporal trajectory dataset D is stored in HDF5 format and contains three datasets: an array of geographic data (n x 6), an array of video frame paths (n x 1), and an array of timestamps (n x 1), which are associated through t i indexing to realize data association of video frame data and spatiotemporal information.

[0070] Specifically, by obtaining multiple frames of video frame data collected by a camera installed on a motion device and spatiotemporal information of the motion device, since the data is collected by using the motion device and the camera installed on the motion device, the data collection efficiency and accuracy are ensured, providing an accurate, real-time, and reliable basis for subsequent fusion of video content and three-dimensional scenes.

[0071] S102, extract image feature points of each frame of the video frame image, and construct a mapping relationship between the image feature points and the geographic data based on the first timestamp and the second timestamp.

[0072] wherein the view feature points refer to key feature points extracted from the video frame image, such as roof markers and wheel edges, which are used for efficient matching and fusion of subsequent video content.

[0073] Specifically, a feature extraction algorithm such as the ORB algorithm can be used to extract feature points corresponding to each frame of video frame image, and then a Gunnar Farneback optical flow algorithm can be used to track the feature points of adjacent frames to generate a feature point motion vector, and then a binary search algorithm can be used to match the first timestamp and the second timestamp, and a corresponding relationship between the matched image feature points and the geographic data is established to form a mapping relationship between the image feature points and the geographic data. It can be understood that, since the mapping relationship between the image feature points and the geographic data is constructed based on the timestamps, the time is accurately aligned, the image feature points can be accurately mapped to the corresponding positions in the three-dimensional scene model, and in a dynamic scene, the mapping relationship can track the changes of the image feature points over time, thereby supporting accurate rendering and fusion of dynamic objects, and solving the problems of low real-time performance and low naturalness of fusion in large-scale scenes and dynamic scenes.

[0074] In one specific embodiment, for a feature point q j with a timestamp of t j , a feature point with a timestamp of <50ms is found in the geographic data by a binary search method. Compared with a sequential search method, the binary search method has a faster search speed, and for a larger data set, the search efficiency can be significantly improved.

[0075] S103, based on the mapping relationship, iteratively solving the pose data of the camera by minimizing the re-projection error, and the target function corresponding to the re-projection error includes a strong constraint optimization term of the geographic data.

[0076] Wherein, the pose data refers to the spatial position and direction of the camera in each frame of video frame image.

[0077] The strong constraint optimization term of the geographic data refers to imposing additional constraints on the pose estimation using geographic data to ensure the accuracy and consistency of the solution.

[0078] The inventors found that the visual method based on the matching of image feature points to solve the pose data is easily affected by light, occlusion, texture missing, etc., and the error will accumulate. Although the geographic data also has noise, the absolute position accuracy in an open environment is higher than that of the image feature points. Therefore, by setting the strong constraint optimization term, the feasible region can be narrowed to the vicinity of the geographic data, thereby suppressing the drift of pure vision and reducing the number of iterations,

[0079] to ensure the global consistency of the final external parameter in a large-scale scene.

[0080] Specifically, a camera projection model can be constructed, which includes an expression of a re-projection error and an expression corresponding to a strong constraint optimization item of geographic data, which are summed as a target function corresponding to the re-projection error, and then the pose data of the camera is iteratively solved with the optimization goal of minimizing the re-projection error. It can be understood that in the embodiment, by including the strong constraint optimization item of geographic data in the target function corresponding to the re-projection error, the accurate data of the vehicle position and attitude provided by the GPS is fully utilized, the error that may be caused by relying only on visual information is reduced, the accuracy and robustness of the camera pose data solving are greatly improved, and the real-time and naturalness of subsequent video fusion are enhanced.

[0081] S104, based on the pose data, determining conversion parameters between a first coordinate system of the video frame image and a second coordinate system of a pre-constructed three-dimensional scene model, obtaining video fusion parameters to fuse the multiple frames of video frame data with the three-dimensional scene model.

[0082] The video fusion parameters include a projection matrix, wherein the projection matrix is a mathematical tool in computer vision and graphics for projecting points in a three-dimensional world onto a two-dimensional image plane, which integrates the internal parameters (such as focal length, principal point position, etc.) and external parameters (such as the position and direction of the camera in the world coordinate system) of the camera, thereby realizing the conversion from three-dimensional coordinates to two-dimensional pixel coordinates.

[0083] Specifically, according to the pose data, a conversion matrix M from the first coordinate system of the camera to the second coordinate system of the pre-constructed three-dimensional scene model is constructed, and then a projection matrix P is calculated by combining the intrinsic matrix K of the camera and the conversion matrix M, which projects points in the world coordinate system onto the image plane. Each vertex in the three-dimensional scene model is converted from the world coordinate system to the two-dimensional pixel coordinate system using the projection matrix P, and the converted two-dimensional pixel coordinates are used to render the three-dimensional scene model in the video frame image. The rendered three-dimensional scene model is fused with the image video frame, ensuring the visual consistency and continuity of the video frame image and the three-dimensional scene model, realizing the spatial alignment and dynamic superposition of the video content and the three-dimensional model, making the fusion of the three-dimensional scene model and the image video frame more real-time and natural, improving the fusion effect of the three-dimensional scene and the video, and thus improving the user immersion and interaction efficiency.

[0084] In summary, the video fusion method based on a three-dimensional scene provided by the embodiment of the application, by acquiring a camera installed on a moving device, collecting multiple frames of video frame data in a target area, each frame of video frame data including a video frame image and a corresponding first timestamp, and acquiring space-time information of the moving device, the space-time information including geographic data and a corresponding second timestamp, ensures data collection efficiency and accuracy, and provides an accurate, real-time and reliable basis for subsequent fusion of video content and a three-dimensional scene; image feature points of each frame of video frame image are extracted, and a mapping relationship between the image feature points and the geographic data is constructed based on the first timestamp and the second timestamp, ensuring accurate time alignment, which can enable the image feature points to be accurately mapped to corresponding positions in the three-dimensional scene model, and meanwhile, in a dynamic scene, the mapping relationship can track changes of the image feature points over time, thereby supporting accurate rendering and fusion of dynamic objects; based on the mapping relationship, pose data of the camera is iteratively solved by minimizing a re-projection error, the target function corresponding to the re-projection error including a strong constraint optimization item of the geographic data, which fully utilizes accurate data of a vehicle position and attitude provided by GPS, reduces errors that can be caused by relying on visual information only, greatly improves accuracy and robustness of solving of the pose data of the camera, and is beneficial to enhancing real-time performance and naturalness of subsequent video fusion; based on the pose data, conversion parameters between a first coordinate system of the video frame image and a second coordinate system of the three-dimensional scene model pre-constructed are determined, video fusion parameters are obtained, and the multiple frames of video frame data are fused with the three-dimensional scene model, ensuring consistency and coherence of the video frame image and the three-dimensional scene model in vision, realizing spatial alignment and dynamic superposition of the video content and the three-dimensional model, making fusion of the three-dimensional scene model and the image video frame more real-time and more natural, and improving fusion effect of the three-dimensional scene and the video.

[0085] In some embodiments of the application, before step S103, the method further comprises:

[0086] S201, determining a penalty item corresponding to the strong constraint of the geographic data;

[0087] S202, obtaining a target function corresponding to the re-projection error based on the penalty item and an expression corresponding to the re-projection error.

[0088] Specifically, in order to determine the target function corresponding to the re-projection error, it is necessary to first calculate a projection equation, which can use the vehicle centroid coordinates, i.e. the coordinates corresponding to the geographic data p i In the three-dimensional scene model coordinate system, is converted into a Cartesian coordinate (x, y, z) through UTM projection X i , Y i ,Z i), i.e. X i , Y i ,Z i ) is the position of the vehicle body center in the Cartesian coordinate system, and the conversion formula is as follows:

[0089]

[0090] wherein , is the central point longitude and the central point latitude of the three-dimensional scene model, is the average latitude, is the earth radius (6371km);

[0091] The conversion matrix of the vehicle body center to the camera coordinate system is as follows:

[0092]

[0093] The projection equation wherein s is a scaling factor; q j is the pixel coordinate on the video frame image; is a 4x4 transformation matrix, including a 3x3 rotation matrix R and a translation vector P t , used to describe the position and direction of the camera relative to the world coordinate system; T is a 3x1 translation vector, and the two together constitute the to-be-solved external parameter, K is the camera intrinsic parameter matrix, and an example is as follows:

[0094]

[0095] Then, the geographic data strong constraint optimization is performed, the target function introduces the GPS weight α=200, and the following equation is constructed:

[0096]

[0097] wherein q i is the pixel coordinate corresponding to the image feature point of the video frame image collected by the camera, which is obtained from the mapping relationship in step 102; p i is the coordinate corresponding to the geographic data; K is the known camera intrinsic parameter matrix; is the translation vector of the first point provided by the GPS, which is the known reference position.

[0098] The target function is obtained by summing the expressions corresponding to the penalty term and the re-projection error by determining the strong constraint of the geographic data, so as to force the iterative optimization to approach the GPS, and prevent the visual mismatch or drift from reducing the parameter calculation accuracy.

[0099] In some embodiments of the present application, step S103 comprises:

[0100] S301, taking the mapping relationship as observation data of an iterative solving process, minimizing the target function as an optimization target, and iteratively solving by using a nonlinear least squares optimization algorithm to obtain the pose data of the camera, wherein a sparse matrix solver is used in each iteration.

[0101] The nonlinear least squares optimization algorithm can be a Levenberg-Marquardt algorithm.

[0102] The inventors have found that the prior art method uses full matrix decomposition to calculate the pose data, without taking advantage of the sparsity of the three-dimensional scene (such as the fact that most areas in a building model are planar structures), wasting computing resources. Therefore, the embodiments of the present application use the SimplicialLDLT solver of the Eigen library for sparse matrix decomposition to improve the solving efficiency and thus improve the fusion real-time performance.

[0103] Specifically, the Levenberg-Marquardt algorithm can be finally used for iterative optimization, the SimplicialLDLT solver of the Eigen library is used for sparse matrix decomposition, and the Cholesky decomposition implementation for efficiently solving sparse linear equations is used to calculate the pose data of the camera.

[0104] In one example, the iterative solving method of the present embodiment is used, and the iteration converges after 3-5 times, the calculation time is <50ms per camera, the position error is ≤20cm, and the attitude angle error is ≤1°.

[0105] In some embodiments of the present application, step S102 comprises:

[0106] S401, extracting feature points of the motion device from each frame of the video frame image by using an improved ORB algorithm, the improved ORB algorithm comprising at least one of adjusting a preset corner point response threshold, introducing a scale pyramid to extract feature points, and retaining feature points by non-maximum suppression, and improving the dimension of feature descriptors.

[0107] S402, tracking the motion of the feature points in adjacent video frame images by using an optical flow method to generate feature point motion vectors, and obtaining the image feature points.

[0108] The optical flow algorithm can be a dense optical flow algorithm, such as a Gunnar Farneback dense optical flow algorithm.

[0109] The inventors have found that existing image feature extraction algorithms, such as the ORB (Oriented FAST and Rotated BRIEF) algorithm, can affect the accuracy of feature point extraction in complex lighting, dynamic scenes, or low-texture areas, thereby leading to low accuracy of feature point matching between video frames and three-dimensional models and large deviations in fusion parameters (such as translation error > 5 cm), which results in poor feature matching robustness. Therefore, in the present embodiment, an improved ORB algorithm is used for feature extraction, and the improvements include at least one of the following: adjusting the response threshold of the preset corner point, introducing a scale pyramid to extract feature points, and retaining feature points through non-maximum suppression, and increasing the dimension of the feature descriptor.

[0110] Specifically, the improved ORB algorithm is used to extract feature points of the motion equipment from each frame of video frame image, and the optical flow method is used to track the motion of the feature points in adjacent video frame images, to generate feature point motion vectors and obtain image feature points, which can enhance the robustness of the feature descriptor to image changes (such as lighting, scale, and rotation) and improve the accuracy of feature matching.

[0111] In one specific embodiment, the OpenCV 4.7.0 library can be used to improve the ORB algorithm as follows: extract a frame of video frame image every 200 ms according to the timestamp ti, a total of 50-50 frames per camera, to ensure that the feature points cover the key positions of the calibration vehicle driving track, then adjust the FAST corner point response threshold to 20 (the traditional value is 10) to reduce false detection points in low-contrast areas; introduce a scale pyramid (4 layers), extract 200 key points from each layer, and retain the main direction feature points through non-maximum suppression; the feature descriptor uses the BRIEF algorithm, and the descriptor dimension is increased from 128 bits to 256 bits to enhance the feature discrimination.

[0112] In one specific embodiment, as shown in Figure 2 the schematic diagram of the image feature point extraction process is as follows:

[0113] First, the video data is taken as input, key frame selection is performed, and key frames are selected from the video for processing. There are two selection methods: fixed interval (such as every 200 ms) selection of a frame for image preprocessing, or dynamic interval adjustment with an error of less than 5 ms. Then, the selected key frames are preprocessed, such as grayscale, denoising, etc. Next, improved ORB feature extraction is performed, that is, the improved ORB algorithm is applied to extract image feature points, such as using the FAST algorithm to detect corners in the image, adaptive threshold < 20, that is, adjusting the threshold of the FAST detection to be less than 20 to adapt to different image conditions, applying non-maximum suppression to ensure that the detected feature points are locally optimal, constructing a scale pyramid of the image for multi-scale feature detection, calculating a 256-bit BRIEF descriptor for each key point for feature matching, and enhancing the descriptor to improve the accuracy and robustness of feature matching. Anti-blurring processing is performed on the feature points to improve the feature extraction effect in blurred images. Feature point verification can also be performed to verify whether the number of extracted feature points reaches 200. If not, new data needs to be collected, and if the number of feature points is sufficient, optical flow tracking is performed to track the motion of the feature points in consecutive frames, and the final feature matching pair is output for subsequent three-dimensional scene fusion or other computer vision tasks.

[0114] In some embodiments of the application, step S102 comprises:

[0115] S501, for each image feature point, based on the first timestamp and the second timestamp, determine the initial matching geographical data to form a plurality of initial mapping pairs;

[0116] S502, based on the initial mapping pair, projection error calculation is performed, and according to the projection error result, the random sample consensus algorithm is used to remove the initial mapping pair that does not meet the matching condition;

[0117] S503, when the number of initial mapping pairs that meet the matching condition is less than the preset number threshold, continue to obtain the image feature points corresponding to the new video frame image, until the number of initial mapping pairs that meet the matching condition is greater than or equal to the preset number threshold, and the mapping relationship is obtained.

[0118] Specifically, for each image feature point, timestamp matching is performed based on the first timestamp and the second timestamp, initial geographic data of matching is determined, a plurality of initial mapping pairs are constituted, then projection error calculation is performed on each initial mapping pair, the projection error calculation can convert the geographic coordinates into image coordinates and compare with the actual image feature point position, the random sample consensus (RANSAC) algorithm is used to remove the initial mapping pairs that do not meet the matching condition, when the number of initial mapping pairs meeting the matching condition is less than a preset number threshold, then the image feature points corresponding to the new video frame image are continuously acquired, the steps of S501-502 are repeated to construct the mapping relationship, until the number of initial mapping pairs meeting the matching condition is greater than or equal to the preset number threshold, and the mapping relationship is generated.

[0119] In one specific embodiment, RANSAC mismatch removal can be used, the number of iterations is set to 500 times, and the threshold is set to 3 pixels (projection error > 3 pixels, considered as mismatch). The specific process is as follows:

[0120] S1: randomly select 4 mapping points, and calculate the basis matrix F;

[0121] S2: use F to calculate the projection error of all mapping points, and keep the point pairs with error < 3 pixels;

[0122] S3: repeat S1-S2, record the time when the number of retained point pairs is the most as the final matching set, and require the number of matching pairs to be ≥ 200 pairs, otherwise reacquire data.

[0123] In some embodiments of the application, step S104 comprises:

[0124] S601, converting the first coordinate system of the video frame image into the Cartesian coordinate system of the three-dimensional scene model through UTM projection based on the pose data;

[0125] S602, determining a scaling factor between mapping the three-dimensional size of the three-dimensional scene model to the two-dimensional image pixel size of the video frame image;

[0126] S603, constructing the attitude rotation matrix corresponding to the pose data according to the Cartesian coordinate system and the scaling factor;

[0127] S604, constructing the projection error equation of the three-dimensional scene model and the video frame image based on the attitude rotation matrix using the optical flow method;

[0128] S605, predicting the optimal pose increment using Kalman filtering according to the projection error equation;

[0129] S606, updating the pose data by superimposing the optimal pose increment on the current pose data.

[0130] S607, solving the conversion parameter by minimizing the projection error equation based on the updated pose data;

[0131] S608, generating the video fusion parameter according to the conversion parameter and the scaling factor.

[0132] Wherein, the scaling factor is determined according to the ratio between the pixel of the video frame image and the size of the motion device.

[0133] Specifically, first, the coordinate conversion matrix is constructed, and the camera geographic coordinates are converted to the three-dimensional scene origin O3D = (X0, Y0, Z0) by UTM projection; the attitude angle is used to construct the rotation matrix R in the order of Z-Y-X:

[0134] where:

[0135]

[0136]

[0137]

[0138] Translation vector ;

[0139] The scaling factor s is calculated by the actual width of the vehicle (such as 2.4m) and the video imaging width (pixels):

[0140]

[0141] Wherein is the physical size of the pixel, unit m / pixel.

[0142] For the coordinate mapping deviation of two-dimensional video frame image and three-dimensional scene model, a differential correction model is introduced to dynamically optimize the projection relationship, and inter-frame motion prediction is performed; Kalman filter is used to predict the next frame camera pose increment , the state vector is 12-dimensional , the process noise covariance matrix, , the differential correction equation is established, and the projection correction model is established based on the optical flow field Δq:

[0143]

[0144] Wherein, q is the original projection pixel coordinate, is the corrected coordinate, and the optical flow error is minimized solving ΔR , ΔT updating the projection matrix in real time P = K [R + ΔR | T + ΔT] -1 to ensure the spatial mapping accuracy of the two-dimensional video frame image and the three-dimensional scene model.

[0145] In one specific embodiment, the video fusion parameter calculation process is as follows:

[0146] Preprocessing data: the process starts with preprocessing data, including GPS / IMU (Inertial Measurement Unit) and video stream.

[0147] Feature extraction: feature points are extracted from the video stream through an improved ORB algorithm.

[0148] Space-time alignment: PTP synchronization technology is used to ensure accurate alignment of timestamps.

[0149] Establishing mapping relationship: the extracted feature points are mapped with GPS / IMU data to construct a projection model.

[0150] Parameter solving: the spatial position and attitude parameters of the camera are obtained through the solving process.

[0151] GPS constraint optimization: the solving results are optimized using GPS data.

[0152] Parameter verification: the parameters obtained by solving are verified, and if they are not qualified, data needs to be re-collected.

[0153] Three-dimensional coordinate conversion: if the parameter verification is qualified, three-dimensional coordinate conversion is performed.

[0154] Double-layer optimization: including cross-modal differential correction, geographical hard constraint and visual soft optimization, to improve the naturalness and real-time performance of fusion.

[0155] Real-time enhancement: CUDA acceleration and dynamic termination strategy are used to improve the real-time performance of calculation.

[0156] Fusion output: the fused video is finally output, in which the video content has realized accurate spatial alignment with the three-dimensional scene.

[0157] In some embodiments of the present application, after step 104, further comprising:

[0158] S701, constructing a first target function and a second target function, the first target function being a first expression of the absolute value of the difference between the optimized position point and the projected theoretical position point, and the second target function being the optical flow error of the video frame image and the structural similarity error of the video frame image;

[0159] S702、based on the first objective function and the second objective function, iteratively optimizing the video fusion parameter.

[0160] Specifically, in order to improve the accuracy of the video fusion parameter, a geographic information hard constraint optimization method combined with a visual consistency soft optimization method can be used to iteratively optimize the video fusion parameter. The process of iterative optimization is as follows:

[0161] The geographic information hard constraint optimization layer, i.e. the first objective function, is:

[0162]

[0163] The constraint condition is:

[0164]

[0165] p is the three-dimensional point obtained by mapping; P GPS is the corresponding GPS measurement position.

[0166] Sparse Cholesky decomposition is used for solving, and the single iteration time is <15 ms, which prioritizes ensuring the projection accuracy of three-dimensional coordinates;

[0167] Next, the visual consistency soft optimization layer is determined, and the TV-L1 optical flow algorithm is used to calculate the optical flow error Eflow. The parameter settings are: the number of pyramid layers: 3 layers, the time step: 0.5, and the smoothing parameter: 0.012.

[0168] The SSIM error Essim is calculated using the skimage library function, the window size is 11x11, different weights are given to the weight channels (such as the red, green, and blue channels of an RGB color image), the target SSIM is ≥0.94, and the function is:

[0169]

[0170] where I represents the original image, represents the image obtained by projection model transformation.

[0171] The comprehensive loss function, i.e. the second objective function, is:

[0172]

[0173] Real-time optimization is performed using Kalman filtering, and the state vector [R, Pt, s] is set to (12 dimensions), the process noise covariance Q = diag (0.01, 0.01,..., 0.01) , and the observation noise covariance R = diag (1, 1,..., 1)The initial prediction error is ≤5%; setting up GPU parallelism allows for sparse matrix decomposition based on CUDA 12.1, improving matrix operation speed and reducing single-frame solution time; when Alternatively, the iteration can be terminated after a preset number of iterations (e.g., 3 times) to further reduce the amount of computation.

[0174] It is worth noting that the real-time performance of video fusion parameter calculation can be further improved through CUDA acceleration and dynamic termination strategies, thereby enhancing the naturalness and real-time performance of video-to-3D scene fusion.

[0175] In some embodiments of the present invention, step 702 includes:

[0176] S501. In each iteration, the first objective function is minimized as the optimization objective, and the first solution result is obtained.

[0177] S502. Using the first solution result as a constraint and minimizing the second objective function as the optimization objective, the video fusion parameters are iteratively optimized.

[0178] Specifically, in each iteration, the first objective function is minimized as the optimization objective to obtain the first solution result. The video fusion parameters are iteratively optimized using the first solution result as a constraint and the second objective function as the optimization objective. This allows for optimization with geographic information hard constraints as the main optimization direction, further reducing visual errors and improving the accuracy of video fusion parameters.

[0179] In a specific implementation, the iterative optimization process for video fusion parameters is as follows:

[0180] Input camera parameters and a 3D scene model, initialize the fusion parameters θ0 and transformation matrix M0, extract feature points from video frames, use an improved ORB algorithm for feature point matching to ensure at least 200 matching pairs, analyze projection errors to provide a basis for subsequent optimization, use geographic information for hard constraint optimization to ensure that the actual geographic location is considered during the optimization process, introduce GPS data for strong constraints, and set the weight α to 200 to ensure the accuracy of position and pose.

[0181] Visual soft optimization: Calculating optical flow error using the TV-L1 optical flow algorithm. E flow Optimize feature point matching.

[0182] Calculate SSIM error using skimage library functions. E ssim To ensure visual quality, with a target SSIM value ≥ 0.94, a comprehensive loss function is constructed. E2 = 0.7E flow +0.3Essim , combined with optical flow error and SSIM error.

[0183] Real-time optimization: dynamic prediction using Kalman filter, setting state vector and process noise covariance matrix, prediction initial value error ≤5%.

[0184] CUDA acceleration: set GPU parallel computing, implement sparse matrix decomposition based on CUDA 12.1, improve matrix operation speed.

[0185] Dynamic termination condition: terminate when the error change of two consecutive iterations is less than 10 -4 or 3 times of iteration, reduce 60% of calculation amount.

[0186] Update matrix and rendering, update transformation matrix M, and perform three-dimensional scene rendering.

[0187] Frame rate check: check if the frame rate reaches 30fps, if yes, perform real-time fusion output; otherwise, adopt dynamic degradation strategy.

[0188] Through the above steps, the fusion parameters θ = (R, T, s) are optimized, ensuring accurate fusion of video and three-dimensional scene, while meeting real-time requirements.

[0189] In order to better implement the video fusion method based on three-dimensional scene in the embodiment of the application, on the basis of the video fusion method based on three-dimensional scene, as shown in Figure 3 , the embodiment of the application also provides a video fusion device based on three-dimensional scene, which comprises:

[0190] An acquisition unit 301 is configured to acquire a camera installed on a motion device, collect a plurality of frames of video frame data in a target area, each frame of the video frame data comprising a video frame image and a corresponding first timestamp, and acquire spatio-temporal information of the motion device, the spatio-temporal information comprising geographic data and a corresponding second timestamp.

[0191] A construction unit 302 is configured to extract image feature points of each frame of the video frame image, and construct a mapping relationship between the image feature points and the geographic data based on the first timestamp and the second timestamp.

[0192] A solving unit 303 is configured to solve pose data of the camera based on the mapping relationship by minimizing a re-projection error, the target function corresponding to the re-projection error comprising a strong constraint optimization term of the geographic data.

[0193] The fusion unit 304 is configured to determine, based on the pose data, a conversion parameter between a first coordinate system of the video frame image and a second coordinate system of a pre-constructed three-dimensional scene model, to obtain a video fusion parameter, so as to fuse the multi-frame video frame data with the three-dimensional scene model.

[0194] The video fusion device based on a three-dimensional scene provided in the above embodiments can implement the technical solutions described in the above video fusion method embodiments based on a three-dimensional scene. The principles of the implementation of the above modules or units can be found in the corresponding content in the above video fusion method embodiments based on a three-dimensional scene, which will not be described here again.

[0195] Correspondingly, the embodiments of the present application further provide a computer readable storage medium, which is configured to store computer readable programs or instructions. When the programs or instructions are executed by a processor, the steps or functions in the video fusion method based on a three-dimensional scene provided in the above embodiments can be implemented.

[0196] Those skilled in the art can understand that all or part of the processes of the above embodiments can be completed by a computer program to instruct related hardware (such as a processor, a controller, etc.) to complete. The computer program can be stored in a computer readable storage medium. The computer readable storage medium includes a magnetic disk, an optical disk, a read-only memory, a random access memory, etc.

[0197] The above describes the video fusion method, device and storage medium based on a three-dimensional scene in detail. The principle and implementation manner of the present application are described by using specific examples. The above embodiment is only used to help understand the method and core idea of the present application. Meanwhile, for those skilled in the art, the specific implementation manner and application range can be changed according to the idea of the present application. The above description should not be understood as a limitation of the present application.

Claims

1. A method of video fusion based on a three-dimensional scene, characterized in that, The method comprises: acquiring a camera mounted on a motion device, collecting a plurality of frames of video frame data in a target area, each frame of the video frame data comprising a video frame image and corresponding first time stamps, and acquiring space-time information of the motion device, the space-time information comprising geographic data and corresponding second time stamps; extracting image feature points of each frame of the video frame image, and constructing a mapping relationship between the image feature points and the geographic data based on the first time stamps and the second time stamps; based on the mapping relationship, iteratively solving pose data of the camera by minimizing a re-projection error, the target function corresponding to the re-projection error comprising a strong constraint optimization term of the geographic data; based on the pose data, determining conversion parameters between a first coordinate system of the video frame image and a second coordinate system of a pre-constructed three-dimensional scene model, obtaining video fusion parameters, and fusing the plurality of frames of video frame data with the three-dimensional scene model.

2. The method of claim 1, wherein, Before the step of based on the mapping relationship, iteratively solving pose data of the camera by minimizing a re-projection error, the method further comprises: determining a penalty term corresponding to the strong constraint of the geographic data; based on the penalty term and an expression corresponding to the re-projection error, obtaining a target function corresponding to the re-projection error.

3. The method of claim 2, wherein, The step of based on the mapping relationship, iteratively solving pose data of the camera by minimizing a re-projection error comprises: using the mapping relationship as observation data in an iterative solving process, minimizing the target function as an optimization objective, and iteratively solving by using a nonlinear least squares optimization algorithm to calculate the pose data of the camera, wherein a sparse matrix decomposer is used in each iteration.

4. The method of claim 1, wherein, The step of extracting image feature points of each frame of the video frame image comprises: extracting feature points of the motion device from each frame of the video frame image by using an improved ORB algorithm, the improved ORB algorithm comprising at least one of adjusting a response threshold of a preset corner point, introducing a scale pyramid to extract feature points, and retaining feature points by non-maximum suppression, and improving the dimension of a feature descriptor; tracking motion of the feature points in adjacent video frame images by using an optical flow method to generate feature point motion vectors, and obtaining the image feature points.

5. The method of claim 1, wherein, The step of constructing a mapping relationship between the image feature points and the geographic data based on the first time stamps and the second time stamps comprises: for each image feature point, determining matching initial geographic data based on the first time stamps and the second time stamps to form a plurality of initial mapping pairs; performing projection error calculation based on the initial mapping pairs, and removing initial mapping pairs that do not meet a matching condition by using a random sample consensus algorithm according to a projection error result; when the number of initial mapping pairs that meet the matching condition is less than a preset number threshold, continue to acquire image feature points corresponding to new video frame images until the number of initial mapping pairs that meet the matching condition is greater than or equal to the preset number threshold, and obtain the mapping relationship.

6. The method of claim 1, wherein, The conversion parameters between the first coordinate system of the video frame image and the second coordinate system of the pre-constructed three-dimensional scene model are determined based on the pose data, and video fusion parameters are obtained, including: The first coordinate system of the video frame image is converted into the Cartesian coordinate system of the three-dimensional scene model through UTM projection based on the pose data; A scaling factor between the three-dimensional size of the three-dimensional scene model and the two-dimensional image pixel size of the video frame image is determined; An attitude rotation matrix corresponding to the pose data is constructed according to the Cartesian coordinate system and the scaling factor; A projection error equation of the three-dimensional scene model and the video frame image is constructed by using an optical flow method based on the attitude rotation matrix; An optimal pose increment is predicted by using a Kalman filtering method according to the projection error equation; The optimal pose increment is added to the current pose data, and the pose data is updated; The conversion parameters are solved by minimizing the projection error equation based on the updated pose data; The video fusion parameters are generated according to the conversion parameters and the scaling factor.

7. The method of claim 1, wherein, After the conversion parameters between the first coordinate system of the video frame image and the second coordinate system of the pre-constructed three-dimensional scene model are determined based on the pose data, and the video fusion parameters are obtained, the method further includes: A first target function and a second target function are constructed, the first target function is a first expression of the absolute value of the difference between the optimized position point and the projected theoretical position point, and the second target function is the optical flow error of the video frame image and the structural similarity error of the video frame image; The video fusion parameters are iteratively optimized based on the first target function and the second target function.

8. The method of claim 7, wherein, The video fusion parameters are iteratively optimized based on the first target function and the second target function, including: In each iteration, a first solving result is obtained by solving the first target function as the optimization objective to minimize the first target function; The video fusion parameters are iteratively optimized by taking the first solving result as a constraint condition and minimizing the second target function as the optimization objective.

9. A video fusion apparatus based on a three-dimensional scene, characterized by, It includes: An acquisition unit is configured to acquire a camera installed on a motion device, collect multiple frames of video frame data in a target area, each frame of the video frame data including a video frame image and a corresponding first timestamp, and acquire spatiotemporal information of the motion device, the spatiotemporal information including geographic data and a corresponding second timestamp; A construction unit is configured to extract image feature points of each frame of the video frame image, and construct a mapping relationship between the image feature points and the geographic data based on the first timestamp and the second timestamp; A calculation unit is configured to iteratively calculate pose data of the camera by minimizing a re-projection error based on the mapping relationship, the target function corresponding to the re-projection error including a strong constraint optimization term of the geographic data; A fusion unit is configured to determine conversion parameters between a first coordinate system of the video frame image and a second coordinate system of a pre-constructed three-dimensional scene model based on the pose data, and obtain video fusion parameters to fuse the multiple frames of video frame data with the three-dimensional scene model.

10. A computer-readable storage medium, characterized in that, A computer readable storage medium storing a program or instructions which, when executed by a processor, implement the steps of the three-dimensional scene-based video fusion method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • High-precision low-drift large-range three-dimensional point cloud map construction and repositioning method

    CN116399354A

  • Video scene mapping and positioning method and device, electronic equipment and storage medium

    CN116485894A