Method for quickly calibrating video stream and projecting video stream to three-dimensional scene for virtual-real fusion

By marking three points in video and three-dimensional scenes and using FOV parameters, the problem of picture splitting and dispersion in video surveillance is solved, and the efficient integration of video streams and three-dimensional scenes is achieved, the calibration efficiency and projection accuracy are improved, and it is suitable for large-scale monitoring systems.

CN120355872APending Publication Date: 2025-07-22上海漂视网络股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510575076.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, there is a problem of too many surveillance images and centralized surveillance in matrix-shaped separation and dispersion, making it difficult to effectively integrate video streams and three-dimensional scenes.

Method used

By marking the same three points A, B, and C in the video picture and the three-dimensional virtual scene, and combining the camera's field of view (FOV) parameters, the three-dimensional model is transformed and aligned, and the video stream is quickly calibrated and projected into the three-dimensional scene.

Benefits of technology

It realizes efficient integration of video streams and three-dimensional scenes, reduces calculation amount, improves calibration efficiency and projection accuracy, is suitable for large-scale monitoring systems, and provides intuitive monitoring experience and more comprehensive information display.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355872A_ABST
    Figure CN120355872A_ABST
Patent Text Reader

Abstract

The invention discloses a method for quickly calibrating a video stream and projecting the video stream to a three-dimensional scene for virtual-real fusion, which comprises the following steps of: S1, selecting a video picture, and displaying a frame of a video as a background; s2, marking points in reality and virtuality: respectively marking three identical points A, B and C in the video picture and the three-dimensional virtual scene; s3, transforming and aligning a three-dimensional model, constructing equal-proportion scaling three-dimensional model transformation, and realizing alignment with a video picture through model offset calculation; and S4, inputting the FOV of the camera, and obtaining the FOV parameter of the monitoring camera to complete the watching distance consistent with the monitoring in the real world. The invention relates to the technical field of video monitoring. According to the method for quickly calibrating the video stream and projecting the video stream to the three-dimensional scene for virtual-real fusion, through a reverse calibration technology, only three corresponding points need to be marked in a video and a three-dimensional model, and the problems that monitoring pictures are too many and centralized monitoring is split and dispersed in a matrix shape are solved in combination with a camera field angle (FOV) parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video surveillance, and particularly to a method for quickly calibrating a video stream and projecting it onto a three-dimensional scene for virtual-real fusion. Background Art

[0002] The rapid progress of computer technology has laid a solid foundation for the birth of the technology of virtual-real fusion in three-dimensional scenes. The increasing maturity of three-dimensional modeling technology and the continuous progress of high-definition video processing technology have provided strong technical support for this technology. The rise of virtual reality and augmented reality has further promoted the development of the technology of fusing three-dimensional models and videos.

[0003] Especially with the acceleration of the urbanization process, the comprehensive management of urban space faces more and more challenges. Advanced graphics processing technology can build accurate three-dimensional models, and at the same time, the continuous improvement of video processing technology has also created conditions for the fusion of the two.

[0004] In the aspect of video surveillance, the technology of virtual-real fusion between video streams and three-dimensional scenes has a wide range of application requirements. It can solve the problems that as the number of surveillance cameras increases, centralized duty personnel face more and more video surveillance windows, the surveillance points are densely distributed and scattered, and the surveillance points and physical locations cannot be intuitively corresponding. In the initial stage of the application of video streams, technically, it is allowed to fuse the positions of video surveillance cameras with three-dimensional scenes, which can solve the problem of the intuitive correspondence between the camera positions and urban spaces, but still faces the problems of too many surveillance images and the matrix-shaped fragmentation and dispersion of centralized surveillance.

[0005] Therefore, the present invention provides a method for quickly calibrating a video stream and projecting it onto a three-dimensional scene for virtual-real fusion. Projecting the video stream onto a three-dimensional scene for virtual-real fusion can provide more comprehensive and accurate information for urban security surveillance. Surveillance personnel can more intuitively understand the overall situation of the city and discover abnormalities in a timely manner. When planning urban spaces, this technology can also assist managers in better planning and managing video surveillance resources in the city. Summary of the Invention

[0006] Aiming at the deficiencies of the prior art, the present invention provides a method for quickly calibrating a video stream and projecting it onto a three-dimensional scene for virtual-real fusion, solving the problems of too many surveillance images and the matrix-shaped fragmentation and dispersion of centralized surveillance.

[0007] To achieve the above object, the present invention is realized through the following technical solutions: A method for quickly calibrating a video stream and projecting it onto a three-dimensional scene for virtual-real fusion, comprising the following steps: S1. Select a video frame and display one frame of the video as the background; S2. Mark points in the real and virtual worlds: Mark the same three points A, B, and C in the video frame and the 3D virtual scene respectively; S3. Transformation and alignment of the 3D model, construct a 3D model transformation with equal proportion scaling, and achieve alignment with the video frame through model offset calculation; S4. Input the camera FOV, and use the FOV (Field of View) parameter of the surveillance camera obtained in advance to complete the viewing distance consistent with the surveillance camera in the real world.

[0008] Preferably, the specific steps of S1 are as follows; S11. First, establish a spatial coordinate system, and first select the origin; S12. Then define the coordinate axes (X-axis, Y-axis, Z-axis); The X-axis is usually defined as the direction horizontally to the right; The Y-axis is usually defined as the direction vertically upward; The Z-axis is usually defined as the direction perpendicular to the XY plane, pointing "upward"; S13. Then determine the positive directions of the coordinate axes; S14. Finally, establish the unit length of the coordinate axes: Preferably, the method for establishing the coordinate system is as follows; First, establish a coordinate system using three mutually perpendicular edges sharing a common vertex. Select a vertex as the origin, and define the three coordinate axes through the three mutually perpendicular edges of this vertex; Then, establish a coordinate system using the perpendicularity of a line and a plane. Select a straight line as the Z-axis, a plane perpendicular to the Z-axis as the XY plane, then select a point on the XY plane as the origin, and the straight line perpendicular to this plane as the Y-axis; Secondly, establish a coordinate system using the perpendicularity of two planes. Select two perpendicular planes, determine the Z-axis through the intersection line of these two planes, then select a point on one plane as the origin, and define the X-axis, and finally determine the Y-axis; Finally, establish a coordinate system using the symmetry relationship in the figure. In a figure with symmetry, the coordinate axes can be defined using the axis of symmetry, and the origin is usually located at the center of symmetry.

[0009] Preferably, the specific steps of S2 are as follows; S21. Mapping of 3D space points to the plane image. First, three points V1, V2, and V3 in space form a plane. According to the three points in space, the plane equation can be determined by the point-normal form: Ax + By + Cz - (Ax0 + By0 + Cz0) = 0, where the normal vector N = (V3 - v1) × (v2 - v1) = (A, B, C), and P0(x0, y0, z0) is any point on the plane; Next, project the point P1(x1, y1, z1) in the current space onto the plane of the surface V1V2V3 to obtain the coordinate Pp. That is, draw a straight line from point P1 along the normal vector N to intersect the plane, and the intersection point is the required projection point. The parametric equation of the drawn straight line is; where (m, n, p) is the direction, N=(A, B, C) is the normal vector of the surface to be projected, P1(x1, y1, z1) is the point to be projected, and t is the parameter of the straight line equation. Therefore, the coordinates of the projection point Pp=(x1+mt, y1+nt, z1+pt). By finding t, the projection point can be obtained. Since the projection point Pp is on the plane of the surface V1V2V3, substitute it into the plane equation; A(x1+mt)+B(y1+nt)+C(z1+pt)=Ax0+By0+Cz0; Expand to get t(Am + Bn + Cp)=(Ax0 + By0 + Cz0)-(Ax1 + By1 + Cz1) Obtain t = ((Ax0 + By0 + Cz0)-(Ax1 + By1 + Cz1)) / (Am + Bn + Cp) Since (m, n, p)=(A, B, C), then t=(N.dot(P0)-N.dot(P1)) / (N.dot(N)). Substitute t into Pp=(x1+mt, y1+nt, z1+pt)=P1+tN to obtain the projection point; S22. The calculation of the projection plane is as follows; Three points determine a plane. Through the self - marked mirror image coordinate points A', B', C' in the three - dimensional space, construct a matrix and use the Gaussian elimination method in linear algebra to determine a plane; Normal vector: First, calculate the vector determined by points A' and B', and then calculate the vector determined by points A' and C'. The cross - product of these two vectors will give the normal vector N of the plane; Plane equation: After obtaining the normal vector N, the plane can be represented by the point - normal form equation: N•(P - A') = 0, where P is any point on the plane and A' is a known point on the plane.

[0010] The calculation of the center point of the two - dimensional plane is as follows; Calculate the center point O of the two - dimensional plane. Use the basic formulas x = width / 2, y = height / 2 in computer image vision to obtain the center point.

[0011] Preferably, the specific steps of S3 are as follows; S31. Transformation of the three - dimensional model; Calculate the relative scaling ratio S = (A - B) / (A' - B'), and generate a flat model in three-dimensional space with length = videowidth, width = videoheight, and height = 0. The center of the model is at the upper left corner; Scale this model by 1 / S times. After creating or importing a three-dimensional model, it is necessary to adjust its size to fit the scene or meet specific visual effects. Scaling transformation can be used to quickly adjust the size of the model. Scaling is divided into one of uniform scaling and non-uniform scaling. Scaling can be expressed by establishing a matrix; Rotate this model. The rotation angles in the X, Y, and Z directions are based on the X, Y, and Z component values of the differences between B' - A' and C' - A'. In the XY, YZ, and XZ planes, use the Pythagorean theorem to find the corresponding angles and calculate the rotation angles using arcsine or arccosine; S32, model offset and alignment; Aligning the video frame with the three-dimensional scene through model transformation (rotation, offset, translation) is based on the following logic: Solve the rotation relationship between two coordinate systems through the rotation between the matrices established by two sets of coordinates, and then solve the scale and translation amount between the two coordinate systems through the translation between the two sets of matrices. If the rotation correspondence between the two matrices is not good enough, it will cause deviations in the aligned coordinate points. If the correspondence is good, there will be no calculation deviations in the aligned coordinate points; Scaling transformation can be represented by a 3x3 matrix. For uniform scaling, the matrix form is as follows: | Sx 0 0 | | 0 Sy 0 | | 0 0 Sz |.

[0012] Preferably, the specific steps of S4 are as follows; S41. Calculation of the viewing distance; Determine the geometric center of the flat model. First, it is necessary to determine the geometric center point of the flat model (such as a screen or a projection plane); Calculate the normal direction. The normal direction is the vertical direction pointing outward from the geometric center point; Determine the field of view (FOV). The field of view is the angle of the visual range captured by the monitoring camera lens. The parameter of FOV is manually entered by the user for the real monitoring camera in "processing step S4; Then apply the Pythagorean theorem to calculate the viewing distance (L), which is the extension distance L in the normal direction of the positive center point of the rectangular flat model. The formula is L = (flat width / 2) / tan(FOV / 2).

[0013] The present invention provides a method for quickly calibrating a video stream and projecting it onto a three-dimensional scene for virtual-real fusion. Compared with the prior art, it has the following beneficial effects: 1. For the method of quickly calibrating a video stream and projecting it onto a three-dimensional scene for virtual-real fusion, through reverse calibration technology, only 3 corresponding points need to be marked in the video and the three-dimensional model. Combining with the camera field of view (FOV) parameters, the projection relationship can be automatically calculated. Compared with the traditional method, there is no need for complex feature matching, the calculation amount is reduced by more than 30%, the calibration efficiency is increased by 50%, and the projection accuracy is higher (error < 0.5%). It is especially suitable for the rapid virtual-real fusion of large-scale monitoring systems and three-dimensional scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a schematic diagram of the projection of the video stream screen of the present invention onto a three-dimensional scene for virtual-real fusion; Figure 2 It is a schematic diagram of the mapping of three-dimensional space points to a planar image of the present invention; Figure 3 It is a schematic diagram of the projection plane P of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] The technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0016] Please refer to Figures 1-3 , the embodiments of the present invention provide a technical solution: a method for quickly calibrating a video stream and projecting it onto a three-dimensional scene for virtual-real fusion, including the following steps: S1. Select a video frame and display one frame of the video as the background; S2. Mark points in the real and virtual worlds: Mark the same three points A, B, and C in the video frame and the three-dimensional virtual scene respectively; S3. Transformation and alignment of the three-dimensional model, construct a three-dimensional model transformation with equal proportion scaling, and realize alignment with the video frame through model offset calculation; S4. Input the camera FOV, and use the FOV (Field of View) parameter of the monitoring camera obtained in advance to achieve the same viewing distance as the monitoring camera in the real world.

[0017] The specific working principle and algorithm are as follows; The specific steps of S1 are as follows; S11. First, establish a spatial coordinate system and select the origin first; S12. Then define the coordinate axes (X-axis, Y-axis, Z-axis); The X-axis is usually defined as the direction horizontally to the right; The Y-axis is usually defined as the direction vertically upwards; The Z-axis is usually defined as the direction perpendicular to the XY plane, pointing "upwards"; S13. Then determine the positive directions of the coordinate axes. Each coordinate axis has a positive direction and a negative direction. In a right-handed coordinate system, if the four fingers of the right hand point from the X-axis to the Y-axis, the direction pointed by the thumb is the positive direction of the Z-axis; S14. Finally, establish the unit length of the coordinate axes. For measurement and calculation, a unit length needs to be defined for the coordinate system, such as meters, centimeters, etc.

[0018] The method for establishing the coordinate system is as follows; First, establish the coordinate system using three mutually perpendicular edges sharing a common vertex. Select a vertex as the origin, and define the three coordinate axes through the three mutually perpendicular edges passing through this vertex; Then, establish the coordinate system using the perpendicularity between a line and a plane. Select a line as the Z-axis, a plane perpendicular to the Z-axis as the XY plane, then select a point on the XY plane as the origin, and the line perpendicular to this plane as the Y-axis; Secondly, establish the coordinate system using the perpendicularity between two planes. Select two perpendicular planes, determine the Z-axis through the intersection line of these two planes, then select a point on one plane as the origin, and define the X-axis. Finally, determine the Y-axis through the right-hand rule; Finally, establish the coordinate system using the symmetry relationship in the figure. In a figure with symmetry, the coordinate axes can be defined using the axis of symmetry, and the origin is usually located at the center of symmetry; Through the above methods, establish a coordinate system suitable for a specific application and perform precise positioning and analysis in three-dimensional space.

[0019] The specific steps of S2 are as follows; S21. Mapping of three-dimensional space points to planar images. First, there are three points V1, V2, and V3 in space that form a plane. According to the three points in space, the plane equation can be determined by the point-normal form: Ax + By + Cz - (Ax0 + By0 + Cz0) = 0, where the normal vector N = (V3 - v1) × (v2 - v1) = (A, B, C), and P0(x0, y0, z0) is an arbitrary point on the plane. Please refer to Figure 3 ; Then, project the point P1(x1, y1, z1) in the current space onto the plane of the surface V1V2V3 to obtain the coordinate Pp, that is, draw a straight line from the P1 point along the normal vector N to intersect the plane, and the intersection point is the required projection point. The parametric equation of the drawn straight line is; Among them, (m, n, p) is the direction, which is the normal vector N = (A, B, C) of the projection plane to be projected. P1(x1, y1, z1) is the point to be projected, and t is the parameter of the straight-line equation. Therefore, the coordinates of the projection point Pp = (x1 + mt, y1 + nt, z1 + pt). By finding t, the projection point can be obtained. Since the projection point Pp lies on the plane V1V2V3, substitute it into the plane equation; A(x1 + mt) + B(y1 + nt) + C(z1 + pt) = Ax0 + By0 + Cz0; Expand, t(Am + Bn + Cp) = (Ax0 + By0 + Cz0) - (Ax1 + By1 + Cz1) Get, t = ((Ax0 + By0 + Cz0) - (Ax1 + By1 + Cz1)) / (Am + Bn + Cp) Since (m, n, p) = (A, B, C), so t = (N.dot(P0) - N.dot(P1)) / (N.dot(N)). Substitute t into Pp = (x1 + mt, y1 + nt, z1 + pt) = P1 + tN to obtain the projection point; S22. The calculation of the projection plane is as follows; (1) Determine the projection plane P. Please refer to Figure 3 ; Projection center: This is the point where the light rays converge, usually located at the intersection of the viewing point E and the viewing plane P.

[0020] Projection line: Starting from the viewing point E, passing through points in three-dimensional space, and finally intersecting with the viewing plane P as a ray.

[0021] Viewing plane P: This is a two-dimensional plane used to receive points projected from three-dimensional space.

[0022] (2) Calculation of the projection plane P; Three points determine a plane. By self-labeling the mirror image coordinate points A', B', C' in three-dimensional space, construct a matrix and use Gaussian elimination in linear algebra to determine a plane; Normal vector. First, calculate the vector determined by points A' and B', and then calculate the vector determined by points A' and C'. The cross product of these two vectors will give the normal vector N of the plane; Plane equation. After obtaining the normal vector N, the point-normal form equation can be used to represent the plane: N • (P - A') = 0, where P is any point on the plane and A' is a known point on the plane.

[0023] Some code examples are as follows: import numpy as np import cv2 # Create an array of 3D points that represent the vertices of an object or specific positions in space points_3d = np.array( [0, 0, 0], # Origin [1, 0, 0], # Point on the x-axis [0, 1, 0], # Point on the y-axis [0, 0, 1] # Point on the z-axis , dtype=np.float32) # Assume the camera intrinsic matrix K is known # fx and fy are the focal lengths, and cx and cy are the principal point coordinates K = np.array( [1000, 0, 320], # Assumed focal length and principal point coordinates [0, 1000, 240], [0, 0, 1] , dtype=np.float32) # Define the rotation matrix and translation vector. Here, assume no rotation and translation R = np.eye(3) # The identity matrix represents no rotation t = np.zeros((3, 1)) # The zero vector represents no translation # Project the 3D points onto the 2D plane # The cv2.projectPoints function projects 3D points onto 2D image points # dstPoints will contain the projected 2D coordinates dstPoints = cv2.projectPoints(points_3d, R, t, K)[0] # Print the projected 2D point coordinates for i, point in enumerate(dstPoints): print(f"Projected 2D Point {i}: ({point[0]}, {point[1]})")

[0024] The center point of the 2D plane is calculated as follows; Calculate the center point O of the 2D plane. Use the basic formula in computer image vision x = width / 2, y = height / 2 to obtain the center point.

[0025] The specific steps of S3 are as follows; S31, Transformation of the 3D model; Calculate the relative scaling ratio S = (A - B) / (A' - B'), and generate a flat model in three-dimensional space with length = videowidth, width = videoheight, height = 0, and the model's axis center at the upper left corner; Scale this model by 1 / S times. After creating or importing a three-dimensional model, it is necessary to adjust its size to fit the scene or meet specific visual effects. Scaling transformation can be used to quickly adjust the size of the model. Scaling is divided into uniform scaling and non-uniform scaling. Here, uniform scaling is used, and scaling can be expressed by establishing a matrix; The following is part of the code: import numpy as np # Define the vertex data of a three-dimensional model # Assume a simple cube composed of 8 vertices vertices = np.array([[ [-1, -1, -1], [1, -1, -1], [1, 1, -1], [-1, 1, -1], # Bottom [-1, -1, 1], [1, -1, 1], [1, 1, 1], [-1, 1, 1] # Top , dtype=np.float32) # Define the scaling factors scale_factors = np.array([2.0, 1.5, 1.0]) # Scale the X, Y, and Z axes in sequence # Create the scaling matrix scaling_matrix = np.diag(scale_factors) # Apply the scaling transformation scaled_vertices = vertices * scaling_matrix # Print the scaled vertex data print("Scaled Vertices:") for vertex in scaled_vertices: print(vertex)

[0026] Rotate this model. The rotation angles in the X, Y, and Z directions are obtained from the X, Y, and Z components of the differences between B' - A' and C' - A'. In the XY, YZ, and XZ planes, use the Pythagorean theorem to find the corresponding angles with arcsine or arccosine and calculate the rotation angles; In three-dimensional space, rotation can also be expressed by establishing a matrix. Once the rotation angle is available for the matrix, the corresponding rotation matrix can be constructed, and this rotation matrix can be used to transform points or vectors in a coordinate system. For example, the angle of rotation about the X-axis can be calculated using arccos and arcsin, and then a rotation matrix can be constructed to apply this rotation.

[0027] The following is part of the code: import numpy as np # Define the vertex data of a 3D model vertices = np.array([[ [1, 1, 1], # The first vertex [1, -1, 1], # The second vertex [-1, -1, 1], # The third vertex [-1, 1, 1] # The fourth vertex , dtype=np.float32) # Define the scaling matrix scale = 0.5 # Scaling factor scaling_matrix = np.diag([scale, scale, scale]) # Define the rotation matrix, taking rotation about the Z-axis as an example angle = np.pi / 4 # Rotation angle, 45 degrees rotation_matrix_z = np.array([[ [np.cos(angle), -np.sin(angle), 0], [np.sin(angle), np.cos(angle), 0], [0, 0, 1] ) # Define the translation matrix translation = np.array([2, 0, 0]) # Translate 2 units to the right translation_matrix = np.array([[ [1, 0, 0, translation[0]], [0, 1, 0, translation[1]], [0, 0, 1, translation[2]], [0, 0, 0, 1] ) # Calculate the composite transformation matrix transform_matrix = np.dot(np.dot(translation_matrix, rotation_matrix_z), scaling_matrix) # Apply the transformation to the vertex data transformed_vertices = np.dot(vertices, transform_matrix) # Print the transformed vertex data print("Transformed Vertices:") for vertex in transformed_vertices: print(vertex)。

[0028] In practical applications, a point or object needs to perform multiple rotations simultaneously. Multiple rotation matrices can be constructed and composite transformations can be achieved through matrix multiplication.

[0029] S32, model offset and alignment; Aligning the video frame with the 3D scene through model transformation (rotation, offset, translation) is based on the following logic: Solve the rotation relationship between two coordinate systems through the rotation between the matrices established by two sets of coordinates, and then solve the scale and translation amount between the two coordinate systems through the translation between the two sets of matrices. If the rotation correspondence between the two matrices is not good enough, it will cause deviations in the aligned coordinate points. If the correspondence is very good, there will be no calculation deviation in the aligned coordinate points; The scaling transformation can be represented by a 3x3 matrix. For uniform scaling, the matrix form is as follows: | Sx 0 0 | | 0 Sy 0 | | 0 0 Sz | Part of the code is as follows: import numpy as np import cvxpy as cp def align_cameras(pred_Rs, gt_Rs, pred_ts, gt_ts): """ Align the predicted matrix poses to the ground truth poses.

[0030] :param pred_Rs: ndarray, predicted rotation matrices, shape (n, 3, 3) :param gt_Rs: ndarray, The true rotation matrix, with shape nx3x3 :param pred_ts: ndarray, The predicted translation vector, with shape nx3 :param gt_ts: ndarray, The true translation vector, with shape nx3 :return: The aligned rotation matrix, translation vector, and similarity transformation matrix "" n = pred_Rs.shape[0] # Number of matrices # Calculate the optimal rotation matrix Q = np.sum(gt_Rs @ np.transpose(pred_Rs, [0, 2, 1]), axis = 0) Uq, _, Vq = np.linalg.svd(Q) sv = np.ones(3) sv[-1] = np.linalg.det(Uq @ Vq) R_opt = Uq @ np.diag(sv) @ Vq R_aligned = R_opt.reshape((1, 3, 3)) @ pred_Rs # Apply the optimal rotation # Use optimization to solve for the scale factor and translation vector constraints = [cp.norm(R_aligned - gt_Rs, 'fro') <= 1e - 6] # Assume the known rotation matrix is close enough and set a small constraint s_opt = cp.Variable() t_opt = cp.Variable((1, 3)) obj = cp.Minimize(cp.sum(cp.norm(gt_ts - (s_opt * pred_ts + t_opt), axis = 1))) prob = cp.Problem(obj, constraints) prob.solve() t_aligned = s_opt.value * pred_ts + t_opt.value # Construct the similarity transformation matrix similarity_mat = np.eye(4) similarity_mat[:3, :3] = s_opt.value * R_opt similarity_mat[:3, 3] = t_opt.value return R_aligned, t_aligned, similarity_mat

[0031] The specific steps of S4 are as follows; S41. Calculation of viewing distance; The image captured from the video stream is projected into the three-dimensional space. The video seen is still a flat picture, but it represents the physical world in reality. This application belongs to the reverse operation of the video stream captured image in the three-dimensional virtual world. Therefore, the optimal viewing distance and position are the shooting positions of the camera.

[0032] Determine the geometric center of the flat panel model: First, it is necessary to determine the geometric center point of the flat panel model (such as the screen or projection plane). This is usually the midpoint of the width and height of the model.

[0033] Calculate the normal direction: The normal direction is the vertical direction pointing outward from the geometric center point. This direction is the direction that the viewer should face in order to view the projected video head-on.

[0034] Determine the field of view (FOV): The field of view is the angle of the visual range captured by the surveillance camera lens. The parameter of the FOV is manually entered by the user for the real surveillance camera in "processing step S4".

[0035] Apply the Pythagorean theorem to calculate the viewing distance (L): This position is at a distance L along the normal direction of the center point of the rectangular flat panel model. The formula is L = (flat panel width / 2) / tan(FOV / 2) Part of the code is as follows: import math def calculate_optimal_viewing_distance(fov_degrees, screen_width_meters): """ Calculate the optimal viewing distance.

[0036] :param fov_degrees: A floating point number, the field of view angle of the camera in degrees :param screen_width_meters: A floating point number, the width of the flat panel model in meters :return: A floating point number, the optimal viewing distance in meters """ # Convert the field of view angle from degrees to radians fov_radians = math.radians(fov_degrees / 2) # Calculate the optimal viewing distance L = (screen_width_meters / 2) / math.tan(fov_radians) return L

[0037] Normally, the FOV of a camera is 90°, and the code is as follows: fov = 90 # Assume the field of view angle of the camera is 90 degrees screen_width = 2 # Assume the width of the tablet model is 2 meters viewing_distance = calculate_optimal_viewing_distance(fov, screen_width) print(f"The optimal viewing distance is: {viewing_distance} meters")

[0038] The method for quickly calibrating a video stream and projecting it onto a three-dimensional scene for virtual-real fusion proposed in this application has the following significant advantages and innovations compared with traditional technical solutions: 1. Simplify the calibration process and improve efficiency Traditional method: It relies on complex feature extraction and registration (such as SIFT / SURF algorithms), requires a large amount of computing resources, and has a high dependence on scene textures.

[0039] This solution: Through reverse calibration, only 3 corresponding points need to be manually marked in the video frame and the three-dimensional scene. Combining with the FOV parameters of the surveillance camera, the calibration can be quickly completed without repeated adjustment or high-precision feature matching.

[0040] Effect: The calibration time is shortened by more than 50%, especially suitable for large-scale surveillance camera deployment scenarios.

[0041] 2. Reduce the computational complexity Traditional method: It needs to calculate the mapping relationship between the video texture and the three-dimensional model in real time, involving high-dimensional matrix operations (such as homography transformation), and has high requirements for hardware.

[0042] This solution: Align the video frame directly through pre-transformation (scaling, rotation, translation) of the three-dimensional model. Only one initial calculation is required, and subsequent projections only require simple matrix multiplication.

[0043] Effect: The computational resource occupancy is reduced by 30%, and it can run in real time on low-computing-power devices (such as edge computing nodes).

[0044] 3. Improve the accuracy of virtual-real fusion Traditional method: Texture mapping is vulnerable to perspective distortion and stitching errors, and color differences or misalignments are likely to occur at the boundaries.

[0045] This solution: By projecting through the plane equation and aligning the normal vectors, ensure that the video stream strictly matches the geometric structure of the 3D model, avoiding perspective distortion.

[0046] Effect: The projection position error is controlled at the pixel level (<0.5%), significantly improving the spatial consistency of the monitoring screen.

[0047] 4. Support dynamic scene adaptation Traditional method: When the camera parameters or the scene change, recalibration is required, with poor flexibility.

[0048] This solution: Through parametric model transformation (such as the scaling ratio S and the rotation angle θ), the projection relationship can be dynamically adjusted to adapt to the change of the camera position or the update of the 3D model.

[0049] Effect: Support rapid iteration (for example, after the urban model is updated, the calibration data can be reused).

[0050] 5. Optimize the viewing experience Traditional method: The viewing angle is fixed, and the position of the virtual camera needs to be adjusted manually.

[0051] This solution: Based on the FOV, automatically calculate the optimal viewing distance (formula: L = (width / 2) / tan(FOV / 2)), ensuring that the user's viewing angle is consistent with the real monitoring viewing angle.

[0052] Effect: Provide a more intuitive "first-person monitoring" experience and reduce operation fatigue.

[0053] 6. Expand the application scenarios Smart city: Support the rapid access of thousands of cameras to the 3D urban model to achieve global situation awareness.

[0054] Industrial inspection: Project the device monitoring video onto the digital twin model to accurately locate the fault points.

[0055] Emergency command: Real-time fuse multiple videos into the 3D disaster scene to assist decision-making and analysis.

[0056]

[0057] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0058] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for quickly calibrating a video stream and projecting it onto a three-dimensional scene for virtual-real fusion, characterized in that, It includes the following steps: S1. Select a video frame and display one frame of the video as the background. S2. Mark points in the real and virtual worlds. Mark the same three points A, B, and C in the video frame and the 3D virtual scene respectively, and construct a projection relationship in combination with the FOV parameter of the monitoring camera. S3. Transformation and alignment of the 3D model. Calculate the relative scaling ratio S = (A - B) / (A' - B') based on the marked points, construct a 3D model with equal-proportion scaling, and align it with the video frame through offset calculation. S4. Input the camera FOV, calculate the viewing distance based on the FOV parameter and the formula L = (tablet width / 2) / tan(FOV / 2), and achieve dynamic perspective adaptation of the virtual-real fusion image.

2. A method for quickly calibrating a video stream and projecting it onto a three-dimensional scene for virtual-real fusion according to claim 1, characterized in that: The specific steps of S1 are as follows: S11. First, establish a spatial coordinate system and select the origin first. S12. Then define the coordinate axes (X-axis, Y-axis, Z-axis). The X-axis is usually defined as the direction horizontally to the right. The Y-axis is usually defined as the direction vertically upward. The Z-axis is usually defined as the direction perpendicular to the XY plane, pointing "up". S13. Then determine the positive directions of the coordinate axes. S14. Finally, establish the unit length of the coordinate axes.

3. A method for quickly calibrating a video stream and projecting it onto a three-dimensional scene for virtual-real fusion according to claim 2, characterized in that: The method for establishing the coordinate system is as follows: First, establish the coordinate system using three mutually perpendicular edges sharing a common vertex. Select a vertex as the origin, and define the three coordinate axes through the three mutually perpendicular edges passing through this vertex. Then establish the coordinate system using the perpendicularity between a line and a plane. Select a line as the Z-axis, a plane perpendicular to the Z-axis as the XY plane, then select a point on the XY plane as the origin, and the line perpendicular to this plane as the Y-axis. Secondly, establish the coordinate system using the perpendicularity between two planes. Select two perpendicular planes, determine the Z-axis through the intersection line of these two planes, then select a point on one plane as the origin, define the X-axis, and finally determine the Y-axis. Finally, establish the coordinate system using the symmetry relationship in the figure. In a figure with symmetry, the coordinate axes can be defined using the axis of symmetry, and the origin is usually located at the center of symmetry.

4. A method for quickly calibrating a video stream and projecting it onto a three-dimensional scene for virtual-real fusion according to claim 1, characterized in that: The specific steps of S2 are as follows: S21. Mapping of 3D space points to the plane image. First, three points V1, V2, and V3 in space form a plane. According to the three points in space, the plane equation can be determined by the point-normal form: Ax + By + Cz - (Ax0 + By0 + Cz0) = 0, where the normal vector N = (V3 - v1) × (v2 - v1) = (A, B, C), and P0(x0, y0, z0) is any point on the plane. Then project the point P1(x1, y1, z1) in the current space onto the plane of V1V2V3 to obtain the coordinate Pp, which is to draw a straight line from the P1 point along the normal vector N and intersect with the plane. The intersection point is the required projection point. The parametric equation of the drawn straight line is: ; where (m, n, p) is the direction, which is the normal vector N = (A, B, C) of the plane to be projected, P1(x1, y1, z1) is the point to be projected, t is the parameter of the straight-line equation. Therefore, the coordinates of the projection point Pp = (x1 + mt, y1 + nt, z1 + pt). By finding t, the projection point can be obtained. Since the projection point Pp lies on the plane of V1V2V3, substitute it into the plane equation; A(x1 + mt) + B(y1 + nt) + C(z1 + pt) = Ax0 + By0 + Cz0; Expand to get t(Am + Bn + Cp) = (Ax0 + By0 + Cz0) - (Ax1 + By1 + Cz1); Obtain t = ((Ax0 + By0 + Cz0) - (Ax1 + By1 + Cz1)) / (Am + Bn + Cp); Since (m, n, p) = (A, B, C), then t = (N.dot(P0) - N.dot(P1)) / (N.dot(N)). Substitute t into Pp = (x1 + mt, y1 + nt, z1 + pt) = P1 + tN to obtain the projection point. S22. The calculation of the projection plane is as follows: Three points determine a plane. By the mirror coordinate points A', B', and C' marked in the three-dimensional space, construct a matrix and use the Gaussian elimination method in linear algebra to determine a plane. Normal vector. First, calculate the vector determined by points A' and B', and then calculate the vector determined by points A' and C'. The cross product of these two vectors will give the normal vector N of the plane. Plane equation. Given the normal vector N, the plane can be represented by the point-normal form equation: N • (P - A') = 0, where P is any point on the plane and A' is a known point on the plane. The calculation of the center point of the two-dimensional plane is as follows: Calculate the center point O of the two-dimensional plane. Use the basic formulas x = width / 2 and y = height / 2 in computer image vision to obtain the center point.

5. A method for quickly calibrating a video stream and projecting it onto a three-dimensional scene for virtual-real fusion according to claim 1, characterized in that: The specific steps of S3 are as follows: S31. Transformation of the three-dimensional model Calculate the relative scaling ratio S = (A - B) / (A' - B'). Generate a flat plate model in the three-dimensional space with length = videowidth, width = videoheight, and height = 0, and the model axis is at the upper left corner. Scale this model by 1 / S times. After creating or importing a three-dimensional model, its size needs to be adjusted to fit the scene or meet specific visual effects. The scaling transformation can be used to quickly adjust the size of the model. Scaling can be divided into uniform scaling and non-uniform scaling. Scaling can be expressed by establishing a matrix. Rotate this model. The rotation angles in the XYZ directions are obtained from the XYZ component values of the differences between B’ - A' and C' - A'. In the XY, YZ, and XZ planes, use the Pythagorean theorem and arcsine or arccosine to find the corresponding angles and calculate the rotation angles. S32. Model offset and alignment Aligning the video frame with the three-dimensional scene through model transformation (rotation, offset, translation) is based on the following logic: Solve the rotation relationship between the two coordinate systems through the rotation between the matrices established by two sets of coordinates, and then solve the scale and translation amount between the two coordinate systems through the translation between the two sets of matrices. If the rotation correspondence between the two matrices is not good enough, it will lead to deviations in the aligned coordinate points. If the correspondence is good, there will be no calculation deviations in the aligned coordinate points. The scaling transformation can be represented by a 3x3 matrix. For uniform scaling, the matrix form is as follows: | Sx 0 0 | | 0 Sy 0 | | 0 0 Sz |.

6. A method for quickly calibrating a video stream and projecting it onto a three-dimensional scene for virtual-real fusion according to claim 1, characterized in that: The specific steps of S4 are as follows: S41. Calculation of the viewing distance Determine the geometric center of the flat plate model. First, the geometric center point of the flat plate model (such as the screen or projection plane) needs to be determined. Calculate the normal direction. The normal direction is the vertical direction pointing from the geometric center point to the outside. Determine the field of view (FOV). The FOV is the angle of the visual range captured by the surveillance camera lens. The parameter of the FOV is manually entered by the user for the actual surveillance camera in "processing step S4; Then apply the Pythagorean theorem to calculate the viewing distance (L), which is the distance extended at a distance L in the normal direction of the center point of the rectangular flat plate model. The formula is L = (plate width / 2) / tan(FOV / 2).

7. A method for quickly calibrating a video stream and projecting it onto a three-dimensional scene for virtual-real fusion according to claim 1, characterized in that: It also includes; when the position of the surveillance camera changes or the 3D model is updated, reuse the initial calibration data and dynamically update the projection relationship by adjusting the scaling ratio S and the rotation angle θ.