Multi-video and three-dimensional scene fusion method, system, equipment and medium
By calibrating camera parameters and cropping the three-dimensional model, identifying overlapping areas, determining the best viewpoint and mapping the video texture, the problem of data redundancy and inaccurate fusion in the construction and update of the three-dimensional model is solved, and efficient 3-dimensional fusion of video is achieved.
Patent Information
- Application Number
- CN202510568466.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-18
AI Technical Summary
The existing three-dimensional model construction and update methods require a large amount of point cloud data, lack of texture information, low processing efficiency, and difficult to meet the needs of intuitive, dynamic and real-time perception. At the same time, the three-dimensional model is not accurate enough when fusion with video, and the segmentation of multiple overlapping areas of video is not accurate.
By reconstructing the three-dimensional model of the target object, calibrating the camera parameters of the video camera, determining the frustum cone plane of the camera, cropping the three-dimensional model, identifying the overlapping areas between the video frame images, clustering and projecting, determining the best viewpoint, segmenting and mapping the video texture onto the three-dimensional model segmentation surface, and generating a video three-dimensional fusion model.
It improves the accuracy of fusion of three-dimensional model and the accuracy of overlapping region segmentation, meeting the needs of intuitive, dynamic and real-time perception of the real world.
Smart Images

Figure CN120339561A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital twins, and in particular, to a method, system, device, and medium for fusing multiple videos with a three-dimensional scene. Background Art
[0002] Compared with existing surveying and mapping geographic information products, three-dimensional real scenes have the advantages of being visually stereoscopic and intuitive, having accurate geometric and spatial relationships, rich texture information, and being easy for humans to understand. In the context of the extensive development of digitization and informatization in the surveying and mapping field, the construction of three-dimensional real scenes plays an important role in improving the visualization of surveying and mapping results and social and economic development, and has become the key basic data for intelligent applications such as smart cities, urban security, augmented reality, and geographic information systems (GIS).
[0003] Three-dimensional models are one of the basic and core components of three-dimensional real scenes. At present, existing methods for constructing and updating three-dimensional models require large volumes of point cloud data with high redundancy. The texture information of point clouds is relatively lacking, and the processing efficiency of point cloud software is low, making it difficult to meet the needs of intuitively, dynamically, and real-time perceiving the real world. At the same time, when fusing with a single video, the fusion of three-dimensional models is not precise enough, and the segmentation of overlapping areas of multiple videos is not precise enough. Summary of the Invention
[0004] In view of this, the present invention provides a method, system, device, and medium for fusing multiple videos with a three-dimensional scene, which solves the technical problems that existing methods for constructing and updating three-dimensional models require large volumes of point cloud data with high redundancy, the texture information of point clouds is relatively lacking, the processing efficiency of point cloud software is low, making it difficult to meet the needs of intuitively, dynamically, and real-time perceiving the real world, and at the same time, when fusing with a single video, the fusion of three-dimensional models is not precise enough, and the segmentation of overlapping areas of multiple videos is not precise enough.
[0005] The first aspect of the present invention provides a method for fusing multiple videos with a three-dimensional scene, including:
[0006] Obtaining two-dimensional image data of a target object and obtaining video frame images of the target object captured by multiple video cameras;
[0007] Reconstructing a three-dimensional model of the target object through the two-dimensional image data and calibrating the camera parameters of each of the video cameras using the three-dimensional model;
[0008] Determining multiple frustum planes of each of the video cameras according to the camera parameters and using the multiple frustum planes to crop the three-dimensional model to obtain multiple three-dimensional model surfaces located inside the frustum planes after cropping;
[0009] Determine the overlapping regions among a plurality of video frame images according to a plurality of three-dimensional model surfaces located inside the plane of the frustum;
[0010] Cluster the plurality of three-dimensional model surfaces corresponding to the overlapping regions, and project the three-dimensional model surfaces onto the frustums of the respective video cameras by using the clustering results, and compare the projected areas to determine the best viewpoints of the overlapping regions; the best viewpoints are the video cameras at the best shooting points corresponding to the overlapping regions;
[0011] Segment the overlapping regions by using the best viewpoints, and map the segmentation results back to the three-dimensional model to determine a plurality of three-dimensional model segmentation surfaces;
[0012] Map the video textures of the plurality of video frame images to the spatial positions of the respective three-dimensional model segmentation surfaces to obtain a video three-dimensional fusion model in which the video frame images are fused with the three-dimensional model.
[0013] Optionally, the reconstructing the three-dimensional model of the target object from the two-dimensional image data and calibrating the camera parameters of the respective video cameras by using the three-dimensional model includes:
[0014] Reconstruct the three-dimensional model of the target object from the two-dimensional image data, extract the first feature points of the two-dimensional image data, match the first feature points among the overlapping two-dimensional image data, and determine the sparse point cloud of the three-dimensional model according to the matched first feature points;
[0015] Extract the second feature points of the video frame images, perform feature matching between the second feature points and the first feature points, and add the second feature points to the sparse point cloud by using the feature matching results to obtain an extended sparse point cloud;
[0016] Determine the observed image coordinates of all the feature points according to the extended sparse point cloud; wherein, the feature points include the first feature points and the second feature points;
[0017] Project the three-dimensional point coordinates corresponding to the feature points onto the image plane according to the initial camera parameters of the video cameras to obtain predicted image coordinates; the initial camera parameters include initial internal parameters and initial external parameters;
[0018] Construct an objective function for bundle adjustment with the goal of minimizing the sum of the squared errors between the observed image coordinates and the predicted image coordinates;
[0019] Perform optimization and solution on the objective function to obtain the optimal solution with the minimum objective function value, and determine the camera parameters of the video cameras according to the optimal solution.
[0020] Optionally, determining multiple frustum planes of each of the video cameras according to the camera parameters, and using the multiple frustum planes to crop the three-dimensional model to obtain multiple three-dimensional model faces located inside the frustum planes after cropping, includes:
[0021] For each of the video cameras, based on the perspective projection principle, determining multiple frustum plane equations of the frustum of the video camera and the frustum planes corresponding to the frustum plane equations;
[0022] Using a three-dimensional transformation matrix to transform all vertices on the three-dimensional model into the camera coordinate system of each of the video cameras, and determining the positional relationship between each vertex of the three-dimensional model in the camera coordinate system and each of the frustum plane equations;
[0023] According to the positional relationship between each vertex and the frustum plane, removing vertices located outside all frustum planes and retaining vertices located inside at least one of the frustum planes to obtain the three-dimensional model faces cropped by the frustum planes under each of the video cameras.
[0024] Optionally, determining the overlapping regions between multiple video frame images according to the multiple three-dimensional model faces located inside the frustum planes, includes:
[0025] According to the three-dimensional model faces cropped under each of the video cameras, determining the bounding boxes of the three-dimensional model faces and determining the regional intersection of the bounding boxes respectively corresponding to each of the video cameras;
[0026] Determining the overlapping regions between the multiple video frame images according to the regional intersection.
[0027] Optionally, clustering the multiple three-dimensional model faces corresponding to the overlapping regions, and using the clustering result to project the three-dimensional model faces onto the frustums of each of the video cameras, and comparing the projected areas to determine the best viewpoints of the overlapping regions, includes:
[0028] Obtaining the multiple three-dimensional model faces within the overlapping regions;
[0029] Clustering the multiple three-dimensional model faces to obtain multiple three-dimensional model face clusters;
[0030] For each of the three-dimensional model face clusters, initializing the maximum projected area of the three-dimensional model face cluster and projecting the three-dimensional model face cluster onto the near planes of the frustums of each of the video cameras to obtain the projected faces respectively corresponding to each of the video cameras;
[0031] Calculate the current projected area of the projected surface corresponding to each of the video cameras respectively, and compare the sizes of each of the current projected areas with the maximum projected area;
[0032] When the current projected area is greater than the maximum projected area, update the current projected area to the maximum projected area, and find the maximum projected area among the current projected areas;
[0033] Use the video camera corresponding to the found maximum projected area as the best viewing point of the 3D model surface cluster, and determine the best viewing point of the overlapping area according to the best viewing points of each 3D model surface cluster.
[0034] Optionally, the step of using the best viewing point to segment the overlapping area and mapping the segmentation result back to the 3D model to determine multiple 3D model segmentation surfaces includes:
[0035] For each 3D model surface cluster, perform edge detection on the projected surface corresponding to the best viewing point;
[0036] Based on the edge detection result, cluster the projected surface according to the gray value feature and color feature of the pixel points, and segment the projected surface;
[0037] Map the segmentation result of each projected surface back to the 3D model to determine the segmentation boundary of the 3D model;
[0038] Segment the 3D model by triangulation according to the segmentation boundary to obtain multiple 3D model segmentation surfaces.
[0039] Optionally, the step of mapping the video texture of multiple video frame images to the spatial positions of each 3D model segmentation surface to obtain a video 3D fusion model in which the video frame images are fused with the 3D model includes:
[0040] According to the positions of all vertices of the 3D model segmentation surface under the best viewing point, determine the video texture coordinates corresponding to each vertex of the 3D model segmentation surface;
[0041] According to each vertex of each 3D model segmentation surface and the video texture coordinates corresponding to each vertex of the 3D model segmentation surface, determine the mapping relationship between the vertex and the video texture coordinates;
[0042] Based on the mapping relationship, map the video texture of multiple video frame images to the vertices of each 3D model segmentation surface to obtain an initial video 3D fusion model;
[0043] Perform a rendering operation on the initial video 3D fusion model to obtain a video 3D fusion model.
[0044] In a second aspect, the present invention provides a multi-video and three-dimensional scene fusion system, comprising:
[0045] A data acquisition module, configured to acquire two-dimensional image data of a target object, and acquire video frame images of the target object captured by a plurality of video cameras;
[0046] A parameter calibration module, configured to reconstruct a three-dimensional model of the target object through the two-dimensional image data, and calibrate the camera parameters of each of the video cameras by using the three-dimensional model;
[0047] A model clipping module, configured to determine a plurality of frustum planes of each of the video cameras according to the camera parameters, and clip the three-dimensional model by using the plurality of frustum planes to obtain a plurality of three-dimensional model surfaces located inside the frustum planes after clipping;
[0048] An overlap recognition module, configured to determine an overlap region between a plurality of video frame images according to the plurality of three-dimensional model surfaces located inside the frustum planes;
[0049] A viewpoint determination module, configured to cluster the plurality of three-dimensional model surfaces corresponding to the overlap region, project the three-dimensional model surfaces onto the frustums of each of the video cameras by using the clustering result, and compare the projected areas of each to determine the best viewpoint of the overlap region; the best viewpoint is the video camera at the best shooting point corresponding to the overlap region;
[0050] A region segmentation module, configured to segment the overlap region by using the best viewpoint, and map the segmentation result back to the three-dimensional model to determine a plurality of three-dimensional model segmentation surfaces;
[0051] A video fusion module, configured to map the video textures of a plurality of video frame images to the spatial positions of each of the three-dimensional model segmentation surfaces to obtain a video three-dimensional fusion model in which the video frame images are fused with the three-dimensional model.
[0052] In a third aspect, the present invention provides an electronic device, the electronic device includes a memory and a processor, and a computer program is stored in the memory. When the computer program is executed by the processor, the processor is caused to execute the steps of the multi-video and three-dimensional scene fusion method as described in the first aspect.
[0053] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the steps of the multi-video and three-dimensional scene fusion method as described in the first aspect are implemented.
[0054] As can be seen from the above technical solutions, the present invention reconstructs the three-dimensional model of the target object, uses it to calibrate the camera parameters of the video camera, determines the frustum plane of the camera, crops the three-dimensional model with these planes to obtain the three-dimensional model surfaces located within the frustum, based on these three-dimensional model surfaces, identifies the overlapping regions between video frame images, clusters the three-dimensional model surfaces in the overlapping regions, projects them onto the camera frustum, compares the projected areas to determine the best viewing point, uses the best viewing point to segment the overlapping regions, maps the results back to the three-dimensional model to determine the three-dimensional model segmentation surface, and maps the texture of the video frame image onto the three-dimensional model segmentation surface, thereby generating a video three-dimensional fusion model that meets the requirements of intuitive, dynamic, and real-time perception of the real world, and improving the accuracy of three-dimensional model fusion and the accuracy of overlapping region segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0056] Figure 1 The application environment of a multi-video and three-dimensional scene fusion method provided by an embodiment of the present invention;
[0057] Figure 2 The flowchart of a multi-video and three-dimensional scene fusion method provided by an embodiment of the present invention;
[0058] Figures 3a - 3b The schematic diagram of the three-dimensional model vertex normalization process provided by an embodiment of the present invention;
[0059] Figure 4 The schematic diagram of texture mapping provided by an embodiment of the present invention;
[0060] Figure 5 The schematic diagram of the structure of a multi-video and three-dimensional scene fusion system provided by an embodiment of the present invention;
[0061] Figure 6 The schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0063] The multi-video and three-dimensional scene fusion method provided by the embodiments of this application can be applied to, for example, Figure 1 the application environment shown in the figure. Among them, the terminal 101 communicates with the server 102 through the network. The data storage system can store the data that the server 102 needs to process. The data storage system can be integrated on the server 102, or can be placed in the cloud or on other network servers. The terminal 101 or the server 102 acquires the two-dimensional image data of the target object and acquires the video frame images of the target object captured by multiple video cameras; reconstructs the three-dimensional model of the target object through the two-dimensional image data, and calibrates the camera parameters of each video camera by using the three-dimensional model; determines multiple frustum planes of each video camera according to the camera parameters, and uses the multiple frustum planes to crop the three-dimensional model to obtain multiple three-dimensional model surfaces located inside the frustum planes after cropping; determines the overlapping areas between the multiple video frame images according to the multiple three-dimensional model surfaces located inside the frustum planes; clusters the multiple three-dimensional model surfaces corresponding to the overlapping areas, and projects the three-dimensional model surfaces to the frustums of each video camera by using the clustering result, and compares the projected areas to determine the best viewing point of the overlapping area; the best viewing point is the video camera at the best shooting point corresponding to the overlapping area; uses the best viewing point to segment the overlapping area, and maps the segmentation result back to the three-dimensional model to determine multiple three-dimensional model segmentation surfaces; maps the video textures of the multiple video frame images to the spatial positions of each three-dimensional model segmentation surface to obtain a video three-dimensional fusion model in which the video frame images are fused with the three-dimensional model.
[0064] The terminal 101 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, etc.
[0065] The server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0066] As Figure 2 shown, the embodiments of this application provide a multi-video and three-dimensional scene fusion method. Taking the method applied to the terminal 101 or the server 102 in Figure 1 as an example, the following steps S1 to S7 are included. Among them:
[0067] Step S1: Obtain the two-dimensional image data of the target object and obtain the video frame images of the target object captured by multiple video cameras.
[0068] Among them, with the continuous development of the unmanned aerial vehicle (UAV) image measurement technology, a large amount of two-dimensional image data of target objects can be easily obtained. The two-dimensional image data can be obtained by means of UAVs, etc. The video frame images are obtained by multiple video cameras taking pictures of the target object from multiple angles.
[0069] Step S2: Reconstruct the three-dimensional model of the target object through the two-dimensional image data, and use the three-dimensional model to calibrate the camera parameters of each video camera.
[0070] Among them, the three-dimensional model of the target object is quickly and automatically constructed by using the three-dimensional reconstruction technology through the two-dimensional image data.
[0071] The camera parameters of the video camera include internal parameters and external parameters. The internal parameters mainly describe the internal geometry and optical characteristics of the camera, such as focal length, optical center position, distortion coefficient, etc.; the external parameters describe the position and attitude of the camera in the world coordinate system.
[0072] Step S3: Determine the multiple frustum planes of each video camera according to the camera parameters, and use the multiple frustum planes to crop the three-dimensional model to obtain multiple three-dimensional model surfaces located inside the frustum planes after cropping.
[0073] Among them, the frustum plane is the plane determined by the viewing frustum of the video camera. The viewing frustum is a geometric body in the camera's viewing space, which determines the spatial range that the camera can see. According to the internal parameters and external parameters in the camera parameters, the viewing frustum of each video camera can be determined, and then its frustum plane can be determined. Using these frustum planes, the three-dimensional model can be cropped to retain the three-dimensional model surfaces located inside the frustum planes.
[0074] Step S4: Determine the overlapping regions among the multiple video frame images according to the multiple three-dimensional model surfaces located inside the frustum planes.
[0075] Among them, by comparing the position relationships of the three-dimensional model surfaces under different video cameras in the camera coordinate system, the overlapping parts between them are determined. These overlapping parts represent the overlapping regions among the multiple video frame images. The determination of the overlapping regions ensures that the video contents from different perspectives can be accurately stitched together to form a coherent three-dimensional scene.
[0076] Step S5: Cluster multiple three-dimensional model faces corresponding to the overlapping region, project the three-dimensional model faces onto the frustums of each video camera using the clustering result, and compare the projected areas to determine the best viewing point for the overlapping region; the best viewing point is the video camera at the best shooting point corresponding to the overlapping region.
[0077] Among them, in the embodiments of the present application, the viewing points of the video cameras corresponding to the model faces in the overlapping region are quantitatively evaluated. In order to improve the final fusion effect, a video camera with the best viewing point is selected for the model faces in the overlapping region.
[0078] In a general example, the selection of the best viewing point is based on the size of the projected area. The larger the projected area, the richer the information of the overlapping region under the viewing angle of this video camera, and it is more suitable as the central viewing angle for fusion. After determining the best viewing point, the overlapping region can be further segmented and processed based on this viewing point to ensure the accuracy and coherence of the video three-dimensional fusion model.
[0079] Step S6: Segment the overlapping region using the best viewing point and map the segmentation result back to the three-dimensional model to determine multiple three-dimensional model segmentation faces.
[0080] Among them, the three-dimensional model faces in the overlapping region are segmented through the viewing angle of the video camera corresponding to the best viewing point. During the segmentation process, the overlapping region can be accurately divided based on the projection information under the best viewing point to ensure that each segmented part corresponds to a clear and continuous three-dimensional model face. After the segmentation is completed, mapping the segmentation result back to the original three-dimensional model can clearly identify multiple three-dimensional model segmentation faces and improve the accuracy of video texture mapping.
[0081] Step S7: Map the video textures of multiple video frame images to the spatial positions of each three-dimensional model segmentation face to obtain a video three-dimensional fusion model in which the video frame images are fused with the three-dimensional model.
[0082] Among them, the pixel information in the video frame image is accurately corresponded to each vertex of the three-dimensional model segmentation face. Through texture mapping technology, the three-dimensional model segmentation face presents a visual effect consistent with the video frame image. Through the precise fusion of the video texture and the three-dimensional model, a video three-dimensional fusion model with a realistic and dynamic effect is generated.
[0083] It should be noted that in the embodiments of the present application, a three-dimensional model of the target object is reconstructed, and it is used to calibrate the camera parameters of the video camera, determine the frustum plane of the camera, crop the three-dimensional model with these planes to obtain the three-dimensional model surfaces located within the frustum, identify the overlapping regions between video frame images based on these three-dimensional model surfaces, cluster the three-dimensional model surfaces in the overlapping regions, project them onto the camera frustum, compare the projected areas to determine the best viewing point, segment the overlapping regions using the best viewing point, map the results back to the three-dimensional model to determine the three-dimensional model segmentation surface, and map the texture of the video frame images onto the three-dimensional model segmentation surface, thereby generating a video three-dimensional fusion model that meets the requirements of intuitive, dynamic, and real-time perception of the real world, and improving the accuracy of three-dimensional model fusion and the accuracy of overlapping region segmentation.
[0084] To project the pixels in a 2D video into a three-dimensional scene, it is necessary to accurately calculate the internal and external parameters of the video camera, that is, to calibrate the parameters of the video camera.
[0085] Suppose a pixel with coordinates (u i , v i ) in a two-dimensional image corresponds to a point with coordinates (X i , Y i , Z i ) in a three-dimensional scene, then the relationship between them can be expressed as Equation (1):
[0086] (1)
[0087] Among them, K is the camera internal parameter matrix, which is defined by the camera itself. Among them are the focal lengths of the camera in the x and y directions respectively, with the unit of pixel; (u0, v0) are the coordinates of the principal point of the image plane, usually located at the center of the image, s represents the scale factor, which is a non-zero constant, and it represents the proportional relationship between the distance from the three-dimensional point to the camera optical center and the distance from the pixel point on the image plane to the principal point. is the camera external parameter matrix, which is used to describe the position and orientation of the camera relative to the world coordinate system. R is the rotation matrix, which is a 3 matrix and is used to describe the rotation angle of the camera; T is the translation vector, which is a 3 matrix and is used to describe the position of the camera origin in the world coordinate system.
[0088] In some embodiments, a three-dimensional model of the target object is reconstructed through two-dimensional image data, and the camera parameters of each video camera are calibrated using the three-dimensional model, including:
[0089] Step S201: Reconstruct the three-dimensional model of the target object from the two-dimensional image data, extract the first feature points of the two-dimensional image data, match the first feature points between the overlapping two-dimensional image data, and determine the sparse point cloud of the three-dimensional model based on the matched first feature points.
[0090] Among them, feature extraction can adopt feature extraction algorithms such as Scale-Invariant Feature Transform (SIFT) and Speeded-Up Robust Features (SURF), and feature point matching can adopt matching methods such as brute-force matching and Approximate Nearest Neighbor Search (FLANN) matching.
[0091] At the same time, a sparse point cloud is constructed to restore the position and attitude of the UAV camera.
[0092] Step S202: Extract the second feature points of the video frame image, perform feature matching between the second feature points and the first feature points, and add the second feature points to the sparse point cloud using the feature matching results to obtain an extended sparse point cloud.
[0093] In specific implementation, first, unify the coordinate system: Before adding the video frame image feature points to the sparse point cloud of the UAV image, it is first necessary to unify the coordinate systems of the two. By establishing an appropriate coordinate transformation relationship, the video frame image and the UAV image are in the same coordinate system.
[0094] Secondly, perform feature matching point association. Using the completed image feature matching results, associate the feature points in the video image with the points in the sparse point cloud of the UAV image. Specifically, by calculating the distance between the matching descriptors, find the point in the sparse point cloud that is closest to the second feature point, thereby establishing a corresponding relationship between the two.
[0095] Finally, add the second feature points to the sparse point cloud, update the sparse point cloud data, and update the sparse point cloud data. In this way, gradually integrate the feature points in the new image into the existing point cloud, making it contain more information from the video frame image and gradually densifying the sparse point cloud.
[0096] Step S203: Determine the observed image coordinates of all feature points according to the extended sparse point cloud; where the feature points include the first feature points and the second feature points.
[0097] Among them, the observed image coordinates are the coordinates of the feature points extracted from the actual image. By traversing each feature point in the extended sparse point cloud, all the video frame images in which each feature point is observed are found, and the coordinates of the feature points in these images are recorded. These observed image coordinates will be used in the subsequent 3D reconstruction and camera parameter calibration processes.
[0098] Step S204: Project the 3D point coordinates corresponding to the feature points onto the image plane according to the initial camera parameters of the video camera; the initial camera parameters include the initial internal parameters and the initial external parameters.
[0099] Among them, based on the initial camera parameters, through the projection transformation formula, the coordinates of the feature points in the 3D space are transformed from the world coordinate system to the camera coordinate system through rotation and translation transformations. Then, the points in the camera coordinate system are projected onto the image plane through the internal parameters of the camera to obtain the predicted image coordinates.
[0100] The predicted image coordinates are the theoretical image coordinates calculated based on the camera parameters and the 3D point coordinates, which reflect the position of the feature points on the image plane under ideal conditions without noise and errors.
[0101] Among them, the projection transformation formula takes into account the internal parameters of the camera (such as focal length, principal point coordinates, etc.) and external parameters (such as rotation matrix, translation vector, etc.), ensuring the accurate mapping of the 3D points to the image plane.
[0102] Among them, the initial internal parameters of the video camera can be initially determined by using the Zhang Zhengyou calibration method, and the external parameters can be initially designed according to the internal parameters. Or, through a drone equipped with sensors such as real-time kinematics and inertial measurement units, the actual position and attitude information of the camera are recorded during shooting, so as to determine the initial external parameters of the video camera.
[0103] Step S205: Construct an objective function for bundle adjustment with the goal of minimizing the sum of the squared errors between the observed image coordinates and the predicted image coordinates.
[0104] Among them, the objective function of bundle adjustment is a function that measures the difference between the observed image coordinates and the predicted image coordinates. By minimizing this function, the camera parameters and the 3D point coordinates can be optimized to more accurately reflect the actual scene. During the bundle adjustment process, the observed image coordinates and the predicted image coordinates of all feature points are considered, and through an iterative optimization algorithm, the camera parameters and the 3D point coordinates are continuously adjusted until the objective function reaches the minimum value or converges below a certain threshold.
[0105] Through bundle adjustment, the accuracy of 3D reconstruction and the precision of camera parameter calibration can be further improved. The optimized camera parameters and 3D point coordinates will be used in the subsequent process of generating the video 3D fusion model to ensure the accuracy and coherence of the fusion effect.
[0106] Step S206: Optimize and solve the objective function to obtain the optimal solution with the minimum objective function value, and determine the camera parameters of the video camera according to the optimal solution.
[0107] Among them, optimization algorithms such as the Newton method and quasi-Newton method can be used to solve the objective function, and at the same time optimize the internal and external parameters of the camera and the 3D point coordinates, so as to obtain more accurate internal and external parameters.
[0108] Among them, the Newton method is an iterative optimization algorithm. Its core idea is to construct a quadratic approximation model of the objective function at the current point, and then solve the minimum value of this quadratic model to obtain the next iteration point. The quasi-Newton method is an improvement of the Newton method because in practical applications, the computational cost of calculating the Hessian matrix and its inverse matrix is very large. The quasi-Newton method reduces the computational cost by approximately calculating the Hessian matrix or its inverse matrix.
[0109] The key to the fusion of video and 3D scene lies in accurately calculating the shooting area of the video camera so as to reproduce the visual effect of the video in the 3D scene, which is similar to the inverse process of shooting a 3D scene. Figures 3a - 3b Shows the normalization process of the vertices of the 3D model in the rendering pipeline and the conversion from the camera coordinate system to the NDC coordinate system. According to the concept of the perspective camera, the 3D model faces within the camera's visible space are always surrounded by the camera's frustum.
[0110] Among them, the NDC (Normalized Device Coordinates) coordinate system is a concept in computer graphics used to map points in a 3D scene to a 2D screen.
[0111] First, use the model view transformation to convert the vertices of the 3D model to the camera coordinate system. According to the perspective projection principle, assume that the coordinates of a point P in the camera frustum are a (P x ,P y ,P z ) in the camera coordinate system, and the calculation formula for the corresponding coordinate value P' in the NDC coordinate system is shown in Equation (2).
[0112] (2)
[0113] Wherein, l, r, b, and t are the coordinate values of the left plane, right plane, lower plane, and upper plane of the frustum in the camera coordinate system, respectively, which are calculated from the internal parameters of the camera; n and f are the coordinate values of the near plane and the far plane, which are generally manually specified according to needs.
[0114] Then, it is converted to normalized device coordinates through perspective or orthographic projection. Its shape is a cube, and the range is [-1, 1]. The surface of the 3D model is usually uneven, and some faces inside the frustum, such as the back face or the occluded face, may not be visible from the camera position, which can lead to incorrect video mapping and affect the rendering efficiency. Therefore, in some embodiments, back-face detection and z-buffer algorithms can be adopted to remove the invisible faces and determine the visible model faces that match the actual shooting range of the camera. According to the z-buffer algorithm, in the NDC coordinate system, z ′ is inversely proportional to the coordinate value of the vertex Pz, as shown in Equation (3):
[0115] (3)
[0116] Wherein, z ′ is called the pseudo-depth, which represents the sequence of the model vertices. The value range is , and Pz corresponds to .
[0117] In the embodiment of the present application, the depth buffer algorithm is implemented in the NDC coordinate system, and the pseudo-depth value is set to the maximum depth of 1.0. Then, the depth of each vertex is calculated one by one and compared with the corresponding value in the depth buffer. If the face depth value is less than the buffer value. Finally, the vertex corresponding to the depth buffer in the NDC coordinate system is restored to the world coordinate system by using the external parameter projection matrix of the camera, and the face corresponding to the visible space of the camera is obtained.
[0118] Finally, to limit the computational amount of multiple videos, it is also necessary to determine the overlapping area within the visible range of the camera. For this purpose, the following introduces the process of determining the multiple frustum planes of each video camera according to the camera parameters and using the multiple frustum planes to clip the 3D model to obtain multiple 3D model faces located inside the frustum plane after clipping, including:
[0119] Step S301: For each video camera, based on the principle of perspective projection, determine the multiple frustum plane equations of the frustum of the video camera and the frustum planes corresponding to the frustum plane equations.
[0120] Wherein, according to the conversion relationship between the camera coordinate system and the world coordinate system, the equations of the frustum planes of each video camera in the world coordinate system are derived. These equations describe the visible range of the camera, that is, which 3D model faces will be captured by the camera.
[0121] For each camera, the equations of the six planes of the frustum are calculated according to its position, orientation and projection parameters. Generally speaking, the six planes of the frustum are the near clipping plane, the far clipping plane, the left clipping plane, the right clipping plane, the upper clipping plane and the lower clipping plane.
[0122] Step S302: using a three-dimensional transformation matrix, convert all vertices on the three-dimensional model into the camera coordinate system of each video camera, and determine the positional relationship between each vertex of the three-dimensional model and each frustum plane in the camera coordinate system according to the equations of each vertex and each frustum plane of the three-dimensional model in the camera coordinate system.
[0123] Among them, all vertices of the 3D model are converted from the world coordinate system to the camera coordinate system of each video camera using a 3D transformation matrix, so as to facilitate subsequent comparison with the video camera frustum plane. For each video camera, the positional relationship between the model vertex and the frustum plane is determined.
[0124] For each camera frustum's plane equation, substitute the vertices of the 3D model in the camera coordinate system into the plane equation. Based on the substitution result, determine which side of the plane the vertex is on. If the value of the vertex substituted into the plane equation is less than 0, the vertex is on the inside of the plane (towards the camera), otherwise it is on the outside.
[0125] Step S303: according to the positional relationship between each vertex and the frustum plane, remove the vertices located outside all frustum planes, retain the vertices located inside at least one frustum plane, and obtain the three-dimensional model surface clipped by the frustum plane under each video camera.
[0126] Among them, according to the positional relationship between each vertex and the frustum plane, the vertices located outside all the frustum planes are eliminated, and the vertices located inside at least one plane are retained, thereby obtaining the three-dimensional model part clipped by the frustum plane under each camera as the three-dimensional model surface clipped by the frustum plane under each video camera.
[0127] The following describes a process of determining an overlapping area between multiple video frame images based on multiple three-dimensional model surfaces located inside the frustum plane, including:
[0128] Step S401: Determine the bounding box of the 3D model surface according to the cropped 3D model surface under each video camera, and determine the area intersection of the bounding boxes corresponding to each video camera.
[0129] Step S402: determining the overlapping area between multiple video frame images according to the area intersection.
[0130] In some embodiments, based on the bounding box method, the overlapping regions of the cropped model parts under multiple video cameras are calculated. That is, the bounding boxes of the cropped models under each video camera are calculated (such as the AABB (Axis-Aligned Bounding Box) bounding box or the Oriented Bounding Box (OBB)), and then the range of the overlapping region is preliminarily determined by calculating the intersection of these bounding boxes. Then, the model vertices located within the intersection bounding box are further judged to determine whether they are truly within all the cropped model parts. Among them, the AABB bounding box is a rectangle (in 2D) or a cuboid (in 3D) aligned with the coordinate axes, and its boundaries are parallel to the coordinate axes.
[0131] When processing video overlapping regions, 3D model faces are often related to multiple video textures. Current techniques usually handle seams by evenly dividing the overlapping region, but this may cause obvious stretching of the edges; while the manual masking method can prevent texture repetition, it requires manual quality evaluation and is difficult to adapt to a large number of video scenarios. To solve these problems, an overlapping region segmentation method based on optimal viewpoint selection is proposed. This method first evaluates the viewpoint angles of multiple video textures on the same model face through a simulation test, then selects the camera parameters of the best view, and finally completes the effective segmentation of the overlapping region.
[0132] For this reason, the following introduces the process of clustering multiple 3D model faces corresponding to the overlapping region, projecting the 3D model faces onto the frustums of each video camera using the clustering results, and comparing the projected areas to determine the best viewpoints of the overlapping region, including:
[0133] Step S501: Obtain multiple 3D model faces within the overlapping region.
[0134] Among them, according to the range of the overlapping region, all 3D model faces located within this overlapping region are extracted from the 3D model.
[0135] Step S502: Cluster the multiple 3D model faces to obtain multiple clusters of 3D model faces.
[0136] Among them, the clustering process is a process of classifying similar 3D model faces into one category. In this embodiment, the K-means clustering or the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm can be used to cluster the 3D model faces within the overlapping region. Through the clustering process, model faces with similar features or attributes can be classified into one category, which is convenient for subsequent analysis and processing. The result of the clustering process is to divide the 3D model faces within the overlapping region into at least one clustering cluster, and each clustering cluster contains a group of similar 3D model faces.
[0137] Step S503: For each 3D model face cluster, initialize the maximum projected area of the 3D model face cluster, and project the 3D model face cluster onto the near plane of the frustum of each video camera to obtain the projected faces corresponding to each video camera respectively.
[0138] Among them, for each 3D model face cluster, first calculate its projected area on the near plane of each video camera frustum. The calculation of the projected area can be achieved by projecting each face in the 3D model face cluster onto the near plane and calculating the area of the projected face. To obtain an accurate projected area, it is necessary to consider the internal parameters of the camera, such as focal length, principal point coordinates, etc., to ensure the accuracy of the projection.
[0139] During the projection process, the projection matrix of the camera can be used to project each vertex in the 3D model face cluster from 3D space onto the 2D plane. The projection matrix contains the internal parameters and external parameters of the camera, and can accurately map 3D points to the image plane. By calculating the coordinates of the projected vertices, the shape and size of the projected face can be obtained, and then the projected area can be calculated.
[0140] Step S504: Calculate the current projected area of the projected faces corresponding to each video camera respectively, and compare the sizes of each current projected area with the maximum projected area.
[0141] Step S505: When the current projected area is greater than the maximum projected area, update the current projected area to the maximum projected area, and find the maximum projected area among each current projected area.
[0142] Among them, initialize the maximum projected area of each 3D model face cluster to 0, then traverse all video cameras, project the 3D model face cluster onto the near plane of each camera frustum, and record the projected area. During the traversal process, update the maximum projected area of each 3D model face cluster, and retain the camera with the largest projected area as the best viewing point for the 3D model face cluster.
[0143] By comparing the projected areas of different 3D model face clusters on each video camera, the best viewing point for each face cluster can be determined.
[0144] The best viewing point is the camera with the largest projected area because it can provide a more complete and clearer view. After determining the best viewing point, the camera parameters of this viewing point can be used to segment the overlapping area to ensure that the segmented area is more visually coherent and consistent.
[0145] Step S506: Use the video camera corresponding to the found maximum projected area as the best viewing point for the 3D model face cluster, and determine the best viewing point of the overlapping area according to the best viewing points of each 3D model face cluster.
[0146] In the embodiments of the present application, the perspective angle of the camera is evaluated by the projected area method. It is not only simple to calculate, but also represents both the video texture angle and the model vertex distance at the same time. In the camera coordinate system, first, project the three-dimensional model surface vertex P(P x ,P y ,P z ) onto the near plane of the frustum, that is, calculate the intersection point of the near plane of the frustum and the line connecting the camera position and the model surface vertex, denoted as point P n ′(x, y, -n). According to the principle of similar triangles, there is the following conversion relationship between P n ′ and P as shown in Formula 4.
[0147] (4)
[0148] Repeat the above steps, project all the vertices of the three-dimensional model surface onto the near plane of the camera frustum in clockwise order, and connect the projection points of adjacent vertices in order of the vertices to obtain a two-dimensional polygon composed of the projection surface. Then calculate the area of the polygon as a quantitative evaluation index of the camera perspective.
[0149] It can be understood that the perspective of the video camera corresponding to the model surface in the overlapping area is quantitatively evaluated. In order to improve the final fusion effect, a video camera with the best perspective is selected for the model surface in the overlapping area. First, traverse the model surfaces in the overlapping area, calculate the perspective evaluation index values of multiple cameras corresponding to a single surface, and then select and record the video camera with the best perspective (that is, the largest projected area, indicating the closest distance).
[0150] The following introduces the process of using the best perspective to segment the overlapping area and mapping the segmentation result back to the three-dimensional model to determine multiple three-dimensional model segmentation surfaces, including:
[0151] Step S601: For each three-dimensional model surface cluster, perform edge detection on the projected surface corresponding to the best perspective.
[0152] Among them, edge detection algorithms such as the Canny operator and the Sobel operator can be applied to the two-dimensional projection map to extract the edge information of the model projection, and these edges are used as the preliminary boundaries for segmentation.
[0153] The Canny operator is a multi-stage edge detection algorithm, which mainly includes four steps: Gaussian smoothing, gradient calculation, non-maximum suppression and double threshold processing. First, the image is smoothed using a Gaussian filter to reduce the impact of noise; then the gradient amplitude and direction of the image are calculated; then, non-maximum suppression is used to retain only the points with the largest local gradient and refine the edges; finally, double threshold processing is used to treat points with gradient amplitudes greater than the high threshold as strong edges, and points between the high threshold and the low threshold and connected to the strong edge as weak edges, thus obtaining the final edge image.
[0154] The Sobel operator is an edge detection operator based on first-order derivatives. It calculates the gradient of the image in the horizontal and vertical directions, and then combines the gradient amplitudes in the two directions to obtain the final edge image. The Sobel operator uses two 3x3 convolution kernels to convolve the image separately, one for detecting horizontal edges and the other for detecting vertical edges.
[0155] Step S602: Based on the edge detection result, the projected surface is clustered according to the gray value characteristics and color characteristics of the pixel points, and the projected surface is segmented.
[0156] It is understandable that edge detection only obtains the rough outline of the model projection. In order to segment the projection image more accurately, it is necessary to further merge adjacent pixels into different areas based on their feature similarities.
[0157] The embodiment of the present application uses K-means clustering to cluster the projected surface, thereby segmenting the projected surface.
[0158] Among them, K-means clustering is an unsupervised clustering algorithm. Its basic idea is to select K initial seed points, assign each pixel to the cluster where the nearest seed point is located, and then update the center point of each cluster. Repeat this process until the stop condition is met (such as the center point no longer changes or the maximum number of iterations is reached). In projection image segmentation, clustering can be performed based on pixel grayscale value, color and other features.
[0159] Step S603: Map the segmentation results of each projected surface back to the three-dimensional model to determine the segmentation boundary of the three-dimensional model.
[0160] Step S604: segment the three-dimensional model according to the segmentation boundaries by using a triangulation method to obtain a plurality of three-dimensional model segmentation surfaces.
[0161] Among them, according to the segmentation results on the 2D projection map, the pixel points corresponding to each segmentation region are found. For each pixel point, through the inverse transformation of the projection matrix, the corresponding vertex in the 3D model is found. The segmentation boundary of the 3D model is determined based on these vertices. Methods such as triangulation can be used to connect these vertices into faces, thus completing the segmentation of the 3D model.
[0162] During the video space restoration process, the cameras corresponding to the faces of the 3D model can be found. For multiple video overlapping regions, by using the optimal viewpoint segmentation method proposed in this paper, the faces of the 3D model can be classified according to different cameras. Therefore, each face of the 3D model within the camera vision space has its own video texture. On the other hand, the fusion of the video and the 3D model requires dynamically attaching the 2D video texture to the surface of the 3D model. This study uses the texture mapping method to achieve a one-to-one association between video pixels and 3D model vertices, and dynamically updates the texture of the 3D model vertices as the video plays. Finally, the rendering of the entire scene is completed based on the rendering pipeline.
[0163] The following introduces the process of mapping the video texture of multiple video frame images to the spatial positions of the segmented faces of each 3D model to obtain a video 3D fusion model in which the video frame images are fused with the 3D model, including:
[0164] Step S701: Determine the video texture coordinates corresponding to each vertex of the 3D model segmentation face according to the positions of all vertices of the 3D model segmentation face under the optimal viewpoint.
[0165] Among them, according to the concept of texture mapping, on the basis of calculating the texture coordinates of the model vertices corresponding to the video, the video texture is dynamically mapped to the surface of the 3D model. After segmenting the 3D model face, according to the optimal viewpoint, the vertex P (P x , P y , P z ) on the model face corresponds to the video texture coordinates P t ′(u, v), as shown in Figure 4 . In the formula, α is the angle between the vertical plane and the perpendicular line from the vertex P to the principal optical axis, and β is the corresponding angle in the horizontal direction. Using the mapping relationship between the texture coordinates and the face vertices, the texture coordinates u and v corresponding to the vertex P can be calculated, as shown in Equation (5):
[0166] (5)
[0167] Among them, fov is the vertical field of view angle of the camera, and aspect is the ratio of the horizontal and vertical field of view angles of the camera. In the camera coordinate system, traverse the 3D model vertices in the visible space of the camera, and derive the texture coordinates according to the model vertex positions.
[0168] In the actual calculation process, this step can be processed during the segmentation stage of the overlapping area to reduce the overall calculation amount.
[0169] Step S702: Determine the mapping relationship between the vertices of each three-dimensional model segmentation surface and the video texture coordinates corresponding to the vertices of each three-dimensional model segmentation surface.
[0170] Step S703: Based on the mapping relationship, map the video textures of multiple video frame images to the vertices of each three-dimensional model segmentation surface to obtain an initial video three-dimensional fusion model.
[0171] Among them, each pixel in the video frame image is corresponded to the vertices on the three-dimensional model segmentation surface one by one to achieve the precise fitting of the video texture and the three-dimensional model. In this process, the texture mapping technology is used to ensure that the video texture can be dynamically updated to the surface of the three-dimensional model as the video plays, so that the fused model not only retains the three-dimensional sense of space but also can present the dynamic effect of the video.
[0172] Step S704: Perform a rendering operation on the initial video three-dimensional fusion model to obtain a video three-dimensional fusion model.
[0173] Among them, according to the three-dimensional model rendering principle, first, the geometric information of the three-dimensional model vertices is processed. This includes steps such as projection transformation, primitive assembly, view frustum culling, and vertex shading. Then, in the rasterization stage, the vertex values in the normalized device coordinates are interpolated and shaded to generate raster pixels or fragments. In the fragment processing stage, texture mapping and hidden surface elimination technologies are applied to enhance visual details and improve rendering efficiency. Finally, the processed fragments are saved in the frame buffer and prepared to be output to the display device, where the texture mapping converges with the fragment processing and geometric drawing pipelines through an independent texture pixel pipeline, thus finally presenting a high-quality 3D image.
[0174] Among them, the projection transformation is the process of converting the three-dimensional model from the model coordinate system to the camera coordinate system. This step is to project the three-dimensional model onto the camera's view plane for subsequent rendering and texture mapping. The projection transformation usually includes two methods: perspective projection and parallel projection. In the embodiment of the present application, perspective projection is adopted because it can simulate the visual effect of the human eye, making distant objects look smaller and near objects look larger, enhancing the realism of the scene.
[0175] Primitive assembly is to assemble the vertices after projection transformation into basic graphic elements, such as points, lines, triangles, etc. These basic graphic elements are the basis for constructing the three-dimensional model.
[0176] View frustum culling is to remove the parts outside the camera frustum to reduce unnecessary computational workload. The frustum is the spatial range that the camera can see and is determined by the position and viewing angle of the camera. Only the parts of the 3D model located within the frustum will be rendered.
[0177] Vertex shading is the process of processing information such as the color and lighting of the vertices of a 3D model. Through vertex shading, lighting effects, shadows, etc. can be added to the 3D model to enhance the realism and three-dimensionality of the scene.
[0178] Rasterization is the process of converting vertex information into pixel information. In the rasterization stage, the color value of each pixel is calculated based on the geometric information and color information of the vertices. This step is the most time-consuming step in 3D rendering because it needs to process each pixel.
[0179] The fragment processing stage is carried out after rasterization and mainly further processes the pixels, such as texture mapping, hidden surface elimination, etc. Texture mapping is the process of mapping a 2D texture image onto the surface of a 3D model, which can increase the details and realism of the model. Hidden surface elimination is to remove the occluded parts and only retain the visible parts to improve the rendering efficiency.
[0180] The processed fragments are finally saved in the frame buffer and wait to be output to the display device. The frame buffer is a buffer for storing image data and it saves the rendering result of the current scene. When all the rendering operations are completed, the image data in the frame buffer will be output to the display device to present the final 3D image.
[0181] It should be noted that the embodiments of the present application use a 3D scene rendering pipeline to fuse a video and a 3D model into a visualization grid. The video space is restored through camera parameters, and a geometric model is constructed by combining vertex positions, normals, and texture information. The video is converted into a dynamic texture, and the face components in the model are interpolated and colored to enhance the realism. Subsequently, the geometry is combined with the video material to form a unified grid structure and is reasonably segmented according to the best viewing point. Finally, through steps such as window transformation and hidden surface elimination, the superimposed fusion rendering of multiple videos and 3D models can be successfully achieved, creating a rich visual experience.
[0182] Based on the same inventive concept, the embodiments of the present application also provide a multi-video and 3D scene fusion system for implementing the multi-video and 3D scene fusion method involved above.
[0183] The implementation solution provided by the system for solving problems is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the multi-video and three-dimensional scene fusion system provided below can refer to the limitations on the multi-video and three-dimensional scene fusion method in the above text, and will not be elaborated here.
[0184] As Figure 5 shown, an embodiment of the present application provides a multi-video and three-dimensional scene fusion system, which includes:
[0185] A data acquisition module 100, configured to acquire two-dimensional image data of a target object, and acquire video frame images of the target object captured by a plurality of video cameras;
[0186] A parameter calibration module 200, configured to reconstruct a three-dimensional model of the target object through the two-dimensional image data, and calibrate the camera parameters of each video camera by using the three-dimensional model;
[0187] A model clipping module 300, configured to determine a plurality of frustum planes of each video camera according to the camera parameters, and clip the three-dimensional model by using the plurality of frustum planes to obtain a plurality of three-dimensional model surfaces located inside the frustum planes after clipping;
[0188] An overlap recognition module 400, configured to determine an overlap region between a plurality of video frame images according to the plurality of three-dimensional model surfaces located inside the frustum planes;
[0189] A viewpoint determination module 500, configured to cluster the plurality of three-dimensional model surfaces corresponding to the overlap region, project the three-dimensional model surfaces onto the frustums of each video camera by using the clustering result, and compare the projected areas to determine the best viewpoint of the overlap region; The best viewpoint is the video camera at the best shooting point corresponding to the overlap region;
[0190] A region segmentation module 600, configured to segment the overlap region by using the best viewpoint, and map the segmentation result back to the three-dimensional model to determine a plurality of three-dimensional model segmentation surfaces;
[0191] A video fusion module 700, configured to map the video textures of a plurality of video frame images to the spatial positions of each three-dimensional model segmentation surface to obtain a video three-dimensional fusion model in which the video frame images are fused with the three-dimensional model.
[0192] In some embodiments, the parameter calibration module 200 is configured to:
[0193] Reconstruct a three-dimensional model of the target object through the two-dimensional image data, extract first feature points of the two-dimensional image data, match the first feature points between the overlapping two-dimensional image data, and determine a sparse point cloud of the three-dimensional model according to the matched first feature points;
[0194] Extracting a second feature point of the video frame image, performing feature matching on the second feature point and the first feature point, and adding the second feature point to the sparse point cloud using the feature matching result to obtain an extended sparse point cloud;
[0195] Determine the observed image coordinates of all feature points according to the extended sparse point cloud; wherein the feature points include a first feature point and a second feature point;
[0196] The three-dimensional point coordinates corresponding to the feature points are projected onto the image plane according to the initial camera parameters of the video camera to obtain the predicted image coordinates; the initial camera parameters include initial internal parameters and initial external parameters;
[0197] The objective function of bundle adjustment is constructed with the goal of minimizing the sum of square errors between the observed image coordinates and the predicted image coordinates.
[0198] The objective function is optimized and solved to obtain the optimal solution with the minimum objective function value, and the camera parameters of the video camera are determined according to the optimal solution.
[0199] In some embodiments, the model tailoring module 300 is used to:
[0200] For each video camera, based on the perspective projection principle, multiple frustum plane equations of the frustum of the video camera and frustum planes corresponding to the frustum plane equations are determined;
[0201] Using the three-dimensional transformation matrix, all vertices on the three-dimensional model are converted into the camera coordinate system of each video camera, and according to the equations of each vertex and each frustum plane of the three-dimensional model in the camera coordinate system, the positional relationship between each vertex and each frustum plane of the three-dimensional model in the camera coordinate system is determined;
[0202] According to the positional relationship between each vertex and the frustum plane, the vertices located outside all frustum planes are removed, and the vertices located inside at least one frustum plane are retained to obtain the three-dimensional model surface clipped by the frustum plane under each video camera.
[0203] In some embodiments, the overlap identification module 400 is used to:
[0204] Determine the bounding box of the 3D model surface according to the cropped 3D model surface under each video camera, and determine the area intersection of the bounding boxes corresponding to each video camera;
[0205] The overlapping area between multiple video frame images is determined according to the area intersection.
[0206] In some embodiments, the viewpoint determination module 500 is configured to:
[0207] Acquire multiple three-dimensional model surfaces in the overlapping area;
[0208] Cluster multiple 3D model faces to obtain multiple clusters of 3D model faces;
[0209] For each cluster of 3D model faces, initialize the maximum projected area of the cluster of 3D model faces, and project the cluster of 3D model faces onto the near plane of the frustum of each video camera to obtain the projected faces corresponding to each video camera respectively;
[0210] Calculate the current projected area of the projected faces corresponding to each video camera respectively, and compare the sizes of the current projected areas with the maximum projected area;
[0211] When the current projected area is greater than the maximum projected area, update the current projected area to the maximum projected area, and find the maximum projected area among the current projected areas;
[0212] Use the video camera corresponding to the found maximum projected area as the best viewing point of the cluster of 3D model faces, and determine the best viewing point of the overlapping area according to the best viewing points of each cluster of 3D model faces.
[0213] In some embodiments, the region segmentation module 600 is used for:
[0214] For each cluster of 3D model faces, perform edge detection on the projected face corresponding to the best viewing point;
[0215] Based on the edge detection results, cluster and segment the projected face according to the gray value features and color features of the pixel points;
[0216] Map the segmentation results of each projected face back to the 3D model to determine the segmentation boundary of the 3D model;
[0217] Segment the 3D model by triangulation according to the segmentation boundary to obtain multiple 3D model segmentation faces.
[0218] In some embodiments, the video fusion module 700 is used for:
[0219] Determine the video texture coordinates corresponding to each vertex of the 3D model segmentation face according to the positions of all vertices of the 3D model segmentation face under the best viewing point;
[0220] Determine the mapping relationship between the vertex and the video texture coordinate according to each vertex of each 3D model segmentation face and the video texture coordinates corresponding to each vertex of the 3D model segmentation face;
[0221] Based on the mapping relationship, map the video texture of multiple video frame images onto the vertices of each 3D model segmentation face to obtain an initial video 3D fusion model;
[0222] Perform a rendering operation on the initial video three-dimensional fusion model to obtain the video three-dimensional fusion model.
[0223] As Figure 6 shown, an embodiment of the present application provides an electronic device. The electronic device 10 includes a memory 20 and a processor 30. A computer program is stored in the memory 20. When the computer program is executed by the processor 30, the processor 30 is caused to execute the steps of the multi-video and three-dimensional scene fusion method in the above-mentioned embodiment.
[0224] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the steps of the multi-video and three-dimensional scene fusion method in the above-mentioned embodiment are implemented.
[0225] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, electronic devices, and computer storage media described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.
[0226] It should be noted that the terms "first", "second", etc. in the description, claims, and drawings of the present invention are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0227] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.
[0228] In several embodiments provided by the present invention, it should be understood that the disclosed systems, electronic devices, computer storage media, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0229] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0230] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0231] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (English full name: Read-Only Memory, English abbreviation: ROM), random access memories (English full name: Random Access Memory, English abbreviation: RAM), magnetic disks, or optical discs that can store program codes.
[0232] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for fusing multiple videos with a three-dimensional scene, characterized in that, Including: Obtaining two-dimensional image data of a target object, and obtaining video frame images of the target object captured by a plurality of video cameras; Reconstructing a three-dimensional model of the target object from the two-dimensional image data, and calibrating the camera parameters of each of the video cameras using the three-dimensional model; Determining a plurality of frustum planes of each of the video cameras according to the camera parameters, and using the plurality of frustum planes to crop the three-dimensional model, so as to obtain a plurality of three-dimensional model surfaces located inside the frustum planes after cropping; Determining an overlapping region between a plurality of video frame images according to the plurality of three-dimensional model surfaces located inside the frustum planes; Clustering the plurality of three-dimensional model surfaces corresponding to the overlapping region, and projecting the three-dimensional model surfaces using the clustering result onto the frustums of each of the video cameras, and comparing the projected areas to determine an optimal viewpoint of the overlapping region; The optimal viewpoint is the video camera at the optimal shooting point corresponding to the overlapping region; Segmenting the overlapping region using the optimal viewpoint, and mapping the segmentation result back to the three-dimensional model to determine a plurality of three-dimensional model segmentation surfaces; Mapping the video textures of the plurality of video frame images to the spatial positions of each of the three-dimensional model segmentation surfaces to obtain a video three-dimensional fusion model in which the video frame images are fused with the three-dimensional model.
2. The multi-video and three-dimensional scene fusion method according to claim 1, characterized in that The reconstructing a three-dimensional model of the target object from the two-dimensional image data, and calibrating the camera parameters of each of the video cameras using the three-dimensional model includes: Reconstructing a three-dimensional model of the target object from the two-dimensional image data, extracting first feature points of the two-dimensional image data, matching the first feature points between the overlapping two-dimensional image data, and determining a sparse point cloud of the three-dimensional model according to the matched first feature points; Extracting second feature points of the video frame images, performing feature matching between the second feature points and the first feature points, and adding the second feature points to the sparse point cloud using the feature matching result to obtain an extended sparse point cloud; Determining the observed image coordinates of all feature points according to the extended sparse point cloud; wherein, the feature points include first feature points and second feature points; Projecting the three-dimensional point coordinates corresponding to the feature points onto an image plane according to the initial camera parameters of the video camera, to obtain predicted image coordinates; the initial camera parameters include initial internal parameters and initial external parameters; Constructing an objective function for bundle adjustment with the goal of minimizing the sum of squared errors between the observed image coordinates and the predicted image coordinates; Performing optimization solution on the objective function to obtain an optimal solution with the minimum objective function value, and determining the camera parameters of the video camera according to the optimal solution.
3. The multi-video and three-dimensional scene fusion method according to claim 1, characterized in that The determining a plurality of frustum planes of each of the video cameras according to the camera parameters, and using the plurality of frustum planes to crop the three-dimensional model, so as to obtain a plurality of three-dimensional model surfaces located inside the frustum planes after cropping includes: For each of the video cameras, based on the perspective projection principle, determining a plurality of frustum plane equations of the frustum of the video camera and frustum planes corresponding to the frustum plane equations; Using a three-dimensional transformation matrix, all vertices on the three-dimensional model are converted into a camera coordinate system of each of the video cameras, and according to equations of the vertices of the three-dimensional model in the camera coordinate system and the planes of the frustum, the positional relationship between the vertices of the three-dimensional model in the camera coordinate system and the planes of the frustum is determined; According to the positional relationship between each of the vertices and the frustum plane, the vertices located outside all the frustum planes are removed, and the vertices located inside at least one of the frustum planes are retained, so as to obtain the three-dimensional model surface clipped by the frustum plane under each of the video cameras.
4. The multi-video and three-dimensional scene fusion method according to claim 1 or 3, wherein The step of determining the overlapping area between the plurality of video frame images according to the plurality of three-dimensional model surfaces located inside the frustum plane comprises: Determine the bounding box of the three-dimensional model surface according to the cropped three-dimensional model surface under each video camera, and determine the area intersection of the bounding boxes corresponding to each video camera; The overlapping area between the multiple video frame images is determined according to the area intersection.
5. The multi-video and three-dimensional scene fusion method according to claim 1, wherein The clustering of the multiple three-dimensional model surfaces corresponding to the overlapping area, and projecting the three-dimensional model surfaces onto the frustum of each video camera using the clustering result, and comparing the projection areas to determine the best viewpoint of the overlapping area, comprises: Acquire multiple three-dimensional model surfaces in the overlapping area; Clustering the plurality of three-dimensional model faces to obtain a plurality of three-dimensional model face clusters; For each of the three-dimensional model surface clusters, the maximum projection area of the three-dimensional model surface cluster is initialized, and the three-dimensional model surface cluster is projected onto the near plane of the frustum of each of the video cameras to obtain the projected surfaces corresponding to each of the video cameras; Calculating the current projection area of the projected surface corresponding to each of the video cameras respectively, and comparing the size of each of the current projection areas with the maximum projection area; When the current projection area is greater than the maximum projection area, the current projection area is updated to the maximum projection area, and the maximum projection area among the current projection areas is found; The video camera corresponding to the found maximum projection area is used as the best viewpoint of the three-dimensional model surface cluster, and the best viewpoint of the overlapping area is determined according to the best viewpoint of each three-dimensional model surface cluster.
6. The multi-video and three-dimensional scene fusion method according to claim 5, wherein The step of segmenting the overlapping area by using the optimal viewpoint and mapping the segmentation result back to the three-dimensional model to determine a plurality of three-dimensional model segmentation surfaces includes: For each of the three-dimensional model face clusters, performing edge detection on the projected face corresponding to the best viewpoint; Based on the edge detection result, clustering the projected surface according to the gray value characteristics and color characteristics of the pixel points, and segmenting the projected surface; Mapping the segmentation results of each of the projected surfaces back to the three-dimensional model to determine the segmentation boundaries of the three-dimensional model; The three-dimensional model is segmented by a triangulation method according to the segmentation boundary to obtain a plurality of segmentation surfaces of the three-dimensional model.
7. The multi-video and three-dimensional scene fusion method according to claim 1, wherein Mapping the video texture of a plurality of video frame images to the spatial positions of the segmentation surfaces of the three-dimensional model to obtain a video three-dimensional fusion model in which the video frame images and the three-dimensional model are fused, includes: Determining the video texture coordinates respectively corresponding to the vertices of the segmentation surface of the three-dimensional model according to the positions of all the vertices of the segmentation surface of the three-dimensional model under the optimal viewing point; Determining the mapping relationship between the vertices and the video texture coordinates according to each vertex of each segmentation surface of the three-dimensional model and the video texture coordinates respectively corresponding to the vertices of the segmentation surface of the three-dimensional model; Based on the mapping relationship, mapping the video texture of a plurality of video frame images to the vertices of each segmentation surface of the three-dimensional model to obtain an initial video three-dimensional fusion model; Performing a rendering operation on the initial video three-dimensional fusion model to obtain a video three-dimensional fusion model.
8. A multi-video and 3D scene fusion system, characterized in that, Includes: A data acquisition module for acquiring two-dimensional image data of a target object and acquiring video frame images of the target object captured by a plurality of video cameras; A parameter calibration module for reconstructing a three-dimensional model of the target object through the two-dimensional image data and calibrating the camera parameters of each of the video cameras by using the three-dimensional model; A model clipping module for determining a plurality of frustum planes of each of the video cameras according to the camera parameters and clipping the three-dimensional model by using the plurality of frustum planes to obtain a plurality of three-dimensional model surfaces located inside the frustum planes after clipping; An overlap recognition module for determining an overlap region between a plurality of video frame images according to the plurality of three-dimensional model surfaces located inside the frustum planes; A viewing point determination module for clustering the plurality of three-dimensional model surfaces corresponding to the overlap region and projecting the three-dimensional model surfaces onto the frustums of each of the video cameras by using the clustering result, and comparing the projected areas to determine the optimal viewing point of the overlap region; The optimal viewing point is the video camera at the optimal shooting point corresponding to the overlap region; A region segmentation module for segmenting the overlap region by using the optimal viewing point and mapping the segmentation result back to the three-dimensional model to determine a plurality of segmentation surfaces of the three-dimensional model; A video fusion module for mapping the video texture of a plurality of video frame images to the spatial positions of the segmentation surfaces of the three-dimensional model to obtain a video three-dimensional fusion model in which the video frame images and the three-dimensional model are fused.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor. When a computer program stored in the memory is executed by the processor, the processor executes the steps of the multi-video and three-dimensional scene fusion method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, the steps of the multi-video and three-dimensional scene fusion method according to any one of claims 1-7 are implemented.
Citation Information
Cited By
Intelligent data transmission system and method applied to video monitoring platform
CN120639938A
Camera point complementing method for global tracking
CN120751279A