3D model imbedding method and system based on pre-recorded 2D video
By performing frame-by-frame three-dimensional point cloud reconstruction and camera motion trajectory inversely pushing the camera, combining 3D grid space and spherical ambient lighting, the complex and inflexible problem of 3D models integrating 2D videos in the prior art is solved, and efficient and automated video synthesis effect is achieved.
Patent Information
- Application Number
- CN202411880223.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is complex and inflexible when integrating 3D models into 2D videos. It depends on specific devices and complex post-synthesis, making it difficult to achieve efficient and automated processing processes.
By performing pixel-level three-dimensional point cloud reconstruction on pre-recorded 2D video frame by frame, inversely pushing the camera motion trajectory, compute the 3D grid space, place the 3D model into the 3D grid space, and perform spherical ambient lighting inference to generate a complete lighting environment.
A more efficient and automated 3D model placement process is realized, and stable and reliable trajectory data is generated, which improves the fidelity and automation of video synthesis effects.
Smart Images

Figure CN119946375A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video processing technology, and in particular to a 3D model placement method based on a pre-recorded 2D video and a 3D model placement system based on a pre-recorded 2D video. Background Art
[0002] Currently, traditional methods of integrating 3D models into 2D videos are usually complicated, not only requiring a lot of manual calibration, but also often relying on specific shooting equipment or complex post-synthesis technology, which limits its practicality and flexibility.
[0003] The related prior art and its defects are as follows:
[0004] (1) Camera trajectory inversion technology based on plane tracking: The current technology that uses local features in the video to perform plane tracking to reconstruct the camera motion trajectory has the main limitation of being highly dependent on the stability of feature points in the image. This dependence makes it very easy for the system to lose tracking points when faced with fast-moving scenes or scenes lacking significant textures, resulting in the failure of camera trajectory reconstruction. In addition, this technology is particularly sensitive to factors such as lighting changes, occlusions, and perspective changes, which further increases the difficulty of tracking. Therefore, in actual applications, a lot of manual correction and intervention are usually required to ensure the accuracy of the tracking results, which not only consumes a lot of human resources, but also limits its degree of automation and the possibility of large-scale application.
[0005] (2) Real-time video recording and 3D model embedding: Currently, the technical solution for embedding three-dimensional objects into recorded videos in real time mainly relies on the built-in horizontal gyroscope of the shooting device to obtain posture information. However, the disadvantage of this method is that it can only work with specific hardware support and its accuracy is limited by the performance of the device itself. More importantly, this method cannot be applied to existing pre-recorded video materials because the latter lacks the necessary dynamic data (such as gyroscope readings), making it impossible for the system to accurately integrate 3D objects into the original video scene. This means that the scope of application of this technology is strictly limited and cannot be operated independently without a dedicated shooting device.
[0006] (3) 3D scene construction and fusion: Although existing methods for creating 3D scenes from video data can simplify some 3D modeling workflows, it is still difficult to achieve direct conversion from raw videos to 3D scenes without manual intervention or pre-existing 3D models. In particular, existing technologies often perform poorly when it comes to complex physical interactions, such as occlusion relationships between objects, light reflection, and projection effects. These technical bottlenecks limit the generation of realistic 3D content and the natural integration of existing elements in the video, thereby affecting the authenticity and immersion of the final visual effect. Summary of the invention
[0007] In response to the above problems, the present invention provides a 3D model placement method and system based on pre-recorded 2D video, which reconstructs the three-dimensional point cloud at the pixel level for each frame of the input video data, reversely infers the motion trajectory of the camera during the shooting process, calculates the 3D grid space according to the image depth information, and places the 3D model in the 3D grid space, providing a more efficient and automated processing flow, which not only simplifies the operation steps, but also can generate more stable and reliable trajectory data, ensuring that the final output video synthesis effect is more in line with the actual logic and visual experience.
[0008] To achieve the above object, the present invention provides a 3D model placement method based on a pre-recorded 2D video, comprising:
[0009] Obtaining a pre-recorded 2D video, and performing 3D point cloud inference on the 2D video frame by frame;
[0010] Based on the three-dimensional point cloud obtained by inference, the motion trajectory of the camera in the 3D environment is inferred;
[0011] Performing depth image inference on the 2D video, and calculating a 3D grid space based on pixel positions of the depth image obtained by inference;
[0012] Determine a coordinate point corresponding to a designated point of the 2D video plane coordinate system in the 3D grid space, and place a 3D model in the 3D grid space based on the coordinate point;
[0013] Spherical environment lighting inference is performed according to the position of the 3D model in the 3D grid space to obtain a 720-degree spherical environment lighting map, the lighting vacant areas in the 720-degree spherical environment lighting map are completed to form a complete lighting environment, and 3D software is used for video stream rendering.
[0014] In the above technical solution, preferably, the motion trajectory of the camera in the 3D environment is inferred based on the three-dimensional point cloud obtained by reasoning, and the specific process includes:
[0015] Performing pairwise comparison on the three-dimensional point clouds generated from all frames of the 2D video;
[0016] Calculate the overlap of shared point cloud data between frames in three-dimensional space, and determine the viewing angle at which the overlap is the highest;
[0017] Determine the camera space transformation increment with the highest credibility between the two frames based on the highest viewing angle;
[0018] The complete motion trajectory of the camera when shooting the video in the 3D environment is calculated according to the camera space transformation increments between all adjacent frames of the 2D video.
[0019] In the above technical solution, preferably, depth image inference is performed on the 2D video, and the 3D grid space is calculated based on the pixel position of the depth image obtained by inference. The specific process includes:
[0020] According to the depth image obtained by inferring the 2D video frame by frame, reading the depth information in the depth image pixel by pixel;
[0021] Convert the coordinates of each pixel in the depth image into coordinates based on the camera in the viewing plane reference system;
[0022] Convert the pixel's view plane coordinates into a vector deflection angle based on the camera's 3D coordinate system;
[0023] The depth information is used as the vector modulus and combined with the vector deflection angle to calculate the coordinate point in the polar coordinate system of the camera 3D grid space;
[0024] The pixels of the 2D video are pushed out along polar coordinates so that each pixel is placed at a reasonable perspective position relative to the camera in the 3D environment, thereby obtaining a 3D grid space.
[0025] In the above technical solution, preferably, the step of determining the coordinate point corresponding to the designated point of the 2D video plane coordinate system in the 3D grid space and placing the 3D model in the 3D grid space based on the coordinate point includes:
[0026] According to a designated point in a plane coordinate system based on a single frame of video, finding a vertex position of the designated point corresponding to the 3D grid space and positions of surrounding vertices;
[0027] An average plane normal vector of the vertex and surrounding vertices is calculated, and a 3D model is placed in the 3D grid space along the plane normal vector passing through the vertex.
[0028] In the above technical solution, preferably, the spherical environment lighting inference is performed according to the position of the 3D model in the 3D grid space to obtain a 720-degree spherical environment lighting map, and the lighting vacant area is completed according to the 720-degree spherical environment lighting map to form a complete lighting environment. The specific process includes:
[0029] After confirming the position of the 3D model in the 3D grid space, a 720-degree spherical environment lighting map is captured using the position of the perspective camera as the origin of the panoramic camera;
[0030] Combined with the reflection vector of a single frame pixel in the camera view in the 3D grid space, the video image is exposed in multiple levels to calculate the possible light source and brightness of the object in the scene;
[0031] The diffusion model is used to fill in the missing areas of the light map to form a complete lighting environment for that perspective.
[0032] The present invention also proposes a 3D model placement system based on pre-recorded 2D video, which uses the 3D model placement method based on pre-recorded 2D video disclosed in any one of the above technical solutions, including:
[0033] A point cloud inference module is used to obtain a pre-recorded 2D video and perform three-dimensional point cloud inference on the 2D video frame by frame;
[0034] The trajectory inversion module is used to infer the motion trajectory of the camera in the 3D environment based on the three-dimensional point cloud obtained by reasoning;
[0035] A spatial reasoning module, used to perform depth image reasoning on the 2D video, and calculate a 3D grid space based on the pixel positions of the depth image obtained by reasoning;
[0036] A coordinate conversion module, used to determine a corresponding coordinate point of a designated point of the 2D video plane coordinate system in the 3D grid space, and place a 3D model in the 3D grid space based on the coordinate point;
[0037] The lighting rendering module is used to perform spherical environment lighting reasoning according to the position of the 3D model in the 3D grid space, obtain a 720-degree spherical environment lighting map, complete the lighting vacant areas in the 720-degree spherical environment lighting map to form a complete lighting environment, and use 3D software to render the video stream.
[0038] In the above technical solution, preferably, the trajectory inversion module is specifically used for:
[0039] Performing pairwise comparison on the three-dimensional point clouds generated from all frames of the 2D video;
[0040] Calculate the overlap of shared point cloud data between frames in three-dimensional space, and determine the viewing angle at which the overlap is the highest;
[0041] Determine the camera space transformation increment with the highest credibility between the two frames based on the highest viewing angle;
[0042] The complete motion trajectory of the camera when shooting the video in the 3D environment is calculated according to the camera space transformation increments between all adjacent frames of the 2D video.
[0043] In the above technical solution, preferably, the spatial reasoning module is specifically used for:
[0044] According to the depth image obtained by inferring the 2D video frame by frame, reading the depth information in the depth image pixel by pixel;
[0045] Convert the coordinates of each pixel in the depth image into coordinates based on the camera in the viewing plane reference system;
[0046] Convert the pixel's view plane coordinates into a vector deflection angle based on the camera's 3D coordinate system;
[0047] The depth information is used as the vector modulus and combined with the vector deflection angle to calculate the coordinate point in the polar coordinate system of the camera 3D grid space;
[0048] The pixels of the 2D video are pushed out along polar coordinates so that each pixel is placed at a reasonable perspective position relative to the camera in the 3D environment, thereby obtaining a 3D grid space.
[0049] In the above technical solution, preferably, the coordinate conversion module is specifically used for:
[0050] According to a designated point in a plane coordinate system based on a single frame of video, finding a vertex position of the designated point corresponding to the 3D grid space and positions of surrounding vertices;
[0051] An average plane normal vector of the vertex and surrounding vertices is calculated, and a 3D model is placed in the 3D grid space along the plane normal vector passing through the vertex.
[0052] In the above technical solution, preferably, the lighting rendering module is specifically used for:
[0053] After confirming the position of the 3D model in the 3D grid space, a 720-degree spherical environment lighting map is captured using the position of the perspective camera as the origin of the panoramic camera;
[0054] Combined with the reflection vector of a single frame pixel in the camera view in the 3D grid space, the video image is exposed in multiple levels to calculate the possible light source and brightness of the object in the scene;
[0055] The diffusion model is used to fill in the missing areas of the light map to form a complete lighting environment for that perspective.
[0056] Compared with the prior art, the present invention has the following beneficial effects:
[0057] (1) Reconstruct the 3D point cloud at the pixel level frame by frame in the input video data, solve the minimum fitting error of the point cloud data set, obtain the scene data with the highest credibility, and further infer the motion trajectory of the camera during the shooting process. Compared with the traditional technology that only relies on feature points in two-dimensional (2D) video to track and solve the camera trajectory, this solution method based on the overall 3D reconstruction of all pixels of the video provides a more efficient and automated processing flow. The calculation model of the present invention not only simplifies the operation steps, but also can generate more stable and reliable trajectory data, ensuring that the final output video synthesis effect is more in line with the actual logic and visual experience.
[0058] (2) With only the original video as input, the system can complete a series of complex processing from video analysis, 3D reconstruction to virtual object embedding, thus realizing an "end-to-end" solution. This method gets rid of the dependence on shooting equipment and reduces the requirements for the working environment, making the application of technology more flexible and convenient, and more user-friendly, so that even non-professionals can easily get started.
[0059] (3) Deep learning technology is applied to the generation of 3D grid space. On this basis, spatial reflection vectors are used to simulate 720-degree panoramic lighting, thereby creating a realistic lighting effect that conforms to physical laws in the virtual environment. This method effectively replaces the tedious and time-consuming steps in the traditional manual 3D modeling process, greatly improving work efficiency and reducing the user's operating burden. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 A schematic flow chart of a method for embedding a 3D model based on a pre-recorded 2D video disclosed in an embodiment of the present invention;
[0061] Figure 2 A schematic diagram of modules of a 3D model embedding system based on pre-recorded 2D video disclosed in an embodiment of the present invention.
[0062] In the figure, the corresponding relationship between each component and the reference numeral is as follows:
[0063] 1. Point cloud reasoning module, 2. Trajectory inference module, 3. Spatial reasoning module, 4. Coordinate transformation module, 5. Lighting rendering module. DETAILED DESCRIPTION
[0064] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0065] The present invention is further described in detail below in conjunction with the accompanying drawings:
[0066] like Figure 1 As shown, a 3D model placement method based on a pre-recorded 2D video provided by the present invention includes:
[0067] Obtain pre-recorded 2D video and perform 3D point cloud inference on the 2D video frame by frame;
[0068] Based on the three-dimensional point cloud obtained by inference, the motion trajectory of the camera in the 3D environment is inferred;
[0069] Perform depth image inference on 2D video and calculate 3D grid space based on the pixel position of the inferred depth image;
[0070] Determine the coordinate point corresponding to the specified point of the 2D video plane coordinate system in the 3D grid space, and place the 3D model in the 3D grid space based on the coordinate point;
[0071] According to the position of the 3D model in the 3D grid space, spherical environmental lighting inference is performed to obtain a 720-degree spherical environmental lighting map. The lighting vacant areas in the 720-degree spherical environmental lighting map are completed to form a complete lighting environment, and 3D software is used for video stream rendering.
[0072] In this implementation, by performing pixel-level 3D point cloud reconstruction on the input video data frame by frame, the motion trajectory of the camera during the shooting process is inferred, the 3D grid space is calculated based on the image depth information, and the 3D model is placed in the 3D grid space, providing a more efficient and automated processing flow, which not only simplifies the operation steps, but also can generate more stable and reliable trajectory data, ensuring that the final output video synthesis effect is more in line with actual logic and visual experience.
[0073] Specifically, a single two-dimensional (2D) video is accepted as input, and on this basis, the seamless fusion of the 2D video and the three-dimensional (3D) model can be achieved without any additional manual intervention or external equipment support. The core of the present invention lies in the end-to-end processing capability, that is, the entire process from receiving the original video input to finally presenting the 3D model in the video frame is completed automatically. The so-called "embedding" is not limited to simply superimposing the 3D model on the 2D video screen, but more emphasizes that the position and properties of the 3D model should be intelligently positioned and adjusted according to the actual environmental logic in the video, such as viewing angle, focal length, and dynamic lighting conditions. This includes ensuring the natural interaction between the 3D model and the video scene, such as correctly handling shadow projection, light source reflection, and the impact of ambient light on the 3D model, so as to achieve visual consistency and coherence.
[0074] In the above implementation, preferably, the motion trajectory of the camera in the 3D environment is inferred based on the three-dimensional point cloud obtained by inference, and the specific process includes:
[0075] Compare the 3D point clouds generated from all frames of the 2D video pairwise;
[0076] Calculate the overlap of shared point cloud data between frames in three-dimensional space, and determine the viewing angle at which the overlap is the highest;
[0077] Determine the camera space transformation increment with the highest credibility between the two frames based on the highest viewing angle;
[0078] The complete motion trajectory of the camera when shooting video in a 3D environment is calculated based on the camera space transformation increments between all adjacent frames of the 2D video.
[0079] Among them, for moving objects in the video, a binary mask is used to ignore them to reduce their impact when fitting the point cloud between two frames.
[0080] In the above implementation, preferably, depth image inference is performed on the 2D video, and the 3D grid space is calculated based on the pixel position of the depth image obtained by inference. The specific process includes:
[0081] After obtaining the motion trajectory of the camera, the depth information in the depth image is read pixel by pixel based on the depth image obtained by reasoning the 2D video frame by frame;
[0082] Convert the coordinates of each pixel in the depth image to the coordinates based on the camera's view plane reference system;
[0083] Convert the pixel's view plane coordinates into a vector deflection angle based on the camera's 3D coordinate system;
[0084] The depth information is used as the vector modulus and combined with the vector deflection angle to calculate the coordinate point in the polar coordinate system of the camera's 3D grid space;
[0085] The pixels of the 2D video are pushed out along the polar coordinates so that each pixel is placed in a reasonable perspective position relative to the camera in the 3D environment, thereby obtaining a 3D grid space composed of pixels of a single frame of the video.
[0086] In the above implementation, preferably, determining the corresponding coordinate point of the designated point of the 2D video plane coordinate system in the 3D grid space, and placing the 3D model in the 3D grid space based on the coordinate point, the specific process includes:
[0087] According to a specified point in a plane coordinate system based on a single frame of video, find the vertex position of the specified point in the 3D grid space and the positions of the surrounding vertices;
[0088] Calculate the average plane normal vector of the vertex and surrounding vertices, and place the 3D model in the 3D grid space along the plane normal vector passing through the vertex.
[0089] In the above implementation, preferably, spherical environment lighting inference is performed according to the position of the 3D model in the 3D grid space to obtain a 720-degree spherical environment lighting map, and the lighting vacant areas in the 720-degree spherical environment lighting map are completed to form a complete lighting environment. The specific process includes:
[0090] After confirming the position of the 3D model in the 3D grid space, take a 720-degree spherical ambient lighting map using the perspective camera’s position as the origin of the panoramic camera;
[0091] Combined with the reflection vector of a single frame pixel in the 3D grid space in the camera view, the video image is exposed in multiple levels to calculate the possible light source and brightness of the object in the scene;
[0092] The diffusion model is used to fill in the missing areas of the light map to form a complete lighting environment for that perspective.
[0093] Furthermore, after placing the 3D model in a 3D environment generated by the video, the environment is rendered in 3D software (such as Blender 4.2.2), and a new video stream in which the 3D model is "placed" in the environment shown in the video can be obtained.
[0094] like Figure 2 As shown, the present invention further proposes a 3D model placement system based on pre-recorded 2D video, which uses the 3D model placement method based on pre-recorded 2D video disclosed in any one of the above embodiments, including:
[0095] Point cloud reasoning module 1, used to obtain pre-recorded 2D video and perform 3D point cloud reasoning on the 2D video frame by frame;
[0096] The trajectory inversion module 2 is used to infer the motion trajectory of the camera in the 3D environment based on the three-dimensional point cloud obtained by inference;
[0097] The spatial reasoning module 3 is used to perform depth image reasoning on the 2D video and calculate the 3D grid space according to the pixel position of the depth image obtained by reasoning;
[0098] A coordinate conversion module 4 is used to determine the corresponding coordinate point of the designated point of the 2D video plane coordinate system in the 3D grid space, and place the 3D model in the 3D grid space based on the coordinate point;
[0099] The lighting rendering module 5 is used to perform spherical environment lighting inference according to the position of the 3D model in the 3D grid space, obtain a 720-degree spherical environment lighting map, complete the lighting vacant areas in the 720-degree spherical environment lighting map to form a complete lighting environment, and use 3D software to render the video stream.
[0100] In this embodiment, a 3D scene is inferred based on a 2D video, and a more accurate motion trajectory of the camera in a 3D environment is obtained by inferring the 3D scene. A depth image is inferred frame by frame based on a 2D video, and a 3D grid space composed of pixels of the current frame is calculated by the depth image. Based on the constructed 3D grid space, a 3D model is placed in the 3D grid space by giving a coordinate point in a screen. Based on the constructed 3D grid space, a 720-degree spherical lighting environment that the 3D model should be exposed to is generated by the position where the object is placed.
[0101] In the above implementation, preferably, the trajectory inversion module 2 is specifically used for:
[0102] Compare the 3D point clouds generated from all frames of the 2D video pairwise;
[0103] Calculate the overlap of shared point cloud data between frames in three-dimensional space, and determine the viewing angle at which the overlap is the highest;
[0104] Determine the camera space transformation increment with the highest credibility between the two frames based on the highest viewing angle;
[0105] The complete motion trajectory of the camera when shooting video in a 3D environment is calculated based on the camera space transformation increments between all adjacent frames of the 2D video.
[0106] In the above implementation, preferably, the spatial reasoning module 3 is specifically used for:
[0107] According to the depth image obtained by inferring the 2D video frame by frame, the depth information in the depth image is read pixel by pixel;
[0108] Convert the coordinates of each pixel in the depth image to the coordinates based on the camera's view plane reference system;
[0109] Convert the pixel's view plane coordinates into a vector deflection angle based on the camera's 3D coordinate system;
[0110] The depth information is used as the vector modulus and combined with the vector deflection angle to calculate the coordinate point in the polar coordinate system of the camera's 3D grid space;
[0111] The pixels of the 2D video are pushed out along the polar coordinates so that each pixel is placed at a reasonable perspective position relative to the camera in the 3D environment, thereby obtaining a 3D grid space.
[0112] In the above implementation, preferably, the coordinate conversion module 4 is specifically used for:
[0113] According to a specified point in a plane coordinate system based on a single frame of video, find the vertex position of the specified point in the 3D grid space and the positions of the surrounding vertices;
[0114] Calculate the average plane normal vector of the vertex and surrounding vertices, and place the 3D model in the 3D grid space along the plane normal vector passing through the vertex.
[0115] In the above implementation, preferably, the lighting rendering module 5 is specifically used for:
[0116] After confirming the position of the 3D model in the 3D grid space, take a 720-degree spherical ambient lighting map using the perspective camera’s position as the origin of the panoramic camera;
[0117] Combined with the reflection vector of a single frame pixel in the 3D grid space in the camera view, the video image is exposed in multiple levels to calculate the possible light source and brightness of the object in the scene;
[0118] The diffusion model is used to fill in the missing areas of the light map to form a complete lighting environment for that perspective.
[0119] According to the 3D model placement method and system based on pre-recorded 2D video disclosed in the above-mentioned embodiment, during the implementation process, the effects of the method and system of the present invention are specifically described through the following examples.
[0120] A chrome ball is placed in the video, which produces correct reflection and projection of the scene shown in the picture (including the invisible part). By comparing the 3D chrome ball before and after it is placed in a certain frame of the video, it can be seen that the chrome ball can produce reasonable lighting interaction with the scene. The smooth chrome ball casts shadows on the surrounding environment and produces pure mirror reflection with a real sense of distance for objects such as colorful flags, stone walls, floor tiles, bicycles, and the sky (off-screen scenes). That is, the visual experience is truly "placed" in the video.
[0121] A box of juice is placed on the object in the video, and the juice maintains a uniform transformation relationship with the object in different frames. The object maintains a uniform transformation relationship with the environment shown in the 2D video in different frames, and it can be seen that the relative rotation, displacement, and size of the object and the juice remain consistent.
[0122] In the scene shown in the video, the projections of objects of different shapes are also different. Under the premise of the same 2D video, it can be seen that the projections of different 3D models are also different, which means that the shape of the physical projection is determined by the shape of the object, which conforms to objective physical logic.
[0123] When the placement of the 3D model is changed in the video, the object projection follows the movement, that is, the object's projection on the scene shown in the video is dynamically related to the scene.
[0124] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A 3D model placement method based on pre-recorded 2D video, characterized in that: include: Obtaining a pre-recorded 2D video, and performing 3D point cloud inference on the 2D video frame by frame; Based on the three-dimensional point cloud obtained by inference, the motion trajectory of the camera in the 3D environment is inferred; Performing depth image inference on the 2D video, and calculating a 3D grid space based on pixel positions of the depth image obtained by inference; Determine a coordinate point corresponding to a designated point of the 2D video plane coordinate system in the 3D grid space, and place a 3D model in the 3D grid space based on the coordinate point; Spherical environment lighting inference is performed according to the position of the 3D model in the 3D grid space to obtain a 720-degree spherical environment lighting map, the lighting vacant areas in the 720-degree spherical environment lighting map are completed to form a complete lighting environment, and 3D software is used for video stream rendering.
2. The 3D model placement method based on pre-recorded 2D video according to claim 1, characterized in that: The motion trajectory of the camera in the 3D environment is obtained by inferring the 3D point cloud obtained by inference. The specific process includes: Performing pairwise comparison on the three-dimensional point clouds generated from all frames of the 2D video; Calculate the overlap of shared point cloud data between frames in three-dimensional space, and determine the viewing angle at which the overlap is the highest; Determine the camera space transformation increment with the highest credibility between the two frames based on the highest viewing angle; The complete motion trajectory of the camera when shooting the video in the 3D environment is calculated according to the camera space transformation increments between all adjacent frames of the 2D video.
3. The 3D model placement method based on pre-recorded 2D video according to claim 1, characterized in that: Performing depth image inference on the 2D video, and calculating the 3D grid space based on the pixel position of the depth image obtained by inference, the specific process includes: According to the depth image obtained by inferring the 2D video frame by frame, reading the depth information in the depth image pixel by pixel; Convert the coordinates of each pixel in the depth image to the coordinates of the camera in the viewing plane reference system; Convert the pixel's view plane coordinates into a vector deflection angle based on the camera's 3D coordinate system; The depth information is used as the vector modulus and combined with the vector deflection angle to calculate the coordinate point in the polar coordinate system of the camera 3D grid space; The pixels of the 2D video are pushed out along polar coordinates so that each pixel is placed at a reasonable perspective position relative to the camera in the 3D environment, thereby obtaining a 3D grid space.
4. The method for embedding a 3D model based on a pre-recorded 2D video according to claim 1, characterized in that: The process of determining the coordinate point corresponding to the designated point of the 2D video plane coordinate system in the 3D grid space and placing the 3D model in the 3D grid space based on the coordinate point specifically includes: According to a designated point in a plane coordinate system based on a single frame of video, finding the vertex position of the designated point corresponding to the 3D grid space and the positions of surrounding vertices; An average plane normal vector of the vertex and surrounding vertices is calculated, and the 3D model is placed in the 3D grid space along the plane normal vector passing through the vertex.
5. The method for embedding a 3D model based on a pre-recorded 2D video according to claim 1, characterized in that: The spherical environment lighting inference is performed according to the position of the 3D model in the 3D grid space to obtain a 720-degree spherical environment lighting map, and the lighting vacant area is completed according to the 720-degree spherical environment lighting map to form a complete lighting environment. The specific process includes: After confirming the position of the 3D model in the 3D grid space, a 720-degree spherical environment lighting map is captured using the position of the perspective camera as the origin of the panoramic camera; Combined with the reflection vector of a single frame pixel in the camera view in the 3D grid space, the video image is exposed in multiple levels to calculate the possible light source and brightness of the object in the scene; The diffusion model is used to fill in the missing areas of the light map to form a complete lighting environment for that perspective.
6. A 3D model placement system based on pre-recorded 2D video, characterized in that: The method for embedding a 3D model based on a pre-recorded 2D video as claimed in any one of claims 1 to 5 comprises: A point cloud inference module is used to obtain a pre-recorded 2D video and perform three-dimensional point cloud inference on the 2D video frame by frame; The trajectory inversion module is used to infer the motion trajectory of the camera in the 3D environment based on the three-dimensional point cloud obtained by reasoning; A spatial reasoning module, used to perform depth image reasoning on the 2D video, and calculate a 3D grid space based on the pixel positions of the depth image obtained by reasoning; A coordinate conversion module, used to determine a corresponding coordinate point of a designated point of the 2D video plane coordinate system in the 3D grid space, and place a 3D model in the 3D grid space based on the coordinate point; The lighting rendering module is used to perform spherical environment lighting reasoning according to the position of the 3D model in the 3D grid space, obtain a 720-degree spherical environment lighting map, complete the lighting vacant areas in the 720-degree spherical environment lighting map to form a complete lighting environment, and use 3D software to render the video stream.
7. The 3D model placement system based on pre-recorded 2D video according to claim 6, characterized in that: The trajectory inversion module is specifically used for: Performing pairwise comparison on the three-dimensional point clouds generated from all frames of the 2D video; Calculate the overlap of shared point cloud data between frames in three-dimensional space, and determine the viewing angle at which the overlap is the highest; Determine the camera space transformation increment with the highest credibility between the two frames based on the highest viewing angle; The complete motion trajectory of the camera when shooting the video in the 3D environment is calculated according to the camera space transformation increments between all adjacent frames of the 2D video.
8. The 3D model placement system based on pre-recorded 2D video according to claim 6, characterized in that: The spatial reasoning module is specifically used for: According to the depth image obtained by inferring the 2D video frame by frame, reading the depth information in the depth image pixel by pixel; Convert the coordinates of each pixel in the depth image to the coordinates of the camera in the viewing plane reference system; Convert the pixel's view plane coordinates into a vector deflection angle based on the camera's 3D coordinate system; The depth information is used as the vector modulus and combined with the vector deflection angle to calculate the coordinate point in the polar coordinate system of the camera 3D grid space; The pixels of the 2D video are pushed out along polar coordinates so that each pixel is placed at a reasonable perspective position relative to the camera in the 3D environment, thereby obtaining a 3D grid space.
9. The 3D model placement system based on pre-recorded 2D video according to claim 6, characterized in that: The coordinate conversion module is specifically used for: According to a designated point in a plane coordinate system based on a single frame of video, finding the vertex position of the designated point corresponding to the 3D grid space and the positions of surrounding vertices; An average plane normal vector of the vertex and surrounding vertices is calculated, and a 3D model is placed in the 3D grid space along the plane normal vector passing through the vertex.
10. The 3D model placement system based on pre-recorded 2D video according to claim 6, characterized in that: The illumination reasoning module and the lighting rendering module are specifically used for: After confirming the position of the 3D model in the 3D grid space, a 720-degree spherical environment lighting map is captured using the position of the perspective camera as the origin of the panoramic camera; Combined with the reflection vector of a single frame pixel in the camera view in the 3D grid space, the video image is exposed in multiple levels to calculate the possible light source and brightness of the object in the scene; The diffusion model is used to fill in the missing areas of the light map to form a complete lighting environment for that perspective.