A Fusion Method of Dynamic Video and 3D Model Based on Image Generation
By introducing NeRF three-dimensional geometric perception generation network and time series geometric information embedding, combined with weighted fusion and geometric consistency optimization, the problem of visual inconsistency and deformation in the fusion of dynamic video and three-dimensional model is solved, and high-precision and high-reality video generation is achieved.
Patent Information
- Application Number
- CN202411573936.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-06
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-11-06
AI Technical Summary
The prior art lacks an in-depth understanding of the three-dimensional dynamic geometric structure in the fusion of dynamic videos and three-dimensional models, resulting in the problems of visual inconsistency, deformation and insufficient real-time performance in complex dynamic scenarios.
The NeRF three-dimensional geometric perception generation network is introduced to process the three-dimensional geometric data in dynamic videos, generate geometric perception images, and capture the movement trajectory, rotation angle and viewing angle changes of objects through the geometric information embedding of time series. A weighted fusion algorithm based on three-dimensional model is used to fuse the image of each frame with the texture information of the three-dimensional model by pixel, and geometric consistency optimization and real-time lighting simulation are performed.
It realizes a dynamic understanding of three-dimensional geometric structure, ensures geometric consistency between frames, avoids frame skipping or object shape changes, and improves the accuracy and visual reality of video generation.
Smart Images

Figure CN119445043B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image generation, and particularly to a method for fusing dynamic videos and three-dimensional models based on image generation. Background Art
[0002] With the continuous development of image generation technology, the fusion of dynamic videos and three-dimensional models has been gradually applied to multiple fields. By seamlessly combining dynamic video content with three-dimensional virtual models, highly realistic visual effects and immersive experiences are created.
[0003] Currently, the mainstream image generation models used for the fusion of dynamic videos and three-dimensional models, such as GAN (Generative Adversarial Network) and VAE (Variational Autoencoder), are mainly used to generate high-quality two-dimensional images. However, such models lack an understanding of three-dimensional geometric structures and mostly only process pixel values in images, ignoring the three-dimensional shapes, positions, and motion trajectories of objects in the scene. As a result, when generating dynamic videos, it is difficult for the model to maintain the consistency of the object's geometric structure, and problems such as object deformation, distortion, or incorrect position often occur.
[0004] Due to the lack of dynamic understanding of geometric structures, it is difficult to guarantee the continuity between generated video frames. Traditional image generation models cannot effectively capture the geometric characteristics of three-dimensional objects changing over time, and this problem is particularly evident in dynamic scenes. Objects in the video may experience discontinuous jumps or morphological changes between different frames, seriously affecting visual fluency.
[0005] To solve the above problems, some solutions introduce technologies such as key frame interpolation and optical flow to estimate the motion trajectories of objects and insert transitions between frames. However, such methods rely too much on motion prediction models and are difficult to handle complex three-dimensional geometric information. When an object rotates or undergoes large-scale motion, the prediction accuracy will drop significantly, affecting the generation effect.
[0006] In addition, there are also some solutions that use texture mapping technology to attach the generated two-dimensional images to the surface of three-dimensional models. However, in dynamic scenes, texture mapping often has problems such as stretching, tearing, or misalignment. Therefore, there is an urgent need for a method for fusing dynamic videos and three-dimensional models based on image generation to solve such problems. Summary of the Invention
[0007] In view of the above existing problems, the present invention is proposed.
[0008] The present invention provides a method for fusing dynamic videos and 3D models based on image generation, which solves the problems of traditional dynamic video generation methods based on image generation and 3D model fusion, such as the lack of in-depth understanding of 3D dynamic geometric structures, and the generated results still having visual inconsistencies, deformations, and insufficient real-time performance in complex dynamic scenarios.
[0009] To solve the above technical problems, the present invention provides the following technical solutions:
[0010] The present invention provides a method for fusing dynamic videos and 3D models based on image generation, which includes,
[0011] Step S1, introducing a 3D geometric perception generation network,
[0012] Using a NeRF (Neural Radiance Fields) generation network to process the 3D geometric data in the dynamic video and generate a geometric perception image.
[0013] Step S2, embedding geometric information of the time series,
[0014] Based on the geometric perception image generated in step S1, introducing geometric change data of the time series to generate a time series geometric perception image.
[0015] Step S3, weighted fusion and geometric consistency optimization,
[0016] Based on the time series geometric perception image generated in step S2, using a weighted fusion algorithm based on the 3D model to fuse each frame of the image with the texture information of the 3D model pixel by pixel.
[0017] Step S4, final output and real-time rendering,
[0018] After the geometric consistency fusion in step S3, output the optimized image sequence and perform real-time rendering.
[0019] Further, in step S1, the 3D geometric data includes 3D geometric coordinates, object shapes, and perspective changes. NeRF generates an image matching the 3D model, that is, a geometric perception image, according to the 3D geometric coordinates, object shapes, and perspective changes in the scene.
[0020] Further, in step S1, the NeRF generation network processes the 3D geometric data to generate a geometric perception image. Specifically:
[0021] For each frame in the dynamic video, use the input of the 3D geometric data to generate a geometric perception image. The geometric data includes 3D coordinates, perspective changes, and object shape information. Define the 3D coordinates as x ∈ R 3 , and the perspective direction as d ∈ R 3, NeRF generates color \(c\in\mathbb{R}\) based on the coordinates of points in the scene and the viewing angles 3 and volume density \(\sigma\in\mathbb{R}\):
[0022] F θ (x, d) = (c, \(\sigma\)), where \(x\) represents the three-dimensional geometric coordinates of a sampling point in the scene, \(d\) represents the viewing direction from the camera to the sampling point, i.e., a vector, \(c\) represents the color value of the point under this viewing angle, the output is a color vector of RGB three channels, and \(\sigma\) represents the volume density of the point, i.e., the opacity of the point in the scene;
[0023] NeRF samples the rays \(r(t)\) emitted from the camera to each pixel in the scene, and solves the weighted sum of the color and opacity when the rays pass through the scene;
[0024] NeRF performs piecewise sampling on each ray to generate a series of sampling points \(x\) i = r(t i ), and uses the network F θ to calculate the color \(c\) i and volume density \(\sigma\) i ,
[0025] Perform weighted accumulation on the sampling points on each ray.
[0026] Furthermore, during the training process of step S1, NeRF optimizes the network parameters \(\theta\) by adjusting based on the mean squared error MSE loss between the real image and the synthetic image:
[0027] where \(L\) represents the total loss function, measuring the difference between the synthetic image and the real image, represents the set of rays, i.e., all the rays in the image, and \(C\) G (r) represents the color value of ray \(r\) in the real image, and \(C(r)\) represents the color value corresponding to this ray generated by the NeRF network,
[0028] Each ray \(r\) outputs its corresponding color \(C(r)\), thus synthesizing the entire geometry-aware image, and each pixel in the image contains the geometric information and illumination information in the scene.
[0029] Furthermore, in step S2, the way to introduce the geometric change data of the time series is:
[0030] Collect the motion trajectory, rotation angle, and viewing angle change information of the object, and dynamically embed the geometric changes in the time series.
[0031] Furthermore, in step S2, the embedding method of the geometric information of the time series is:
[0032] Based on the geometric perception image generated in step S1, introduce the geometric change data of the object in the time series, including the movement trajectory, rotation angle, and perspective change of the object. Set the three-dimensional position of the object at time t as x(t) ∈ R 3 , the rotation matrix of the object is R(t) ∈ SO(3), then the change of the position and orientation of the object over time is expressed as:
[0033]
[0034] where x(t) represents the three-dimensional position of the object at time t, x0 represents the initial position of the object, that is, the position at time t = 0, v(t) represents the instantaneous velocity vector of the object, that is, the movement trajectory of the object at each time point, R(t) represents the rotation matrix of the object at time t, that is, the rotation state of the object, R0 represents the initial rotation matrix of the object, and ω(t) represents the instantaneous angular velocity vector of the object, that is, the rotation change of the object.
[0035] Use the position and rotation data x(t) and R(t) of the object in the time series to construct the geometric perception image of each frame. Specifically, the geometric perception image I(t) at time t is generated based on the geometric information transformation of the previous frame I(t - Δt):
[0036] where I(t) represents the geometric perception image at time t, represents the geometric transformation function, which is responsible for transforming the previous frame image I(t - Δt) according to the displacement x(t) and rotation matrix R(t) of the object. I(t - Δt) represents the geometric perception image of the previous frame, and Δt represents the time interval between two frames.
[0037] Furthermore, in step S2, the embedding method of the geometric information of the time series also includes:
[0038] Re-calculate the light in the scene by combining the time series motion data of the object. The path r(t) of the light is dynamically adjusted according to the displacement and rotation of the object:
[0039] r(t) = o(t) + t·R(t)·d(t), where r(t) represents the light path at time t, o(t) represents the dynamic position of the camera optical center at time t, d(t) represents the camera viewing direction at time t, and is updated based on the movement and rotation of the object.
[0040] After combining the movement trajectory and perspective change of the object, each frame of the geometric perception image I(t) at time t is calculated by the ray integral generated by the NeRF network. The ray color C(t) generated by NeRF depends on the geometric information of the time series, and the formula is:
[0041] Among them, C(t) represents the color corresponding to the light ray r(t) at time t. Based on the object position and rotation state in the scene, T i (t) represents the cumulative transmittance at the i-th sampling point, defined as Among them, σ j (t) and δ j (t) represent the density and distance of the sampling point respectively. As time changes, σ i (t) represents the volume density of the object, which is updated over time. δ i (t) represents the distance between adjacent sampling points, defined as t i+1 (t) - t i (t), c i (t) represents the color value of the sampling point at time t. Combining the geometric change data in the time series, a geometrically consistent and continuous geometric perception image is generated for each frame compared with the previous frame.
[0042] Furthermore, in step S3, a geometric consistency optimization mechanism is introduced, and real-time lighting simulation is performed simultaneously to optimize the image authenticity.
[0043] In step S3, the fusion process adopts GPU acceleration and parallel computing technology.
[0044] Furthermore, in step S3, the weighted fusion and geometric consistency optimization methods are as follows:
[0045] Perform weighted fusion processing on each frame of image I(t), and combine it with the texture information of the 3D model. Assuming the texture information of the 3D model is T(x, y, z), the texture value of each point is given in the 3D space. The mapping between each pixel in the image and the 3D space point is:
[0046] I fused (t, u, v) = α·I(t, u, v) + (1 - α)·T(x(u, v), y(u, v), z(u, v)), where I(t, u, v) represents the color value of the image at pixel coordinates (u, v) at time t, T(x(u, v), y(u, v), z(u, v)) represents the texture value corresponding to the coordinates (x(u, v), y(u, v), z(u, v)) in the 3D model. This value corresponds to the mapping of the pixel position (u, v) in the image in the 3D space. α represents the fusion weight coefficient, which controls the weighted ratio of the original image and the texture information, and the value range is 0 ≤ α ≤ 1.
[0047] Introduce a geometric consistency optimization mechanism, compare the geometric information at the same position in the previous and subsequent frames for optimization. Let the same pixel position at time t and t - 1 be (u, v):
[0048] Lgeo = Σ (u,v) |I fused (t, u, v) - I fused (t - 1, u, v)| 2 , where L geo represents the geometric consistency loss function, that is, the color difference of the same pixel at times t and t - 1. I fused (t, u, v) represents the color value of the fused image at time t at pixel (u, v). I fused (t - 1, u, v) represents the color value of the fused image at time t - 1 at pixel (u, v).
[0049] Furthermore, in step S3, the weighted fusion and geometric consistency optimization method also includes:
[0050] Introduce light simulation to calculate the final light intensity of the pixel points:
[0051] I final (t, u, v) = I fused (t, u, v) · (k a · L a + k d · (n · l) + k s · (r · v) m ), where I final (t, u, v) represents the final light intensity at pixel (u, v) at time t, k a represents the ambient light coefficient, L a represents the ambient light intensity, k d represents the diffuse reflection coefficient, n represents the normal vector of the object surface, l represents the incident light direction vector, k s represents the specular reflection coefficient, r represents the reflected light direction vector, r = 2(n · l)n - l, v represents the camera viewing direction vector, m represents the specular highlight exponent to control the sharpness of the specular highlight. By combining real-time light simulation with ambient light, diffuse reflection, and specular reflection, calculate the final light intensity of each pixel point, thereby enhancing the authenticity and visual effect of the image.
[0052] The beneficial effects of the present invention are:
[0053] In the present invention, a NeRF three-dimensional geometric perception generation network is introduced to directly process three-dimensional geometric data, enabling each generated image to contain the geometric information of the scene, avoiding object deformation, distortion, or incorrect position. The three-dimensional perception ability of NeRF enables the shape and position of the object to always match the actual geometric structure in the scene, thereby achieving a dynamic understanding of the geometric structure during the generation process. At the same time, geometric information of the time series is embedded, which can effectively capture the movement trajectory, rotation angle, and perspective change of the object, ensuring geometric consistency between frames and avoiding frame skipping or sudden changes in object morphology.
[0054] In the present invention, the dynamic position and rotation matrix of the object are introduced into the time series, which can effectively capture and update the geometric state of the object in the scene, dynamically adjust the light path so that the movement trajectory of the object is completely consistent with the changes in the scene, effectively cope with large-scale movements and complex rotations, enhance the geometric capture ability, and improve the accuracy of video generation.
[0055] In the present invention, based on a weighted fusion algorithm of three-dimensional models, each frame of the image is fused with the texture information in the three-dimensional model pixel by pixel, and the mapping position of the three-dimensional model is calculated pixel by pixel, so that the texture information of each frame is closely combined with the surface of the three-dimensional geometric model, reducing the misalignment and stretching of the texture in the dynamic scene.
[0056] In the present invention, real-time lighting simulation is introduced, and ambient light, diffuse reflection, and specular reflection are incorporated into the rendering process, making the generated images not only geometrically consistent but also more realistic in lighting effects, enabling the objects in the video scene to exhibit natural light and shadow changes under different lighting conditions, increasing the visual realism and three-dimensional sense. Description of the Drawings
[0057] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0058] Figure 1 It is a schematic flowchart of the fusion method of the dynamic video and three-dimensional model based on image generation of the present invention. Detailed Embodiments
[0059] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will give a detailed description of the specific embodiments of the present invention in conjunction with the drawings in the specification.
[0060] In the following description, numerous specific details are set forth to provide a thorough understanding of the present invention. However, the present invention may be practiced in other ways different from those described herein. Persons skilled in the art can make similar extensions without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0061] Secondly, as used herein, "one embodiment" or "an embodiment" refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other.
[0062] Example 1, referring to Figure 1 , this embodiment provides a method for fusing dynamic video generated based on images with a three-dimensional model, including the following steps:
[0063] Step S1, introduce a three-dimensional geometric perception generation network,
[0064] Adopt a NeRF (Neural Radiance Fields) generation network to process the three-dimensional geometric data in the dynamic video and generate a geometric perception image.
[0065] In step S1, the three-dimensional geometric data includes three-dimensional geometric coordinates, object shape, and perspective change. NeRF generates an image matching the three-dimensional model according to the three-dimensional geometric coordinates, object shape, and perspective change in the scene, that is, a geometric perception image.
[0066] In step S1, the NeRF generation network processes the three-dimensional geometric data to generate a geometric perception image. Specifically:
[0067] For each frame in the dynamic video, use the input of the three-dimensional geometric data to generate a geometric perception image. The geometric data includes three-dimensional coordinates, perspective change, and object shape information. Define the three-dimensional coordinates as x ∈ R 3 , the perspective direction as d ∈ R 3 , and NeRF generates a color c ∈ R 3 and a volume density σ ∈ R according to the coordinates of the points in the scene and the perspective:
[0068] F θ (x, d) = (c, σ), where x represents the three-dimensional geometric coordinates of a sampling point in the scene, d represents the perspective direction from the camera to the sampling point, that is, a vector, c represents the color value of the point under this perspective, and the output is a color vector with three RGB channels, and σ represents the volume density of the point, that is, the opacity of the point in the scene.
[0069] NeRF samples the rays r(t) emitted from the camera to each pixel in the scene, and solves the weighted sum of the color and opacity when the rays pass through the scene;
[0070] NeRF performs piecewise sampling on each ray to generate a series of sampling points x i = r(t i ), and uses the network F θ to calculate the color c i and the volume density σ i ,
[0071] Perform weighted accumulation on the sampling points on each ray;
[0072] During the training process of step S1, NeRF optimizes the network parameters θ based on the mean squared error MSE loss between the real image and the synthetic image:
[0073] where L represents the total loss function, measuring the difference between the synthetic image and the real image, represents the set of rays, that is, all the rays in the image, C G (r) represents the color value of the ray r in the real image, and C(r) represents the color value corresponding to this ray generated by the NeRF network,
[0074] Each ray r outputs its corresponding color C(r), thus synthesizing the entire geometric perception image. Each pixel in the image contains geometric information and lighting information in the scene,
[0075] Specifically, the NeRF three-dimensional geometric perception generation network can process relatively complex three-dimensional geometric data, including the three-dimensional coordinates, shapes, and perspective changes of objects, so that the geometric information contained in each generated frame of the image can effectively reflect the actual object form in the scene. Combining density and color, each pixel corresponds one-to-one with a point in three-dimensional space.
[0076] Step S2, embedding of geometric information in the time series,
[0077] Based on the geometric perception image generated in step S1, introduce the geometric change data in the time series to generate the geometric perception image of the time series,
[0078] In step S2, the way to introduce the geometric change data in the time series is:
[0079] Collect the motion trajectory, rotation angle, and perspective change information of the object, and dynamically embed the geometric changes in the time series,
[0080] In step S2, the way to embed the geometric information in the time series is:
[0081] Based on the geometric perception image generated in step S1, introduce the geometric change data of the object in the time series, including the object's motion trajectory, rotation angle, and perspective change. Set the three-dimensional position of the object at time t as x(t) ∈ R 3 , and the rotation matrix of the object as R(t) ∈ SO(3). Then the change of the object's position and orientation over time is expressed as:
[0082]
[0083] where x(t) represents the three-dimensional position of the object at time t, x0 represents the initial position of the object, i.e., the position at time t = 0, v(t) represents the instantaneous velocity vector of the object, i.e., the motion trajectory of the object at each time point, R(t) represents the rotation matrix of the object at time t, i.e., the rotation state of the object, R0 represents the initial rotation matrix of the object, and ω(t) represents the instantaneous angular velocity vector of the object, i.e., the rotation change of the object.
[0084] Use the position and rotation data x(t) and R(t) of the object in the time series to construct the geometric perception image of each frame. Specifically, the geometric perception image I(t) at time t is generated based on the geometric information transformation of the previous frame I(t - Δt):
[0085] where I(t) represents the geometric perception image at time t, represents the geometric transformation function, which is responsible for transforming the previous frame image I(t - Δt) according to the displacement x(t) and rotation matrix R(t) of the object. I(t - Δt) represents the geometric perception image of the previous frame, and Δt represents the time interval between two frames.
[0086] In step S2, the geometric information embedding method of the time series also includes:
[0087] Re - calculate the light in the scene by combining the time - series motion data of the object. The path of the light r(t) is dynamically adjusted according to the displacement and rotation of the object:
[0088] r(t) = o(t)+t·R(t)·d(t), where r(t) represents the light path at time t, o(t) represents the dynamic position of the camera optical center at time t, and d(t) represents the camera viewing direction at time t, which is updated based on the motion and rotation of the object.
[0089] After combining the object's motion trajectory and perspective change, the geometric perception image I(t) of each frame at time t is calculated by the ray integral generated by the NeRF network. The ray color C(t) generated by NeRF depends on the geometric information of the time series, and the formula is:
[0090] Among them, C(t) represents the color corresponding to the light ray r(t) at time t. Based on the object position and rotation state in the scene, T i (t) represents the cumulative transmittance at the i-th sampling point, defined as Among them, σ j (t) and δ j (t) represent the density and distance of the sampling point respectively. As time changes, σ i (t) represents the volume density of the object, which is updated over time. δ i (t) represents the distance between adjacent sampling points, defined as t i+1 (t) - t i (t), c i (t) represents the color value of the sampling point at time t. Combining the geometric change data in the time series, a geometrically consistent and continuous geometric perception image is generated for each frame compared with the previous frame.
[0091] Specifically, by embedding the geometric information of the time series, the geometric continuity and consistency between frames are greatly improved. The motion trajectory, rotation angle, and perspective change of the object are introduced to dynamically adjust the geometric information in each frame of the image, ensuring the continuity and smoothness of the images in the time series. The generated video will not show jumps or distortions in the geometric structure, making the objects in the dynamic scene maintain morphological consistency at different time points.
[0092] Step S3, weighted fusion and geometric consistency optimization,
[0093] Based on the geometric perception images of the time series generated in step S2, a weighted fusion algorithm based on a 3D model is adopted, and the image of each frame is fused with the texture information of the 3D model pixel by pixel.
[0094] In step S3, a geometric consistency optimization mechanism is introduced, and real-time lighting simulation is carried out simultaneously to optimize the authenticity of the image.
[0095] In step S3, GPU acceleration and parallel computing technologies are used in the fusion process.
[0096] In step S3, the method of weighted fusion and geometric consistency optimization is as follows:
[0097] Perform weighted fusion processing on each frame of the image I(t), and combine it with the texture information of the 3D model. Assuming the texture information of the 3D model is T(x, y, z), the texture value of each point is given in the 3D space. The mapping between each pixel in the image and the 3D space point is:
[0098] I fused(t, u, v) = α·I(t, u, v) + (1 - α)·T(x(u, v), y(u, v), z(u, v)), where I(t, u, v) represents the color value of the image at pixel coordinates (u, v) at time t, T(x(u, v), y(u, v), z(u, v)) represents the texture value corresponding to the coordinates (x(u, v), y(u, v), z(u, v)) in the 3D model, this value corresponds to the mapping of the pixel position (u, v) in the image in 3D space, α represents the fusion weight coefficient, controlling the weighted ratio of the original image and texture information, and the value range is 0 ≤ α ≤ 1.
[0099] Introduce a geometric consistency optimization mechanism, compare the geometric information at the same position in the previous and subsequent frames for optimization, and assume that the same pixel position at time t and t - 1 is (u, v):
[0100] L geo = ∑ (u,v) |I fused (t, u, v) - I fused (t - 1, u, v)| 2 , where L geo represents the geometric consistency loss function, that is, the color difference of the same pixel at time t and t - 1, I fused (t, u, v) represents the color value of the fused image at pixel (u, v) at time t, I fused (t - 1, u, v) represents the color value of the fused image at pixel (u, v) at time t - 1.
[0101] Specifically, pixel - by - pixel weighted fusion of the generated image with the texture information of the 3D model can effectively combine the generated geometric - aware image with the existing 3D texture data, making the texture information of the generated image in 3D space match the actual scene. The geometric consistency optimization mechanism compares the geometric information at the same position in adjacent frames, maximizes the geometric consistency between frames, and reduces the probability of geometric distortion or frame skipping.
[0102] In step S3, the weighted fusion and geometric consistency optimization methods also include:
[0103] Introduce light simulation to calculate the final light intensity of the pixel point:
[0104] I final (t, u, v) = I fused (t, u, v)·(k a ·L a + k d ·(n·l)+ k s ·(r·v) m )), where I final(t, u, v) represents the final light intensity at pixel (u, v) at time t, and k a represents the ambient light coefficient, and L a represents the ambient light intensity, and k d represents the diffuse reflection coefficient, n represents the normal vector of the object surface, l represents the incident light direction vector, and k s represents the specular reflection coefficient, r represents the reflected light direction vector, r = 2(n·l)n - l, v represents the camera view direction vector, m represents the specular highlight exponent, controls the sharpness of the specular highlight. By combining ambient light, diffuse reflection, and specular reflection through real-time lighting simulation, the final light intensity of each pixel is calculated, thereby enhancing the authenticity and visual effect of the image.
[0105] Specifically, ambient light, diffuse reflection, and specular reflection are introduced for simulation, significantly improving the lighting effect of the generated image and enhancing the realism of the final rendered image. Combining the changes in the normal vector of the object surface and the reflected light, a more realistic lighting effect is generated, adding more visual layers to the image and making the light and shadow changes in the dynamic scene more natural and realistic.
[0106] Step S4, final output and real-time rendering.
[0107] After the geometric consistency fusion in step S3, an optimized image sequence is output and real-time rendering is performed.
[0108] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for fusing dynamic video and three-dimensional model based on image generation, characterized in that: include, Step S1, introducing a 3D geometry perception generation network, using NeRF generation network, processing 3D geometry data in dynamic video, and generating geometry perception images; Step S2, embedding geometric information of the time series, based on the geometric perception image generated in step S1, introducing the geometric change data of the time series, and generating a time series geometric perception image; In step S2, the geometric information of the time series is embedded as follows: Based on the geometric perception image generated in step S1, the geometric change data of the object in the time series is introduced, including the object's motion trajectory, rotation angle and viewing angle change, and the three-dimensional position of the object at time t is set to x(t)∈R 3 , the rotation matrix of the object is R(t)∈SO(3), then the change of the position and direction of the object over time is expressed as: Among them, x(t) represents the three-dimensional position of the object at time t, x0 represents the initial position of the object, that is, the position at time t=0, v(t) represents the instantaneous velocity vector of the object, that is, the motion trajectory of the object at each time point, R(t) represents the rotation matrix of the object at time t, that is, the rotation state of the object, R0 represents the initial rotation matrix of the object, ω(t) represents the instantaneous angular velocity vector of the object, that is, the rotation change of the object, The geometric perception image of each frame is constructed using the position and rotation data x(t) and R(t) of the object in the time series. Specifically, the geometric perception image I(t) at time t is generated based on the geometric information transformation of the previous frame I(t-Δt): Among them, I(t) represents the geometric perception image at time t, represents the geometric transformation function, which is responsible for transforming the previous frame image I(t-Δt) according to the displacement x(t) and rotation matrix R(t) of the object. I(t-Δt) represents the geometric perception image of the previous frame, and Δt represents the time interval between two frames. Other ways to embed geometric information of time series include: The light in the scene is recalculated based on the time series motion data of the object, and the path r(t) of the light is dynamically adjusted according to the displacement and rotation of the object: r(t)=o(t)+t·R(t)·d(t), where r(t) represents the light path at time t, o(t) represents the dynamic position of the camera's optical center at time t, and d(t) represents the camera's viewing direction at time t, which is updated based on the object's motion and rotation. After combining the object's motion trajectory and the change in viewing angle, each frame of geometric perception image I(t) at time t is calculated by the light integral generated by the NeRF network. The light color C(t) generated by NeRF depends on the geometric information of the time series, and the formula is: Where C(t) represents the color corresponding to the light ray r(t) at time t, based on the position and rotation state of the objects in the scene, T i (t) represents the cumulative transmittance of the i-th sampling point, defined as Among them, σ j (t) and δ j (t) represent the density and distance of the sampling points, respectively, and change with time, σ i (t) represents the volume density of the object, which is updated over time, δ i (t) represents the distance between adjacent sampling points, defined as t i+1 (t)-t i (t), c i (t) represents the color value of the sampling point at time t; Step S3, weighted fusion and geometric consistency optimization, based on the time series geometric perception image generated in step S2, a weighted fusion algorithm based on a three-dimensional model is used to fuse the image of each frame with the texture information of the three-dimensional model pixel by pixel; Step S4, final output and real-time rendering, after the geometric consistency fusion in step S3, output the optimized image sequence and perform real-time rendering.
2. The method for fusing dynamic video and three-dimensional model based on image generation according to claim 1, characterized in that: In step S1, the 3D geometric data includes 3D geometric coordinates, object shape and viewing angle changes. NeRF generates an image matching the 3D model, namely, a geometrically perceived image, based on the 3D geometric coordinates, object shape and viewing angle changes in the scene.
3. The method for fusing dynamic video and three-dimensional model based on image generation according to claim 2, characterized in that: In step S1, NeRF generates a network and processes 3D geometric data to generate a geometrically perceived image. Specifically: For each frame in the dynamic video, a geometric perception image is generated using the input of three-dimensional geometric data, which includes three-dimensional coordinates, perspective changes, and object shape information. The three-dimensional coordinate is defined as x∈R 3 , the viewing direction is d∈R 3 , NeRF generates color c∈R according to the coordinates and viewing angle of the point in the scene 3 And the volume density σ∈R: F θ (x, d) = (c, σ), where x represents the three-dimensional geometric coordinates of a sampling point in the scene, d represents the viewing direction from the camera to the sampling point, i.e., a vector, c represents the color value of the point at the viewing angle, and the output is a color vector of RGB three channels, and σ represents the volume density of the point, i.e., the opacity of the point in the scene; NeRF samples the light along the ray r(t) emitted from the camera to each pixel in the scene, and solves the weighted sum of the color and opacity of the light as it passes through the scene. NeRF samples each ray segment by segment to generate a series of sampling points x i = r(t i ), and use the network F for each sampling point θ Calculate color c i and volume density σ i , Perform weighted accumulation on the sampling points on each ray.
4. The method for fusing dynamic video and three-dimensional model based on image generation according to claim 3, characterized in that: During the training process of step S1, NeRF optimizes the network parameters θ based on the mean square error MSE loss between real images and synthetic images: Among them, L represents the total loss function, which measures the difference between the synthetic image and the real image. represents the light set, that is, all the light in the image, C G (r) represents the color value of light r in the real image, C(r) represents the color value corresponding to the light generated by the NeRF network, Each ray r outputs its corresponding color C(r).
5. The method for fusing dynamic video and three-dimensional model based on image generation according to claim 4, characterized in that: In step S2, the method of introducing the geometric change data of the time series is: Collect the object's motion trajectory, rotation angle and perspective change information, and dynamically embed the geometric changes in the time series.
6. The method for fusing dynamic video and three-dimensional model based on image generation according to claim 5, characterized in that: In step S3, a geometric consistency optimization mechanism is introduced and real-time lighting simulation is performed. In step S3, the fusion process uses GPU acceleration and parallel computing technology.
7. The method for fusing dynamic video and three-dimensional model based on image generation according to claim 6, characterized in that: In step S3, the weighted fusion and geometric consistency optimization method is: Perform weighted fusion processing on each frame image I(t) and combine it with the texture information of the 3D model. Assuming that the texture information of the 3D model is T(x, y, z), the texture value of each point in the 3D space is given, and the mapping between each pixel in the image and the 3D space point is: I fused (t,u,v)=α·I(t,u,v)+(1-α)·T(x(u,v),y(u,v),z(u,v)), where I(t,u,v) represents the color value of the image at the pixel coordinate (u,v) at time t, T(x(u,v),y(u,v),z(u,v)) represents the texture value corresponding to the coordinate (x(u,v),y(u,v),z(u,v)) in the three-dimensional model, which corresponds to the mapping of the pixel position (u,v) in the image in the three-dimensional space, and α represents the fusion weight coefficient, which controls the weighted proportion of the original image and texture information, and its value range is 0≤α≤1. A geometric consistency optimization mechanism is introduced to compare the geometric information of the same position in the previous and next frames for optimization. Suppose the same pixel position at time t and time t-1 is (u, v): L geo =∑ (u,v) |I fused (t,u,v)-I fused (t-1,u,v)| 2 , where L geo represents the geometric consistency loss function, that is, the color difference of the same pixel at time t and t-1, I fused (t,u,v) represents the color value of the pixel (u,v) in the fused image at time t, I fused (t-1,u,v) represents the color value of the pixel (u,v) in the fused image at time t-1.
8. The method for fusing dynamic video and three-dimensional model based on image generation according to claim 7, characterized in that: In step S3, the weighted fusion and geometric consistency optimization method further includes: Introduce lighting simulation to calculate the final lighting intensity of the pixel: I final (t,u,v)=I fused (t,u,v)·(k a ·L a +k d ·(n·l)+k s ·(r·v) m ), where I final (t,u,v) represents the final light intensity at pixel (u,v) at time t, k a Indicates the ambient light coefficient, L a Indicates the ambient light intensity, k d represents the diffuse reflection coefficient, n represents the normal vector of the object surface, l represents the direction vector of the incident light, k s represents the mirror reflection coefficient, r represents the direction vector of the reflected light, r=2(n·l)nl, v represents the direction vector of the camera viewing angle, and m represents the highlight index, which controls the highlight sharpness.
Citation Information
Patent Citations
Dynamic three-dimensional human body rendering synthesis method based on Hash coding of intrinsic coordinates
CN116109757A
Monocular video dynamic human body three-dimensional reconstruction method based on attitude optimization
CN116681838A
Cited By
User motion video image enhancement method and system based on machine vision
CN120635477A
A user motion video image enhancement method and system based on machine vision
CN120635477B