Video dynamic target three-dimensional reconstruction method based on artificial intelligence
Patent Information
- Application Number
- CN202510178063.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing three-dimensional reconstruction methods face dynamic targets with large depth estimation errors, target occlusion and rapid motion lead to loss of depth information, and lack of adaptability to multi-objective types.
The three-dimensional reconstruction method of dynamic targets based on artificial intelligence is adopted to achieve high-precision three-dimensional reconstruction of dynamic targets through the combination of image optical flow calculation, weighted processing, deep learning optimization, timing regression and generative adversarial network.
The accuracy of dynamic target depth estimation is improved, the depth information loss caused by occlusion and rapid motion is supplemented, the system's adaptability to different target types is enhanced, and efficient and accurate three-dimensional modeling is achieved.
Smart Images

Figure CN120107474A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and in particular to a method for three-dimensional reconstruction of dynamic objects in videos based on artificial intelligence. Background Art
[0002] With the rapid development of computer vision and artificial intelligence technology, video-based 3D reconstruction of dynamic targets has gradually become an important research direction in the fields of computer graphics, virtual reality and augmented reality. Video dynamic target 3D reconstruction technology is usually used to convert targets in dynamic scenes from 2D images into 3D models, which are then used in animation production, object interaction in virtual environments, and robot path planning.
[0003] Existing 3D reconstruction methods usually face the following challenges when facing dynamic targets:
[0004] Depth estimation error: Traditional depth estimation methods are mostly based on static scenes and assume that the target moves slowly. Therefore, the accuracy of depth estimation is reduced in the case of fast-moving dynamic targets and complex backgrounds. Due to the changes in the speed and movement angle of dynamic targets, the accuracy of optical flow calculation is easily affected, resulting in errors and incompleteness of depth information during 3D reconstruction.
[0005] Target occlusion and rapid motion problems: In practical applications, the target may lose depth information due to relative motion and occlusion by other objects in the scene. The lost depth information will affect subsequent 3D modeling, resulting in incomplete target models and reconstruction errors.
[0006] Lack of adaptability to multiple target types: Most existing technologies are optimized for specific types of dynamic targets and cannot handle the 3D reconstruction tasks of different types of targets. In complex dynamic scenes, how to quickly and accurately perform 3D modeling of different target types has become a technical challenge.
[0007] In order to overcome the above technical bottlenecks, technicians in this field provide a method for three-dimensional reconstruction of dynamic targets in videos based on artificial intelligence to solve the above problems. Summary of the invention
[0008] In view of the deficiencies in the prior art, the present invention provides a method for three-dimensional reconstruction of dynamic objects in videos based on artificial intelligence to solve the problems raised in the above-mentioned background technology.
[0009] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for three-dimensional reconstruction of dynamic objects in video based on artificial intelligence, comprising:
[0010] Step 1: Image optical flow calculation: Obtain each frame of the video, calculate the optical flow value of the pixel in each frame by extracting the spatial gradient of each pixel in the image and the motion information of the target, and obtain the displacement and motion trajectory of the target between video frames by analyzing the optical flow value;
[0011] Step 2: Optical flow weighted processing: Based on the optical flow calculation, the optical flow value of each pixel is weighted. The weighted optical flow information can reflect the motion characteristics of dynamic targets, which is convenient for depth estimation in subsequent steps.
[0012] Step 3: Depth estimation and optimization: Using the calculated weighted optical flow information, the depth of each pixel in the video frame is estimated through a deep learning model. The depth estimation result is optimized according to the target motion information in the input image. This process is based on the Bayesian optimization method and integrates the image observation data with the prior depth information.
[0013] Step 4, 3D reconstruction mapping: Based on the optimized depth estimation, the camera's internal and external parameters and perspective projection model are used to map the 2D coordinates in each frame image to the 3D space, and each pixel of the dynamic target is converted into the corresponding 3D coordinates to provide spatial data for the 3D reconstruction of the target;
[0014] Step 5: Time series regression and depth information supplementation: When performing 3D reconstruction, the motion changes of dynamic targets are taken into account and the depth information of the target is predicted using the time series regression method. When the target moves quickly or is occluded, the depth changes of consecutive frames are modeled and combined with the generative adversarial network technology to supplement the missing depth area;
[0015] Step 6: 3D model repair and optimization: After generating the preliminary 3D model, use graph neural network to repair the model details and optimize the incomplete areas caused by occlusion and depth estimation errors;
[0016] Step 7: 3D model rendering and output: Render the repaired 3D model using a combination of texture mapping and lighting rendering techniques. Through the rendering process, a realistic 3D model is provided for virtual reality and animation applications in computer graphics.
[0017] Preferably, the image optical flow calculation in step 1 includes using the spatial gradient And the target motion vector v(x, y, t), calculate the optical flow value of the pixel in each frame image, and solve it through the following optical flow equation:
[0018]
[0019] Where I(x, y, t) represents the pixel intensity of the image at position (x, y) and time t. is the spatial gradient of the image at position (x, y) and time t, and v(x, y, t) is the displacement vector of the pixel.
[0020] Preferably, the optical flow weighting processing in step 2 includes dynamically adjusting the weighting factor W(x, y, t) based on the target's motion speed and angle to improve the accuracy of weighted optical flow calculation in the dynamic target area. The weighted optical flow equation is:
[0021]
[0022] Among them, W(x, y, t) is the weighting factor, I(x, y, t) represents the pixel intensity of the image at position (x, y) and time t, is the spatial gradient of the image at position (x, y) and time t, and v(x, y, t) is the displacement vector of the pixel.
[0023] Preferably, the depth estimation and optimization in step 3 includes processing weighted optical flow information using a convolutional neural network to extract spatial features of dynamic targets in the video, and optimizing the depth estimation based on the Bayesian optimization method by the following formula:
[0024]
[0025] Where p(D|I) is the posterior probability of the depth estimate D given the image data I, p(I|D) is the likelihood of the image data I given the depth D, p(D) is the prior distribution of the depth, and p(I) is the marginal probability of the image data.
[0026] Preferably, the three-dimensional reconstruction mapping in step 4 includes mapping the two-dimensional coordinates to the three-dimensional space using the following perspective projection formula based on the optimized depth estimation through the internal and external parameters of the camera and the perspective projection model:
[0027]
[0028] Among them, Z is the depth of the target, f is the focal length of the camera, and x ′ is the coordinate of the target in the image, x 0 is the initial coordinate of the target, X is the three-dimensional coordinate of the target, X 0 Represents the original position in world coordinates.
[0029] Preferably, the temporal regression and depth information supplementation in step 5 include modeling the depth change in the video frame through a long short-term memory network, predicting the depth information of the next frame according to the motion state of the target, and the video dynamic target 3D reconstruction method is combined with a generative adversarial network to supplement the missing depth area through the following formula:
[0030]
[0031] Among them, G is the generator, D is the discriminator, x is the real image data, z is the noise input to the generator, E represents the mathematical expectation calculation of the data distribution, x~p data (x) represents the real data distribution p data (x) is the real image data x sampled, D(x) represents the probability that the real image data is the real data, z~p 2 (z) represents the prior distribution p 2 (z) is the noise z sampled in, G(z) is the forged data generated by the generator G with the noise z as input, and D(G(z)) is the output of the discriminator D on the data G(z) generated by the generator G.
[0032] Preferably, the three-dimensional model repair and optimization in step 6 includes optimizing the generated three-dimensional mesh model using a graph neural network, especially performing surface repair on incomplete areas caused by occlusion and depth estimation errors, wherein the graph neural network optimizes the three-dimensional structure of the target by the following formula:
[0033] y=∑ j∈N(i) A ij ·x j ,
[0034] Among them, y is the output feature of the target node, x j is the feature of the adjacent nodes, N(i) represents the set of adjacent nodes of node i, A ij Represents the connection relationship between node i and node j.
[0035] Preferably, the 3D model rendering and output in step 7 includes performing final rendering of the 3D model in combination with texture mapping and lighting rendering technology. During the rendering process, the 3D model is realistically rendered based on the camera light source and the lighting conditions in the scene, and specifically the following physical model is used for lighting calculation:
[0036] I=I ambient +I diffuse +I specular ,
[0037] Where I represents the total illumination received by the target surface, I ambient is the ambient light, I diffuse is diffuse lighting, I specular Specular lighting.
[0038] Preferably, the three-dimensional model rendering and outputting step includes rendering the surface features of the three-dimensional target in detail in combination with shadow and reflection processing, wherein the shadow and reflection processing is based on the following formula:
[0039] Itotal =I direct +α·I reflected ,
[0040] Among them, I total is the final lighting result, I direct is the illumination generated by direct light source, α is the reflection coefficient, I reflected For reflected light.
[0041] Preferably, the video dynamic target 3D reconstruction method can be adaptively optimized in different target types and complex environments through transfer learning, so that different types of dynamic targets can be quickly and effectively reconstructed in 3D. This process is achieved by introducing a multi-task learning framework and using the following loss function for optimization:
[0042] L total =∑ t (L depth (t)+λ 1 L texture (t)+λ 2 L appearance (t)),
[0043] Among them, L total is the total loss function, L depth (t) represents the depth estimation loss, L texture (t) represents texture loss, L appearance (t) represents the appearance loss, λ 1 and λ 2 is the regularization parameter and t is the time step.
[0044] The present invention provides a method for 3D reconstruction of dynamic objects in video based on artificial intelligence.
[0045] Beneficial effects:
[0046] 1. The present invention introduces weighted processing in image optical flow calculation and adjusts the weighting factor according to the movement speed and angle of the dynamic target, thereby improving the accuracy of depth estimation of dynamic targets in fast motion and complex backgrounds, obtaining accurate depth data, providing depth information for subsequent three-dimensional reconstruction, and solving the common error problem in depth estimation of dynamic targets.
[0047] 2. The present invention combines the time series regression method with the generative adversarial network to effectively predict the depth changes of the target, and supplement the depth loss caused by occlusion and rapid motion, thereby supplementing and repairing the target depth information, obtaining a complete and accurate three-dimensional model, and solving the problem of depth information loss caused by rapid motion and occlusion of dynamic targets.
[0048] 3. Through transfer learning and multi-task learning framework, the present invention can perform adaptive adjustments according to different types of dynamic targets, achieve rapid adaptation to 3D modeling tasks of different target types, obtain an efficient and universal 3D reconstruction method, enhance the adaptability and generalization ability of the system, and cope with the 3D modeling needs of various complex dynamic targets. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION
[0050] In order to make the technical personnel in the technical field understand the scheme of the present invention, the technical scheme in the embodiment of the present invention will be clearly and completely described below in combination with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a partial embodiment of the present invention, not a complete embodiment. Based on the embodiment of the present invention, other embodiments obtained by ordinary technicians in the field without creative work should fall within the scope of protection of the present invention.
[0051] The present invention is described in detail below in conjunction with the accompanying drawings:
[0052] Example:
[0053] Please refer to the attached Figure 1 The embodiment of the present invention provides a method for 3D reconstruction of dynamic objects in video based on artificial intelligence, comprising:
[0054] Step 1: Image optical flow calculation: Obtain each frame of the video, calculate the optical flow value of the pixel in each frame by extracting the spatial gradient of each pixel in the image and the motion information of the target, and obtain the displacement and motion trajectory of the target between video frames by analyzing the optical flow value;
[0055] Step 2: Optical flow weighted processing: Based on the optical flow calculation, the optical flow value of each pixel is weighted. The weighted optical flow information can reflect the motion characteristics of dynamic targets, which is convenient for depth estimation in subsequent steps.
[0056] Step 3: Depth estimation and optimization: Using the calculated weighted optical flow information, the depth of each pixel in the video frame is estimated through a deep learning model. The depth estimation result is optimized according to the target motion information in the input image. This process is based on the Bayesian optimization method and integrates the image observation data with the prior depth information.
[0057] Step 4, 3D reconstruction mapping: Based on the optimized depth estimation, the camera's internal and external parameters and perspective projection model are used to map the 2D coordinates in each frame image to the 3D space, and each pixel of the dynamic target is converted into the corresponding 3D coordinates to provide spatial data for the 3D reconstruction of the target;
[0058] Step 5: Time series regression and depth information supplementation: When performing 3D reconstruction, the motion changes of dynamic targets are taken into account and the depth information of the target is predicted using the time series regression method. When the target moves quickly or is occluded, the depth changes of consecutive frames are modeled and combined with the generative adversarial network technology to supplement the missing depth area;
[0059] Step 6: 3D model repair and optimization: After generating the preliminary 3D model, use graph neural network to repair the model details and optimize the incomplete areas caused by occlusion and depth estimation errors;
[0060] Step 7: 3D model rendering and output: Render the repaired 3D model using a combination of texture mapping and lighting rendering technology. Through the rendering process, a realistic 3D model is provided for virtual reality and animation applications in computer graphics.
[0061] Benefits of step 1: By calculating the optical flow information between image frames, the target's motion trajectory and pixel displacement information can be accurately obtained. This provides motion features for subsequent depth estimation, so that the 3D reconstruction of dynamic targets can reflect the changes of the target in the time dimension.
[0062] Benefits of step 2: Introducing weighted processing based on optical flow calculation makes optical flow estimation accurate, especially in the case of fast target motion and complex background interference. The weighted optical flow information can highlight the motion characteristics of dynamic targets, reduce noise in the depth estimation process, and improve the stability of depth calculation.
[0063] Benefits of step 3: Use deep learning models to process optical flow information and combine it with Bayesian optimization methods to improve the accuracy of depth estimation. By fusing observation data and prior information, the error in depth estimation is reduced, and the authenticity and detail expression of the reconstructed model are improved.
[0064] Benefits of step 4: Through the perspective projection model, the pixels in the two-dimensional image are mapped to the three-dimensional space, providing basic data for three-dimensional reconstruction, ensuring the accurate calculation of the three-dimensional coordinates of the dynamic target, enabling the model to match the spatial information of the real world, and improving the geometric accuracy of the three-dimensional reconstruction.
[0065] Benefits of step 5: The depth changes of dynamic targets are modeled through the temporal regression method, which can predict the depth information of future frames. When the target moves quickly or is occluded, the missing depth data can be supplemented by generating an adversarial network, which can effectively improve the integrity of the 3D model.
[0066] Benefits of step 6: The 3D model is optimized through graph neural networks, especially the incomplete areas caused by occlusion and depth estimation errors are repaired, which effectively enhances the surface details of the 3D model, makes the reconstructed 3D model smooth and coherent, and reduces defects and errors in the reconstruction process.
[0067] Benefits of step 7: Combine texture mapping and lighting rendering techniques to perform final visualization of the optimized 3D model, so that the model presents a realistic appearance in computer graphics applications. The rendering process can simulate the real lighting environment and improve the realism of the 3D model.
[0068] In summary, the present invention achieves high-precision three-dimensional reconstruction of dynamic targets by combining optical flow calculation, deep learning, Bayesian optimization, temporal regression, generative adversarial networks, graph neural networks, and high-quality rendering technology.
[0069] The image optical flow calculation in step 1 includes using the spatial gradient And the target motion vector v(x, y, t), calculate the optical flow value of the pixel in each frame image, and solve it through the following optical flow equation:
[0070]
[0071] Where I(x, y, t) represents the pixel intensity of the image at position (x, y) and time t. is the spatial gradient of the image at position (x, y) and time t, and v(x, y, t) is the displacement vector of the pixel.
[0072] The optical flow weighting processing in step 2 includes dynamically adjusting the weighting factor W(x, y, t) based on the target's motion speed and angle to improve the accuracy of weighted optical flow calculation in the dynamic target area. The weighted optical flow equation is:
[0073]
[0074] Among them, W(x, y, t) is the weighting factor, I(x, y, t) represents the pixel intensity of the image at position (x, y) and time t, is the spatial gradient of the image at position (x, y) and time t, and v(x, y, t) is the displacement vector of the pixel.
[0075] The present invention accurately extracts the motion trajectory of dynamic targets by calculating the optical flow information between video frames, so that the motion state of the target can be accurately described in the subsequent depth estimation and three-dimensional reconstruction process. The optical flow calculation is based on the pixel-level spatial gradient and target motion vector, which can effectively capture the displacement information of the target between video frames and provide basic data for subsequent depth estimation. Especially in the case of high-speed moving targets and complex background environments, this method can ensure the stable extraction of motion information and provide input for depth calculation and three-dimensional mapping.
[0076] Based on the optical flow calculation, the present invention dynamically adjusts the weighting factor to make the optical flow estimation accurate under different motion states. Weighted optical flow processing can optimize the optical flow calculation according to the target's motion speed and angle. Especially in the case of fast motion speed and complex background, this method can reduce the error of optical flow estimation and improve the depth estimation accuracy of dynamic targets. Through this step, the optical flow distortion caused by motion blur and background interference can be effectively reduced, thereby improving the accuracy of subsequent depth calculation and providing input for three-dimensional reconstruction.
[0077] In summary, the present invention realizes accurate motion feature extraction of dynamic targets in video frames by combining optical flow calculation and weighted optical flow processing.
[0078] The depth estimation and optimization in step 3 includes using a convolutional neural network to process the weighted optical flow information, extracting the spatial features of dynamic targets in the video, and optimizing the depth estimation based on the Bayesian optimization method through the following formula:
[0079]
[0080] Where p(D|I) is the posterior probability of the depth estimate D given the image data I, p(I|D) is the likelihood of the image data I given the depth D, p(D) is the prior distribution of the depth, and p(I) is the marginal probability of the image data.
[0081] The present invention combines convolutional neural networks with Bayesian optimization methods to achieve efficient depth estimation of weighted optical flow information and improve the accuracy and robustness of the estimation results. Convolutional neural networks can automatically extract the spatial features of dynamic targets in videos, avoid information loss caused by traditional manual feature extraction methods, and make depth estimation accurate.
[0082] The Bayesian optimization method constructs the optimal posterior probability model of depth estimation by fusing image observation data with the prior distribution of depth information, so that the depth calculation can effectively adapt to the changes of targets under different motion states and lighting conditions. Especially in complex scenes, Bayesian optimization can dynamically adjust the depth estimation parameters, reduce calculation errors, and improve the accuracy of 3D modeling.
[0083] In summary, the present invention extracts the spatial features of dynamic targets through convolutional neural networks and optimizes depth estimation with Bayesian optimization methods to ensure the accuracy and stability of 3D reconstruction. Compared with traditional depth estimation methods, this solution can provide depth information in complex environments with fast movement, occlusion, and lighting changes of dynamic targets, providing input data for the generation of 3D models.
[0084] The 3D reconstruction mapping in step 4 includes mapping the 2D coordinates to the 3D space using the following perspective projection formula based on the optimized depth estimation through the camera's internal and external parameters and the perspective projection model:
[0085]
[0086] Among them, Z is the depth of the target, f is the focal length of the camera, and x ′ is the coordinate of the target in the image, x 0 is the initial coordinate of the target, X is the three-dimensional coordinate of the target, X 0 Represents the original position in world coordinates.
[0087] The present invention combines the internal and external parameters of the camera with the perspective projection model, based on the optimized depth estimation, to accurately map the two-dimensional image coordinates to the three-dimensional space, and realize the accurate three-dimensional reconstruction of dynamic targets. The perspective projection formula can establish the precise relationship between the image pixels and the three-dimensional world coordinates, so that the reconstructed target can maintain the real spatial proportion and position relationship in the three-dimensional space.
[0088] 1. Improve the geometric accuracy of 3D reconstruction: This method uses the perspective projection formula for depth mapping to ensure that the reconstructed 3D coordinates are consistent with the real world and reduce the geometric distortion caused by projection errors.
[0089] 2. Adapt to different perspective changes: The internal and external parameter correction of the camera can compensate for the perspective deviation caused by changes in camera position and angle, so that the three-dimensional structure of the dynamic target remains consistent between different frames, thereby improving the stability of modeling.
[0090] 3. Suitable for complex motion scenes: For high-speed moving targets, perspective projection combined with depth optimization can avoid the position drift of the target's three-dimensional coordinates due to motion blur and projection distortion, thereby improving the three-dimensional modeling quality of dynamic targets.
[0091] In summary, this invention combines the internal and external parameters of the camera with the perspective projection model and uses the optimized depth information to achieve accurate mapping of two-dimensional images to three-dimensional space, ensuring accurate restoration of the spatial structure of dynamic targets. Compared with traditional three-dimensional reconstruction methods, it can accurately adapt to complex moving targets and improve the stability and geometric accuracy of three-dimensional modeling in dynamic scenes.
[0092] The temporal regression and depth information supplementation in step 5 include modeling the depth changes in the video frame through the long short-term memory network, predicting the depth information of the next frame according to the motion state of the target, and combining the video dynamic target 3D reconstruction method with the generative adversarial network to supplement the missing depth area through the following formula:
[0093]
[0094] Among them, G is the generator, D is the discriminator, x is the real image data, z is the noise input to the generator, E represents the mathematical expectation calculation of the data distribution, x~pdata (x) represents the real data distribution p data (x) is the real image data x sampled, D(x) represents the probability that the real image data is the real data, z~p 2 (z) represents the prior distribution p 2 (z) is the noise z sampled in, G(z) is the forged data generated by the generator G with the noise z as input, and D(G(z)) is the output of the discriminator D on the data G(z) generated by the generator G.
[0095] The present invention models the depth changes in video frames through a long short-term memory network and combines a generative adversarial network to supplement the missing depth areas, thereby improving the integrity and accuracy of three-dimensional reconstruction of dynamic targets.
[0096] Improve the temporal consistency of depth estimation: The long short-term memory network can analyze the depth change trend between consecutive frames, predict the depth information of the next frame, make the depth estimation smooth in the temporal dimension, and reduce the depth discontinuity problem caused by rapid target motion.
[0097] Improve the integrity of the 3D model: By combining the long short-term memory network and the generative adversarial network, this method can predict the depth of future frames, effectively fill the depth loss caused by occlusion and illumination changes, avoid deformation, fracture and missing areas in the 3D modeling process, and improve the integrity of the reconstructed model.
[0098] In summary, the present invention uses a long short-term memory network to perform temporal modeling of depth changes, and combines a generative adversarial network to supplement the missing depth information, thereby ensuring the stability of depth estimation of dynamic targets in the temporal dimension.
[0099] The 3D model repair and optimization in step 6 includes optimizing the generated 3D mesh model using a graph neural network, especially repairing the surface of incomplete areas caused by occlusion and depth estimation errors. The graph neural network optimizes the target 3D structure through the following formula:
[0100] y=∑ j∈N(i) A ij ·x j ,
[0101] Among them, y is the output feature of the target node, x j is the feature of the adjacent nodes, N(i) represents the set of adjacent nodes of node i, A ij Represents the connection relationship between node i and node j.
[0102] The present invention optimizes the generated three-dimensional mesh model through a graph neural network, and especially performs surface repair on incomplete areas caused by occlusion and depth estimation errors, thereby improving the integrity and realism of the three-dimensional reconstruction.
[0103] 1. Graph neural networks can use the information of adjacent nodes to update features, making the local geometric structure of the target more coherent. Through the optimization formula, the features of the target node are updated by the adjacent nodes, so that the areas affected by occlusion and depth estimation errors can be automatically repaired to restore the integrity of the mesh.
[0104] 2. Traditional 3D reconstruction methods are prone to surface breaks and unevenness in areas with large occlusion and depth estimation errors. This invention uses the feature propagation capability of graph neural networks and optimizes the adjacency matrix to make the 3D surface structure smoother and eliminate surface discontinuities caused by noise and missing data.
[0105] 3. The feature update mechanism of the graph neural network can optimize the 3D structure at different scales, making it rich in details. Especially in the 3D reconstruction of complex surfaces, this method can effectively restore details and improve the visual quality of the 3D model.
[0106] In summary, the present invention optimizes the three-dimensional grid structure through graph neural network to solve the problems of model missing and surface fracture caused by occlusion and depth error.
[0107] The 3D model rendering and output in step 7 includes the final rendering of the 3D model by combining texture mapping and lighting rendering technology. During the rendering process, the 3D model is realistically rendered based on the camera light source and the lighting conditions in the scene. Specifically, the following physical model is used for lighting calculation:
[0108] I=I ambient +I diffuse +I specular ,
[0109] Where I represents the total illumination received by the target surface, I ambient is the ambient light, I diffuse is diffuse lighting, I specular Specular lighting.
[0110] The 3D model rendering and output steps include detailed rendering of the surface features of the 3D target in combination with shadow and reflection processing, which is based on the following formula:
[0111] I total =I direct +α·I reflected ,
[0112] Among them, I total is the final lighting result, I direct is the illumination generated by direct light source, α is the reflection coefficient, I reflected For reflected light.
[0113] The present invention combines texture mapping with lighting rendering technology to perform final visual optimization on the three-dimensional model to make it realistic, especially through lighting calculation, shadow and reflection processing, to ensure that the model maintains a high degree of realism under different environmental lighting conditions.
[0114] 1. Through the illumination calculation model, ensure that the surface of the three-dimensional target can correctly receive and reflect the changes in ambient light, diffuse light and specular light, so that the model can maintain consistent visual effects under different lighting conditions.
[0115] 2. By combining shadow and reflection processing, this method can generate dynamic shadows and light reflections on the three-dimensional model, so that the model can interact realistically with the ambient light source, thereby improving the three-dimensional sense and detail expression of the model.
[0116] 3. Control the lighting response of different materials through the reflection coefficient so that different surfaces can present the correct optical effects.
[0117] In summary, the present invention ensures the realism and adaptability of the three-dimensional model under different lighting conditions through the illumination calculation model, shadow and reflection processing. Compared with the traditional three-dimensional modeling method, this solution improves the illumination consistency and detail performance of the model by combining ambient light, diffuse light and specular light calculation, combined with shadow and reflection optimization.
[0118] The video dynamic target 3D reconstruction method can be adaptively optimized in different target types and complex environments through transfer learning, so that different types of dynamic targets can be quickly and effectively reconstructed in 3D. This process is achieved by introducing a multi-task learning framework and using the following loss function for optimization:
[0119] L total =∑ t (L depth (t)+λ 1 L texture (t)+λ 2 L appearance (t)),
[0120] Among them, L total is the total loss function, L depth (t) represents the depth estimation loss, L texture (t) represents texture loss, L appearance (t) represents the appearance loss, λ 1 and λ 2 is the regularization parameter and t is the time step.
[0121] The present invention uses transfer learning and multi-task learning frameworks to achieve adaptive optimization of different types of dynamic targets, significantly improving the efficiency and accuracy of 3D reconstruction of dynamic targets. By adopting this framework, the system can quickly and effectively adapt to different target types and complex environments, ensuring 3D modeling in various scenarios.
[0122] 1. The introduction of transfer learning enables the model to borrow existing knowledge and features, quickly adapt to different types of targets, reduce training time, and improve modeling efficiency.
[0123] 2. Through the multi-task learning framework, the system can simultaneously learn depth estimation, texture and appearance features, so that the feature sharing and collaborative optimization between tasks can accelerate the 3D reconstruction process. The collaboration between different tasks can reduce error propagation, improve reconstruction accuracy, and effectively handle the diversity of target types in complex scenes.
[0124] 3. This loss function optimizes depth estimation, texture loss and appearance loss comprehensively, so that the generation of the three-dimensional model depends on single depth information, and can integrate the visual features and appearance information of the target, thereby improving the expressiveness and stability of the model in complex environments.
[0125] In summary, the present invention uses transfer learning and a multi-task learning framework to enable the 3D reconstruction method to achieve adaptive optimization under different target types and complex environments, thereby improving the efficiency, accuracy, and adaptability of 3D reconstruction.
[0126] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for 3D reconstruction of dynamic objects in video based on artificial intelligence, characterized in that: include: Step 1: Image optical flow calculation: Obtain each frame of the video, calculate the optical flow value of the pixel in each frame by extracting the spatial gradient of each pixel in the image and the motion information of the target, and obtain the displacement and motion trajectory of the target between video frames by analyzing the optical flow value; Step 2: Optical flow weighted processing: Based on the optical flow calculation, the optical flow value of each pixel is weighted. The weighted optical flow information can reflect the motion characteristics of dynamic targets, which is convenient for depth estimation in subsequent steps. Step 3: Depth estimation and optimization: Using the calculated weighted optical flow information, the depth of each pixel in the video frame is estimated through a deep learning model. The depth estimation result is optimized according to the target motion information in the input image. This process is based on the Bayesian optimization method and integrates the image observation data with the prior depth information. Step 4, 3D reconstruction mapping: Based on the optimized depth estimation, the camera's internal and external parameters and perspective projection model are used to map the 2D coordinates in each frame image to the 3D space, and each pixel point of the dynamic target is converted into the corresponding 3D coordinate to provide spatial data for the 3D reconstruction of the target; Step 5: Time series regression and depth information supplementation: When performing 3D reconstruction, the motion changes of dynamic targets are taken into account and the depth information of the target is predicted using the time series regression method. When the target moves quickly or is occluded, the depth changes of consecutive frames are modeled and combined with the generative adversarial network technology to supplement the missing depth area; Step 6: 3D model repair and optimization: After generating the preliminary 3D model, use graph neural network to repair the model details and optimize the incomplete areas caused by occlusion and depth estimation errors; Step 7: 3D model rendering and output: Render the repaired 3D model using a combination of texture mapping and lighting rendering technology. Through the rendering process, a realistic 3D model is provided for virtual reality and animation applications in computer graphics.
2. The method for 3D reconstruction of dynamic objects in video based on artificial intelligence according to claim 1, characterized in that: The image optical flow calculation in step 1 includes using spatial gradient And the target motion vector v(x, y, t), calculate the optical flow value of the pixel in each frame image, and solve it through the following optical flow equation: Where I(x, y, t) represents the pixel intensity of the image at position (x, y) and time t. is the spatial gradient of the image at position (x, y) and time t, and v(x, y, t) is the displacement vector of the pixel.
3. The method for 3D reconstruction of dynamic objects in video based on artificial intelligence according to claim 2, characterized in that: The optical flow weighting processing in step 2 includes dynamically adjusting the weighting factor W(x, y, t) based on the target's motion speed and angle to improve the accuracy of weighted optical flow calculation in the dynamic target area. The weighted optical flow equation is: Among them, W(x, y, t) is the weighting factor, I(x, y, t) represents the pixel intensity of the image at position (x, y) and time t, is the spatial gradient of the image at position (x, y) and time t, and v(x, y, t) is the displacement vector of the pixel.
4. The method for 3D reconstruction of dynamic objects in video based on artificial intelligence according to claim 1, characterized in that: The depth estimation and optimization in step 3 includes processing the weighted optical flow information using a convolutional neural network, extracting the spatial features of dynamic targets in the video, and optimizing the depth estimation based on the Bayesian optimization method using the following formula: Where p(D|I) is the posterior probability of the depth estimate D given the image data I, p(I|D) is the likelihood of the image data I given the depth D, p(D) is the prior distribution of the depth, and p(I) is the marginal probability of the image data.
5. The method for 3D reconstruction of dynamic objects in video based on artificial intelligence according to claim 1, characterized in that: The three-dimensional reconstruction mapping in step 4 includes mapping the two-dimensional coordinates to the three-dimensional space using the following perspective projection formula based on the optimized depth estimation through the internal and external parameters of the camera and the perspective projection model: Among them, Z is the depth of the target, f is the focal length of the camera, and x ′ is the coordinate of the target in the image, x0 is the initial coordinate of the target, X is the three-dimensional coordinate of the target, and X0 represents the original position in the world coordinate system.
6. The method for 3D reconstruction of dynamic objects in video based on artificial intelligence according to claim 1, characterized in that: The temporal regression and depth information supplementation in step 5 include modeling the depth change in the video frame through the long short-term memory network, predicting the depth information of the next frame according to the motion state of the target, and the video dynamic target 3D reconstruction method combined with the generative adversarial network to supplement the missing depth area through the following formula: Among them, G is the generator, D is the discriminator, x is the real image data, z is the noise input to the generator, E represents the mathematical expectation calculation of the data distribution, x~p data (x) represents the real data distribution p data (x), D(x) represents the probability that the real image data is the real data, z~p2(z) represents the noise z sampled in the prior distribution p2(z), G(z) is the forged data generated by the generator G with the noise z as input, and D(G(z)) is the output of the discriminator D for discriminating the data G(z) generated by the generator G.
7. The method for 3D reconstruction of dynamic objects in video based on artificial intelligence according to claim 1, characterized in that: The three-dimensional model repair and optimization in step 6 includes optimizing the generated three-dimensional mesh model using a graph neural network, especially performing surface repair on incomplete areas caused by occlusion and depth estimation errors. The graph neural network optimizes the three-dimensional structure of the target using the following formula: and=∑ j∈N(i) TO ij ·x j , Among them, y is the output feature of the target node, x j is the feature of the adjacent nodes, N(i) represents the set of adjacent nodes of node i, A ij Represents the connection relationship between node i and node j.
8. The method for 3D reconstruction of dynamic objects in video based on artificial intelligence according to claim 1, characterized in that: The 3D model rendering and output in step 7 includes performing final rendering of the 3D model by combining texture mapping and lighting rendering technology. During the rendering process, the 3D model is realistically rendered based on the camera light source and the lighting conditions in the scene. Specifically, the following physical model is used for lighting calculation: I=I ambient +I diffuse +I specular , Where I represents the total illumination received by the target surface, I ambient is the ambient light, I diffuse is diffuse lighting, I specular Specular lighting.
9. The method for 3D reconstruction of dynamic objects in video based on artificial intelligence according to claim 8, characterized in that: The three-dimensional model rendering and output step includes carefully rendering the surface features of the three-dimensional target in combination with shadow and reflection processing, wherein the shadow and reflection processing is based on the following formula: I total =I direct +α·I reflected , Among them, I total is the final lighting result, I direct is the illumination generated by direct light source, α is the reflection coefficient, I reflected For reflected light.
10. The method for 3D reconstruction of dynamic objects in video based on artificial intelligence according to claim 1, characterized in that: The video dynamic target 3D reconstruction method can be adaptively optimized in different target types and complex environments through transfer learning, so that different types of dynamic targets can be quickly and effectively reconstructed in 3D. This process is achieved by introducing a multi-task learning framework and using the following loss function for optimization: L total =∑ t (L depth (t)+λ1L texture (t)+λ2L appearance (t)), Among them, L total is the total loss function, L depth (t) represents the depth estimation loss, L texture (t) represents texture loss, L appearance (t) represents the appearance loss, λ1 and λ2 are regularization parameters, and t is the time step.
Citation Information
Cited By
Multi-target visual identification method and system in dynamic scene
CN120580650A
A multi-target visual recognition method and system in a dynamic scene
CN120580650B
Dynamic video generation method and device, storage medium and electronic equipment
CN120634897A
Unmanned aerial vehicle dynamic target speed measurement method based on optical flow self-motion compensation
CN122066736A
An unmanned aerial vehicle dynamic target speed measurement method based on optical flow self-motion compensation
CN122066736B