Video fusion method based on live-action three-dimensional scene and electronic equipment
Through three-dimensional light field modeling based on Passon equation and dynamic target modeling of deep learning models, combined with Lambert reflection model and multi-resolution light and shadow optimization, the limitations of light field modeling and dynamic target processing in the existing technology are solved, and the video fusion effect with high precision and high real-time is achieved.
Patent Information
- Application Number
- CN202510152926.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has limitations in light field modeling, dynamic target processing and real-time rendering, and it is difficult to meet the needs of high precision and high real-time, especially in complex geometric scenarios.
The three-dimensional light field modeling technology based on Pasone equation and variational optimization is adopted, and dynamic target segmentation and three-dimensional modeling is combined with deep learning models. The lighting properties of the dynamic target are adjusted through the Lambertian reflection model, and real-time rendering is achieved through multi-resolution light and shadow optimization and GPU accelerated calculation.
It realizes accurate calculation of scene lighting distribution, high-precision integration of dynamic targets and three-dimensional scenes, ensuring light and shadow consistency and real-time performance, and is suitable for complex scenes and dynamic application fields.
Smart Images

Figure CN120088439A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of virtual reality technology, and specifically provides a video fusion method and an electronic device based on a real - scene three - dimensional scene. Background Art
[0002] With the rapid development of virtual reality, augmented reality, and three - dimensional modeling technologies, the fusion technology of dynamic video and real - scene three - dimensional scenes has been widely applied in fields such as film production and virtual tours. However, there are still many limitations in existing technologies in aspects such as light field modeling, dynamic target processing, and real - time rendering, making it difficult to meet the requirements of high precision and high real - time performance.
[0003] Most existing light field modeling methods adopt simplified light source models and cannot accurately simulate the lighting characteristics of multiple light sources, reflection, and occlusion in complex scenes. This deficiency leads to unrealistic lighting effects, especially in complex geometric scenes, lacking the ability to describe physically consistent light fields.
[0004] In terms of dynamic target modeling, traditional methods mostly rely on multi - view synchronous capture or simple segmentation techniques, making it difficult to balance accuracy and efficiency. Monocular depth estimation methods are vulnerable to light and background interference in fast - changing dynamic scenes, resulting in large errors in the three - dimensional modeling of targets and scene matching.
[0005] In addition, the lighting and shadow processing of dynamic targets mostly adopts static simulation or texture overlay methods, making it difficult to achieve deep interaction between targets and scenes. The unrealistic projection and occlusion relationships make the fusion effect have an obvious visual abruptness, making it difficult to meet the requirements of complex scenes.
[0006] In terms of real - time performance, the lighting calculations of complex scenes mostly rely on offline processing, with a large amount of calculation and high rendering latency. Real - time optimization methods often sacrifice accuracy and cannot meet the requirements of real - time performance and visual quality simultaneously. In response to the above problems, the present invention proposes a solution with high lighting accuracy and dynamic interactivity. Summary of the Invention
[0007] Aiming at the deficiencies of the existing technology, the present invention provides a video fusion method and an electronic device based on a real - scene three - dimensional scene, which solve the problem of high - precision fusion of dynamic video targets and real - scene three - dimensional scenes in terms of lighting consistency, geometric matching, and real - time performance.
[0008] To achieve the above objectives, the present invention is realized through the following technical solutions: A video fusion method based on a real - scene three - dimensional scene, including the following steps:
[0009] S1. Obtain the three - dimensional data of the real - scene through a data acquisition device, and generate a three - dimensional scene model using point cloud reconstruction or multi - view geometry methods;
[0010] S2. Optimize the 3D scene model, including mesh simplification, texture mapping, and light source information initialization, to obtain the spatial position and radiation intensity distribution of the scene light source;
[0011] S3. Construct a light distribution model for the 3D scene based on the Poisson equation, and optimize it through variational methods to obtain the light distribution function, where the scene light distribution is affected by the light source position, radiation intensity, and light diffusion law;
[0012] S4. Use a deep learning model to segment dynamic objects in the video frame, extract the object's contour, depth information, and motion trajectory, and map the object to the corresponding position in the 3D scene;
[0013] S5. Adjust the lighting properties of the dynamic object based on the Lambert reflection model to make the surface light and shadow characteristics of the object consistent with the light field of the 3D scene;
[0014] S6. Calculate the projection area of the dynamic object, and dynamically update the light and shadow changes in the scene according to the object's spatial position, including the determination of occlusion relationships and the correction of light attenuation;
[0015] S7. Perform coupled calculations on the light field of the 3D scene and the light and shadow of the dynamic object through an optimization model, and achieve seamless integration of the dynamic object and the 3D scene in real-time rendering.
[0016] Preferably, the Poisson equation is used to describe the light field distribution in 3D space. When establishing the scene light distribution model, it is solved based on the following optimization functional:
[0017] By constraining the light field distribution to satisfy the intensity distribution and spatial diffusion law of the scene light source, and using regularization constraints to optimize the smoothness of the light field.
[0018] Preferably, the dynamic object segmentation and depth estimation steps include:
[0019] Use a deep learning model to segment dynamic objects in the video frame to generate the foreground contour of the dynamic object;
[0020] Generate the spatial coordinates of the object in the 3D scene based on the object depth map, and perform the conversion from pixel coordinates to 3D coordinates through the camera intrinsic matrix.
[0021] Preferably, the light and shadow adjustment steps are implemented as follows:
[0022] Adjust the surface lighting of the dynamic object based on the Lambert reflection model so that the incident light intensity is affected by the global light field distribution;
[0023] Calculate the reflected light intensity according to the angle between the normal vector of the dynamic object and the light direction, and dynamically adjust the brightness value of the object surface.
[0024] Preferably, the calculation of the projection area of the dynamic target includes the following steps:
[0025] Determine the projection range of the target on the surface of the three-dimensional scene according to the target depth map and the motion trajectory;
[0026] Calculate the light and shadow changes of the target occlusion, and dynamically correct the light intensity of the occlusion area based on the scene light field distribution.
[0027] Preferably, the light and shadow coupling calculation between the dynamic target and the three-dimensional scene is realized in the following way:
[0028] By constructing an optimization objective function, constrain the surface illumination characteristics of the dynamic target to be consistent with the scene light field, and make the light and shadow changes in the projection area conform to the law of light attenuation;
[0029] Based on the motion trajectory of the dynamic target and the scene light source information, update the input parameters of the optimization objective function in real time.
[0030] Preferably, the real-time rendering includes the following steps:
[0031] Use the multi-resolution calculation method to partition the scene light field distribution, and preferentially calculate the light and shadow changes within the user's viewing angle range;
[0032] Adopt GPU acceleration to optimize the solution of the Poisson equation and the dynamic light and shadow adjustment process to improve the real-time performance.
[0033] An apparatus for video fusion based on a real-scene three-dimensional scene, comprising:
[0034] A data acquisition module, configured to acquire three-dimensional data of the real-scene and generate a three-dimensional scene model;
[0035] A lighting modeling module, configured to construct a scene lighting distribution model based on the Poisson equation and solve the lighting distribution function through variational optimization;
[0036] A video processing module, configured to segment, depth estimate, and three-dimensional model the dynamic target in the input video to obtain the contour, depth map, and motion trajectory of the dynamic target;
[0037] A light and shadow adjustment module, configured to adjust the light and shadow of the dynamic target based on the Lambert reflection model, and calculate its projection area and occlusion relationship;
[0038] A coupling calculation module, configured to jointly optimize the light field of the three-dimensional scene and the light and shadow of the dynamic target to ensure light and shadow consistency;
[0039] A rendering module, configured to perform real-time rendering on the fusion result based on multi-resolution calculation and GPU acceleration technology.
[0040] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned method is implemented.
[0041] Preferably, the present invention also provides a storage medium, in which a program is stored. After the program is loaded into the electronic device, the electronic device is enabled to execute the above-mentioned method.
[0042] The present invention provides a video fusion method and an electronic device based on a real-scene three-dimensional scene, having the following beneficial effects:
[0043] 1. The present invention adopts a three-dimensional light field modeling technology based on the Poisson equation and variational optimization, achieving the technical effect of accurately calculating the scene light distribution. Compared with the existing solutions that only use simple light source models or preset light fields, it solves the problems of uneven light distribution and lack of physical authenticity, and performs particularly well in multi-light source scenarios and complex geometric structures.
[0044] 2. The present invention combines deep learning object segmentation and monocular depth estimation methods to achieve high-precision three-dimensional modeling and motion trajectory extraction of dynamic targets in videos, achieving the technical effect of accurately corresponding dynamic targets to three-dimensional scene geometric information. Compared with traditional multi-viewpoint or manual annotation modeling methods, it solves the problems of low target modeling accuracy, cumbersome calculation, and inadaptability to dynamic changes.
[0045] 3. The present invention adopts a dynamic light and shadow adjustment technology. Through the dynamic correction of the Lambert reflection model and the projection area, it achieves the effect of physical consistency between the dynamic target light and shadow and the three-dimensional scene light field. Compared with the existing solutions that directly superimpose dynamic target textures, it solves the technical problems of abrupt, untrue light and shadow and insufficient interaction with scene light sources, and is particularly outstanding in projection and occlusion processing.
[0046] 4. The present invention realizes the real-time rendering of complex scenes and dynamic targets by introducing multi-resolution light and shadow optimization and GPU accelerated computing technologies, achieving the technical effect of combining high precision and high performance. Compared with the existing solutions that require a large amount of offline calculation, it solves the problems of poor real-time performance and difficulty in meeting the requirements of interactive applications, and is particularly suitable for dynamic scene application fields such as virtual reality and augmented reality. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic diagram of the step flow of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the specification of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] Embodiment:
[0050] Please refer to the attached Figure 1 , the present invention discloses a video fusion method and an electronic device based on a real - scene three - dimensional scene. This method solves the problems of illumination consistency, light and shadow interaction, and real - time rendering by modeling the real - scene three - dimensional scene and calculating the light field distribution, and dynamically fusing the targets in the dynamic video with the three - dimensional scene.
[0051] Modeling of the real - scene three - dimensional scene
[0052] As the first step of the present invention, this step aims to obtain high - precision three - dimensional geometric structures and their light source distribution information by collecting and modeling three - dimensional data of the real scene, providing reliable basic data for subsequent light distribution modeling. This step is closely related to the subsequent steps. By generating an optimized three - dimensional scene model, the subsequent light field distribution calculation and the fusion of video targets have high geometric consistency and spatial reference.
[0053] First, use a data acquisition device to scan or photograph the real scene.
[0054] Generally, data acquisition is the first step in constructing a three - dimensional scene model. In the present invention, one of the following two data acquisition methods is adopted, and the specific selection depends on the size, complexity of the scene, and application requirements:
[0055] Method 1: Laser Radar (LiDAR) scanning. Obtain the point cloud data of the real - scene through a LiDAR scanner. The three - dimensional coordinates of each point in the point cloud data are (x, y, z), where:
[0056] x, y, and z respectively represent the position coordinates of a certain point in the three - dimensional space in the point cloud;
[0057] The point cloud data has the characteristic of high precision and is suitable for indoor, building, or large - scale outdoor scene modeling.
[0058] Method 2: Multi - view photogrammetry. As an option, static images of the scene can be taken from multiple angles, and a three - dimensional model of the scene can be generated through image matching and three - dimensional reconstruction algorithms. Generally, the pixel coordinate points (u, v) of each image are converted into three - dimensional point cloud coordinates through the internal parameter matrix K of the camera:
[0059]
[0060] Wherein:
[0061] X, Y, Z: Three-dimensional spatial coordinates of the scene points after reconstruction;
[0062] u, v: Pixel coordinates in the scene image;
[0063] D(u, v): Depth value of the pixel;
[0064] K: Camera internal parameter matrix.
[0065] The multi-view photo measurement method is applicable to medium and small-sized scenes, and the accuracy of its results is mainly affected by the image resolution and the camera angle.
[0066] After the data acquisition is completed, it is necessary to process the point cloud or photo data to generate a 3D model. Specifically, in some embodiments, a point cloud to mesh conversion algorithm is used to perform 3D modeling on the point cloud data, or a 3D scene geometry model is generated by processing multi-view images through a photo reconstruction algorithm.
[0067] As an option, the point cloud to mesh algorithm usually adopts a triangulation method, such as Delaunay triangulation. Its main principle is:
[0068] Connect the points in the point cloud into non-overlapping triangular patches;
[0069] The vertex coordinates of the triangular surface are (x, y, z) in the point cloud data.
[0070] In another possible implementation, the multi-view photo reconstruction algorithm adopts a combination of feature point matching and depth reconstruction. Specifically, the three-dimensional coordinates of the scene points are calculated through the feature point matching relationship between images.
[0071] In some embodiments, there may be redundant points or geometric noise in the generated 3D model, so it is necessary to optimize the model. Specifically, it includes the following contents:
[0072] Mesh simplification: Generally, mesh simplification techniques achieve the lightweight of the model by reducing the number of polygons while retaining the geometric details of the model. The optimized mesh vertex coordinates are (x′, y′, z′), satisfying the following constraints:
[0073] min||(x, y, z)-(x′, y′, z′)||
[0074] Wherein:
[0075] (x, y, z): Original mesh vertex coordinates;
[0076] (x′, y′, z′): Optimized grid vertex coordinates.
[0077] Texture mapping: As an option, a texture image can be mapped to the surface of a 3D model. By mapping the vertices of each mesh patch to the pixel coordinates (u, v) of the corresponding image, the visualization of the model is enhanced.
[0078] Light source initialization: In the present invention, the radiation intensity distribution of the light source in the 3D scene provides input parameters for subsequent light field modeling. The specific content of light source initialization includes:
[0079] Determination of the light source position (x s , y s , z s );
[0080] Calculation of the distribution of the radiation intensity ρ(x, y, z).
[0081] Specifically, the radiation intensity of the light source can be calculated by measuring the emission power of the light source and combining it with a spatial diffusion model. The diffusion characteristics of the light source satisfy the following relationship:
[0082]
[0083] Where:
[0084] P: Emission power of the light source;
[0085] r: Distance from the light source to the scene point (x, y, x)
[0086] The 3D model processed through the above steps finally includes the following:
[0087] Optimized scene geometric model, described as a grid structure;
[0088] Scene light source position and radiation intensity distribution;
[0089] Mapping relationship between the texture image and the grid surface.
[0090] The 3D scene model is stored in a standard 3D file format (such as OBJ or PLY) for subsequent light field modeling and video object fusion
[0091] Modeling of the light field distribution in the 3D scene
[0092] Modeling of the scene light field distribution is one of the core steps of the present invention. After completing the modeling of the 3D scene, the accurate calculation of the light field distribution is of great significance for subsequent dynamic video object fusion and light and shadow interaction. In this step, a light field distribution model based on the Poisson equation is constructed, combined with the geometric structure of the scene and the light source parameters, to generate the illumination intensity distribution at each point in the scene, providing necessary input for subsequent dynamic light and shadow adjustment.
[0093] In general, the light field distribution is determined by the position of light sources, radiation intensity, and the laws of light diffusion and attenuation in the scene. In this embodiment, each link of light field distribution modeling is described in detail, including formula description and numerical solution methods, and various adaptive adjustment methods of the model are also provided.
[0094] The light distribution in a three-dimensional scene is determined by the position of light sources and radiation intensity. By physically modeling to calculate the light field in the scene, the core is to establish a light distribution model based on the Poisson equation. The light distribution is described by the following formula:
[0095]
[0096] Where:
[0097] L(x, y, z) represents the light intensity at point (x, y, z) in the three-dimensional scene, with the unit of W / m 2 ;
[0098] ρ(x, y, z) represents the radiation source intensity at point (x, y, z), with the unit of W / m 3
[0099] is the Laplace operator, and its specific definition is:
[0100]
[0101] In some embodiments, the light sources in the scene are modeled as point light sources or surface light sources. Generally, the radiation intensity distribution of a point light source can be expressed by the following formula:
[0102]
[0103] Where:
[0104] P represents the total power of the light source, with the unit of W;
[0105] r is the distance from the light source to the scene point (x, y, z), with the unit of m;
[0106] 4πr 2 is the diffusion area of light on the spherical surface.
[0107] When there is a surface light source in the scene, the surface light source can be discretized into a combination of multiple point light sources, and the total radiation intensity of the surface light source is obtained by superimposing the radiation intensities of each point light source.
[0108] To accurately calculate the light field distribution, a light field optimization model based on the variational method is constructed in this embodiment. The optimization functional of the light field distribution is defined as:
[0109]
[0110] Among them:
[0111] Ω is the domain of the three-dimensional scene, representing the geometric volume of the scene;
[0112] The first term Constrains the light field distribution to satisfy the Poisson equation;
[0113] The second term Is the regularization term, used to constrain the smoothness of the light field and avoid excessive changes in the illumination gradient;
[0114] λ is the regularization parameter, used to balance the weights of the light field distribution constraint and the smoothness constraint.
[0115] In a possible implementation, the smoothness regularization of the light field can be further optimized by adaptively adjusting the value of λ to improve the accuracy of light field calculation.
[0116] Generally, numerical solution of the light field distribution requires discretization of the three-dimensional scene. In this embodiment, a voxel grid is used to divide the three-dimensional space. The center point (x i , y j , z k ) of each voxel represents a discrete sampling point, and its illumination intensity L i,j,k Satisfies the following discretized Poisson equation:
[0117]
[0118] Among them:
[0119] L i+1,j,k And L i-1,j,k Are the illumination intensities of the adjacent voxels in the x direction at the point (x i , x j , z k );
[0120] L i,j+1,k And L i,j-1,k Are the illumination intensities of the adjacent voxels in the y direction at the point (x i , y j , z k );
[0121] L i,j,k+1 And L i,j,k-1 Are the illumination intensities of the adjacent voxels in the z direction at the point (x i , y j , z k );
[0122] ρ i,j,k Is at (x i , yj , z k Radiation source intensity of the point.
[0123] As a possible implementation, during the numerical solution process, the gradient descent method is used to optimize the optical field functional. The gradient update formula is:
[0124]
[0125] Where:
[0126] t represents the number of iterations;
[0127] η is the learning rate, used to control the step size of each update;
[0128] The expression of the gradient is:
[0129]
[0130] Where:
[0131] Respectively represent the square terms of the gradients in the x, y, and z directions.
[0132] In some embodiments, after the solution is completed, the optical field distribution result is stored in the form of a three-dimensional matrix, and each matrix element corresponds to the illumination intensity value L of the three-dimensional voxel i,j,k . As an option, the optical field distribution data can also be further processed, for example:
[0133] Perform multi-resolution storage on the result, and adjust the resolution of the optical field according to the refinement requirements of the user's concerned area;
[0134] Bind the optical field result to the three-dimensional model for real-time rendering and light and shadow calculation.
[0135] Segmentation and Modeling of Video Targets
[0136] After completing the modeling of the optical field distribution in the three-dimensional scene, it is necessary to segment and model the dynamic targets in the input video. The main purpose of this step is to associate the geometric information, motion trajectory, and illumination characteristics of the dynamic targets with the three-dimensional scene. By obtaining the three-dimensional coordinates and depth information of the targets and combining their dynamic changes in the scene, it provides a necessary basis for subsequent light and shadow adjustment and fusion.
[0137] Generally, the segmentation and modeling of video targets rely on deep learning techniques and camera projection models, and can adapt to various target types and different video scenes. In some embodiments, it can also be refined by combining the shape, texture, and light and shadow characteristics of the targets to improve the accuracy of the targets.
[0138] In this embodiment, a deep learning model is used to perform dynamic target segmentation on the input video frames. Generally, a segmentation model based on a convolutional neural network (such as MODNet or Deeplab) is selected. These models can efficiently extract the foreground regions of the targets in the video frames.
[0139] As a possible implementation, the segmentation result generates a binary contour map of the target. The value of each pixel in this contour map indicates whether the pixel belongs to the dynamic target. Let represent the segmentation result of the pixel point (u, v):
[0140]
[0141] Where:
[0142] u and v represent the horizontal and vertical coordinates of the pixel;
[0143] The value of is a binary result.
[0144] As an option, the segmentation result can be further refined. For example, by combining the motion features or texture features of the target, the edge regions are corrected to improve the segmentation accuracy.
[0145] In some embodiments, the depth information of the target is generated by a deep learning model. Specifically, an algorithm based on monocular depth estimation (such as MiDaS or DPT) is used to calculate the depth value D(u, v) for each pixel point in the video frame. This depth value reflects the distance between the target pixel and the camera, with the unit of meter (m).
[0146] Specifically, the depth estimation result generates a depth map D(u, v), where:
[0147] u and v are pixel coordinates;
[0148] D(u, v) is the depth value of this pixel.
[0149] As a possible implementation, the depth map can be smoothed through the post-processing stage of the model to remove noise. Gaussian filtering or bilateral filtering can be used in the smoothing process to enhance the stability of the depth map.
[0150] is to map the target pixel from the two-dimensional coordinate system to the three-dimensional scene coordinate system through the depth map and the internal parameter matrix of the camera. The mapping relationship is described by the following formula:
[0151]
[0152] Where:
[0153] X, Y, Z are the coordinates of the target point in the three-dimensional space respectively;
[0154] u and v are the pixel coordinates of the target point;
[0155] D(u, v) is the depth value of the target point;
[0156] K is the camera internal parameter matrix, defined as:
[0157]
[0158] Where:
[0159] f x and f y are the focal lengths of the camera;
[0160] c x and c y are the pixel coordinates of the camera optical center.
[0161] Generally, this three-dimensional coordinate conversion process calculates independently for each pixel point to generate a point cloud representation of the target in three-dimensional space
[0162] Extract the motion path of the target through the change trajectory of the target in consecutive video frames. The motion trajectory can be expressed as a set of time series data {(X t , Y t , Z t )}, where:
[0163] t is the time frame;
[0164] (X t , Y t , Z t ) are the three-dimensional space coordinates of the target at the t-th frame.
[0165] To improve the stability of the trajectory, in this embodiment, Kalman filtering is used to smooth the trajectory data. The state transition equation of Kalman filtering is:
[0166] x t = Fx t-1 + w t
[0167] Where:
[0168] x t = [X t , Y t , Z t T is the current state of the target;
[0169] F is the state transition matrix;
[0170] w t is the process noise.
[0171] This filtering method can reduce the mutation phenomenon in the trajectory and improve the predictability of the trajectory.
[0172] In addition to segmentation, depth, and trajectory information, the shape and texture of the target can also be modeled in this embodiment. Generally, the shape can be represented by generating a mesh through the triangulation of the target point cloud, and the texture is extracted from the original video frame and mapped onto the mesh surface.
[0173] Specifically, the triangular mesh of the target can be generated by the following method:
[0174] Using the Delaunay triangulation algorithm, connect adjacent points in the target point cloud into triangles;
[0175] Use the generated triangular mesh for dynamic light and shadow calculation and visual enhancement.
[0176] Dynamic light and shadow adjustment
[0177] After completing the segmentation and modeling of the video target, it is necessary to adjust the lighting attributes of the dynamic target to make it consistent with the global light field distribution of the three-dimensional scene. The main purpose of this step is to calculate the light and shadow characteristics of the dynamic target by physical light and shadow modeling, combined with the spatial position, shape, and surface normal of the target, and at the same time process its projection and occlusion relationships in the three-dimensional scene.
[0178] Generally, the light and shadow adjustment of the dynamic target depends on the physical laws of light, such as the Lambert reflection model and the diffusion law of projection light and shadow. In some embodiments, the texture characteristics of the target can also be introduced to enhance the authenticity of the light and shadow effect.
[0179] The light intensity on the target surface is calculated based on the Lambert reflection model. The Lambert reflection model assumes that the target surface is a uniform diffuse reflection surface, and the intensity I of the reflected light on its surface is related to the direction and intensity of the incident light and the direction of the surface normal vector. The specific expression is:
[0180] I = k d Lcosθ
[0181] Where:
[0182] I is the intensity of the reflected light on the target surface, with the unit of W / m 2 ;
[0183] k d is the diffuse reflection coefficient of the target surface, and its value range is 0 ≤ k d ≤ 1;
[0184] L is the intensity of the incident light at the point on the target surface in the scene light field distribution, with the unit of W / m 2 ;
[0185] cosθ represents the angle between the light direction and the normal vector of the target surface, and its calculation is as follows:
[0186] cosθ = n·l
[0187] Where:
[0188] n is the normal vector of the target surface point, a unitized three-dimensional vector;
[0189] l is the direction vector of the incident light, a unitized three-dimensional vector.
[0190] Specifically, for each surface point in the dynamic target, the reflected light intensity I needs to be calculated point by point, considering the dynamic changes of the scene light field distribution L and the normal vector n.
[0191] In a possible implementation, a specular reflection term can be further added to the calculation of the illumination intensity. At this time, the formula for the reflected light intensity I is extended to:
[0192] I = k d Lcosθ + k s L(cosφ) n
[0193] Where:
[0194] k s is the specular reflection coefficient;
[0195] cosφ represents the angle between the reflected light direction and the viewing direction;
[0196] n is the highlight index of specular reflection, controlling the distribution range of the highlight.
[0197] The projection area of the dynamic target is jointly determined by the depth information of the target and the three-dimensional scene geometry.
[0198] Specifically, the calculation steps of the target projection area include the following:
[0199] Generally, each surface point of the target will form a projection point on the scene surface. Assume the three-dimensional coordinates of the target point are (X, Y, Z), and the point on the scene surface is (x, y, z). According to the projection direction of the light, the position of the projection point corresponding to the target point on the scene surface can be calculated.
[0200] As an option, the depth comparison method can be used to determine whether the target occludes the scene surface. The depth relationship between the target point (X, Y, Z) and the scene point (x, y, z) satisfies:
[0201]
[0202] Where:
[0203] Dscene (x, y, z) is the depth of the scene point;
[0204] D target (X, Y, Z) is the depth of the target point.
[0205] For a determined projection area, its light intensity needs to be adjusted according to the light attenuation caused by occlusion. Specifically, the light intensity of the projection area is adjusted to:
[0206] L shadow = L - ΔL
[0207] Where:
[0208] L shadow is the light field intensity of the projection area;
[0209] L is the light field intensity when not occluded;
[0210] ΔL is the attenuation amount of light.
[0211] In some embodiments, a dynamic target may simultaneously occlude multiple scene objects. In this embodiment, a multi-layer occlusion detection algorithm is used to judge the occlusion relationship of the target layer by layer.
[0212] Specifically, the visibility of each scene point (x, y, z) is determined by the depth sorting of the target. The visibility calculation formula is:
[0213]
[0214] Where:
[0215] N is the number of all occluding targets in the scene;
[0216] Occlusion i (x, y, z) represents the occlusion relationship of the i-th target to the scene point (x, y, z).
[0217] To meet the requirements of real-time calculation, a multi-resolution processing method is introduced for dynamic light and shadow adjustment. Specifically, the light and shadow calculation of the target is divided into a high-resolution area and a low-resolution area.
[0218] Generally, the area closer to the user's perspective of the dynamic target is assigned for high-resolution calculation. The remaining areas are quickly estimated for light and shadow changes through low-resolution calculation. In addition, the matrix operations in light and shadow calculation are optimized for parallel calculation, and the GPU is used to improve the processing speed.
[0219] As a possible implementation method, a time-related light and shadow smoothing mechanism is also added in this embodiment. The purpose of light and shadow smoothing is to reduce the jitter of light and shadow changes of dynamic targets in consecutive frames. The specific formula for light and shadow smoothing is:
[0220]
[0221] Wherein:
[0222] is the smooth light field intensity of the current frame;
[0223] L t is the original light field intensity of the current frame;
[0224] L t-1 is the smooth light field intensity of the previous frame;
[0225] α is the smoothing coefficient, and its usual value range is 0 ≤ α ≤ 1.
[0226] Dynamic Light and Shadow Coupling and Real-Time Rendering
[0227] After the light and shadow adjustment of the dynamic target is completed, it is necessary to further couple the light and shadow relationship between the target and the scene. This step realizes the lighting consistency and dynamic interaction between the target and the scene through the joint optimization of the scene light field and the dynamic target light and shadow. In addition, to ensure that the system meets the real-time rendering requirements, this step adopts multi-resolution processing and parallel computing methods to optimize the coupling calculation and rendering process.
[0228] Generally, dynamic light and shadow coupling needs to consider the distribution of the scene global light field, the real-time motion information of the target, and the user's interaction operations at the same time. In some embodiments, the coupling calculation can also dynamically adjust the regional resolution to improve efficiency.
[0229] The coupling calculation of the scene light field and the dynamic target light and shadow is carried out by constructing a joint optimization objective function. The specific form of the optimization objective function is:
[0230]
[0231] Wherein:
[0232] L represents the global light field distribution of the scene;
[0233] I represents the surface illumination intensity of the dynamic target;
[0234] k d is the diffuse reflection coefficient of the target surface;
[0235] cosθ is the angle between the incident light and the surface normal, representing the surface light and shadow characteristics;
[0236] L shadow,j is the light field intensity of the projection area;
[0237] ΔL j is the light attenuation value, representing the light and shadow change caused by occlusion;
[0238] β is a weight parameter used to balance the light and shadow adjustment of the dynamic target and the light field consistency of the projection area.
[0239] Specifically, the first term of the optimization objective function constrains the light and shadow characteristics of the target surface to be consistent with the scene light field. The second term constrains the light field distribution in the projection area to match the light attenuation.
[0240] In a possible implementation, the projection and occlusion areas are dynamically updated in this embodiment. Generally, the projection area changes in real time with the movement of the target. The update process of the projection includes the following steps:
[0241] First, calculate the current spatial position of the target according to the movement trajectory {(X t , Y t , Z t )} of the target;
[0242] Then, calculate the projection range of the target on the scene surface through geometric projection;
[0243] Finally, update the light field distribution L shadow of the projection area, and the specific adjustment formula is:
[0244]
[0245] Where:
[0246] (x, y, z) are the three-dimensional coordinates of the point on the scene surface;
[0247] L(x, y, z) is the original light intensity of this point in the scene light field;
[0248] ΔL is the amount of light attenuation caused by the projection.
[0249] As an option, the boundary of the projection area can be smoothed to avoid sudden changes in light and shadow.
[0250] In this embodiment, a multi-resolution calculation method is introduced to optimize the area of the scene light field and the light and shadow calculation of the dynamic target. Specifically, the scene is divided into multiple resolution areas according to the user's perspective and the movement range of the dynamic target.
[0251] Generally, the area closer to the user's perspective is assigned as a high-resolution calculation area, while other areas use low-resolution to quickly estimate the light and shadow changes. The specific steps of the multi-resolution calculation include:
[0252] First, define the high-resolution calculation area according to the user's perspective V and the dynamic target position P;
[0253] Then, accurately calculate the light field and light and shadow adjustment within the high-resolution region;
[0254] Finally, quickly estimate the light and shadow changes within the low-resolution region through interpolation methods.
[0255] Specifically, bilinear interpolation algorithm can be adopted for light field interpolation, and its formula is:
[0256]
[0257] Where:
[0258] L′(x,y) is the light field intensity after interpolation;
[0259] L i,j 、L i+1,j 、L i,j+1 、L i+1,j+1 are the light field intensities of the four surrounding grid points respectively;
[0260] α x and α y are the interpolation weights in the x and y directions respectively.
[0261] To meet the requirements of real-time rendering, in this embodiment, the GPU parallelization method is adopted to accelerate the light and shadow coupling and rendering tasks. Specifically, the coupling calculation is split into multiple parallel tasks, and each task independently processes the light and shadow changes of a spatial grid.
[0262] Generally, the parallel processing of the GPU can significantly improve the calculation speed, especially when there are many dynamic objects or the scene is complex. As a possible implementation, the memory allocation strategy of the GPU can be further optimized to reduce the overhead of data transmission.
[0263] In a possible implementation, to avoid the discontinuity of light and shadow changes between consecutive frames, a light and shadow smoothing mechanism is introduced in this embodiment. The specific formula for light and shadow smoothing is:
[0264]
[0265] Where:
[0266] is the smoothed light field intensity of the current frame;
[0267] L t is the original light field intensity of the current frame;
[0268] L t-1 is the smoothed light field intensity of the previous frame;
[0269] α is the smoothing coefficient, and its usual value range is 0 ≤ α ≤ 1.
[0270] Generally, smooth light and shadow can effectively reduce the jitter of the light and shadow changes of dynamic objects and improve the visual coherence of rendering.
[0271] A device for video fusion based on a real-scene three-dimensional scene described below can be correspondingly referred to the method for video fusion based on a real-scene three-dimensional scene described above.
[0272] The device of this embodiment can be used to execute the above method embodiment, and its principle and technical effect are similar, so they will not be elaborated here.
[0273] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it executes the above method.
[0274] The present invention also provides a storage medium, in which a program is stored. After the program is loaded into the electronic device, the electronic device is enabled to execute the above method.
[0275] Among them, the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM for short), electrically erasable programmable read-only memory (EEPROM for short), erasable programmable read-only memory (EPROM for short), programmable read-only memory (PROM for short), read-only memory (ROM for short), magnetic memory, flash memory, magnetic disk or optical disc.
[0276] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A video fusion method based on real three-dimensional scene, characterized in that: The following steps are involved: S1. Acquire three-dimensional data of a real scene through a data acquisition device, and generate a three-dimensional scene model using point cloud reconstruction or multi-view geometry method; S2. Optimize the three-dimensional scene model, including mesh simplification, texture mapping and light source information initialization, and obtain the spatial position and radiation intensity distribution of the scene light source; S3. Construct a lighting distribution model for a three-dimensional scene based on the Passon equation, and obtain the lighting distribution function by optimizing the variational method, where the scene lighting distribution is affected by the position of the light source, the radiation intensity, and the diffusion law of light; S4. Use the deep learning model to segment the dynamic targets in the video frame, extract the contour, depth information and motion trajectory of the target, and map the target to the corresponding position in the three-dimensional scene; S5. Adjust the illumination properties of the dynamic target based on the Lambertian reflection model so that the surface light and shadow characteristics of the target are consistent with the light field of the three-dimensional scene; S6. Calculate the projection area of the dynamic target and dynamically update the light and shadow changes in the scene according to the spatial position of the target, including the determination of occlusion relationship and the correction of light attenuation; S7. The light field of the three-dimensional scene and the light and shadow of the dynamic target are coupled calculated by optimizing the model, and the seamless integration of the dynamic target and the three-dimensional scene is achieved in real-time rendering.
2. The video fusion method based on real three-dimensional scene according to claim 1, characterized in that: The Passon equation is used to describe the light field distribution in three-dimensional space. When establishing a scene illumination distribution model, it is solved based on the following optimization functional: The light field distribution is constrained to meet the intensity distribution and spatial diffusion law of the scene light source, and the smoothness of the light field is optimized by regularization constraints.
3. The video fusion method based on real three-dimensional scene according to claim 1, characterized in that: The dynamic target segmentation and depth estimation steps include: Use deep learning models to segment dynamic targets in video frames and generate foreground contours of dynamic targets; The spatial coordinates of the target in the three-dimensional scene are generated based on the target depth map, and the pixel coordinates are converted to three-dimensional coordinates through the camera intrinsic parameter matrix.
4. The video fusion method based on real three-dimensional scene according to claim 1, characterized in that: The light and shadow adjustment step is achieved by: The surface illumination of dynamic targets is adjusted based on the Lambertian reflection model so that the incident light intensity is affected by the global light field distribution. The reflected light intensity is calculated based on the angle between the normal vector of the dynamic target and the direction of the light, and the brightness value of the target surface is dynamically adjusted.
5. The video fusion method based on real three-dimensional scene according to claim 1, characterized in that: The projection area calculation of the dynamic target comprises the following steps: Determine the projection range of the target on the surface of the three-dimensional scene according to the target depth map and motion trajectory; Calculate the light and shadow changes of the target occlusion, and dynamically correct the light intensity of the occluded area based on the scene light field distribution.
6. The video fusion method based on real three-dimensional scene according to claim 1, characterized in that: The light and shadow coupling calculation of the dynamic target and the three-dimensional scene is achieved by the following method: By constructing an optimization objective function, the surface illumination characteristics of the dynamic target are constrained to be consistent with the scene light field, and the light and shadow changes in the projection area are made to conform to the law of light attenuation; The input parameters of the objective function are updated in real time based on the motion trajectory of the dynamic target and the scene light source information.
7. The video fusion method based on real three-dimensional scene according to claim 1, characterized in that: The real-time rendering comprises the following steps: The multi-resolution calculation method is used to partition the scene light field distribution, and the light and shadow changes within the user's viewing angle are calculated first; GPU acceleration is used to optimize the solution of the Passon equation and the dynamic light and shadow adjustment process to improve real-time performance.
8. A device for video fusion based on real-life three-dimensional scenes, characterized in that: The method for video fusion based on a real three-dimensional scene according to any one of claims 1 to 7 comprises: A data acquisition module is used to obtain three-dimensional data of the real scene and generate a three-dimensional scene model; The illumination modeling module is used to construct a scene illumination distribution model based on the Passon equation and solve the illumination distribution function through variational optimization; The video processing module is used to segment, estimate the depth and perform 3D modeling of dynamic targets in the input video, and obtain the contour, depth map and motion trajectory of the dynamic targets; The light and shadow adjustment module is used to adjust the light and shadow of the dynamic target based on the Lambertian reflection model and calculate its projection area and occlusion relationship; The coupled calculation module is used to jointly optimize the light field of the 3D scene and the light and shadow of the dynamic target to ensure the consistency of light and shadow; The rendering module is used to perform real-time rendering of the fusion results based on multi-resolution calculation and GPU acceleration technology.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A storage medium, characterized in that: The storage medium stores a program, and after the program is loaded into the electronic device, the electronic device executes the method according to any one of claims 1 to 7.
Citation Information
Cited By
Multi-light-source real-time role rendering system and method applied to animation game
CN122435108A