A multi-view three-dimensional virtual-real fusion rendering method and related device

By constructing the geometric and appearance fields of the four-dimensional spatiotemporal domain and using the Immersive Transformer module for multi-hop inference, the problems of viewpoint selection and rendering quality in dynamic scenes are solved, achieving efficient free-viewpoint 3D rendering and improving the quality and automation level of the rendered images.

CN121600228BActive Publication Date: 2026-04-21CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CENT SOUTH UNIV
Filing Date
2026-01-28
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies lack unified spatiotemporal flow field modeling in 3D reconstruction and rendering of dynamic scenes, lack advanced reasoning mechanisms for viewpoint selection and camera language, lack aesthetic and task-oriented joint optimization of rendering targets, and have fragmented technology stacks across multiple scenes, making it difficult to achieve automated selection of important viewpoints and high-quality rendering.

Method used

By acquiring multi-time frame, multi-view video sequences for 3D reconstruction, geometric field, appearance field, and spatiotemporal flow field in the four-dimensional spatiotemporal domain are constructed. The immersion Transformer module is used for multi-hop inference to calculate the spatiotemporal importance distribution. Based on this, the geometric field and appearance field are rendered. A combination of high-resolution and low-resolution rendering is used to generate free-view 3D rendering results.

Benefits of technology

Without increasing overall computational overhead, it improves rendering quality, solves the problem of fragmented multi-scene rendering, and enables automated selection of important viewpoints and high-quality rendering of dynamic scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121600228B_ABST
    Figure CN121600228B_ABST
Patent Text Reader

Abstract

The application provides a multi-view three-dimensional virtual-real fusion rendering method and related equipment, and relates to the field of computer vision. By constructing a geometric field, an appearance field and a space-time flow field on a four-dimensional space-time domain, a target scene is divided into a space-time flow line or a local space-time block, fragmentation caused by pure frame-by-frame rendering and scene-specific rules is avoided, a multi-scene fragmentation problem is solved, a space-time importance distribution under each time and each virtual view is calculated by using an updated feature vector and an attention weight corresponding to the updated feature vector to render the geometric field and the appearance field on the four-dimensional space-time domain, and the rendering picture quality is improved without significantly increasing the overall computing overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a multi-view 3D virtual-real fusion rendering method and related equipment. Background Technology

[0002] With the development of technologies such as multi-view camera arrays, 3D reconstruction, and neural rendering, 3D playback formats such as "free-viewpoint" and "volumetric video" have gradually emerged in many situations. A typical approach involves deploying multiple cameras around the site, reconstructing the 3D scene using multi-view geometry or methods based on Neural Radiance Fields (NeRF) or Gaussian Splatting, and then rendering and displaying it from any viewpoint using a virtual camera. However, existing technologies still have the following shortcomings in the rendering process:

[0003] 1. Lack of unified spatiotemporal flow field modeling for dynamic scenes: Existing free-viewpoint systems mostly focus on static geometry reconstruction and frame-by-frame rendering. For example, for dynamic targets such as players and balls in sports scenes, or pills and equipment in pharmaceutical production lines, they are usually only implicitly included in volume representation or feature fields, lacking explicit descriptions of "spatiotemporal flow field" or "motion trajectory network". This makes it difficult to gather spatiotemporally continuous perspectives to analyze overall actions, complete process flow and multiple causal links, and is also not conducive to subsequent perspective scheduling and key event tracking.

[0004] 2. Lack of advanced reasoning mechanisms in perspective selection and camera language: Existing sports free-view products and industrial virtual-real fusion systems rely heavily on manual scripts or simple heuristic rules for camera switching and perspective trajectory. While they can achieve "viewing from any angle," they struggle to automatically prioritize "views and time periods with tactical or educational value." For the traceability of defective pills in pharmaceutical production lines, existing systems typically lack a unified multi-hop reasoning and decision-making mechanism, making it impossible to automatically filter out truly important segments and perspectives in the spatiotemporal dimension.

[0005] 3. Rendering targets are mainly reconstruction errors, lacking joint optimization of aesthetics and tasks: Existing neural rendering systems mostly use indicators such as reconstruction error and frame rate as the main optimization targets. For key business indicators such as the visual tension and aesthetic effect of sports movements, and the recognizability of pill defects in the virtual-real fusion view, there is a lack of systematic mathematical modeling and joint optimization mechanisms. In practical applications, it is often necessary to rely on post-production manual editing, manual parameter adjustment and empirical rules to make up for it, which increases labor costs and weakens the intelligence level of the system.

[0006] 4. Fragmented technology stacks across multiple scenarios, with no closed loop between drug detection and virtual-real fusion rendering: In the pharmaceutical industry, defect detection and classification of tablets are usually completed by independent 2D machine vision systems, while 3D or virtual-real fusion visualization systems focus more on process display and scene restoration. There is a lack of a unified optimization framework between the detection model and the rendering system, making it difficult to fully utilize the stereoscopic effect and interactivity of virtual-real fusion to enhance quality inspection results. Moreover, different scenarios often use incompatible algorithms and systems, resulting in fragmented technology stacks, high costs of redundant system construction and maintenance, and difficulty in sharing underlying modeling and inference capabilities. Summary of the Invention

[0007] This invention provides a multi-view 3D virtual-real fusion rendering method and related equipment, the purpose of which is to improve the quality of rendered images while solving the problem of fragmentation in multiple scenes.

[0008] To achieve the above objectives, this invention provides a multi-view 3D virtual-real fusion rendering method, including:

[0009] Step 1: Acquire multi-time frame, multi-view video sequence in the target scene to obtain image data, and calibrate the internal and external parameters of the acquisition device to obtain calibration results;

[0010] Step 2: Based on the image data and calibration results, the target scene is reconstructed in three dimensions to obtain the geometric field, appearance field and spatiotemporal flow field in the four-dimensional spatiotemporal domain. The geometric field and appearance field are discretized through the spatiotemporal flow field to obtain the four-dimensional Gaussian primitive. The spatiotemporal flow field is used to describe the motion trajectory of the target in the target scene over time.

[0011] Step 3: Integrate the central trajectory in the four-dimensional Gaussian primitive to obtain the spacetime streamlines. Divide the spacetime streamlines and spacetime flow field into multiple spacetime markers, and extract features from each spacetime marker to obtain the feature vector corresponding to each spacetime marker.

[0012] Step 4: Input the feature vector corresponding to each spatiotemporal marker into the trained Immersive Transformer module for multi-hop inference to obtain the updated feature vector and the attention weights corresponding to the updated feature vector.

[0013] Step 5: Calculate the spatiotemporal importance distribution at each time point and from each virtual viewpoint based on the updated feature vector and the attention weights corresponding to the updated feature vector, and render the geometric field and appearance field in the four-dimensional spatiotemporal domain based on the spatiotemporal importance distribution to obtain the free-view 3D rendering result of the target scene.

[0014] Furthermore, spatiotemporal markers include position, velocity, curvature, texture descriptors, object category labels, and scene identifiers.

[0015] Furthermore, the expression for the feature vector corresponding to each spatiotemporal marker, obtained by feature extraction for each spatiotemporal marker, is as follows:

[0016] ;

[0017] in, Indicates the relationship with the first Feature vectors corresponding to spatiotemporal markers Represents the feature encoder. Indicates the first A spatiotemporal marker.

[0018] Furthermore, step 4 includes:

[0019] For each feature vector corresponding to a spatiotemporal marker, two sets of learnable projection matrices are introduced for each feature vector to obtain a query vector and a key vector;

[0020] Affinity strength is calculated based on query vector and key vector to obtain the affinity matrix;

[0021] The affinity matrix is ​​gated sparsity by a learnable threshold and temperature to obtain a gated matrix;

[0022] After fusing the affinity matrix and the gating matrix, Softmax is applied to obtain the attention weights corresponding to the updated feature vectors, and strong affinity subgraphs are induced based on the attention weights.

[0023] After constructing the graph Laplacian matrix on the strong affinity subgraph, it is coupled with all eigenvectors to obtain the updated eigenvectors.

[0024] Furthermore, the elements in the gating matrix are calculated using either soft or hard gates, as expressed by:

[0025] or ;

[0026] in, Represents the first in the gating matrix Line number Column elements, Indicates a soft door. Indicates a hard door. Represents spacetime markers and The level of friendliness, Indicates the learnable threshold. Indicates temperature.

[0027] Furthermore, after constructing the graph Laplacian matrix on the strong affinity subgraph and coupling it with all eigenvectors, the expression for the updated eigenvectors is obtained as follows:

[0028] ;

[0029] in, This represents the updated feature vector. Represents the eigenvector. Indicates diffusion intensity. This represents the graph Laplace matrix.

[0030] Furthermore, based on the spatiotemporal importance distribution, the geometric and appearance fields in the four-dimensional spatiotemporal domain are rendered to obtain the free-viewpoint 3D rendering results of the target scene, including:

[0031] Based on the spatiotemporal importance distribution, the geometric field and appearance field in the four-dimensional spatiotemporal domain are divided into regions, resulting in high importance regions and low importance regions.

[0032] High-importance regions are rendered using high resolution, high sampling rate, and complex shading models to obtain the first rendering result;

[0033] For low-importance regions, a second rendering result is obtained by rendering with low resolution, low sampling rate, and a simplified shading model.

[0034] The first and second rendering results are combined using either microrasterization or volumetric rendering to obtain a free-view 3D rendering result of the target scene.

[0035] The present invention also provides a multi-view 3D virtual-real fusion rendering device, comprising:

[0036] The acquisition module is used to acquire multi-time frame, multi-view video sequences in the target scene, obtain image data, and calibrate the internal and external parameters of the acquisition device to obtain calibration results.

[0037] The reconstruction module is used to perform three-dimensional reconstruction of the target scene based on image data and calibration results, and obtain the geometric field, appearance field and spatiotemporal flow field in the four-dimensional spatiotemporal domain. The geometric field and appearance field are discretized through the spatiotemporal flow field to obtain the four-dimensional Gaussian primitive. The spatiotemporal flow field is used to describe the motion trajectory of the target in the target scene over time.

[0038] The feature extraction module is used to integrate the central trajectory in the four-dimensional Gaussian primitive to obtain the spatiotemporal streamlines, divide the spatiotemporal streamlines and spatiotemporal flow field into multiple spatiotemporal markers, and extract features for each spatiotemporal marker to obtain the feature vector corresponding to each spatiotemporal marker.

[0039] The inference module is used to input the feature vector corresponding to each spatiotemporal marker into the trained Immersive Transformer module for multi-hop inference, so as to obtain the updated feature vector and the attention weights corresponding to the updated feature vector.

[0040] The rendering module is used to calculate the spatiotemporal importance distribution at each time point and from each virtual viewpoint based on the updated feature vector and the attention weights corresponding to the updated feature vector, and to render the geometric field and appearance field in the four-dimensional spatiotemporal domain based on the spatiotemporal importance distribution to obtain the free-view 3D rendering result of the target scene.

[0041] The present invention also provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a multi-view three-dimensional virtual-real fusion rendering method.

[0042] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements a multi-view 3D virtual-real fusion rendering method.

[0043] The above-described solution of the present invention has the following beneficial effects:

[0044] This invention acquires multi-timeframe, multi-view video sequences in a target scene to obtain image data. The acquisition device's intrinsic and extrinsic parameters are calibrated, and the calibration results are used to perform 3D reconstruction of the target scene, obtaining the geometric field, appearance field, and spatiotemporal flow field in the four-dimensional spatiotemporal domain. The geometric field and appearance field are discretized using the spatiotemporal flow field to obtain a four-dimensional Gaussian primitive. The center trajectory in the four-dimensional Gaussian primitive is integrated to obtain spatiotemporal streamlines. The spatiotemporal streamlines and spatiotemporal flow field are divided into multiple spatiotemporal markers, and features are extracted from each spatiotemporal marker to obtain a feature vector corresponding to each marker. The feature vector corresponding to each spatiotemporal marker is input into a trained immersion Transformer module for multi-hop inference to obtain an updated feature vector and attention weights corresponding to the updated feature vector. Based on the updated... The spatiotemporal importance distribution at each time point and virtual viewpoint is calculated using the feature vector and the attention weights corresponding to the updated feature vector. Based on the spatiotemporal importance distribution, the geometric field and appearance field in the four-dimensional spatiotemporal domain are rendered to obtain the free-view 3D rendering result of the target scene. Compared with the prior art, this invention constructs the geometric field, appearance field, and spatiotemporal flow field in the four-dimensional spatiotemporal domain, dividing the target scene into spatiotemporal streamlines or local spatiotemporal blocks. This avoids the fragmentation caused by simple frame-by-frame rendering and scene-specific rules, and solves the problem of multi-scene fragmentation. By calculating the spatiotemporal importance distribution at each time point and virtual viewpoint using the updated feature vector and the attention weights corresponding to the updated feature vector, the geometric field and appearance field in the four-dimensional spatiotemporal domain are rendered, improving the quality of the rendered image without significantly increasing the overall computational overhead.

[0045] Other beneficial effects of the present invention will be described in detail in the following detailed description section. Attached Figure Description

[0046] Figure 1 This is a flowchart illustrating an embodiment of the present invention;

[0047] Figure 2 This is a schematic diagram of multi-view acquisition and spatiotemporal flow lines in a sports scene according to an embodiment of the present invention;

[0048] Figure 3 This is a structural diagram of the multi-view 3D virtual-real fusion rendering device in an embodiment of the present invention;

[0049] Figure 4 This is a schematic diagram of the structure of the terminal device in an embodiment of the present invention. Detailed Implementation

[0050] To make the technical problems, solutions, and advantages of this invention clearer, a detailed description will be provided below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0051] In the description of this invention, it should be noted that the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0052] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a locking connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0053] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0054] This invention addresses existing problems by providing a multi-view 3D virtual-real fusion rendering method and related equipment.

[0055] like Figure 1 As shown, an embodiment of the present invention provides a multi-view 3D virtual-real fusion rendering method, including:

[0056] Step 1: Acquire multi-time frame, multi-view video sequence in the target scene to obtain image data, and calibrate the internal and external parameters of the acquisition device to obtain calibration results;

[0057] Step 2: Based on the image data and calibration results, the target scene is reconstructed in three dimensions to obtain the geometric field, appearance field and spatiotemporal flow field in the four-dimensional spatiotemporal domain. The geometric field and appearance field are discretized through the spatiotemporal flow field to obtain the four-dimensional Gaussian primitive. The spatiotemporal flow field is used to describe the motion trajectory of the target in the target scene over time.

[0058] Step 3: Integrate the central trajectory in the four-dimensional Gaussian primitive to obtain the spacetime streamlines. Divide the spacetime streamlines and spacetime flow field into multiple spacetime markers, and extract features from each spacetime marker to obtain the feature vector corresponding to each spacetime marker.

[0059] Step 4: Input the feature vector corresponding to each spatiotemporal marker into the trained Immersive Transformer module for multi-hop inference to obtain the updated feature vector and the attention weights corresponding to the updated feature vector.

[0060] Step 5: Calculate the spatiotemporal importance distribution at each time point and from each virtual viewpoint based on the updated feature vector and the attention weights corresponding to the updated feature vector, and render the geometric field and appearance field in the four-dimensional spatiotemporal domain based on the spatiotemporal importance distribution to obtain the free-view 3D rendering result of the target scene.

[0061] In this embodiment of the invention, the target scenarios include, but are not limited to, live sports broadcasts and replays, pharmaceutical production and quality inspection, conferences and teaching, home care monitoring, and live product broadcasts.

[0062] Step 1 in this embodiment of the invention specifically includes:

[0063] A set of cameras and / or depth sensors are deployed in the target scene to form a multi-view acquisition array, thereby obtaining a multi-view video sequence with multiple time frames. ;

[0064] The intrinsic and extrinsic parameters of each camera were calibrated, and the calibration results were obtained, including the intrinsic parameter matrix. and external parameters The intrinsic parameter matrix includes focal length, principal point, pixel scale, etc., while the extrinsic parameters include... This represents a 3×3 rotation matrix from the world coordinate system to the camera coordinate system, used to describe the camera orientation. This represents a translation vector from the world coordinate system to the camera coordinate system, with a size of 3×1, used to describe the camera position and construct a unified world coordinate system.

[0065] In this embodiment of the invention, the sparse point cloud or initial depth value of each time frame is estimated by combining structured light and SfM / SLAM or a depth sensor as a geometric prior.

[0066] Specifically, step 2 includes:

[0067] Given the intrinsic parameter matrices of each camera and external parameters Under these conditions, for each time frame or selected key frames, the multi-view video sequence is... The reconstruction is performed, and the resulting 3D reconstruction is used to establish a 4D spatiotemporal domain within the same world coordinate system. geometric field on , appearance and spacetime flow field Spatiotemporal flow field is used to describe the motion trajectory of a target in a target scene over time;

[0068] Discretizing the geometric and appearance fields using the spatiotemporal flow field yields a four-dimensional Gaussian primitive. For the ... A four-dimensional Gaussian primitive ,in, Indicates the center trajectory, Describing covariance, Represents a color or appearance vector. Indicates opacity.

[0069] In this embodiment of the invention, the specific process of reconstructing the multi-view video sequence for each time frame or selected key frames to obtain the three-dimensional reconstruction result is as follows:

[0070] First, the multi-view video sequence Time synchronization and alignment are performed, and distortion correction and radiometric correction (e.g., exposure, white balance, or gamma uniformization) are applied to images from each camera based on the intrinsic parameter matrix to reduce the impact of cross-camera imaging differences on matching and reconstruction.

[0071] Then, foreground segmentation or ROI clipping is performed on the target region to obtain a mask, which is used to suppress background interference in subsequent matching and 3D fusion.

[0072] Then in the same time frame The process involves extracting features from video sequences from various perspectives to obtain local feature points or descriptors, and establishing matching pairs between different cameras. ;

[0073] Then, limit constraints are introduced when establishing the correspondence:

[0074] Calculate arbitrary camera pairs using external parameters. The basic matrix The matching point must satisfy This eliminates false matches;

[0075] For multi-view matching points that pass the consistency check, the known projection matrix is ​​used. Triangulation is performed to obtain three-dimensional points (space points) in the world coordinate system. This forms the sparse point cloud at that moment. ,here Represents a set of points / point clouds, with subscripts. Indicates the first Point cloud corresponding to each time frame, superscript The term "sparse" point cloud is usually obtained by feature matching and triangulation (fewer points but more reliable). It is used for initialization, pose optimization, and constraint of subsequent dense reconstruction. It calculates the reprojection error of 3D points in the world coordinate system and performs threshold filtering to ensure the reliability of the point cloud.

[0076] With the goal of minimizing reprojection error, a joint optimization is performed on the 3D points in the world coordinate system and the camera's extrinsic parameters, expressed as:

[0077] ;

[0078] in, This means "finding the minimum / optimizing", which means minimizing the overall error by adjusting unknown parameters. The set of all reconstructed 3D points (variables) Time frame index (the first) Frame). The first frame A number of spatial points (feature points / corner points / texture points) that are tracked / matched. Indicates the camera number (number) (one camera / viewpoint) For all "cameras" "with all "space points" The observations are accumulated (which can also be understood as summing up all observation errors). This represents the square of the L2 norm (square of the Euclidean distance), which here represents the square of the reprojection error. Representing the Frame number The spatial point at the th The observed pixel coordinates on a camera image (derived from feature matching / tracking), typically two-dimensional pixel coordinates. (Homogeneous form can also be used) , It is a projection function, that is, using perspective division to transform homogeneous coordinates to pixel coordinates, for example when Obtain homogeneous pixels Time (here) , , (For the next pixel coordinates):

[0079] ;

[0080] This transforms the homogeneous coordinates back into true pixel coordinates. This process further reduces geometric errors caused by calibration residuals and matching noise, thereby improving the accuracy of 3D reconstruction. The overall physical / geometric formula is achieved by adjusting the positions of 3D points. And optionally, fine-tuning the camera's extrinsic parameters so that the positions of these 3D points projected onto the images of each camera are as close as possible to the actual observed pixel positions. This allows for a more accurate 3D structure and camera pose.

[0081] Finally, after obtaining sparse point clouds Then, dense depth maps are estimated for each viewpoint. Dense depth can be obtained through multi-view stereo (MVS). Given the camera pose, depth is searched across multiple viewpoints based on photometric and geometric consistency, ensuring consistent pixel intensity of the same 3D point projected onto each viewpoint. When structured light or a depth sensor is present, the output of the structured light / depth sensor can be directly used as... Initial values ​​or strong constraints are used to improve the reconstructability of weakly textured and reflective regions; depth maps from each camera are used to... The back projection is used to obtain a point cloud in the camera coordinate system, and then the extrinsic parameters are used to transform it to the world coordinate system. Finally, the point clouds from all viewpoints are fused to obtain a dense point cloud. superscript The point cloud is represented as "dense" and fused to obtain a dense 3D result. The fusion can be performed using TSDF / voxel fusion or weighted point cloud fusion. Outliers and floating points are removed based on observation consistency to obtain a stable 3D surface.

[0082] After obtaining dense point clouds Then, surface reconstruction is performed to generate a triangular mesh. (e.g., Poisson reconstruction or voxel-based isosurface extraction), and project the colors of multi-view images onto mesh vertices or texture maps based on visibility to achieve appearance reconstruction. When sampling colors from multiple viewpoints for the same surface point, weighted fusion is performed according to the angle of view, distance and occlusion relationship to reduce stitching and color difference.

[0083] Finally, each time frame is obtained. The 3D reconstruction results include, but are not limited to: sparse point clouds Dense point clouds Triangular mesh And its corresponding texture / color information.

[0084] Specifically, integrating the central trajectory in the four-dimensional Gaussian primitive yields the expression for the spacetime streamline:

[0085] ;

[0086] in, Represents spacetime streamlines.

[0087] In embodiments of the present invention, the spatiotemporal marker includes, but is not limited to, one or more of the following information:

[0088] Three-dimensional spatial position and its velocity and acceleration at any time;

[0089] Three-dimensional normals, curvature, or local deformation parameters;

[0090] Local appearance or texture features extracted by a convolutional neural network;

[0091] Category labels for action categories, object categories, or process steps.

[0092] Ideally, the spatiotemporal markers include position, velocity, curvature, texture descriptor, object category label, and scene identifier.

[0093] Specifically, feature extraction is performed on each spatiotemporal marker, and the expression for the feature vector corresponding to each spatiotemporal marker is as follows:

[0094] ;

[0095] in, Indicates the relationship with the first Feature vectors corresponding to spatiotemporal markers Represents the feature encoder. Indicates the first A spatiotemporal marker.

[0096] Specifically, step 4 includes:

[0097] For the feature vector corresponding to each spatiotemporal marker For each feature vector Introducing two sets of learnable projection matrices, we obtain the query vector and the key vector, calculated as follows:

[0098] ;

[0099] ;

[0100] in, Represents the query vector. Represents the key vector. , Both represent learnable projection matrices;

[0101] Based on query vector and key vector The affinity was calculated to obtain the affinity matrix used to characterize the correlation between the 4e lines of the spatiotemporal flow. The calculation expression is:

[0102] ;

[0103] in, Indicates the degree of affinity;

[0104] By using learnable thresholds and temperature to perform gated sparsity on the affinity matrix, a gated matrix is ​​obtained, thereby achieving controllable sparsity of the connections.

[0105] After fusing the affinity matrix and the gating matrix, Softmax is applied to obtain the attention weights corresponding to the updated feature vectors, and strong affinity subgraphs are induced based on the attention weights.

[0106] After constructing the graph Laplacian matrix on the strong affinity subgraph, it is coupled with all eigenvectors to obtain the updated eigenvectors.

[0107] Specifically, the elements in the gating matrix are calculated using either soft or hard gates, and the expression is as follows:

[0108] or ;

[0109] in, Represents the first in the gating matrix Line number Column elements, Indicates a soft door. Indicates a hard door. Represents spacetime markers and The level of friendliness, Indicates the learnable threshold. Indicates temperature.

[0110] The embodiments of the present invention further include the following before fusing the affinity matrix and the gating matrix:

[0111] A Top-K strategy is adopted to retain the K connections with the highest affinity in each row or column of the gating matrix, so as to improve the sparsity and interpretability of the attention graph structure.

[0112] In this embodiment of the invention, the expression for fusing the affinity matrix and the gating matrix is ​​as follows:

[0113] ;

[0114] in, Represents the fusion matrix. , All represent the fusion weights.

[0115] In this embodiment of the invention, the expression for obtaining the strong affinity subgraph based on attention weight induction is:

[0116] ;

[0117] in, Represents a strong affinity subgraph. This represents the attention weight.

[0118] Specifically, after constructing the graph Laplacian matrix on the strong affinity subgraph and coupling it with all eigenvectors, the expression for the updated eigenvectors is obtained as follows:

[0119] ;

[0120] in, This represents the updated feature vector. Represents the eigenvector. , Indicates diffusion intensity, used to control the smoothness of information between spatiotemporal streamlines. Represents the graph Laplace matrix. , ,in, This represents a graph weight matrix constructed based on attention weights or affinity. This represents the degree matrix of the graph weight matrix.

[0121] Before rendering the geometric and appearance fields in the four-dimensional spatiotemporal domain based on the spatiotemporal importance distribution, this embodiment of the invention further includes:

[0122] Virtual camera perspectives are planned based on their spatiotemporal importance distribution. Examples include motion-following perspectives and keyframe slow-motion perspectives for live sports broadcasts and replays; defect close-up perspectives for pharmaceutical quality inspections; optimal speaker-screen combination perspectives for teaching; priority perspectives for care recipients in home monitoring; and optimized product-host composition perspectives for live product broadcasts.

[0123] Specifically, based on the spatiotemporal importance distribution, the geometric and appearance fields in the four-dimensional spatiotemporal domain are rendered to obtain free-viewpoint 3D rendering results of the target scene, including:

[0124] Based on the spatiotemporal importance distribution, the geometric field and appearance field in the four-dimensional spatiotemporal domain are divided into regions, resulting in high importance regions and low importance regions.

[0125] High-importance regions are rendered using high resolution, high sampling rate, and complex shading models to obtain the first rendering result;

[0126] For low-importance regions, a second rendering result is obtained by rendering with low resolution, low sampling rate, and a simplified shading model.

[0127] The first and second rendering results are combined using either microrasterization or volumetric rendering to obtain a free-view 3D rendering result of the target scene.

[0128] In this embodiment of the invention, the complex shading model and the simplified shading model are used to map the appearance parameters of the four-dimensional Gaussian primitives at a given time and a given virtual viewpoint to pixel colors, so as to achieve differentiated rendering of high-importance regions and low-importance regions.

[0129] The simplified shading model is a view representation that is weakly or independently of the viewpoint, and is preferably used in low importance regions to reduce computational overhead and maintain real-time performance.

[0130] Preferably, the simplified shading model includes at least one of the following implementations:

[0131] (1) Direct Color Model: For the first A four-dimensional Gaussian primitive, at time... The color vector is given by the appearance field or Gaussian appearance parameters. The color vector represents the linear mapping result of the RGB color or appearance vector of the Gaussian primitive at that moment;

[0132] (2) Low-order spherical harmonic model: For the th A four-dimensional Gaussian primitive at time... Storing low-order spherical harmonics And based on the viewing direction of the corresponding pixel in the virtual viewpoint. Calculate colors:

[0133] ;

[0134] in, Represents a four-dimensional Gaussian primitive (the first) The index number of (Gaussian) units; Represents time index (the first) (frame / moment); color vector Representing the Gauss at time 1 The color vector (usually RGB three channels) is independent of or weakly related to the viewpoint in the "direct color model"; Representing the Gauss at time 1 The lower-order spherical harmonic coefficients, where It is the number of the spherical harmonic basis function, used to express "the appearance that changes with the direction of the line of sight" (e.g., the appearance of a highlight that changes with the viewing angle). Indicates the first Each spherical harmonic basis function in the direction The values ​​that can be taken on (mathematical basis functions, not physical quantities). The line of sight is represented by a unit vector pointing from the camera / virtual viewpoint toward a spatial point or pixel, thus expressing certain viewpoint-related effects while maintaining simplified calculations.

[0135] Furthermore, camera-specific exposure coefficients and color gains can be introduced to linearly correct the rendering results, thereby improving cross-camera appearance consistency.

[0136] Complex shading models are material representations of decomposable appearances and are preferred for use in high-importance areas to achieve greater realism and stronger recognizability.

[0137] More preferably, the complex coloring model is for each four-dimensional Gaussian primitive at time... Configure material elements, including but not limited to: normals Roughness , metallicity Diffuse color and mirror parameters It also introduces ambient lighting parameters (such as spherically harmonic ambient light) and one or more master light source parameters to modulate diffuse and specular reflection components;

[0138] Complex coloring models can be decomposed into diffuse and specular reflections using microsurface reflection:

[0139] ;

[0140] in, The subscript indicates diffuse reflection. The subscript indicates specular reflection. Used to characterize the inherent color and diffuse component of a material. Used to characterize the specular and specular components related to the line of sight;

[0141] Furthermore, approximate shadow or occlusion terms and indirect light approximations can be introduced to improve cross-view consistency and reduce flicker. Ambient lighting parameters approximate "background lighting from all directions" using a small number of parameters, often used for overall indoor and outdoor lighting consistency. Principal light source parameters describe the direction, intensity, color, etc., of one or more principal light sources (directional light / point light / spotlight). The aforementioned camera-by-camera exposure coefficient and color gain are scalar / vector correction quantities for cross-camera color alignment (compensating for differences in exposure and white balance between different cameras).

[0142] In this embodiment of the invention, combining the first rendering result and the second rendering result by differential rasterization or volume rendering means: at a given time... Based on the spatial and temporal importance distribution of a given virtual perspective, high-importance regions and low-importance regions are divided. The first rendering result obtained by using a complex shading model for the high-importance regions and the second rendering result obtained by using a simplified shading model for the low-importance regions are then combined in a consistent and differentiable manner at the pixel level to output a free-viewpoint 3D rendering result.

[0143] Preferably, the synthesis process includes at least one of the following implementations:

[0144] (1) Gaussian Splatting synthesis method

[0145] Preferably, the camera parameters for a given virtual viewpoint are: At any moment , will the The center of a four-dimensional Gaussian primitive Transform to the camera coordinate system and project onto the pixel plane to obtain the screen space center. At the same time, covariance By projecting the Jacobian onto screen space, we obtain two-dimensional elliptic parameters, which are used to describe the coverage of the Gaussian primitive in the pixel plane.

[0146] More preferably, for pixels (Within the coverage area of ​​the two-dimensional ellipse) Calculate the coverage weight:

[0147] ;

[0148] in, It is obtained from the inverse matrix of the screen space covariance; and pixel-level transparency is obtained by combining it with the opacity. :

[0149] ;

[0150] Furthermore, pixel colors are obtained by performing forward composition on the visible four-dimensional Gaussian primitives according to depth from near to far:

[0151] ;

[0152] in, The model is given by a complex shading model corresponding to a high importance region or a simplified shading model corresponding to a low importance region. The high importance region and the low importance region are constrained by a region partitioning mask generated by the spatiotemporal importance distribution, so that the high importance region outputs the first rendering result and the low importance region outputs the second rendering result, and consistent synthesis is used at the boundary to ensure smooth transition. Thus, the free-view 3D rendering result of the target scene at this moment and at this virtual viewpoint is obtained.

[0153] (2) Volume rendering method

[0154] More preferably, for each pixel in a given virtual viewpoint, construct a ray along the viewing direction. ,in The ray origin is usually the camera center / virtual viewpoint optical center; The unit vector of the ray direction, which is also the viewing direction of that pixel. The distance parameter (scalar) along the ray, and the set of sampling points along the ray direction { Integral synthesis is performed on};

[0155] At any moment The four-dimensional Gaussian primitives are represented as volume density and volume color, where volume density... Defined as:

[0156] ;

[0157] in, It is opacity; Therefore For the mean, The three-dimensional Gaussian distribution of covariance at point The value (a mathematical function used to convert Gaussian primitives into continuum density contributions), i.e. It is a three-dimensional Gaussian kernel;

[0158] Body color The coloring model is given by a complex coloring model for high-importance regions or a simplified coloring model for low-importance regions.

[0159] Furthermore, based on transmittance Perform discrete volume rendering to synthesize pixel colors, where... No. Step size of each sampling segment (distance between adjacent sampling points):

[0160] ;

[0161] in, It represents the volume density / extinction coefficient, used to characterize "the intensity of light being absorbed / blocked by the medium as it propagates along the light path". At any moment Along pixel ray The Each sampling point location The density value at that location. The larger the size, the thicker and less transparent the area, resulting in stronger light attenuation. The smaller the value, the sparser and more transparent the location, and the weaker the light attenuation. It approximates the light from the camera to the... "Absorption / Optical Thickness" on the path before each sampling point; Then give the transmittance That is, reaching the first At each sampling point, what percentage of the light remains unabsorbed or unblocked?

[0162] The region division mask obtained from the spatiotemporal importance distribution determines the shading and sampling configuration of the first rendering result or the second rendering result in different regions, thereby realizing the differentiable synthesis of the first rendering result and the second rendering result to obtain the free-view 3D rendering result.

[0163] In order to improve the rendering effect, this embodiment of the invention establishes a joint loss function to optimize the provided rendering method. The expression of the joint loss function is as follows:

[0164] ;

[0165] in, This represents the reconstruction loss, applied at the image / feature level, used to constrain the rendering result to be consistent with the original video / depth. This indicates task loss, which varies depending on the scenario (such as quality inspection accuracy, action recognition, defect monitoring, etc.). The aesthetic loss is indicated by scores from an aesthetic evaluation network based on aspects such as camera language, composition, and movement. This represents the constraints on the smoothness and physical consistency of the spatiotemporal flow field, preventing jitter and unreasonable distortion. This indicates that the parameters are regularized.

[0166] It should be noted that the aesthetic evaluation network can be trained based on a large number of excellent video replays, exemplary teaching videos, on-site examples from compliant pharmaceutical factories, and high-quality live product examples, and output comprehensive scores on camera movement, composition, action expressiveness, and information readability.

[0167] This invention provides a specific example of the method provided in a sports setting, and the process is as follows:

[0168] like Figure 2 As shown, multiple synchronous cameras, such as 32 to 64, are deployed around the basketball court to cover key areas of the court in a circular manner. The images captured by each camera are used to solve for intrinsic and extrinsic parameters through the calibration module in the camera, and are unified to the world coordinate system. They are then combined with traditional multi-view geometry or depth sensors to obtain sparse three-dimensional point clouds as multi-view video sequences.

[0169] The player skeleton is restored using a three-dimensional human pose estimation and tracking algorithm, and the basketball position is tracked in three dimensions. For each time sample point, the geometric field and appearance field are discretized in its vicinity to obtain a four-dimensional Gaussian primitive. The position sequence of each player's joints and the basketball in time is represented as several spatiotemporal streamlines, thereby forming a spatiotemporal flow field that describes the player's movements and trajectories.

[0170] The aforementioned spatiotemporal streamlines and spatiotemporal flow fields are divided into several spatiotemporal segments, each segment forming a spatiotemporal identifier to describe action segments such as "breakthrough", "jump", "hang in the air", "shoot", and "landing". A feature encoding network is used to extract the three-dimensional position, velocity, acceleration, joint angle changes, action category label, and local texture features of each spatiotemporal identifier, generating a feature vector corresponding to each spatiotemporal identifier.

[0171] The feature vector corresponding to each spatiotemporal marker is input into the trained Immersion Transformer module. The query vector and key vector are calculated using a learnable projection matrix, and the affinity strength is obtained in the form of a dot product. Then, the affinity strength is gated and sparsified using a learnable threshold and temperature to obtain a gating matrix. A Top-K sparsity strategy is used to filter the connections. Then, the affinity matrix and the gating matrix are fused and Softmax is applied to obtain the attention weights corresponding to the updated feature vectors. A strong affinity subgraph is induced based on the attention weights. After constructing a graph Laplacian matrix on the strong affinity subgraph, it is coupled with all feature vectors to realize the coordination within the action chain and the long-term dependency modeling in multi-round matches, thereby automatically highlighting action streamlines with high visual appeal and tactical value.

[0172] The spatiotemporal importance distribution at each moment and from each virtual perspective is calculated based on the updated feature vector and the attention weights corresponding to the updated feature vector. Virtual camera trajectories are automatically generated for each key action, such as the player's shoulder view, the basket overhead view, and the defender's view. Smooth interpolation is used to ensure that the virtual camera moves naturally.

[0173] Based on the spatiotemporal importance distribution, the geometric and appearance fields in the four-dimensional spatiotemporal domain are divided into high-importance and low-importance regions. High-importance regions are rendered using high resolution, high sampling rate, and complex shading models to achieve visual effects such as slow motion, depth of field, and halo. Low-importance regions are rendered using low resolution, low sampling rate, and simplified shading models. The first and second rendering results are then combined using differentiable rasterization or volumetric rendering to obtain a free-viewpoint 3D rendering result of the target scene.

[0174] This invention acquires multi-view video sequences across multiple time frames in a target scene to obtain image data. The acquisition device's intrinsic and extrinsic parameters are calibrated, and the calibration results are used for 3D reconstruction of the target scene, yielding the geometric field, appearance field, and spatiotemporal flow field in the four-dimensional spatiotemporal domain. The geometric field and appearance field are discretized using the spatiotemporal flow field to obtain a four-dimensional Gaussian primitive. The center trajectory in the four-dimensional Gaussian primitive is integrated to obtain a spatiotemporal streamline. The spatiotemporal streamline and spatiotemporal flow field are divided to obtain multiple spatiotemporal markers, and features are extracted from each marker to obtain a corresponding feature vector. The feature vector corresponding to each spatiotemporal marker is input into a trained immersion Transformer module for multi-hop inference to obtain an updated feature vector and attention weights corresponding to the updated feature vector. Based on... The updated feature vector and the attention weights corresponding to the updated feature vector are used to calculate the spatiotemporal importance distribution at each time point and virtual viewpoint. Based on the spatiotemporal importance distribution, the geometric field and appearance field in the four-dimensional spatiotemporal domain are rendered to obtain the free-view 3D rendering result of the target scene. Compared with the prior art, this invention constructs the geometric field, appearance field and spatiotemporal flow field in the four-dimensional spatiotemporal domain, dividing the target scene into spatiotemporal streamlines or local spatiotemporal blocks. This avoids the fragmentation caused by simple frame-by-frame rendering and scene-specific rules, and solves the problem of multi-scene fragmentation. By calculating the spatiotemporal importance distribution at each time point and virtual viewpoint through the updated feature vector and the attention weights corresponding to the updated feature vector, the geometric field and appearance field in the four-dimensional spatiotemporal domain are rendered, which improves the rendering quality without significantly increasing the overall computational overhead.

[0175] Corresponding to the multi-view 3D virtual-real fusion rendering method described in the above embodiments, such as Figure 3 As shown, this embodiment of the invention also provides a multi-view 3D virtual-real fusion rendering device 100, which includes:

[0176] The acquisition module 101 is used to acquire multi-view video sequences of multiple time frames in the target scene, obtain image data, and calibrate the internal and external parameters of the acquisition device to obtain calibration results.

[0177] The reconstruction module 102 is used to perform three-dimensional reconstruction of the target scene based on image data and calibration results, to obtain the geometric field, appearance field and spatiotemporal flow field in the four-dimensional spatiotemporal domain, and to discretize the geometric field and appearance field through the spatiotemporal flow field to obtain the four-dimensional Gaussian primitive. The spatiotemporal flow field is used to describe the motion trajectory of the target in the target scene over time.

[0178] The feature extraction module 103 is used to integrate the central trajectory in the four-dimensional Gaussian primitive to obtain the spatiotemporal streamline, divide the spatiotemporal streamline and spatiotemporal flow field into multiple spatiotemporal markers, and extract features for each spatiotemporal marker to obtain the feature vector corresponding to each spatiotemporal marker.

[0179] The inference module 104 is used to input the feature vector corresponding to each spatiotemporal marker into the trained Immersive Transformer module for multi-hop inference to obtain the updated feature vector and the attention weights corresponding to the updated feature vector.

[0180] The rendering module 105 is used to calculate the spatiotemporal importance distribution at each time point and each virtual viewpoint based on the updated feature vector and the attention weight corresponding to the updated feature vector, and to render the geometric field and appearance field in the four-dimensional spatiotemporal domain based on the spatiotemporal importance distribution to obtain the free-view 3D rendering result of the target scene.

[0181] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0182] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0183] This invention also provides a terminal device, such as... Figure 4 As shown, the terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 4 The diagram shows only one processor, a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100. When the processor D100 executes the computer program D102, it implements the above-described multi-view 3D virtual-real fusion rendering method.

[0184] The terminal device D10 can be a desktop computer, laptop, handheld computer, server, server cluster, or cloud server, etc. This terminal device may include, but is not limited to, a processor D100 and a memory D101. Those skilled in the art will understand that... Figure 4 This is merely an example of terminal device D10 and does not constitute a limitation on terminal device D10. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.

[0185] The processor D100 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0186] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may be an external storage device of the terminal device D10, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal device D10. Furthermore, the memory D101 may include both internal and external storage units of the terminal device D10. The memory D101 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory D101 can also be used to temporarily store data that has been output or will be output.

[0187] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0188] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0189] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements a multi-view 3D virtual-real fusion rendering method.

[0190] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a building device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks.

[0191] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A multi-view 3D virtual-real fusion rendering method, characterized in that, include: Step 1: Acquire multi-time frame, multi-view video sequence in the target scene to obtain image data, and calibrate the internal and external parameters of the acquisition device to obtain calibration results; Step 2: Based on the image data and the calibration results, perform 3D reconstruction of the target scene to obtain the geometric field, appearance field, and spatiotemporal flow field in the four-dimensional spatiotemporal domain. Then, discretize the geometric field and appearance field using the spatiotemporal flow field to obtain a four-dimensional Gaussian primitive. For the first... A four-dimensional Gaussian primitive ,in, Indicates the center trajectory, Describing covariance, Represents a color or appearance vector. Indicates opacity. The time frame is represented, and the spatiotemporal flow field is used to describe the motion trajectory of the target in the target scene over time. Step 3: Integrate the center trajectory in the four-dimensional Gaussian primitive to obtain the spatiotemporal streamline. Divide the spatiotemporal streamline and the spatiotemporal flow field into multiple spatiotemporal markers, and extract features from each spatiotemporal marker to obtain the feature vector corresponding to each spatiotemporal marker. The spatiotemporal marker includes position, velocity, curvature, texture descriptor, object category label and scene identifier. Step 4: Input the feature vector corresponding to each spatiotemporal marker into the trained Immersive Transformer module for multi-hop inference to obtain the updated feature vector and the attention weights corresponding to the updated feature vector, including: For each feature vector corresponding to a spatiotemporal marker, two sets of learnable projection matrices are introduced for each feature vector to obtain a query vector and a key vector; The affinity strength is calculated based on the query vector and the key vector to obtain the affinity matrix; The affinity matrix is ​​gated sparsity by a learnable threshold and temperature to obtain a gated matrix; After fusing the affinity matrix and the gating matrix, Softmax is applied to obtain the attention weights corresponding to the updated feature vectors, and a strong affinity subgraph is induced based on the attention weights. After constructing a graph Laplacian matrix on the strong affinity subgraph and coupling it with all eigenvectors, the expression for the updated eigenvectors is obtained as follows: ; in, This represents the updated feature vector. Represents the eigenvector. Indicates diffusion intensity. Represents the graph Laplace matrix; Step 5: Calculate the spatiotemporal importance distribution at each time point and each virtual viewpoint based on the updated feature vector and the attention weights corresponding to the updated feature vector, and render the geometric field and appearance field in the four-dimensional spatiotemporal domain based on the spatiotemporal importance distribution to obtain the free-view 3D rendering result of the target scene.

2. The multi-view 3D virtual-real fusion rendering method according to claim 1, characterized in that, Feature extraction is performed on each spatiotemporal marker, and the expression for the feature vector corresponding to each spatiotemporal marker is as follows: ; in, Indicates the relationship with the first Feature vectors corresponding to spatiotemporal markers Represents the feature encoder. Indicates the first A spatiotemporal marker.

3. The multi-view 3D virtual-real fusion rendering method according to claim 1, characterized in that, Based on the spatiotemporal importance distribution, the geometric and appearance fields in the four-dimensional spatiotemporal domain are rendered to obtain the free-viewpoint 3D rendering result of the target scene, including: Based on the spatiotemporal importance distribution, the geometric field and appearance field in the four-dimensional spatiotemporal domain are divided into regions, resulting in high importance regions and low importance regions. The highly important regions are rendered using high resolution, high sampling rate, and a complex shading model to obtain the first rendering result; The low-importance regions are rendered using low resolution, low sampling rate, and a simplified shading model to obtain a second rendering result; The first rendering result and the second rendering result are combined by means of microrasterization or volume rendering to obtain a free-view 3D rendering result of the target scene.

4. A multi-view 3D virtual-real fusion rendering device, characterized in that, include: The acquisition module is used to acquire multi-time frame, multi-view video sequences in the target scene, obtain image data, and calibrate the internal and external parameters of the acquisition device to obtain calibration results. The reconstruction module is used to perform three-dimensional reconstruction of the target scene based on the image data and the calibration results, obtaining the geometric field, appearance field, and spatiotemporal flow field in the four-dimensional spatiotemporal domain. The geometric field and appearance field are then discretized using the spatiotemporal flow field to obtain a four-dimensional Gaussian primitive. For the first... A four-dimensional Gaussian primitive ,in, Indicates the center trajectory, Describing covariance, Represents a color or appearance vector. Indicates opacity. The time frame is represented, and the spatiotemporal flow field is used to describe the motion trajectory of the target in the target scene over time. The feature extraction module is used to integrate the central trajectory in the four-dimensional Gaussian primitive to obtain the spatiotemporal streamline, divide the spatiotemporal streamline and the spatiotemporal flow field into multiple spatiotemporal markers, and extract features for each spatiotemporal marker to obtain the feature vector corresponding to each spatiotemporal marker. The spatiotemporal marker includes position, velocity, curvature, texture descriptor, object category label and scene identifier. The inference module is used to input the feature vector corresponding to each spatiotemporal marker into the trained immersion Transformer module for multi-hop inference, obtaining the updated feature vector and the attention weights corresponding to the updated feature vector, including: For each feature vector corresponding to a spatiotemporal marker, two sets of learnable projection matrices are introduced for each feature vector to obtain a query vector and a key vector; The affinity strength is calculated based on the query vector and the key vector to obtain the affinity matrix; The affinity matrix is ​​gated sparsity by a learnable threshold and temperature to obtain a gated matrix; After fusing the affinity matrix and the gating matrix, Softmax is applied to obtain the attention weights corresponding to the updated feature vectors, and a strong affinity subgraph is induced based on the attention weights. After constructing a graph Laplacian matrix on the strong affinity subgraph and coupling it with all eigenvectors, the expression for the updated eigenvectors is obtained as follows: ; in, This represents the updated feature vector. Represents the eigenvector. Indicates diffusion intensity. Represents the graph Laplace matrix; The rendering module is used to calculate the spatiotemporal importance distribution at each time point and each virtual viewpoint based on the updated feature vector and the attention weight corresponding to the updated feature vector, and to render the geometric field and appearance field in the four-dimensional spatiotemporal domain based on the spatiotemporal importance distribution to obtain the free-view 3D rendering result of the target scene.

5. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the multi-view three-dimensional virtual-real fusion rendering method as described in any one of claims 1 to 3.

6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the multi-view three-dimensional virtual-real fusion rendering method as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Reconstruction method of three-dimensional reconstruction model based on two-dimensional Gaussian splashing

    CN120374867A

  • Animatable character generation using 3D representations

    US20250157114A1