Scene structure perceived real-time three-dimensional scene view synthesis method, medium and equipment

By combining SfM technology and hybrid pre-trained models to optimize point cloud distribution and using hash grids and view-aware networks to generate features, the shortcomings of the new three-dimensional scene synthesis method in the existing technology in large-scale scenes and complex lighting conditions are solved, and high-quality three-dimensional scene view synthesis and real-time rendering are achieved.

CN120107525APending Publication Date: 2025-06-06SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510113900.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing new view synthesis method of three-dimensional scenes has problems such as point cloud loss, noise, redundancy and inefficient modeling when dealing with large-scale scenes, especially in complex lighting conditions, it is difficult to achieve natural lighting changes and detail rendering.

Method used

By combining SfM technology and scene geometric prior information, curvature values ​​and view-sensitive features provided by hybrid pre-trained models, point cloud distribution is optimized and high-quality three-dimensional scene views are generated. The specific steps include: generating the initial point cloud, optimizing the point cloud through pre-training models, performing densification processing, using a hash grid and a view-aware network to generate hash features and view-aware features, and finally generating Gaussian parameters through a multi-layer perceptron for rendering.

Benefits of technology

It significantly improves the detailed expression and rendering quality of three-dimensional scenes, improves the naturalness and rendering efficiency of lighting changes, and is suitable for real-time processing tasks of large-scale complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107525A_ABST
    Figure CN120107525A_ABST
Patent Text Reader

Abstract

The invention provides a scene structure perceived real-time three-dimensional scene view synthesis method, a medium and equipment. The method comprises the following steps: acquiring an initial point cloud by using an SfM technology; optimizing the initial point cloud by using scene geometric prior information provided by the hybrid pre-training model and generating an anchor point cloud; carrying out joint densification on the anchor points in the low curvature value area in the anchor point cloud and the area close to the scene surface; interpolating the densified point cloud position of the anchor point through a Hash grid to obtain Hash features; meanwhile, obtaining visual angle sensitive features through a visual angle sensing network by using the relative position of the point cloud camera; and inputting into a multi-layer perceptron to generate Gaussian parameters, and obtaining a three-dimensional scene view through rasterization rendering. According to the method, the geometric structure and the illumination information in the three-dimensional space in the visual field range of the camera can be efficiently generated, so that natural and high-quality real-time rendering is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional scene image processing, and more specifically, to a scene structure-aware real-time three-dimensional scene view synthesis method, medium and device. Background Art

[0002] Existing new view synthesis methods can be roughly divided into two categories according to the representation they are based on: 1) new view synthesis methods for 3D scenes based on NeRF (Neural Radiant Field); 2) new view synthesis methods for 3D scenes based on 3DGS (3D Gaussian Sputtering).

[0003] NeRF-based methods are limited by spectral bias when learning complex scene details. The model tends to prioritize learning low-frequency information, but lacks performance on high-frequency details (such as complex geometry or texture), affecting modeling accuracy. In addition, they rely on sampling a large number of pixels along the light to predict color and density. Even if the sampling strategy is improved, the computational requirements still lead to slow rendering speeds, making it difficult to meet the needs of real-time applications.

[0004] The 3DGS-based 3D scene new view synthesis method uses an explicit 3D Gaussian ellipsoid to represent the scene, thereby achieving real-time rendering speed and higher rendering quality. However, this method still has the following unresolved issues: 1) The initial SfM (Structure from Motion) input point cloud usually results in missing points and noise due to the presence of low-texture areas and viewpoint overlap areas in the input image, which is particularly noticeable in large-scale scenes; 2) The original 3DGS-based method only densifies according to the gradient of each Gaussian ellipsoid, ignoring the spatial structure information of the scene, which in turn leads to redundancy of point clouds and inefficiency in scene modeling to some extent, especially when the scene is large and complex; 3) For large outdoor scenes, which naturally have more complex and changeable lighting conditions, existing methods find it difficult to model natural lighting changes and usually render blurred scene structures. Summary of the invention

[0005] In order to overcome the shortcomings and deficiencies in the prior art, the purpose of the present invention is to provide a real-time three-dimensional scene view synthesis method, medium and device with scene structure perception; this method can efficiently generate geometric structure and lighting information in the three-dimensional space within the camera's field of view, thereby achieving natural, high-quality real-time rendering.

[0006] In order to achieve the above object, the present invention is implemented by the following technical solution: a real-time three-dimensional scene view synthesis method with scene structure perception, comprising the following steps: S1. Use SfM technology to obtain the initial point cloud P for the input multi-view RGB image. SfM; S2. Use hybrid pre-trained model M scene The scene geometry prior information provided by the initial point cloud P SfM Optimize and generate anchor point cloud P initial ; S3, the anchor point cloud P initial The anchor points in the low curvature value area in the scene and the area close to the scene surface are jointly densified to obtain the densified anchor point cloud; S4. Interpolate the densified anchor point cloud position through the hash grid H to obtain the hash feature f h ; At the same time, the relative position of the point cloud camera is used through the view perception network to obtain the view sensitive feature φ; S5, obtain the densified anchor point cloud, curvature value C, point cloud camera relative position x p , hash feature f h The 3D scene view is obtained by connecting it with the view-sensitive feature φ and inputting it into the multi-layer perceptron MLPs to generate Gaussian parameters, and then rendering it through rasterization.

[0007] Preferably, the step S2 comprises the following sub-steps: S21, by mixing pre-trained model M scene Generate scene geometry prior information and obtain the scene surface; S22, by the initial point cloud P SfM The query point set Q is generated by sampling and spatial perturbation; the scene geometry prior information is used, and the signed distance function and occupancy field are combined to screen and optimize the query point set to generate an anchor point cloud; S23, perform sparse processing and iterative optimization on the anchor point cloud, and finally generate an optimized anchor point cloud P initial .

[0008] Preferably, the step S22 is: by SfM The query point set Q is generated by random sampling and spatial perturbation of the sampling points; where the spatial perturbation range is [−R d , R d ]; then the distance from the query point to the scene surface is calculated using the signed distance function, retaining the points that meet the deviation range [−R s , R s ] query points, and the remaining query points are removed as noise points; The step S23 is to merge the retained query points into the initial point cloud P SfM In the process, the point cloud distribution is optimized through sparse operation; the parameter R is gradually adjusted through multiple rounds of iterations. s Until the set point cloud density requirements are met, the optimized anchor point cloud P is finally generated.initial .

[0009] Preferably, the step S3 comprises the following sub-steps: S31, based on the surface perception mechanism, through the scene geometry prior information, in the anchor point cloud P initial Generate new anchor points in the area close to the scene surface; S32, based on the curvature perception mechanism, using the anchor point cloud P initial The curvature value of the anchor point cloud P is identified initial The area with insufficient density is detected and new anchor points are generated; S33. Integrate the new anchor points generated based on the surface perception mechanism and the curvature perception mechanism; adjust the point cloud distribution through iterative optimization to obtain a densified anchor point cloud.

[0010] Preferably, the step S31 is to calculate a mask M gradient Used to identify the anchor points that need to be grown according to the cumulative gradient threshold; for N anchor points, the mask M of the i-th anchor point gradient for: ; Among them, M gradient ∈{0,1} N ; ∇θ i (t) represents the gradient of the i-th anchor point in the t-th iteration; T represents the cumulative number of iterations; τ is the cumulative gradient threshold for judging growth; Based on the surface densification, we first use the hybrid pre-trained model M scene The obtained signed distance function identifies the current anchor point cloud P initial The N anchor points of the scene are effective in describing the surface structure, and the surface dense mask M is obtained. surface ; ; Among them, M surface ∈{0,1} N ;ø S ( i ) represents the signed distance function value of the i-th anchor point, 1 represents valid, 0 represents invalid; ϵ represents the parameter of the allowable deviation range; Optimize the mask M′ through logical operations: ; Among them, ∨ represents logical OR; The step S32 is to calculate the curvature value C of each anchor point in the anchor point cloud: ; Among them, λ p(p=1,2,3) are the eigenvalues ​​of the covariance matrix, arranged in descending order; According to the curvature value C, the spatial structure expression of the anchor point is judged to obtain the curvature mask M curvature : ; The step S33 is to: transform the curvature mask M curvature Combined with mask M′, the final mask M is obtained final : .

[0011] Preferably, the step S4 comprises the following sub-steps: S41. Interpolate the densified anchor point cloud position through the hash grid H to obtain the hash feature f h ; S42, combined with the point cloud camera relative position x p , using multi-layer perceptron MLP feature Generate view-sensitive features φ to dynamically model changes in lighting and shadows.

[0012] Preferably, the step S41 is to interpolate the anchor point position x a On the hash grid H, obtain the corresponding hash feature f h : ; ; Among them, θ l i represents the i-th vector in the l-th layer hash grid; D h Represents the vector θ l i Dimension; T l is the table size of the l-th layer hash grid; L is the total number of layers of the hash grid; The step S42 is to define a spatial posture correction parameter α and a ray deviation correction parameter β: ; ; Among them, x p Represents the relative position of the point cloud camera; MLP α 、MLP β They represent multi-layer perceptrons respectively; Through multi-layer perceptron MLP feature Generate view-sensitive features φ in high-dimensional parameter space: ; Where ∘ is the Hadamard product.

[0013] Preferably, the step S5 is to generate k Gaussian primitives for each anchor point of the densified anchor point cloud, and the position of each Gaussian primitive is calculated by the following formula: ; Among them, {O 0 ,…,O k−1} is the learnable offset, l a is the distance scaling factor; The curvature value C and anchor point position x are input through multi-layer perceptrons MLPs. a , point cloud camera relative position x p , hash feature f h And the view-sensitive feature φ, calculate the Gaussian parameter A 0 , …, A k-1 : ; According to the Gaussian parameters and the position of the Gaussian primitives, the anchor points are rendered using rasterization technology to obtain a three-dimensional scene view.

[0014] A readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the scene structure-aware real-time three-dimensional scene view synthesis method.

[0015] A computer device comprises a processor and a memory for storing a program executable by the processor. When the processor executes the program stored in the memory, the real-time three-dimensional scene view synthesis method with scene structure perception is implemented.

[0016] Compared with the prior art, the present invention has the following advantages and beneficial effects: 1. Improvement of detail expression and rendering quality: By introducing geometric priors, curvature information, view-sensitive features, hash features, and camera posture correction parameters of the pre-trained scene reconstruction surface, the present invention can accurately capture the spatial structure and illumination changes of the three-dimensional scene, significantly improve detail expression, and improve the blur and unnatural illumination transition problems in scene reconstruction and rendering; 2. Efficiency and scalability: The hash grid is used to optimize spatial queries, significantly reducing the overhead of anchor point storage and calculation while maintaining high access efficiency, which is suitable for real-time processing tasks of large-scale three-dimensional scenes; 3. Enhanced noise-resistant performance: Combining the geometric information of the pre-trained scene surface and multiple input features (such as curvature, Gaussian position, camera posture, etc.), it significantly reduces noise absorption under complex lighting conditions and improves model stability and rendering accuracy; 4. Adaptability to complex scenes: Through the surface densification and curvature densification strategies guided by pre-trained surface information, the present invention demonstrates excellent adaptability in outdoor scenes with complex geometry and lighting conditions, effectively improving the reconstruction quality of low-density areas and making the rendering results more natural and realistic.

[0017] 5) Real-time and computational efficiency: The pre-trained surface geometry information and multiple features are combined to update the Gaussian primitive parameters, optimize the modeling process of viewpoint sensitivity and dynamic lighting changes, significantly improve real-time rendering performance, and reduce computational redundancy.

[0018] In summary, the present invention further improves the accuracy and efficiency of three-dimensional scene modeling and rendering through pre-trained scene reconstruction surface technology, and is applicable to the fields of virtual reality, augmented reality, autonomous driving perception, film and television production, etc., and has important practical value. By accurately capturing scene geometry and lighting changes, the present invention achieves high-quality, low-cost scene reconstruction and rendering, reflecting excellent technical effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a flow chart of a method for synthesizing real-time three-dimensional scene views with scene structure perception according to the present invention; Figure 2 is a schematic diagram of the structure-aware joint densification strategy of the present invention; Figure 3 It is a schematic diagram of a hash grid-assisted view-sensitive feature enhancement and modulation module of the present invention; Figure 4 This is a visualization result diagram of the present invention and other existing methods on the Tanks&Temple dataset; Figure 5 This is a visualization result diagram of the present invention and other existing methods on the Mill-19 and MatrixCity datasets. DETAILED DESCRIPTION

[0020] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0021] Embodiment 1 This embodiment provides a real-time 3D scene view synthesis method with scene structure perception, which is widely applicable to large-scale 3D scene modeling, virtual reality (VR), augmented reality (AR), autonomous driving perception system, and film and television rendering. By combining a variety of innovative technologies and optimization strategies, this embodiment significantly improves the rendering accuracy, processing efficiency, and adaptability in complex scenes. The method flow is as follows Figure 1 As shown, the following steps are included: S1. For the input multi-view RGB image, use SfM technology (such as COLMAP software) to obtain the initial point cloud PSfM .

[0022] S2. Geometry-guided anchor initialization: using a hybrid pre-trained model M scene The scene geometry prior information provided by (Occ-SDF Hybrid) is used to calculate the initial point cloud P SfM Optimize and generate anchor point cloud P initial .

[0023] The geometric prior knowledge of the pre-trained hybrid scene surface reconstruction model (Occ-SDF Hybrid) is used to optimize the distribution of anchor points in the three-dimensional scene to improve the quality of scene reconstruction and rendering. This method generates an optimized anchor point set by combining geometric prior information and the characteristics of sparse point clouds to solve the problems of missing anchor points, noise point interference and uneven distribution in weak texture areas in traditional methods. Step S2 includes the following sub-steps: S21, by mixing pre-trained model M scene Generate scene geometry prior information and obtain the scene surface.

[0024] S22, by the initial point cloud P SfM The query point set Q is generated by sampling and spatial perturbation; the scene geometry prior information is used, and the signed distance function (SDF) and occupancy field are combined to screen and optimize the distribution of the query point set to generate an anchor point cloud with higher coverage and more uniform distribution; Specifically, by SfM The query point set Q is generated by random sampling and spatial perturbation of the sampling points; where the spatial perturbation range is [−R d , R d ] to improve the coverage of weak texture areas; then the distance from the query point to the scene surface is calculated by the signed distance function (SDF), retaining the deviation range [−R s , R s ] query points, and the remaining query points are removed as noise points.

[0025] S23, perform sparse processing and iterative optimization on the anchor point cloud, and finally generate an optimized anchor point cloud P initial , to further improve the distribution quality of the anchor point cloud; Specifically, the retained query points are merged into the initial point cloud P SfM and through the sparse operation (for each point R s The point cloud distribution is optimized by removing the clustered points in the area to avoid excessive aggregation of anchor points; the parameter R is gradually adjusted through multiple rounds of iterations. s (Gradually increase) until the set point cloud density requirements are met, and finally generate the optimized anchor point cloud Pinitial .

[0026] S3, structure-aware joint densification strategy, such as Figure 2 As shown: The anchor point cloud P initial The anchor points in the low curvature value areas in the scene and the areas close to the scene surface are jointly densified to obtain the densified anchor point cloud. It is used to optimize the distribution of point clouds in three-dimensional scenes and improve the accuracy and detail performance of complex scene rendering. This strategy combines the two mechanisms of surface perception and curvature perception to jointly densify the anchor point cloud and achieve structural optimization of the anchor point distribution; this strategy effectively improves the distribution uniformity and geometric expression ability of the anchor points, significantly enhances the detail restoration effect of large-scale three-dimensional scenes in rendering, and provides important technical support for high-precision scene reconstruction.

[0027] Step S3 includes the following sub-steps: S31, based on the surface perception mechanism, through the scene geometry prior information (signed distance function SDF 0 value surface), in the anchor point cloud P initial New anchor points are generated in the area close to the scene surface, thereby enhancing the coverage of the geometric structure.

[0028] For the previous 3DGS method, a method based only on training gradients is usually used to determine the Gaussian ellipsoid to be densified. gradient Used to identify the anchor points that need to be grown according to the cumulative gradient threshold; for N anchor points, the mask M of the i-th anchor point gradient for: ; Among them, M gradient ∈{0,1} N ∇θ i (t) represents the gradient of the i-th anchor point in the t-th iteration; T represents the cumulative number of iterations; τ is the cumulative gradient threshold for judging growth; when the cumulative gradient of the i-th anchor point is greater than or equal to τ, M gradient (i) Set to 1, indicating that the point needs to be grown; otherwise, it is set to 0; this method ensures that anchor points with significant gradient changes are optimized and grown first.

[0029] Although this method alleviates the problems caused by the sparse point cloud generated by COLMAP to a certain extent, it still has significant limitations: 1) uneven distribution of anchor points; 2) lack of perception of the scene surface structure, resulting in reduced rendering accuracy. To address these problems, the present invention proposes a joint densification that combines the curvature of the scene point cloud and the perception of the surface structure. Specifically, the method includes two parts: a) surface-based densification and b) curvature-based densification, which together ensure that the newly added anchor points are distributed near the scene surface.

[0030] Based on the surface densification, we first use the hybrid pre-trained model M scene The obtained signed distance function (SDF) identifies the current anchor point cloud P initial The N anchor points of the scene are effective in describing the surface structure, and the surface dense mask M is obtained. surface ; ; Among them, M surface ∈{0,1} N ;ø S ( i ) represents the signed distance function value of the i-th anchor point, 1 represents valid and 0 represents invalid. Considering that the anchor point describing the structure may deviate from the 0 isosurface of the signed distance function value, ϵ represents the parameter of the allowable deviation range.

[0031] Since there are areas in the mask generated by the original heuristic densification method that fail to accurately describe the surface structure of the scene, the present invention combines these areas with the surface densification mask M. surface = 1 as the target of anchor point densification; optimize the mask M′ through logical operation: ; Among them, ∨ represents logical OR; The optimized mask M′ reflects the real surface structure of the scene more accurately and provides a more reliable basis for subsequent rendering.

[0032] In the previous section, although surface densification increases the number of anchor points near the scene surface, it is difficult to capture details in areas with low density and significant geometric changes. Studies have shown that curvature can reflect regional density and characterize the ability to express details. To this end, we propose a curvature-based densification method to add anchor points in areas with low curvature and insufficient spatial information to better capture scene details.

[0033] S32, based on the curvature perception mechanism, using the anchor point cloud P initial The curvature value of the anchor point cloud P is identified initial It can also filter out areas with insufficient density and generate new anchor points to accurately capture complex detail changes.

[0034] Specifically, calculate the curvature value C of each anchor point in the anchor point cloud: ; Among them, λ p (p=1,2,3) are the eigenvalues ​​of the covariance matrix, arranged in descending order; the curvature value C reflects the ability of the anchor point p to capture geometric details in its neighborhood; the nearest k=10 neighbors are selected for calculation; According to the curvature value C, the spatial structure expression of the anchor point is judged to obtain the curvature mask Mcurvature : ; The step S33 is to: transform the curvature mask M curvature Combined with mask M′, the final mask M is obtained final : ; Through the final mask M final , which solves the inherent defects of previous densification methods and further improves the quality of rendering details.

[0035] S33. Integrate the new anchor points generated based on the surface perception mechanism and the curvature perception mechanism; adjust the point cloud distribution through iterative optimization to obtain a densified anchor point cloud, ensuring that the newly added anchor points not only fit the structure surface but also evenly cover the low-density area.

[0036] S4, hash grid-assisted view-sensitive feature enhancement and modulation module, such as Figure 3 As shown: In order to achieve a natural transition of lighting conditions in outdoor scenes, the present invention interpolates the densified anchor point cloud positions through the hash grid H to obtain the hash feature f h ; At the same time, the relative position of the point cloud camera is used through the perspective perception network to obtain the perspective-sensitive feature φ. This module significantly improves the detail performance and lighting transition effect of the three-dimensional scene in a complex lighting environment, providing strong support for high-precision real-time rendering.

[0037] Step S4 includes the following sub-steps: S41. Interpolate the densified anchor point cloud position through the hash grid H to obtain the hash feature f h ; Establish spatial interaction relationships between anchor points and their neighborhoods to fully capture local geometry and texture information; S42, combined with the point cloud camera relative position x p , using multi-layer perceptron MLP feature Generate view-sensitive features φ to dynamically model changes in lighting and shadows.

[0038] Step S41 refers to: for the densified anchor point cloud, interpolating the anchor point position x a On the hash grid H, obtain the corresponding hash feature f h : ; ; Among them, θ l i represents the i-th vector in the l-th layer hash grid; D h Represents the vector θ li Dimension; T l is the table size of the l-th layer hash grid; L is the total number of layers of the hash grid; hash feature f h It not only effectively represents the interaction relationship between anchor points in 3D space, but is also used for subsequent Gaussian primitive refinement, parameter update and rasterization, thereby improving computational efficiency while enhancing the structural expression and rendering quality of 3D scenes. Although the hash feature f h Although the relationship between anchor points is represented, it is still insufficient when synthesizing new perspective images under complex lighting conditions, resulting in unnatural expression of lighting changes. To this end, we enhance the model's expression of the viewpoint-sensitive characteristics of the Gaussian primitives by capturing the intrinsic connection between their spatial position and the camera pose.

[0039] Step S42 is to define a spatial attitude correction parameter α and a ray deviation correction parameter β: ; ; Among them, x p Represents the relative position of the point cloud camera; MLP α 、MLP β They represent multi-layer perceptrons respectively; The spatial attitude correction parameter α and the ray offset correction parameter β are both composed of multi-layer perceptrons with the same structure but independent parameters. They take the relative camera coordinates as input to more accurately model the perspective changes of lighting and shadows.

[0040] Based on this, we use the obtained correction parameters to adjust the posture of the camera view relative to the anchor point position. In addition, through the multi-layer perceptron MLP feature Generate view-sensitive features φ in a high-dimensional parameter space to more deeply reveal the impact of view changes on the appearance of rendered images: ; where ∘ is the Hadamard product (element-wise multiplication); in this way, we can effectively model the complex lighting and illumination changes in large-scale outdoor scenes, thereby achieving more natural lighting transition effects in rendered images.

[0041] S5. Neural Gaussian parameter update: To solve the limitations of existing methods in large-scale outdoor scenes in terms of illumination changes and detail expression, the present invention dynamically adjusts the properties of Gaussian anchor points in three-dimensional scenes through various features obtained in the previous steps; this method can achieve high-precision geometric structure modeling and detail rendering at the anchor point level, significantly improving the visual expression and rendering consistency of complex scenes, and providing important technical support for efficient real-time rendering. p , hash feature fh The 3D scene view is obtained by connecting it with the view-sensitive feature φ and inputting it into the multi-layer perceptron MLPs to generate Gaussian parameters, and then rendering it through rasterization.

[0042] The densified anchor point cloud, curvature value, point cloud camera relative position, hash feature and view-sensitive feature are fused to construct high-dimensional features to fully characterize the geometry, lighting and texture properties of the anchor point under different viewpoints; the updated Gaussian parameters are generated using multi-layer perceptrons (MLPs), including the Gaussian position μ∈R 3 , rotation posture q∈R 4 , scaling factor s∈R 3 , opacity α∈R and color c∈R 3 , to ensure the adaptability of Gaussian anchor points to details and lighting changes; dynamically adjust parameters through optimization mechanism to make the anchor point distribution more consistent with the scene surface, while improving detail presentation and lighting transition effects.

[0043] Specifically, step S5 refers to: generating k Gaussian primitives for each anchor point in the densified anchor point cloud, and the position of each Gaussian primitive is calculated by the following formula: ; Among them, {O 0 ,…,O k−1} is the learnable offset, l a is the distance scaling factor; The curvature value C and anchor point position x are input through multi-layer perceptrons MLPs. a , point cloud camera relative position x p , hash feature f h And the view-sensitive feature φ, calculate the Gaussian parameter A 0 , …, A k-1 : ; According to the Gaussian parameters and the position of the Gaussian primitives, the anchor points are rendered using rasterization technology to obtain a three-dimensional scene view.

[0044] Each multi-layer perceptron refers to a multi-layer perceptron that performs parameter optimization after performing loss calculation between the three-dimensional scene view obtained by rasterization rendering and the real image.

[0045] The technical advantages of the present invention include: Influenced by the respective characteristics and mutual optimization of occupancy and signed distance field (SDF) in the scene surface reconstruction algorithm, the present invention introduces the scene geometry prior information provided by the hybrid pre-training model for the first time to guide the initialization of anchor points; it can eliminate noise anchor points far away from the structure surface, and can also generate new high-density anchor points in weak texture and repeated texture areas, thereby significantly optimizing the accuracy of geometric representation. In addition, in response to the deviation problem introduced by the adaptive encryption strategy in the current 3DGS method (such as uneven distribution of anchor points and lack of structural perception), the present invention proposes a structure-aware joint densification strategy, combining surface encryption and curvature densification methods, densifying anchor points near the structure surface, and increasing the point density in low curvature areas, effectively solving the problems of low Gaussian densification efficiency and insufficient number of point clouds in large-scale scenes, and further optimizing the uniformity and structural integrity of point cloud distribution. In order to improve the detail rendering effect and realize the natural lighting transition in outdoor scenes, we propose a perspective-sensitive feature enhancement and modulation scheme based on hash grid assistance, which effectively captures lighting changes and achieves a more natural lighting transition effect. The organic combination of these technological innovations has significantly improved rendering quality and efficiency, providing a more accurate and efficient real-time representation method for large-scale outdoor scenes.

[0046] In order to verify the technical effect of the present invention, the present invention is compared with other existing methods (for example, Mip-NerF360, iNPG, Plenoxels, 3DGS, GaMeS, Compact3DGs, Seaffold-Gs) on the Tanks&Temple dataset, as well as the Mill-19 and MatrixCity datasets. The results are shown in Tables 1 and 2; the visualization results are shown in Figure 2. Figure 4 and Figure 5 shown.

[0047] Table 1 Experimental results on the Tanks&Temple dataset

[0048] Table 2 Experimental results on Mill-19 and MatrixCity datasets

[0049] The experimental results show that the proposed method surpasses all existing methods in all three indicators on the Tanks&Temple dataset, and shows a very competitive advantage over existing methods on the Mill-19 and MatrixCity datasets, especially in the LPIPS indicator. The rendering results also show that the proposed method has significant advantages in processing color gradients, scene edges, and texture details.

[0050] Embodiment 2 This embodiment provides a readable storage medium, wherein the readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the scene structure-aware real-time three-dimensional scene view synthesis method described in the first embodiment.

[0051] Embodiment 3 This embodiment provides a computer device, including a processor and a memory for storing a program executable by the processor. When the processor executes the program stored in the memory, the scene structure-aware real-time three-dimensional scene view synthesis method described in the first embodiment is implemented.

[0052] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.

Claims

1. A scene structure-aware real-time 3D scene view synthesis method, characterized by: The steps include: S1. Use SfM technology to obtain the initial point cloud P for the input multi-view RGB image. SfM ; S2. Use hybrid pre-trained model M scene The scene geometry prior information provided by the initial point cloud P SfM Optimize and generate anchor point cloud P initial ; S3, the anchor point cloud P initial The anchor points in the low curvature value area in the scene and the area close to the scene surface are jointly densified to obtain the densified anchor point cloud; S4. Interpolate the densified anchor point cloud position through the hash grid H to obtain the hash feature f h ; At the same time, the relative position of the point cloud camera is used through the view perception network to obtain the view sensitive feature φ; S5, obtain the densified anchor point cloud, curvature value C, point cloud camera relative position x p , hash feature f h The 3D scene view is obtained by connecting it with the view-sensitive feature φ and inputting it into the multi-layer perceptron MLPs to generate Gaussian parameters, and then rendering it through rasterization.

2. The scene structure-aware real-time 3D scene view synthesis method according to claim 1, characterized in that: The step S2 comprises the following sub-steps: S21, by mixing pre-trained model M scene Generate scene geometry prior information and obtain the scene surface; S22, by the initial point cloud P SfM The sampling and spatial perturbation of generate the query point set Q; Using scene geometry prior information, combined with signed distance function and occupancy field, the query point set is screened and distribution optimized to generate anchor point cloud; S23, perform sparse processing and iterative optimization on the anchor point cloud, and finally generate an optimized anchor point cloud P initial .

3. The scene structure-aware real-time 3D scene view synthesis method according to claim 2, characterized in that: The step S22 is: by SfM The query point set Q is generated by random sampling and spatial perturbation of the sampling points; where the spatial perturbation range is [−R d , R d ]; then the distance from the query point to the scene surface is calculated using the signed distance function, retaining the points that meet the deviation range [−R s , R s ] query points, and the remaining query points are removed as noise points; The step S23 is to merge the retained query points into the initial point cloud P SfM In the process, the point cloud distribution is optimized through sparse operation; the parameter R is gradually adjusted through multiple rounds of iterations. s Until the set point cloud density requirements are met, the optimized anchor point cloud P is finally generated. initial .

4. The scene structure-aware real-time 3D scene view synthesis method according to claim 1, characterized in that: The step S3 comprises the following sub-steps: S31, based on the surface perception mechanism, through the scene geometry prior information, in the anchor point cloud P initial Generate new anchor points in the area close to the scene surface; S32, based on the curvature perception mechanism, using the anchor point cloud P initial The curvature value of the anchor point cloud P is identified initial The area with insufficient density is selected and new anchor points are generated; S33. Integrate the new anchor points generated based on the surface perception mechanism and the curvature perception mechanism; adjust the point cloud distribution through iterative optimization to obtain a densified anchor point cloud.

5. The scene structure-aware real-time 3D scene view synthesis method according to claim 4, characterized in that: The step S31 is to calculate a mask M gradient Used to identify the anchor points that need to be grown according to the cumulative gradient threshold; for N anchor points, the mask M of the i-th anchor point gradient for: ; Among them, M gradient ∈{0,1} N ∇θ i (t) represents the gradient of the i-th anchor point in the t-th iteration; T represents the cumulative number of iterations; τ is the cumulative gradient threshold for judging growth; Based on the surface densification, we first use the hybrid pre-trained model M scene The obtained signed distance function identifies the current anchor point cloud P initial The N anchor points of the scene are effective in describing the surface structure, and the surface dense mask M is obtained. surface ; ; Among them, M surface ∈{0,1} N ;ø S ( i ) represents the signed distance function value of the i-th anchor point, 1 represents valid, 0 represents invalid; ϵ represents the parameter of the allowable deviation range; Optimize the mask M′ through logical operations: ; Among them, ∨ represents logical OR; The step S32 is to calculate the curvature value C of each anchor point in the anchor point cloud: ; Among them, λ p (p=1,2,3) are the eigenvalues ​​of the covariance matrix, arranged in descending order; According to the curvature value C, the spatial structure expression of the anchor point is judged to obtain the curvature mask M curvature : ; The step S33 is to: transform the curvature mask M curvature Combined with mask M′, the final mask M is obtained final : 。 6. The scene structure-aware real-time 3D scene view synthesis method according to claim 1, characterized in that: The step S4 comprises the following sub-steps: S41. Interpolate the densified anchor point cloud position through the hash grid H to obtain the hash feature f h ; S42, combined with the point cloud camera relative position x p , using multi-layer perceptron MLP feature Generate view-sensitive features φ to dynamically model changes in lighting and shadows.

7. The scene structure-aware real-time 3D scene view synthesis method according to claim 6, characterized in that: The step S41 is to interpolate the anchor point position x a On the hash grid H, obtain the corresponding hash feature f h : ; ; Among them, θ l i represents the i-th vector in the l-th layer hash grid; D h Represents the vector θ l i Dimension; T l is the table size of the l-th layer hash grid; L is the total number of layers of the hash grid; The step S42 is to define a spatial posture correction parameter α and a ray deviation correction parameter β: ; ; Among them, x p Represents the relative position of the point cloud camera; MLP α 、MLP β They represent multi-layer perceptrons respectively; Through multi-layer perceptron MLP feature Generate view-sensitive features φ in high-dimensional parameter space: ; Where ∘ is the Hadamard product.

8. The scene structure-aware real-time 3D scene view synthesis method according to claim 1, characterized in that: The step S5 is to generate k Gaussian primitives for each anchor point of the densified anchor point cloud, and the position of each Gaussian primitive is calculated by the following formula: ; Among them, {O0,…,O k−1 } is the learnable offset, l a is the distance scaling factor; The curvature value C and anchor point position x are input through multi-layer perceptrons MLPs. a , point cloud camera relative position x p , hash feature f h and the view-sensitive feature φ, calculate the Gaussian parameters A0,…,A k-1 : ; According to the Gaussian parameters and the position of the Gaussian primitives, the anchor points are rendered using rasterization technology to obtain a three-dimensional scene view.

9. A readable storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, enables the processor to execute the scene structure-aware real-time three-dimensional scene view synthesis method according to any one of claims 1 to 8.

10. A computer device comprising a processor and a memory for storing a program executable by the processor, characterized in that: When the processor executes the program stored in the memory, the real-time three-dimensional scene view synthesis method with scene structure perception described in any one of claims 1 to 8 is implemented.