Method, device and equipment for generating three-dimensional gaussian representation based on scene data
By generating scene point clouds, camera poses, and rendered images to produce 3D Gaussian representations, the problem of low efficiency and low accuracy in 3D scene reconstruction in existing technologies is solved, and automated and high-precision 3D reconstruction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU QUNHE INFORMATION TECHNOLOGIES CO LTD
- Filing Date
- 2025-12-16
- Publication Date
- 2026-05-08
AI Technical Summary
In existing 3D scene reconstruction technologies, manual reconstruction is inefficient, while hardware reconstruction is not accurate in low-texture areas. It is also difficult to accurately establish the correspondence between images and point clouds, which affects training accuracy.
The system generates scene point clouds, camera poses, and rendered images based on 3D scene data. It generates 3D Gaussian representations using a 3D Gaussian algorithm, including point cloud generation, pose generation, and image generation modules. The system utilizes 3D transformation matrices and a rendering engine for automated processing.
It improves the processing efficiency and accuracy of 3D reconstruction, enhances data consistency, and improves the accuracy of 3D Gaussian representation and scene reconstruction quality.
Smart Images

Figure CN121392155B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to the fields of 3D Gaussian representation generation and 3D modeling technology. Background Technology
[0002] 3D scene generation and reconstruction primarily rely on manual reconstruction and hardware reconstruction. Manual reconstruction involves manually constructing the scene and camera positions using 3D modeling software and generating rendered images, resulting in low modeling efficiency. Hardware reconstruction utilizes methods such as structured light and laser scanning to acquire point cloud data, combined with traditional multi-view geometry techniques to reconstruct the 3D scene. However, hardware reconstruction suffers from inconsistent training data, exhibiting significant errors in low-texture areas, and making it difficult to accurately establish the correspondence between images, point clouds, and cameras, thus affecting training accuracy. Summary of the Invention
[0003] This disclosure provides a method, apparatus, and device for generating three-dimensional Gaussian representations based on scene data, to solve or alleviate one or more technical problems in the prior art.
[0004] Firstly, this disclosure provides a method for generating a 3D Gaussian representation based on scene data, including:
[0005] Generating scene point clouds based on the triangular mesh corresponding to the model in 3D scene data includes: obtaining the triangular mesh and 3D transformation matrix corresponding to the model; scaling the triangular mesh corresponding to the model based on the scaling value of the 3D transformation matrix; and sampling based on the scaled triangular mesh to obtain the scene point cloud.
[0006] The camera pose is generated based on the apartment layout in the 3D scene data.
[0007] A rendered image is generated based on the camera pose and the 3D scene data;
[0008] A 3D Gaussian representation is generated based on the scene point cloud, the camera pose, and the rendered image.
[0009] Secondly, this disclosure provides a three-dimensional Gaussian representation generation apparatus based on scene data, comprising:
[0010] The point cloud generation module is used to generate scene point clouds based on the triangular mesh corresponding to the model in the 3D scene data. The point cloud generation module includes: an acquisition submodule, used to acquire the triangular mesh and 3D transformation matrix corresponding to the model; a scaling submodule, used to scale the triangular mesh corresponding to the model based on the scaling value of the 3D transformation matrix; and a point cloud submodule, used to sample based on the scaled triangular mesh to obtain the scene point cloud.
[0011] The pose generation module is used to generate camera poses based on the apartment layout in the 3D scene data.
[0012] An image generation module is used to generate rendered images based on the camera pose and the 3D scene data.
[0013] The Gaussian generation module is used to generate a 3D Gaussian representation based on the scene point cloud, the camera pose, and the rendered image.
[0014] Thirdly, an electronic device is provided, comprising:
[0015] At least one processor; and
[0016] The memory is communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.
[0018] Fourthly, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of the present disclosure.
[0019] Fifthly, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of the present disclosure.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0021] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments provided according to this disclosure and should not be construed as limiting the scope of this disclosure.
[0022] Figure 1 This is a flowchart illustrating a method for generating a 3D Gaussian representation based on scene data according to an embodiment of the present disclosure;
[0023] Figure 2 This is a flowchart illustrating a method for generating a 3D Gaussian representation based on scene data according to another embodiment of the present disclosure;
[0024] Figure 3 This is a flowchart illustrating a method for generating a 3D Gaussian representation based on scene data according to another embodiment of the present disclosure;
[0025] Figure 4 This is a flowchart illustrating a method for generating a 3D Gaussian representation based on scene data according to another embodiment of the present disclosure;
[0026] Figure 5 This is a flowchart illustrating a method for generating a 3D Gaussian representation based on scene data according to another embodiment of the present disclosure;
[0027] Figure 6 This is a flowchart illustrating a method for generating a 3D Gaussian representation based on scene data according to another embodiment of the present disclosure;
[0028] Figure 7 This is a flowchart illustrating the overall process structure for automatically generating 3D Gaussian scenes based on input 3D scene data.
[0029] Figure 8 This is a schematic diagram of the structure of a three-dimensional Gaussian representation generation device based on scene data according to an embodiment of the present disclosure;
[0030] Figure 9 This is a schematic diagram of the structure of a three-dimensional Gaussian representation generation device based on scene data according to another embodiment of the present disclosure;
[0031] Figure 10 This is a block diagram of an electronic device used to implement embodiments of the present disclosure. Detailed Implementation
[0032] The present disclosure will now be described in further detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.
[0033] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.
[0034] Figure 1 This is a flowchart illustrating a method for generating a 3D Gaussian representation based on scene data according to an embodiment of the present disclosure. The method may include:
[0035] S110. Generate scene point cloud based on the triangular mesh corresponding to the model in the 3D scene data;
[0036] S120. Generate camera pose based on the apartment layout in the 3D scene data;
[0037] S130. Generate a rendered image based on the camera pose and the 3D scene data;
[0038] S140. Generate a 3D Gaussian representation based on the scene point cloud, the camera pose, and the rendered image.
[0039] In this embodiment of the disclosure, three-dimensional scene data may also be referred to as three-dimensional data, 3D scene data, 3D data, etc. Three-dimensional scene data may include data such as models, materials, lighting, and floor plan structure. In some examples, models may include: architectural models, such as models of walls, doors, windows, stairs, balconies, etc. of a residence; furniture and appliance models, such as models of sofas, dining tables, beds, refrigerators, air conditioners, televisions, etc.; and decorative items, such as models of paintings, plants, ornaments, carpets, curtains, etc. In some examples, materials may include wall materials, floor materials, furniture materials, door and window glass materials, etc. In some examples, lighting data may include the color temperature, brightness, and illumination range of main light sources, auxiliary light sources, ambient light sources, etc. In some examples, floor plan structure may include spatial layout, such as the division and area dimensions of the living room, bedroom, kitchen, and bathroom; structural details, such as the location and thickness of load-bearing walls, and the dimensions of door and window openings; and spatial relationships, such as the connection method of rooms, floor height, and the location and height of beams.
[0040] In this embodiment, by parsing a model in a 3D scene data, several triangular meshes corresponding to that model can be obtained. A triangular mesh is the geometric representation of a 3D model. By decomposing the model surface into countless non-overlapping, seamlessly connected triangular facets—i.e., triangular meshes—the three-dimensional shape of the model is precisely constructed using vertex coordinates, facet connection relationships, and normal directions. Each triangular facet is defined by three spatial vertices, and all triangular facets combine to form the surface contour of the model. For example, the curved surfaces of a furniture model and the edges of a building wall can be refined using triangular meshes. Sampling points can be obtained by sampling several triangular meshes of the model. A point cloud of the model can be generated based on all the sampled points of the model. A scene point cloud can be generated based on the point clouds of all models in the scene.
[0041] In this embodiment, by parsing the floor plan structure in the 3D scene data, information such as the room outlines and heights within the floor plan structure can be obtained. Through a preset camera pose configuration strategy, multiple camera poses can be generated in each room of the floor plan structure. A camera pose can include camera position and camera angle. The camera position represents the spatial location of the camera. The camera angle can include the camera's orientation and posture, such as horizontal forward, vertical downward, tilted upward by 30°, or rotated 90° around its own axis. Different camera poses can correspond to different camera viewpoints. The camera viewpoint represents the camera's shooting direction and range. The same camera position may correspond to multiple camera viewpoints. Based on a single camera viewpoint and the materials, lighting, etc., in the 3D scene data, a corresponding rendered image can be generated. Based on all camera poses within the scene, several multi-view rendered images can be generated.
[0042] In this embodiment, a 3D Gaussian representation of the 3D scene data can be calculated using a 3D Gaussian algorithm based on the scene point cloud, camera pose, and rendered image generated from the 3D scene data. The 3D Gaussian representation can be used to express the geometric shape and appearance information of a 3D scene. It can be applied to various scenarios such as 3D reconstruction, virtual reality, augmented reality, film and animation production, and industrial design and manufacturing.
[0043] According to embodiments of this disclosure, scene point clouds, camera poses, and rendered images are automatically generated based on 3D scene data, and a 3D Gaussian representation is automatically generated. This can improve processing efficiency, enhance the data consistency of point clouds, rendered images, and camera poses, thereby improving the accuracy of 3D Gaussian representation data and the accuracy of 3D reconstruction, etc.
[0044] Figure 2 This is a flowchart illustrating a method for generating a 3D Gaussian representation based on scene data according to another embodiment of the present disclosure. The method may include one or more features of the aforementioned method for generating a 3D Gaussian representation based on scene data. In one embodiment, step S110 generates a scene point cloud based on the triangular mesh corresponding to the model in the 3D scene data, including:
[0045] S210. Obtain the triangular mesh and 3D transformation matrix corresponding to the model;
[0046] S220. Scale the triangular mesh corresponding to the model based on the scaling value of the three-dimensional transformation matrix;
[0047] S230. Sampling is performed based on the scaled triangular mesh to obtain the scene point cloud.
[0048] In this embodiment of the disclosure, the triangular mesh and 3D transformation matrix of each model can be obtained by parsing from the 3D scene model. The triangular mesh includes geometric information such as vertices and faces (composed of triangles), and the 3D transformation matrix can be used to describe transformations such as translation, rotation, and scaling of the model. The 3D transformation matrix can include various scaling values, various translation values, and various rotation angles. For example, the scaling value can be 1:2, 1:3, 3:1, or 4:3, etc.; the rotation angle can be a 30° rotation to the right, a 70° rotation to the left, etc.
[0049] In this embodiment, scaling the model changes the area of the triangular mesh, thus affecting the sampling density. Therefore, the triangular mesh can be sampled separately based on different scaling values of the model. After loading the triangular mesh, the model can be divided into different scaling groups according to different scaling values. Within each scaling group, the triangular mesh is first scaled according to its corresponding scaling value. For example, the scaling values are 1:2 and 1:3, and the scene data includes the original models A, B, and C. Among them, models A1, B1, and C1 scaled at a scaling value of 1:2 belong to scaling group 1, and models A2, B2, and C2 scaled at a scaling value of 1:3 belong to scaling group 2. By sampling the models of different scaling groups according to the set sampling method, point clouds of different scaling groups can be obtained. The point clouds of all scaling groups can be combined to form the overall scene point cloud.
[0050] In this embodiment of the disclosure, a multi-process approach can be used to sample and process multiple scaled triangular meshes in parallel, and finally the sampling results of all scaled triangular meshes are summarized and merged into a scene point cloud.
[0051] According to embodiments of this disclosure, by scaling and sampling triangular meshes, different accuracy requirements can be adapted to improve the efficiency of scene point cloud generation.
[0052] Figure 3 This is a flowchart illustrating a method for generating a 3D Gaussian representation based on scene data according to another embodiment of the present disclosure. The method may include one or more features of the aforementioned method for generating a 3D Gaussian representation based on scene data. In one embodiment, step S230 involves sampling based on a scaled triangular mesh to obtain a scene point cloud, including:
[0053] S310. Based on the surface area of the scaled triangular mesh and the given point spacing, perform uniform surface sampling on the scaled triangular mesh to obtain multiple base points;
[0054] S320. Based on these multiple base points, multiple new points are obtained by calculating according to the rotation parameters and / or translation parameters in the three-dimensional transformation matrix;
[0055] S330. Based on the multiple base points and the multiple newly added points, obtain the point cloud of the scaling group corresponding to the scaling value; wherein, the point cloud of all scaling groups of a model constitutes the point cloud of the model, and the point cloud of all models in the 3D scene data constitutes the point cloud of the scene.
[0056] In this embodiment, the given point spacing is a preset sampling parameter, which can be used to adjust the density of sampling points. The area of a triangular mesh differs from that of its scaled-down counterpart. Uniform surface sampling can be performed based on the surface areas of the original and scaled-down triangular meshes and the given point spacing. The core objective of uniform surface sampling is to generate a set of discrete points with uniform distribution and consistent density on the surface of a 3D model, with sampling points spaced as evenly as possible and without obvious clustering or gaps. Multiple base points of the triangular mesh can be obtained through uniform sampling. For example, if the surface area of the original triangular mesh is S1, uniform surface sampling at a given spacing yields 10 sampling points. Similarly, if the surface area of the original triangular mesh is S2 after being scaled down, uniform surface sampling at a given spacing yields 6 sampling points. Furthermore, if the surface area of the original triangular mesh is S3 after being scaled up, uniform surface sampling at a given spacing yields 25 sampling points. Sampling points from all scaling groups of a model can constitute base points.
[0057] In this embodiment, the three-dimensional transformation matrix may include parameters such as rotation and translation of the model. After sampling the base points of the triangular mesh, the base points of the triangular mesh can be calculated according to the rotation parameters and / or translation parameters in the three-dimensional transformation matrix to obtain multiple newly added sampling points (hereinafter referred to as new points). The translation parameter T = (Tx, Ty, Tz) in the three-dimensional transformation matrix allows the base point (x, y, z) to be calculated according to the corresponding values in the translation vector to obtain the new point (x+Tx, y+Ty, z+Tz). For example, if Tx=5, Ty=3, and Tz=7, the base point moves 5 units in the positive X-axis direction, 3 units in the positive Y-axis direction, and 7 units in the positive Z-axis direction. The rotation parameter in the three-dimensional transformation matrix is a 30° rotation around its own axis, which can be used to calculate the position of the base point after the model has rotated 30°.
[0058] In this embodiment, multiple base points and multiple new points of the triangular mesh in the current scaling group of a model are merged to obtain the point cloud of the model in that scaling group. Based on the point clouds of all scaling groups of the model, the point cloud of the model can be obtained. Merging the point clouds of all models can yield the scene point cloud. Alternatively, the base points and new points of all models belonging to a scaling group can be merged to obtain the point cloud of that scaling group, and then the point clouds of all scaling groups can be merged into the scene point cloud.
[0059] For example, models A, B, and C belong to scaling group 0, models A1, B1, and C1 belong to scaling group 1, and models A2, B2, and C2 belong to scaling group 2. Among them, A1 and A2 are obtained by scaling model A, B1 and B2 are obtained by scaling model B, and C1 and C2 are obtained by scaling model C.
[0060] In one approach, the base points and newly added points of model A in scaling group 0 can be merged to obtain the point cloud of model A in scaling group 0. Similarly, the base points and newly added points of model A1 in scaling group 1 can be merged to obtain the point cloud of model A in scaling group 1, and the base points and newly added points of model A2 in scaling group 2 can be merged to obtain the point cloud of model A in scaling group 2. The point clouds of model A in scaling groups 0, 1, and 2 can then be merged to obtain the point cloud of model A. Likewise, the point clouds of model B and model C can be obtained. Finally, the point clouds of models A, B, and C can be merged to obtain the scene point cloud.
[0061] In another approach, the point clouds of scaling groups A, B, and C are merged at their base points and new points in scaling group 0 to obtain the point cloud of scaling group 0. Similarly, the point clouds of scaling groups A1, B1, and C1 are merged at their base points and new points in scaling group 1 to obtain the point cloud of scaling group 1. Finally, the point clouds of scaling groups 0, 1, and 2 are merged to obtain the scene point cloud.
[0062] According to the embodiments of this disclosure, new points are obtained by calculating the base points using the parameters in the three-dimensional transformation matrix. This allows for the rapid acquisition of more sampling points, reduces redundant sampling points, improves sampling efficiency and accuracy, and ultimately enhances overall processing efficiency.
[0063] Figure 4 This is a flowchart illustrating a method for generating a 3D Gaussian representation based on scene data according to another embodiment of the present disclosure. The method may include one or more features of the aforementioned method for generating a 3D Gaussian representation based on scene data. In one embodiment, step S120 generates a camera pose based on the floor plan structure in the 3D scene data, including:
[0064] S410. Based on the camera spacing and the volume of the rooms in the apartment layout, obtain the number of camera points required for the room.
[0065] S420. Perform Poisson sampling in the room to determine the camera locations in the room;
[0066] S430. Generate the camera viewpoint for each camera point to obtain the camera pose.
[0067] In this embodiment, camera pose can be acquired using a spatial strategy. The floor plan structure may include data such as room outlines, room heights, and room volumes in the 3D scene. The camera spacing can be a preset parameter used to determine the distance between each camera point. Based on the camera spacing and the room volumes in the floor plan structure, the number of camera points required in each room can be calculated. The room volume is calculated using the room's 3D geometry (e.g., cuboid, irregular shape). The number of camera points required for the current room is determined based on the room volume data. For example, if the room volume is 5 square meters, and one camera point can cover 0.5 square meters of room area, then the room requires 10 camera points.
[0068] In this embodiment, Poisson Disk Sampling is a highly uniform spatial sampling algorithm. Its core feature is ensuring that the distance between sampling points is not less than a preset threshold, generating a uniformly distributed point set without clusters or gaps. Performing Poisson sampling within a room can evenly distribute a predetermined number of camera points across the room's spatial locations.
[0069] In this embodiment of the disclosure, one or more camera viewpoints can be generated at each camera location, and a camera location and a first camera viewpoint can correspond to a camera pose. For example, camera location P1 has corresponding camera poses in six different viewpoints: the camera pose of P1 in the front view, the camera pose of P1 in the rear view, the camera pose of P1 in the left view, the camera pose of P1 in the right view, the camera pose of P1 in the top view, and the camera pose of P1 in the bottom view. By determining the camera location and the corresponding camera viewpoint, the camera's shooting range can be determined.
[0070] According to embodiments of this disclosure, determining camera positions by camera spacing ensures coverage within a given space. Determining camera positions through Poisson sampling ensures uniform camera distribution, avoiding data redundancy caused by overlapping viewpoints. Compared to manual placement, automatic camera pose arrangement improves placement efficiency.
[0071] Figure 5 This is a flowchart illustrating a method for generating a 3D Gaussian representation based on scene data according to another embodiment of the present disclosure. The method may include one or more features of the aforementioned method for generating a 3D Gaussian representation based on scene data. In one embodiment, step S120 generates a camera pose based on the floor plan structure in the 3D scene data, including:
[0072] S510. Based on the planar outline of the room in the apartment structure, obtain the annular curve of the room;
[0073] S520. Based on the height of the room, the annular curve is copied to obtain a multi-layer annular curve;
[0074] S530. Based on the camera spacing on the multi-layered circular curve, determine the required camera positions for the room;
[0075] S540. Based on the curve direction of each layer of the annular curve, determine the camera viewpoint of each camera point on the annular curve of that layer, and obtain the camera pose.
[0076] In this embodiment, the camera pose can be obtained using a surround strategy. The planar outline of a room in the apartment layout can be determined based on the projected boundary lines of the room's floor (walls, ceiling) on a horizontal plane. The planar outline can accurately describe the room's planar shape, wall layout, and door and window positions. Based on the planar outline, shrinking inward by a preset distance, such as 1 cm or 1 decimeter, yields the room's circular curve.
[0077] In this embodiment, the room height can be obtained by acquiring the apartment layout. The room height can be the vertical distance from the indoor floor to the ceiling (or the lower edge of the suspended ceiling). Based on the room height, multiple copies of the circular curve can be made, allowing the circular curves to be distributed evenly throughout the room. For example, if the room height is 3 meters, the circular curve can be copied twice, for a total of three layers, with each layer spaced one meter apart. Similarly, if the room height is 6 meters, the circular curve can be copied three times, for a total of four layers, with each layer spaced 1.5 meters apart. The specific number of layers and spacing can be flexibly set according to the accuracy, computing power, and other requirements of the actual application scenario.
[0078] In this embodiment of the disclosure, the spacing between cameras can be preset on the circular curve. This spacing can also be referred to as camera spacing, camera point spacing, point spacing, preset spacing, etc. The required camera points in the room can be determined by the camera spacing. For example, if the total length of the circular curve is 50m and the camera spacing is 1 meter, then there are 50 preset camera points on the first-floor circular curve; if the camera spacing is 2 meters, then there are 25 preset camera points on the first-floor circular curve.
[0079] In this embodiment, the direction of the annular curve can be either clockwise or counterclockwise. When arranging cameras on the annular curve, the direction of each layer of the annular curve can be considered. The shooting order and viewpoint coverage of the cameras are determined based on the curve direction. For example, if the cameras are arranged clockwise, images are captured one by one clockwise, starting with the first camera. Alternatively, if the cameras are arranged clockwise, images are captured starting with the first camera, facing the inner side of the annular curve (towards the center), and the shooting direction is consistent with the clockwise direction of the annular curve. The camera viewpoint of each camera point on the annular curve is determined by the curve direction, thereby generating the camera pose corresponding to each camera point.
[0080] After determining the camera positions on the circular curve, the camera angle corresponding to each camera position can be determined based on the position of the circular curve. For example, the pitch angle of each camera position can be set: the upper layer looks down (set a top view or a downward view), and the lower layer looks up (set a bottom view or a downward view).
[0081] According to embodiments of this disclosure, camera positions and viewing angles are determined by camera spacing and planar contour lines, resulting in uniform camera distribution, avoiding data redundancy caused by overlapping viewing angles, and improving camera coverage. Compared to manual placement, automatic camera pose arrangement can improve placement efficiency.
[0082] In the embodiments of this disclosure, the spatial strategy and the orbiting strategy can be used individually or in combination. For example, the individual camera poses determined by the spatial strategy and the orbiting strategy can be combined to obtain the overall camera pose.
[0083] Figure 6 This is a flowchart illustrating a method for generating a 3D Gaussian representation based on scene data according to another embodiment of the present disclosure. The method may include one or more features of the aforementioned method for generating a 3D Gaussian representation based on scene data. In one embodiment, step S130 generates a rendered image based on the camera pose and the 3D scene data, including:
[0084] S610, Animation sequence of a camera synthesized from multiple camera poses;
[0085] S620. Construct the scene to be rendered based on the animation sequence and the 3D scene data;
[0086] S630 uses the rendering engine to render the scene to be rendered frame by frame, obtaining a rendered image corresponding to each camera pose.
[0087] In this embodiment, the camera animation sequence refers to the smooth change of camera pose over time to generate a continuous viewpoint, which is then processed and output as an animation sequence, also known as an image sequence. By interpolating the camera's position and orientation along the time axis, multiple camera poses determined by spatial and / or orbital strategies can generate multiple animation sequences. For example, the trajectory of the animation sequence can include basic trajectories such as straight lines (translation scanning), loops (rotation around a target), arcs (orbiting and ascending / descending), polylines (multi-target switching), complex trajectories such as spirals (ascending / descending around a target), Bézier curves (free curve motion), and following trajectories (moving with an object), etc. The camera pose changes in the animation sequence can include fixed orientation, such as always pointing towards a target (e.g., the center of a room, an exhibit) during camera movement; dynamic orientation, such as adjusting the orientation synchronously with the trajectory (e.g., pointing towards the inside of the loop during looping motion, slowly turning towards a new target during straight-line motion); and viewpoint scaling, such as combining focal length adjustment (zooming in to magnify details, zooming out to show the whole picture), etc.
[0088] In this embodiment, a scene to be rendered can be constructed based on the generated animation sequence and material and lighting data in the 3D scene data. Specifically, the camera's animation sequence can be bound to the camera's viewpoint in the rendering engine. The material and lighting conditions to be rendered in the scene are determined using the material and lighting data in the 3D scene data. For example, material conditions may include material properties (such as roughness and reflectivity) that need to match real-world objects (e.g., matte walls, highly reflective metal furniture), and the materials of different objects in the 3D scene should also be clearly distinguishable (e.g., wood grain on the floor, latex paint on the walls, transparency of glass). For example, lighting conditions may include the color temperature, brightness, and illumination range of the key light, fill light, back light, etc.
[0089] In this embodiment, the rendering engine can be a real-time rendering engine or an offline rendering engine. The rendering engine parses the constructed scene to be rendered, calculating the pixel color of each frame according to physical or non-physical rules based on the scene's 3D model, materials, camera animation, lighting, and other data, ultimately outputting an animation sequence. Based on the animation sequence, a rendered image can be obtained frame-by-frame according to each camera pose. If a camera position corresponds to multiple camera views, rendered images from multiple views of that position can be obtained.
[0090] According to embodiments of this disclosure, constructing a scene to be rendered using animation sequences and 3D scene data allows for accurate scene reproduction, reduces errors, and maintains the controllability and consistency of the rendering style. Obtaining the rendered image through a rendering engine enhances visual realism and improves scene rendering efficiency.
[0091] In one implementation, generating a 3D Gaussian representation based on the scene point cloud, the camera pose, and the rendered image includes:
[0092] The initial position is determined using the scene point cloud, the viewpoint of the predicted image is determined using the camera pose, and the rendered image is used as the target image for optimization. A three-dimensional Gaussian representation is generated based on the three-dimensional Gaussian sputtering algorithm.
[0093] In this embodiment, the initial pose of the camera can be locked in a pre-constructed point cloud map or scene point cloud using a scene point cloud combination algorithm. The camera pose can be used to determine parameters such as the viewing direction and field of view of the predicted image, thereby determining the viewpoint of the predicted image. A rendered image is used as the optimized target image for the predicted image. 3D Gaussian Splatting (3DGS) is an explicit 3D scene representation and neural rendering technique that uses a 3D Gaussian model to explicitly represent a 3D scene based on point clouds. A 3D Gaussian representation can be generated using the 3D Gaussian sputtering algorithm, replacing traditional discrete point or implicit volume rendering models to describe the 3D scene.
[0094] According to embodiments of this disclosure, a 3D Gaussian representation can be generated by using scene point clouds, camera poses, rendered images, and a 3D Gaussian algorithm. This can automate the process, adapt to the needs of multiple scenarios, and improve the efficiency of 3D scene reconstruction.
[0095] The 3D Gaussian representation generation method based on scene data proposed in this disclosure includes a method for automatically generating a 3D Gaussian scene based on input 3D scene data (including model, material, lighting, and scene floor plan structure). The overall process structure is as follows: Figure 7 As shown, the method includes the following steps:
[0096] 1.1 Overall Process
[0097] S701, Obtain raw scene data;
[0098] S702, Generate scene point cloud;
[0099] S703, Generate camera pose;
[0100] S704, Generate a renderable scene;
[0101] S705, Batch rendering of multiple views;
[0102] S706, Training Gaussian Scenario;
[0103] 1.2 Scene Point Cloud Generation
[0104] The process involves parsing all model data from the original 3D scene to obtain the triangular meshes and their 3D transformation matrices corresponding to the models. Multi-process parallel processing is used to sample and generate point clouds from each mesh. The 3D transformation matrix represents the globally localized geometric information of the model and can be converted into rotation, scaling, and translation transformation attributes. Since scaling changes the mesh area and sampling density, different scaling values require separate sampling of the mesh. After loading each mesh, the transformations of the model instance are grouped according to scaling.
[0105] Each scaling group first scales the mesh and estimates the number of points based on the surface area and the given point spacing. It then performs uniform surface sampling to obtain a batch of base points. Next, it applies rotation and translation to each instance in the group to reuse the sampled points and avoid duplicate sampling. Finally, it combines all the meshes into a scene point cloud.
[0106] 1.3 Camera Pose Generation
[0107] Based on the input scene floor plan (including room outlines and height information), a batch of camera poses are generated according to the configured strategy parameters. Camera position generation includes a combination of "space" and "surround" strategies.
[0108] 1. Space: Based on the camera spacing and room volume, calculate the number of camera points required for each room, and perform Poisson sampling within the space to ensure uniformity and randomness. Generate a six-way (six-sided cube) viewpoint for each camera point to ensure full coverage from multiple perspectives.
[0109] 2. Surround: Based on the room's planar outline, a circular curve is created by shrinking inwards by a certain distance. Several layers are then replicated according to the room's height, and a tilt angle is set for each layer (upper layers look down, lower layers look up). Camera points are evenly generated along each curve based on the camera spacing. The line of sight towards the room is calculated for each point based on the curve direction, and combined with the tilt angle, the final camera pose is generated.
[0110] 1.4 Obtaining the Rendered Image
[0111] All camera poses are synthesized into an animation sequence of the camera, and a renderable scene is constructed by combining the original scene data. The scene is then rendered frame by frame using an offline rendering engine to obtain multi-view rendered images that correspond one-to-one with the camera poses.
[0112] 1.5 3D Gaussian Training
[0113] The 3DGS (3D Gaussian Splatting) algorithm is implemented based on gsplat (an open-source CUDA-accelerated Gaussian rasterization library). The generated point cloud, camera pose, and rendered image are used as training inputs to automatically construct an optimization process and iteratively train to obtain a 3D Gaussian scene representation, thereby achieving high-quality 3D scene reconstruction.
[0114] Beneficial effects:
[0115] Fully automated pipeline: No human intervention required, from 3D scenes to trainable multi-view data is generated automatically.
[0116] High data consistency: The model, point cloud, camera, and rendered image strictly maintain spatial registration, improving the accuracy of Gaussian training.
[0117] Highly controllable: Point cloud sampling and camera distribution can be flexibly adjusted by setting parameters such as density and camera spacing.
[0118] High-quality training data: The offline rendering engine can generate training images with high resolution and realistic lighting effects, improving scene reproduction.
[0119] Figure 8 This is a flowchart illustrating a scene data-based 3D Gaussian representation generation apparatus 800 according to an embodiment of the present disclosure. The apparatus may include:
[0120] The point cloud generation module 810 is used to generate scene point clouds based on the triangular mesh corresponding to the model in the 3D scene data.
[0121] The pose generation module 820 is used to generate camera poses based on the apartment layout in the 3D scene data.
[0122] Image generation module 830 is used to generate a rendered image based on the camera pose and the 3D scene data;
[0123] Gaussian generation module 840 is used to generate a 3D Gaussian representation based on the scene point cloud, the camera pose, and the rendered image.
[0124] Figure 9 This is a schematic diagram of a three-dimensional Gaussian representation generation device 900 based on scene data according to another embodiment of the present disclosure. The device 900 includes: a point cloud generation module 910, a pose generation module 920, an image generation module 930, and a Gaussian generation module 940. The functions of these modules are the same as those of the modules in the three-dimensional Gaussian representation generation device based on scene data in the above embodiment. In one embodiment, the point cloud generation module 910 includes:
[0125] Submodule 911 is used to obtain the triangular mesh and 3D transformation matrix corresponding to the model.
[0126] The scaling submodule 912 is used to scale the triangular mesh corresponding to the model based on the scaling value of the 3D transformation matrix;
[0127] Point cloud submodule 913 is used to sample based on a scaled triangular mesh to obtain scene point clouds.
[0128] In one implementation, the point cloud submodule 913 is used to uniformly sample the surface of the scaled triangular mesh according to the surface area of the scaled triangular mesh and the given point spacing to obtain multiple base points; based on the multiple base points, multiple new points are obtained by calculating according to the rotation parameters and / or translation parameters in the three-dimensional transformation matrix; based on the multiple base points and the multiple new points, the point cloud of the scaling group corresponding to the scaling value is obtained; wherein, the point clouds of all scaling groups of a model constitute the point cloud of the model, and the point clouds of all models in the three-dimensional scene data constitute the scene point cloud.
[0129] In one embodiment, the pose generation module 920 includes:
[0130] The first acquisition submodule 921 is used to obtain the number of camera points required for the room based on the camera spacing and the volume of the room in the apartment structure;
[0131] The first determination submodule 922 is used to perform Poisson sampling in the room to determine the camera positions in the room.
[0132] The first pose module 923 is used to generate the camera view for each camera point and obtain the camera pose.
[0133] In one embodiment, the pose generation module 920 includes:
[0134] The second acquisition submodule 924 is used to acquire the annular curve of the room based on the planar outline of the room in the apartment structure.
[0135] The copy submodule 925 is used to copy the annular curve based on the height of the room to obtain multiple annular curves;
[0136] The second determining submodule 926 is used to determine the required camera positions for the room based on the camera spacing on the multi-layered circular curve;
[0137] The second pose submodule 927 is used to determine the camera viewpoint of each camera point on the annular curve layer based on the curve direction of each annular curve layer, and obtain the camera pose.
[0138] In one embodiment, the image generation module 930 includes:
[0139] Composition submodule 931 is used to synthesize camera animation sequences based on multiple camera poses;
[0140] Submodule 932 is constructed to build the scene to be rendered based on the animation sequence and the 3D scene data.
[0141] The rendering submodule 933 is used to render the scene to be rendered frame by frame using the rendering engine to obtain a rendered image corresponding to each camera pose.
[0142] In one implementation, the Gaussian generation module 940 is used to determine the initial position using the scene point cloud, determine the viewpoint of the predicted image using the camera pose, use the rendered image as the optimization target image, and generate a three-dimensional Gaussian representation based on a three-dimensional Gaussian sputtering algorithm.
[0143] Figure 10 This is a structural block diagram of an electronic device according to an embodiment of the present disclosure. Figure 10 As shown, the electronic device includes a memory 1010 and a processor 1020. The memory 1010 stores a computer program that can run on the processor 1020. The number of memories 1010 and processors 1020 can be one or more. The memory 1010 can store one or more computer programs, which, when executed by the electronic device, cause the electronic device to perform the methods provided in the above-described method embodiments. The electronic device may also include a communication interface 1030 for communicating with external devices and performing data exchange and transmission.
[0144] If the memory 1010, processor 1020, and communication interface 1030 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 10 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0145] Optionally, in a specific implementation, if the memory 1010, processor 1020 and communication interface 1030 are integrated on a single chip, then the memory 1010, processor 1020 and communication interface 1030 can communicate with each other through an internal interface.
[0146] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Machines (ARM) architecture.
[0147] Further, optionally, the aforementioned memory may include read-only memory and random access memory, and may also include non-volatile random access memory. The memory may be volatile or non-volatile, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. Many forms of RAM are available by way of example, but not limitation. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate Synchronous DRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct RAMBUS RAM (DR RAM).
[0148] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this disclosure is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line, DSL) or wireless (e.g., infrared, Bluetooth, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., Solid State Disks (SSDs)). It is worth noting that the computer-readable storage media mentioned in this disclosure can be non-volatile storage media; in other words, it can be non-transient storage media.
[0149] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0150] In the description of the embodiments of this disclosure, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.
[0151] In the description of the embodiments disclosed herein, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone.
[0152] In the description of embodiments of this disclosure, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more.
[0153] The above description is merely an exemplary embodiment of this disclosure and is not intended to limit this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the protection scope of this disclosure.
Claims
1. A method for generating a 3D Gaussian representation based on scene data, comprising: Generating a scene point cloud based on the triangular mesh corresponding to the model in 3D scene data includes: obtaining the triangular mesh corresponding to the model and a 3D transformation matrix; scaling the triangular mesh corresponding to the model based on the scaling value of the 3D transformation matrix; and sampling based on the scaled triangular mesh to obtain the scene point cloud. The camera pose is generated based on the apartment layout in the 3D scene data. A rendered image is generated based on the camera pose and the 3D scene data; A 3D Gaussian representation is generated based on the scene point cloud, the camera pose, and the rendered image; The scene point cloud is obtained by sampling based on the scaled triangular mesh, including: Based on the surface area of the scaled triangular mesh and the given point spacing, uniform surface sampling is performed on the scaled triangular mesh to obtain multiple base points; Based on the aforementioned base points, multiple new points are obtained by calculating according to the rotation parameters and / or translation parameters in the three-dimensional transformation matrix; Based on the multiple base points and the multiple newly added points, the point cloud of the scaling group corresponding to the scaling value is obtained; wherein, the point cloud of all scaling groups of a model constitutes the point cloud of the model, and the point cloud of all models in the 3D scene data constitutes the scene point cloud.
2. The method according to claim 1, wherein, Generate camera pose based on the apartment layout in the 3D scene data, including: Based on the camera spacing and the volume of the rooms in the apartment layout, the number of camera points required for the room is determined. Poisson sampling is performed within the room to determine the camera locations within the room; Generate a camera viewpoint for each camera location to obtain the camera pose.
3. The method according to claim 1, wherein, Generate camera pose based on the apartment layout in the 3D scene data, including: Based on the planar outline of the room in the apartment structure, obtain the annular curve of the room; Based on the height of the room, the annular curve is copied to obtain a multi-layer annular curve; Based on the camera spacing on the multi-layered circular curve, determine the required camera positions for the room; Based on the curve direction of each layer of the annular curve, the camera viewpoint of each camera point on that layer of the annular curve is determined, and the camera pose is obtained.
4. The method according to any one of claims 1 to 3, generating a rendered image based on the camera pose and the three-dimensional scene data, comprising: Animation sequences of cameras are synthesized based on the poses of multiple cameras; A scene to be rendered is constructed based on the animation sequence and the 3D scene data; The rendering engine is used to render the scene to be rendered frame by frame, resulting in a rendered image corresponding to each camera pose.
5. The method according to any one of claims 1 to 3, wherein, Generating a 3D Gaussian representation based on the scene point cloud, the camera pose, and the rendered image includes: The initial position is determined using the scene point cloud, the viewpoint of the predicted image is determined using the camera pose, the rendered image is used as the target image for optimization, and a three-dimensional Gaussian representation is generated based on the three-dimensional Gaussian sputtering algorithm.
6. A three-dimensional Gaussian representation generation device based on scene data, comprising: The point cloud generation module is used to generate scene point clouds based on the triangular mesh corresponding to the model in the 3D scene data. The point cloud generation module includes: an acquisition submodule for acquiring the triangular mesh and 3D transformation matrix corresponding to the model; a scaling submodule for scaling the triangular mesh corresponding to the model based on the scaling value of the 3D transformation matrix; and a point cloud submodule for sampling based on the scaled triangular mesh to obtain the scene point cloud. The pose generation module is used to generate camera poses based on the apartment layout in the 3D scene data. The image generation module is used to generate a rendered image based on the camera pose and the 3D scene data; The Gaussian generation module is used to generate a three-dimensional Gaussian representation based on the scene point cloud, the camera pose, and the rendered image; The point cloud submodule is used to uniformly sample the surface of the scaled triangular mesh according to the surface area of the scaled triangular mesh and the given point spacing to obtain multiple base points; based on the multiple base points, multiple new points are obtained by calculating according to the rotation parameters and / or translation parameters in the three-dimensional transformation matrix; based on the multiple base points and the multiple new points, the point cloud of the scaling group corresponding to the scaling value is obtained; wherein, the point clouds of all scaling groups of a model constitute the point cloud of the model, and the point clouds of all models in the three-dimensional scene data constitute the scene point cloud.
7. The apparatus according to claim 6, wherein, The pose generation module includes: The first acquisition submodule is used to acquire the number of camera points required for the room based on the camera spacing and the volume of the room in the apartment structure. The first determination submodule is used to perform Poisson sampling in the room to determine the camera positions in the room; The first pose module is used to generate the camera viewpoint for each camera point to obtain the camera pose.
8. The apparatus according to claim 6, wherein, The pose generation module includes: The second acquisition submodule is used to acquire the annular curve of the room based on the planar outline of the room in the apartment structure; The copying submodule is used to copy the annular curve based on the height of the room to obtain multiple annular curves; The second determining submodule is used to determine the required camera positions for the room based on the camera spacing on the multi-layered circular curve; The second pose submodule is used to determine the camera viewpoint of each camera point on the annular curve layer based on the curve direction of each annular curve layer, thereby obtaining the camera pose.
9. The apparatus according to any one of claims 6 to 8, wherein the image generation module comprises: The compositing submodule is used to synthesize camera animation sequences based on multiple camera poses; A construction submodule is used to construct the scene to be rendered based on the animation sequence and the 3D scene data; The rendering submodule is used to render the scene to be rendered frame by frame using the rendering engine to obtain a rendered image corresponding to each camera pose.
10. The apparatus according to any one of claims 6 to 8, wherein, The Gaussian generation module is used to determine the initial position using the scene point cloud, determine the viewpoint of the predicted image using the camera pose, use the rendered image as the optimization target image, and generate a three-dimensional Gaussian representation based on the three-dimensional Gaussian sputtering algorithm.
11. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-5.
13. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-5.