Visual optimization method and device for three-dimensional model

By assigning potential features to the three-dimensional Gaussian body model and training the semantic recognition model, the accuracy of the three-dimensional Gaussian body model in recognition and classification is solved, improving the authenticity of the rendering effect and improving the user experience.

CN120070716APending Publication Date: 2025-05-30BEIJING LANGJING INNOVATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411959477.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The three-dimensional Gaussian splashing technology is difficult to achieve high accuracy when identifying and classifying the categories of physical objects in scenes, resulting in greater differences in the rendering effect of the three-dimensional Gaussian body model than the real image, reducing the user's immersive experience.

Method used

By assigning potential features to each of the three-dimensional Gaussian bodies in the three-dimensional Gaussian body models of the target scene and determining projection information based on multiple shooting angles, the semantic recognition model is trained to optimize the semantic feature set, thereby improving the accuracy of recognition and classification.

Benefits of technology

It improves the accuracy of the three-dimensional Gaussian body model in recognition and classification, reduces the difference between rendering effects and real images, and improves the user's immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070716A_ABST
    Figure CN120070716A_ABST
Patent Text Reader

Abstract

The invention discloses a visual optimization method and device for a three-dimensional model. The method comprises the following steps: acquiring a three-dimensional Gaussian body model of a target scene; distributing a corresponding potential feature to each three-dimensional Gaussian body in the three-dimensional Gaussian body model; determining projection information of the three-dimensional Gaussian body model under each shooting view angle according to the plurality of shooting view angles; wherein the projection information of any shooting view angle comprises a projection image of the three-dimensional Gaussian body model under the shooting view angle and a potential feature set; inputting the projection image corresponding to each shooting view angle and the potential feature set into a semantic recognition model to obtain a prediction semantic graph corresponding to each shooting view angle; according to the prediction semantic graph and the reference semantic graph corresponding to each shooting view angle, optimizing the semantic recognition model to obtain an optimized semantic recognition model; and according to the optimized semantic recognition model, determining a semantic feature set corresponding to the three-dimensional Gaussian body model for a prediction semantic graph of the projection image corresponding to each shooting view angle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of three-dimensional modeling technology, and more specifically, to a method and device for visual optimization of three-dimensional models. Background Art

[0002] As a flexible and efficient representation method, 3D Gaussian Splatting (3DGS) can represent a scene more accurately by optimizing Gaussian coefficients, and combines the advantages of neural network-based optimization and explicit structured data storage. It can achieve faster training and real-time performance, especially for complex scenes and high-resolution outputs.

[0003] The 3DGS technology mainly relies on the geometric information of the three-dimensional Gaussian volume model to perceive and understand the three-dimensional scene. However, the geometric information can only provide the display attribute information such as the shape, size, and spatial position of the entity objects in the scene, and cannot provide the semantic features such as the category and function of the entity objects in the scene. This will lead to inaccurate classification when identifying and classifying the categories of entity objects in the scene through the three-dimensional Gaussian volume model of the scene, and further lead to a large difference between the rendering effect of the three-dimensional Gaussian volume model and the real image, reducing the immersive experience of users. Summary of the Invention

[0004] An object of an embodiment of the present invention is to provide a visual optimization method for a three-dimensional model, which can add the semantic features corresponding to each three-dimensional Gaussian body in the three-dimensional Gaussian volume model of the target scene, thereby improving the accuracy of identifying and classifying the categories to which the three-dimensional Gaussian bodies belong, improving the effect of the rendered image based on the three-dimensional Gaussian volume model, and improving the immersive experience of users.

[0005] According to a first aspect of the present invention, there is provided a method, which includes:

[0006] Obtain a three-dimensional Gaussian volume model of a target scene;

[0007] Assign potential features to each three-dimensional Gaussian body in the three-dimensional Gaussian volume model;

[0008] Determine the projection information of the three-dimensional Gaussian volume model at each of the plurality of shooting perspectives; wherein, the projection information of any one of the shooting perspectives includes the projection image of the three-dimensional Gaussian volume model at the shooting perspective and a set of potential features, and the set of potential features includes a plurality of potential features corresponding to a plurality of three-dimensional Gaussian bodies visible in the three-dimensional Gaussian volume model at the shooting perspective;

[0009] Input the projection image and the set of potential features corresponding to each shooting perspective into a semantic recognition model to obtain a predicted semantic map corresponding to each shooting perspective;

[0010] Optimize the semantic recognition model according to the predicted semantic map and the reference semantic map corresponding to each shooting perspective, to obtain an optimized semantic recognition model;

[0011] Determine the semantic feature set corresponding to the three-dimensional Gaussian volume model according to the predicted semantic map of the projection image corresponding to each shooting perspective by the optimized semantic recognition model.

[0012] Optionally, the step of inputting the projection image and the latent feature set corresponding to each shooting perspective into the semantic recognition model to obtain the predicted semantic map corresponding to each shooting perspective includes:

[0013] For any shooting perspective, determine the feature fusion image of the shooting perspective according to the projection image and the latent feature set of the shooting perspective, to obtain the feature fusion image of each shooting perspective;

[0014] Input the feature fusion image of each shooting perspective into the semantic recognition model to obtain the predicted semantic map corresponding to each shooting perspective.

[0015] Optionally, the step of determining the feature fusion image of the shooting perspective according to the projection image and the latent feature set of the shooting perspective includes:

[0016] For each pixel point in the projection image of the shooting perspective, determine the latent feature of the pixel point according to the latent features of at least one three-dimensional Gaussian volume corresponding to the pixel point in the latent feature set of the shooting perspective;

[0017] Associate the latent feature corresponding to each pixel point in the projection image of the shooting perspective with the pixel point to obtain the feature fusion image of the shooting perspective.

[0018] Optionally, the predicted semantic map includes multiple predicted semantic labels, and the reference semantic map includes multiple reference semantic labels. The step of optimizing the semantic recognition model according to the predicted semantic map and the reference semantic map corresponding to each shooting perspective to obtain an optimized semantic recognition model includes:

[0019] Construct a semantic loss function according to the multiple predicted semantic labels included in the predicted semantic map of each shooting perspective and the multiple reference semantic labels included in the reference semantic map; wherein, the semantic loss function is used to characterize the difference between the multiple predicted semantic labels included in the predicted semantic map of each shooting perspective and the multiple reference semantic labels included in the reference semantic map;

[0020] Optimize the semantic recognition model according to the semantic loss function to obtain an optimized semantic recognition model.

[0021] Optionally, determining the semantic feature set corresponding to the three-dimensional Gaussian volume model according to the predicted semantic map of the projection image corresponding to each shooting perspective by the optimized semantic recognition model includes:

[0022] Inputting the projection image and the latent feature set corresponding to each shooting perspective into the optimized semantic recognition model to obtain the predicted semantic map for each shooting perspective; wherein, the predicted semantic map includes the predicted semantic label corresponding to each pixel point.

[0023] For each pixel point in the predicted semantic map of each shooting perspective, determining at least one semantic feature of the three-dimensional Gaussian volume corresponding to the pixel point according to the predicted semantic label of the pixel point.

[0024] Obtaining the semantic feature set corresponding to the three-dimensional Gaussian volume model according to the semantic features corresponding to each three-dimensional Gaussian volume in the three-dimensional Gaussian volume model.

[0025] Optionally, the method further includes:

[0026] Inputting the projection image and the latent feature set corresponding to each shooting perspective into the light recognition model to obtain the predicted light map corresponding to each shooting perspective.

[0027] Optimizing the light recognition model according to the predicted light map and the reference light map corresponding to each shooting perspective to obtain an optimized light recognition model.

[0028] Determining the light feature set of the three-dimensional Gaussian volume model according to the light recognition result of the projection image corresponding to each shooting perspective by the optimized light recognition model.

[0029] Optionally, the light recognition model includes a diffuse reflection recognition model and a specular highlight recognition model. The step of inputting the projection image and the latent feature set corresponding to each shooting perspective into the light recognition model to obtain the predicted light map corresponding to each shooting perspective includes:

[0030] Inputting the projection image and the latent feature set corresponding to each shooting perspective into the diffuse reflection recognition model to obtain the predicted diffuse reflection map corresponding to each shooting perspective.

[0031] Inputting the projection image and the latent feature set corresponding to each shooting perspective into the specular highlight recognition model to obtain the predicted specular highlight map corresponding to each shooting perspective.

[0032] Determining the predicted light map corresponding to each shooting perspective according to the predicted specular highlight map and the predicted diffuse reflection map corresponding to each shooting perspective.

[0033] Optimizing the light recognition model according to the predicted light map and the reference light map corresponding to each shooting angle to obtain an optimized light recognition model includes:

[0034] Optimizing the diffuse reflection recognition model and the specular highlight recognition model according to the predicted light map and the reference light map corresponding to each shooting angle to obtain an optimized diffuse reflection recognition model and the specular highlight recognition model.

[0035] Optionally, inputting the projection image and the latent feature set corresponding to each shooting angle into the specular highlight recognition model to obtain the predicted specular highlight map corresponding to each shooting angle, including:

[0036] For any shooting angle, determining the direction head corresponding to the shooting angle according to the projection image and the latent feature set corresponding to the shooting angle to obtain the direction head of each shooting angle;

[0037] Inputting the direction head and the projection image corresponding to each shooting angle into the specular highlight recognition model to obtain the predicted specular highlight map corresponding to each shooting angle.

[0038] Optionally, determining the direction head corresponding to the shooting angle according to the projection image and the latent feature set corresponding to the shooting angle includes:

[0039] For each pixel point in the projection image, determining the latent feature of the pixel point according to the latent features of at least one three-dimensional Gaussian body corresponding to the pixel point in the latent feature set of the shooting angle;

[0040] Determining the direction head of the shooting angle according to the latent features corresponding to each pixel point in the projection image.

[0041] According to the second aspect of the present invention, there is also provided a visual optimization device for a three-dimensional model, which includes a memory and a processor, the memory is used to store executable instructions; the processor is used to operate according to the control of the instructions to execute the method described in the first aspect of the present invention.

[0042] One beneficial effect of the present invention is that by assigning latent features to each three-dimensional Gaussian body in the three-dimensional Gaussian body model of the target scene, the three-dimensional Gaussian body can be more comprehensively described, thereby deepening the system's understanding of the target scene. Furthermore, when rendering a two-dimensional image of a certain perspective based on the three-dimensional Gaussian body model, the rendering accuracy of the two-dimensional image can be improved. By determining the projection information of the three-dimensional Gaussian body model at each of the multiple shooting perspectives, and thus training the semantic recognition model, the accuracy of the semantic recognition of the image by the semantic recognition model can be improved. By determining the semantic feature set corresponding to the three-dimensional Gaussian body model according to the predicted semantic map of the projection image corresponding to each shooting perspective by the optimized semantic recognition model, the problem of inaccurate classification when identifying and classifying the categories of entity objects in the scene by the three-dimensional Gaussian body model of the scene can be avoided, the accuracy of semantic classification can be improved, and furthermore, the difference between the rendering effect of the three-dimensional Gaussian body model and the real image can be reduced, improving the user's immersive experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The drawings incorporated in and constituting a part of this specification illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.

[0044] Figure 1 is a schematic hardware structure diagram of a visual optimization device for a three-dimensional model according to an embodiment of the present invention;

[0045] Figure 2 is a schematic flowchart of a visual optimization method for a three-dimensional model according to an embodiment of the present invention;

[0046] Figure 3 is a schematic structural diagram of a visual optimization device for a three-dimensional model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] Various exemplary embodiments of the present invention will now be described in detail with reference to the drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present invention.

[0048] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended as a limitation on the present invention or its application or use.

[0049] Well-known technologies, methods, and devices in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be regarded as part of the specification.

[0050] In all the examples shown and discussed here, any specific values should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.

[0051] It should be noted that like reference numerals and letters denote like items in the following figures, and thus, once an item is defined in one figure, further discussion thereof is not required in subsequent figures.

[0052] <Hardware Configuration>

[0053] Figure 1 is a schematic structural diagram of a visual optimization device 100 for a three-dimensional model according to an embodiment of the present invention.

[0054] As Figure 1 shown, the visual optimization device 100 for a three-dimensional model may be any electronic device, such as a PC, a laptop, a server, etc.

[0055] In this embodiment, referring to Figure 1 shown, the visual optimization device 100 for a three-dimensional model may include a processor 1100, a memory 1200, an interface device 1300, a communication device 1400, a display device 1500, an input device 1600, a speaker 1700, a microphone 1800, and so on.

[0056] The processor 1100 may be a mobile version processor. The memory 1200 includes, for example, a ROM (Read Only Memory), a RAM (Random Access Memory), a non-volatile memory such as a hard disk, etc. The interface device 1300 includes, for example, a USB interface, a headphone interface, etc. The communication device 1400 can perform wired or wireless communication, for example. The communication device 1400 may include a short-range communication device, for example, any device that performs short-range wireless communication based on short-range wireless communication protocols such as the Hilink protocol, WiFi (IEEE 802.11 protocol), Mesh, Bluetooth, ZigBee, Thread, Z-Wave, NFC, UWB, LiFi, etc. The communication device 1400 may also include a remote communication device, for example, any device that performs WLAN, GPRS, 2G / 3G / 4G / 5G remote communication. The display device 1500 is, for example, a liquid crystal display screen, a touch display screen, etc. The display device 1500 is used to display the collected remote sensing images. The input device 1600 may include, for example, a touch screen, a keyboard, etc. The user can input / output voice information through the speaker 1700 and the microphone 2800.

[0057] In this embodiment, the memory 1200 of the three-dimensional model visual optimization device 100 is used to store instructions for controlling the processor 1100 to operate to at least execute the three-dimensional model visual optimization method according to any embodiment of the present invention. Those skilled in the art can design the instructions according to the solutions disclosed in the present invention. How the instructions control the processor to operate is well known in the art and will not be described in detail herein.

[0058] Although Figure 1 shows multiple devices of the three-dimensional model visual optimization device 100, however, the present invention may only relate to some of the devices. For example, the three-dimensional model visual optimization device 100 only relates to the memory 1200, the processor 1100, and the display device 1500.

[0059] In this embodiment, the three-dimensional model visual optimization device 100 implements the method according to any embodiment of the present invention based on the three-dimensional Gaussian volume model of the target scene to obtain a semantically optimized three-dimensional Gaussian volume model.

[0060] <Method Embodiment>

[0061] Figure 2 is a schematic flowchart of the three-dimensional model visual optimization method according to an embodiment of the present invention, and this method can be implemented by the three-dimensional model visual optimization device 2000.

[0062] According to Figure 2 shown, the three-dimensional model visual optimization method of this embodiment may include the following steps S2100 to S2600:

[0063] Step S2100, obtain the three-dimensional Gaussian volume model of the target scene.

[0064] In this embodiment, the target scene may be an urban scene, a natural landscape, an indoor environment, a historical site, etc., which is not limited herein.

[0065] The three-dimensional Gaussian volume model of the target scene may be constructed by the three-dimensional model visual optimization device or provided by a third-party platform, which is not limited herein.

[0066] The three-dimensional Gaussian volume model includes multiple three-dimensional Gaussian volumes, and each three-dimensional Gaussian volume in the multiple three-dimensional Gaussian volumes corresponds to a set of Gaussian volume parameters. This set of Gaussian volume parameters includes: the Gaussian volume center position coordinates, the scale parameter of the Gaussian volume, the color information of the Gaussian volume, and the rotation matrix of the Gaussian volume. Among them, the Gaussian volume center position coordinates are used to represent the position of the Gaussian volume in space. The scale parameter of the Gaussian volume is used to represent the width, height, and depth of the Gaussian volume, that is, the "size" of the Gaussian volume in three-dimensional space. The rotation matrix of the Gaussian volume is used to represent the direction of the Gaussian volume relative to the coordinate system. The color information of the Gaussian volume is used to represent the color of the Gaussian volume.

[0067] In some embodiments, the steps for the visual optimization device of the three-dimensional model to construct the three-dimensional Gaussian volume model of the target scene are as follows. That is, obtaining the three-dimensional Gaussian volume model of the target scene in step S2100 includes: steps S2100.1 to S2100.2.

[0068] Step S2100.1: Obtain a plurality of first images captured for the target scene.

[0069] In this embodiment, the plurality of first images are obtained by capturing the target scene from a plurality of shooting perspectives, and each shooting perspective in the plurality of shooting perspectives can correspond to at least one first image. Among them, the plurality of first images can be images captured for the target scene in advance, or images obtained through a third-party platform, and no limitation is made here.

[0070] The shooting perspective can be the perspective information of the camera when shooting the target scene.

[0071] In some examples, the perspective information includes camera pose information.

[0072] In some other examples, the perspective information can include information such as camera pose information, field of view angle, focal length, aperture, exposure parameters, etc.

[0073] Those skilled in the art should understand that no specific limitation is made on the perspective information here.

[0074] Step S2100.2: Construct the three-dimensional Gaussian volume model of the target scene according to the plurality of first images.

[0075] In this embodiment, the target scene is described by the three-dimensional Gaussian volume model obtained from the plurality of first images.

[0076] In some embodiments, constructing the three-dimensional Gaussian volume model of the target scene according to the plurality of first images in step S2100.2 includes: steps S2100.21 to S2100.22.

[0077] Step S2100.21: Construct a sparse point cloud of the target scene according to the camera pose information matched with each first image in the plurality of first images.

[0078] In this embodiment, the camera pose information can be the position and orientation of the camera in three-dimensional space, which is represented by a position vector (representing the position of the camera in the world coordinate system) and a rotation matrix (or Euler angles, quaternions, etc., representing the direction of the camera relative to the world coordinate system). The camera pose information is crucial for determining how the camera observes and records the target scene.

[0079] The sparse point cloud of the target scene is a set of key points extracted from multiple first images for characterizing the target scene.

[0080] This step can be implemented through feature point matching and Structure from Motion (SfM) algorithm. That is, the SfM algorithm can extract feature points from each of the multiple first images and calculate the camera pose information corresponding to the first image. Then, based on the camera pose information matched for each of the multiple first images, a three-dimensional sparse point cloud of the target scene is constructed.

[0081] Step S2100.21, construct a three-dimensional Gaussian volume model of the target scene according to the sparse point cloud.

[0082] In this embodiment, a three-dimensional Gaussian volume is initialized with the position of each point in the sparse point cloud as the center position of the three-dimensional Gaussian volume. And, in the process of initializing a three-dimensional Gaussian volume at the position of each point in the sparse point cloud, information such as the scale parameter of the Gaussian volume, the color information of the Gaussian volume, the selection attribute of the Gaussian volume, and the opacity of the Gaussian volume also needs to be considered.

[0083] The color information of the Gaussian volume can be obtained according to the color information of the point corresponding to the center position of the Gaussian volume in the sparse point cloud, or according to the color information extracted from the first image, or calculated according to a certain color estimation algorithm.

[0084] The scale parameter of the Gaussian volume is used to characterize the width, height, and depth of the Gaussian volume, that is, the "size" of the Gaussian volume in three-dimensional space. The scale parameter of the Gaussian volume can be fixed or dynamically adjusted according to the local density of the points in the sparse point cloud. For example, if the points in the sparse point cloud are relatively dense, a smaller-scale Gaussian volume can be used to represent the local structure more precisely. If the points in the sparse point cloud are relatively sparse, a larger-scale Gaussian volume can be used to cover a wider area.

[0085] The selection attribute of the Gaussian volume can be whether the user or the algorithm selects to include this Gaussian volume in visualization or further processing. In some applications, it may be necessary to decide whether to retain the Gaussian volume according to specific criteria (such as the scale parameter of the Gaussian volume, the center position of the Gaussian volume, or the degree of overlap with other Gaussian volumes).

[0086] The opacity of the Gaussian volume can be the opposite of the transparency of the Gaussian volume in visualization, which determines the visibility of the Gaussian volume during rendering. The opacity can be used to represent the uncertainty or importance of the Gaussian volume. For example, a Gaussian volume with a lower opacity may indicate a lower confidence level or importance.

[0087] After initializing a three-dimensional Gaussian volume at the position of each point in the sparse point cloud, a three-dimensional Gaussian volume model of the target scene can be obtained.

[0088] Step S2200: Assign latent features to each three-dimensional Gaussian volume in the three-dimensional Gaussian volume model.

[0089] In this embodiment, the three-dimensional Gaussian volume model obtained in step S2100 includes a plurality of three-dimensional Gaussian volumes, and each three-dimensional Gaussian volume includes a set of three-dimensional Gaussian volume parameters. This set of three-dimensional Gaussian volume parameters characterizes the explicit features of the three-dimensional Gaussian volume (for example, scale parameter, center position coordinates, rotation matrix, color information, etc.). These explicit features are directly and clearly defined features, which can be obtained through direct measurement or extraction from the three-dimensional model.

[0090] To enhance the comprehensive description of the target scene by the three-dimensional Gaussian volume model, latent features can be assigned to each three-dimensional Gaussian volume in the set. Among them, the latent features can be the intrinsic attributes and deep information of the entity objects in the target scene. They are the implicit and unobvious features in the three-dimensional Gaussian volume, and are related to the physical attributes or visual manifestations of the three-dimensional Gaussian volume.

[0091] The latent features can be the surface normal information, depth information, material properties (such as roughness, glossiness, reflectivity), occlusion and visibility between entity objects, shape features, functional information, etc. of the entity objects.

[0092] Among them, the surface normal information can describe the direction of the surface of the entity object at each point. It can be estimated through three-dimensional reconstruction technology or from a high-resolution depth map, or predicted from low-resolution data by training a depth network. The depth information can be the distance between the surface of the entity object and the camera or a certain reference plane, which can be predicted from the images taken of the target scene by a deep learning model. The material properties can be the physical properties of the surface of the entity object, such as roughness, glossiness, reflectivity, scattering rate, etc. These material properties can be extracted from multiple images of the captured target scene by using a machine learning model. The occlusion and visibility between entity objects can be the occlusion and visibility information between entity objects in the target scene, which can be simulated by graphics rendering technology or inferred from complex scenes by using a deep learning model. The shape features can be the geometric shape features of the entity object, such as curvature, concavity and convexity, etc. They are analyzed from the three-dimensional model of the target scene through geometric processing algorithms or extracted from the images taken of the target scene by using a deep learning model. The functional information can be the functional description of the entity object. For example, a "chair" is used for "sitting". It requires combining domain knowledge and semantic understanding and may involve natural language processing (NLP) technology.

[0093] The latent features include multiple types of information and are a high-dimensional feature.

[0094] In one example, the latent features can be a 32-dimensional vector.

[0095] Step S2300: Determine the projection information of the three-dimensional Gaussian volume model at each of the multiple shooting perspectives.

[0096] In this embodiment, the three-dimensional Gaussian volume model is projected onto the image plane of each shooting perspective among the multiple shooting perspectives to obtain the projection information of the three-dimensional Gaussian volume model at each shooting perspective.

[0097] The process of projecting the three-dimensional Gaussian volume model onto the image plane of each shooting perspective among the multiple shooting perspectives is essentially to convert the three-dimensional spatial coordinates of the multiple three-dimensional Gaussian volumes included in the three-dimensional Gaussian volume model into corresponding two-dimensional pixel coordinates.

[0098] Here, the process of converting the three-dimensional spatial coordinates of a three-dimensional Gaussian volume into the two-dimensional pixel coordinates of this three-dimensional Gaussian volume is as follows: First, through the camera pose information (i.e., the extrinsic parameters of the camera), the three-dimensional space of the three-dimensional Gaussian volume is converted from the world coordinate system to the camera coordinate system to obtain the coordinates of the three-dimensional Gaussian volume in the camera coordinate system. Then, through the intrinsic parameters of the camera, the coordinates of the three-dimensional Gaussian volume in the camera coordinate system are projected onto the image plane to obtain the two-dimensional pixel coordinates of this three-dimensional Gaussian volume.

[0099] After a three-dimensional Gaussian volume model is projected onto the image plane of a shooting perspective, the projection information corresponding to this shooting perspective can be obtained. Among them, the projection information of any shooting perspective includes the projection image of the three-dimensional Gaussian volume model at this shooting perspective and a set of latent features. The set of latent features includes multiple latent features corresponding to the multiple three-dimensional Gaussian volumes visible in this shooting perspective.

[0100] For any shooting perspective, there is the following relationship between the projection image corresponding to this shooting perspective and the set of latent features: Any latent feature included in the set of latent features of this shooting perspective corresponds to a three-dimensional Gaussian volume, and the center position of this three-dimensional Gaussian volume corresponds to a pixel point on the projection image of this shooting perspective. Moreover, the pixel points on the projection image where the center positions of different three-dimensional Gaussian volumes are located can be coincident or non-coincident. That is, for each pixel point on the projection image of any shooting perspective, it can correspond to at least one latent feature in the set of latent features of this shooting perspective.

[0101] In one example, the three-dimensional Gaussian volume model can be projected onto the image plane of a shooting perspective through 3DGS technology to obtain the projection information of this shooting perspective.

[0102] In some embodiments, determining the projection information of the three-dimensional Gaussian volume model at each of the multiple shooting perspectives in step S2300 includes: step S2300.1 and step S2300.2.

[0103] Step S2300.1: Determine the projected image of the three-dimensional Gaussian volume model at each of the shooting perspectives through the first rendering technique.

[0104] Step S2300.2: Determine the set of potential features of the three-dimensional Gaussian volume model at each of the shooting perspectives through the second rendering technique.

[0105] In this embodiment, first, the set of projected Gaussian volumes of the three-dimensional Gaussian volume model at each shooting perspective can be determined through the second rendering technique. Among them, the set of projected Gaussian volumes of any shooting perspective includes multiple three-dimensional Gaussian volumes visible in this shooting perspective. Then, according to the potential features corresponding to each three-dimensional Gaussian volume in the set of projected Gaussian volumes of any shooting perspective, the set of potential features of this shooting perspective is determined.

[0106] The first rendering technique can be, for example, the three-dimensional Gaussian splatting (3DGS) technique, and the second rendering technique can be, for example, the Alpha Rendering technique. The Alpha Rendering technique can refer to using the alpha channel (transparency channel) to process the transparency and semi-transparency effects of objects during three-dimensional rendering. Through this Alpha Rendering technique, the computational amount during rendering can be reduced. Through the 3DGS rendering technique, the quality of the projected image can be improved. By rendering in these two ways, it is possible to reduce the computational amount during the rendering process while improving the quality of the projected image.

[0107] Step S2400: Input the projected image and the set of potential features corresponding to each shooting perspective into the semantic recognition model to obtain the predicted semantic map corresponding to each shooting perspective.

[0108] In this embodiment, the semantic recognition model can be, for example, an untrained CNN network. The semantic recognition model can perform semantic recognition on the projected image, that is, identify and distinguish the categories (such as people, vehicles, buildings, etc.) to which each pixel point in the projected image belongs, and output the predicted semantic map. Among them, the resolution of the predicted semantic map is the same as that of the projected image, that is, the pixel points of the projected image and the pixel points of the predicted semantic map are completely aligned.

[0109] Each pixel point in the predicted semantic map corresponds to a predicted semantic label, and the predicted semantic label corresponding to any pixel point is used to represent the predicted category of the semantic recognition model for this pixel point.

[0110] In some embodiments, in step S2400, inputting the projected image and the set of potential features corresponding to each shooting perspective into the semantic recognition model to obtain the predicted semantic map corresponding to each shooting perspective includes steps S2400.1 to S2400.2.

[0111] Step S2400.1. For any shooting perspective, determine the feature fusion image of the shooting perspective according to the projection image and the set of potential features of the shooting perspective, so as to obtain the feature fusion image of each shooting perspective.

[0112] For example, for the A shooting perspective, fuse the set of potential features corresponding to this shooting perspective into the projection image corresponding to this shooting perspective to obtain the feature fusion image of this shooting perspective.

[0113] In some embodiments, in step S2400.1, determining the feature fusion image of the shooting perspective according to the projection image and the set of potential features of the shooting perspective includes: step S2400.11 and step S2400.12.

[0114] Step S2400.11. For each pixel point in the projection image of the shooting perspective, determine the potential feature of the pixel point according to the potential features of at least one three-dimensional Gaussian body corresponding to the pixel point in the set of potential features of the shooting perspective.

[0115] In this embodiment, each pixel point in the projection image corresponds to at least one three-dimensional Gaussian body, and the at least one three-dimensional Gaussian body corresponds to at least one potential feature. That is to say, each pixel point in the projection image corresponds to at least one potential feature. At this time, if a pixel point corresponds to one potential feature, then use this potential feature as the potential feature of the pixel point. If a pixel point corresponds to multiple potential features, then the multiple potential features can be weighted and averaged to obtain an average potential feature as the potential feature of the pixel point.

[0116] Step S2400.12. Associate the potential feature corresponding to each pixel point in the projection image of the shooting perspective to obtain the feature fusion image of the shooting perspective.

[0117] In this embodiment, the potential feature can be associated with the pixel point of its corresponding projection image to obtain the feature fusion image. That is to say, associate the potential feature corresponding to each pixel point in the projection image to obtain the feature fusion image of this shooting perspective.

[0118] Step S2400.2. Input the feature fusion image of each shooting perspective into the semantic recognition model to obtain the predicted semantic map corresponding to each shooting perspective.

[0119] In this embodiment, the semantic recognition model predicts the category to which each pixel point in the input feature fusion map belongs to obtain the predicted semantic map.

[0120] Step S2500. Optimize the semantic recognition model according to the predicted semantic map and the reference semantic map corresponding to each shooting perspective to obtain the optimized semantic recognition model.

[0121] In this embodiment, the reference semantic map of any shooting perspective is an image that shoots the target scene from this shooting perspective and labels the actual categories to which the captured entity objects belong. That is to say, the reference semantic map includes the reference semantic labels corresponding to each pixel point, and the reference semantic label corresponding to any pixel point is used to represent the actual category to which this pixel point belongs. The predicted semantic map of any shooting perspective includes the predicted semantic labels of each pixel point, and the predicted semantic label is determined according to the predicted semantic map of this pixel point by the semantic recognition model. Moreover, the pixel distributions in the predicted semantic map and the reference semantic map corresponding to any shooting perspective are aligned, that is, the pixel points in the predicted semantic map of any shooting perspective and the pixel points in the reference semantic map are in one-to-one correspondence.

[0122] By optimizing the semantic recognition model, the difference between the predicted semantic map and the reference semantic map of each shooting perspective can be reduced, so as to obtain an optimized semantic recognition model.

[0123] In some embodiments, the predicted semantic map includes multiple predicted semantic labels. Among them, for each pixel point in the predicted semantic map, there corresponds a predicted semantic label. The reference semantic map includes multiple reference semantic labels. Among them, for each pixel point in the reference semantic map, there corresponds a reference semantic label. The pixel points of the reference semantic map and the predicted semantic map are in one-to-one correspondence.

[0124] In these embodiments, in step S2500, according to the predicted semantic map and the reference semantic map corresponding to each shooting perspective, optimizing the semantic recognition model to obtain an optimized semantic recognition model includes: step S2500.1 and step S2500.2.

[0125] Step S2500.1, according to the multiple predicted semantic labels included in the predicted semantic map of each shooting perspective and the multiple reference semantic labels included in the reference semantic map, construct a semantic loss function.

[0126] In this embodiment, the semantic loss function is used to represent the difference between the multiple predicted semantic labels included in the predicted semantic map of each shooting perspective and the multiple reference semantic labels included in the reference semantic map.

[0127] The process of constructing the semantic loss functions of multiple shooting perspectives can be: according to the multiple predicted semantic labels included in the predicted semantic map of any shooting perspective and the multiple reference semantic labels included in the reference semantic map, construct the semantic loss function corresponding to this shooting perspective. Then add the semantic loss functions corresponding to multiple shooting perspectives to obtain the total semantic loss function of multiple shooting perspectives as the final semantic loss function.

[0128] In one example, the semantic loss function can be the Cross-Entropy Loss function. It can measure the difference between the probability distribution of the predicted semantic labels of the semantic recognition model and the probability distribution of the reference semantic labels.

[0129] Since the entity objects in the target scene may include multiple types (e.g., people, cars, animals, etc.), that is, the recognition of their semantics belongs to a multi-classification problem. At this time, the cross-entropy loss can be calculated for each entity object category and then averaged as the semantic loss function.

[0130] In one example, the calculation formula for the cross-entropy loss of multiple entity object categories is as follows:

[0131]

[0132] where C is the number of categories of entity objects, y is the reference semantic label of the i-th entity object category (1 if it is the i-th category, otherwise 0), and p i is the probability that the model predicts the i-th entity object category.

[0133] Step S2500.2, optimize the semantic recognition model according to the semantic loss function to obtain an optimized semantic recognition model.

[0134] In this embodiment, during the process of optimizing the semantic recognition model, according to the value of the semantic loss function calculated each time, the gradient of the semantic loss function with respect to the model parameters of the semantic recognition model is calculated through the backpropagation algorithm. According to the calculated gradient, the optimization algorithm (such as gradient descent) is used to update the model parameters to optimize the semantic recognition model.

[0135] Step S2600, determine the semantic feature set corresponding to the three-dimensional Gaussian volume model according to the predicted semantic map of the projection image corresponding to each shooting view by the optimized semantic recognition model.

[0136] In this embodiment, through the optimized semantic recognition model, the projection images of each shooting view among multiple shooting views are re-recognized to obtain a predicted semantic map (also called an optimized semantic map) corresponding to each shooting view. For each pixel point in the predicted semantic map (also called an optimized semantic map) of each shooting view, according to the predicted semantic label of the pixel point, the semantic feature of the three-dimensional Gaussian volume corresponding to the pixel point in the three-dimensional Gaussian volume model is determined, so as to obtain the semantic feature set corresponding to the three-dimensional Gaussian volume model. Among them, the semantic feature set includes the semantic features corresponding to each three-dimensional Gaussian volume in the three-dimensional Gaussian volume model.

[0137] In some embodiments, in step S2600, determining the semantic feature set corresponding to the three-dimensional Gaussian volume model according to the predicted semantic map of the projection image corresponding to each shooting view by the optimized semantic recognition model includes: step S2600.1 and step S2600.3.

[0138] Step S2600.1, inputting the projection image and the latent feature set corresponding to each shooting view into the optimized semantic recognition model to obtain the predicted semantic map of each shooting view.

[0139] In this embodiment, the predicted semantic map includes the predicted semantic label corresponding to each pixel point.

[0140] This step is basically the same as the above step S2400 and will not be elaborated here.

[0141] Step S2600.2, for each pixel point in the predicted semantic map of each shooting view, determining at least one semantic feature of the three-dimensional Gaussian volume corresponding to the pixel point according to the predicted semantic label of the pixel point.

[0142] For example, if the predicted semantic label of a certain pixel point in the predicted semantic map is a vehicle, the semantic features of the two three-dimensional Gaussian volumes corresponding to this pixel point are vehicles.

[0143] Step S2600.3, obtaining the semantic feature set corresponding to the three-dimensional Gaussian volume model according to the semantic features corresponding to each three-dimensional Gaussian volume in the three-dimensional Gaussian volume model.

[0144] In this embodiment, after performing semantic annotation on each three-dimensional Gaussian volume in the three-dimensional Gaussian volume model according to the predicted semantic maps corresponding to multiple shooting views, the semantic feature set corresponding to the three-dimensional Gaussian volume model can be obtained.

[0145] In some examples, the semantic feature set can be encoded into the three-dimensional Gaussian volume model.

[0146] In this example, the semantic features corresponding to each three-dimensional Gaussian volume in the three-dimensional Gaussian volume model can be encoded to obtain a semantic three-dimensional Gaussian volume model. The semantic three-dimensional Gaussian volume model not only contains geometric information but also rich semantic features, enabling the three-dimensional Gaussian volume model to more comprehensively describe the entity objects in the target scene. Furthermore, when performing semantic segmentation or entity object recognition, the recognition ability and accuracy of the three-dimensional Gaussian model can be improved.

[0147] In some examples, the semantic feature set may not be encoded into the three-dimensional Gaussian volume model

[0148] In this example, it is only necessary to extract semantic features from a three-dimensional Gaussian volume model (i.e., a three-dimensional Gaussian model), without directly adding the semantic features to the three-dimensional Gaussian model. In this case, the semantic features can be used for auxiliary tasks such as navigation, path planning, or scene understanding, without changing the structure of the model.

[0149] According to an embodiment of the present application, by assigning latent features to each three-dimensional Gaussian volume in the three-dimensional Gaussian volume model of the target scene, the three-dimensional Gaussian volume can be more comprehensively described, thereby deepening the system's understanding of the target scene. Furthermore, when rendering a two-dimensional image from a certain perspective based on the three-dimensional Gaussian volume model, the rendering accuracy of the two-dimensional image can be improved. By determining the projection information of the three-dimensional Gaussian volume model at each of the multiple shooting perspectives, and thus training a semantic recognition model, the accuracy of the semantic recognition of the image by the semantic recognition model can be improved. By determining the semantic feature set corresponding to the three-dimensional Gaussian volume model according to the predicted semantic map of the projection image corresponding to each shooting perspective by the optimized semantic recognition model, the problem of inaccurate classification when identifying and classifying the categories of entity objects in the scene by the three-dimensional Gaussian volume model of the scene can be avoided, the accuracy of semantic classification can be improved, and furthermore, the difference between the rendering effect of the three-dimensional Gaussian volume model and the real image can be reduced, improving the user's immersive experience.

[0150] In order to further improve the rendering effect of the three-dimensional Gaussian volume model and obtain high-quality rendered images, the inventors also provide an embodiment of how to determine the illumination feature set corresponding to the three-dimensional Gaussian volume model, which is as follows:

[0151] In some embodiments, the method further includes: step S3100 to step S3300.

[0152] Step S3100, input the projection image corresponding to each shooting perspective and the latent feature set into an illumination recognition model to obtain the predicted illumination map corresponding to each shooting perspective.

[0153] In this embodiment, the illumination recognition model can be a diffuse reflection recognition model, a specular highlight recognition model, a refractive index recognition model, etc.

[0154] Those skilled in the art should understand that the illumination recognition model can be any model used to characterize the interaction between light and entity objects. For example, a transparency recognition model, a roughness recognition model, etc., are not limited here.

[0155] Correspondingly, the predicted illumination map can be a diffuse reflection map, a specular highlight map, a refractive index map, a transparency map, etc., which are also not limited here.

[0156] In some embodiments, the lighting features include diffuse reflection features and specular highlight features. Correspondingly, the lighting recognition model includes a diffuse reflection recognition model and a specular highlight recognition model.

[0157] In this embodiment, the diffuse reflection features can be used to characterize the color, texture, and surface roughness of the surface of the three-dimensional Gaussian volume. The specular highlight features can be used to characterize the specular reflection characteristics of the three-dimensional Gaussian volume.

[0158] In these embodiments, in step S3100, inputting the projection image and the potential feature set corresponding to each of the shooting perspectives into the lighting recognition model to obtain the predicted lighting map corresponding to each shooting perspective includes:

[0159] Step S3100.1, inputting the projection image and the potential feature set corresponding to each of the shooting perspectives into the diffuse reflection recognition model to obtain the predicted diffuse reflection map corresponding to each shooting perspective.

[0160] In this embodiment, the predicted diffuse reflection map of any shooting perspective refers to the diffuse reflection effect diagram of the three-dimensional Gaussian volume model predicted by the diffuse reflection recognition model at this shooting perspective.

[0161] The pixel points of the predicted diffuse reflection map and the pixel points of the projection image are in one-to-one correspondence, and the pixel value of each pixel point in the predicted diffuse reflection map is the predicted diffuse reflection value of this pixel point.

[0162] Diffuse reflection refers to the phenomenon that after light irradiates the surface of an object, it is scattered evenly in all directions.

[0163] The diffuse reflection recognition model can be, for example, a diffuse reflection U-Net network.

[0164] Those skilled in the art should understand that the diffuse reflection U-Net network is a relatively existing network and will not be elaborated here.

[0165] In some embodiments, in step S3100.1, inputting the projection image and the potential feature set corresponding to each of the shooting perspectives into the diffuse reflection recognition model to obtain the predicted diffuse reflection map corresponding to each shooting perspective includes: steps S3100.11 to S3100.12.

[0166] Step S3100.11, for any shooting perspective, determining the feature fusion image of this shooting perspective according to the projection image and the potential feature set of this shooting perspective to obtain the feature fusion image corresponding to each shooting perspective.

[0167] This step is the same as step S2400.1 above and will not be elaborated here.

[0168] Step S3100.12: Input the feature fusion image of each shooting perspective into the diffuse reflection recognition model to obtain the predicted diffuse reflection map corresponding to each shooting perspective.

[0169] In this embodiment, the pixel numbers of the feature fusion image input to the diffuse reflection recognition model and the predicted diffuse reflection map are the same.

[0170] Step S3100.2: Input the projection image and the potential feature set corresponding to each shooting perspective into the specular highlight recognition model to obtain the predicted specular highlight map corresponding to each shooting perspective.

[0171] In this embodiment, the pixel points of the predicted specular highlight map and the projection image are in one-to-one correspondence, and the pixel value of each pixel point in the predicted specular highlight map is the predicted specular highlight intensity value corresponding to this pixel point. The specular highlight intensity value represents the specular reflection effect of the three-dimensional Gaussian body corresponding to this pixel point under specific lighting conditions, which is related to the glossiness of the surface of the three-dimensional Gaussian body and the lighting direction. The higher the predicted specular highlight intensity value, the stronger the specular highlight effect of this pixel point.

[0172] The specular highlight recognition model can be, for example, a specular highlight CNN network. The specular highlight recognition model is used to identify the specular highlight area in the input image.

[0173] In some embodiments, in step S3100.2, inputting the projection image and the potential feature set corresponding to each shooting perspective into the specular highlight recognition model to obtain the predicted specular highlight map corresponding to each shooting perspective includes: step S3100.21 and step S3100.22.

[0174] Step S3100.21: For any shooting perspective, determine the direction head corresponding to the shooting perspective according to the projection image and the potential feature set corresponding to the shooting perspective, to obtain the direction head of each shooting perspective.

[0175] In this embodiment, the direction head of the shooting perspective is used to characterize the potential features corresponding to each pixel point in the projection image under this shooting perspective.

[0176] In some embodiments, in step S3100.21, determining the direction head corresponding to the shooting perspective according to the projection image and the potential feature set corresponding to the shooting perspective includes: step SA1 and step SA2.

[0177] Step SA1: For each pixel point in the projection image, determine the potential feature of the pixel point according to the potential features of at least one three-dimensional Gaussian body corresponding to the pixel point in the potential feature set of the shooting perspective.

[0178] This step is the same as the above step S2400.11 and will not be elaborated here.

[0179] Step SA2: Determine the direction vector of the shooting angle according to the potential features corresponding to each pixel point in the projection image.

[0180] In this embodiment, for any shooting angle, cascade the potential features corresponding to each pixel point in the projection image of this shooting angle with this shooting angle to obtain the direction vector of this shooting angle.

[0181] Step S3100.22: Input the direction vector and the projection image corresponding to each shooting angle into the specular highlight recognition model to obtain the predicted specular highlight map corresponding to each shooting angle.

[0182] Step S3100.3: Determine the predicted illumination map corresponding to each shooting angle according to the predicted specular highlight map and the predicted diffuse reflection map corresponding to each shooting angle.

[0183] In this embodiment, since the pixel points of the predicted specular highlight map and the pixel points of the projection image are in one-to-one correspondence, and the pixel points of the predicted diffuse reflection map and the pixel points of the projection image are in one-to-one correspondence, the pixel points of the predicted specular highlight map and the pixel points of the predicted diffuse reflection map are also in one-to-one correspondence. Therefore, the pixel values of the corresponding pixel points of the predicted diffuse reflection map and the predicted specular highlight map can be added to obtain the predicted illumination map. That is, for the corresponding pixel points in the predicted specular highlight map and the predicted diffuse reflection map of any shooting angle, according to the predicted diffuse reflection value and the predicted specular highlight intensity value of this pixel point, obtain the predicted illumination value of this pixel point, so as to obtain the predicted illumination map of this shooting angle. Among them, the predicted illumination value of each pixel point in the predicted illumination map is the sum of the predicted specular highlight intensity value corresponding to this pixel point and the predicted diffuse reflection value corresponding to this pixel point.

[0184] Step S3200: Optimize the illumination recognition model according to the predicted illumination map and the reference illumination map corresponding to each shooting angle to obtain the optimized illumination recognition model.

[0185] In this embodiment, an illumination loss function can be constructed according to the predicted illumination map and the reference illumination map corresponding to each shooting angle. Then, according to this illumination loss function, optimize the illumination recognition model to obtain the optimized illumination recognition model.

[0186] In an embodiment where the illumination recognition model includes a diffuse reflection recognition model and a specular highlight recognition model, step S3200 of optimizing the illumination recognition model according to the predicted illumination map and the reference illumination map corresponding to each shooting angle to obtain the optimized illumination recognition model includes:

[0187] Optimize the diffuse reflection recognition model and the specular highlight recognition model according to the predicted illumination map and the reference illumination map corresponding to each shooting angle to obtain the optimized diffuse reflection recognition model and the specular highlight recognition model.

[0188] In this embodiment, the light loss function can be the total loss function of the L1 pixel loss and the SSIM structural similarity index. Among them, the L1 pixel loss, that is, the L1 loss (also known as the absolute error), represents the sum of the absolute values of the differences in pixel values between the reference light map and the predicted light map. By reducing the L1 loss, the pixel values of the predicted light map can be made as close as possible to those of the reference light map. The SSIM structural similarity index is an index for measuring the structural similarity between the predicted light map and the reference light map based on the brightness, contrast, and structural information of the image.

[0189] Use the backpropagation algorithm to calculate the gradient of the total loss function with respect to the model parameters (including the model parameters of the specular highlight recognition model and the diffuse reflection recognition model). Update the model parameters of the specular highlight recognition model and the diffuse reflection recognition model according to the calculated gradient and the optimization algorithm (such as gradient descent).

[0190] Step S3300, determine the light feature set of the three-dimensional Gaussian volume model according to the light recognition result of the optimized light recognition model for the projection image corresponding to each shooting angle.

[0191] In this embodiment, through the optimized light recognition model, re-recognize the projection images of each shooting angle among multiple shooting angles to obtain the predicted light map (also called the optimized light map) corresponding to each shooting angle. For each pixel point in the predicted light map (also called the optimized light map) of each shooting angle, determine the light feature of the three-dimensional Gaussian volume corresponding to the pixel point in the three-dimensional Gaussian volume model according to the predicted light value of the pixel point, so as to obtain the light feature set corresponding to the three-dimensional Gaussian volume model. Among them, the light feature set includes the light features corresponding to each three-dimensional Gaussian volume in the three-dimensional Gaussian volume model.

[0192] In some examples, the light feature set can be encoded into the three-dimensional Gaussian volume model.

[0193] In some other examples, the light feature set may not be encoded into the three-dimensional Gaussian volume model.

[0194] In order to verify the effect of the method in the above embodiments of the present application after the visual optimization of the three-dimensional model, the inventor verified the method of an example of the present application through the Replica dataset. Among them, Replica is a dataset for highly realistic 3D indoor scene reconstruction containing 18 rooms and buildings. Each scene consists of a dense grid, high-resolution high-dynamic range (HDR) textures, semantic class and instance information for each primitive, planar mirrors, and glass reflectors. The Replica dataset is dedicated to the study of world generation models that rely on visually, geometrically, and semantically realistic scenes.

[0195] According to an example of a method, the implementation steps on the Replica dataset are as follows:

[0196] Step S1. Initialize a three-dimensional Gaussian volume model using the sparse point cloud provided by the Replica dataset.

[0197] Step S2. Assign potential features to each three-dimensional Gaussian volume in the three-dimensional Gaussian volume model.

[0198] Step S3. Determine the projection information of the three-dimensional Gaussian volume model at each shooting angle according to multiple shooting angles, where the projection information of any shooting angle includes the projection image and the set of potential features of the three-dimensional Gaussian volume model at this shooting angle.

[0199] Step S4. For any shooting angle, determine the direction head corresponding to this shooting angle according to the projection image and the set of potential features corresponding to this shooting angle, and obtain the direction head of each shooting angle.

[0200] Step S5. Input the direction head and the projection image corresponding to each shooting angle into the highlight recognition model to obtain the predicted highlight map corresponding to each shooting angle.

[0201] Step S6. Input the projection image and the set of potential features corresponding to each shooting angle into the diffuse reflection recognition model to obtain the predicted diffuse reflection map corresponding to each shooting angle.

[0202] Step S7. Determine the predicted illumination map corresponding to each shooting angle according to the predicted highlight map and the predicted diffuse reflection map corresponding to each shooting angle.

[0203] Step S8. Input the projection image and the set of potential features corresponding to each shooting angle into the semantic recognition model to obtain the predicted semantic map corresponding to each shooting angle.

[0204] Step S9. Construct an illumination loss function according to the predicted illumination map and the original image in the dataset, and update the highlight recognition model and the diffuse reflection recognition model according to the illumination loss function to obtain the optimized diffuse reflection recognition model and highlight recognition model.

[0205] Step S10. Determine the diffuse reflection features and highlight features of the three-dimensional Gaussian volume model according to the recognition results of the optimized diffuse reflection recognition model and highlight recognition model for the projection images of multiple shooting angles.

[0206] Step S11. Construct a semantic loss function according to the predicted semantic map and the original semantic map, and update the semantic recognition model according to the semantic loss function to obtain the optimized semantic recognition model.

[0207] Step S12. Determine the semantic features of the three-dimensional Gaussian volume model according to the recognition results of the optimized semantic recognition model for the projection images of multiple shooting perspectives.

[0208] Step S13. Encode the semantic features, diffuse reflection features, and specular highlights features of the three-dimensional Gaussian volume model into the three-dimensional Gaussian volume model to obtain the encoded three-dimensional Gaussian volume model.

[0209] Step S14. Render the encoded three-dimensional Gaussian volume model through the test set perspective to obtain the test render image.

[0210] Step S15. Input the test render image into the optimized semantic recognition model to obtain the predicted semantic map.

[0211] In this step, after obtaining the predicted semantic map, the semantic segmentation effect can be evaluated in combination with the original semantic map in the test set. The mIoU (mean Intersection over Union) index of the predicted semantic segmentation map (i.e., the predicted semantic map) is 48.3%: it calculates the average of the IoU values of all entity object categories. The IoU value can be the degree of overlap between the predicted segmentation region and the true annotation region.

[0212] The mAcc (mean Accuracy) index is 65.8%, that is to say, the overall classification accuracy of the semantic recognition model on the test set is 65.8%, indicating that the semantic recognition model performs well in identifying the object categories in the image.

[0213] Step S16. Input the test render image into the optimized diffuse reflection recognition model to obtain the predicted diffuse reflection map.

[0214] Step S17. Input the test render image into the optimized specular highlights recognition model to obtain the predicted specular highlights map.

[0215] Step S18. Add the pixel values of the predicted diffuse reflection map obtained in Step S16 and the predicted specular highlights map obtained in Step S17 to obtain the predicted illumination map.

[0216] The SSIM of the predicted illumination map obtained in this step is 0.88, the mIoU index of the semantic map is 48.3, the mAcc index is 65.8, and the PSNR index is 32.37.

[0217] The PSNR value (Peak Signal-to-Noise Ratio) can be the reconstruction quality or distortion degree of the image. It is calculated by comparing the original image and the predicted illumination map in the test set. The higher the PSNR value, the better the quality of the predicted illumination map and the smaller the distortion.

[0218] A PSNR value of 32.37 indicates that the difference between the predicted illumination map and the original image in the test set is within a certain range.

[0219] <Embodiment of the storage medium>

[0220] An embodiment of the present application provides a readable storage medium, on which a program or instructions are stored, and when the program or instructions are executed by a processor, the steps of the method described in any of the above method embodiments are implemented.

[0221] <Embodiment of the device>

[0222] Figure 3 It is a structural block diagram of a visual optimization device 3000 for a three-dimensional model according to an embodiment of the present invention.

[0223] In this embodiment, as Figure 3 shown, the visual optimization device 3000 for a three-dimensional model includes a memory 3001 and a processor 3002. The memory 3001 is used to store executable instructions, and the processor 3002 is used to operate under the control of the instructions to execute the method described in any of the above embodiments.

[0224] The present invention may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0225] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0226] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0227] The computer program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages - such as Smalltalk, C++, etc., and conventional procedural programming languages - such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., via an Internet service provider through the Internet). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present invention.

[0228] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0229] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions for implementing various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0230] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0231] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction may include one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive boxes may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box of the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system for performing the specified functions or acts, or may be implemented by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that implementation via hardware, implementation via software, and implementation via a combination of software and hardware are equivalent.

[0232] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.

Claims

1. A visual optimization method for a three-dimensional model, characterized in that: The method comprises: Obtain a three-dimensional Gaussian model of the target scene; assigning potential features to each three-dimensional Gaussian volume in the three-dimensional Gaussian volume model; Determine projection information of the three-dimensional Gaussian model at each shooting angle according to multiple shooting angles; wherein the projection information of any shooting angle includes a projection image of the three-dimensional Gaussian model at the shooting angle and a potential feature set, and the potential feature set includes a plurality of potential features corresponding to a plurality of three-dimensional Gaussian bodies visible in the three-dimensional Gaussian model at the shooting angle; Inputting the projection image and the potential feature set corresponding to each shooting angle into a semantic recognition model to obtain a predicted semantic map corresponding to each shooting angle; Optimizing the semantic recognition model according to the predicted semantic map and the reference semantic map corresponding to each shooting angle to obtain an optimized semantic recognition model; According to the predicted semantic graph of the projection image corresponding to each shooting angle by the optimized semantic recognition model, a semantic feature set corresponding to the three-dimensional Gaussian volume model is determined.

2. The method according to claim 1, characterized in that The step of inputting the projection image and the potential feature set corresponding to each shooting angle into the semantic recognition model to obtain the predicted semantic graph corresponding to each shooting angle includes: For any shooting angle, determining a feature fusion image of the shooting angle according to the projection image of the shooting angle and the potential feature set, and obtaining the feature fusion image of each shooting angle; The feature fusion image of each shooting angle is input into the semantic recognition model to obtain the predicted semantic map corresponding to each shooting angle.

3. The method according to claim 2, characterized in that The step of determining the feature fused image of the shooting angle according to the projection image of the shooting angle and the potential feature set comprises: For each pixel in the projection image of the shooting angle of view, determining a potential feature of the pixel according to a potential feature of at least one three-dimensional Gaussian volume corresponding to the pixel in the potential feature set of the shooting angle of view; Each pixel point in the projection image of the shooting angle of view is associated with a potential feature corresponding to the pixel point to obtain a feature fused image of the shooting angle of view.

4. The method according to claim 1, characterized in that: The predicted semantic graph includes a plurality of predicted semantic labels, the reference semantic graph includes a plurality of reference semantic labels, and the semantic recognition model is optimized according to the predicted semantic graph and the reference semantic graph corresponding to each shooting angle to obtain the optimized semantic recognition model, including: Constructing a semantic loss function according to the multiple predicted semantic labels included in the predicted semantic graph of each shooting angle and the multiple reference semantic labels included in the reference semantic graph; wherein the semantic loss function is used to characterize the difference between the multiple predicted semantic labels included in the predicted semantic graph of each shooting angle and the multiple reference semantic labels included in the reference semantic graph; According to the semantic loss function, the semantic recognition model is optimized to obtain an optimized semantic recognition model.

5. The method according to claim 1, characterized in that The step of determining a semantic feature set corresponding to the three-dimensional Gaussian volume model according to a predicted semantic map of the projection image corresponding to each shooting angle according to the optimized semantic recognition model comprises: Inputting the projection image and the potential feature set corresponding to each shooting angle into the optimized semantic recognition model to obtain a predicted semantic map of each shooting angle; wherein the predicted semantic map includes a predicted semantic label corresponding to each pixel point; For each pixel in the predicted semantic map of each shooting angle, determine a semantic feature of at least one three-dimensional Gaussian volume corresponding to the pixel according to the predicted semantic label of the pixel; According to the semantic features corresponding to each three-dimensional Gaussian volume in the three-dimensional Gaussian volume model, a semantic feature set corresponding to the three-dimensional Gaussian volume model is obtained.

6. The method according to claim 1, characterized in that The method further comprises: Inputting the projection image and the potential feature set corresponding to each of the shooting angles into the illumination recognition model to obtain a predicted illumination map corresponding to each of the shooting angles; Optimizing the illumination recognition model according to the predicted illumination map and the reference illumination map corresponding to each shooting angle to obtain an optimized illumination recognition model; The illumination feature set of the three-dimensional Gaussian model is determined according to the illumination recognition result of the optimized illumination recognition model for the projection image corresponding to each shooting angle.

7. The method according to claim 6, characterized in that The illumination recognition model includes a diffuse reflection recognition model and a highlight recognition model. The projection image and the potential feature set corresponding to each shooting angle are input into the illumination recognition model to obtain a predicted illumination map corresponding to each shooting angle, including: Inputting the projection image and the potential feature set corresponding to each shooting angle into the diffuse reflection recognition model to obtain the predicted diffuse reflection map corresponding to each shooting angle; Inputting the projection image and the potential feature set corresponding to each shooting angle into the highlight recognition model to obtain a predicted highlight map corresponding to each shooting angle; Determining a predicted illumination map corresponding to each shooting angle according to the predicted highlight map and the predicted diffuse reflection map corresponding to each shooting angle; The step of optimizing the illumination recognition model according to the predicted illumination map and the reference illumination map corresponding to each shooting angle to obtain an optimized illumination recognition model includes: According to the predicted illumination map and the reference illumination map corresponding to each shooting angle, the diffuse reflection recognition model and the highlight recognition model are optimized to obtain optimized diffuse reflection recognition model and highlight recognition model.

8. The method according to claim 7, characterized in that The step of inputting the projection image and the potential feature set corresponding to each shooting angle into the highlight recognition model to obtain the predicted highlight map corresponding to each shooting angle includes: For any shooting angle, according to the projection image and the potential feature set corresponding to the shooting angle, determine the direction head corresponding to the shooting angle, and obtain the direction head of each shooting angle; The direction head and the projection image corresponding to each shooting angle are input into the highlight recognition model to obtain the predicted highlight image corresponding to each shooting angle.

9. The method according to claim 8, characterized in that The step of determining the direction head corresponding to the shooting angle according to the projection image and the potential feature set corresponding to the shooting angle comprises: For each pixel in the projection image, determine a potential feature of the pixel according to a potential feature of at least one three-dimensional Gaussian volume corresponding to the pixel in the potential feature set of the shooting perspective; The direction of the shooting angle of view is determined according to the potential features corresponding to each pixel in the projection image.

10. A visual optimization device for a three-dimensional model, comprising a memory and a processor, wherein the memory is used to store executable instructions; and the processor is used to operate according to the control of the instructions to execute the method as described in any one of claims 1 to 9.