Model training method, spatial editing method and device applied to Gaussian scene
By employing multi-view encoding and model training methods, full-scene-level editing of 3D Gaussian scenes was achieved, solving the problems of low editing efficiency and poor quality in existing technologies and improving both editing efficiency and quality.
Patent Information
- Application Number
- CN202611123150.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-27
- Publication Date
- 2026-08-25
AI Technical Summary
Existing 3D Gaussian scene editing methods are insufficient to support full-scene editing of more than 100,000 Gaussian elements, and 2D editing methods result in problems such as poor geometric consistency and artifact drift between viewpoints.
Multi-view encoding is used to decouple the initial scene image into a unified 3D latent representation. The 3D editing module and task adaptation module in the scene editing model to be trained are used to perform full-scene-level editing. The efficiency and quality of editing are improved through model training.
It enables 3D Gaussian scene editing without viewpoint shifts, significantly improving editing efficiency, avoiding problems such as poor geometric consistency and artifact drift between viewpoints, and ensuring editing quality and user experience.
Smart Images

Figure CN122637136A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to a model training method, a spatial editing method and apparatus for Gaussian scenes. Background Technology
[0002] 3D Gaussian Splatting (3DGS) can represent complete scenes in Gaussian primitives and supports real-time rendering. In applications such as home decoration, virtual reality, and digital twins, users often need to edit existing 3D scenes represented by 3DGS, such as adding or deleting objects, changing materials, and adjusting lighting.
[0003] However, most current 3DGS scene editing methods adopt a "2D editing - 3D optimization" process, that is, first use a 2D diffusion model to modify the rendered image frame by frame according to the text instructions, and then perform scene-by-scene iterative optimization on the 3D scene represented by 3DGS to fit the editing result. Although a few methods attempt to directly map the input to a 3D Gaussian, these editing methods are mostly geared towards object reconstruction or generation, and are difficult to support the editing needs of full-scene-level editing of more than 100,000 Gaussian units. Summary of the Invention
[0004] This disclosure provides a model training method, a spatial editing method and apparatus for Gaussian scenes, to solve or alleviate one or more technical problems in the prior art.
[0005] Firstly, this disclosure provides a model training method for editing models applied to 3D Gaussian scenes, including: Obtain target training samples, wherein the target training samples include initial editing instruction information, initial scene images corresponding to each of the N viewpoints, and reference scene images corresponding to each initial scene image; the initial editing instruction information is used to instruct the modification of the appearance parameters of the initial 3D Gaussian scene corresponding to the initial scene image; the initial scene images corresponding to the N viewpoints are used to describe the initial 3D Gaussian scene under different viewpoints, and the reference scene images corresponding to the initial scene images are used to represent the labeled scene image after the appearance parameters of the initial 3D Gaussian scene are modified based on the initial editing instruction information; N is an integer greater than or equal to 2; The initial scene images corresponding to the N viewpoints are subjected to multi-view encoding processing to obtain the three-dimensional latent representation to be edited; wherein, the three-dimensional latent representation to be edited is a unified representation that maps the initial three-dimensional Gaussian scene to the latent space and integrates the N viewpoints. The three-dimensional latent representation to be edited and the initial editing semantic vector representing the initial editing instruction information are input into the scene editing model to be trained. The three-dimensional editing module in the scene editing model to be trained is used to edit the three-dimensional latent representation to be edited based on the initial editing semantic vector to obtain the initial three-dimensional latent representation. Then, the task adaptation module in the scene editing model to be trained is used to correct the initial three-dimensional latent representation based on the initial editing semantic vector to obtain the target three-dimensional latent representation. Based on the target's three-dimensional latent representation, estimated scene images corresponding to each of the N viewpoints are obtained; Based on the estimated scene images corresponding to each viewpoint and the reference scene images corresponding to each viewpoint, the target loss value is obtained. Based on the target loss value, the 3D editing module and the task adaptation module in the scene editing model to be trained are trained to obtain the target scene editing model.
[0006] Secondly, this disclosure provides a model training device for editing models applied to 3D Gaussian scenes, comprising: A sample acquisition unit is used to acquire target training samples, wherein the target training samples include initial editing instruction information, initial scene images corresponding to each of the N viewpoints, and reference scene images corresponding to each initial scene image; the initial editing instruction information is used to instruct the modification of the appearance parameters of the initial 3D Gaussian scene corresponding to the initial scene image; the initial scene images corresponding to the N viewpoints are used to describe the initial 3D Gaussian scene under different viewpoints, and the reference scene images corresponding to the initial scene images are used to represent the labeled scene image after the appearance parameters of the initial 3D Gaussian scene are modified based on the initial editing instruction information; N is an integer greater than or equal to 2; The model training unit is used to perform multi-view encoding processing on the initial scene images corresponding to the N viewpoints to obtain the 3D latent representation to be edited; wherein, the 3D latent representation to be edited is a unified representation that maps the initial 3D Gaussian scene to the latent space and integrates the N viewpoints; it is used to input the 3D latent representation to be edited and the initial editing semantic vector representing the initial editing instruction information into the scene editing model to be trained, so as to use the 3D editing module in the scene editing model to edit the 3D latent representation to be edited based on the initial editing semantic vector to obtain the initial 3D latent representation; and, using the task adaptation module in the scene editing model to be trained, and based on the initial editing semantic vector, to correct the initial 3D latent representation to obtain the target 3D latent representation; based on the target 3D latent representation, the estimated scene images corresponding to each of the N viewpoints are obtained; based on the estimated scene images corresponding to each viewpoint and the reference scene images corresponding to each viewpoint, the target loss value is obtained, so as to train the 3D editing module and the task adaptation module in the scene editing model to be trained based on the target loss value to obtain the target scene editing model.
[0007] Thirdly, this disclosure provides a spatial editing method for three-dimensional Gaussian scenes, including: Obtain the scene images to be processed corresponding to each of the M viewpoints, as well as the target editing instruction information; wherein, the scene images to be processed corresponding to the M viewpoints are used to describe the target 3D Gaussian scene under different viewpoints; the target editing instruction information is used to instruct the modification of the appearance parameters of the target 3D Gaussian scene corresponding to the scene images to be processed; M is an integer greater than or equal to 2; The scene images to be processed corresponding to the M viewpoints are subjected to multi-view encoding processing to obtain a first three-dimensional latent representation; wherein, the first three-dimensional latent representation is a unified representation of the target three-dimensional Gaussian scene mapped to the latent space and integrating the M viewpoints; The first three-dimensional latent representation and the target editing semantic vector representing the target editing instruction information are input into the target scene editing model to obtain the second three-dimensional latent representation; wherein, the target scene editing model is obtained after training using the above-described model training method; Based on the second three-dimensional latent representation, the target three-dimensional Gaussian scene after editing based on the target editing instruction information is obtained.
[0008] Fourthly, this disclosure provides a spatial editing apparatus for three-dimensional Gaussian scenes, comprising: The data input unit is used to acquire the scene images to be processed corresponding to each of the M viewpoints, as well as the target editing instruction information; wherein, the scene images to be processed corresponding to the M viewpoints are used to describe the target 3D Gaussian scene under different viewpoints; the target editing instruction information is used to instruct the modification of the appearance parameters of the target 3D Gaussian scene corresponding to the scene images to be processed; M is an integer greater than or equal to 2; A spatial editing unit is used to perform multi-view encoding processing on the scene images to be processed corresponding to the M viewpoints to obtain a first three-dimensional latent representation; wherein, the first three-dimensional latent representation is a unified representation of the target three-dimensional Gaussian scene mapped to the latent space and integrating the M viewpoints; the first three-dimensional latent representation and the target editing semantic vector representing the target editing instruction information are input into the target scene editing model to obtain a second three-dimensional latent representation; wherein, the target scene editing model is obtained after training using the above-mentioned model training method; based on the second three-dimensional latent representation, the target three-dimensional Gaussian scene edited based on the target editing instruction information is obtained.
[0009] Fifthly, an electronic device is provided, comprising: At least one processor; and The memory is communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in the present disclosure.
[0010] In a sixth aspect, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods according to embodiments of the present disclosure.
[0011] In a seventh aspect, a computer program product is provided, including a computer program that, when executed by a processor, implements any of the methods according to embodiments of the present disclosure.
[0012] The beneficial effects of the technical solution provided in this disclosure include at least the following: In this way, the proposed solution decouples the initial scene images corresponding to N viewpoints into a viewpoint-independent 3D latent representation to be edited. Combined with a scene editing model to be trained and an initial editing semantic vector representing the initial editing instructions, it performs full-scene-level editing (i.e., unified editing across multiple viewpoints) of the 3D latent representation. This editing process eliminates the need for viewpoint-by-view editing of the 3D Gaussian scene, significantly improving the editing efficiency of the 3D Gaussian scene. Simultaneously, it effectively avoids problems such as poor geometric consistency and artifact drift between viewpoints caused by 2D viewpoint-by-view editing, thereby improving the quality of scene editing. Furthermore, the proposed solution first utilizes the 3D editing module in the scene editing model to obtain the initial 3D latent representation, and then uses the task adaptation module in the scene editing model to further refine the initial 3D latent representation output by the 3D editing module based on the initial editing semantic vector, further improving the accuracy of the obtained target 3D latent representation, making it closer to the requirements of full-scene-level editing. Thus, while achieving viewpoint-free scene editing and improving editing efficiency, it further ensures the editing quality of the 3D Gaussian scene, thereby enhancing the user experience.
[0013] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0014] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments provided according to this disclosure and should not be construed as limiting the scope of this disclosure.
[0015] Figure 1 This is an illustrative flowchart of a model training method for an editing model applied to a 3D Gaussian scene according to an embodiment of this application. Figure 1 ; Figure 2 This is an illustrative diagram of the training logic of a scene editing model to be trained according to an embodiment of this application. Figure 1 ; Figure 3 This is an illustrative flowchart of a model training method for an editing model applied to a 3D Gaussian scene according to an embodiment of this application. Figure 2 ; Figure 4 This is an illustrative diagram illustrating the processing logic of a scene editing model to be trained according to an embodiment of this application; Figure 5This is an illustrative diagram of the processing logic of the 3D editing module in a scene editing model to be trained according to an embodiment of this application. Figure 1 ; Figure 6 This is an illustrative diagram of the processing logic of the 3D editing module in a scene editing model to be trained according to an embodiment of this application. Figure 2 ; Figure 7 This is an illustrative flowchart of a model training method for an editing model applied to a 3D Gaussian scene according to an embodiment of this application. Figure 3 ; Figure 8 This is an illustrative diagram of the training logic of a scene editing model to be trained according to an embodiment of this application. Figure 2 ; Figure 9 This is an illustrative flowchart of a model training method for an editing model applied to a 3D Gaussian scene according to an embodiment of this application. Figure 4 ; Figure 10 This is an illustrative diagram of the training logic of a scene editing model to be trained according to an embodiment of this application. Figure 3 ; Figure 11 This is a schematic flowchart of a spatial editing method applied to a 3D Gaussian scene according to an embodiment of this application; Figure 12 This is a flowchart illustrating a spatial editing method applied to a 3D Gaussian scene according to an embodiment of this application in a specific example; Figure 13 This is a schematic diagram of the structure of a model training device for an editing model applied to a 3D Gaussian scene according to an embodiment of this application; Figure 14 This is a schematic diagram of a spatial editing device applied to a three-dimensional Gaussian scene according to an embodiment of this application; Figure 15 This is a block diagram of an electronic device used to implement the model training method for an editing model applied to a 3D Gaussian scene or the spatial editing method applied to a 3D Gaussian scene according to the embodiments of this disclosure. Detailed Implementation
[0016] The present disclosure will now be described in further detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.
[0017] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.
[0018] This disclosure provides a model training method for an editing model applied to a 3D Gaussian scene. The method first encodes the observed initial scene images from multiple perspectives (i.e., N perspectives) into a unified 3D voxel latent representation (corresponding to the 3D latent representation to be edited). Then, this unified 3D voxel latent representation and an initial editing semantic vector representing the initial editing instructions are input into the scene editing model to be trained. The 3D editing module in the scene editing model is then used to edit the 3D latent representation to be edited based on the initial editing semantic vector, thus obtaining the initial 3D latent representation. Furthermore, the task adaptation module in the scene editing model is used to edit the 3D latent representation based on the initial editing instructions. The initial 3D latent representation is modified using sense vectors to obtain a fully edited 3D latent representation (corresponding to the target 3D latent representation). This process supports full-scene-level editing without iterative optimization of the 3D Gaussian scene from each viewpoint, thus effectively improving the efficiency of scene editing. Furthermore, based on the fully edited 3D latent representation, estimated scene images corresponding to each of the N viewpoints are obtained. Loss calculations are performed on the estimated scene images and reference scene images corresponding to each viewpoint to obtain the target loss. This target loss is then used to train the 3D editing module and the task adaptation module in the scene editing model to obtain the target scene editing model. Thus, the obtained target scene editing model possesses editing capabilities that are not subject to scene-by-scene optimization, laying the foundation for efficient full-scene-level editing of 3D Gaussian scenes.
[0019] Specifically, Figure 1 This is an illustrative flowchart of a model training method for an editing model applied to a 3D Gaussian scene according to an embodiment of this application. Figure 1 This method can be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.
[0020] Furthermore, the method includes at least a portion of the following: For example... Figure 1 As shown, it includes: Step S101: Obtain the target training sample.
[0021] Here, the target training samples include initial editing instruction information, initial scene images corresponding to each of the N (N is an integer greater than or equal to 2) viewpoints, and reference scene images corresponding to each initial scene image.
[0022] Furthermore, the initial editing instruction information is used to instruct the modification of the appearance parameters (such as lighting, materials, etc.) of the initial 3D Gaussian scene corresponding to the initial scene image.
[0023] Furthermore, in one example, the initial editing instruction information may specifically be descriptive text describing changes to the appearance parameters of the initial 3D Gaussian scene, or it may specifically be an example image showing the effect of the changes to the appearance parameters of the initial 3D Gaussian scene, or it may be both descriptive text and an example image. This disclosure does not impose any specific limitations on this.
[0024] Furthermore, the initial scene images corresponding to the N viewpoints are also the initial scene images corresponding to each of the N viewpoints; the initial three-dimensional Gaussian scenes under different viewpoints can be described by the initial scene images corresponding to the N viewpoints. In other words, the initial scene images corresponding to each viewpoint are used to describe the initial three-dimensional Gaussian scenes under the corresponding viewpoints.
[0025] Furthermore, the reference scene image corresponding to the initial scene image is used to represent the labeled scene image after the appearance parameters of the initial 3D Gaussian scene have been modified based on the initial editing instruction information. Here, each of the N initial scene images corresponds to a labeled scene image. This provides a basis for the subsequent calculation of the target loss value.
[0026] Here, the initial 3D Gaussian scene is represented by a set of 3D Gaussian volumes. Furthermore, the set of 3D Gaussian volumes contains multiple 3D Gaussian volumes (also called Gaussian units), and each 3D Gaussian volume can be used to represent local geometry, appearance, or density distribution in the 3D scene.
[0027] Step S102: Perform multi-view encoding processing on the initial scene images corresponding to the N viewpoints to obtain the three-dimensional latent representation to be edited.
[0028] Here, the three-dimensional latent representation to be edited is a unified representation that maps the initial three-dimensional Gaussian scene to the latent space and integrates the N perspectives.
[0029] Step S103: Input the three-dimensional latent representation to be edited and the initial editing semantic vector representing the initial editing instruction information into the scene editing model to be trained, so as to use the three-dimensional editing module in the scene editing model to be trained and, based on the initial editing semantic vector, to edit the three-dimensional latent representation to be edited to obtain the initial three-dimensional latent representation; and, use the task adaptation module in the scene editing model to be trained and, based on the initial editing semantic vector, to correct the initial three-dimensional latent representation to obtain the target three-dimensional latent representation.
[0030] In one example, the initial editing semantic vector representing the initial editing instruction information can be obtained as follows: The initial editing instruction information is encoded using a modal encoder (such as a text encoder) to obtain the initial editing semantic vector representing the initial editing instruction information. This provides constraints for subsequent scene editing using the scene editing model to be trained.
[0031] Step S104: Based on the target's three-dimensional latent representation, obtain the estimated scene image corresponding to each of the N viewpoints.
[0032] Step S105: Based on the estimated scene image corresponding to each viewpoint and the reference scene image corresponding to each viewpoint, obtain the target loss value. Based on the target loss value, train the 3D editing module and the task adaptation module in the scene editing model to be trained to obtain the target scene editing model.
[0033] Here, the reference scene image corresponding to each viewpoint can be specifically the reference scene image corresponding to each initial scene image among the N initial scene images corresponding to the viewpoints.
[0034] For example, in one example, the above-mentioned method of obtaining the target loss value based on the estimated scene image corresponding to each viewpoint and the reference scene image corresponding to each viewpoint can specifically include: for the i-th viewpoint among N viewpoints, a first loss value can be obtained based on the estimated scene image corresponding to the i-th viewpoint (i takes the value of an integer from 1 to N) among N viewpoints and the reference scene image corresponding to the i-th viewpoint. For N viewpoints, N first loss values can be obtained, and then the target loss value can be obtained based on the N first loss values.
[0035] In other words, the present scheme encodes the initial scene images corresponding to N viewpoints into a unified 3D latent representation to be edited from multiple perspectives. Then, it inputs the 3D latent representation to be edited and the initial editing semantic vector representing the initial editing instruction information into the scene editing model to be trained to obtain the target 3D latent representation. Based on the target 3D latent representation, it obtains the estimated scene image corresponding to each of the N viewpoints. At this time, the reference scene image corresponding to each initial scene image is used as the label data for training the scene editing model to be trained. Then, based on the estimated scene image corresponding to each viewpoint and the reference scene image corresponding to each viewpoint, the target loss value is obtained. Using the target loss value, the parameters of the 3D editing module and the task adaptation module in the scene editing model to be trained are fine-tuned to obtain the target scene editing model.
[0036] In this way, the proposed solution decouples the initial scene images corresponding to N viewpoints into a viewpoint-independent 3D latent representation to be edited. Combined with a scene editing model to be trained and an initial editing semantic vector representing the initial editing instructions, it performs full-scene-level editing (i.e., unified editing across multiple viewpoints) of the 3D latent representation. This editing process eliminates the need for viewpoint-by-view editing of the 3D Gaussian scene, significantly improving the editing efficiency of the 3D Gaussian scene. Simultaneously, it effectively avoids problems such as poor geometric consistency and artifact drift between viewpoints caused by 2D viewpoint-by-view editing, thereby improving the quality of scene editing. Furthermore, the proposed solution first utilizes the 3D editing module in the scene editing model to obtain the initial 3D latent representation, and then uses the task adaptation module in the scene editing model to further refine the initial 3D latent representation output by the 3D editing module based on the initial editing semantic vector, further improving the accuracy of the obtained target 3D latent representation, making it closer to the requirements of full-scene-level editing. Thus, while achieving viewpoint-free scene editing and improving editing efficiency, it further ensures the editing quality of the 3D Gaussian scene, thereby enhancing the user experience.
[0037] Furthermore, this disclosed solution trains the model by using the loss value between the estimated scene image corresponding to each viewpoint and the reference scene image corresponding to each viewpoint, significantly improving the robustness and stability of the model during full-scene-level editing. Thus, the trained target scene editing model possesses the capability for full-scene-level editing of 3D Gaussian scenes without per-view editing, laying the foundation for efficient editing of 3D Gaussian scenes using the target scene editing model and improving the user's scene editing efficiency and experience.
[0038] Furthermore, the scene editing model to be trained in this disclosed scheme includes a 3D editing module for coarse-grained editing of the 3D latent representation to be edited, and a task adaptive module for fine-grained editing of the editing results output by the 3D editing module. This significantly enhances the accuracy, realism, and controllability of the model in editing 3D Gaussian scenes, thereby achieving high-precision semantic editing while realizing scene editing without perspective shifts, thus providing favorable support for interactive editing of complex 3D Gaussian scenes.
[0039] Furthermore, in a specific example, the 3D latent representation to be edited can be obtained in the following manner; specifically, the multi-view encoding processing of the initial scene images corresponding to the N viewpoints described above to obtain the 3D latent representation to be edited (e.g., step S102) can specifically include: Step S102-1: Extract features from the initial scene images corresponding to each viewpoint to obtain feature maps corresponding to each initial scene image.
[0040] For example, in one example, an image encoder is used to extract features from the initial scene images corresponding to each viewpoint in order to obtain feature maps corresponding to each initial scene image.
[0041] The above is merely an example; in practical applications, other feature extraction methods may also be used, and this disclosure does not limit such methods.
[0042] Step S102-2: Project the preset three-dimensional voxel grid onto the feature map corresponding to each initial scene image, and sample the feature blocks corresponding to the voxels in the preset three-dimensional voxel grid on each feature map.
[0043] Step S102-3: Perform fusion processing (such as visibility weighted aggregation processing) on the feature blocks corresponding to the voxels in the preset three-dimensional voxel grid on each feature map to obtain the fused features corresponding to each voxel in the preset three-dimensional voxel grid.
[0044] Step S102-4: Process the fusion features corresponding to each voxel in the preset three-dimensional voxel mesh to obtain the three-dimensional latent representation to be edited.
[0045] For example, in one instance, the fused features corresponding to each voxel in a preset 3D voxel grid are processed by a 3D refinement network, such as a 3D convolutional network, to obtain the 3D latent representation to be edited.
[0046] Thus, this disclosure provides a specific scheme for multi-view encoding processing of initial scene images corresponding to N viewpoints. This scheme can fuse the initial scene images corresponding to N viewpoints into a unified latent representation of a 3D Gaussian scene, thereby decoupling the number of 3D Gaussian volumes in the 3D Gaussian scene from the number of input viewpoints. This lays the foundation for enabling the model to edit 3D Gaussian scenes without viewpoints and to achieve full-scene-level editing of 3D Gaussian scenes.
[0047] For example, such as Figure 2As shown, the target training samples include initial editing instruction information, initial scene images corresponding to N viewpoints, and reference scene images corresponding to each initial scene image. First, a modal encoder (such as a text encoder) is used to encode the initial editing instruction information to obtain an initial editing semantic vector representing the initial editing instruction information. Then, multi-view encoding processing is performed on the initial scene images corresponding to the N viewpoints to decouple them into viewpoint-independent 3D latent representations to be edited. The specific processing logic of the above multi-view encoding processing is as follows: First, an image encoder is used to extract features from the initial scene images corresponding to each viewpoint to obtain each... The initial scene image is used to generate a feature map, and a pre-constructed 3D voxel grid is projected onto the feature map corresponding to each initial scene image to sample the feature blocks corresponding to the voxels in the pre-constructed 3D voxel grid on each feature map. Then, the features corresponding to the voxels in the pre-constructed 3D voxel grid on each feature map are aggregated and refined. For example, the feature blocks corresponding to the voxels in the pre-constructed 3D voxel grid on each feature map are aggregated by visibility weighting to obtain the fused features corresponding to each voxel in the pre-constructed 3D voxel grid. Then, a 3D convolutional network is used to process the fused features corresponding to each voxel in the pre-constructed 3D voxel grid to obtain the 3D latent representation to be edited.
[0048] Furthermore, such as Figure 2 As shown, the 3D latent representation to be edited and the initial editing semantic vector representing the initial editing instructions are input into the scene editing model to be trained for 3D latent space editing. For example, firstly, the 3D editing module in the scene editing model to be trained is used to edit the 3D latent representation to be edited based on the initial editing semantic vector to obtain the initial 3D latent representation. Then, the task adaptation module in the scene editing model to be trained is used to correct the initial 3D latent representation based on the initial editing semantic vector to obtain the target 3D latent representation. At this point, the estimated scene images corresponding to each viewpoint are obtained according to the target 3D latent representation. Based on the estimated scene images corresponding to each viewpoint and the reference scene images corresponding to each viewpoint (i.e., the reference scene images corresponding to each initial scene image), the target loss value is obtained. Then, based on the target loss value, the parameters of the 3D editing module and the task adaptation module in the scene editing model to be trained are fine-tuned to obtain the target scene editing model. In this way, the trained target scene editing model has the ability to edit 3D Gaussian scenes without alternating viewpoints, thus laying the foundation for efficient editing of 3D Gaussian scenes using the target scene editing model and improving the user's scene editing efficiency and experience.
[0049] Figure 3 This is an illustrative flowchart of a model training method for an editing model applied to a 3D Gaussian scene according to an embodiment of this application. Figure 2This method can be optionally applied to electronic devices, such as personal computers, servers, and server clusters. It is understood that the above... Figure 1 and Figure 2 The methods shown can also be applied to this example, and the related content will not be elaborated further in this example.
[0050] Furthermore, the method includes at least a portion of the following: For example... Figure 3 As shown, it includes: Step S301: Obtain the target training sample.
[0051] Here, the target training samples include initial editing instruction information, initial scene images corresponding to each of the N (N is an integer greater than or equal to 2) viewpoints, and reference scene images corresponding to each initial scene image.
[0052] Furthermore, the initial editing instruction information is used to instruct the modification of the appearance parameters (such as lighting, materials, etc.) of the initial 3D Gaussian scene corresponding to the initial scene image.
[0053] Furthermore, the initial scene images corresponding to the N viewpoints are used to describe the initial three-dimensional Gaussian scene under different viewpoints.
[0054] Furthermore, the reference scene image corresponding to the initial scene image is used to represent the label scene image after the appearance parameters of the initial three-dimensional Gaussian scene have been modified based on the initial editing instruction information.
[0055] It should be noted that the information regarding the initial editing instructions and the initial 3D Gaussian scene can be found in the example above, and will not be repeated here.
[0056] Step S302: Perform multi-view encoding processing on the initial scene images corresponding to the N viewpoints to obtain the three-dimensional latent representation to be edited.
[0057] Here, the three-dimensional latent representation to be edited is a unified representation that maps the initial three-dimensional Gaussian scene to the latent space and integrates the N perspectives.
[0058] It should be noted that the specific processing logic for multi-view encoding can be found in the example above, and will not be repeated here.
[0059] Step S303: Utilize the low-resolution editing sub-network in the 3D editing module of the scene editing model to be trained, and combine the initial editing semantic vector representing the initial editing instruction information and the 3D latent representation to be edited to obtain voxel change features, and perform fusion processing based on the voxel change features and the 3D latent representation to be edited to obtain a fused 3D latent representation.
[0060] Here, the low-resolution editing subnetwork is used to perform global editing of the three-dimensional latent representation from a low-resolution dimension.
[0061] For example, in one example, for a 54×54×54 resolution 3D latent representation to be edited, before inputting this resolution 3D latent representation to the low-resolution editing sub-network, the 3D latent representation to be edited at this resolution is first downsampled. For example, using a convolutional network with downsampling processing capability in the 3D editing module, the 3D latent representation to be edited at this resolution is downsampled to half of the original resolution, resulting in a low-resolution (i.e., 27×27×27) 3D latent representation to be edited. Then, the low-resolution editing sub-network is applied to the low-resolution 3D latent representation to be edited, and combined with the initial editing semantic vector, to achieve global editing at the low-resolution level.
[0062] Here, the 3D latent representation to be edited with a resolution of 54×54×54 can be specifically understood as the 3D latent representation to be edited represented by a 54×54×54 voxel mesh.
[0063] Here, in one example, the low-resolution editing subnetwork may be specifically a U-net network, or it may be specifically other network structures with global editing for the three-dimensional latent representation, which is not specifically limited in this disclosure.
[0064] Furthermore, the voxel variation features represent the spatial distribution and amount of editing variation of the voxels to be edited in the three-dimensional latent representation to be edited.
[0065] Step S304: Using the high-resolution editing sub-network in the 3D editing module, the local appearance information in the fused 3D latent representation is restored to obtain the initial 3D latent representation.
[0066] Here, the high-resolution editing subnetwork is used to recover local appearance information in the three-dimensional latent representation from a high-resolution dimension.
[0067] For example, continuing with the 54×54×54 resolution 3D latent representation to be edited, before using the high-resolution editing sub-network to restore the local appearance information in the low-resolution fused 3D latent representation output by the low-resolution editing sub-network, the low-resolution fused 3D latent representation is first upsampled. For example, using a convolutional network with upsampling capabilities in the 3D editing module, the low-resolution fused 3D latent representation is upsampled to restore its resolution to the original resolution (or high resolution), resulting in a 54×54×54 resolution fused 3D latent representation. Then, the high-resolution editing sub-network is applied to the fused 3D latent representation at this resolution to restore its internal local appearance information in the high-resolution dimension.
[0068] Furthermore, in one example, the low-resolution editing subnetwork and the high-resolution editing subnetwork are isomorphic networks; that is, in this example, the low-resolution editing subnetwork and the high-resolution editing subnetwork correspond to each other in terms of network structure design and have similar shapes, and belong to two instances of the same network architecture.
[0069] Here, the "isomorphic network" referred to in this example means a neural network that is similar or identical in topology, such as the type and number of network layers, the arrangement of neurons, and the parameter configuration. In short, it means a network with a similar or identical skeleton.
[0070] Furthermore, in one example, the high-resolution editing subnetwork may also be a U-net network, and the U-net networks of both are similar or identical in architecture.
[0071] Alternatively, in one example, the high-resolution editing subnetwork may be a different network structure from the low-resolution editing subnetwork and has the ability to restore local appearance; this disclosure does not impose any specific limitations on this.
[0072] Step S305: Using the task adaptation module in the scene editing model to be trained, and based on the initial editing semantic vector, the initial three-dimensional latent representation is corrected to obtain the target three-dimensional latent representation.
[0073] Step S306: Based on the target's three-dimensional latent representation, obtain the estimated scene image corresponding to each of the N viewpoints.
[0074] Step S307: Based on the estimated scene image corresponding to each viewpoint and the reference scene image corresponding to each viewpoint, obtain the target loss value. Based on the target loss value, train the 3D editing module and the task adaptation module in the scene editing model to be trained to obtain the target scene editing model.
[0075] For example, such as Figure 4As shown, when the 3D latent representation to be edited and the semantic vector of the text instruction (corresponding to the initial editing semantic vector representing the initial editing instruction information mentioned above) are input into the scene editing model to be trained: First, the 3D latent representation to be edited (with a resolution of 54×54×54) is downsampled to obtain a low-resolution (e.g., 27×27×27) 3D latent representation to be edited. Then, the low-resolution editing sub-network in the 3D editing module of the scene editing model to be trained is used, and combined with the semantic vector and the low-resolution 3D latent representation to be edited, the voxel change features are obtained. Based on the voxel change features and the low-resolution 3D latent representation to be edited, the voxel change features are obtained. The 3D latent representation is edited and fused to obtain a low-resolution fused 3D latent representation. Then, the low-resolution fused 3D latent representation is upsampled to obtain a high-resolution (e.g., 54×54×54) fused 3D latent representation. Then, the high-resolution editing sub-network in the 3D editing module is used to restore the local appearance information in the high-resolution fused 3D latent representation to obtain an initial 3D latent representation. Further, the task adaptation module in the scene editing model to be trained is used, and based on the initial editing semantic vector, the initial 3D latent representation is corrected to obtain the target 3D latent representation.
[0076] Thus, this disclosed solution, through a "low-resolution-high-resolution" editing architecture, effectively balances the consistency of global semantics with the fidelity of local details. Specifically, the low-resolution editing sub-network enables unified global appearance modification across viewpoints, avoiding geometric drift in multi-view editing; while the high-resolution editing sub-network specifically restores lost local appearance information such as texture and edges in the fused 3D latent representation, effectively suppressing the inherent detail blurring defects in 3D latent space editing. In this way, while completing the editing according to the editing intent, the fine visual texture of the original scene is maintained, effectively solving the problem of "inaccurate global modification and unclear local details" in existing scene editing methods. This significantly improves the accuracy and realism of the model when editing scenes, making the obtained target 3D latent representation closer to the requirements of full-scene-level editing. Furthermore, it enables the training of a target scene editing model capable of non-viewpoint-by-view editing of 3D Gaussian scenes, laying the foundation for subsequent full-scene-level editing of 3D Gaussian scenes using the target scene editing model.
[0077] Furthermore, in a specific example, the 3D editing module in the scene editing model to be trained further includes a first cross-attention layer; the method further includes: The first cross-attention result is obtained by utilizing the first cross-attention layer in the 3D editing module and combining the initial editing semantic vector and the 3D latent representation to be edited.
[0078] Here, in one example, before obtaining the first cross-attention result using the first cross-attention layer, the 3D latent representation to be edited may specifically be a low-resolution 3D latent representation to be edited obtained through downsampling. That is, the 3D latent representation to be edited processed by the first cross-attention layer may specifically be a low-resolution 3D latent representation to be edited obtained through downsampling.
[0079] For example, when a low-resolution 3D latent representation to be edited is obtained through downsampling, a first cross-attention layer is used to perform cross-attention processing on the query vector (Query, Q) corresponding to each voxel in the low-resolution 3D latent representation to be edited, the key vector (Key, K) corresponding to each element in the initial editing semantic vector, and the value vector (Value, V) to obtain the first cross-attention result. In this way, a coarse-grained dynamic association between the editing semantic vector and the 3D latent representation to be edited is realized, providing a clear direction for subsequent coarse-grained global editing.
[0080] In one example, the above-described method of utilizing the low-resolution editing sub-network in the 3D editing module, combined with the initial editing semantic vector and the 3D latent representation to be edited, to obtain voxel variation features, and then fusing these voxel variation features with the 3D latent representation to be edited to obtain a fused 3D latent representation (e.g., step S303), can specifically include: Step S303-1: The first cross-attention result is fused with the three-dimensional latent representation to be edited (or the low-resolution three-dimensional latent representation to be edited) to obtain the three-dimensional latent representation to be edited carrying the first cross-attention result.
[0081] For example, in one example, the first cross-attention result is fused with the 3D latent representation to be edited (or a low-resolution 3D latent representation to be edited) using a weighted fusion process (or a connection process) to obtain the 3D latent representation to be edited carrying the first cross-attention result.
[0082] Step S303-2: Using the low-resolution editing sub-network in the 3D editing module, and combining the 3D latent representation to be edited carrying the first cross-attention result and the initial editing semantic vector, voxel change features are obtained. Based on the voxel change features, the 3D latent representation to be edited (or the low-resolution 3D latent representation to be edited) is fused to obtain a fused 3D latent representation.
[0083] For example, such as Figure 5As shown, before obtaining the fused 3D latent representation using the low-resolution editing sub-network in the 3D editing module of the scene editing model to be trained, the first cross-attention layer in the 3D editing module is used first, and combined with the 3D latent representation to be edited (or the low-resolution 3D latent representation to be edited) and the semantic vector of the text instruction, to obtain the first cross-attention result. The first cross-attention result is then fused with the 3D latent representation to be edited (or the low-resolution 3D latent representation to be edited) to obtain the 3D latent representation to be edited carrying the first cross-attention result. Then, the low-resolution editing sub-network is used, and combined with the 3D latent representation to be edited carrying the first cross-attention result and the semantic vector of the text instruction, to obtain the voxel change features. The voxel change features are then fused with the 3D latent representation to be edited (or the low-resolution 3D latent representation to be edited) to obtain the fused 3D latent representation.
[0084] In this way, the proposed solution establishes a coarse-grained dynamic relationship between the editing semantic vector and the 3D latent representation by introducing a first cross-attention layer, and obtains the first cross-attention result. The result obtained by fusing the first cross-attention result with the 3D latent representation to be edited is then input into the low-resolution editing sub-network. This ensures that subsequent global editing is no longer a blind overall offset, but a precise semantic-driven adjustment. Thus, through the "first locate, then edit" mechanism, the risk of erroneous modification of irrelevant areas is reduced. This improves the accuracy of global appearance modification while further suppressing cross-view geometric drift, ensuring the accuracy of scene-wide editing. This provides favorable support for the subsequent realization of view-free scene editing.
[0085] It should be noted that, in one example, where the low-resolution editing sub-network in the 3D editing module contains multiple serially connected network layers (such as a first encoding layer and a first decoding layer), the above-described method of utilizing the low-resolution editing sub-network in the 3D editing module, combined with the 3D latent representation to be edited carrying the first cross-attention result and the initial editing semantic vector, to obtain voxel change features, and then fusing the voxel change features with the 3D latent representation to be edited to obtain a fused 3D latent representation, can specifically include: The 3D latent representation to be edited, carrying the cross-attention result -11 (i.e., the first cross-attention result), is input into the first encoding layer in the low-resolution editing sub-network to obtain the first encoding result. At this time, the first cross-attention layer in the 3D editing module can be used, combined with the initial editing semantic vector and the first encoding result, to obtain the cross-attention result -12. The cross-attention result -12 is then fused with the first encoding result to obtain the first encoding result carrying the cross-attention result -12. Further, the first encoding result carrying the cross-attention result -12 is input into the first decoding layer in the low-resolution editing sub-network: first, the voxel change features are obtained, and then the voxel change features and the 3D latent representation to be edited (such as the low-resolution 3D latent representation to be edited) are fused to obtain the fused 3D latent representation.
[0086] Here, the network structure and processing logic of the low-resolution editing subnetwork described above are merely an example. As long as it can achieve coarse-grained global modification of the potential 3D representation to be edited, this disclosure does not impose any specific restrictions on it.
[0087] Furthermore, in a specific example, the 3D editing module in the scene editing model to be trained further includes a second cross-attention layer; the method further includes: The second cross-attention layer in the 3D editing module is used, and the initial editing semantic vector and the fused 3D latent representation are combined to obtain the second cross-attention result.
[0088] Here, in one example, before utilizing the second cross-attention layer to obtain the second cross-attention result, the fused 3D latent representation used can specifically be a high-resolution fused 3D latent representation obtained by upsampling the fused 3D latent representation output by the low-resolution editing subnetwork. That is, the fused 3D latent representation processed by the second cross-attention layer can specifically be a high-resolution fused 3D latent representation obtained by upsampling.
[0089] For example, after upsampling to obtain a high-resolution fused 3D latent representation, a second cross-attention layer is used to perform cross-attention processing on the query vector corresponding to each voxel in the high-resolution fused 3D latent representation and the key vector and value vector corresponding to each element in the initial edit semantic vector to obtain the second cross-attention result. In this way, a fine-grained dynamic association between the edit semantic vector and the fused 3D latent representation is realized, providing a clear direction for subsequent high-resolution local appearance information recovery.
[0090] In one example, the above-described method of using the high-resolution editing sub-network in the 3D editing module to restore the local appearance information in the fused 3D latent representation to obtain the initial 3D latent representation (e.g., step S304) can specifically include: Step S304-1: The second cross-attention result is fused with the fused three-dimensional latent representation (or the high-resolution fused three-dimensional latent representation) to obtain a fused three-dimensional latent representation carrying the second cross-attention result.
[0091] It should be noted that the implementation method of the fusion process in step S304-1 can be referred to the example in step S303-1 above, and will not be repeated here.
[0092] Step S304-2: Using the high-resolution editing sub-network in the 3D editing module, and combining the fused 3D latent representation carrying the second cross-attention result with the initial editing semantic vector, the initial 3D latent representation is obtained.
[0093] For example, such as Figure 6 As shown, before obtaining the initial 3D latent representation using the high-resolution editing sub-network in the 3D editing module of the scene editing model to be trained, the second cross-attention layer in the 3D editing module is first used, and the fused 3D latent representation (or the high-resolution fused 3D latent representation) and the semantic vector of the text instruction are combined to obtain the second cross-attention result. The second cross-attention result is then fused with the fused 3D latent representation (or the high-resolution fused 3D latent representation) to obtain the fused 3D latent representation carrying the second cross-attention result. Then, the high-resolution editing sub-network is used, and the fused 3D latent representation carrying the second cross-attention result and the semantic vector of the text instruction are combined to obtain the initial 3D latent representation.
[0094] Thus, this disclosed solution achieves semantically guided fine-grained local feature refocusing through a second cross-attention layer, combined with voxels from the fused 3D latent representation and the initial editing semantic vector. This process ensures that the high-resolution editing sub-network remains constrained by editing instructions while restoring local details such as texture and edges, avoiding semantic deviations or artifacts caused by indiscriminate editing. Furthermore, this disclosed solution inputs the result of fusing the second cross-attention layer with the fused 3D latent representation into the high-resolution editing sub-network, preserving the consistency of low-resolution global editing while accurately restoring high-frequency local appearance information, achieving semantically driven detail restoration. This effectively solves the problem of local detail loss in 3D latent space editing, thereby improving visual fidelity and further ensuring semantic alignment of local appearances across different viewpoints, thus enhancing the accuracy of full-scene-level editing.
[0095] It should be noted that, in one example, where the high-resolution editing sub-network in the 3D editing module contains multiple serially connected network layers (such as a second encoding layer and a second decoding layer), the above-described method of utilizing the high-resolution editing sub-network in the 3D editing module, combined with the fused 3D latent representation carrying the second cross-attention result and the initial editing semantic vector, to obtain the initial 3D latent representation, may specifically include: The fused 3D latent representation carrying cross-attention result-21 (i.e., the second cross-attention result) is input into the second encoding layer in the high-resolution editing sub-network to obtain the second encoding result. At this time, the second cross-attention layer in the 3D editing module can be used, combined with the initial editing semantic vector and the second encoding result, to obtain cross-attention result-22. The cross-attention result-22 is then fused with the second encoding result to obtain the second encoding result carrying cross-attention result-22. Further, the second encoding result carrying cross-attention result-22 is input into the second decoding layer in the high-resolution editing sub-network to obtain the second decoding result. At this time, the second decoding result is the initial 3D latent representation.
[0096] Here, the network structure and processing logic of the high-resolution editing subnetwork described above are merely illustrative examples. As long as the restoration of local appearance information in the fused three-dimensional latent representation can be achieved, this disclosure does not impose any specific restrictions on it.
[0097] It should also be noted that, in one example, the 3D editing module further includes an adaptive normalization layer. In this case, before using the aforementioned cross-attention layer (such as the first cross-attention layer or the second cross-attention layer) for cross-attention processing, the adaptive normalization layer can be used to normalize the initial editing semantic vector representing the initial editing instruction information to obtain the normalized initial editing semantic vector. The normalized initial editing semantic vector is then used to participate in the processing performed by the aforementioned cross-attention layer.
[0098] Figure 7 This is an illustrative flowchart of a model training method for an editing model applied to a 3D Gaussian scene according to an embodiment of this application. Figure 3 This method can be optionally applied to electronic devices, such as personal computers, servers, and server clusters. It is understood that the above... Figures 1 to 6 The methods shown can also be applied to this example, and the related content will not be elaborated further in this example.
[0099] Furthermore, the method includes at least a portion of the following: For example... Figure 7 As shown, it includes: Step S701: Obtain the target training sample.
[0100] Here, the target training samples include initial editing instruction information, initial scene images corresponding to each of the N (N is an integer greater than or equal to 2) viewpoints, and reference scene images corresponding to each initial scene image.
[0101] Furthermore, the initial editing instruction information is used to instruct the modification of the appearance parameters (such as lighting, materials, etc.) of the initial 3D Gaussian scene corresponding to the initial scene image.
[0102] Furthermore, the initial scene images corresponding to the N viewpoints are used to describe the initial three-dimensional Gaussian scene under different viewpoints.
[0103] Furthermore, the reference scene image corresponding to the initial scene image is used to represent the label scene image after the appearance parameters of the initial three-dimensional Gaussian scene have been modified based on the initial editing instruction information.
[0104] It should be noted that the information regarding the initial editing instructions and the initial 3D Gaussian scene can be found in the example above, and will not be repeated here.
[0105] Step S702: Perform multi-view encoding processing on the initial scene images corresponding to the N viewpoints to obtain the three-dimensional latent representation to be edited.
[0106] Here, the three-dimensional latent representation to be edited is a unified representation that maps the initial three-dimensional Gaussian scene to the latent space and integrates the N perspectives.
[0107] It should be noted that the specific processing logic for multi-view encoding can be found in the example above, and will not be repeated here.
[0108] Step S703: Using the 3D editing module in the scene editing model to be trained, and based on the initial editing semantic vector representing the initial editing instruction information, the 3D latent representation to be edited is edited to obtain the initial 3D latent representation.
[0109] It should be noted that the relevant content regarding the 3D editing module can be found in the example above, and will not be repeated here.
[0110] Step S704: Based on the edit type corresponding to the initial edit semantic vector, call at least one initial task header.
[0111] Here, the task adaptation module in the scene editing model to be trained includes a local appearance editing task head and a global illumination editing task head; further, the initial task head is one of the local appearance editing task head and the global illumination editing task head.
[0112] It should be noted that, in order to further improve the editing quality of 3D Gaussian scenes, in practical applications, the local appearance editing task head can be divided into multiple task heads for specific local appearance editing tasks, such as object addition / deletion task heads and material change task heads. The above is only an example; other editing task heads can also be divided for editing local appearances. This disclosure does not impose specific restrictions on the division method and number of local appearance editing task heads.
[0113] Step S705: Using at least one invoked initial task header and based on the initial edit semantic vector, perform a correction process on the initial three-dimensional latent representation to obtain the target three-dimensional latent representation.
[0114] Step S706: Based on the target's three-dimensional latent representation, obtain the estimated scene image corresponding to each of the N viewpoints.
[0115] Step S707: Based on the estimated scene image corresponding to each viewpoint and the reference scene image corresponding to each viewpoint, obtain the target loss value. Based on the target loss value, train the 3D editing module and the task adaptation module in the scene editing model to be trained to obtain the target scene editing model.
[0116] For example, such as Figure 8 As shown, given the 3D latent representation to be edited and the initial editing semantic vector representing the initial editing instruction information, the 3D editing module in the scene editing model to be trained, combined with the initial editing semantic vector, is used to edit the 3D latent representation to be edited to obtain the initial 3D latent representation. Before using the task adaptation module in the scene editing model to be trained, at least one initial task head (such as a local appearance editing task head and / or a global illumination editing task head) is called according to the editing type corresponding to the initial editing semantic vector. Then, the initial 3D latent representation is corrected based on the called initial task head and the initial editing semantic vector to obtain the target 3D latent representation. Further, the estimated scene image corresponding to each viewpoint is obtained using the target 3D latent representation. Based on the estimated scene image corresponding to each viewpoint and the reference scene image corresponding to each viewpoint (i.e., the reference scene image corresponding to each initial scene image), the target loss value is obtained. Finally, based on the target loss value, the parameters of the 3D editing module and the task adaptation module in the scene editing model to be trained are fine-tuned to obtain the target scene editing model.
[0117] In this way, the proposed solution achieves "on-demand" fine-grained correction through the dual editing task heads of local appearance and global illumination set in the task adaptation module. Moreover, it automatically identifies the editing type based on the initial editing semantic vector and then dynamically calls the task head corresponding to the editing type to correct the initial 3D latent representation. This effectively solves the problem of excessive divergence in the 3D editing module when performing global editing, significantly improves the model's adaptability to different editing intentions, and enhances the model's generalization and stability in full-scene-level editing of complex 3D Gaussian scenes.
[0118] The following sections describe the specific processing details for obtaining the target's 3D latent representation, based on the type of initial task header invoked: (1) The initial task header called is the local appearance editing task header. In a specific example, the target three-dimensional latent representation can be obtained in the following manner; specifically, the above-described method of using at least one invoked initial task header and, based on the initial edit semantic vector, refining the initial three-dimensional latent representation to obtain the target three-dimensional latent representation (e.g., step S705) can specifically include: Step S7051-1: When the initial task head called is a local appearance editing task head, the editing residual is obtained by using the local appearance editing task head and based on the initial scene image corresponding to each viewpoint and the reference scene image corresponding to the initial scene image.
[0119] Here, the edit residual is used to quantify the "modifications to be made" between the initial scene image and the corresponding reference scene image. This provides a clear direction for subsequent correction of the initial 3D latent representation using the task adaptation module.
[0120] Step S7051-2: Generate spatial gating features based on the edit residual and the initial edit semantic vector.
[0121] For example, in one instance, prior features of the 3D editing region are obtained from 2D image differences (i.e., editing residuals). Specifically, for each of the N viewpoints, a pixel difference map between the initial scene image and the reference scene image is calculated. Then, the voxel centers of a pre-defined voxel grid are projected onto the pixel difference map corresponding to each viewpoint, and samples are collected on the pixel difference map corresponding to each viewpoint and aggregated across viewpoints to obtain low-resolution prior features of the 3D editing region. In this way, 2D differences are improved into 3D voxel-level priors, i.e., 3D editing region prior features, through multi-view backprojection and voxel aggregation.
[0122] Furthermore, based on the prior features of the 3D editing region and the initial editing semantic vector, 3D spatial gating features are generated. Specifically, the prior features of the 3D editing region and the initial editing semantic vector are first input into the local appearance editing task head. Using this local appearance editing task head, the initial editing semantic vector is mapped to a global bias term. The 3D editing region prior features are transformed and fused with the bias term, and then activated by an activation function to obtain low-resolution spatial gating features. The low-resolution spatial gating features are then interpolated and upsampled to the same resolution as the initial 3D latent representation to obtain full-resolution spatial gating features.
[0123] Here, the spatial gating feature represents the editable region in the three-dimensional latent representation to be edited, the editing intensity of voxels in the editable region, and the non-editable region.
[0124] Furthermore, in one example, the edit region represents the region containing one or more voxels in the three-dimensional latent representation to be edited that need to be edited; correspondingly, the non-edit region represents the region containing one or more voxels in the three-dimensional latent representation to be edited that do not need to be edited.
[0125] Step S7051-3: Based on the spatial gating features, restore the non-editable regions in the initial three-dimensional latent representation to obtain the target three-dimensional latent representation.
[0126] It should be noted that although the 3D editing module in this disclosure can complete the editing of the 3D latent representation to be edited based on the initial editing instructions, its divergent nature makes it prone to "collateral damage" when editing the 3D latent representation: while editing the voxels that need editing, it generates unexpected perturbations on voxels that do not need editing, which obviously deviates from the user's editing intention. Therefore, this disclosure introduces editing residuals to quantify "content to be modified" during the model training phase, and generates spatial gating features in conjunction with the initial editing semantic vector. Then, based on the spatial gating features, the initial 3D latent representation is modified differentially. For example, unexpected changes in non-editable regions are specifically restored, while the modification effects in editable regions are retained. In this way, the divergent behavior of the model is constrained at the training level to guide the model to learn the ability to edit only in editable regions, thereby significantly improving the accuracy, controllability, and fidelity of the model's full-scene-level editing.
[0127] Thus, this disclosed solution introduces editing residuals to construct a precise correction mechanism of "pixel difference-guided latent space gating." This mechanism utilizes generated spatial gating features to endow the model with fine-grained spatial discrimination capabilities, accurately locating non-editable and editable regions in the initial 3D latent representation. It then restores the "content" in the non-editable regions and preserves the "content" in the editable regions. This effectively overcomes the pain points of "over-modification" or "false positives" in semantic-driven editing. While ensuring the accurate implementation of editing instructions, it preserves as much detail as possible in the original scene, achieving a high degree of consistency between geometric information and semantic intent. This significantly improves the accuracy and controllability of the model's full-scene-level editing, making the obtained target 3D latent representation more closely match the requirements of full-scene-level editing.
[0128] Furthermore, in a specific example, the above-described process of restoring the non-editable regions in the initial 3D latent representation based on the spatial gating features to obtain the target 3D latent representation (e.g., step S7051-3) may specifically include: Step S7051-3-1: Based on the spatial gating features, obtain the three-dimensional mask matrix.
[0129] For example, in one example, for the edit region represented by the spatial gating feature, a mask value can be set for each voxel based on its editing intensity within the edit region; for instance, the editing intensity of the voxels in the edit region can be directly used as the mask value for that voxel. For the non-edit region represented by the spatial gating feature, the mask value of each voxel in the non-edit region is set to a fixed value (such as "0") to obtain a three-dimensional mask matrix containing the mask values of each voxel. This provides a basis for subsequently focusing editing or updates within the edit region while simultaneously suppressing drift phenomena in the non-edit region.
[0130] The above is merely an example. In practical applications, for the editing area, the mask value of the voxel can be set to another fixed value (such as "1") based on the editing intensity of the voxel in the editing area. In this way, a three-dimensional mask matrix containing 0 and 1 can be obtained. This disclosure does not impose any specific limitations on this.
[0131] Step S7051-3-2: Based on the three-dimensional mask matrix, the initial three-dimensional latent representation and the three-dimensional latent representation to be edited are weighted to restore the non-edited region and obtain the target three-dimensional latent representation.
[0132] For example, in one example, let the 3D mask matrix be M, the 3D latent representation to be edited be z_src, and the initial 3D latent representation be z_edit_raw. If the target 3D latent representation is denoted as z_edit_final, then the target 3D latent representation z_edit_final can be obtained using the following formula: In this way, the proposed solution uses a three-dimensional mask matrix to weight the initial three-dimensional latent representation and the three-dimensional latent representation to be edited. This concentrates the editing changes in the editing area, while the content in the non-editing area retains the original scene state. This effectively overcomes the pain points of "over-modification" or "false positives" in semantic-driven editing and significantly improves the controllability of the model's full-scene-level editing.
[0133] (2) The initial task header called is the global illumination editing task header. In a specific example, the target three-dimensional latent representation can be obtained in the following manner; specifically, the above-described method of using at least one invoked initial task header and, based on the initial edit semantic vector, refining the initial three-dimensional latent representation to obtain the target three-dimensional latent representation (e.g., step S705) can specifically include: Step S7052-1: If the initial task header being invoked is a global illumination editing task header, the global color features are obtained by using the global illumination editing task header and combining it with the initial editing instruction information.
[0134] Step S7052-2: Based on the global color features, process the initial three-dimensional latent representation to obtain the target three-dimensional latent representation.
[0135] For example, in one example, the global color features are first mapped to affine transformation parameters, which include channel scaling factors for controlling illumination intensity and channel offsets for controlling illumination color; then, based on the affine transformation parameters, the initial 3D latent representation is subjected to affine transformation processing to obtain the target 3D latent representation.
[0136] For example, let the initial 3D latent representation be z_edit_raw, the channel scaling factor used to control illumination intensity in the affine transformation parameters be scale, and the channel offset used to control illumination color in the affine transformation parameters be bias. Then, if the target 3D latent representation is... Then for the target's three-dimensional latent representation For a voxel, the target's three-dimensional latent representation The affine transformation result corresponding to this voxel (which can be denoted as) The following formula can be used to obtain it: here, This represents the element-wise multiplication operator.
[0137] Furthermore, for the target's three-dimensional latent representation For each voxel in the model, after completing the affine transformation of all voxels, the three-dimensional latent representation of the target can be obtained. .
[0138] In this way, the disclosed solution achieves professional modeling and unified control of lighting attributes through a global illumination editing task head. Specifically, the global illumination editing task head extracts global color features from the initial editing semantic vector and applies them to the initial 3D latent representation, ensuring the consistency and rationality of lighting changes. Moreover, compared to local appearance editing, the global illumination editing task head focuses on low-frequency global attributes such as ambient light and tone mapping, effectively avoiding light and shadow fragmentation or color discontinuity caused by interference from local details. Thus, while maintaining the geometric structure and local texture of the 3D Gaussian scene, it achieves natural and stable lighting style transfer, significantly improving the model's full-scene-level lighting editing capabilities, thereby making the obtained target 3D latent representation closer to the requirements of full-scene-level editing.
[0139] It should be noted that, in one example, the global illumination editing task head and the local appearance editing task head can be invoked simultaneously according to the editing type corresponding to the initial editing semantic vector. In this case, in one example, the above-described process of using at least one invoked initial task head and, based on the initial editing semantic vector, to correct the initial 3D latent representation to obtain the target 3D latent representation can specifically include: using the local appearance editing task head and, based on the initial editing semantic vector, correcting the initial 3D latent representation to obtain the locally corrected initial 3D latent representation; and using the global illumination editing task head and, based on the initial editing semantic vector, correcting the locally corrected initial 3D latent representation to obtain the target 3D latent representation.
[0140] In the above example, besides calling the local appearance editing task head first and then the global illumination editing task head, the present disclosure can also call the global illumination editing task head first and then the local appearance editing task head, or call the two task heads in parallel, and perform fusion processing based on the processing results of each task head to obtain the target's three-dimensional latent representation. The present disclosure does not impose specific limitations on this.
[0141] Figure 9 This is an illustrative flowchart of a model training method for an editing model applied to a 3D Gaussian scene according to an embodiment of this application. Figure 3This method can be optionally applied to electronic devices, such as personal computers, servers, and server clusters. It is understood that the above... Figures 1 to 8 The methods shown can also be applied to this example, and the related content will not be elaborated further in this example.
[0142] Furthermore, the method includes at least a portion of the following: For example... Figure 9 As shown, it includes: Step S901: Obtain the target training sample.
[0143] Here, the target training samples include initial editing instruction information, initial scene images corresponding to each of the N (N is an integer greater than or equal to 2) viewpoints, and reference scene images corresponding to each initial scene image.
[0144] Furthermore, the initial editing instruction information is used to instruct the modification of the appearance parameters (such as lighting, materials, etc.) of the initial 3D Gaussian scene corresponding to the initial scene image.
[0145] Furthermore, the initial scene images corresponding to the N viewpoints are used to describe the initial three-dimensional Gaussian scene under different viewpoints.
[0146] Furthermore, the reference scene image corresponding to the initial scene image is used to represent the label scene image after the appearance parameters of the initial three-dimensional Gaussian scene have been modified based on the initial editing instruction information.
[0147] It should be noted that the information regarding the initial editing instructions and the initial 3D Gaussian scene can be found in the example above, and will not be repeated here.
[0148] Step S902: Perform multi-view encoding processing on the initial scene images corresponding to the N viewpoints to obtain the three-dimensional latent representation to be edited.
[0149] Here, the three-dimensional latent representation to be edited is a unified representation that maps the initial three-dimensional Gaussian scene to the latent space and integrates the N perspectives.
[0150] It should be noted that the specific processing logic for multi-view encoding can be found in the example above, and will not be repeated here.
[0151] Step S903: Input the three-dimensional latent representation to be edited and the initial editing semantic vector representing the initial editing instruction information into the scene editing model to be trained, so as to use the three-dimensional editing module in the scene editing model to be trained and, based on the initial editing semantic vector, edit the three-dimensional latent representation to be edited to obtain the initial three-dimensional latent representation; and, use the task adaptation module in the scene editing model to be trained and, based on the initial editing semantic vector, correct the initial three-dimensional latent representation to obtain the target three-dimensional latent representation.
[0152] It should be noted that the relevant content regarding the 3D editing module and task adaptation module in the scene editing model to be trained can be referred to the above example, and will not be repeated here.
[0153] Step S904: Using a preset Gaussian decoding model, perform Gaussian decoding on the target three-dimensional latent representation to obtain the initial three-dimensional Gaussian scene after editing based on the initial editing instruction information.
[0154] Here, in one example, the preset Gaussian decoding model may specifically be a decoder used to decode a 3D latent representation into a 3D Gaussian scene, or it may specifically be other models used to convert a 3D latent representation into a 3D Gaussian scene. This disclosure does not impose any specific limitations on this.
[0155] Furthermore, in one example, the initial three-dimensional Gaussian scene edited based on the initial editing instruction information can specifically be the scene obtained by editing the initial three-dimensional Gaussian scene corresponding to the initial scene image using the above steps S902 to S904 and in combination with the initial editing instruction information.
[0156] Step S905: Based on the edited initial 3D Gaussian scene, perform rendering processing to obtain the estimated scene images corresponding to each viewpoint.
[0157] For example, in one example, the rendering engine is invoked, and the initial 3D Gaussian scene is rendered based on the edited version to obtain estimated scene images for each viewpoint.
[0158] Step S906: Based on the estimated scene image corresponding to each viewpoint and the reference scene image corresponding to each viewpoint, obtain the target loss value. Based on the target loss value, train the 3D editing module and the task adaptation module in the scene editing model to be trained to obtain the target scene editing model.
[0159] In this way, after obtaining the target 3D latent representation, the present solution uses a preset Gaussian decoding model to decode the target 3D latent representation into an edited initial 3D Gaussian scene, and performs rendering processing based on the edited initial 3D Gaussian scene to obtain the estimated scene image corresponding to each viewpoint. Then, the target loss value between the estimated scene image and the reference scene image corresponding to each viewpoint is calculated. In this way, the target loss value can be used to back-optimize the ability of the scene editing model to be trained to edit the 3D Gaussian scene, which significantly improves the robustness and stability of full-scene-level editing during model training, thus providing favorable support for subsequent use of the target scene editing model to achieve viewpoint-free editing of 3D Gaussian scenes.
[0160] Furthermore, in a specific example, the above-described model training based on the target loss value for the 3D editing module and the task adaptation module in the scene editing model to be trained can specifically include: Based on the target loss value, the preset Gaussian decoding model, as well as the 3D editing module and task adaptation module in the scene editing model to be trained, are jointly trained to obtain the target scene editing model and the target Gaussian decoding model.
[0161] In this way, the proposed solution can utilize the target loss value to train the preset Gaussian decoding model and the 3D editing module and task adaptation module in the scene editing model to be trained simultaneously. This effectively eliminates the scene representation bias caused by the insufficient scene decoding performance of the Gaussian decoding model, significantly improves the stability and robustness of model training, and efficiently obtains a target scene editing model with the ability to edit 3D Gaussian scenes without perspective shifts. This provides favorable support for the subsequent realization of 3D Gaussian scene editing without perspective shifts and improves the user's editing efficiency and experience.
[0162] For example, continue with Figure 2 Taking the training logic of the scene editing model shown as an example, as follows: Figure 10As shown, when the scene editing model outputs a target 3D latent representation, a preset Gaussian decoding model can be used to perform Gaussian decoding on the output target 3D latent representation to obtain an edited initial 3D Gaussian scene. Further, the edited initial 3D Gaussian scene is differentiable to obtain estimated scene images corresponding to each viewpoint. Then, based on the estimated scene images corresponding to each viewpoint and the reference scene images corresponding to each viewpoint (i.e., the reference scene images corresponding to each initial scene image), the target loss value can be calculated. Using the obtained target loss value, the parameters of the preset Gaussian decoding model, as well as the 3D editing module and task adaptation module in the scene editing model to be trained, are fine-tuned to obtain the target scene editing model and the target Gaussian decoding model. This effectively eliminates the scene representation bias caused by insufficient scene decoding performance of the Gaussian decoding model, significantly improves the stability and robustness of the model during training, and provides favorable support for subsequent full-scene-level editing using the model.
[0163] Figure 11 This is a schematic flowchart illustrating a spatial editing method applied to a 3D Gaussian scene according to an embodiment of this application. This method can be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.
[0164] Furthermore, the method includes at least a portion of the following: For example... Figure 11 As shown, it includes: Step S1101: Obtain the scene image to be processed corresponding to each of the M viewpoints, as well as the target editing instruction information.
[0165] Here, the scene images to be processed corresponding to the M viewpoints are used to describe the target 3D Gaussian scene under different viewpoints; the target editing instruction information is used to instruct the modification of the appearance parameters of the target 3D Gaussian scene corresponding to the scene images to be processed.
[0166] Here, M is an integer greater than or equal to 2; furthermore, in one example, M and N can be the same or different.
[0167] It should be noted that the relevant content regarding target editing command information and target 3D Gaussian scene can be found in the examples of initial editing command information and initial 3D Gaussian scene mentioned above, and will not be repeated here.
[0168] Step S1102: Perform multi-view encoding processing on the scene images to be processed corresponding to the M viewpoints to obtain the first three-dimensional latent representation.
[0169] Here, the first three-dimensional latent representation is a unified representation of the target three-dimensional Gaussian scene mapped to the latent space and integrating the M perspectives.
[0170] It should be noted that the processing logic for multi-view encoding can be referred to the above example, and will not be repeated here.
[0171] Step S1103: Input the first three-dimensional latent representation and the target editing semantic vector representing the target editing instruction information into the target scene editing model to obtain the second three-dimensional latent representation.
[0172] Here, the target scene editing model is obtained by training the scene editing model to be trained using any of the above embodiments.
[0173] It should be noted that the model architecture and processing logic of the target scene editing model can be referred to the model structure and processing logic of the scene editing model to be trained mentioned above, and will not be repeated here.
[0174] Step S1104: Based on the second three-dimensional latent representation, obtain the target three-dimensional Gaussian scene after editing based on the target editing instruction information.
[0175] For example, in one instance, the second three-dimensional latent representation can be decoded using a target Gaussian decoding model to obtain a target three-dimensional Gaussian scene edited based on the target editing instruction information.
[0176] Here, the target 3D Gaussian scene edited based on the target editing instruction information can specifically be the scene obtained by editing the target 3D Gaussian scene corresponding to the scene image to be processed using the above steps S1102 to S1104 and in combination with the target editing instruction information.
[0177] In this way, the present invention first decouples the target scene images corresponding to M viewpoints into a first three-dimensional latent representation independent of the viewpoint. Then, using the target scene editing model and combining it with the target editing semantic vector representing the target editing instruction information, the first three-dimensional latent representation is edited to obtain the edited second three-dimensional latent representation, and then decoded to obtain the edited target three-dimensional Gaussian scene. Thus, the present invention eliminates the need for viewpoint-by-view editing of the three-dimensional Gaussian scene, significantly improving the efficiency of scene editing. At the same time, it effectively avoids the problems of poor geometric consistency and artifact drift between viewpoints caused by two-dimensional viewpoint editing, thereby improving the quality of editing. In addition, it greatly reduces the computing resources required for scene editing, thus meeting the real-time or near-real-time business needs of online interaction, batch home decoration scheme generation, etc.
[0178] Furthermore, the target scene editing model in this disclosed solution is equipped with task heads for various editing types (such as global illumination, object addition / deletion, and material replacement). Thus, compared to the existing approach of using the same workflow to handle different editing types of tasks, this disclosed solution uniformly supports multiple editing types. In other words, the editing effect presented is more stable when dealing with different editing types of tasks, and users do not need to switch editing workflows, thereby significantly improving the efficiency of users performing full-scene-level editing and enhancing the user experience.
[0179] Furthermore, since this disclosed solution decouples the target scene images corresponding to M viewpoints into a first three-dimensional latent representation independent of the viewpoint, the capacity of the three-dimensional Gaussian volume in the three-dimensional Gaussian scene in this disclosed solution is no longer limited by the number of input viewpoints or the fixed token length. Compared with existing reconstruction-guided models, this disclosed solution can stably generate a three-dimensional scene containing hundreds of thousands of three-dimensional Gaussian volumes, and can support full-scene-level editing of hundreds of thousands of three-dimensional Gaussian volumes, thereby meeting the user's editing needs.
[0180] For example, Figure 12 This is a schematic diagram of the scene editing method of this disclosure in a specific example; such as Figure 12 As shown, firstly, a modal encoder is used to encode the target editing instructions, obtaining a target editing semantic vector representing the target editing instruction information. Then, multi-view encoding processing is performed on the scene images to be processed corresponding to each of the M viewpoints to obtain a first 3D latent representation. Secondly, the first 3D latent representation and the target editing semantic vector representing the target editing instruction information are input into a target scene editing model (including a 3D editing module and a task adaptation module) to edit the first 3D latent representation according to the target editing instruction information, thereby obtaining a second 3D latent representation. Finally, a target Gaussian decoding model is used to perform Gaussian decoding processing on the second 3D latent representation to obtain the edited target 3D Gaussian scene. In this way, there is no need to edit the 3D Gaussian scene from one viewpoint at a time, significantly improving the user's editing efficiency and thus enhancing the user experience.
[0181] It should be noted that, in practical applications, when calling the local appearance editing task header in the task adaptation module, this disclosed solution can, after the local appearance editing task header outputs the second 3D latent representation, overlay the 3D voxel space view corresponding to the second 3D latent representation onto the target 3D Gaussian scene as a semi-transparent voxel thermal layer to obtain a 3D Gaussian scene that can be previewed by the user (which can be called a preview 3D Gaussian scene). This allows the user to view the editing range of the scene from different perspectives. Furthermore, if it is determined that the user has confirmed the preview of the 3D Gaussian scene, the Gaussian decoding process for the second 3D latent representation is further executed.
[0182] Furthermore, in one example, for local appearance editing involving adding or deleting objects, this disclosed solution, while displaying a preview of a 3D Gaussian scene, allows users to set object placement anchor points within the editing area of that preview scene. This significantly enhances the user's editing freedom. For instance, in one example, for local appearance editing involving changing materials, this disclosed solution, while displaying a preview of a 3D Gaussian scene, allows users to select an object surface within the editing area of that preview scene to replace the material on the selected object surface. Alternatively, in another example, for global illumination editing, this disclosed solution, in addition to being able to invoke the global illumination editing task head for lighting (or color) editing based on editing command information (such as text commands), also provides adjustable sliders for color temperature, brightness, etc., on the front-end editing interface. Furthermore, when the user adjusts these sliders, the slider changes can be mapped to global color transformation parameters in real time, and the global color transformation parameters can be quickly applied to the second 3D latent representation, thus meeting the user's customized editing needs.
[0183] This disclosure provides a model training device for editing models applied to 3D Gaussian scenes, such as... Figure 13 As shown, it includes: The sample acquisition unit 1301 is used to acquire target training samples, wherein the target training samples include initial editing instruction information, initial scene images corresponding to each of the N viewpoints, and reference scene images corresponding to each initial scene image; the initial editing instruction information is used to instruct the modification of the appearance parameters of the initial 3D Gaussian scene corresponding to the initial scene image; the initial scene images corresponding to the N viewpoints are used to describe the initial 3D Gaussian scene under different viewpoints, and the reference scene images corresponding to the initial scene images are used to represent the label scene image after the appearance parameters of the initial 3D Gaussian scene are modified based on the initial editing instruction information; N is an integer greater than or equal to 2; The model training unit 1302 is used to perform multi-view encoding processing on the initial scene images corresponding to the N viewpoints to obtain a three-dimensional latent representation to be edited; wherein, the three-dimensional latent representation to be edited is a unified representation that maps the initial three-dimensional Gaussian scene to the latent space and integrates the N viewpoints; it is used to input the three-dimensional latent representation to be edited and the initial editing semantic vector representing the initial editing instruction information into the scene editing model to be trained, so as to use the three-dimensional editing module in the scene editing model to be trained and, based on the initial editing semantic vector, to edit the three-dimensional latent representation to be edited to obtain the initial three-dimensional latent representation; and, using the task adaptation module in the scene editing model to be trained and, based on the initial editing semantic vector, to correct the initial three-dimensional latent representation to obtain the target three-dimensional latent representation; based on the target three-dimensional latent representation, to obtain the estimated scene image corresponding to each of the N viewpoints; based on the estimated scene image corresponding to each viewpoint and the reference scene image corresponding to each viewpoint, to obtain the target loss value, so as to train the three-dimensional editing module and the task adaptation module in the scene editing model to be trained based on the target loss value, to obtain the target scene editing model.
[0184] In a specific example of the scheme disclosed herein, the 3D editing module in the scene editing model to be trained includes a low-resolution editing subnetwork for global editing of the 3D latent representation, and a high-resolution editing subnetwork for recovering local appearance information in the 3D latent representation; Specifically, the model training unit is used for: Using the low-resolution editing sub-network in the 3D editing module, and combining the initial editing semantic vector and the 3D latent representation to be edited, voxel variation features are obtained. Based on the voxel variation features and the 3D latent representation to be edited, a fused 3D latent representation is obtained. The voxel variation features represent the spatial distribution and editing variation of the voxels to be edited in the 3D latent representation to be edited. The high-resolution editing sub-network in the 3D editing module is used to restore the local appearance information in the fused 3D latent representation to obtain the initial 3D latent representation.
[0185] In a specific example of the scheme disclosed herein, the 3D editing module in the scene editing model to be trained further includes a first cross-attention layer; The model training unit is further configured to: The first cross-attention layer in the 3D editing module is used, and the initial editing semantic vector and the 3D latent representation to be edited are combined to obtain the first cross-attention result; The first cross-attention result is fused with the three-dimensional latent representation to be edited to obtain the three-dimensional latent representation to be edited carrying the first cross-attention result; Using the low-resolution editing sub-network in the 3D editing module, and combining the 3D latent representation to be edited carrying the first cross-attention result with the initial editing semantic vector, voxel change features are obtained. Based on the voxel change features and the 3D latent representation to be edited, a fused 3D latent representation is obtained.
[0186] In a specific example of the scheme disclosed herein, the 3D editing module in the scene editing model to be trained further includes a second cross-attention layer; The model training unit is further configured to: The second cross-attention layer in the 3D editing module is used, and the initial editing semantic vector and the fused 3D latent representation are combined to obtain the second cross-attention result; The second cross-attention result is fused with the fused 3D latent representation to obtain a fused 3D latent representation carrying the second cross-attention result; The initial three-dimensional latent representation is obtained by utilizing the high-resolution editing sub-network in the three-dimensional editing module and combining it with the fused three-dimensional latent representation carrying the second cross-attention result and the initial editing semantic vector.
[0187] In a specific example of the scheme disclosed herein, the task adaptation module in the scene editing model to be trained includes a local appearance editing task head and a global illumination editing task head; Specifically, the model training unit is used for: Based on the editing type corresponding to the initial editing semantic vector, at least one initial task head is invoked; wherein, the initial task head is one of the local appearance editing task head and the global lighting editing task head; Using at least one invoked initial task header and based on the initial edit semantic vector, the initial 3D latent representation is modified to obtain the target 3D latent representation.
[0188] In a specific example of the scheme disclosed herein, the model training unit is specifically used for: When the initial task head called is a local appearance editing task head, the editing residual is obtained by using the local appearance editing task head and based on the initial scene image corresponding to each view and the reference scene image corresponding to the initial scene image. Based on the edit residual and the initial edit semantic vector, spatial gating features are generated, wherein the spatial gating features represent the edit region in the three-dimensional latent representation to be edited, the edit intensity of voxels in the edit region, and the non-edit region; Based on the spatial gating features, the non-editable regions in the initial three-dimensional latent representation are restored to obtain the target three-dimensional latent representation.
[0189] In a specific example of the scheme disclosed herein, the model training unit is specifically used for: Based on the aforementioned spatial gating features, a three-dimensional mask matrix is obtained; Based on the three-dimensional mask matrix, the initial three-dimensional latent representation and the three-dimensional latent representation to be edited are weighted to restore the non-editable region and obtain the target three-dimensional latent representation.
[0190] In a specific example of the scheme disclosed herein, the model training unit is specifically used for: When the initial task header invoked is a global illumination editing task header, the global color features are obtained by using the global illumination editing task header and combining it with the initial editing instruction information; Based on the global color features, the initial three-dimensional latent representation is processed to obtain the target three-dimensional latent representation.
[0191] In a specific example of the scheme disclosed herein, the model training unit is specifically used for: Using a preset Gaussian decoding model, the target three-dimensional latent representation is subjected to Gaussian decoding to obtain an initial three-dimensional Gaussian scene after editing based on the initial editing instruction information; Based on the edited initial 3D Gaussian scene, rendering is performed to obtain estimated scene images corresponding to each viewpoint.
[0192] In a specific example of the scheme disclosed herein, the model training unit is specifically used for: Based on the target loss value, the preset Gaussian decoding model, as well as the 3D editing module and task adaptation module in the scene editing model to be trained, are jointly trained to obtain the target scene editing model and the target Gaussian decoding model.
[0193] In a specific example of the scheme disclosed herein, the model training unit is specifically used for: Feature extraction is performed on the initial scene images corresponding to each viewpoint to obtain the feature maps corresponding to each initial scene image; The preset three-dimensional voxel grid is projected onto the feature map corresponding to each initial scene image, and the feature blocks corresponding to the voxels in the preset three-dimensional voxel grid on each feature map are sampled. The feature blocks corresponding to the voxels in the preset three-dimensional voxel grid on each feature map are fused to obtain the fused features corresponding to each voxel in the preset three-dimensional voxel grid. The fusion features corresponding to each voxel in the preset three-dimensional voxel mesh are processed to obtain the three-dimensional latent representation to be edited.
[0194] This disclosure also provides a spatial editing device for three-dimensional Gaussian scenes, such as... Figure 14 As shown, it includes: The data input unit 1401 is used to acquire the scene images to be processed corresponding to each of the M viewpoints, as well as the target editing instruction information; wherein, the scene images to be processed corresponding to the M viewpoints are used to describe the target three-dimensional Gaussian scene under different viewpoints; the target editing instruction information is used to instruct the modification of the appearance parameters of the target three-dimensional Gaussian scene corresponding to the scene images to be processed; M is an integer greater than or equal to 2; The spatial editing unit 1402 is used to perform multi-view encoding processing on the scene images to be processed corresponding to the M viewpoints to obtain a first three-dimensional latent representation; wherein, the first three-dimensional latent representation is a unified representation of the target three-dimensional Gaussian scene mapped to the latent space and integrating the M viewpoints; the first three-dimensional latent representation and the target editing semantic vector representing the target editing instruction information are input into the target scene editing model to obtain a second three-dimensional latent representation; wherein, the target scene editing model is obtained after training using any of the above model training methods; based on the second three-dimensional latent representation, the target three-dimensional Gaussian scene edited based on the target editing instruction information is obtained.
[0195] For a description of the specific functions and examples of each unit of the apparatus in this disclosure embodiment, please refer to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be repeated here.
[0196] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0197] Figure 15 This is a structural block diagram of an electronic device according to an embodiment of the present disclosure. Figure 15 As shown, the electronic device includes a memory 1510 and a processor 1520. The memory 1510 stores a computer program that can run on the processor 1520. The number of memories 1510 and processors 1520 can be one or more. The memory 1510 can store one or more computer programs, which, when executed by the electronic device, cause the electronic device to perform the methods provided in the above-described method embodiments. The electronic device may also include a communication interface 1530 for communicating with external devices and performing data exchange and transmission.
[0198] If the memory 1510, processor 1520, and communication interface 1530 are implemented independently, they can be interconnected via a bus to communicate with each other. This bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 15 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0199] Optionally, in a specific implementation, if the memory 1510, processor 1520 and communication interface 1530 are integrated on a single chip, the memory 1510, processor 1520 and communication interface 1530 can communicate with each other through an internal interface.
[0200] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Machines (ARM) architecture.
[0201] Further, optionally, the aforementioned memory may include read-only memory and random access memory, and may also include non-volatile random access memory. The memory may be volatile or non-volatile, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. Many forms of RAM are available by way of example, but not limitation. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct RAMBUS RAM (DR RAM).
[0202] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line, DSL) or wireless (e.g., infrared, Bluetooth, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer, or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., Solid State Disks (SSDs)). It is worth noting that the computer-readable storage media mentioned in this disclosure can be non-volatile storage media; in other words, it can be non-transient storage media.
[0203] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0204] In the description of the embodiments of this disclosure, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.
[0205] In the description of the embodiments disclosed herein, unless otherwise stated, " / " means "or". For example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone.
[0206] In the description of embodiments of this disclosure, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more.
[0207] The above description is merely an exemplary embodiment of this disclosure and is not intended to limit this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the protection scope of this disclosure.
Claims
1. A model training method for editing models applied to 3D Gaussian scenes, comprising: Obtain target training samples, wherein the target training samples include initial editing instruction information, initial scene images corresponding to each of the N viewpoints, and reference scene images corresponding to each initial scene image; the initial editing instruction information is used to instruct the modification of the appearance parameters of the initial 3D Gaussian scene corresponding to the initial scene image; the initial scene images corresponding to the N viewpoints are used to describe the initial 3D Gaussian scene under different viewpoints, and the reference scene images corresponding to the initial scene images are used to represent the labeled scene image after the appearance parameters of the initial 3D Gaussian scene are modified based on the initial editing instruction information; N is an integer greater than or equal to 2; The initial scene images corresponding to the N viewpoints are subjected to multi-view encoding processing to obtain the three-dimensional latent representation to be edited; wherein, the three-dimensional latent representation to be edited is a unified representation that maps the initial three-dimensional Gaussian scene to the latent space and integrates the N viewpoints. The three-dimensional latent representation to be edited and the initial editing semantic vector representing the initial editing instruction information are input into the scene editing model to be trained. The three-dimensional editing module in the scene editing model to be trained is used to edit the three-dimensional latent representation to be edited based on the initial editing semantic vector to obtain the initial three-dimensional latent representation. Then, the task adaptation module in the scene editing model to be trained is used to correct the initial three-dimensional latent representation based on the initial editing semantic vector to obtain the target three-dimensional latent representation. Based on the target's three-dimensional latent representation, estimated scene images corresponding to each of the N viewpoints are obtained; Based on the estimated scene images corresponding to each viewpoint and the reference scene images corresponding to each viewpoint, the target loss value is obtained. Based on the target loss value, the 3D editing module and the task adaptation module in the scene editing model to be trained are trained to obtain the target scene editing model.
2. The method according to claim 1, wherein, The 3D editing module in the scene editing model to be trained includes a low-resolution editing sub-network for global editing of the 3D latent representation, and a high-resolution editing sub-network for recovering local appearance information in the 3D latent representation; The step of using the 3D editing module in the scene editing model to be trained, and based on the initial editing semantic vector, to edit the 3D latent representation to be edited to obtain the initial 3D latent representation includes: Using the low-resolution editing sub-network in the 3D editing module, and combining the initial editing semantic vector and the 3D latent representation to be edited, voxel variation features are obtained. Based on the voxel variation features and the 3D latent representation to be edited, a fused 3D latent representation is obtained. The voxel variation features represent the spatial distribution and editing variation of the voxels to be edited in the 3D latent representation to be edited. The high-resolution editing sub-network in the 3D editing module is used to restore the local appearance information in the fused 3D latent representation to obtain the initial 3D latent representation.
3. The method according to claim 2, wherein, The 3D editing module in the scene editing model to be trained further includes a first cross-attention layer; the method further includes: The first cross-attention layer in the 3D editing module is used, and the initial editing semantic vector and the 3D latent representation to be edited are combined to obtain the first cross-attention result; Specifically, the process of utilizing the low-resolution editing sub-network in the 3D editing module, combining the initial editing semantic vector and the 3D latent representation to be edited, to obtain voxel variation features, and then fusing these voxel variation features with the 3D latent representation to be edited to obtain a fused 3D latent representation, includes: The first cross-attention result is fused with the three-dimensional latent representation to be edited to obtain the three-dimensional latent representation to be edited carrying the first cross-attention result; Using the low-resolution editing sub-network in the 3D editing module, and combining the 3D latent representation to be edited carrying the first cross-attention result with the initial editing semantic vector, voxel change features are obtained. Based on the voxel change features and the 3D latent representation to be edited, a fused 3D latent representation is obtained.
4. The method according to claim 2, wherein, The 3D editing module in the scene editing model to be trained further includes a second cross-attention layer; the method further includes: The second cross-attention layer in the 3D editing module is used, and the initial editing semantic vector and the fused 3D latent representation are combined to obtain the second cross-attention result; The step of using the high-resolution editing sub-network in the 3D editing module to restore the local appearance information in the fused 3D latent representation to obtain the initial 3D latent representation includes: The second cross-attention result is fused with the fused 3D latent representation to obtain a fused 3D latent representation carrying the second cross-attention result; The initial three-dimensional latent representation is obtained by utilizing the high-resolution editing sub-network in the three-dimensional editing module and combining it with the fused three-dimensional latent representation carrying the second cross-attention result and the initial editing semantic vector.
5. The method according to any one of claims 1-4, wherein, The task adaptation module in the scene editing model to be trained includes a local appearance editing task head and a global illumination editing task head; The step of utilizing the task adaptation module in the scene editing model to be trained, and based on the initial editing semantic vector, to correct the initial 3D latent representation to obtain the target 3D latent representation includes: Based on the editing type corresponding to the initial editing semantic vector, at least one initial task head is invoked; wherein, the initial task head is one of the local appearance editing task head and the global lighting editing task head; Using at least one invoked initial task header and based on the initial edit semantic vector, the initial 3D latent representation is modified to obtain the target 3D latent representation.
6. The method according to claim 5, wherein, The step of using at least one invoked initial task header and, based on the initial edit semantic vector, refining the initial 3D latent representation to obtain the target 3D latent representation includes: When the initial task head called is a local appearance editing task head, the editing residual is obtained by using the local appearance editing task head and based on the initial scene image corresponding to each view and the reference scene image corresponding to the initial scene image. Based on the edit residual and the initial edit semantic vector, spatial gating features are generated, wherein the spatial gating features represent the edit region in the three-dimensional latent representation to be edited, the edit intensity of voxels in the edit region, and the non-edit region; Based on the spatial gating features, the non-editable regions in the initial three-dimensional latent representation are restored to obtain the target three-dimensional latent representation.
7. The method according to claim 6, wherein, The process of restoring the non-editable regions in the initial 3D latent representation based on the spatial gating features to obtain the target 3D latent representation includes: Based on the aforementioned spatial gating features, a three-dimensional mask matrix is obtained; Based on the three-dimensional mask matrix, the initial three-dimensional latent representation and the three-dimensional latent representation to be edited are weighted to restore the non-editable region and obtain the target three-dimensional latent representation.
8. The method according to claim 5, wherein, The step of using at least one invoked initial task header and, based on the initial edit semantic vector, refining the initial 3D latent representation to obtain the target 3D latent representation includes: When the initial task header invoked is a global illumination editing task header, the global color features are obtained by using the global illumination editing task header and combining it with the initial editing instruction information; Based on the global color features, the initial three-dimensional latent representation is processed to obtain the target three-dimensional latent representation.
9. The method according to any one of claims 1-4, wherein, The process of obtaining estimated scene images corresponding to each of the N viewpoints based on the target's three-dimensional latent representation includes: Using a preset Gaussian decoding model, the target three-dimensional latent representation is subjected to Gaussian decoding to obtain an initial three-dimensional Gaussian scene after editing based on the initial editing instruction information; Based on the edited initial 3D Gaussian scene, rendering is performed to obtain estimated scene images corresponding to each viewpoint.
10. The method according to claim 9, wherein, The step of training the 3D editing module and the task adaptation module in the scene editing model to be trained based on the target loss value includes: Based on the target loss value, the preset Gaussian decoding model, as well as the 3D editing module and task adaptation module in the scene editing model to be trained, are jointly trained to obtain the target scene editing model and the target Gaussian decoding model.
11. The method according to claim 1, wherein, The process of performing multi-view encoding on the initial scene images corresponding to the N viewpoints to obtain the 3D latent representation to be edited includes: Feature extraction is performed on the initial scene images corresponding to each viewpoint to obtain the feature maps corresponding to each initial scene image; The preset three-dimensional voxel grid is projected onto the feature map corresponding to each initial scene image, and the feature blocks corresponding to the voxels in the preset three-dimensional voxel grid on each feature map are sampled. The feature blocks corresponding to the voxels in the preset three-dimensional voxel grid on each feature map are fused to obtain the fused features corresponding to each voxel in the preset three-dimensional voxel grid. The fusion features corresponding to each voxel in the preset three-dimensional voxel mesh are processed to obtain the three-dimensional latent representation to be edited.
12. A spatial editing method applied to a 3D Gaussian scene, comprising: Obtain the scene images to be processed corresponding to each of the M viewpoints, as well as the target editing instruction information; wherein, the scene images to be processed corresponding to the M viewpoints are used to describe the target 3D Gaussian scene under different viewpoints; the target editing instruction information is used to instruct the modification of the appearance parameters of the target 3D Gaussian scene corresponding to the scene images to be processed; M is an integer greater than or equal to 2; The scene images to be processed corresponding to the M viewpoints are subjected to multi-view encoding processing to obtain a first three-dimensional latent representation; wherein, the first three-dimensional latent representation is a unified representation of the target three-dimensional Gaussian scene mapped to the latent space and integrating the M viewpoints; The first three-dimensional latent representation and the target editing semantic vector representing the target editing instruction information are input into the target scene editing model to obtain the second three-dimensional latent representation; wherein, the target scene editing model is obtained by training using the model training method described in any one of claims 1 to 11; Based on the second three-dimensional latent representation, the target three-dimensional Gaussian scene after editing based on the target editing instruction information is obtained.
13. A model training device for editing models applied to 3D Gaussian scenes, comprising: A sample acquisition unit is used to acquire target training samples, wherein the target training samples include initial editing instruction information, initial scene images corresponding to each of the N viewpoints, and reference scene images corresponding to each initial scene image; the initial editing instruction information is used to instruct the modification of the appearance parameters of the initial 3D Gaussian scene corresponding to the initial scene image; the initial scene images corresponding to the N viewpoints are used to describe the initial 3D Gaussian scene under different viewpoints, and the reference scene images corresponding to the initial scene images are used to represent the labeled scene image after the appearance parameters of the initial 3D Gaussian scene are modified based on the initial editing instruction information; N is an integer greater than or equal to 2; The model training unit is used to perform multi-view encoding processing on the initial scene images corresponding to the N viewpoints to obtain the 3D latent representation to be edited; wherein, the 3D latent representation to be edited is a unified representation that maps the initial 3D Gaussian scene to the latent space and integrates the N viewpoints; it is used to input the 3D latent representation to be edited and the initial editing semantic vector representing the initial editing instruction information into the scene editing model to be trained, so as to use the 3D editing module in the scene editing model to edit the 3D latent representation to be edited based on the initial editing semantic vector to obtain the initial 3D latent representation; and, using the task adaptation module in the scene editing model to be trained, and based on the initial editing semantic vector, to correct the initial 3D latent representation to obtain the target 3D latent representation; based on the target 3D latent representation, the estimated scene images corresponding to each of the N viewpoints are obtained; based on the estimated scene images corresponding to each viewpoint and the reference scene images corresponding to each viewpoint, the target loss value is obtained, so as to train the 3D editing module and the task adaptation module in the scene editing model to be trained based on the target loss value to obtain the target scene editing model.
14. The apparatus according to claim 13, wherein, The 3D editing module in the scene editing model to be trained includes a low-resolution editing sub-network for global editing of the 3D latent representation, and a high-resolution editing sub-network for recovering local appearance information in the 3D latent representation; Specifically, the model training unit is used for: Using the low-resolution editing sub-network in the 3D editing module, and combining the initial editing semantic vector and the 3D latent representation to be edited, voxel variation features are obtained. Based on the voxel variation features and the 3D latent representation to be edited, a fused 3D latent representation is obtained. The voxel variation features represent the spatial distribution and editing variation of the voxels to be edited in the 3D latent representation to be edited. The high-resolution editing sub-network in the 3D editing module is used to restore the local appearance information in the fused 3D latent representation to obtain the initial 3D latent representation.
15. The apparatus according to claim 13 or 14, wherein, The task adaptation module in the scene editing model to be trained includes a local appearance editing task head and a global illumination editing task head; Specifically, the model training unit is used for: Based on the editing type corresponding to the initial editing semantic vector, at least one initial task head is invoked; wherein, the initial task head is one of the local appearance editing task head and the global lighting editing task head; Using at least one invoked initial task header and based on the initial edit semantic vector, the initial 3D latent representation is modified to obtain the target 3D latent representation.
16. The apparatus according to claim 15, wherein, The model training unit is specifically used for: When the initial task head called is a local appearance editing task head, the editing residual is obtained by using the local appearance editing task head and based on the initial scene image corresponding to each view and the reference scene image corresponding to the initial scene image. Based on the edit residual and the initial edit semantic vector, spatial gating features are generated, wherein the spatial gating features represent the edit region in the three-dimensional latent representation to be edited, the edit intensity of voxels in the edit region, and the non-edit region; Based on the spatial gating features, the non-editable regions in the initial three-dimensional latent representation are restored to obtain the target three-dimensional latent representation.
17. The apparatus according to claim 15, wherein, The model training unit is specifically used for: When the initial task header invoked is a global illumination editing task header, the global color features are obtained by using the global illumination editing task header and combining it with the initial editing instruction information; Based on the global color features, the initial three-dimensional latent representation is processed to obtain the target three-dimensional latent representation.
18. A spatial editing device for a three-dimensional Gaussian scene, comprising: The data input unit is used to acquire the scene images to be processed corresponding to each of the M viewpoints, as well as the target editing instruction information; wherein, the scene images to be processed corresponding to the M viewpoints are used to describe the target 3D Gaussian scene under different viewpoints; the target editing instruction information is used to instruct the modification of the appearance parameters of the target 3D Gaussian scene corresponding to the scene images to be processed; M is an integer greater than or equal to 2; A spatial editing unit is used to perform multi-view encoding processing on the scene images to be processed corresponding to the M viewpoints to obtain a first three-dimensional latent representation; wherein, the first three-dimensional latent representation is a unified representation of the target three-dimensional Gaussian scene mapped to the latent space and integrating the M viewpoints; the first three-dimensional latent representation and the target editing semantic vector representing the target editing instruction information are input into the target scene editing model to obtain a second three-dimensional latent representation; wherein, the target scene editing model is obtained after training using the model training method described in any one of claims 1 to 11; based on the second three-dimensional latent representation, the target three-dimensional Gaussian scene edited based on the target editing instruction information is obtained.
19. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-12.
20. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-12.
21. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-12.