Stylized three-dimensional virtual scene generation method, electronic device and storage medium

By introducing a large model to generate style reference images in 3D scene reconstruction and fine-tuning the 3DGS model, the problems of style inconsistency and visual artifacts in 3D scene style transfer are solved, and style unification and stable conversion under multiple perspectives are achieved.

CN120339528BActive Publication Date: 2025-09-23SHENZHEN XGRIDS-INNOVATION CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510825393.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-23
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Existing 3D scene style transfer technology suffers from style inconsistency and visual artifacts under multiple perspectives, and is overly dependent on the original image, resulting in unstable style transfer effects.

Method used

A large model is used to generate a style baseline image based on the user-selected target 3DGS rendered image and target style file. Style transfer is performed directly on the 3DGS model through the large model, and the 3DGS model is fine-tuned using the rendered image after style transfer to optimize scene parameters to ensure style consistency.

Benefits of technology

It achieves consistency in style transfer of three-dimensional scenes from multiple perspectives, reduces dependence on the original image, improves the stability and efficiency of the style transfer process, and ensures the unity of the output image in color, texture and artistic style.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339528B_ABST
    Figure CN120339528B_ABST
Patent Text Reader

Abstract

The present application relates to the field of three-dimensional reconstruction technology and discloses a method, electronic device, and storage medium for generating a stylized three-dimensional virtual scene. The method comprises: utilizing a 3DGS model to render multiple 3DGS-rendered images of a three-dimensional scene according to camera parameters from multiple perspectives; inputting a target 3DGS-rendered image and a target style file into a large model, and migrating the target style into the target 3DGS-rendered image via the large model to obtain a style reference image; inputting multiple 3DGS-rendered images and style reference images into the large model, and migrating the style of the style reference image to each 3DGS-rendered image via the large model to obtain multiple stylized rendered images; and using the multiple stylized rendered images as fine-tuning datasets to fine-tune the 3DGS model and update the scene parameters in the 3DGS model. The present application achieves global style unification for stylized three-dimensional scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of three-dimensional reconstruction technology, and specifically to a method for generating a stylized three-dimensional virtual scene, an electronic device, and a storage medium. Background Art

[0002] In recent years, advances in computer vision and graphics have spurred advancements in 3D scene reconstruction and style transfer technologies. Traditional 2D image style transfer methods leverage convolutional neural networks to extract features and incorporate normalization to achieve style transfer. While these methods have achieved significant success in the image field, direct application to 3D scenes can lead to inconsistent styles and visual artifacts due to differences in multi-view data. Regarding 3D reconstruction, new technologies such as NeRF and 3DGS can accurately capture geometry and texture, restoring high-quality 3D scenes, but they fail to fully address the issue of uniformly transferring artistic styles.

[0003] Existing approaches for 3D scene style transfer first capture and preprocess the original image, then transfer the style frame by frame, and then reconstruct the 3D scene based on the stylized image. However, this approach is highly dependent on the original image and processes each frame independently, resulting in unstable styles. This leads to problems such as style drift, visual artifacts, and inconsistencies across multiple views. Manual editing and correction are also costly and inefficient. Therefore, achieving globally unified and visually consistent style transfer while ensuring geometrically faithful restoration is a key challenge that needs to be addressed. Summary of the Invention

[0004] In view of the above problems, embodiments of the present application provide a stylized three-dimensional virtual scene generation method, electronic device, and storage medium, which are used to solve the problem of inconsistent styles when stylizing three-dimensional scenes in the prior art.

[0005] According to one aspect of an embodiment of the present application, a method for generating a stylized three-dimensional virtual scene is provided, the method comprising:

[0006] Using the 3DGS model, multiple 3DGS rendered images of the three-dimensional scene are obtained according to the camera parameters of multiple perspectives;

[0007] Determining a target 3DGS rendered image and obtaining a target style file, wherein the target 3DGS rendered image is one of the multiple 3DGS rendered images, and the target style file is used to express the target style;

[0008] Inputting the target 3DGS rendered image and the target style file into a large model, and migrating the target style to the target 3DGS rendered image through the large model to obtain a style reference image;

[0009] Inputting the plurality of 3DGS rendered images and the style reference image into the large model, and transferring the style of the style reference image to each 3DGS rendered image through the large model to obtain a plurality of stylized rendered images;

[0010] Training step: using the plurality of stylized rendered images as a fine-tuning dataset, performing fine-tuning training on the 3DGS model to update scene parameters in the 3DGS model, thereby obtaining an optimized 3DGS model.

[0011] Optionally, after the training step, the method further includes:

[0012] Utilizing the optimized 3DGS model, rendering according to camera parameters of multiple perspectives to obtain multiple stylized 3DGS rendered images of the three-dimensional scene;

[0013] Inputting the plurality of stylized 3DGS rendered images into the large model, and performing anomaly detection on a stylized three-dimensional virtual scene using the large model, wherein the stylized three-dimensional virtual scene is composed of the plurality of stylized 3DGS rendered images;

[0014] If the large model detects that an abnormality exists in the stylized three-dimensional virtual scene, generating a new 3DGS rendered image that can cover the abnormal area in the stylized three-dimensional virtual scene through the large model;

[0015] Migrating the target style to the new 3DGS rendered image based on the new 3DGS rendered image and the target style file using the large model to obtain a stylized new 3DGS rendered image;

[0016] Adding the stylized new 3DGS rendered image to the fine-tuning dataset;

[0017] The training step and subsequent steps are repeatedly performed until the large model detects that there is no abnormality in the stylized three-dimensional virtual scene.

[0018] Optionally, determining a target 3DGS rendered image and obtaining a target style file includes:

[0019] In response to a user selecting an image from the plurality of 3DGS rendered images, determining the target 3DGS rendered image;

[0020] Get the target style file entered by the user.

[0021] Optionally, the target style file is one or more of image, text, and voice.

[0022] Optionally, the camera parameters include intrinsic parameters and extrinsic parameters of the virtual camera.

[0023] Optionally, the scene parameters include the color and texture of the Gaussian sphere of the 3DGS model.

[0024] Optionally, performing anomaly detection on the stylized three-dimensional virtual scene using the large model includes:

[0025] The large model is used to detect whether there are floating objects, defects or local inconsistencies in the stylized three-dimensional virtual scene.

[0026] Optionally, if the large model detects that an abnormality exists in the stylized three-dimensional virtual scene, generating a new 3DGS rendered image that can cover the abnormal area in the stylized three-dimensional virtual scene through the large model includes:

[0027] If the large model detects that the stylized three-dimensional virtual scene has an abnormality, determining an abnormal region in the stylized three-dimensional virtual scene where the abnormality exists;

[0028] Determining a virtual camera viewing angle capable of covering the abnormal area;

[0029] Generate a new 3DGS rendering image under the perspective of the virtual camera.

[0030] According to another aspect of an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the stylized three-dimensional virtual scene generation method as described above.

[0031] According to another aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the stylized three-dimensional virtual scene generation method as described above is implemented.

[0032] In this embodiment, a large model is used to generate a style baseline image based on a user-selected target 3DGS rendered image and a target style file, serving as a unified style baseline. This style baseline image is then used to perform style transfer directly on the high-quality 3DGS rendered image generated by an existing 3DGS model. The multi-view stylized rendered images after style transfer are then used as a fine-tuning dataset to fine-tune the 3DGS model and optimize it. Finally, the stylized rendered image is directly generated using the 3DGS model. This approach ensures consistency in color, texture, and artistic style across all viewpoints of the style-transferred output image. Furthermore, since style transfer does not require the original image to be styled, strict reliance on the original captured data is avoided, resulting in a more stable style transfer process.

[0033] The above description is only an overview of the technical solutions of the embodiments of the present application. In order to more clearly understand the technical means of the embodiments of the present application, they can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The accompanying drawings are only used to illustrate the embodiments and are not to be considered as limiting the present application. In addition, the same reference symbols are used to represent the same components throughout the drawings. In the drawings:

[0035] Figure 1 A schematic diagram of an application scenario of an embodiment of the present application is shown;

[0036] Figure 2 A flowchart of a method for generating a stylized three-dimensional virtual scene provided by an embodiment of the present application is shown;

[0037] Figure 3 A flowchart of another method for generating a stylized three-dimensional virtual scene provided by an embodiment of the present application is shown;

[0038] Figure 4 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0039] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited to the embodiments set forth herein.

[0040] In recent years, the rapid development of computer vision and computer graphics has driven continuous progress in 3D scene reconstruction and style transfer technologies. Traditional 2D image style transfer methods utilize convolutional neural networks (such as VGG19) to extract content and style features, and then achieve artistic style transfer through feature fusion and normalization techniques (such as AdaIN). These methods have achieved remarkable results in the image domain, but when directly applied to 3D scenes, they are prone to style inconsistencies and visual artifacts due to differences in multi-view data.

[0041] In the area of ​​3D reconstruction, new implicit representation methods such as Neural Radiance Fields (NeRF) and point-based rendering techniques such as 3D Gaussian Splatting (3DGS) can accurately capture the scene's geometry and texture information, recovering high-quality 3D scenes from multi-view images and achieving high-fidelity conversion from 2D images to 3D models. However, traditional 3D scene reconstruction methods primarily focus on restoring the geometric accuracy of the real physical world and have not fully considered the issue of uniformly transferring artistic styles.

[0042] In one 3D scene style transfer scheme, original images are first captured and preprocessed. RGB images are acquired from multiple perspectives, and camera pose information is obtained through image filtering, resolution adjustment, and feature extraction. Style transfer is then performed frame by frame. A pretrained VGG19 convolutional neural network is used to extract content and style features from each captured original image frame (the style is represented by the statistical relationship between activation maps from multiple convolutional layers in VGG19). Multi-scale features are then fused using a feature pyramid network. A style transfer network based on methods such as AdaIN is then employed to transform the original image into an image with the target artistic style, resulting in a stylized image for each original frame. 3D scene reconstruction is then performed based on these multiple stylized images. Using these stylized images as a fine-tuning dataset, the original 3D scene is optimized and reconstructed using NeRF or 3DGS techniques, ultimately outputting a stylized 3D scene.

[0043] This method directly performs style conversion on the collected original image, and then integrates the style conversion results into the three-dimensional scene through the reconstruction module. This reconstruction process is highly dependent on the original image, relying on the consistency of style details in the images from each perspective to ensure the consistency of the scene in different areas. However, sometimes the original image may have been lost and unavailable, and the original image is often affected by factors such as the shooting environment, lighting conditions, noise, and shooting angle, which may cause unnecessary interference and lead to unstable style conversion effects. Therefore, the excessive reliance on the original image makes it difficult for the entire process to ensure high consistency of the final output stylized scene when faced with uneven image quality.

[0044] Because each frame of image is processed using independent style transfer, the style features extracted by the VGG19 convolutional neural network are unstable and lack a unified style benchmark. As a result, the stylized images generated in different regions / different perspectives have differences in color, texture, and artistic expression, resulting in style drift or visual artifacts, making it difficult to ensure multi-perspective style consistency. This multi-perspective inconsistency problem directly affects the overall beauty and coherence of the three-dimensional scene.

[0045] How to achieve artistic style conversion with unified global style and consistent visual effects while ensuring the true restoration of the three-dimensional scene geometric structure has become a key issue that needs to be urgently solved in the current technical field.

[0046] This application provides a method for generating a stylized three-dimensional virtual scene. The method uses a large model to generate a style reference image based on a user-selected target 3DGS rendered image and a target style file as a unified style reference. The style reference image is used to directly perform style conversion on the high-quality 3DGS rendered image generated by an existing 3DGS model. The multi-perspective stylized rendered images after style conversion are used as a fine-tuning dataset to fine-tune the 3DGS model to optimize the 3DGS model. Finally, the stylized rendered image is directly obtained using the 3DGS model. Through the above method, the consistency of the color, texture, and artistic style of the output image after style transfer from all perspectives is ensured. Since there is no need to perform style conversion on the original image, strict dependence on the original collected data is avoided, and the style transfer process is more stable.

[0047] Figure 1 A schematic diagram of an application scenario of an embodiment of the present application is shown. As shown in the figure, in a three-dimensional scene, the three-dimensional scene is scanned by a three-dimensional scanning and reconstruction device 1 to obtain point cloud data and image data. For example, the three-dimensional scanning and reconstruction device 1 can be provided with a laser radar, a camera, and a computing unit, wherein the laser radar collects point cloud data of the three-dimensional scene, the camera obtains images of the three-dimensional scene from different perspectives, and the computing unit fuses the point cloud data and the image to achieve three-dimensional reconstruction. The three-dimensional scanning and reconstruction device 1 can be a handheld device or a non-handheld device. If the three-dimensional scanning and reconstruction device 1 is a handheld device, the user holds the three-dimensional scanning and reconstruction device 1 and moves it in the three-dimensional scene to collect point cloud data and image data of the three-dimensional scene. If the three-dimensional scanning and reconstruction device 1 is a non-handheld device, the three-dimensional scanning and reconstruction device 1 can be moved autonomously in the three-dimensional scene to collect point cloud data and image data of the three-dimensional scene.

[0048] Terminal device 2 is communicatively connected to 3D scanning and reconstruction device 1, receiving and displaying the 3D reconstruction results obtained by 3D scanning and reconstruction device 1 after 3D reconstruction of the 3D scene. Alternatively, 3D scanning and reconstruction device 1 may simply collect raw point cloud data and image data and transmit them to terminal device 2, which then fuses and processes the data to perform 3D reconstruction, obtaining and displaying the 3D reconstruction results. Terminal device 2 may be a mobile phone, tablet computer, laptop computer, desktop computer, server, or other device; a mobile phone is used as an example in the figure. Terminal device 2 and 3D scanning and reconstruction device 1 may be connected via wired or wireless communication.

[0049] In other scenarios, the lidar and camera can also be separate devices and not integrated into the three-dimensional scanning and reconstruction device 1. They send the data collected by each of them to the terminal device 2, which performs data fusion processing to obtain and display the three-dimensional reconstruction results.

[0050] The above-mentioned three-dimensional reconstruction result can be a 3DGS model. The 3DGS model can embed the target style features, and the stylized three-dimensional virtual scene can be displayed by opening the 3DGS model. The virtual scene is a three-dimensional environment generated by the above-mentioned electronic device, which usually contains information such as geometric shapes, textures and lighting, and can be interactively browsed by observers from various perspectives. The virtual scene can be derived from the reconstruction of a real scene or an artificial synthetic scene. The virtual scene output by this application is a reconstruction of a real scene, and has been converted in visual style (such as converted into a cartoon style, oil painting style, etc.), so it is called a stylized three-dimensional virtual scene.

[0051] Figure 2 The flowchart of the method for generating a stylized three-dimensional virtual scene provided by an embodiment of the present application is shown, and the method is executed by an electronic device. The electronic device may be Figure 1 The three-dimensional scanning and reconstruction device 1 or terminal device 2 in the embodiment. Figure 2 As shown, the method includes the following steps:

[0052] S101 , using a 3DGS model, rendering according to camera parameters of multiple viewing angles to obtain multiple 3DGS rendered images of a three-dimensional scene.

[0053] 3D scene reconstruction is the process of recovering the 3D geometry and texture appearance of a scene from multi-view 2D image data. Common methods include point cloud reconstruction based on Structure from Motion (SfM) and Multi-View Stereo (MVS), NeRF implicit modeling, and point primitive representation methods such as 3DGS. The reconstructed results can take the form of sparse or dense point clouds, voxel meshes, triangular mesh models, or volume rendering models implicitly represented by neural networks.

[0054] 3DGS is a three-dimensional scene reconstruction and rendering technology. Its core idea is to use three-dimensional Gaussian functions to represent points in the scene, and project these Gaussian functions onto a two-dimensional image plane for rendering. 3DGS combines the advantages of explicit point cloud representation and continuous field representation, and can achieve efficient real-time rendering while ensuring rendering quality. This application uses the 3DGS model as a three-dimensional reconstruction module to generate stylized three-dimensional virtual scenes. The 3DGS model can support rendering and observation from any new perspective. The 3DGS model refers to a software module that reconstructs a three-dimensional scene using 3D Gaussian sphere technology, that is, a module that converts the original image or point cloud taken by the camera into a renderable 3D scene.

[0055] This step obtains multiple 3DGS rendered images based on the existing 3DGS model, without relying on the original images captured by the original camera. The 3DGS model has the ability to accurately reconstruct geometry and textures, ensuring that images from each perspective truly reflect the scene structure.

[0056] The camera parameters used for rendering include the virtual camera's intrinsic and extrinsic parameters. Intrinsic parameters describe the camera's internal imaging characteristics, such as focal length, principal point coordinates, and lens distortion. These parameters are primarily related to the camera's hardware structure and are inherent properties of the camera itself. Extrinsic parameters describe the camera's position and posture (i.e., pose information) in the world coordinate system. These parameters include rotation matrices and translation vectors, reflecting the camera's external geometric relationship with respect to the world coordinate system. Using the camera parameters from multiple viewpoints, multiple 3DGS rendered images can be generated from the corresponding viewpoints of the 3D scene.

[0057] The number of 3DGS rendered images obtained by rendering an indoor scene is approximately 1,000 to 2,000.

[0058] S102: Determine a target 3DGS rendering image and obtain a target style file.

[0059] The target 3DGS rendered image is one of the multiple 3DGS rendered images, and the target style file is used to express the target style (the style that the user wants to generate in the end). After the image style transfer is performed in this application, the resulting 3D virtual scene is a scene with the target style.

[0060] Image Style Transfer is a technology that combines the content of one image with the artistic style of another image to generate an image with a new style. Here, "content" usually refers to the high-level semantics or macro structure of the image, and "style" refers to the visual patterns such as color, texture, and brushstrokes in the image. Classic neural style transfer uses convolutional neural networks to extract features of content images and style images, and generates a synthetic image that retains the content and presents the target style through optimization. In the subsequent steps, this application uses a large model to convert real images into images of any artistic style, providing consistent style input from multiple perspectives for subsequent three-dimensional virtual scene style transfer.

[0061] Step S102 includes the following steps:

[0062] S1021 : In response to a user selecting an image from a plurality of 3DGS rendered images, determining a target 3DGS rendered image.

[0063] The user can select an image from the multiple 3DGS rendered images generated in step S101 that can show most of the three-dimensional scene in the largest viewing angle as the target 3DGS rendered image, which is used as a scene description reference to express the types of objects in the scene.

[0064] S1022: Obtain the target style file input by the user.

[0065] Users can select any style as their target style, such as cartoon or oil painting. The target style file can be one or more of an image, text, or audio. For example, users can directly input an image of the target style, a text description of the target style, an audio description of the target style, or a combination of these.

[0066] S103 , inputting the target 3DGS rendered image and the target style file into the large model, and migrating the target style into the target 3DGS rendered image through the large model to obtain a style reference image.

[0067] The large model used in this application is an advanced artificial intelligence model with multimodal generation and understanding capabilities. It can simultaneously process multiple input forms such as text, images, and audio, and has multiple output forms such as text, images, and audio. This large model has the following functions:

[0068] Joint image and text understanding: supports multimodal input, inputs arbitrary images and / or text descriptions, and extracts high-level semantic representations;

[0069] Image generation and style transfer: capable of image stylization, image completion, and image diffusion reconstruction;

[0070] Consistent modeling capability: Ability to generate images with consistent style across viewpoints.

[0071] Large models that can be used include but are not limited to GPT-4 Omni (GPT4o) and the like.

[0072] The user selects a target 3DGS rendered image from among multiple generated 3DGS rendered images as a content reference, and also provides a target style file. Both are then fed into the large model, which leverages its fusion and generation capabilities to generate a unified style baseline image based on these two inputs. This style baseline image serves as a global style reference, guiding subsequent style transfer across all viewpoints and ensuring consistency in color, texture, and artistic style across the output images.

[0073] S104: Input multiple 3DGS rendered images and style reference images into the large model, and transfer the style of the style reference image to each 3DGS rendered image through the large model to obtain multiple stylized rendered images.

[0074] This step uses the multiple 3DGS rendered images obtained in S101 and the style reference image generated in S103 as dual inputs, which are then fed into the large model. The large model then fuses these two inputs, achieving direct style transfer for each rendered image from each perspective, generating a stylized rendered image. This process directly transfers style from the 3DGS rendered images based on a unified style reference image, migrating all 3DGS rendered images to the style of the baseline style reference image. This ensures a consistent style across multiple perspectives, while preserving the geometric details of the original scene and reducing reliance on the original captured data. This results in a more stable and reliable data source for the resulting stylized 3D virtual scene.

[0075] The existing technology uses the VGG19+AdaIN method, and its style features are statistical (such as Gram matrix), which cannot achieve global control and consistent abstraction of style transfer. This application uses large models such as GPT-4o. Different from the statistical features of style in the existing technology, the large model can understand artistic style and abstract expression, and can ensure the uniformity of the style of the converted images under multiple perspectives while maintaining the geometric details of the original scene. Based on steps S103 and S104, a large model is introduced to generate a unified style baseline image, and the baseline is used as a reference, and then the large model is used to realize direct style transfer of images rendered from all perspectives, ensuring that the output images are highly consistent in style.

[0076] S105 , using the multiple stylized rendered images as fine-tuning datasets, fine-tuning the 3DGS model to update the scene parameters in the 3DGS model, and obtaining an optimized 3DGS model.

[0077] This step is the training step. The 3DGS model is fine-tuned using the multiple stylized rendered images generated in S104 as a fine-tuning dataset (i.e., supervised data with a consistent image style) to optimize the scene parameters within the 3DGS model. This optimization process is equivalent to feeding back the style transfer effect from S104 to the 3DGS model, allowing the 3DGS model to embed the characteristics of the target style when generating subsequent rendered images, thereby improving the overall style consistency and visual quality of the 3D scene.

[0078] Scene parameters include the color and texture of the Gaussian spheres in the 3DGS model. A 3DGS model contains thousands of Gaussian spheres, each with attributes such as position, orientation, color, size, and texture (i.e., transparency). Of these attributes, color and texture directly influence the style and detail of the rendered image. 3DGS renderings based on the target style optimize the color and texture of the 3DGS model, ensuring that the target style is automatically reflected in subsequent rendered images generated from the 3DGS model.

[0079] During the training process, the 3DGS model is first used to generate rendered images from multiple perspectives. The "multiple perspectives" here correspond one-to-one to the multiple perspectives used to generate the multiple 3DGS rendered images in step S101. A style loss is then calculated for each generated rendered image compared to the stylized rendered images from the same perspective in the fine-tuning dataset. Backpropagation is used to continuously optimize the color and transparency of the Gaussian spheres, gradually approaching the target style. That is, during the training process, the 3DGS model refers to multiple 3DGS rendered images that have been stylized in the target style, fine-tunes the color, texture, and other parameters of each Gaussian sphere in the 3DGS model, and obtains an optimized 3DGS model. Users can open the 3DGS model file through software and browse the stylized three-dimensional virtual scene.

[0080] The 3DGS model obtained through steps S101-S105 embeds the target style features. When the 3DGS model file is opened using software, a stylized 3D virtual scene is displayed. However, such a 3D virtual scene may contain floating objects, defects, or local inconsistencies. Existing technologies for addressing these issues require manual editing, which increases costs and reduces overall system efficiency. In some embodiments, a large model is used to perform anomaly detection on the stylized 3D virtual scene rendered from the optimized 3DGS model for floating objects, defects, and local inconsistencies. Based on the detection results, the 3DGS model is iteratively optimized until the large model returns normal detection results.

[0081] Figure 3 A flowchart of another method for generating a stylized three-dimensional virtual scene provided by an embodiment of the present application is shown, and the method is executed by an electronic device. The electronic device may be Figure 1The three-dimensional scanning and reconstruction device 1 or terminal device 2 in the embodiment. Figure 3 As shown, the method includes the following steps:

[0082] S201 , using a 3DGS model, rendering according to camera parameters of multiple viewing angles to obtain multiple 3DGS rendered images of a three-dimensional scene.

[0083] S202, determining a target 3DGS rendering image and obtaining a target style file.

[0084] S203 , inputting the target 3DGS rendered image and the target style file into the large model, and migrating the target style into the target 3DGS rendered image through the large model to obtain a style reference image.

[0085] S204 , multiple 3DGS rendered images and style reference images are input into the large model, and the style of the style reference image is transferred to each 3DGS rendered image through the large model to obtain multiple stylized rendered images.

[0086] S205 , using the multiple stylized rendered images as fine-tuning datasets, fine-tuning the 3DGS model to update the scene parameters in the 3DGS model, and obtaining an optimized 3DGS model.

[0087] Steps S201 to S205 are similar to the aforementioned steps S101 to S105 and will not be repeated here.

[0088] S206 , using the optimized 3DGS model, rendering according to camera parameters of multiple perspectives to obtain multiple stylized 3DGS rendered images.

[0089] When the optimized 3DGS model generates multi-view images, the generated multi-view images will have the target style, and the visual style of each image will be consistent. The 3D scene in this step is the same as the 3D scene in S101, but the perspective can be different from that in S101.

[0090] In step S207, multiple stylized 3DGS rendered images are input into the large model, and anomaly detection is performed on the stylized three-dimensional virtual scene through the large model to determine whether there is an anomaly in the stylized three-dimensional virtual scene. If so, step S208 is executed; otherwise, the process ends.

[0091] The stylized three-dimensional virtual scene is composed of multiple stylized 3DGS rendered images. Anomaly detection is to detect whether there are floating objects, defects or local inconsistencies in the stylized three-dimensional virtual scene. The large model in the embodiment of the present application can use image understanding capabilities to find floating objects, defects, local inconsistencies, etc. in the image. If the large model detects that there are any of the floating objects, defects or local inconsistencies in the stylized three-dimensional virtual scene, it is determined that there is an anomaly in the stylized three-dimensional virtual scene; if the large model detects that there are no floating objects, defects and local inconsistencies in the stylized three-dimensional virtual scene, it is determined that there is no anomaly in the stylized three-dimensional virtual scene.

[0092] S208 , generating a new 3DGS rendering image that can cover the abnormal area in the stylized three-dimensional virtual scene through the large model.

[0093] When the large model detects an anomaly in the stylized 3D virtual scene from a certain perspective, it automatically generates a new stylized 3DGS rendered image from that perspective as supplementary data, which is then added to the fine-tuning dataset to optimize and train the 3DGS model again. Step S208 includes the following steps:

[0094] S2081: If the large model detects that the stylized three-dimensional virtual scene has an abnormality, determine an abnormal area in the stylized three-dimensional virtual scene where the abnormality exists.

[0095] S2082: Determine a virtual camera viewing angle that can cover the abnormal area.

[0096] S2083, generating a new 3DGS rendering image from the perspective of the virtual camera.

[0097] The new 3DGS rendering images regenerated from the large model do not have the aforementioned floating objects, defects and local inconsistencies.

[0098] S209 , migrating the target style to the new 3DGS rendered image based on the new 3DGS rendered image and the target style file through the large model to obtain a stylized new 3DGS rendered image.

[0099] S210: Add the stylized new 3DGS rendered image to the fine-tuning dataset, and go to step S205.

[0100] The previously optimized 3DGS model is retrained using a new, stylized 3DGS rendered image free of floating objects, defects, and local inconsistencies. This reduces the likelihood of these issues appearing when the newly optimized 3DGS model renders the scene. This process is repeated several times until no anomalies are detected in the large model, completing the optimization of the 3DGS model and resulting in a usable final 3DGS model.

[0101] After the process is completed, the iteratively optimized 3DGS model is obtained. Users can open the 3DGS model file through the software and browse the stylized 3D virtual scene.

[0102] In the above embodiment, the image generation and intelligent detection capabilities of the large model are utilized to automatically identify floating objects, defects, and local inconsistencies in the stylized three-dimensional virtual scene rendered by the 3DGS model, generate supplementary data from a specific perspective, correct the current style transfer results, and feed the corrected data back to the 3DGS model to further automatically optimize the parameters of the 3DGS model until the anomaly detection results of the large model meet the preset standards (i.e., no anomalies are detected). Finally, the feedback optimization process is terminated, thereby improving the overall visual effect of the 3DGS model rendering.

[0103] Apply the above Figure 2 and Figure 3 The electronic device of the embodiment of the stylized 3D virtual scene generation method comprises the following software modules:

[0104] 3DGS rendering module: also known as the 3DGS model, renders the pre-processed multi-view data according to the camera parameters of multiple viewpoints to obtain multiple 3DGS rendered images of the three-dimensional scene, ensuring the accurate restoration of the geometry and texture information of the original three-dimensional scene.

[0105] Style transfer module: Using a large model, it outputs a unified style baseline image based on the user's selected 3DGS rendered image and the target style file, providing a reference for global style transfer, and generates a stylized rendered image based on the aforementioned multiple 3DGS rendered images and style baseline images.

[0106] 3DGS model optimization module: Use the generated stylized rendered image as a fine-tuning dataset to update the color and texture parameters in the 3DGS model to obtain the optimized 3DGS model.

[0107] Apply the above Figure 3 The electronic device of the embodiment of the stylized 3D virtual scene generation method further comprises the following software modules:

[0108] Intelligent Detection Module: Based on the detection capabilities of large models, it automatically identifies floating objects, defects, and local inconsistencies generated during the 3DGS model style conversion process, and generates supplementary data to correct the output, ensuring that the entire system output is free of floating objects or visual artifacts.

[0109] Figure 4 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. The specific embodiment of the present application does not limit the specific implementation of the electronic device.

[0110] like Figure 4As shown, the electronic device 300 may include a processor 302 and a memory 304 .

[0111] The memory 304 is used to store a computer program 306. The memory 304 may include a high-speed RAM memory, or may also include a non-volatile memory, such as at least one disk memory. The computer program 306 may include computer-executable instructions.

[0112] The processor 302 is configured to execute the computer program 306 to implement the above-mentioned embodiment of the method for generating a stylized three-dimensional virtual scene.

[0113] The processor 302 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the electronic device 300 may be processors of the same type, such as one or more CPUs, or may be processors of different types, such as one or more CPUs and one or more ASICs.

[0114] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the embodiment of the method for generating a stylized three-dimensional virtual scene is implemented.

[0115] An embodiment of the present application provides a computer program that can be executed by a processor to implement the above-mentioned embodiment of the method for generating a stylized three-dimensional virtual scene.

[0116] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the computer program implements the above-mentioned embodiment of the method for generating a stylized three-dimensional virtual scene.

[0117] In the several embodiments provided in this application, if any function is implemented in the form of a software function module / unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the technical solution of this application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server or other electronic device) to execute all or part of the steps of the method described in each embodiment of this application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store computer program code.

[0118] The algorithm or demonstration provided here are not inherently relevant to any particular computer, virtual system or other equipment. Various general purpose systems can also be used together with the teachings based on this. According to the above description, it is obvious that the structure required for constructing this type of system. In addition, the present application embodiment is not directed to any specific programming language yet. It should be understood that various programming languages ​​can be utilized to realize the content of the present application described here, and the above description of specific languages ​​is for the purpose of disclosing the best mode of implementation of the present application.

[0119] It should be noted that the above embodiments illustrate rather than limit the present application, and that a person skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between brackets should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In claims that list several means, several units or modules of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names. The steps in the above embodiments should not be understood as limiting the order of execution unless otherwise specified.

[0120] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for generating a stylized three-dimensional virtual scene, characterized in that: The method comprises: Using the 3DGS model, multiple 3DGS rendered images of the three-dimensional scene are obtained according to the camera parameters of multiple perspectives; Determining a target 3DGS rendered image and obtaining a target style file, wherein the target 3DGS rendered image is one of the multiple 3DGS rendered images, and the target style file is used to express the target style; Inputting the target 3DGS rendered image and the target style file into a large model, and migrating the target style to the target 3DGS rendered image through the large model to obtain a style reference image; Inputting the plurality of 3DGS rendered images and the style reference image into the large model, and transferring the style of the style reference image to each 3DGS rendered image through the large model to obtain a plurality of stylized rendered images; Training step: using the plurality of stylized rendered images as a fine-tuning dataset, performing fine-tuning training on the 3DGS model to update scene parameters in the 3DGS model, thereby obtaining an optimized 3DGS model.

2. The method according to claim 1, characterized in that After the training step, the method further comprises: Utilizing the optimized 3DGS model, rendering according to camera parameters of multiple perspectives to obtain multiple stylized 3DGS rendered images of the three-dimensional scene; Inputting the plurality of stylized 3DGS rendered images into the large model, and performing anomaly detection on a stylized three-dimensional virtual scene using the large model, wherein the stylized three-dimensional virtual scene is composed of the plurality of stylized 3DGS rendered images; If the large model detects that an abnormality exists in the stylized three-dimensional virtual scene, generating a new 3DGS rendered image that can cover the abnormal area in the stylized three-dimensional virtual scene through the large model; Migrating the target style to the new 3DGS rendered image based on the new 3DGS rendered image and the target style file using the large model to obtain a stylized new 3DGS rendered image; Adding the stylized new 3DGS rendered image to the fine-tuning dataset; The training step and subsequent steps are repeatedly performed until the large model detects that there is no abnormality in the stylized three-dimensional virtual scene.

3. The method according to claim 1, characterized in that Determining the target 3DGS rendering image and obtaining the target style file includes: In response to a user selecting an image from the plurality of 3DGS rendered images, determining the target 3DGS rendered image; Get the target style file entered by the user.

4. The method according to claim 1, wherein The target style file is one or more of image, text, and voice.

5. The method according to claim 1, wherein The camera parameters include intrinsic parameters and extrinsic parameters of the virtual camera.

6. The method according to claim 1, characterized in that The scene parameters include the color and texture of the Gaussian sphere of the 3DGS model.

7. The method according to claim 2, characterized in that The performing anomaly detection on the stylized three-dimensional virtual scene by using the large model includes: The large model is used to detect whether there are floating objects, defects or local inconsistencies in the stylized three-dimensional virtual scene.

8. The method according to claim 2, characterized in that If the large model detects that an abnormality exists in the stylized three-dimensional virtual scene, generating a new 3DGS rendered image that can cover the abnormal area in the stylized three-dimensional virtual scene through the large model includes: If the large model detects that the stylized three-dimensional virtual scene has an abnormality, determining an abnormal region in the stylized three-dimensional virtual scene where the abnormality exists; Determining a virtual camera viewing angle capable of covering the abnormal area; Generate a new 3DGS rendering image under the perspective of the virtual camera.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the stylized three-dimensional virtual scene generation method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the stylized three-dimensional virtual scene generation method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Three-dimensional scene style migration method and device, equipment and storage medium

    CN116934936A