Stylized three-dimensional virtual scene generation method, electronic equipment and storage medium
Through the big model, the style benchmark image is generated and the 3DGS model is fine-tuned and trained, the problem of inconsistent style of three-dimensional scenes is solved, style consistency and stability are achieved in multiple perspectives, dependence on the original image is reduced, and efficiency and quality of stylization of three-dimensional scenes are improved.
Patent Information
- Application Number
- CN202510825393.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The existing three-dimensional scene style transfer technology has problems such as inconsistency in style, visual artifacts and inconsistency in multiple perspectives, which leads to unstable stylization effects of three-dimensional scenes and is too high to rely on the original image, making it difficult to achieve global style unity and visual consistency.
The large model is used to generate the style benchmark image based on the target 3DGS selected by the user, and the 3DGS model is fine-tuned and trained through the large model to ensure style consistency, and the large model is used for abnormal detection and optimization to generate a stylized three-dimensional virtual scene.
The consistency of the color, texture and artistic style of style conversion from multiple perspectives is achieved, the dependence on the original image is reduced, the stability and efficiency of the style transfer process is improved, and the cost of manual editing is avoided.
Smart Images

Figure CN120339528A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of three-dimensional reconstruction technology, and specifically to a stylized three-dimensional virtual scene generation method, electronic device, and storage medium. Background Art
[0002] In recent years, the development of computer vision and graphics has promoted the progress of 3D scene reconstruction and style transfer technology. Traditional 2D image style transfer methods use convolutional neural networks to extract features and integrate normalization to achieve style transfer. They have achieved remarkable results in the field of images, but when directly used in 3D scenes, they are prone to style inconsistency and visual artifacts due to differences in multi-view data. In terms of 3D reconstruction, new technologies such as NeRF and 3DGS can accurately capture geometry and texture and restore high-quality 3D scenes, but they do not fully consider the problem of unified artistic style transfer.
[0003] An existing 3D scene style transfer scheme first collects and preprocesses the original image, transfers the style frame by frame, and then reconstructs the 3D scene based on the stylized image. However, this method is highly dependent on the original image and each frame is processed independently, resulting in unstable style. There are problems such as style drift, visual artifacts, and multi-view inconsistency. Manual editing and correction are costly and inefficient. Therefore, achieving global style uniformity and visual consistency based on geometric true restoration is a key issue that needs to be solved urgently. Summary of the invention
[0004] In view of the above problems, the embodiments of the present application provide a stylized three-dimensional virtual scene generation method, an electronic device and a storage medium, which are used to solve the problem of inconsistent styles when stylizing three-dimensional scenes in the prior art.
[0005] According to one aspect of an embodiment of the present application, a method for generating a stylized three-dimensional virtual scene is provided, the method comprising: Using the 3DGS model, multiple 3DGS rendering images of the three-dimensional scene are obtained according to the camera parameters of multiple viewing angles; Determine a target 3DGS rendered image and obtain a target style file, wherein the target 3DGS rendered image is one of the multiple 3DGS rendered images, and the target style file is used to express the target style; Inputting the target 3DGS rendered image and the target style file into a large model, and migrating the target style to the target 3DGS rendered image through the large model to obtain a style reference image; Inputting the plurality of 3DGS rendered images and the style reference image into the large model, and transferring the style of the style reference image to each 3DGS rendered image through the large model to obtain a plurality of stylized rendered images; Training steps: Use the multiple stylized rendering images as a fine-tuning dataset to fine-tune the 3DGS model to update the scene parameters in the 3DGS model, obtaining an optimized 3DGS model.
[0006] Optionally, after the training steps, the method further includes: Using the optimized 3DGS model, render multiple stylized 3DGS rendering images of the three-dimensional scene according to the camera parameters of multiple perspectives; Input the multiple stylized 3DGS rendering images into the large model, and perform anomaly detection on the stylized three-dimensional virtual scene through the large model, where the stylized three-dimensional virtual scene is composed of the multiple stylized 3DGS rendering images; If the large model detects an anomaly in the stylized three-dimensional virtual scene, generate a new 3DGS rendering image that can cover the abnormal area in the stylized three-dimensional virtual scene through the large model; Transfer the target style to the new 3DGS rendering image through the large model based on the new 3DGS rendering image and the target style file to obtain a stylized new 3DGS rendering image; Supplement the stylized new 3DGS rendering image to the fine-tuning dataset; Repeat the training steps and subsequent steps until the large model detects that there is no anomaly in the stylized three-dimensional virtual scene.
[0007] Optionally, determining the target 3DGS rendering image and obtaining the target style file includes: In response to the user's operation of selecting an image from the multiple 3DGS rendering images, determine the target 3DGS rendering image; Obtain the target style file input by the user.
[0008] Optionally, the target style file is one or more of an image, text, and speech.
[0009] Optionally, the camera parameters include the internal and external parameters of the virtual camera.
[0010] Optionally, the scene parameters include the color and texture of the Gaussian sphere of the 3DGS model.
[0011] Optionally, the performing anomaly detection on the stylized three-dimensional virtual scene through the large model includes: Detect whether there are floating objects, defects, or local inconsistency problems in the stylized three-dimensional virtual scene through the large model.
[0012] Optionally, if the large model detects an anomaly in the stylized 3D virtual scene, a new 3DGS rendering image that can cover the anomalous area in the stylized 3D virtual scene is generated by the large model, including: If the large model detects an anomaly in the stylized 3D virtual scene, determine the anomalous area where the anomaly exists in the stylized 3D virtual scene; Determine the virtual camera view that can cover the anomalous area; Generate a new 3DGS rendering image from the virtual camera view.
[0013] According to another aspect of the embodiments of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the stylized 3D virtual scene generation method as described above.
[0014] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the stylized 3D virtual scene generation method as described above is implemented.
[0015] In the embodiments of the present application, a large model is used to generate a style reference image based on the target 3DGS rendering image and the target style file selected by the user as a unified style reference. The style reference image is directly used to perform style conversion on the high-quality 3DGS rendering image generated by the existing 3DGS model, and the stylized rendering images from multiple perspectives after style conversion are used as a fine-tuning dataset to fine-tune and train the 3DGS model to optimize the 3DGS model. Finally, the stylized rendering image is directly obtained using the 3DGS model. In the above manner, it is ensured that the output images after style transfer in all perspectives are consistent in color, texture, and artistic style, and since there is no need to perform style conversion on the original image, strict dependence on the original collected data is avoided, and the style transfer process is more stable.
[0016] The above description is only an overview of the technical solutions of the embodiments of the present application. In order to be able to understand the technical means of the embodiments of the present application more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features, and advantages of the embodiments of the present application more obvious and understandable, the specific embodiments of the present application are specifically given below. Description of the Drawings
[0017] The drawings are only used to illustrate the embodiments and are not considered to be a limitation of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings: Figure 1 A schematic diagram of the application scenario of the embodiments of the present application is shown; Figure 2The flowchart of the method for generating a stylized three-dimensional virtual scene provided by the embodiments of the present application is shown; Figure 3 The flowchart of another method for generating a stylized three-dimensional virtual scene provided by the embodiments of the present application is shown; Figure 4 The structural schematic diagram of the electronic device provided by the embodiments of the present application is shown. Detailed implementation manners
[0018] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein.
[0019] In recent years, the rapid development of computer vision and computer graphics has promoted the continuous progress of three-dimensional scene reconstruction and style transfer technologies. Traditional two-dimensional image style transfer methods use convolutional neural networks (such as VGG19) to extract content and style features, and achieve artistic style conversion through feature fusion and normalization techniques (such as AdaIN). These methods have achieved remarkable results in the field of images, but when directly applied to three-dimensional scenes, due to the differences between multi-view data, problems such as inconsistent styles and visual artifacts are likely to occur.
[0020] In terms of three-dimensional reconstruction, new implicit representation methods such as Neural Radiance Fields (NeRF) and point-based rendering techniques such as 3D Gaussian Splatting (3DGS) can accurately capture the geometric structure and texture information of the scene, recover high-quality three-dimensional scenes from multi-view images, and achieve high-fidelity conversion from two-dimensional images to three-dimensional models. However, traditional 3D scene reconstruction methods mainly focus on restoring the geometric accuracy of the real physical world and have not fully considered the problem of unified transfer of artistic styles.
[0021] In a 3D scene style transfer scheme, the original image is first collected and preprocessed, RGB images are collected from multiple perspectives, and camera posture information is obtained through image screening, resolution adjustment and feature extraction. Then, style transfer processing is performed frame by frame. The content features and style features of each collected original image frame are extracted using the pre-trained VGG19 convolutional neural network (the statistical relationship between the activation maps of multiple convolutional layers in VGG19 is used to represent the style). Then, the multi-scale features are fused through the feature pyramid network, and the style transfer network based on methods such as AdaIN is used to achieve the effect of converting the original image into the target artistic style image, and a stylized image is obtained after each frame of the original image is stylized. Then, the 3D scene is reconstructed based on multiple stylized images. Multiple stylized images are used as fine-tuning data sets, combined with NeRF or 3DGS technology, the original 3D scene is optimized and reconstructed, and finally a stylized 3D scene is output.
[0022] This method directly performs style transfer on the collected original image, and then integrates the style transfer result into the three-dimensional scene through the reconstruction module. The reconstruction process is highly dependent on the original image, which relies on the consistency of the style details of the images from each perspective to ensure the consistency of the scene in different areas. However, sometimes the original image may have been lost and may be unavailable, and the original image is often affected by factors such as the shooting environment, lighting conditions, noise, and shooting angle, resulting in unnecessary interference and unstable style transfer effects. Therefore, the excessive reliance on the original image makes it difficult for the entire process to ensure high consistency in the final output stylized scene when faced with uneven image quality.
[0023] Since each frame of the image is processed using independent style transfer, the style features extracted by the VGG19 convolutional neural network are unstable and lack a unified style benchmark, resulting in differences in color, texture, and artistic expression in the stylized images generated in different regions / different perspectives, resulting in style drift or visual artifacts, making it difficult to ensure multi-perspective style consistency. This multi-perspective inconsistency problem directly affects the overall beauty and coherence of the three-dimensional scene.
[0024] How to achieve artistic style conversion with unified global style and consistent visual effects while ensuring the true restoration of the geometric structure of the three-dimensional scene has become a key issue that needs to be urgently solved in the current technical field.
[0025] The present application provides a method for generating a stylized three-dimensional virtual scene. A large model is used to generate a style reference image based on the target 3DGS rendering image and the target style file selected by the user as a unified style reference. The style reference image is directly used to perform style conversion on the high-quality 3DGS rendering images generated by the existing 3DGS model. The stylized rendering images from multiple perspectives after style conversion are used as a fine-tuning dataset to fine-tune and train the 3DGS model to optimize the 3DGS model. Finally, the stylized rendering images are directly obtained using the 3DGS model. In this way, the consistency of the output images after style transfer in terms of color, texture, and artistic style is ensured for all perspectives. Moreover, since there is no need to perform style conversion on the original images, the strict dependence on the original acquisition data is avoided, and the style transfer process is more stable.
[0026] Figure 1 FIG. shows a schematic diagram of an application scenario of an embodiment of the present application. As shown in the figure, in a three-dimensional scene, point cloud data and image data are obtained by scanning the three-dimensional scene using a three-dimensional scanning and reconstruction device 1. For example, the three-dimensional scanning and reconstruction device 1 can be equipped with a lidar, a camera, and a computing unit. The lidar collects the point cloud data of the three-dimensional scene, the camera acquires images of different perspectives of the three-dimensional scene, and the computing unit performs fusion processing on the point cloud data and the images to achieve three-dimensional reconstruction. The three-dimensional scanning and reconstruction device 1 can be a handheld device or a non-handheld device. If the three-dimensional scanning and reconstruction device 1 is a handheld device, the user holds the three-dimensional scanning and reconstruction device 1 and moves it in the three-dimensional scene to collect the point cloud data and image data of the three-dimensional scene. If the three-dimensional scanning and reconstruction device 1 is a non-handheld device, the three-dimensional scanning and reconstruction device 1 can move autonomously in the three-dimensional scene to collect the point cloud data and image data of the three-dimensional scene.
[0027] The terminal device 2 is communicatively connected to the three-dimensional scanning and reconstruction device 1, receives the three-dimensional reconstruction result obtained after the three-dimensional scanning and reconstruction device 1 performs three-dimensional reconstruction on the three-dimensional scene, and displays it. Of course, the three-dimensional scanning and reconstruction device 1 can also only collect the original point cloud data and image data and send them to the terminal device 2, and the terminal device 2 performs fusion processing on the data to perform three-dimensional reconstruction, obtain, and display the three-dimensional reconstruction result. The terminal device 2 can be a device such as a mobile phone, a tablet computer, a laptop computer, a desktop computer, a server, etc. In the figure, a mobile phone is used as an example. The terminal device 2 and the three-dimensional scanning and reconstruction device 1 can be connected by means of wired communication or wireless communication.
[0028] In some other scenarios, the lidar and the camera can also be separate devices and not integrated into the three-dimensional scanning and reconstruction device 1. They send the data they collect to the terminal device 2, and the terminal device 2 performs data fusion processing to obtain the three-dimensional reconstruction result and display it.
[0029] The above three-dimensional reconstruction result can be a 3DGS model. The 3DGS model can embed target style features, and opening the 3DGS model can display a stylized three-dimensional virtual scene. The virtual scene is a three-dimensional environment generated by the above electronic device, usually including information such as geometric shapes, textures, and lighting, and can be interactively browsed by an observer from various perspectives. The virtual scene can be sourced from the reconstruction of a real scene or an artificial synthetic scene. The virtual scene output by this application is a reconstruction of a real scene and has been transformed in visual style (such as transformed into a cartoon style, an oil painting style, etc.), so it is called a stylized three-dimensional virtual scene.
[0030] Figure 2 The flowchart of the method for generating a stylized three-dimensional virtual scene provided by an embodiment of this application is shown. This method is executed by an electronic device. The electronic device can be Figure 1 the three-dimensional scanning and reconstruction device 1 or the terminal device 2 in Figure 2 As shown, this method includes the following steps: S101, using the 3DGS model, render multiple 3DGS rendered images of the three-dimensional scene according to the camera parameters of multiple perspectives.
[0031] Three-dimensional scene reconstruction is the process of restoring the three-dimensional geometric structure and texture appearance of a scene based on two-dimensional image data from multiple perspectives. Common methods include point cloud reconstruction based on Structure from Motion (SfM) and Multi-View Stereo (MVS), NeRF implicit modeling, and point primitive representation methods such as 3DGS. The reconstruction result can be in the form of sparse or dense point clouds, voxel grids, triangular mesh models, or volume rendering models implicitly represented by neural networks.
[0032] 3DGS is a three-dimensional scene reconstruction and rendering technology. Its core idea is to represent points in a scene with three-dimensional Gaussian functions and project these Gaussian functions onto a two-dimensional image plane for rendering. 3DGS combines the advantages of explicit point cloud representation and continuous field representation, and can achieve efficient real-time rendering while ensuring rendering quality. This application uses the 3DGS model as a three-dimensional reconstruction module to generate a stylized three-dimensional virtual scene. The 3DGS model supports rendering and observing from any new perspective. The 3DGS model refers to a software module that reconstructs a three-dimensional scene using 3D Gaussian sphere technology, that is, a module that turns the original images or point clouds captured by a camera into a renderable 3D scene.
[0033] This step obtains multiple 3DGS rendered images based on the existing 3DGS model, rather than relying on the original pictures captured by the original camera. The 3DGS model has precise geometric and texture reconstruction capabilities, ensuring that the images from each perspective truly reflect the scene structure.
[0034] The camera parameters used during rendering include the internal and external parameters of the virtual camera. Among them, the internal parameters are the parameters that describe the internal imaging characteristics of the camera, such as focal length, principal point coordinates, lens distortion, etc., which are mainly related to the hardware structure of the camera itself and are inherent attributes of the camera. The external parameters are the parameters that describe the position and pose (i.e., the pose information) of the camera in the world coordinate system, including the rotation matrix and the translation vector, which reflect the external geometric relationship of the camera relative to the world coordinate system. According to the camera parameters of multiple perspectives, multiple 3DGS rendered images corresponding to the perspectives of the three-dimensional scene can be rendered.
[0035] The number of 3DGS rendered images obtained by rendering an indoor scene is approximately 1000 - 2000.
[0036] S102. Determine the target 3DGS rendered image and obtain the target style file.
[0037] Among them, the target 3DGS rendered image is one of the multiple 3DGS rendered images, and the target style file is used to express the target style (the style that the user hopes to finally generate). After the image style transfer in this application, the obtained three-dimensional virtual scene is a scene with the target style.
[0038] Image Style Transfer is a technology that combines the content of one image with the artistic style of another image to generate an image with a new style. Among them, "content" usually refers to the high-level semantics or macro structure of the image, and "style" refers to the visual patterns such as colors, textures, and brushstrokes in the image. Classical neural style transfer uses a convolutional neural network to extract the features of the content image and the style image, and generates a synthetic image that retains both the content and presents the target style through optimization. In the subsequent steps of this application, a large model is used to convert real images into images of any artistic style, providing a consistent style input from multiple perspectives for the subsequent three-dimensional virtual scene style transfer.
[0039] Step S102 includes the following steps: S1021. Respond to the user's operation of selecting one image from multiple 3DGS rendered images and determine the target 3DGS rendered image.
[0040] The user can select an image from the multiple 3DGS rendered images generated in step S101 that can see most of the scenes in the three-dimensional scene within the largest viewing angle range as the target 3DGS rendered image, which is used as a reference for scene description to express what types of objects are in the scene.
[0041] S1022. Obtain the target style file input by the user.
[0042] The user can select any style as the target style, such as the cartoon style, the oil painting style, etc. The target style file is one or more of images, texts, and voices. For example, the user can directly input a target style image, or input a text describing the target style, or input a voice describing the target style, or a combination of the above files.
[0043] S103, input the target 3DGS rendered image and the target style file into the large model, and transfer the target style to the target 3DGS rendered image through the large model to obtain a style reference image.
[0044] The large model adopted in this application is an advanced artificial intelligence model with multi-modal generation and understanding capabilities, which can simultaneously process various input forms such as texts, images, and audios, and also has output forms in various modes such as texts, images, and audios. The large model has the following functions: Joint understanding of text and image: Support multi-modal input, input any image and / or text description, and extract high-level semantic representations; Image generation and style transfer: Have the capabilities of image stylization, image completion, and image diffusion reconstruction; Consistency modeling ability: Can generate images with consistent styles across perspectives.
[0045] The large model can use, including but not limited to, GPT-4 Omni (GPT4o), etc.
[0046] The user selects one of the generated multiple 3DGS rendered images as the target 3DGS rendered image for content reference, and at the same time provides the target style file. Input both of them into the large model, and use the fusion and generation capabilities of the large model to generate a unified style reference image based on the above two-way input. This style reference image is used as a global style reference to guide subsequent style conversions for each perspective, ensuring the consistency of the output images in terms of color, texture, and artistic style.
[0047] S104, input the multiple 3DGS rendered images and the style reference image into the large model, and transfer the style of the style reference image to each 3DGS rendered image through the large model to obtain multiple stylized rendered images.
[0048] In this step, multiple 3DGS rendered images obtained in S101 and the style reference image generated in S103 are used as dual inputs and fed into the large model simultaneously. The large model fuses the two inputs to achieve direct style transfer for the rendered images of each perspective, generating stylized rendered images. This process directly performs style transfer on the 3DGS rendered images based on the unified style reference image, migrating all the 3DGS rendered images to the style of the style reference image serving as the benchmark, ensuring the unity of the converted image styles under multiple perspectives, while maintaining the geometric details of the original scene and reducing the dependence on the original acquisition data, making the finally generated stylized three-dimensional virtual scene more stable and reliable in terms of data sources.
[0049] The prior art uses the VGG19+AdaIN method, whose style features are statistical (such as the Gram matrix), and it cannot achieve global control and consistency abstraction of style transfer. This application uses a large model such as GPT-4o. Different from the statistical features of the style in the prior art, the large model can understand artistic styles and abstract expressions, and can ensure the unity of the converted image styles under multiple perspectives while maintaining the geometric details of the original scene. Based on steps S103 and S104, a large model is introduced to generate a unified style reference image, and this reference is used as a reference. Then, the large model is used to achieve direct style transfer for the rendered images of all perspectives, ensuring a high degree of consistency in the style of the output images.
[0050] S105: Use multiple stylized rendered images as a fine-tuning dataset to fine-tune and train the 3DGS model to update the scene parameters in the 3DGS model, obtaining an optimized 3DGS model.
[0051] This step is a training step. By using multiple stylized rendered images generated in S104 as a fine-tuning dataset (i.e., supervised data with unified image styles), the 3DGS model is fine-tuned and trained to optimize the scene parameters in the 3DGS model. This optimization process is equivalent to feeding back the style transfer effect in S104 to the 3DGS model, enabling the 3DGS model to embed the features of the target style when generating subsequent rendered images, thereby improving the overall style consistency and visual quality of the three-dimensional scene.
[0052] The scene parameters include the color and texture of the Gaussian spheres in the 3DGS model. The 3DGS model contains thousands of Gaussian spheres, and each point has attributes such as position, direction, color, size, and texture (i.e., transparency). Among these attributes, color and texture directly affect the style and details of the rendered image. Optimize the color and texture of the 3DGS model based on the 3DGS rendered images of the target style, so that the 3DGS model can automatically reflect the target style when generating subsequent rendered images.
[0053] In the above training process, first, the 3DGS model is used to generate rendered images from multiple perspectives, where the "multiple perspectives" correspond one-to-one to the multiple perspectives based on which multiple 3DGS rendered images are generated in step S101. Then, the style loss is calculated for each generated rendered image and the stylized rendered image of the same perspective in the fine-tuning dataset, and the color and transparency of the Gaussian sphere are continuously optimized through backpropagation, gradually approaching the target style. That is, during the training process, the 3DGS model refers to multiple 3DGS rendered images that have been stylized by the target style, and fine-tunes the parameters such as the color and texture of each Gaussian sphere in the 3DGS model to obtain an optimized 3DGS model. The user can open the 3DGS model file through software and view the stylized three-dimensional virtual scene.
[0054] The 3DGS model obtained through steps S101 - S105 embeds the target style features, and what is seen when opening the 3DGS model file with software is a stylized three-dimensional virtual scene. However, floating objects, defects, or local inconsistencies may occur in such a three-dimensional virtual scene. The prior art relies on manual editing to address the problems of floating objects, defects, or local inconsistencies in the three-dimensional virtual scene, which increases costs and reduces the overall system efficiency. In some embodiments, a large model is used to perform anomaly detection on the floating objects, defects, and local inconsistencies in the stylized three-dimensional virtual scene rendered from the optimized 3DGS model, and the 3DGS model is iteratively optimized according to the detection results until the detection results of the large model are normal.
[0055] Figure 3 The flowchart of another method for generating a stylized three-dimensional virtual scene provided by an embodiment of the present application is shown. This method is executed by an electronic device. The electronic device can be Figure 1 the three-dimensional scanning and reconstruction device 1 or the terminal device 2 in Figure 3 As shown, this method includes the following steps: S201, using the 3DGS model, render multiple 3DGS rendered images of the three-dimensional scene according to the camera parameters of multiple perspectives.
[0056] S202, determine the target 3DGS rendered image and obtain the target style file.
[0057] S203, input the target 3DGS rendered image and the target style file into the large model, and transfer the target style to the target 3DGS rendered image through the large model to obtain a style reference image.
[0058] S204, input multiple 3DGS rendered images and the style reference image into the large model, and transfer the style of the style reference image to each 3DGS rendered image through the large model to obtain multiple stylized rendered images.
[0059] S205. Use multiple stylized rendered images as a fine-tuning data set to fine-tune the 3DGS model to update the scene parameters in the 3DGS model, and obtain an optimized 3DGS model.
[0060] Steps S201 to S205 are similar to the aforementioned steps S101 to S105, and will not be elaborated here.
[0061] S206. Use the optimized 3DGS model to render multiple stylized 3DGS rendered images according to the camera parameters of multiple perspectives.
[0062] When the optimized 3DGS model generates multi-perspective images, the generated multi-perspective images will carry the target style by themselves, and the images are consistent in visual style. The three-dimensional scene in this step is the same scene as the three-dimensional scene in S101, but the perspective can be different from the perspective in S101.
[0063] S207. Input multiple stylized 3DGS rendered images into a large model, and use the large model to perform anomaly detection on the stylized three-dimensional virtual scene to determine whether there is an anomaly in the stylized three-dimensional virtual scene. If so, execute step S208; otherwise, end the process.
[0064] Among them, the stylized three-dimensional virtual scene is composed of multiple stylized 3DGS rendered images. Anomaly detection is to detect whether there are floating objects, defects, or local inconsistency problems in the stylized three-dimensional virtual scene. The large model in the embodiments of the present application can use the picture understanding ability to find floating objects, defects, local inconsistencies, etc. in the pictures. If the large model detects any of the floating objects, defects, or local inconsistency problems in the stylized three-dimensional virtual scene, it is determined that there is an anomaly in the stylized three-dimensional virtual scene; if the large model detects none of the floating objects, defects, and local inconsistency problems in the stylized three-dimensional virtual scene, it is determined that there is no anomaly in the stylized three-dimensional virtual scene.
[0065] S208. Use the large model to generate new 3DGS rendered images that can cover the abnormal areas in the stylized three-dimensional virtual scene.
[0066] When the large model detects that there is an anomaly in the stylized three-dimensional virtual scene when viewed from a certain perspective, the large model automatically generates a new stylized 3DGS rendered image from this perspective as supplementary data, and supplements it to the fine-tuning data set to optimize and train the 3DGS model again. Step S208 includes the following steps: S2081. If the large model detects that there is an anomaly in the stylized three-dimensional virtual scene, determine the abnormal area where there is an anomaly in the stylized three-dimensional virtual scene.
[0067] S2082. Determine the virtual camera perspective that can cover the abnormal area.
[0068] S2083, generate a new 3DGS rendered image from the perspective of a virtual camera.
[0069] The new 3DGS rendered image regenerated by the large model does not have the aforementioned floating objects, defects, and local inconsistency problems.
[0070] S209, transfer the target style to the new 3DGS rendered image based on the new 3DGS rendered image and the target style file through the large model to obtain a stylized new 3DGS rendered image.
[0071] S210, supplement the stylized new 3DGS rendered image to the fine-tuning dataset, and go to step S205.
[0072] Use the stylized new 3DGS rendered image without floating objects, defects, and local inconsistency problems to optimize and train the 3DGS model optimized in the previous time again, reducing the probability of floating objects, defects, and local inconsistency problems occurring when the 3DGS model renders the scene. When the loop is repeated multiple times until the large model detects no anomalies, the optimization of the 3DGS model is completed, and the final available 3DGS model is obtained, ending the process.
[0073] After the process ends, an iteratively optimized 3DGS model is obtained. The user can open the 3DGS model file through the software and view the stylized three-dimensional virtual scene.
[0074] In the above embodiments, by utilizing the image generation and intelligent detection capabilities of the large model, it is possible to automatically identify floating objects, defects, and local inconsistency problems in the stylized three-dimensional virtual scene rendered by the 3DGS model, generate supplementary data from a specific perspective, correct the current style transfer result, and re-feed the corrected data to the 3DGS model to further automatically optimize the parameters of the 3DGS model until the anomaly detection result of the large model meets the preset standard (that is, no anomalies are detected), and finally end the feedback optimization process, improving the overall visual effect of the 3DGS model rendering.
[0075] Apply the above Figure 2 and Figure 3 An electronic device implementing the method for generating a stylized three-dimensional virtual scene of the above embodiment has the following software modules: 3DGS rendering module: that is, the 3DGS model, which renders multiple 3DGS rendered images of the three-dimensional scene according to the camera parameters of multiple perspectives for the preprocessed multi-perspective data, ensuring the accurate restoration of the geometric and texture information of the initial three-dimensional scene.
[0076] Style transfer module: Using a large model, based on the 3DGS rendered images selected by the user and the target style file, output a unified style reference image to provide a reference for global style transfer, and generate stylized rendered images based on the aforementioned multiple 3DGS rendered images and the style reference image.
[0077] 3DGS model optimization module: Using the generated stylized rendered images as a fine-tuning data set, update the color and texture parameters in the 3DGS model to obtain an optimized 3DGS model.
[0078] Apply the above Figure 3 An electronic device implementing the stylized three-dimensional virtual scene generation method of the above embodiment further includes the following software modules: Intelligent detection module: Based on the detection ability of the large model, automatically identify floating objects, defects, and local inconsistencies generated during the style conversion of the 3DGS model, and generate supplementary data to correct the output to ensure that the entire system output has no floating objects or visual artifacts.
[0079] Figure 4 FIG. shows a schematic structural diagram of an electronic device provided by an embodiment of the present application. The specific implementation of the electronic device is not limited in the specific embodiments of the present application.
[0080] As Figure 4 shown, the electronic device 300 may include: a processor 302 and a memory 304.
[0081] Among them, the memory 304 is used to store a computer program 306. The memory 304 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory. The computer program 306 may include computer-executable instructions.
[0082] The processor 302 is used to execute the computer program 306 to implement the above embodiment of the stylized three-dimensional virtual scene generation method.
[0083] The processor 302 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the electronic device 300 may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0084] An embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the embodiment of the above-mentioned stylized three-dimensional virtual scene generation method.
[0085] An embodiment of the present application provides a computer program that can be executed by a processor to implement the embodiment of the above-mentioned stylized three-dimensional virtual scene generation method.
[0086] An embodiment of the present application provides a computer program product, which includes a computer program that when executed by a processor, implements the embodiment of the above-mentioned stylized three-dimensional virtual scene generation method.
[0087] In several embodiments provided by the present application, if any function is implemented in the form of a software functional module / unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the technical solution of the present application can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for causing a computer device (which can be an electronic device such as a personal computer or a server) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store computer program codes.
[0088] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings provided herein. The structure required to construct such systems will be apparent from the above description. In addition, the embodiments of the present application are not directed to any specific programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the description of the specific language above is to disclose the best implementation mode of the present application.
[0089] It should be noted that the above embodiments illustrate the present application rather than limit the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a claim enumerating several devices, several units or modules of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words may be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.
[0090] The above embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A method for generating a stylized three-dimensional virtual scene, characterized in that, The method includes: Using a 3DGS model to render multiple 3DGS rendered images of a three-dimensional scene according to camera parameters from multiple perspectives; Determining a target 3DGS rendered image and obtaining a target style file, where the target 3DGS rendered image is one of the multiple 3DGS rendered images, and the target style file is used to express the target style; Inputting the target 3DGS rendered image and the target style file into a large model, and migrating the target style to the target 3DGS rendered image through the large model to obtain a style reference image; Inputting the multiple 3DGS rendered images and the style reference image into the large model, and migrating the style of the style reference image to each 3DGS rendered image through the large model to obtain multiple stylized rendered images; Training step: Using the multiple stylized rendered images as a fine-tuning data set to fine-tune and train the 3DGS model to update the scene parameters in the 3DGS model, and obtaining an optimized 3DGS model.
2. The method according to claim 1, wherein After the training step, the method further includes: Using the optimized 3DGS model to render multiple stylized 3DGS rendered images of a three-dimensional scene according to camera parameters from multiple perspectives; Inputting the multiple stylized 3DGS rendered images into the large model, and performing anomaly detection on the stylized three-dimensional virtual scene through the large model, where the stylized three-dimensional virtual scene is composed of the multiple stylized 3DGS rendered images; If the large model detects an anomaly in the stylized three-dimensional virtual scene, generating a new 3DGS rendered image that can cover the abnormal area in the stylized three-dimensional virtual scene through the large model; Migrating the target style to the new 3DGS rendered image through the large model based on the new 3DGS rendered image and the target style file to obtain a stylized new 3DGS rendered image; Supplementing the stylized new 3DGS rendered image to the fine-tuning data set; Repeating the training step and subsequent steps until the large model detects that there is no anomaly in the stylized three-dimensional virtual scene.
3. The method according to claim 1, wherein The determining the target 3DGS rendered image and obtaining the target style file includes: Responding to an operation where the user selects an image from the multiple 3DGS rendered images to determine the target 3DGS rendered image; Obtaining the target style file input by the user.
4. The method according to claim 1, wherein The target style file is one or more of an image, text, and speech.
5. The method according to claim 1, wherein The camera parameters include the internal and external parameters of a virtual camera.
6. The method according to claim 1, characterized in that, The scene parameters include the color and texture of the Gaussian sphere of the 3DGS model.
7. The method according to claim 2, wherein The performing anomaly detection on the stylized three-dimensional virtual scene through the large model includes: Detecting whether there are floating objects, defects, or local inconsistency problems in the stylized three-dimensional virtual scene through the large model.
8. The method according to claim 2, wherein The if the large model detects an anomaly in the stylized three-dimensional virtual scene, generating a new 3DGS rendered image that can cover the abnormal area in the stylized three-dimensional virtual scene through the large model includes: If the large model detects an anomaly in the stylized three-dimensional virtual scene, determine the anomalous area with anomalies in the stylized three-dimensional virtual scene; Determine a virtual camera view that can cover the anomalous area; Generate a new 3DGS rendering image from the virtual camera view.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the stylized three-dimensional virtual scene generation method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the stylized three-dimensional virtual scene generation method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Rendering method from three-dimensional model to two-dimensional image based on deep learning
CN110211192A
Three-dimensional scene rendering method, device and equipment
CN115311395A
Three-dimensional scene style migration method and device, equipment and storage medium
CN116934936A
3D scene model generation method and device, electronic equipment and storage medium
CN119399373A
Three-dimensional scene style migration method, electronic equipment and storage medium
CN119991910A
Cited By
Artwork illumination effect cross-medium migration rendering system and method thereof
CN120580341A
Controllable generation method and device of three-dimensional effect picture, equipment and storage medium
CN121353552A