Three-dimensional scene stylization processing method and device, equipment and storage medium
By performing viewing angle grouping and feature extraction methods on three-dimensional scene images, combined with diffusion model processing, local consistency and information leakage problems in stylization of three-dimensional scenes are solved, and a three-dimensional scene model with more consistent stylization effects is generated.
Patent Information
- Application Number
- CN202510584857.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-19
AI Technical Summary
The prior art is difficult to guarantee local consistency in three-dimensional scene stylized processing, and there is a problem of information leakage.
By grouping the images of three-dimensional scenes according to the perspective, and performing feature extraction and diffusion model processing based on the images of adjacent perspectives, combining style features and content features, stylized images are generated, and parameter adjustments are made to the three-dimensional scene model.
The consistency of local features is achieved, information leakage during the stylization process is avoided, and the local consistency of the rendering effect is improved.
Smart Images

Figure CN120510264A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, apparatus, device, and storage medium for stylizing a three-dimensional scene. Background Art
[0002] 3D stylization is a technique that applies a specific artistic style or visual effect to a 3D model or scene. Through a series of complex algorithms and models, the geometry, color, texture, and other aspects of the 3D model are adjusted and optimized to give it a specific artistic style.
[0003] Existing technologies can employ stylization methods based on Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS). NeRF-based stylization achieves more precise color control by incorporating a loss function designed for image style transfer during training, or performs geometric stylization by mimicking a reference style to achieve consistent stylization across new perspectives. 3DGS stylization offers fast construction, real-time rendering, and high-quality rendering.
[0004] However, when implementing 3D stylization based on existing technologies, not only is local consistency difficult to ensure, but there is also the problem of information leakage. For example, during the stylization process, content information may leak into style control, or style information may leak into content information, resulting in poor stylization effects. Summary of the Invention
[0005] The purpose of this application is to address the deficiencies in the above-mentioned prior art and to provide a method, apparatus, device, and storage medium for stylizing a three-dimensional scene, so as to solve the problems in the prior art of difficulty in ensuring local consistency and information leakage during the stylization process.
[0006] To achieve the above objectives, the technical solutions adopted in this application are as follows:
[0007] In a first aspect, the present application provides a method for stylizing a three-dimensional scene, the method comprising:
[0008] Acquire a style image and a plurality of to-be-processed images of the to-be-processed 3D scene model, wherein each of the to-be-processed images is an image of the to-be-processed 3D scene model at a different perspective;
[0009] Performing encoding processing on the style image and the plurality of images to be processed respectively to obtain style features corresponding to the style image and content features corresponding to the plurality of images to be processed;
[0010] Grouping the plurality of images to be processed according to viewing angles to obtain a plurality of adjacent viewing angle image groups, each adjacent viewing angle image group including a plurality of images to be processed having adjacent viewing angles;
[0011] Each of the adjacent view image groups, the content features, and the style features is input into a diffusion model for stylization processing to obtain a stylized image corresponding to each of the images to be processed, and parameters of the three-dimensional scene model to be processed are adjusted according to each of the stylized images to obtain a stylized target three-dimensional scene model.
[0012] In a second aspect, the present application provides a stylized processing device for a three-dimensional scene, the device comprising:
[0013] An acquisition module, configured to acquire a style image and a plurality of to-be-processed images of the to-be-processed 3D scene model, wherein each of the to-be-processed images is an image of the to-be-processed 3D scene model at a different perspective;
[0014] An encoding module, configured to encode the style image and the plurality of images to be processed, respectively, to obtain style features corresponding to the style image and content features corresponding to the plurality of images to be processed;
[0015] a grouping module, configured to group the plurality of images to be processed according to viewing angles to obtain a plurality of adjacent viewing angle image groups, each of the adjacent viewing angle images comprising a plurality of images to be processed having adjacent viewing angles;
[0016] A processing module is configured to input each of the adjacent viewpoint image groups, the content features, and the style features into a diffusion model for stylization processing to obtain a stylized image corresponding to each of the images to be processed, and to adjust parameters of the 3D scene model to be processed based on each of the stylized images to obtain a stylized target 3D scene model.
[0017] In a third aspect, an embodiment of the present application further provides an electronic device comprising: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of a three-dimensional scene stylization processing method as described in any one of the first aspects.
[0018] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of a three-dimensional scene stylization processing method as described in any one of the first aspects are executed.
[0019] The beneficial effects of this application are: by grouping the images to be processed according to adjacent viewpoints and fusing the features of the images from these adjacent viewpoints, the extracted content features can more comprehensively represent the contextual information of the adjacent viewpoints, thereby ensuring the consistency of local features. By extracting features from the style image and the image to be processed separately, the style feature and content feature extraction processes are decoupled, thereby solving the problem of information leakage during the stylization process.
[0020] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 A schematic diagram of an application scenario provided by an embodiment of the present application is shown;
[0023] Figure 2 A flowchart of a stylized processing method for a three-dimensional scene provided in an embodiment of the present application is shown;
[0024] Figure 3 A flowchart of obtaining an image to be processed provided by an embodiment of the present application is shown;
[0025] Figure 4 A flowchart of extracting content features and style features provided by an embodiment of the present application is shown;
[0026] Figure 5 A flow chart of extracting content features provided by an embodiment of the present application is shown;
[0027] Figure 6 A flowchart of another method for extracting content features provided by an embodiment of the present application is shown;
[0028] Figure 7 The following is an overall flow chart of a stylized processing of a three-dimensional scene provided by an embodiment of the present application;
[0029] Figure 8 A flowchart of obtaining a stylized image provided by an embodiment of the present application is shown;
[0030] Figure 9 A flowchart of obtaining image content features provided by an embodiment of the present application is shown;
[0031] Figure 10 A flowchart of obtaining initial image features provided by an embodiment of the present application is shown;
[0032] Figure 11 A flowchart of another method for obtaining a stylized image provided by an embodiment of the present application is shown;
[0033] Figure 12 A flowchart of stylized processing based on a stylized image provided by an embodiment of the present application is shown;
[0034] Figure 13 A flow chart for determining a target loss value provided by an embodiment of the present application is shown;
[0035] Figure 14 A schematic structural diagram of a three-dimensional scene stylization processing device provided in an embodiment of the present application is shown;
[0036] Figure 15 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work fall within the scope of protection of the present application.
[0038] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.
[0039] In the prior art, in order to achieve stylized processing of three-dimensional scenes, commonly used processing methods include: implicit methods based on neural radiation fields, 3D Gaussian splashing methods, and diffusion-based 2D stylization methods.
[0040] Among them, the implicit method based on NeRF achieves more precise control of color during training by combining a specific image style transfer loss. However, this approach usually faces the problems of long training time and slow rendering speed.
[0041] The 3D Gaussian splattering method has the advantages of fast construction, real-time rendering, and high-quality rendering. However, although this method has many improvements in speed, the stylized effect often suffers from local color and texture inconsistencies.
[0042] Diffusion-based 2D stylization methods are prone to leakage of style and content information when applied to 3D stylization tasks, and this method also has the problem of poor local consistency.
[0043] In summary, existing 3D stylization methods have problems such as poor local consistency of rendering results and easy information leakage during the stylization process.
[0044] Based on this, this application proposes a method for stylizing 3D scenes. By grouping 3D scene images by viewpoint and ensuring local color and texture consistency in the stylized results based on images from adjacent viewpoints, and by decoupling the extraction of style features from image content features, this method effectively reduces the mutual interference between style and content information, thereby preventing information leakage. This method thus addresses the issue of information leakage during the stylization process while ensuring local consistency in the rendering effect.
[0045] First, let's explain 3D scenes. A 3D scene uses virtualization technology to realistically simulate and digitally represent various physical forms and spatial relationships in the real world. A 3D scene includes at least one 3D model, such as an object model or terrain model, as well as material and texture information, light source information, and animation effect information.
[0046] 3D scene stylization can be to process the 3D model, material and texture information, light source information, animation effect information and other contents in the 3D scene based on a specified style, so that the contents in the 3D scene are rendered according to the specified style, and a 3D scene after stylization is obtained.
[0047] The method of this application can be applied to stylize a three-dimensional scene. Specifically, refer to Figure 1 , is a schematic diagram of an application scenario provided in this application. The user can determine the 3D scene to be stylized and select the style by selecting a style image. At the same time, the user can also enter some auxiliary instructions to indicate which models in the 3D scene need to be stylized. For example, "Stylize the 3D scene, and refer to the selected style image for the stylized content." In this way, the electronic device can stylize the 3D scene based on the style image selected by the user and the auxiliary instructions, thereby obtaining a stylized 3D scene.
[0048] Next, combine Figure 2, further describes the stylized processing method of the three-dimensional scene in this application. The execution subject of this method can be an electronic device, such as Figure 2 As shown, the method includes:
[0049] S201 , obtaining a style image and a plurality of to-be-processed images of a to-be-processed 3D scene model, wherein each to-be-processed image is an image of the to-be-processed 3D scene model at a different perspective.
[0050] Among them, the three-dimensional scene model to be processed can be the aforementioned three-dimensional scene, which includes multiple object models, material and texture information, light source information, animation effect information, etc. The three-dimensional scene model to be processed is stylized, and the object models, material and texture information, light source information, animation effect information, etc. in the three-dimensional scene model can be rendered into a specified style.
[0051] Before stylization processing is performed, a style image may be obtained. For example, multiple style images may be provided for the user to select, and the style image may be determined based on the user's selection instruction. In another example, the user may enter a style instruction to describe the desired image style, and the electronic device may push multiple images that match the description based on the user's text description for the user to select.
[0052] A style image is an image that provides a specific visual style in the image style transfer task. It contains features such as color, texture, and lines. These features are extracted and applied to other images to achieve style transfer.
[0053] The image to be processed is the image for which style transfer is performed. It contains structural and semantic information about objects and scenes that the user wishes to preserve. For example, the image to be processed can be a 2D projection of the 3D scene model to be processed from different perspectives. The image to be processed can represent the content information within the 3D scene model to be processed.
[0054] As a possible implementation method, the 3D scene model to be processed can be a 3DGS scene. 3DGS scene is a technology that uses a Gaussian distributed point cloud model to represent and render a 3D scene. The objects and environment in the scene are described by a set of 3D Gaussian points, each of which has attributes such as position, shape, color, and transparency. These Gaussian points can be efficiently combined to generate a complete 3D scene image observed from any angle. In this application, the 3D scene model to be processed can be projected to obtain images to be processed at different perspectives.
[0055] As another possible implementation method, the 3D scene model to be processed may be screenshotted at different perspectives to obtain multiple images to be processed, or the images to be processed at different perspectives may be obtained by controlling camera parameters and rendering.
[0056] S202 : Encode the style image and the multiple images to be processed respectively to obtain style features corresponding to the style image and content features corresponding to the multiple images to be processed.
[0057] Optionally, by encoding the style image, a feature vector that can describe the style can be extracted to obtain a style feature. By encoding the image to be processed, a feature vector that can describe the image content can be obtained to obtain a content feature.
[0058] Among them, style features can be features extracted from style images that can describe their visual style. Style features are related to visual elements such as color, texture, and lines of the image, but do not involve the specific content and structure of the image.
[0059] Content features are features extracted from the image to be processed that can describe the structure and semantic information of objects, scenes, etc. Content features are related to the high-level semantics of the image, such as the shape, position, and category of the object.
[0060] S203 , grouping the multiple images to be processed according to viewing angles to obtain multiple adjacent viewing angle image groups, where each adjacent viewing angle image group includes multiple images to be processed with adjacent viewing angles.
[0061] The multiple images to be processed refer to a group of images that need to be processed. These images come from different viewpoints and have different viewpoint characteristics and content information. The multiple images to be processed can be a collection of images of the same scene captured from multiple angles, or a sequence of images captured from different positions and angles.
[0062] Optionally, grouping by viewpoint may refer to classifying multiple images to be processed according to viewpoint features of the images to be processed. The viewpoint features may include information such as shooting angle, direction, and position, which determine how objects or scenes in the image are presented.
[0063] An adjacent-view image group can be a collection of images with adjacent or similar perspective characteristics. In these image groups, each image is continuous or close in perspective to the other images. The images in the adjacent-view image group may capture different parts or angles of the same scene or object, but the perspective changes are relatively small.
[0064] By grouping the images to be processed according to perspective, images with similar or continuous perspective features can be grouped together so that subsequent processing can consider the perspective relationship and spatial continuity between images, which helps to utilize the perspective correlation between images in the processing process and provide more consistent and continuous information.
[0065] S204: Input each adjacent view image group, content feature, and style feature into a diffusion model for stylization processing to obtain a stylized image corresponding to each image to be processed, and adjust parameters of the 3D scene model to be processed based on each stylized image to obtain a stylized target 3D scene model.
[0066] The diffusion model is a generative model that gradually adds noise to data and then learns how to reverse this process to generate new data. In this application, the diffusion model can effectively combine content and style features in stylization processing, generating images that retain the original content while also retaining the target style.
[0067] Optionally, the diffusion model can use adjacent view image groups as input information, with content features and style features as conditional input information. The diffusion model can combine the image to be processed in each adjacent view image group with other images in the adjacent view image group to ensure consistency in local color and texture. Furthermore, the diffusion model can combine content features and style features to generate a stylized image that is consistent with the style image and the semantic structure of the image to be processed.
[0068] After obtaining the stylized image, in one possible implementation, the 3D scene model to be processed can be stylized based on the stylized image, and the parameters of the 3D scene model to be processed can be adjusted according to the 3D scene model after stylization, so as to obtain the stylized target 3D scene model.
[0069] In an embodiment of the present application, a style image and multiple images to be processed of a three-dimensional scene model to be processed are obtained, each image to be processed is an image of the three-dimensional scene model to be processed at a different perspective, the style image and the multiple images to be processed are respectively encoded and processed to obtain style features corresponding to the style image and content features corresponding to the multiple images to be processed, the multiple images to be processed are grouped according to perspective to obtain multiple adjacent perspective image groups, each adjacent perspective image includes multiple images to be processed with adjacent perspectives, each adjacent perspective image group, content features and style features are input into a diffusion model for stylization processing to obtain stylized images corresponding to each image to be processed, and parameters of the three-dimensional scene model to be processed are adjusted according to each stylized image to obtain a stylized target three-dimensional scene model.
[0070] By grouping the processed images according to adjacent viewpoints and fusing their features, the extracted content features can more comprehensively represent the contextual information of adjacent viewpoints, thus ensuring the consistency of local features. By extracting features from the style image and the processed image separately, the style and content feature extraction processes are decoupled, thereby resolving the information leakage problem during the stylization process.
[0071] The following is a further description of obtaining multiple images to be processed of a 3D scene model to be processed. Figure 3 As shown, the above step S201 includes:
[0072] S301: Obtain parameter information of each Gaussian basis element in the 3D scene model to be processed.
[0073] Optionally, the 3D scene model includes multiple Gaussian primitives, each of which describes a point or region in the scene using a 3D Gaussian distribution. Each Gaussian primitive is defined by a set of parameters that determine the position, shape, orientation, and appearance characteristics of the Gaussian primitive in 3D space.
[0074] Optionally, the parameter information of the Gaussian basis element includes: position, covariance matrix, color, and transparency. The position can be the position coordinates of the center point of the Gaussian basis element in three-dimensional space. A single Gaussian basis element G(x) can be expressed as the following formula (1).
[0075]
[0076] Among them, x represents a point in space, μ represents the position information of the Gaussian basis element, and the covariance matrix Σ is used to describe the diffusion degree and direction of the Gaussian basis element in three dimensions, which determines the shape and direction of the Gaussian basis element.
[0077] S302 : Perform projection processing based on parameter information of each Gaussian basis to obtain a plurality of images to be processed.
[0078] During the projection process, the parameter information of each Gaussian primitive is used to calculate its projection on the 2D plane. Specifically, the position and covariance matrix of the Gaussian primitive are used to determine its center position and spread on the 2D plane, while the color and transparency information are used to determine the color and transparency of the projection on the image.
[0079] As a possible implementation, the 3D Gaussian primitives can be projected onto the 2D image plane using the 3DGS rendering process. The projection process involves transforming the covariance matrix to accommodate changes in viewing angle. The covariance matrix is converted to a new covariance matrix using the following equation (2):
[0080] Σ′=JWEW T JT (2)
[0081] Where W is the view transformation matrix and J is the Jacobian matrix of the affine approximation of the projection transformation. To support differentiable optimization, the covariance matrix Σ is decomposed into a scaling matrix S and a rotation matrix R as shown in Equation (3).
[0082] Σ=RSS T R T (3)
[0083] Through projection processing, multiple two-dimensional images to be processed can be generated from different perspectives. Each image to be processed corresponds to the result of observing a three-dimensional scene from a single perspective.
[0084] The following is a further description of encoding the style image and the multiple images to be processed to obtain the style features corresponding to the style image and the content features corresponding to the multiple images to be processed. Figure 4 As shown, the above step S202 includes:
[0085] S401: Input the style image into the style extraction model to obtain style features.
[0086] The style extraction model is used to extract and transform style features from the style image. For example, the style extraction model can be a style projection module in the CSGO (Cross-Scale Grid Attention) architecture.
[0087] The style projection module includes a graphics encoder. The style image is first encoded by the graphics encoder, and then the encoded image is input into the style projection module for feature extraction to obtain style features.
[0088] S402: Input a plurality of images to be processed and style features into a content extraction model to obtain content features corresponding to each image to be processed.
[0089] The content extraction model extracts and transforms content features from the image to be processed, extracting high-level semantic features from the image through a multi-layer neural network. The content extraction model can also process the image in blocks and extract and process features within each block based on style features, thereby better controlling local features and avoiding problems such as loss of detail and structural distortion during the stylization process.
[0090] Exemplarily, the content extraction model can be composed of the content projection module and Tile ControlNet of the CSGO architecture.
[0091] The following is a further explanation of the above-mentioned input of multiple images to be processed and style features into the content extraction model to obtain the content features corresponding to each image to be processed, such as Figure 5 Shown, including:
[0092] S501 : Encode the image to be processed to obtain an image coding sequence.
[0093] Optionally, the content extraction model may include a graphics encoder, and the graphics encoder may be used to encode the image to be processed to obtain an image encoding sequence.
[0094] An image coding sequence is a series of numerical values or vectors obtained after the encoding process. These values or vectors represent the main features and information of the image. The image coding sequence can be regarded as the internal representation of the image in the coding model, which retains the key information of the image.
[0095] S502: Extract features from the image coding sequence to obtain a first content feature.
[0096] Among them, the first content feature can be a high-level feature extracted from the image coding sequence. The high-level feature is related to semantic information such as the shape, position, category, etc. of the object in the image, and contains basic semantic information in the image to be processed.
[0097] S503: Obtain content features corresponding to the image to be processed according to the first content feature, the image coding sequence, and the style feature.
[0098] It should be noted that the style information of the content image may be leaked during the content feature extraction process, causing the style of the generated image to deviate from the target style. Taking the extraction of local content features from the processed image as an example, by using style features as conditions in the content feature extraction process, the model can suppress the content image style when processing local features in the processed image and ensure that the local features are consistent with the target style, avoiding style inconsistencies or leakage of the content image style during local processing.
[0099] Furthermore, the above process of obtaining the content feature corresponding to the image to be processed according to the first content feature, the image coding sequence and the style feature is as follows: Figure 6 Shown, including:
[0100] S601: Using the style feature as a conditional feature, perform feature extraction on the image coding sequence to obtain a second content feature.
[0101] Optionally, the style feature can be used as a conditional feature and input into the feature extraction model together with the image coding sequence. When extracting the image content feature, the content extraction model will consider the influence of the style feature to obtain a second content feature.
[0102] The second content features are extracted from the image coding sequence while taking into account the style features. These features not only contain the semantic information of the image, but also incorporate the influence of the style features, making the generated image consistent with the original image in content and consistent with the style image in style.
[0103] S602: Concatenate the first content feature and the second content feature to obtain a content feature corresponding to the image to be processed.
[0104] Alternatively, the first content feature and the second content feature can be concatenated based on the following formula (4) to obtain a content feature. In this case, the content feature not only includes the semantic structure information of the image to be processed, but also combines the influence of the style feature, thereby providing a more comprehensive feature representation for the subsequent stylization process.
[0105]
[0106] in, The image to be processed I c The encoded sequence at time step t, represents the first content feature, Represents the second content feature, Indicates content characteristics, D s represents the style feature, and ε represents the graph encoder.
[0107] In an embodiment of the present application, by using style features as input conditions in the content feature extraction process, the model can suppress the content image style when processing local features in the image to be processed, and ensure that the local features are consistent with the target style, thereby avoiding style inconsistency or leakage of content image style during local processing.
[0108] Figure 7 This is a schematic diagram of the overall process of stylization in this application, refer to Figure 7 ,The diffusion model includes : content control module and style control module.
[0109] Among them, the content extraction module extracts content features from the image to be processed, and the style extraction module extracts style features from the style image. The content features are used as input conditions for the content control module, and the style features are used as input conditions for the style control module. After the adjacent perspective image groups of the image to be processed are input into the diffusion model, the diffusion model can perform stylized processing to obtain a stylized image.
[0110] The above process of inputting each adjacent view image group, content feature and style feature into the diffusion model for stylization processing to obtain the stylized image corresponding to each image to be processed is as follows: Figure 8 Shown, including:
[0111] S801: Input the adjacent view image group and content features into the content control module, and the content control module performs feature fusion processing to obtain image content features of each to-be-processed image in the adjacent view image group.
[0112] For each image to be processed in the input adjacent view image group, the content control module can perform feature extraction and combine it with other images to be processed in the adjacent view image group to obtain the context information of the image to be processed under multiple view angles, and fuse the context information and content features of the image to be processed to obtain a more comprehensive feature representation.
[0113] The image content features include the semantic features of the image to be processed and incorporate contextual information from adjacent viewpoint image groups.
[0114] S802: Input the image content features and style features into the style control module, and the style control module performs feature fusion processing to obtain stylized images corresponding to each to-be-processed image in the adjacent view image group.
[0115] Optionally, the style control module can perform feature fusion processing on the input image content features and style features, thereby combining the image content features and style features to generate a feature representation that retains the original content while also having the target style.
[0116] After feature fusion processing, the style control module can generate stylized images corresponding to each image to be processed in the adjacent view image group. The stylized images are consistent with the image to be processed in content and consistent with the style image in visual style.
[0117] Optionally, refer to Figure 7 ,The content control module includes : adjacent view attention layer and cross attention layer.
[0118] like Figure 9 As shown, the process of the above-mentioned content control module performing feature fusion processing to obtain image content features includes:
[0119] S901: Acquire a current image in an adjacent-view image group, where the current image is any image to be processed in the adjacent-view image group.
[0120] It should be noted that after the adjacent view image group is input into the diffusion model, the content control module may sequentially perform feature fusion processing on each to-be-processed image in the adjacent view image group to obtain image content features of each to-be-processed image.
[0121] S902. The adjacent view attention layer performs feature extraction on the current image and each other image in the adjacent view image group to obtain initial image features corresponding to the current image.
[0122] The adjacent view attention layer is used to perform attention calculation and feature extraction on the current image and other images in the adjacent view image group to obtain initial image features. This allows the image features of other views in the adjacent view image group to be integrated into the current image, giving the current image contextual information from other views and ensuring consistency in local information between the current image and images from other views.
[0123] S903: The cross-attention layer performs feature fusion processing on the content features and the initial image features to obtain the image content features corresponding to the current image.
[0124] Optionally, the cross-attention layer can perform feature fusion processing on the content features and the initial image features to obtain a richer content feature representation.
[0125] Further, such as Figure 10 As shown, the adjacent view attention layer extracts features from the current image and each other image in the adjacent view image group to obtain the initial image features corresponding to the current image, including the following process:
[0126] S1001. Obtain a query vector of a current image and a key vector matrix and a value vector matrix of an adjacent view image group.
[0127] Among them, the key vector matrix of the adjacent view image group can be defined as the following formula (5), and the value vector matrix can be defined as the following formula (6).
[0128] K NV =[K1,K2,...,K N ] T (5)
[0129] V NV =[V1,V2,...,V N ] T (6)
[0130] Where N is the number of neighborhood views in the adjacent view image group.
[0131] S1002. Perform attention calculation on the query vector, key vector matrix, and value vector matrix to obtain a target attention weight matrix.
[0132] Optionally, the attention calculation method can be shown in the following formula (7). represents the attention weight matrix, Q i Represents the query vector.
[0133]
[0134] S1003. Extract features of the current image according to the target attention weight matrix to obtain initial image features.
[0135] Optionally, a matrix multiplication operation can be performed on the attention matrix and the value vector matrix of the current image to obtain a weighted feature matrix. This weighted feature matrix is then normalized and aggregated to obtain the initial image features. The initial image features fuse the image features of multiple adjacent viewpoints, thus ensuring the consistency of local information across multiple adjacent viewpoints.
[0136] Continue to refer to Figure 7 ,The style control module includes: adjacent view attention layer and cross attention layer.
[0137] The style control module performs feature fusion processing to obtain the stylized image corresponding to each image to be processed in the adjacent view image group, as shown in the following example: Figure 11 Shown, including:
[0138] S1101. The adjacent view attention layer performs feature extraction based on the initial image features of the current image and the initial image features of each image in the adjacent view image group to obtain intermediate image features.
[0139] Optionally, the style control module may extract features from the current initial image features and the initial image features of other images in the adjacent view image group to which the initial image features belong, and further fuse the features of the adjacent views to obtain intermediate image features.
[0140] S1102: The cross-attention layer performs feature fusion processing on the style features and the intermediate image features to obtain a stylized image.
[0141] Optionally, the cross-attention layer can take the style features as conditions to perform feature fusion processing on the style features and the intermediate image features, and the resulting stylized image has the same style as the style image.
[0142] It is worth noting that, referring to Figure 7 The content control module and style control module each include multiple sequentially connected adjacent view attention layers and cross attention layers. The processing steps of each layer in the content control module are the same as steps S902-S903 above, and the processing steps of each layer in the style control module are the same as steps S1101-S1102 above. This application will not repeat them here.
[0143] Alternatively, the process of stylized image prediction using the diffusion model can be expressed as the following equations (8) and (9).
[0144]
[0145] Among them, e t represents the predicted noise, α t Represents the scheduling coefficient in the DDIM scheduler, and the new features integrated with adjacent perspectives are expressed as Then the above equations (8) and (9) can be further expressed as the following equation (10).
[0146]
[0147] After obtaining the stylized images, the parameters of the 3D scene model to be processed can be adjusted according to each stylized image to obtain the stylized target 3D scene model, such as Figure 12 As shown, the above step S204 includes:
[0148] S1201: Extract features from the stylized image to obtain style image features.
[0149] The style image features describe the visual style information of the stylized image, such as color, texture, and lines. For example, the style image features can be obtained by extracting features using a pre-trained convolutional neural network.
[0150] S1202 : Perform stylization processing on the 3D scene model to be processed according to the style image features to obtain an initial 3D scene model.
[0151] As one possible implementation, the 3D scene model to be processed is represented as a set of Gaussian primitives, each defined by parameters such as position, covariance matrix, color, and transparency. By mapping the parametric information of the style image features to the parameters of the model, the 3D scene model to be processed can be stylized.
[0152] For example, the color information in the style features can be mapped to the color parameters of the model to achieve stylized processing of the color; the texture information can be mapped to the texture parameters of the model to achieve stylized processing of the texture; according to the shape information in the style image features, the covariance matrix of the model is adjusted to achieve stylized processing of the shape; the transparency parameters of the style image features are mapped to the transparency parameters of the model to achieve transparency adjustment of the model.
[0153] S1203: Determine a target loss value based on the initial 3D scene model and the stylized image.
[0154] As a possible implementation method, the initial three-dimensional scene model can be projected to obtain multiple projected images, and the loss value between the projected image and the stylized image can be calculated to obtain the target loss value.
[0155] S1204: Determine whether the target loss value meets a preset condition; if not, update the Gaussian parameters of the 3D scene model to be processed based on the target loss value to obtain a new 3D scene model to be processed.
[0156] Repeat the above steps S1202 to S1204 until the target loss value meets the preset conditions, and use the current round of to-be-processed 3D scene model as the target 3D scene model.
[0157] The target loss value satisfies a preset condition, which can be either convergence or the number of iterations reaching a preset number. If the loss value does not meet the preset condition, the Gaussian parameters of the 3D scene model to be processed can be updated based on the target loss value, and then stylization processing can be performed again based on the style image features until the target loss value meets the preset condition.
[0158] The following is a further explanation of the above-mentioned determination of the target loss value based on the initial 3D scene model and the stylized image. Figure 13 Shown, including:
[0159] S1301: Determine a first loss value based on color information of each pixel in the initial three-dimensional scene model and color information of each pixel in the stylized image.
[0160] As a possible implementation method, the initial three-dimensional scene model may be rendered to obtain a corresponding plurality of rendered images, each rendered image having the same viewing angle as the stylized image.
[0161] Optionally, the first loss value is used to represent the difference between the predicted value and the true value of each pixel in the initial 3D scene model on the RGB (Red, Green, Blue) three color channels. The true value can be the RGB value of the pixel in the stylized image. As an example, the first loss value can be based on the loss function L of the splatfacto model. splatfacto Sure.
[0162] S1302: Determine a second loss value based on the feature vector of the initial three-dimensional scene model at each position and the feature vector of the stylized image at each position.
[0163] Optionally, the feature vector at a position may represent the style features of the image at that position, including color, texture, lines, etc. By calculating the loss value for the feature vector at the same position, a second loss value may be obtained based on the loss values at each position.
[0164] Among them, the second loss value L NNFM The calculation method of can be shown in the following formula (11). M is the number of pixels of the rendered image corresponding to the initial 3D scene model, F r (i,j),FS (i, j) represents the feature vector of the rendered image and the stylized image at position (i, j), respectively.
[0165]
[0166] S1303: Determine a target loss value according to the first loss value and the second loss value.
[0167] Optionally, the sum of the first loss value and the second loss value may be used as the target loss value, or the first loss value and the second loss value may be weighted to obtain the target loss value.
[0168] Target loss value L fine It can be expressed as the following formula (12).
[0169] L fine =L NNFM +L splatfacto
[0170] By minimizing the above loss function and adjusting the 3D Gaussian scene parameters, the rendered scene image is closer to the stylized image in content and style, and finally a stylized 3D Gaussian scene is obtained.
[0171] Based on the same inventive concept, the embodiments of the present application also provide a stylized processing device for a three-dimensional scene corresponding to the stylized processing method for a three-dimensional scene. Since the principle of solving the problem by the device in the embodiments of the present application is similar to the stylized processing method for a three-dimensional scene in the embodiments of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0172] Figure 14 A schematic structural diagram of a stylized processing of a three-dimensional scene provided in an embodiment of the present application is shown. The apparatus includes an acquisition module 1401 , an encoding module 1402 , a grouping module 1403 and a processing module 1404 .
[0173] An acquisition module 1401 is configured to acquire a style image and a plurality of to-be-processed images of a to-be-processed 3D scene model, wherein each to-be-processed image is an image of the to-be-processed 3D scene model at a different perspective;
[0174] An encoding module 1402 is configured to encode the style image and the plurality of images to be processed, respectively, to obtain style features corresponding to the style image and content features corresponding to the plurality of images to be processed;
[0175] A grouping module 1403 is configured to group the plurality of images to be processed according to viewing angles to obtain a plurality of adjacent viewing angle image groups, each of which includes a plurality of images to be processed having adjacent viewing angles;
[0176] Processing module 1404 is used to input each adjacent perspective image group, content features and style features into the diffusion model for stylization processing to obtain stylized images corresponding to each image to be processed, and adjust the parameters of the 3D scene model to be processed based on each stylized image to obtain a stylized target 3D scene model.
[0177] In a feasible implementation, the encoding module 1402 is specifically configured to:
[0178] Input the style image into the style extraction model to obtain style features;
[0179] Multiple images to be processed and style features are input into the content extraction model to obtain the content features corresponding to each image to be processed.
[0180] In a feasible implementation, the encoding module 1402 is specifically configured to:
[0181] Perform encoding processing on the image to be processed to obtain an image coding sequence;
[0182] Performing feature extraction on the image coding sequence to obtain a first content feature;
[0183] A content feature corresponding to the image to be processed is obtained according to the first content feature, the image coding sequence and the style feature.
[0184] In a feasible implementation, the encoding module 1402 is specifically configured to:
[0185] Using style features as conditional features, perform feature extraction on the image coding sequence to obtain the second content feature;
[0186] The first content feature and the second content feature are concatenated to obtain a content feature corresponding to the image to be processed.
[0187] In a feasible embodiment, the diffusion model includes: a content control module and a style control module;
[0188] The processing module 1404 is specifically configured to:
[0189] The adjacent view image group and content features are input into the content control module, which performs feature fusion processing to obtain image content features of each image to be processed in the adjacent view image group;
[0190] The image content features and style features are input into the style control module, which performs feature fusion processing to obtain stylized images corresponding to each image to be processed in the adjacent perspective image group.
[0191] In one feasible embodiment, the content control module includes: an adjacent view attention layer and a cross attention layer;
[0192] The processing module 1404 is specifically configured to:
[0193] Acquire a current image in the adjacent view image group, where the current image is any image to be processed in the adjacent view image group;
[0194] The adjacent view attention layer extracts features from the current image and other images in the adjacent view image group to obtain the initial image features corresponding to the current image;
[0195] The cross-attention layer performs feature fusion processing on the content features and the initial image features to obtain the image content features corresponding to the current image.
[0196] In one feasible embodiment, the processing module 1404 is specifically configured to:
[0197] Get the query vector of the current image and the key vector matrix and value vector matrix of the adjacent view image group;
[0198] Perform attention calculation on the query vector, key vector matrix, and value vector matrix to obtain the target attention weight matrix;
[0199] The current image is feature extracted according to the target attention weight matrix to obtain the initial image features.
[0200] In one feasible embodiment, the style control module includes: a neighboring view attention layer and a cross attention layer;
[0201] The processing module 1404 is specifically configured to:
[0202] The adjacent view attention layer extracts features based on the initial image features of the current image and the initial image features of each image in the adjacent view image group to obtain intermediate image features;
[0203] The cross-attention layer performs feature fusion processing on the style features and the intermediate image features to obtain a stylized image.
[0204] In one feasible embodiment, the processing module 1404 is specifically configured to:
[0205] A. Extract features from the stylized image to obtain style image features;
[0206] B. Stylize the 3D scene model to be processed according to the style image features to obtain an initial 3D scene model;
[0207] C. Determine the target loss value based on the initial 3D scene model and the stylized image;
[0208] D. Determine whether the target loss value meets the preset conditions. If not, update the Gaussian parameters of the 3D scene model to be processed based on the target loss value to obtain a new 3D scene model to be processed;
[0209] Repeat steps B to D above until the target loss value meets the preset conditions, and use the current round of processed 3D scene model as the target 3D scene model.
[0210] In one feasible embodiment, the processing module 1404 is specifically configured to:
[0211] determining a first loss value based on color information of each pixel in the initial three-dimensional scene model and color information of each pixel in the stylized image;
[0212] Determining a second loss value based on the feature vector of the initial three-dimensional scene model at each position and the feature vector of the stylized image at each position;
[0213] A target loss value is determined according to the first loss value and the second loss value.
[0214] In one feasible embodiment, the processing module 1404 is specifically configured to:
[0215] Obtain parameter information of each Gaussian basis element in the three-dimensional scene model to be processed;
[0216] Projection processing is performed based on the parameter information of each Gaussian basis element to obtain multiple images to be processed.
[0217] In this embodiment, by grouping the images to be processed according to adjacent viewpoints and fusing their features, the extracted content features can more comprehensively represent the contextual information of adjacent viewpoints, thereby ensuring the consistency of local features. By extracting features from the style image and the image to be processed separately, the style feature and content feature extraction processes are decoupled, thereby resolving the information leakage problem during the stylization process.
[0218] Figure 15 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application is shown, including: a processor 1501, a storage medium 1502, and a bus 1503. The storage medium 1502 stores machine-readable instructions executable by the processor 1501. When the electronic device executes a stylization processing method for a three-dimensional scene in the embodiment, the processor 1501 communicates with the storage medium 1502 via the bus 1503. The processor 1501 executes the machine-readable instructions. The processor 1501 performs the preamble of the method item to perform the following steps:
[0219] Acquire multiple to-be-processed images of the to-be-processed three-dimensional scene model, each to-be-processed image being an image of a perspective of the to-be-processed three-dimensional scene model;
[0220] Encoding the style image and the multiple images to be processed respectively to obtain style features corresponding to the style image and content features corresponding to the multiple images to be processed;
[0221] Grouping the multiple images to be processed according to viewing angles to obtain multiple adjacent viewing angle image groups, each adjacent viewing angle image group including multiple images to be processed with adjacent viewing angles;
[0222] Each adjacent view image group, content feature, and style feature are input into the diffusion model for stylization processing to obtain a stylized image corresponding to each image to be processed. The parameters of the 3D scene model to be processed are adjusted according to each stylized image to obtain a stylized target 3D scene model.
[0223] In a feasible implementation, when the processor 1501 performs encoding processing on the style image and the multiple images to be processed respectively to obtain style features corresponding to the style image and content features corresponding to the multiple images to be processed, it is specifically configured to:
[0224] Input the style image into the style extraction model to obtain style features;
[0225] Multiple images to be processed and style features are input into the content extraction model to obtain the content features corresponding to each image to be processed.
[0226] In a feasible implementation, when the processor 1501 inputs a plurality of images to be processed and style features into a content extraction model to obtain content features corresponding to each image to be processed, the processor 1501 is specifically configured to:
[0227] Perform encoding processing on the image to be processed to obtain an image coding sequence;
[0228] Performing feature extraction on the image coding sequence to obtain a first content feature;
[0229] A content feature corresponding to the image to be processed is obtained according to the first content feature, the image coding sequence and the style feature.
[0230] In a feasible implementation manner, when the processor 1501 obtains the content feature corresponding to the image to be processed according to the first content feature, the image coding sequence, and the style feature, it is specifically configured to:
[0231] Using style features as conditional features, perform feature extraction on the image coding sequence to obtain the second content feature;
[0232] The first content feature and the second content feature are concatenated to obtain a content feature corresponding to the image to be processed.
[0233] In a feasible embodiment, the diffusion model includes: a content control module and a style control module;
[0234] When the processor 1501 inputs each adjacent view image group, content feature, and style feature into the diffusion model for stylization processing to obtain a stylized image corresponding to each image to be processed, it is specifically configured to:
[0235] The adjacent view image group and content features are input into the content control module, which performs feature fusion processing to obtain image content features of each image to be processed in the adjacent view image group;
[0236] The image content features and style features are input into the style control module, which performs feature fusion processing to obtain stylized images corresponding to each image to be processed in the adjacent perspective image group.
[0237] In one feasible embodiment, the content control module includes: an adjacent view attention layer and a cross attention layer;
[0238] When the processor 1501 executes the content control module to perform feature fusion processing to obtain image content features, it is specifically used to:
[0239] Acquire a current image in the adjacent view image group, where the current image is any image to be processed in the adjacent view image group;
[0240] The adjacent view attention layer extracts features from the current image and other images in the adjacent view image group to obtain the initial image features corresponding to the current image;
[0241] The cross-attention layer performs feature fusion processing on the content features and the initial image features to obtain the image content features corresponding to the current image.
[0242] In a feasible implementation, the processor 1501, when executing the adjacent view attention layer to extract features from the current image and each other image in the adjacent view image group to obtain the initial image features corresponding to the current image, is specifically configured to:
[0243] Get the query vector of the current image and the key vector matrix and value vector matrix of the adjacent view image group;
[0244] Perform attention calculation on the query vector, key vector matrix, and value vector matrix to obtain the target attention weight matrix;
[0245] The current image is feature extracted according to the target attention weight matrix to obtain the initial image features.
[0246] In one feasible embodiment, the style control module includes: a neighboring view attention layer and a cross attention layer;
[0247] When the processor 1501 executes the feature fusion processing by the style control module to obtain the stylized image corresponding to each to-be-processed image in the adjacent view image group, it is specifically configured to:
[0248] The adjacent view attention layer extracts features based on the initial image features of the current image and the initial image features of each image in the adjacent view image group to obtain intermediate image features;
[0249] The cross-attention layer performs feature fusion processing on the style features and the intermediate image features to obtain a stylized image.
[0250] In a feasible implementation, when the processor 1501 adjusts parameters of the 3D scene model to be processed according to each stylized image to obtain a stylized target 3D scene model, the processor 1501 is specifically configured to:
[0251] A. Extract features from the stylized image to obtain style image features;
[0252] B. Stylize the 3D scene model to be processed according to the style image features to obtain an initial 3D scene model;
[0253] C. Determine the target loss value based on the initial 3D scene model and the stylized image;
[0254] D. Determine whether the target loss value meets the preset conditions. If not, update the Gaussian parameters of the 3D scene model to be processed based on the target loss value to obtain a new 3D scene model to be processed;
[0255] Repeat steps B to D above until the target loss value meets the preset conditions, and use the current round of processed 3D scene model as the target 3D scene model.
[0256] In a feasible implementation, when determining the target loss value based on the initial 3D scene model and the stylized image, the processor 1501 is specifically configured to:
[0257] determining a first loss value based on color information of each pixel in the initial three-dimensional scene model and color information of each pixel in the stylized image;
[0258] Determining a second loss value based on the feature vector of the initial three-dimensional scene model at each position and the feature vector of the stylized image at each position;
[0259] A target loss value is determined according to the first loss value and the second loss value.
[0260] In a feasible implementation manner, when executing the process of acquiring a plurality of images to be processed of a three-dimensional scene model to be processed, the processor 1501 is specifically configured to:
[0261] Obtain parameter information of each Gaussian basis element in the three-dimensional scene model to be processed;
[0262] Projection processing is performed based on the parameter information of each Gaussian basis element to obtain multiple images to be processed.
[0263] In this embodiment, by grouping the images to be processed according to adjacent viewpoints and fusing their features, the extracted content features can more comprehensively represent the contextual information of adjacent viewpoints, thereby ensuring the consistency of local features. By extracting features from the style image and the image to be processed separately, the style feature and content feature extraction processes are decoupled, thereby resolving the information leakage problem during the stylization process.
[0264] An embodiment of the present application further provides a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium. The computer program is executed when a processor is run, and the processor performs the following steps:
[0265] Acquire multiple to-be-processed images of the to-be-processed three-dimensional scene model, each to-be-processed image being an image of a perspective of the to-be-processed three-dimensional scene model;
[0266] Encoding the style image and the multiple images to be processed respectively to obtain style features corresponding to the style image and content features corresponding to the multiple images to be processed;
[0267] Grouping the multiple images to be processed according to viewing angles to obtain multiple adjacent viewing angle image groups, each adjacent viewing angle image group including multiple images to be processed with adjacent viewing angles;
[0268] Each adjacent view image group, content feature, and style feature are input into the diffusion model for stylization processing to obtain a stylized image corresponding to each image to be processed. The parameters of the 3D scene model to be processed are adjusted according to each stylized image to obtain a stylized target 3D scene model.
[0269] In a feasible embodiment, when the processor performs encoding processing on the style image and the multiple images to be processed respectively to obtain style features corresponding to the style image and content features corresponding to the multiple images to be processed, the processor is specifically configured to:
[0270] Input the style image into the style extraction model to obtain style features;
[0271] Multiple images to be processed and style features are input into the content extraction model to obtain the content features corresponding to each image to be processed.
[0272] In a feasible embodiment, when the processor inputs a plurality of images to be processed and style features into a content extraction model to obtain content features corresponding to each image to be processed, the processor is specifically configured to:
[0273] Perform encoding processing on the image to be processed to obtain an image coding sequence;
[0274] Performing feature extraction on the image coding sequence to obtain a first content feature;
[0275] A content feature corresponding to the image to be processed is obtained according to the first content feature, the image coding sequence and the style feature.
[0276] In a feasible implementation manner, when the processor obtains the content feature corresponding to the image to be processed based on the first content feature, the image coding sequence, and the style feature, it is specifically configured to:
[0277] Using style features as conditional features, perform feature extraction on the image coding sequence to obtain the second content feature;
[0278] The first content feature and the second content feature are concatenated to obtain a content feature corresponding to the image to be processed.
[0279] In a feasible embodiment, the diffusion model includes: a content control module and a style control module;
[0280] When the processor inputs each adjacent view image group, content feature, and style feature into the diffusion model for stylization processing to obtain a stylized image corresponding to each image to be processed, the processor is specifically configured to:
[0281] The adjacent view image group and content features are input into the content control module, which performs feature fusion processing to obtain image content features of each image to be processed in the adjacent view image group;
[0282] The image content features and style features are input into the style control module, which performs feature fusion processing to obtain stylized images corresponding to each image to be processed in the adjacent perspective image group.
[0283] In one feasible embodiment, the content control module includes: an adjacent view attention layer and a cross attention layer;
[0284] When the processor executes the content control module to perform feature fusion processing to obtain image content features, it is specifically used to:
[0285] Acquire a current image in the adjacent view image group, where the current image is any image to be processed in the adjacent view image group;
[0286] The adjacent view attention layer extracts features from the current image and other images in the adjacent view image group to obtain the initial image features corresponding to the current image;
[0287] The cross-attention layer performs feature fusion processing on the content features and the initial image features to obtain the image content features corresponding to the current image.
[0288] In one feasible embodiment, when the processor executes the adjacent view attention layer to extract features from the current image and each other image in the adjacent view image group to obtain the initial image features corresponding to the current image, it is specifically configured to:
[0289] Get the query vector of the current image and the key vector matrix and value vector matrix of the adjacent view image group;
[0290] Perform attention calculation on the query vector, key vector matrix, and value vector matrix to obtain the target attention weight matrix;
[0291] The current image is feature extracted according to the target attention weight matrix to obtain the initial image features.
[0292] In one feasible embodiment, the style control module includes: a neighboring view attention layer and a cross attention layer;
[0293] When the processor executes the style control module to perform feature fusion processing to obtain a stylized image corresponding to each to-be-processed image in the adjacent view image group, the processor is specifically configured to:
[0294] The adjacent view attention layer extracts features based on the initial image features of the current image and the initial image features of each image in the adjacent view image group to obtain intermediate image features;
[0295] The cross-attention layer performs feature fusion processing on the style features and the intermediate image features to obtain a stylized image.
[0296] In a feasible embodiment, when the processor adjusts parameters of the 3D scene model to be processed according to each stylized image to obtain a stylized target 3D scene model, the processor is specifically configured to:
[0297] A. Extract features from the stylized image to obtain style image features;
[0298] B. Stylize the 3D scene model to be processed according to the style image features to obtain an initial 3D scene model;
[0299] C. Determine the target loss value based on the initial 3D scene model and the stylized image;
[0300] D. Determine whether the target loss value meets the preset conditions. If not, update the Gaussian parameters of the 3D scene model to be processed based on the target loss value to obtain a new 3D scene model to be processed;
[0301] Repeat steps B to D above until the target loss value meets the preset conditions, and use the current round of processed 3D scene model as the target 3D scene model.
[0302] In one feasible embodiment, when determining the target loss value based on the initial 3D scene model and the stylized image, the processor is specifically configured to:
[0303] determining a first loss value based on color information of each pixel in the initial three-dimensional scene model and color information of each pixel in the stylized image;
[0304] Determining a second loss value based on the feature vector of the initial three-dimensional scene model at each position and the feature vector of the stylized image at each position;
[0305] A target loss value is determined according to the first loss value and the second loss value.
[0306] In a feasible implementation manner, when executing the process of acquiring a plurality of images to be processed of a three-dimensional scene model to be processed, the processor is specifically configured to:
[0307] Obtain parameter information of each Gaussian basis element in the three-dimensional scene model to be processed;
[0308] Projection processing is performed based on the parameter information of each Gaussian basis element to obtain multiple images to be processed.
[0309] In this embodiment, by grouping the images to be processed according to adjacent viewpoints and fusing their features, the extracted content features can more comprehensively represent the contextual information of adjacent viewpoints, thereby ensuring the consistency of local features. By extracting features from the style image and the image to be processed separately, the style feature and content feature extraction processes are decoupled, thereby resolving the information leakage problem during the stylization process.
[0310] In the embodiment of the present application, the computer program can also execute other machine-readable instructions when run by the processor to execute other methods described in the embodiment. For the specific execution method steps and principles, please refer to the description of the embodiment and will not be repeated here.
[0311] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0312] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0313] In addition, each functional unit in the embodiments provided in the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0314] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0315] It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are only used to distinguish the description and are not to be understood as indicating or implying relative importance.
[0316] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The scope of protection of the present application is not limited thereto. Although the present application has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed in the present application, or make equivalent replacements for some of the technical features thereof. However, these modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present application. They should all be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for stylizing a three-dimensional scene, characterized in that: include: Acquire a style image and a plurality of to-be-processed images of the to-be-processed 3D scene model, wherein each of the to-be-processed images is an image of the to-be-processed 3D scene model at a different perspective; performing encoding processing on the style image and the plurality of images to be processed respectively to obtain style features corresponding to the style image and content features corresponding to the plurality of images to be processed; Grouping the plurality of images to be processed according to viewing angles to obtain a plurality of adjacent viewing angle image groups, each adjacent viewing angle image group including a plurality of images to be processed having adjacent viewing angles; Each of the adjacent view image groups, the content features, and the style features is input into a diffusion model for stylization processing to obtain a stylized image corresponding to each of the images to be processed, and parameters of the three-dimensional scene model to be processed are adjusted according to each of the stylized images to obtain a stylized target three-dimensional scene model.
2. The method according to claim 1, characterized in that The encoding process is performed on the style image and the plurality of images to be processed respectively to obtain style features corresponding to the style image and content features corresponding to the plurality of images to be processed, including: Inputting the style image into a style extraction model to obtain the style features; The plurality of images to be processed and the style features are input into a content extraction model to obtain content features corresponding to each of the images to be processed.
3. The method according to claim 2, characterized in that Inputting the plurality of images to be processed and the style features into a content extraction model to obtain content features corresponding to each of the images to be processed includes: Performing encoding processing on the image to be processed to obtain an image encoding sequence; Performing feature extraction on the image coding sequence to obtain a first content feature; A content feature corresponding to the image to be processed is obtained according to the first content feature, the image coding sequence and the style feature.
4. The method according to claim 3, characterized in that The obtaining, based on the first content feature, the image coding sequence, and the style feature, a content feature corresponding to the image to be processed includes: Using the style feature as a conditional feature, performing feature extraction on the image coding sequence to obtain a second content feature; The first content feature and the second content feature are concatenated to obtain a content feature corresponding to the image to be processed.
5. The method according to claim 1, wherein The diffusion model includes: a content control module and a style control module; The step of inputting each of the adjacent view image groups, the content features, and the style features into a diffusion model for stylization processing to obtain a stylized image corresponding to each of the to-be-processed images includes: Inputting the adjacent view image group and the content features into the content control module, and having the content control module perform feature fusion processing to obtain image content features of each to-be-processed image in the adjacent view image group; The image content features and the style features are input into the style control module, and the style control module performs feature fusion processing to obtain stylized images corresponding to each to-be-processed image in the adjacent view image group.
6. The method according to claim 5, characterized in that The content control module includes: an adjacent view attention layer and a cross attention layer; The process of the content control module performing feature fusion processing to obtain image content features includes: Acquire a current image in the adjacent-view image group, where the current image is any image to be processed in the adjacent-view image group; The adjacent view attention layer performs feature extraction on the current image and each other image in the adjacent view image group to obtain initial image features corresponding to the current image; The cross attention layer performs feature fusion processing on the content features and the initial image features to obtain image content features corresponding to the current image.
7. The method according to claim 6, characterized in that The process of extracting features of the current image and each other image in the adjacent view image group by the adjacent view attention layer to obtain initial image features corresponding to the current image includes: Obtaining a query vector of the current image and a key vector matrix and a value vector matrix of the adjacent view image group; Performing attention calculation on the query vector, the key vector matrix, and the value vector matrix to obtain a target attention weight matrix; Feature extraction is performed on the current image according to the target attention weight matrix to obtain the initial image features.
8. The method according to claim 5, characterized in that The style control module includes: a neighboring view attention layer and a cross attention layer; The process of the style control module performing feature fusion processing to obtain a stylized image corresponding to each to-be-processed image in the adjacent view image group includes: The adjacent view attention layer performs feature extraction based on the initial image features of the current image and the initial image features of each image in the adjacent view image group to obtain intermediate image features; The cross-attention layer performs feature fusion processing on the style features and the intermediate image features to obtain the stylized image.
9. The method according to claim 1, characterized in that The step of adjusting parameters of the to-be-processed three-dimensional scene model according to each of the stylized images to obtain a stylized target three-dimensional scene model includes: A. extracting features from the stylized image to obtain style image features; B. performing stylization processing on the to-be-processed 3D scene model according to the style image features to obtain an initial 3D scene model; C. determining a target loss value based on the initial 3D scene model and the stylized image; D. Determine whether the target loss value satisfies a preset condition; if not, update the Gaussian parameters of the to-be-processed 3D scene model based on the target loss value to obtain a new to-be-processed 3D scene model; Repeat steps B to D above until the target loss value meets the preset conditions, and use the current round of to-be-processed 3D scene model as the target 3D scene model.
10. The method according to claim 9, characterized in that The determining of a target loss value according to the initial three-dimensional scene model and the stylized image includes: determining a first loss value according to color information of each pixel in the initial three-dimensional scene model and color information of each pixel in the stylized image; determining a second loss value according to the feature vector of the initial three-dimensional scene model at each position and the feature vector of the stylized image at each position; The target loss value is determined according to the first loss value and the second loss value.
11. The method according to claim 1, characterized in that The step of obtaining a plurality of images to be processed of the three-dimensional scene model to be processed includes: Obtaining parameter information of each Gaussian basis element in the three-dimensional scene model to be processed; Projection processing is performed based on the parameter information of each Gaussian basis element to obtain the multiple images to be processed.
12. A stylized processing device for a three-dimensional scene, characterized in that: include: An acquisition module, configured to acquire a style image and a plurality of to-be-processed images of the to-be-processed 3D scene model, wherein each of the to-be-processed images is an image of the to-be-processed 3D scene model at a different perspective; An encoding module, configured to encode the style image and the plurality of images to be processed, respectively, to obtain style features corresponding to the style image and content features corresponding to the plurality of images to be processed; a grouping module, configured to group the plurality of images to be processed according to viewing angles to obtain a plurality of adjacent viewing angle image groups, each of the adjacent viewing angle images comprising a plurality of images to be processed having adjacent viewing angles; A processing module is configured to input each of the adjacent viewpoint image groups, the content features, and the style features into a diffusion model for stylization processing to obtain a stylized image corresponding to each of the images to be processed, and to adjust parameters of the 3D scene model to be processed based on each of the stylized images to obtain a stylized target 3D scene model.
13. An electronic device, characterized in that: include: A processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the storage medium communicate via the bus, and the processor executes the machine-readable instructions to perform the steps of the stylized processing method for a three-dimensional scene as described in any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for stylizing a three-dimensional scene as claimed in any one of claims 1 to 11 are executed.