Scene style migration method and device, equipment and storage medium

By constructing three-dimensional scenes and using technical means such as multi-dimensional probability density function, VGG encoder and decoder, the problem of reconstruction error in 3D style migration is solved, achieving a more efficient and consistent style migration effect.

CN120070159APending Publication Date: 2025-05-30XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510119129.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing 3D style transfer methods are prone to reconstruction errors during the explicit reconstruction process, which affects the style transfer effect.

Method used

By acquiring scene images and camera poses, the three-dimensional scene is constructed, the multi-dimensional probability density function is called to map the target image and scene images to the N-dimensional space, and projected back to the three-dimensional scene. The features are extracted using the pre-trained VGG encoder, the attention is calculated and the characteristics are fused, the three-dimensional scene is updated through the VGG decoder, and the color distribution is adjusted through the principal component analysis method.

Benefits of technology

It effectively reduces the reconstruction error during the style transfer process, improves the effect and consistency of style transfer, and ensures the retention of the geometric structure of the scene and the content information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070159A_ABST
    Figure CN120070159A_ABST
Patent Text Reader

Abstract

The invention provides a scene style migration method and device, equipment and a storage medium, and the method comprises the steps: firstly obtaining a scene image and a camera pose, and constructing a three-dimensional scene based on the scene image and the camera pose; then, obtaining a target image, calling a multi-dimensional probability density function to map the target image and the scene image to an N-dimensional space, and projecting the target image and the scene image back to the three-dimensional scene; then, calling a pre-trained VGG encoder to perform feature extraction on the target image and the scene image to generate a target style feature and a scene feature, and calculating attention between the target style feature and the scene feature; fusing the standard style features and the scene features based on the attention, and updating the three-dimensional scene through a pre-trained VGG decoder; and finally, processing the scene image and the target image at different visual angles through a principal component analysis method, and adjusting the color distribution of the updated three-dimensional scene. The problem that reconstruction errors occur in the style migration process is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine vision, and particularly to a method, device, equipment and storage medium for scene style transfer. Background Art

[0002] Scene style transfer is an important task in the field of computer vision and has broad application prospects in multiple application scenarios such as virtual reality, augmented reality, video post-production, and game design. The goal of this task is to effectively transfer the style features of the target image into the 3D scene while maintaining the geometric structure and content information of the scene. Existing style transfer methods can be divided into two categories: 2D style transfer and 3D style transfer.

[0003] In the field of 2D style transfer, the research has been relatively mature, mainly including style transfer methods based on optimization and generative network methods. For example, some methods perform style transfer by optimizing the content loss and style loss between the source image and the target image. However, these methods may not effectively retain the content structure of the source image during the transfer process. In addition, the generative network method directly generates stylized images. Although it simplifies the transfer process, there are still certain limitations in multi-scale feature extraction and retaining image content.

[0004] Compared with 2D style transfer, 3D style transfer faces more challenges. Existing 3D style transfer methods usually rely on the geometric reconstruction of the scene. However, reconstruction errors will inevitably occur during the explicit reconstruction process, thus affecting the final style transfer effect.

[0005] In view of this, the present application is proposed. Summary of the Invention

[0006] The present invention discloses a method, device, equipment and storage medium for scene style transfer, aiming to solve the problem of reconstruction errors occurring during the style transfer process.

[0007] The first embodiment of the present invention provides a method for scene style transfer, including: Obtain a scene image and a camera pose, and construct a three-dimensional scene based on the scene image and the camera pose; Obtain a target image, and call a multi-dimensional probability density function to map the target image and the scene image to the N-dimensional space and project them back to the three-dimensional scene; Call a pre-trained VGG encoder to extract features from the target image and the scene image to generate target style features and scene features, calculate the attention between the target style features and the scene features, fuse the target style features and the scene features based on the attention, and update the three-dimensional scene through a pre-trained VGG decoder; Process the scene images and target images from different perspectives through principal component analysis, and adjust the color distribution of the updated three-dimensional scene.

[0008] Preferably, the steps of obtaining the target image, calling the multi-dimensional probability density function to map the target image and the scene image to the N-dimensional space, and projecting them back to the three-dimensional scene are as follows: Through the projection matrix Map the target image and the scene image into the N-dimensional space, and its expression is: , where 、 Are the mapping coordinates of the scene image and the target image in the N-dimensional feature space respectively; Calculate 、 The marginal probability density functions of , and project them onto all axes To obtain the marginal probability density function ; among them, for each axis, an interpolation algorithm is used to calculate a one-dimensional transformation to match And , project the In the Dimension space And project it back to the three-dimensional color space through the rotation matrix.

[0009] Preferably, after obtaining the target image, calling the multi-dimensional probability density function to map the target image and the scene image to the N-dimensional space, and projecting them back to the three-dimensional scene, the following steps are further included: Maintain the gradient information of the scene image by optimizing the loss function, so that the color features of the optimized generated image are similar to those of the target image. Among them, the expression of the loss function is:

[0010] Where Is the weight factor of the gradient preservation term, Is the weight factor of the color similarity term, Is the gradient of the optimized image J at the pixel position (i, j), Is the gradient of the scene image I at the pixel position (i, j), Is the color value of the optimized image at the pixel position (i, j), Is the color value obtained by mapping the target image I at the pixel position (i, j) to the N-dimensional feature space and then mapping it back to the RGB space.

[0011] Preferably, the pre-trained VGG encoder is called to extract features from the target image and the scene image to generate target style features and scene features, and the attention between the target style features and the scene features is calculated. Based on the attention, the target style features and the scene features are fused and the three-dimensional scene is updated through the pre-trained VGG decoder, specifically as follows: Call the pre-trained VGG encoder to extract features from the target image and the scene image. Among them, the feature map of the scene image is represented as and the feature map of the target image is represented as , After normalizing the feature map of the scene image and the feature map of the target image, perform two convolutional operations to transform them into two new feature spaces f and g, and calculate the attention between the feature maps to generate two feature maps with different resolutions , and the expression for calculating the attention is: ; where represents the value after attention calculation between the feature of the content image at the th layer and the two feature maps of the target image, represents the value after mean-variance channel normalization of the feature map of the content image at the th layer, represents the value after mean-variance channel normalization of the feature map of the target image at the th layer, represents the value after the feature of the target image at the th layer passes through the convolutional transformation; Fuse the two feature maps with different resolutions to generate a fused feature. Call the pre-trained VGG decoder to synthesize the fused feature into a style image and update the three-dimensional scene. The operation expression of the fused feature is:

[0012] where is the fused feature, , are two 1*1 convolutional transformations on the attention features of the fourth and fifth layers.

[0013] Preferably, the scene images and target images from different perspectives are processed by the principal component analysis method, and the color distribution of the updated three-dimensional scene is adjusted, specifically as follows: Combine the image sets from different perspectives and represent them as , and represent the target image as , calculate the mean and variance of the image set and the target image respectively, and perform a singular value decomposition operation on the covariance matrix. The decomposition results are as follows:

[0014] After normalizing and scaling the decomposition results, calculate the final rotation and translation matrices to correct the color distribution of the scene:

[0015] The rotation matrix obtained through calculation and the translation vector are used to adjust the color distribution of the 3D scene, where and are diagonal matrices, is the color mean of the rendered image set, is the mean of the target image, is the color covariance of the rendered image set, is the covariance of the target image.

[0016] The second embodiment of the present invention provides a scene style transfer device, including: A 3D scene construction unit, configured to obtain a scene image and a camera pose, and construct a 3D scene based on the scene image and the camera pose; A mapping unit, configured to obtain a target image, and call a multi-dimensional probability density function to map the target image and the scene image to an N-dimensional space and project them back to the 3D scene; An update unit, configured to call a pre-trained VGG encoder to extract features from the target image and the scene image to generate a target style feature and a scene feature, calculate the attention between the target style feature and the scene feature, fuse the target style feature and the scene feature based on the attention, and update the 3D scene through a pre-trained VGG decoder; An adjustment unit, configured to process the scene image and the target image from different perspectives through the principal component analysis method, and adjust the color distribution of the updated 3D scene.

[0017] The third embodiment of the present invention provides a downhill assist device for an electric vehicle, including a memory and a processor. The memory stores a computer program that can be executed by the processor to implement a downhill assist method for an electric vehicle as described in any one of the above.

[0018] The fourth embodiment of the present invention provides a computer-readable storage medium storing a computer program, which can be executed by a processor of the device where the computer-readable storage medium is located to implement the downhill assist method for a pure electric vehicle as described in any one of the above.

[0019] The present invention discloses a method, device, equipment and storage medium for scene style migration. First, a scene image and a camera pose are obtained, and a three-dimensional scene is constructed based on the scene image and the camera pose. Then, a target image is obtained, and a multi-dimensional probability density function is called to map the target image and the scene image to an N-dimensional space and project them back to the three-dimensional scene. Next, a pre-trained VGG encoder is called to extract features from the target image and the scene image to generate target style features and scene features, and the attention between the target style features and the scene features is calculated. The target style features and the scene features are fused based on the attention and the three-dimensional scene is updated through a pre-trained VGG decoder. Finally, the scene images and target images from different perspectives are processed by principal component analysis, and the color distribution of the updated three-dimensional scene is adjusted. The problem of reconstruction error occurring during style migration is solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a flowchart of a method for scene style migration provided by the first embodiment of the present invention; Figure 2 is a schematic diagram of the overall framework of the model provided by the present invention; Figure 3 and 4 is a schematic diagram of data comparison provided by the present invention; Figure 5 is a visualization comparison diagram of scene style migration based on a radiation field provided by the present invention; Figure 6 is a module diagram of a device for scene style migration provided by the second embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0022] For a better understanding of the technical solutions of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0023] It should be clear that the described embodiments are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative work belong to the scope of protection of the present invention.

[0024] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0025] It should be understood that the term "and / or" used herein is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.

[0026] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detected (stated condition or event)" or "in response to detecting (stated condition or event)".

[0027] The "first / second" mentioned in the embodiments is only to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged in a specific order or sequence when permitted. It should be understood that the objects distinguished by "first / second" can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.

[0028] The following will give a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings.

[0029] The present invention discloses a method, device, equipment and storage medium for scene style migration, aiming to solve the problem of reconstruction error occurring during the process of style migration.

[0030] Please refer to Figure 1 , the first embodiment of the present invention provides a method for scene style migration, which can be executed by a scene style migration device (hereinafter referred to as the migration device), specifically, by one or more processors in the migration device, to at least achieve the following steps: S101. Obtain a scene image and a camera pose, and construct a three-dimensional scene based on the scene image and the camera pose; In this embodiment, the migration device may be a terminal with data processing capabilities such as a server, a desktop computer, or a laptop computer. The corresponding operating system and application software can be installed in the migration device, and the functions required in this embodiment can be realized through the combination of the operating system and the application software.

[0031] It should be noted that in this embodiment, please combine Figure 2 When obtaining the scene image, corresponding strategies are adopted for different data sources. If the data comes from a publicly available structured dataset, since it already contains preprocessed scene pictures and corresponding camera poses, the processing can directly construct a three-dimensional scene based on the Neural Radiance Field (NeRF). On the other hand, to improve the applicability of the method, it also supports extracting scene data from user-made videos. In this case, to deal with the large number of redundant frames that may exist in the video, a video frame extraction and duplicate removal algorithm based on image hash similarity is adopted. By calculating the hash value of each frame image and comparing the similarity of its adjacent frames, duplicate frames can be efficiently removed, and only key frames with significant changes are retained. Subsequently, images are uniformly extracted from the deduplicated key frames to ensure the reasonable distribution of the input data in time and space.

[0032] Next, to obtain the camera pose corresponding to the scene image, the camera parameters, including the camera intrinsics and extrinsics, are inferred from the extracted scene images through the SfM module. Specifically, the SfM module first restores the 3D point cloud and the initial pose of the camera based on the image feature point matching relationship using the triangulation method, and then gradually adds images through an incremental optimization strategy until the reconstruction and optimization of all camera poses are finally completed.

[0033] Finally, based on the obtained scene image and camera pose, this embodiment uses the Neural Radiance Field (NeRF) technology to achieve high-quality reconstruction of the three-dimensional scene. As an advanced three-dimensional scene representation method, NeRF can learn the radiance and transparency of each three-dimensional spatial point through a neural network based on the scene image and the camera pose. Specifically, first, the camera pose and the scene image are input into the neural network to generate rays for rendering and perform multi-point sampling on each ray; subsequently, information such as the position and viewing direction of the sampling points is input into a multi-layer perceptron (MLP) to predict the color and volume density of the sampling points; finally, based on the volume rendering formula, the outputs of all sampling points are aggregated into image pixel values, thereby performing end-to-end optimization training on the network. The trained NeRF model can generate a three-dimensional scene with rich details and realistic lighting effects, while retaining the geometric information and material characteristics of the scene.

[0034] S102. Obtain the target image, and call the multi-dimensional probability density function to map the target image and the scene image to the N-dimensional space and project them back to the three-dimensional scene; Specifically, in this embodiment, by constructing a projection matrix, the target image and the scene image are mapped from the original three-dimensional color space to the N-dimensional feature space. In this process, the expression of the projection matrix can be , where and are the mapping coordinates of the scene image and the target image in the N-dimensional feature space respectively, and the pixel points of each image are transformed into high-dimensional feature representations. The introduction of the high-dimensional feature space can capture the semantic and structural information of the image more fully.

[0035] Next, in the N-dimensional space, calculate the marginal probability density functions of the target image and the scene image . Then perform the feature alignment process. For each axis, use an interpolation algorithm to calculate a one-dimensional transformation to and for matching, and project the in the -dimensional space to ; After the feature alignment is completed, project the matching result in the high-dimensional space back to the three-dimensional color space through a rotation matrix to restore the final three-dimensional scene representation. Specifically, the rotation matrix is used to adjust the direction and coordinates when mapping back to the three-dimensional space to ensure that the restored image maintains a sense of reality and consistency. At the same time, this mapping and projection process can effectively retain the detailed information of the target image and avoid unnecessary distortion or blurring during the fusion process. Among them, the expression of the rotation matrix is: , repeat the above steps until convergence in all projection cases, which is the finally color-transferred image. However, the result after the above projection transformation is often dull and contains artifacts. Therefore, in order to reduce the artifacts in the image and retain the gradient information of the image as much as possible, in the post-processing stage of this embodiment, the following loss function is optimized to reduce the artifacts in the image:

[0036] Where is the weight factor of the gradient preservation term, is the weight factor of the color similarity term, is the gradient of the optimized image J at the pixel position (i, j), is the gradient of the scene image I at the pixel position (i, j), is the color value of the optimized image at the pixel position (i, j), is the color value obtained by mapping the target image I at the pixel position (i, j) to the N-dimensional feature space and then mapping it back to the RGB space; It should be noted that through the optimization of the above loss function, the generated image not only retains the gradient information of the scene image, such as the sharpness of the edges and the fineness of the texture, but also achieves a high degree of consistency with the target image in terms of color features. In practice, it is found that it can effectively avoid the common problems of over-smoothing or structural distortion in traditional color transfer methods, making the generated image achieve a balance in terms of detail retention and color consistency. At the same time, the optimized image can significantly enhance the realism and immersion of the visual effect when applied to scenarios such as 3D scene construction and virtual reality presentation.

[0037] S103, call the pre-trained VGG encoder to extract features from the target image and the scene image to generate target style features and scene features, calculate the attention between the target style features and the scene features, fuse the target style features and the scene features based on the attention, and update the 3D scene through the pre-trained VGG decoder; Specifically, in this embodiment, the pre-trained VGG encoder is called to extract features from the target image and the scene image. Among them, the feature map of the scene image is expressed as and the feature map of the target image is expressed as , After normalizing the feature maps of the scene image and the target image to reduce the influence of the dynamic range difference between different images on the feature matching accuracy, two new feature spaces f and g are obtained through two convolutional transformations, and the attention between the feature maps is calculated to generate two feature maps with different resolutions , and the expression for calculating the attention is: ; Among them, represents the value obtained by calculating the attention between the feature of the content image at the th layer and the two feature maps of the target image, represents the value of the feature map of the content image at the th layer after mean-variance channel normalization, represents the value of the feature map of the target image at the th layer after mean-variance channel normalization, represents the feature of the target image at the th layer after convolutional transformation; The two feature maps with different resolutions Fuse them to generate fused features, call the pre-trained VGG decoder to synthesize the style image from the fused features, and update the 3D scene. The operation of the fused features is expressed as:

[0038] where is the fused feature, and are two 1×1 convolutional transformations on the attention features of the fourth and fifth layers.

[0039] The fusion strategy takes into account the detail expressiveness of feature maps with different resolutions and the overall style features, ensuring that the generated image can not only maintain the local details of the scene image but also fully reflect the style features of the target image. The role of the decoder is to map the fused features back to the pixel space to generate an image with the target style features.

[0040] S104, Process the scene images and target images from different perspectives through principal component analysis, and adjust the color distribution of the updated 3D scene.

[0041] Specifically: Combine the image sets from different perspectives and represent them as and represent the target image as . First, calculate the mean and variance of the scene image set and the target image respectively, and further perform singular value decomposition on the covariance matrix of the color distribution. The decomposition of the covariance matrix can reveal the main axis direction of the image in the color feature space, providing a basis for subsequent color adjustment. The decomposition results are as follows:

[0042] Calculate the final rotation and translation matrices after normalizing and scaling the decomposition results to correct the color distribution of the scene:

[0043] The calculated rotation matrix and translation vector are used to adjust the color distribution of the 3D scene to make the adjusted color distribution closer to the style features of the target image. Among them, and are diagonal matrices, is the color mean of the rendered image set, is the mean of the target image, is the color covariance of the rendered image set, is the covariance of the target image.

[0044] The optimized color distribution is applied to update the texture map of the 3D scene, thereby achieving the stylized rendering of the scene. It should be noted that this method based on principal component analysis and singular value decomposition can effectively retain the detailed information in the scene while reducing color distortion. For example, when the target image has complex color gradients, singular value decomposition can accurately capture its dominant hues and texture features, and the rotation matrix calculated through normalization ensures the consistency of color features between the scene and the target image.

[0045] Please combine Figures 3 - 5 , this embodiment is a hierarchical scene stylization method that gradually transfers image features to the scene, enabling style transfer between any scene and target image. While embedding image features into the scene as much as possible, it ensures the content structure and perspective consistency of the scene.

[0046] This embodiment is evaluated on two public datasets Figure 3 and Figure 4 and is compared with the current open-source state-of-the-art method for scene style transfer based on radiance fields and some classic 2D style transfer methods. The comparison metrics are the image similarity LPIPS and the root mean square error RMSE, mainly measuring the multi-view consistency of the stylized scene. Short-range refers to the angle between the perspectives of two pictures taken in the same scene being between 0 and 15°, and long-range means the angle between the perspectives is between 15 and 45°.

[0047] Figure 5 Comparing the visualization results of the scene style transfer method based on radiance fields, it can be seen from Tables 4 and 5 that 2D style transfer methods are difficult to ensure multi-view consistency. Therefore, this embodiment focuses on the advantages and disadvantages of 3D scene style transfer methods based on radiance fields. The stylization metrics of this embodiment are far superior to other scene style transfer methods based on radiance fields. From the visualization results in the second column and the last column of the comparison Figure 6 , it is not difficult to find that the stylization result of this embodiment is far superior to StyleRF, and the scene style is more similar to the target image. At the same time, it can also be found that the images rendered from the stylized scene of this embodiment have more accurate and complete depth maps.

[0048] Please refer to Figure 6 , the second embodiment of the present invention provides a scene style transfer device, including: A 3D scene construction unit 201 for obtaining a scene image and a camera pose and constructing a 3D scene based on the scene image and the camera pose; A mapping unit 202 for obtaining a target image and calling a multi-dimensional probability density function to map the target image and the scene image to an N-dimensional space and project them back to the 3D scene; An update unit 203, configured to call a pre-trained VGG encoder to extract features from the target image and the scene image, so as to generate a target style feature and a scene feature, calculate the attention between the target style feature and the scene feature, fuse the target style feature and the scene feature based on the attention, and update the three-dimensional scene through a pre-trained VGG decoder; An adjustment unit 204, configured to process the scene image and the target image from different perspectives by using the principal component analysis method, and adjust the color distribution of the updated three-dimensional scene.

[0049] A third embodiment of the present invention provides a downhill assist device for an electric vehicle, including a memory and a processor. A computer program is stored in the memory and can be executed by the processor to implement a downhill assist method for an electric vehicle as described in any one of the above.

[0050] A fourth embodiment of the present invention provides a computer-readable storage medium storing a computer program, which can be executed by a processor of a device where the computer-readable storage medium is located to implement a downhill assist method for an electric vehicle as described in any one of the above.

[0051] The present invention discloses a method, device, equipment and storage medium for scene style transfer. First, a scene image and a camera pose are obtained, and a three-dimensional scene is constructed based on the scene image and the camera pose. Then, a target image is obtained, and a multi-dimensional probability density function is called to map the target image and the scene image to an N-dimensional space and project them back to the three-dimensional scene. Next, a pre-trained VGG encoder is called to extract features from the target image and the scene image to generate a target style feature and a scene feature, calculate the attention between the target style feature and the scene feature, fuse the target style feature and the scene feature based on the attention, and update the three-dimensional scene through a pre-trained VGG decoder. Finally, the scene image and the target image from different perspectives are processed by using the principal component analysis method, and the color distribution of the updated three-dimensional scene is adjusted. The problem of reconstruction error occurring during the style transfer process is solved.

[0052] Exemplarily, the computer program described in the third and fourth embodiments of the present invention can be divided into one or more modules. The one or more modules are stored in the memory and executed by the processor to complete the present invention. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in a device for implementing a scene style transfer. For example, the device described in the second embodiment of the present invention.

[0053] The so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the migration method of the one scenario style, and uses various interfaces and lines to connect the whole to implement various parts of the migration method of the one scenario style.

[0054] The memory can be used to store the computer programs and / or modules. The processor realizes various functions of the migration method of the one scenario style by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, a text conversion function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, text message data, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0055] Among them, if the implemented module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0056] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0057] As described above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A scene style migration method, characterized in that: include: Acquire a scene image and a camera pose, and construct a three-dimensional scene based on the scene image and the camera pose; Acquire a target image, and call a multidimensional probability density function to map the target image and the scene image to an N-dimensional space, and project them back to the three-dimensional scene; Calling a pre-trained VGG encoder to perform feature extraction on the target image and the scene image to generate a target style feature and a scene feature, and calculating the attention between the target style feature and the scene feature, fusing the target style feature and the scene feature based on the attention, and updating the three-dimensional scene through a pre-trained VGG decoder; The scene images and target images at different viewing angles are processed by principal component analysis, and the color distribution of the updated three-dimensional scene is adjusted.

2. The scene style migration method according to claim 1, characterized in that: The acquiring of the target image, calling a multidimensional probability density function to map the target image and the scene image to an N-dimensional space, and projecting the result back to the three-dimensional scene, is specifically as follows: Through the projection matrix The target image and the scene image are mapped into N-dimensional space, and the expression is: ,in, , are the mapping coordinates of the scene image and the target image in the N-dimensional feature space respectively; calculate , The marginal probability density function of , and projected onto all axes The marginal probability density function is obtained by ; For each axis, an interpolation algorithm is used to calculate a one-dimensional transformation. and To match, In dimensional space Projection , and projected back to the three-dimensional color space through the rotation matrix.

3. The scene style transfer method according to claim 1, characterized in that: After acquiring the target image, calling the multidimensional probability density function to map the target image and the scene image to the N-dimensional space, and projecting them back to the three-dimensional scene, the method further includes: The gradient information of the scene image is maintained by optimizing the loss function so that the color features of the image generated after optimization are similar to those of the target image. The expression of the loss function is: in, is the weight factor of the gradient preservation term, is the weight factor of the color similarity term, is the gradient of the optimized image J at the pixel position (i, j), is the gradient of the scene image I at the pixel position (i, j), is the color value of the optimized image at the pixel position (i, j), It is the color value obtained by mapping the target image I at the pixel position (i, j) to the N-dimensional feature space and then mapping it back to the RGB space.

4. The scene style transfer method according to claim 1, characterized in that: The calling of the pre-trained VGG encoder to extract features from the target image and the scene image to generate target style features and scene features, and calculating the attention between the target style features and the scene features, fusing the target style features and the scene features based on the attention, and updating the three-dimensional scene through the pre-trained VGG decoder, specifically: The pre-trained VGG encoder is called to extract features from the target image and the scene image, wherein the feature map of the scene image is represented as , the feature map of the target image is expressed as , After normalizing the feature map of the scene image and the feature map of the target image, they are transformed into two new feature spaces f and g through two convolutions, and the attention between the feature maps is calculated to generate two feature maps with different resolutions. , the expression for calculating attention is: ; in, Indicates the content image The value of the two-layer feature map of the layer feature and the target image after attention calculation, Indicates the content image The feature map of the layer is the value after mean-variance channel normalization, The target image The feature map of the layer is the value after mean-variance channel normalization, Represents the target image The characteristics of the layer The value after convolution transformation; Two feature maps of different resolutions Fusion is performed to generate fusion features, the pre-trained VGG decoder is called to synthesize the fusion features into a style image, and the three-dimensional scene is updated. The operation expression of the fusion features is: in, To fusion features, , Two 1*1 convolution transformations are performed on the attention features of the fourth and fifth layers.

5. The scene style transfer method according to claim 1, characterized in that: The method of processing the scene image and the target image at different viewing angles by the principal component analysis method and adjusting the color distribution of the updated three-dimensional scene is specifically as follows: The images from different perspectives are combined and represented as , and the target image is represented as , respectively calculate the mean and variance of the image set and the target image, and perform a singular value decomposition operation on the covariance matrix. The decomposition results are as follows: The decomposition results are normalized and scaled to calculate the final rotation and translation matrix to correct the color distribution of the scene: The calculated rotation matrix and translation vectors is used to adjust the color distribution of the 3D scene, where and is a diagonal matrix, is the color mean of the rendered image set, is the mean of the target image, is the color covariance of the rendered image set, is the covariance of the target image.

6. A scene style migration device, characterized in that: include: A three-dimensional scene construction unit, used to obtain a scene image and a camera pose, and construct a three-dimensional scene based on the scene image and the camera pose; A mapping unit, used to acquire a target image, and call a multidimensional probability density function to map the target image and the scene image to an N-dimensional space, and project them back to the three-dimensional scene; An updating unit, configured to call a pre-trained VGG encoder to perform feature extraction on the target image and the scene image to generate a target style feature and a scene feature, calculate the attention between the target style feature and the scene feature, fuse the target style feature and the scene feature based on the attention, and update the three-dimensional scene through a pre-trained VGG decoder; The adjustment unit is used to process the scene image and the target image at different viewing angles by using a principal component analysis method, and to adjust the color distribution of the updated three-dimensional scene.

7. A downhill assist device for a pure electric vehicle, characterized in that: It comprises a memory and a processor, wherein the memory stores a computer program, and the computer program can be executed by the processor to implement a downhill assist method for a pure electric vehicle as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: A computer program is stored, and the computer program can be executed by a processor of the device where the computer-readable storage medium is located to implement a downhill assist method for a pure electric vehicle as described in any one of claims 1 to 5.